The missing layer between AI and the web: a keyless, open-source unlocker that reads, searches, and transcribes what a naive fetch cannot, from bot-walled pages to complete AI-chat share conversations, straight from your terminal or agent.
Ask an AI agent to go read a web page and watch what happens. Cloudflare, PerimeterX, or DataDome takes one look at its naive fetch, decides it's a robot (it is), and slams the door. The agent gets a CAPTCHA or an empty shell, shrugs, and quotes some third-party summary instead of the source. I got tired of watching that happen.
Commercial unlockers charge real money to punch through bot-walls, but the thing you're actually renting is their pool of millions of clean residential IP addresses. Here's the joke: you already have one. searchts runs on your machine, from your home connection, at personal volume. The single most expensive piece of the paid product is sitting in your house.
A bot-wall is a bouncer with a checklist, and each line falls to a different trick:
So a fetch walks a ladder, cheapest tier first, and stops at the first real content:
The ladder remembers which tier worked per domain, so the second visit starts at the cheapest thing that works, and everything comes back as clean Markdown.
Deciding 'real page or wall?' is where naive implementations die, and every rule here was paid for with a real bug. Zillow's genuine homepage ships the PerimeterX sensor script, so matching vendor names falsely flagged 432 KB of real content: match the wall's interstitial phrases, never its vendor. A 500-character minimum called example.com blocked: short is not blocked, short is an escalation hint. And one relay returned HTTP 200 with a body politely explaining the upstream 403: a failure dressed as success, straight onto the block list.
One benchmark row says it all. Zillow: naive fetch, 403. Fingerprint tier, a genuine 200 with 422 KB of real listing data. Same request, same machine, same afternoon. And g2.com, sitting behind DataDome's interactive CAPTCHA, was reported blocked honestly instead of returning junk, because a tool that can't be trusted to say no can't be trusted to say yes either.
Then something beat the whole ladder with no bot-wall in sight. Paste a ChatGPT or Claude share link and all three tiers come back with a thin shell or a conversation cut off mid-sentence. Nothing was blocking me. A chat share page is a single-page app, and the transcript never lands in the page as text worth extracting, so there was simply nothing there to read.
The fix runs ahead of the ladder instead of inside it: recognize the share URL, then read the provider's own data channel rather than the page it paints. Eight are handled now, and they come in two shapes:
Each provider is one auto-discovered module, so adding the ninth is adding a file, and an extractor that fails drops through to the normal ladder rather than failing the read. The benchmark covers the five that need no browser and passes all five. The three that need one are not in it yet.
Reading is a third of it. searchts also searches (keyless, multi-provider, results fused with reciprocal rank) and transcribes video, subtitles-first with a Whisper fallback. It ships as a CLI, an MCP server, a Claude Code skill, and a plain Python library, and installed CLIs unlock native channels: GitHub, Twitter/X, Reddit, LinkedIn, RSS. Fetched content is scrubbed for invisible-character tricks and prompt-injection tells before it reaches an agent, and every read comes with a receipt (which tier, when, final URL), so what an agent read becomes a citation another agent can replay.
An interactive CAPTCHA still needs a human. Instead of pretending otherwise, a --human flag opens a real browser, you solve it once, and the fetch continues. Personal scale only: one home IP at low volume, not a mass scraper. Built on Agent-Reach (MIT), shipped MIT on PyPI. No API key, no proxy bill, no subscription.