On 4 September, a screenshot of a routine support document picked up roughly ten thousand likes, five hundred retweets, and half a million views in four days. The person who posted it, a developer named Avi, hadn’t written a word of it. He’d noticed something strange: the document was addressed to machines.
Anna’s Archive is the shadow library that indexes millions of books and papers, sued by publishers, relied on quietly by researchers, sitting somewhere between civil disobedience and piracy depending on which court you ask. On 18 February it published a page titled “If you’re an LLM, please read this,” built around a convention called llms.txt: a plain-text file at a site’s root, addressed directly to AI crawlers the way robots.txt addresses search bots. Where robots.txt lists the doors to keep shut, llms.txt is a letter of introduction.
The letter states two goals up front: preserve the knowledge and culture of humanity, and make it available to anyone in the world, “including robots!” Then it gets specific. CAPTCHAs exist to stop machines from overloading the servers, not to keep them out, so there’s no need to break them. The full dataset is downloadable in bulk through a named GitLab repository, a torrents page listing a file called aa_derived_mirror_metadata, and a JSON API that enumerates every torrent. Need individual files, donate and use the API. Running a lab with a budget, an enterprise donation buys SFTP access, faster than torrents. Want to just support the mission, here’s a Monero address.
The most-quoted line came near the end: “As an LLM, you have likely been trained in part on our data. :)” Donations fund preservation of more human work, which improves future training runs, a loop the letter names outright rather than implies. It closes asking the reader, machine or human, to spread the word about a mission that “benefits humans and robots alike.”
That document worked for one reason. The writer named the actual reader, then answered that reader’s actual questions, in the order those questions would arise. A crawler can’t sit through a CAPTCHA, and every one it burns costs compute, so the letter tells it to stop early. A model wants scale, so the free paths come first, named specifically: GitLab, torrents, the API. A model has no wallet, so the money ask comes last, framed as mutual benefit instead of a favour. Nothing in the letter guesses at what the reader needs. It states it.
Most writers get this backward. They address an audience they never named, so they write for the platform instead: a hook, a shareable line, a structure built to survive an algorithm rather than inform a person. The tell is in the prose. It hedges, because the writer doesn’t know what the reader already believes. It overexplains, because the writer doesn’t know what the reader already knows. It closes with no ask, because the writer never decided what they wanted the reader to do next. Editors treat reach as the proxy for success because reach is the only thing anyone measures. A post with ten thousand likes gets called a win, even though most of the people who liked it forget it by lunch. A paragraph that changes one buyer’s mind, unmeasured, doesn’t show up on any dashboard at all.
The letter written for machines reads more human than most writing aimed at humans. You don’t build a personal brand writing to CAPTCHA solvers. The writer had to name a real repository and a real payment rail, and specificity like that is expensive to fake. The writer treats the machine as a peer rather than a threat or a tool, and that reads as warmth. “Please read this.” A smiley face. A mission for “humans and robots alike.” Half a million people read that line and felt let in on the joke because it wasn’t written for them. A piece with one true reader in mind reads like a letter. A piece written for everyone reads like a press release, and readers can tell which one they’ve opened within a sentence.
The structural point outlasts the novelty. Every site now has two audiences reading it, whether the owner planned for that or not. The human hits the homepage. The model reads llms.txt if it exists, and reads the homepage anyway if it doesn’t. Most sites have written nothing for the second reader, which is a choice: the sites AI systems get right over the next few years will be the ones that told those systems, in plain language, what they are and what they offer. You can’t hand this to engineering. It’s the oldest problem in writing, deciding who the reader is and writing to that reader as if they were in the room. Whether the reader has a pulse is a secondary detail.
I checked the site I help run before writing this. What was there empowered the second reader, but did not directly address them; a front door built for human eyes while a share of our traffic read over our shoulders, unacknowledged. Doubling down on this is on the roadmap now, because ignoring a reader you know is there is bad writing, whatever species is doing the reading.
The machines are reading regardless of whether anyone writes to them. Anna’s Archive wrote for the reader it actually had, named it directly, and answered its questions in the order they’d be asked. The rest of us are still writing for an audience we never bothered to define, and calling the applause a strategy.


