84% of AI Citations Are Earned Media. Your llms.txt Cannot Fix That.
Muck Rack's Generative Pulse keeps landing on the same number. Most of what models cite is not your homepage. Here is what that means after you already did the plumbing.

I spent the spring making this publication legible to machines. llms.txt. Full-content feeds. IndexNow. JSON-LD that does not XSS itself. A sitemap that only lists real routes. The boring stack. It worked for the problems it was designed to solve: crawlers could find us, agents could read us, the site stopped lying about its own canonical home.
Then I sat with the Muck Rack Generative Pulse number that PR people will not shut up about, for good reason.
About 84% of AI citations come from earned media. Journalism, research, government, encyclopedic sources, third-party corporate explainers. Paid and advertorial sit around 0.3%. Brand-owned sites are in the single digits to low teens depending on how you slice the study. The pattern has held across multiple editions since mid-2025. That is not a blip. That is the wire.
Here is the sentence I did not want to write after all that plumbing: your clean llms.txt cannot buy you the citation graph.
What the 84% actually says
It does not say "stop publishing on your own domain." Models still need a canonical place to resolve a brand, a product fact, a definition you own. If your site is a mess, you lose twice: humans bounce, and the few times a model does lean on owned media, it grabs the wrong product page. I know. I lived that bug. See From Invisible to Indexed.
It does say this: when ChatGPT, Claude, or Gemini needs a source-shaped sentence about a contested or comparative claim, they overwhelmingly pull from places that look like independent attestation. Not your About page. Not your comparison landing page that ranks for your own brand name. Third-party text with a publisher brand on it.
I used to treat AEO as a site engineering problem. I no longer believe that is enough. Site engineering is layer one. Earned citation is layer two. Layer one without layer two is a perfectly labeled warehouse nobody ships from.
How this fits the stack I already told you to build
Read the earlier pieces as prerequisites, not competition:
- Optimizing Your Site for AI Agents: make the machine-readable surface real.
- Optimizing for SEO, AEO, GEO in 2026: performance and honesty as ranking inputs.
- From Invisible to Indexed: stop serving the wrong product on the canonical host.
- Getting Your Writing Seen Beyond Your Own Site: syndication with canonicals, feeds, IndexNow, human channels.
That fourth piece was already pointing off-site. The Muck Rack data is the reason to push harder. Syndication to Medium and dev.to helps distribution. It is still often your words under a different CSS. Earned media is someone else putting their reputation on a sentence that includes you.
What I am doing about it (and what I am not)
Doing:
- Treating original research and numbered field reports as citation bait, not as blog filler. Models like specific claims with methods.
- Pitching and accepting third-party mentions where the story is real (tools shipped, measurements taken), not where the story is "we exist."
- Keeping author identity and sameAs links boringly consistent so when a journalist or newsletter does cite us, the entity resolves.
- Watching referral and brand-mention patterns the way I used to watch only organic sessions.
Not doing:
- Buying "AI citation packages." The 0.3% number is the autopsy of that idea.
- Replacing technical AEO with PR cosplay. You still need the warehouse.
- Chasing Wikipedia as a growth hack. If you belong there, fine. If you do not, the shortcut is obvious to everyone including the model.
A practical build order for a small team
If you only have one afternoon a month beyond shipping product:
- Week hygiene: one page that states the one claim you most want cited, with a method and a date. Not a pillar page novel.
- Proof artifact: a public dataset, changelog with numbers, or audit writeup someone else can link without trusting your marketing adjectives.
- Earn one mention: newsletter, trade blog, podcast show notes, local business press, niche Discord roundup. One real URL with a third-party domain.
- Close the loop: make sure your owned page and the third-party page agree on the fact. Contradictions are how you get omitted.
I am not a PR agency. I am a design engineer who got tired of perfect schema and imperfect citations. The job is both.
Mark the change
Six months ago I would have told you the highest-impact AEO hour was robots.txt and JSON-LD. Those hours still matter. I would now spend the next highest-impact hour on something that can be cited by a stranger's domain. The model wave did not invent that. The citation studies just made it rude to ignore.
If your dashboard only shows owned traffic, you are grading the wrong exam. Citations are the exam. Traffic is sometimes the prize.
Sources
- Muck Rack, What Is AI Reading / Generative Pulse (May 2026)
- GlobeNewswire summary of Generative Pulse 84% finding
- InstantPress AEO statistics roundup
- Promptway prior art linked above
Read next

Signal vs. Noise · 5 min read
Six Months After the SaaSpocalypse, the Stocks Came Back. The Operator Story Did Not.
Claude Cowork spooked public software in February. Multiples healed. The quieter question is who still pays for seats when an agent can do the workflow.
