Getting Cited By Perplexity: A Step By Step Breakdown
What llms.txt Proposes It is a proposed convention: a file at your root offering a curated, plain text guide to your site for language model consumers, pointing at the documents you consider authoritative.
Put someone's name against this. Crawler rules sit between marketing, development and whoever administers the content delivery network, which in most organisations means nobody checks them. The failures documented here are not difficult to find, they are simply nobody's job, and a quarterly review taking half an hour prevents the most complete form of invisibility available.
How to Prioritise When Everything Is Slow This work is slower than anything else in the discipline, so sequencing matters. Start with sources you can edit directly, since claiming and correcting listings is nearly free and takes effect within weeks.
What robots.txt Controls It is a request, honoured by mainstream crawlers, that certain user agents avoid certain paths. It has no enforcement behind it and it does not secure anything, but the major providers respect it.
Ahrefs found in July 2025, across 15,000 long-tail prompts and four assistants, that around 80 percent of cited pages did not rank for the original query at all. If citation and ranking were the same thing, that number would be close to zero. ai search visibility
And in a fast moving category where competitors are actively publishing, monthly can miss a shift. Even then, keep the full set monthly and run a small subset more frequently rather than expanding everything.
How to Handle Published Statistics Every figure you repeat should carry its publisher, sample size and date. This is not pedantry, it is self protection, because figures in this field get repeated until nobody remembers the sample.
Every usability study for thirty years has said readers scan, look for the relevant section, and want the conclusion before the reasoning. Extraction wants the same thing for different reasons. When somebody claims that writing for machines requires sacrificing readability, they are usually describing keyword stuffing, which is a separate and obsolete practice.
The idea is reasonable and adoption is inconsistent. Support varies by provider and no major system currently treats it as required. Treat it as a cheap and speculative addition rather than a deliverable worth paying much for.
The Blocks Nobody Chose Most blocking discovered during audits was never a decision. A disallow copied from a template. A staging rule that survived a migration. A security plugin with an aggressive default. A content delivery network setting labelled bot protection with a switch nobody has looked at since launch.
What You Can Do Legitimately More than most teams assume. Claim every profile that allows it and complete it properly. Correct factual errors on platforms that accept corrections, which most do when you have evidence. Respond to reviews, including critical ones, since an unanswered complaint reads as inattention.
If you must change the prompt set, add new prompts as a separate cohort and keep the original series running unchanged. Editing the instrument retrospectively destroys the comparison you have been building.
Deciding Whether to Block Anything There is a legitimate argument for restricting training crawlers, particularly for publishers whose archive is the product. That is a commercial and editorial decision and it deserves a real discussion rather than a default.
Make Sure It Can Fetch You Check that your robots.txt permits the relevant crawler, and check your server logs for what it actually receives. Bot management products frequently serve challenge pages to legitimate retrieval agents, which produces total invisibility with no error anyone sees.
Keep a record of what you predicted as well as what you measured. Writing down at the start of a quarter what you expect to move, and then reading it back at the end, is the cheapest way to find out whether your model of this channel is any good. Most teams never do it, which is why the same confident explanations survive for years without ever being tested.
Freshness Counts More Than You Expect Because retrieval happens at answer time, a page published or updated this week can be cited this week. This is a meaningful difference from ranking systems where authority accrues slowly.
Where Third Party Coverage Fits Even after all of the above, most citations in a commercial category will point somewhere other than your site. That is not a failure of your optimisation, it is how the system weighs self interested sources.
What Not to Do, and Why It Backfires Fabricated reviews, seeded forum threads under false identities, and paid placements presented as independent all exist and all fail on the same axis. Detection has improved, platforms enforce against it, and the reputational cost when it surfaces exceeds anything the visibility was worth.
Then listen for language. When prospects begin describing your business using phrasing you did not write and your competitors do not use, that phrasing came from somewhere, and generated answers are an increasingly likely source. It is anecdotal, it is not a number, and it is often the earliest indication that anything is working.