How Llms.txt And Robots.txt Affect AI Crawlers
One final practical check costs nothing. Ask for a client reference in a category structurally similar to yours rather than a famous name, and when you speak to them ask what the agency got wrong rather than what went well. References are chosen to be positive, so the useful information is in how candidly they describe the difficult parts.
These pages are cited heavily and are frequently thin, because most are assembled purely to capture the search phrase. A genuinely useful one that says which alternative suits which situation, including cases where staying put is correct, will outperform a dozen keyword driven versions.
The other habit worth building is writing down the number rather than the impression. Teams know their typical lead time, their price band and the size of job they decline, and almost never publish any of it, because a range feels like a commitment. It is a commitment, and it is also the only part of the page a machine can use, which makes it the difference between a page that gets cited and one that does not.
What Padding Looks Like Screenshots of favourable answers with no indication of how many runs produced them. Industry news summaries that could have been written without opening your account. A rising score with no methodology. Traffic charts from unrelated channels included to fill space.
Blocking these is therefore not one decision. Turning away a training crawler is a defensible editorial position. Turning away the agent that fetches pages at answer time removes you from answers entirely, and the two are frequently confused.
The prompt set is the instrument, and almost every weak measurement programme in this field has a weak prompt set at the bottom of it. Get this wrong and everything downstream measures the wrong thing with great precision.
What Makes a Comparison Page Quotable Most vendor comparison pages are unusable, because they are arguments dressed as comparisons. Every row favours the publisher and the conclusion was written first, which is transparent to a reader and produces nothing a model can lift as an impartial claim.
Citation happens at the level of a passage, not a page. A model attaches a source to a specific claim it lifted, which means the real unit of work is a paragraph that stays true and useful once it has been removed from everything around it.
Weight toward the commercial tiers. Roughly a third on buying intent, a quarter on evaluation, a quarter on problem framing and the remainder split between definitional and branded is a reasonable starting distribution.
A practical editing pass makes this concrete. Take a published page and highlight every sentence that could be quoted on its own and still be both true and useful. On most brand pages the highlighted portion is under a tenth of the text. Getting it to a third, without adding length, is usually achievable by moving conclusions forward and replacing three vague sentences with one specific one.
Absence is not disqualifying on its own, since their category is crowded and they may serve a niche. But they should have an interesting answer, and the answer should not be defensive. A practitioner who has run this test on themselves will have thought about it and will tell you what they found.
Where to Get Real Language Four sources, all of which you already own. Sales call notes, where prospects describe their problem before anyone corrects their terminology. Support tickets, where customers describe things going wrong in their own words.
Good versions read like this: mention rate on evaluation prompts rose from two in fifteen to six in fifteen, which we attribute to the three directory corrections completed in week two, though a competitor also stopped publishing during the same period.
Keep a small number of deliberately hostile prompts in the set permanently. Questions asking whether you are expensive, slow or suitable only for large clients reveal what the system believes about your reputation, and the belief is often traceable to one specific source. Nobody enjoys reading those answers, and they generate more actionable work than the flattering prompts do.
They will not quote statistics without sources, and they will not present a tool's sampled estimate as a count of what happened. If none of these boundaries come up unprompted, ask directly and listen for whether the answer sounds rehearsed or considered.
One warning worth stating plainly: none of this means writing for machines. Content that reads as if it were assembled for extraction tends to get treated as low quality by both readers and systems. The goal is writing that a person would find unusually clear and direct, which happens to be exactly what a model can quote. ai seo agency
What robots.txt Controls It is a request, honoured by mainstream crawlers, that certain user agents avoid certain paths. It has no enforcement behind it and it does not secure anything, but the major providers respect it.