Built to be quoted
The short version: Quotability is not a formatting trick. Answer engines repeat what is original, specific, and attributed, and pass over what only looks tidy. A page can carry every structural signal in the current playbook and still hand a machine nothing worth lifting, because structure only helps a claim that was already worth repeating travel further.
Why does structure get the credit that substance actually earned?
The current advice for getting quoted by an AI answer engine reads like a formatting checklist: a bolded answer block, a question-shaped heading, a tidy list, a sprinkle of schema markup. None of it is wrong, exactly. It is aimed at the wrong layer of the problem, and the evidence for that is more direct than most of the advice admits.
A page can check every box on that list and still say nothing a machine has not already read somewhere else: a bolded answer restating the obvious, a heading shaped like a question nobody actually asks, a list of tips gathered from other lists of tips. The wrapper is not the problem. What is inside it is.
Ahrefs tested roughly 1,885 pages that added schema markup on its own, no new claim, no new number, just a machine-readable label wrapped around content that already existed, and found no reliable citation lift from the addition alone (Search Engine Roundtable, to re-verify). Schema is close to pure structure: a wrapper with nothing new inside it. If structure alone were what answer engines rewarded, this is exactly where the effect should have shown up plainly. It did not.
What did the first real study of this actually find?
A Princeton, Georgia Tech, and IIT Delhi team published one of the first empirical studies of what actually moves a source's standing inside a generated answer, presented at KDD 2024. Working through a range of ways a page might be changed, the result was specific rather than vague: adding cited statistics and direct quotations lifted a source's visibility inside the generated answer more than formatting changes did (arXiv 2311.09735, to re-verify).
Read plainly, that finding puts content ahead of its container. The lift came from giving the answer engine something to attribute: a number with a source behind it, a quotation with a name attached, not a cleaner shell around the same recycled claim. Wrapper changes were the smaller effect. Content changes were the larger one. That is the reverse of the order most guidance still assumes, and it is worth taking seriously precisely because it is one of the few academic studies to test the question directly, rather than infer it from correlation.
That distinction matters because most of what circulates about AI visibility comes from vendors with a product to sell, or a single dashboard describing its own results. An academic study, tested across many pages rather than one client's outcome, is a rarer kind of evidence, and it happens to point away from the wrapper and toward the words inside it.
Why does formatting still get most of the credit?
Partly because it is the easier half to sell and standardize: a template, an audit, a checklist item that either exists on a page or does not. Structure is visible, teachable, and finishable in an afternoon, which makes it a comfortable place to focus.
And partly because formatting and substance tend to travel together, which makes them easy to confuse for one another. A team with the discipline to run original research, name a source, and state a precise number is also, usually, the kind of team that bothers to structure the page well. The structure did not cause the citation. It rode along with the substance that did, and it is easy to mistake the passenger for the driver.
None of this makes structure worthless. It has a real, narrower job: making a claim that is already worth lifting easy for a machine to find and extract without guessing. How to get cited by AI covers that mechanic in full: the answer block, the question-shaped heading, the source attached to every figure. What structure does not do, and was never going to do, is manufacture a claim worth lifting out of one that was not there to begin with.
What actually makes a claim worth lifting?
Strip the formatting question away and three things are left, and each is a property of the claim itself, not the page wrapped around it.
- Originality. A number, a finding, or a stance that exists nowhere else, so an engine composing an answer has nowhere else to get it from.
- Specificity. A precise claim rather than a restated consensus, worded so a machine does not have to guess whether it is the real answer or one of a dozen near-identical phrasings already circulating.
- Attribution. A name, a source, and a date attached to the claim, so it can be repeated without the engine taking on the risk of unverified authority.
A page can have all the formatting in the world and still fail on all three. A page with none of the formatting and all three still tends to get found, quoted, and paraphrased anyway, because the machine's actual job is composing a trustworthy answer, not rewarding a template.
Why do most citations point to pages a brand does not own?
The pattern holds once the frame widens past a brand's own website. Muck Rack's review of AI citation behavior found that the large majority, about 94 percent, of citations point to sources a brand does not own (Muck Rack, to re-verify): press coverage, independent analysis, forum and community discussion.
That figure is difficult to square with a formatting theory of citation. A brand cannot add schema markup to a journalist's article or a stranger's forum comment. What it can do is say something specific, original, and checkable enough that a journalist, an analyst, or a commenter repeats it unprompted, on a page the brand never touched. Once a claim is being repeated independently, in multiple places, an engine composing an answer has more confirmation to draw on than the brand's own page could ever supply alone. The formatting on that one page was never going to decide the other 94 percent.
What does this mean for what gets published next?
Visibility is no longer ranking makes the case for why citation replaced ranking as the finish line, and how a brand should measure that shift once it accepts the premise. This piece has a narrower job: insisting on where the credit belongs once a brand decides to chase that finish line at all.
The reliable move is not a smarter template. It is a truer claim: specific enough to be worth stating, original enough that no one else has already said it, and attributed clearly enough that repeating it costs an engine nothing in credibility. Structure can carry that claim further once it exists. It has never been able to invent it, and no amount of formatting discipline will change that order.
A brand chasing formatting alone tends to plateau: tidy pages, modest results, no clear reason either way. A brand chasing a truer claim has somewhere to go, because an original claim keeps earning fresh mentions each time someone else decides it is worth repeating on a page of their own.
Frequently asked questions
Only at the margins. Formatting helps a machine find and extract a claim that is already worth repeating; it does not manufacture that claim. Ahrefs tested roughly 1,885 pages that added schema markup alone and found no reliable citation lift, which is close to a controlled test of formatting by itself (to re-verify).
A Princeton, Georgia Tech, and IIT Delhi team found that adding cited statistics and direct quotations lifted a source's visibility inside a generated answer more than formatting changes did. Content outperformed its container, which is the opposite order most quotability advice still assumes (arXiv 2311.09735, to re-verify).
Not useless, but not sufficient on its own. Schema helps a machine parse what is already on the page correctly; it does not add a new claim. Ahrefs' test of pages that added schema in isolation found no reliable citation lift, which suggests schema pays off only when it labels content already strong enough to be worth citing (to re-verify).
Three things: originality, a fact or stance that exists nowhere else; specificity, a precise claim rather than restated consensus; and attribution, a name, source, and date attached so the claim can be repeated without risk. A page can have every formatting best practice in place and still fail on all three.
Muck Rack's review of AI citation behavior found that about 94 percent of citations point to sources a brand does not own: press coverage, independent analysis, community discussion (to re-verify). A brand cannot format someone else's article. What it can do is say something specific and original enough that other people repeat it unprompted, which is what an engine ends up citing.
Because a machine still has to find and extract a claim before it can repeat it, and a clean structure is what makes that possible without guessing. Structure does not create the value of a claim; it delivers it. A well-structured page built on a weak claim still fails, and a strong claim buried in a wall of unstructured text still gets missed.
It can help a page get considered, since many answer engines draw from the same index search uses. But ranking well does not guarantee a citation. What decides the citation is whether the page contains a claim original, specific, and attributed enough to repeat safely. Rank gets a page noticed. Substance gets it quoted.
It has to actually exist. Originality here means a number, a finding, or a stance nobody else has published, not a rewritten version of something already circulating. A brand can manufacture the appearance of originality with clever phrasing, but an answer engine comparing many sources tends to expose a restated claim rather than reward it.
Treating quotability as a template to install rather than a claim to earn. Adding an answer block, a question-shaped heading, and schema markup around a restated industry consensus checks every formatting box and still gives a machine nothing worth lifting, because none of those boxes create the original, specific, attributed substance an engine is actually looking for.
No, it means the opposite: cite real sources, attach names and dates, and add original findings wherever a brand genuinely has them. Third-party citation is itself a form of attribution that makes a claim safer to repeat. The argument is against empty formatting, not against evidence. It asks for more of it, not less.