专为引用打造
The short version: Quotability is not a formatting trick. Answer engines repeat what is original, specific, and attributed, and pass over what only looks tidy. A page can carry every structural signal in the current playbook and still hand a machine nothing worth lifting, because structure only helps a claim that was already worth repeating travel further.
Why does structure get the credit that substance actually earned?
The current advice for getting quoted by an AI answer engine reads like a formatting checklist: a bolded answer block, a question-shaped heading, a tidy list, a sprinkle of schema markup. None of it is wrong, exactly. It is aimed at the wrong layer of the problem, and the evidence for that is more direct than most of the advice admits.
A page can check every box on that list and still say nothing a machine has not already read somewhere else: a bolded answer restating the obvious, a heading shaped like a question nobody actually asks, a list of tips gathered from other lists of tips. The wrapper is not the problem. What is inside it is.
Ahrefs tested roughly 1,885 pages that added schema markup on its own, no new claim, no new number, just a machine-readable label wrapped around content that already existed, and found no reliable citation lift from the addition alone (Search Engine Roundtable, to re-verify). Schema is close to pure structure: a wrapper with nothing new inside it. If structure alone were what answer engines rewarded, this is exactly where the effect should have shown up plainly. It did not.
What did the first real study of this actually find?
A Princeton, Georgia Tech, and IIT Delhi team published one of the first empirical studies of what actually moves a source's standing inside a generated answer, presented at KDD 2024. Working through a range of ways a page might be changed, the result was specific rather than vague: adding cited statistics and direct quotations lifted a source's visibility inside the generated answer more than formatting changes did (arXiv 2311.09735, to re-verify).
Read plainly, that finding puts content ahead of its container. The lift came from giving the answer engine something to attribute: a number with a source behind it, a quotation with a name attached, not a cleaner shell around the same recycled claim. Wrapper changes were the smaller effect. Content changes were the larger one. That is the reverse of the order most guidance still assumes, and it is worth taking seriously precisely because it is one of the few academic studies to test the question directly, rather than infer it from correlation.
That distinction matters because most of what circulates about AI visibility comes from vendors with a product to sell, or a single dashboard describing its own results. An academic study, tested across many pages rather than one client's outcome, is a rarer kind of evidence, and it happens to point away from the wrapper and toward the words inside it.
Why does formatting still get most of the credit?
Partly because it is the easier half to sell and standardize: a template, an audit, a checklist item that either exists on a page or does not. Structure is visible, teachable, and finishable in an afternoon, which makes it a comfortable place to focus.
And partly because formatting and substance tend to travel together, which makes them easy to confuse for one another. A team with the discipline to run original research, name a source, and state a precise number is also, usually, the kind of team that bothers to structure the page well. The structure did not cause the citation. It rode along with the substance that did, and it is easy to mistake the passenger for the driver.
None of this makes structure worthless. It has a real, narrower job: making a claim that is already worth lifting easy for a machine to find and extract without guessing. How to get cited by AI covers that mechanic in full: the answer block, the question-shaped heading, the source attached to every figure. What structure does not do, and was never going to do, is manufacture a claim worth lifting out of one that was not there to begin with.
What actually makes a claim worth lifting?
Strip the formatting question away and three things are left, and each is a property of the claim itself, not the page wrapped around it.
- Originality. A number, a finding, or a stance that exists nowhere else, so an engine composing an answer has nowhere else to get it from.
- Specificity. A precise claim rather than a restated consensus, worded so a machine does not have to guess whether it is the real answer or one of a dozen near-identical phrasings already circulating.
- Attribution. A name, a source, and a date attached to the claim, so it can be repeated without the engine taking on the risk of unverified authority.
A page can have all the formatting in the world and still fail on all three. A page with none of the formatting and all three still tends to get found, quoted, and paraphrased anyway, because the machine's actual job is composing a trustworthy answer, not rewarding a template.
Why do most citations point to pages a brand does not own?
The pattern holds once the frame widens past a brand's own website. Muck Rack's review of AI citation behavior found that the large majority, about 94 percent, of citations point to sources a brand does not own (Muck Rack, to re-verify): press coverage, independent analysis, forum and community discussion.
That figure is difficult to square with a formatting theory of citation. A brand cannot add schema markup to a journalist's article or a stranger's forum comment. What it can do is say something specific, original, and checkable enough that a journalist, an analyst, or a commenter repeats it unprompted, on a page the brand never touched. Once a claim is being repeated independently, in multiple places, an engine composing an answer has more confirmation to draw on than the brand's own page could ever supply alone. The formatting on that one page was never going to decide the other 94 percent.
What does this mean for what gets published next?
Visibility is no longer ranking makes the case for why citation replaced ranking as the finish line, and how a brand should measure that shift once it accepts the premise. This piece has a narrower job: insisting on where the credit belongs once a brand decides to chase that finish line at all.
The reliable move is not a smarter template. It is a truer claim: specific enough to be worth stating, original enough that no one else has already said it, and attributed clearly enough that repeating it costs an engine nothing in credibility. Structure can carry that claim further once it exists. It has never been able to invent it, and no amount of formatting discipline will change that order.
A brand chasing formatting alone tends to plateau: tidy pages, modest results, no clear reason either way. A brand chasing a truer claim has somewhere to go, because an original claim keeps earning fresh mentions each time someone else decides it is worth repeating on a page of their own.
常见问题
只能起到边际作用。良好的格式能帮助机器找到并提取一个本身就值得被复述的观点,但它无法凭空制造出这个观点。Ahrefs测试了约1,885个仅仅添加了结构化数据标记(schema markup)的页面,发现引用率并没有可靠的提升,这已经接近于一次单独针对格式本身的对照测试(待核实)。
普林斯顿大学、佐治亚理工学院和印度理工学院德里分校组成的团队发现,相比调整格式,加入有出处的统计数据和直接引语,更能提升一个信息来源在AI生成答案中的可见性。内容的作用超过了承载内容的容器,这与大多数“可引用性”建议默认的先后顺序恰好相反(arXiv 2311.09735,待核实)。
并非毫无用处,只是单靠它还不够。结构化数据标记能帮助机器正确解析页面上已有的内容,但它不会增加新的观点。Ahrefs对仅添加了结构化数据标记的页面进行的测试,并未发现引用率有可靠的提升,这说明结构化数据标记只有在为本身就足够有分量、值得被引用的内容做标注时,才能真正发挥作用(待核实)。
三件事:原创性,即别处找不到的事实或立场;具体性,即精确的论断,而非对共识的又一次复述;以及来源标注,即附上名称、出处和日期,让这个观点可以被安全地复述。一个页面可以完全遵循每一条格式最佳实践,却在这三点上全部落空。
Muck Rack对AI引用行为的研究发现,大约94%的引用指向的是品牌并不拥有的信息来源:媒体报道、独立分析、社区讨论(待核实)。品牌无法去调整别人文章的格式。它能做的,是说出足够具体、足够原创的内容,让其他人自发地复述它,而这正是引擎最终会引用的内容。
因为机器在复述一个观点之前,仍然需要先找到并提取这个观点,而清晰的结构正是让机器无需靠猜测就能做到这一点的原因。结构本身并不创造观点的价值,它只是负责传递价值。一个结构精良、但立论薄弱的页面依然会失败,而一个被淹没在毫无结构的大段文字中的有力观点,同样会被机器错过。
有帮助,因为许多问答引擎调用的正是与搜索引擎相同的索引,排名靠前的页面更容易被纳入考虑范围。但排名靠前并不能保证被引用。真正决定是否被引用的,是页面所包含的观点是否足够原创、足够具体、来源标注是否足够清晰,以至于可以被安全地复述。排名让页面被看到,内容本身才让它被引用。
原创性必须真实存在。这里所说的原创性,是指别人尚未发布过的数据、发现或立场,而不是把已经流传的内容换一种说法重写一遍。品牌可以靠巧妙的措辞制造出貌似原创的表象,但当一个问答引擎对比多个信息来源时,往往会识破这种改写,而不是给予奖励。
把“可引用性”当作一个可以直接套用的模板,而不是一个需要靠内容去赢得的东西。围绕一段不过是复述行业共识的文字,加上一个答案区块、一个问句式标题,再配上结构化数据标记,格式上的每一项要求都满足了,但机器依然找不到任何值得摘取的内容,因为这些格式要素,没有一项真正创造出引擎实际在寻找的原创、具体、有来源标注的实质内容。
不,恰恰相反:引用真实的信息来源,附上名称和日期,并在确实拥有原创发现的地方加入这些发现。引用第三方本身就是一种来源标注,能让一个观点被复述时更加安全可靠。这个论点反对的是空洞的形式主义,而不是证据本身。它要求的是更多证据,而不是更少。