Cited answers · 2026-07-03 · 8 min read

Cited answers should make every important claim checkable

What a useful citation must show, how to test it and why a source link alone does not make an AI answer trustworthy.

A superscript number is easy to add to an AI answer. It is much harder to make that number useful. If it opens the first page of a 60-page document, points to a passage that only shares a keyword, or cites a source the reader cannot access, the answer is decorated rather than supported.

Take the claim, “The retry policy changed after the March incident and only affects enterprise imports.” There are at least three facts to verify: the policy changed, the incident caused the change, and the scope is limited to enterprise imports. An incident review may support the reason, while the current specification supports the scope. One link at the end cannot honestly stand behind all three.

Citations work best at claim level. The reader should be able to open the exact passage, see a little context around it, identify the source and its date, and return to the answer without losing their place. Verification should take seconds, not a second research session.

The surrounding context matters because a matching sentence can still be misleading. It may sit under a heading marked “rejected proposal,” refer to an old product version or be followed by an exception. A useful citation shows enough of the record to catch those details.

Conflicting sources should remain visible. If an approved policy and a newer draft disagree, the answer can explain the conflict and identify which record appears authoritative. Quietly blending both into a neat conclusion removes precisely the information a decision-maker needs.

Permission handling is a separate test, not a footnote. Retrieval must exclude material the current user cannot open before ranking or generation begins. Otherwise a restricted title, snippet or inferred fact can leak even if the final link is hidden.

We also expect a cited system to stop. When the sources answer only half the question, the response should separate what is supported from what is missing. “The specification confirms the current limit; no source in this workspace explains why it was chosen” is a useful result.

A practical evaluation needs only a small set of real questions. Ask reviewers to verify each important claim and record where they hesitate. Broken deep links, irrelevant passages, outdated records, repeated citations and unsupported connective language become obvious quickly.

Measure verification time as well as answer quality. A concise answer that takes ten minutes to check has moved the search work around, not removed it. A slightly less polished answer with precise passages and an explicit evidence gap may be far more useful.

Citations cannot guarantee that a generated answer is correct. They can make its claims inspectable, its uncertainty visible and its failures specific enough to fix. That is the standard company knowledge tools should meet before their summaries influence product, support or operational decisions.