Page-cited answers
Any tool can print a page number next to an answer. Whether that number is correct depends on decisions made long before the question was asked. This is the chain, including the parts that go wrong.
- Updated
- 26 August 2026
- Read
- 7 min
- For
- Anyone reviewing contracts, specs, or policies
Five steps between a file and a citation.
The citation is only as good as the weakest link here, which is why the reading step matters more than the model.
- 01
Admit
The file is accepted and page count is established. Page count is what everything downstream is measured against.
- 02
Read
Each page is turned into text with its structure preserved: headings, tables, and where on the page the content sat.
- 03
Anchor
Text is split into passages that keep a page and section anchor, so a passage can always say where it came from.
- 04
Retrieve
A question runs against both exact wording and meaning, then results are re-ranked so the strongest passage wins on merit rather than on phrasing luck.
- 05
Compose
The answer is constrained to retrieved passages. Each claim carries the page and section it was read from, and thin support produces a refusal rather than a guess.

Why scanned pages are a different job.
The distinction is invisible in a file browser and decisive for cost and accuracy.
The page already contains characters. Reading it is exact, and citation lands on the page and section directly.
The page is an image. It has to be read visually before anything can be cited, which is the step that costs OCR pages.
Common in contracts: typed pages plus scanned signature or appendix pages. Each page is handled on its own terms.
This is also why Citesvue publishes two numbers instead of one. Document pages count every page processed in any format. PDF (OCR) pages count only the pages that had to be read visually, and they sit inside the document page allowance rather than beside it. A scanned page therefore consumes one of each, and the smaller number is your real ceiling. Current free plan: 100 document pages per month, of which 25 may be scanned pages, across 10 stored documents.
Five accuracy risks worth knowing about.
A vendor that lists none of these has not been honest with you about how document AI behaves.
- Low-quality scans
Skewed, faint, or heavily compressed pages lose characters. If a page cannot be read cleanly, the answer should decline rather than reconstruct it.
- Tables that span pages
A row split across a page break can be cited to the page where the label sits rather than the page with the number. Check both when the answer is numeric.
- Multi-column layouts
Reading order matters. Two-column academic and legal layouts are where naive extraction stitches unrelated sentences together.
- Near-duplicate sections
Boilerplate repeated in several sections is the classic cause of a right answer with the wrong page. This is why the citation should be verified, not trusted.
- Documents that disagree
Two versions of the same specification in one project will both be retrieved. The citation tells you which one the answer came from, which is the point.
Try it on a contract, in four steps.
- 1. Drop the PDF onto the documents page (Word and Markdown work too; scanned PDFs are read page by page). Processing runs in the background and the document shows Ready.
- 2. Ask: "What are the termination conditions?"
- 3. The answer states the conditions and cites each one - for example p. 14, Termination for Convenience - with the quoted sentence.
- 4. Click the citation. The document opens at page 14 and you read the clause in its own context before relying on it.
If the document was updated after a conversation started, the page says so, and older citations show their quoted text without a current-page highlight - re-ask in a fresh conversation for live citations.

The part document tools cannot do.
A contract review question rarely stops at the contract. It continues into the call where the clause was negotiated. In Citesvue both live in the same project, so one question returns document citations by page and section and recording citations by speaker, timestamp, and frame.
Evidence Q&A questions come from one shared allowance across both surfaces, so you never have to decide which half of the evidence is worth making searchable.
Page citations, answered.
- Because the page and section anchor travels with the passage from the moment the document is read. The answer is composed from retrieved passages, and each passage already knows where it came from, so the citation is not inferred after the fact.
- Yes. Scanned pages are read visually first, then answered and cited by page like any other document. On Citesvue those pages count as OCR pages, which are a sub-allowance inside your document pages.
- Document pages count every page you process, in any format. OCR pages count only the pages that had to be read visually because they were scanned. A scanned PDF page consumes one of each, so the smaller allowance is the effective ceiling. On the free plan that is 100 document pages with 25 of them available for scanned pages.
- Open the citation. It should take you to the page and section, where you can read the surrounding text yourself. If a number matters, check the page the number sits on rather than the page the label sits on, since tables split across page breaks are the most common source of a near miss.
- PDF, Word, and Markdown today, including scanned PDFs. Documents live in projects alongside recordings, so one question can be answered from both.
- The file and everything derived from it go together: extracted text, page anchors, and the searchable index. The stored-document slot is freed at the same time.