checklist
Document pipeline evaluation checklist
Score an extraction pipeline on your own pages before you trust it. Fixtures, metrics, cost, provenance, and the questions a reviewer will ask.
Most document pipelines are judged on a demo PDF and a good feeling. This checklist is the alternative: a short set of checks you can run on your own pages, with Verdant or any other extractor, before the output feeds a graph, a retrieval system, or an agent.
1. Pick the sample
- Choose 10 to 30 documents that represent the hard cases, not the clean ones: scans, rotated pages, multi-column layouts, tables that span pages, charts, forms, handwriting.
- Include at least one document per source system or vendor template you expect in production.
- Record where each document came from and who is allowed to see it. If a document cannot leave your environment, the evaluation cannot either.
2. Write the truth
- For each sample, write down what a correct extraction contains: headings, tables as rows and columns, key fields, figures and their captions.
- Keep the truth in version control next to the pipeline. Verdant stores these as truth fixtures and scores runs against them.
- Mark which parts of the truth matter most. A wrong total in a table usually matters more than a missing footer.
3. Run more than one option
- Run at least two configurations: a different model, a different prompt, or a different parser for the same pages.
- Keep the configuration with each run. A score without its configuration cannot be reproduced.
- Trace each run so you can open any page and see what the model was shown and what it returned.
4. Score what the next step needs
- Structure: did headings, lists, and tables survive as structure, or were they flattened to prose?
- Fidelity: are numbers, names, dates, and identifiers exact? Spot-check against the page image.
- Completeness: did any page come back empty or truncated? Silent blank pages are the most common failure.
- Provenance: can every extracted fact be traced to a page? Downstream, GraphForge can attach that evidence to each claim only if the pipeline kept it.
5. Count the cost
- Record pages per minute and cost per thousand pages for each configuration.
- Decide the quality floor first, then choose the cheapest option above it. The best model is rarely the right default.
6. Decide, and write it down
- State the configuration you chose, the score it achieved on the sample, the cost, and the failure modes you accept.
- Name the trigger for re-running the evaluation: a new document source, a model deprecation, or a quality complaint.
If you want help running this on a document set, talk to us. A pilot is this checklist, done with you, on your pages.