Verdant
Knowledge extraction from visually rich documents, scored on your own pages and run in your own environment.
Verdant is the extraction half of the Curate Labs stack. It does visually rich document understanding, the step the field calls VRDU: reading scans, tables, charts, forms, and long PDFs and turning them into structured text that people, agents, and a graph can use. It runs on your own machines, with the models you choose.
Why it exists
Retrieval and agent pilots work on clean text and fail on the documents that matter. OCR and generic extractors drop structure, invent layout chrome, and flatten diagrams. No single parser wins on every document, and the best models change every few months. The useful question is which approach works on your pages, at what cost, and whether you can prove it.
What it does today
- Digests images and PDFs into structured Markdown, with optional HTML, using a vision model you choose. Each page is rasterized, read, and stitched back into one document.
- Rebuilds charts and diagrams as Mermaid, so the picture survives as something a model can read.
- Serves people and agents. A web UI and a CLI for your team. An MCP server so coding agents can digest and read documents directly.
- Scores itself on your documents. Truth fixtures, Langfuse evals, and OpenObserve traces come in the box, so you compare models and prompts with numbers instead of impressions.
- Starts with one command.
docker compose upbrings up the whole stack, self-hosted. Nothing is sent to a third-party service by default.
Where it is going
The direction is a pluggable framework for the steps of document understanding and information extraction, from reading a page to extracting entities and relationships, so a person or an agent can test the options on a document set and keep the best and cheapest one. The current release implements the digestion pipeline well; the pluggable framework is in progress. We describe the idea as the direction and the digestion pipeline as what works today.
Verdant is open source under AGPL-3.0. It is designed to be self-hosted and is not offered as a hosted service.
From extraction to a graph
When you need the entities and relationships inside the documents, the structured text Verdant produces is the input to GraphForge. That pairing is the Curate Labs path from a pile of PDFs to a graph a reviewer can check.
Three ways to use it
- Open source: clone it and run it.
- Services: have us test pipeline options on your documents and hand you the one that wins, with the evaluation.
- Hosted Verdant is not on offer. If you want it run for you inside your own cloud account, that is a services engagement.
Research behind it
Next steps
- Get Verdant on GitHub
- Talk to Curate Labs about finding the right pipeline for your documents