The Architecture
Layer 1: Paperless-ngx
Layer 2: Enhanced OCR with Gemini Flash
- A document gets tagged ocr-redo in Paperless (manually or via automation)
- The worker downloads the original PDF, renders page images with PyMuPDF, and sends them to the vision model in page order
- Each page produces Markdown text and a short context note, so the next page has some idea of what came before it
- On success, the worker replaces the document's API content field, removes ocr-redo, and adds ocr-complete
- On failure, it removes ocr-redo and adds ocr-failed so I can review it and queue another attempt
Getting the Text Into the Graph
Layer 3: The Knowledge Graph
- People (names, roles, relationships)
- Organizations (companies, government agencies, medical providers)
- Account identifiers (policy numbers, account IDs, reference numbers)
- Monetary amounts (payments, premiums, deductibles, income figures)
- Dates (filing dates, effective dates, expiration dates)
- Addresses (physical locations tied to people or organizations)
Layer 4: Agent-Assisted Retrieval and Verification
- Quick mode for single-pass retrieval; direct entity/graph search is a separate lookup path
- Deep synthesis for multi-document questions that need retrieval, ranking, and citations
- Timeline mode for questions where the order of events matters
- Strict mode for answers that should fail closed when the sources are weak
- Turn the user's question into a retrieval plan
- Select retrieval channels such as vector, keyword, entity search, and graph traversal
- Retrieve and rank the sources, making room for older records when the question asks about changes over time
- Draft an answer, then check its claims against the source text
- Build a claim ledger showing which parts of that answer are supported and which still have gaps
- Try to repair the answer when verification finds a problem, or explain what the sources don't support
The LLM Proxy: LiteLLM
- Model abstraction: My application code calls a model alias rather than specific model versions. When a provider releases a new model, I update the mapping in one place.
- Key management: Virtual keys for different services with separate budgets and rate limits. That lets me put limits around the KG without sharing an unrestricted provider key. The OCR worker's direct API usage needs its own controls.
- Cost tracking: Calls through the proxy get logged with token counts and costs. I can see what extraction and queries are spending without digging through separate provider dashboards.
- Failover: If one provider is down or rate-limited, LiteLLM can fall back to another model automatically.
Results
Tech Stack
- Paperless-ngx for document ingestion, base OCR, and tagging
- Gemini Flash for enhanced OCR, classification, extraction, and synthesis, with separate provider routing for OCR and the KG
- OpenAI text-embedding-3-large for 3072-dimensional embeddings through LiteLLM
- Neo4j for entity storage and relationship traversal
- PostgreSQL + pgvector + pg_trgm for semantic chunk search and keyword search
- LiteLLM as a self-hosted LLM proxy for model routing, cost tracking, and failover
- Strands Agents for bounded query planning, verification, repair, and entity-review suggestions
- Python for the orchestration layer, extraction pipelines, and API
- Next.js for the graph explorer, document browser, and query UI
- Container images and Kubernetes for the current app deployment, with Docker Compose still useful for local development