Charlotte, NC
BlogApril 6, 2026

Building a Personal Knowledge Graph Over Your Document Archive

Blake McCarn
Building a Personal Knowledge Graph Over Your Document Archive
I have a problem that I think a lot of people share: too many documents, not enough understanding of what's actually in them. Tax returns, insurance policies, contracts, medical records, receipts, bank statements. Over the years I've scanned and dumped hundreds of these into Paperless-ngx, which is a fantastic open-source document management system. It handles ingestion, OCR, tagging, and basic full-text search really well. But here's the thing. Keyword search only gets you so far. I can search for my insurance provider and find the right documents. But I can't ask "what's my deductible across all my insurance policies?" or "what estimated tax payments did I make last year?" Those questions require understanding the content of multiple documents and reasoning across them. That's a fundamentally different problem than search. So I built a knowledge graph on top of it. The system has five main stages. Documents that need better OCR take an extra trip through the vision model, but everything comes back through Paperless before it reaches the graph: The KG's model calls and embeddings route through a self-hosted LiteLLM proxy. The separate OCR worker uses provider SDKs directly. The query path also uses Strands Agents for bounded planning and verification, but the agents don't replace the retrieval system. They sit around it. Let me walk through each layer. This is the foundation. Documents come in through a consumption directory (scanned, emailed, or manually uploaded), and Paperless handles the initial OCR with Tesseract, assigns tags, and stores everything in a searchable archive. I'm not going to spend much time here because Paperless-ngx is well-documented and there are plenty of guides on setting it up. The important thing is that it gives me a stable, API-accessible document store with decent baseline OCR. Every document gets an ID, metadata, and extracted text that downstream systems can pull from. My archive covers everything from tax returns and insurance policies to medical records and contracts. Paperless is still where the documents live. The graph and search indexes are layers I can rebuild from it, not replacements for the archive. This is where it gets interesting. Tesseract does a solid job on clean, typed documents. But real-world paperwork isn't always clean. Handwritten notes on forms, complex table layouts in financial statements, checkboxes on medical intake forms. Traditional OCR struggles with these. I built a separate OCR enhancement container that runs Google's Gemini Flash model against documents that need better extraction. The workflow is tag-based:
  1. A document gets tagged ocr-redo in Paperless (manually or via automation)
  2. The worker downloads the original PDF, renders page images with PyMuPDF, and sends them to the vision model in page order
  3. Each page produces Markdown text and a short context note, so the next page has some idea of what came before it
  4. On success, the worker replaces the document's API content field, removes ocr-redo, and adds ocr-complete
  5. On failure, it removes ocr-redo and adds ocr-failed so I can review it and queue another attempt
The key design decision here was keeping this as a separate layer rather than replacing Paperless's built-in OCR. That way the base system stays vanilla and upgradeable. I can update Paperless without worrying about breaking my custom OCR pipeline, and the enhanced layer evolves independently. One detail worth calling out: this replaces the searchable text, not the original PDF. The worker updates Paperless's content field and tags through the API. That gives both Paperless search and the knowledge graph the improved text without changing the underlying document. The worker processes several documents at a time, but pages within each document stay in order so that context carries forward. I use Gemini Flash as the primary OCR model, with configurable Anthropic and xAI fallbacks. Unlike the KG, this worker calls the providers directly, so its usage is tracked separately from LiteLLM. Of course, cleaner text isn't automatically correct text. If a number or checkbox looks questionable, I still go back to the scan. The two services don't talk directly to each other. The OCR worker updates Paperless, and the KG picks up that text through the Paperless API when sync runs. Documents don't need an ocr-complete tag to get indexed; the base OCR is perfectly usable for many of them. If I want to hold a document out of the graph, the default skip tag is needs-review. Tagging it for re-OCR alone doesn't do that. Sync needs to know more than whether a PDF changed. I might correct its title, change a tag, or improve the extraction rules without touching the file. So alongside a hash of the OCR text, I keep a fingerprint of the metadata, model, and processing rules. If those change, the next sync knows the old extraction needs another pass. I also check which document IDs are present in Paperless, Neo4j, and the vector store, along with their processing records. Just comparing counts isn't enough. Two databases can both say they have 100 documents and still be missing different ones. This is the core of the system. Once documents have good text extraction (either from base OCR or the enhanced pipeline), the system classifies each document and then runs entity extraction with a prompt tailored to that document type. That classification step matters more than I expected. A tax form, an insurance policy, a medical bill, and a contract all contain dates, amounts, names, and identifiers, but they do not mean the same thing. Treating every document as generic text made the early graph noisy. Type-specific extraction gives the LLM a narrower job and keeps the resulting entities closer to the real document structure. Long documents get broken into overlapping sections rather than squeezed into one enormous prompt. Each section goes through five steps: extract metadata, identify entities, check those entities against the text, propose relationships, and check the relationships. The extracted facts need quotes from the source to back them up. If a response gets cut off, the system can split that section into smaller pieces and retry it. It doesn't mark the document finished just because the first few pages worked. The pipeline extracts entities and document-specific facts such as:
  • People (names, roles, relationships)
  • Organizations (companies, government agencies, medical providers)
  • Account identifiers (policy numbers, account IDs, reference numbers)
  • Monetary amounts (payments, premiums, deductibles, income figures)
  • Dates (filing dates, effective dates, expiration dates)
  • Addresses (physical locations tied to people or organizations)
These entities and their relationships get stored in a Neo4j graph database. Document chunks and entity embeddings are stored in PostgreSQL with pgvector and pg_trgm indexes. I keep the actual OCR text separate from the metadata used to help search. A generated summary might help find the right document, but it can't serve as proof for another generated answer. That proof has to come from the source text. Here's what makes the graph powerful compared to flat search: relationships. A keyword search for your utility company gives you every document that mentions them. The graph tells you that the utility company is connected to a specific account number, which is connected to a specific address, which is connected to payment records across multiple years. You can traverse those relationships to answer questions that span documents. The graph can also surface connections you wouldn't think to search for. But getting those connections right is more important than making lots of them. Entity resolution is deliberately conservative. Similar names or embeddings can suggest a possible duplicate, but they don't get to merge anything. Linking a mention to an existing entity needs a supported identity match or an alias backed by the source. Merging two existing entities goes through human review, and a decision to keep them separate needs to survive later processing. The entity steward helps by suggesting things to review, not by silently cleaning up the graph. With personal documents, I'd rather miss a merge than combine two different people or accounts. There's also the less glamorous problem of writing to two databases. Neo4j and PostgreSQL don't share a transaction, so I prepare the extraction and embeddings before replacing a document's indexed data, then record completion last. If something fails halfway through, the missing completion record tells the next run there's work to recover. The first version had two query modes: fast graph search and deep AI synthesis. That was a good starting point, but it was too blunt. Some questions just need a quick entity lookup. Some need a timeline. Some need an answer only if the evidence is strong enough. Lumping all of that into "fast" and "deep" made the interface simpler than the actual problem. The current version still keeps the fast path, but the deeper path is more structured:
  • Quick mode for single-pass retrieval; direct entity/graph search is a separate lookup path
  • Deep synthesis for multi-document questions that need retrieval, ranking, and citations
  • Timeline mode for questions where the order of events matters
  • Strict mode for answers that should fail closed when the sources are weak
This is where Strands Agents fit in. I use them for planning and review around the Python query engine. Together, the steps look like this:
  1. Turn the user's question into a retrieval plan
  2. Select retrieval channels such as vector, keyword, entity search, and graph traversal
  3. Retrieve and rank the sources, making room for older records when the question asks about changes over time
  4. Draft an answer, then check its claims against the source text
  5. Build a claim ledger showing which parts of that answer are supported and which still have gaps
  6. Try to repair the answer when verification finds a problem, or explain what the sources don't support
The important boundary is that the agents do not get to rummage through the database directly or invent their own sources. The Python query engine still owns retrieval, ranking, source packing, and citations. Strands helps plan and critique the work, but the system keeps the evidence path explicit. Timeline mode uses the dates in the final verified answer, keeping each observation attached to its source. That avoids having one model tell a story in the answer and another tell a slightly different one in the timeline. It also keeps an important distinction visible: the most recent document I found isn't necessarily proof of what's true today. Sometimes only part of an answer survives verification. In that case, the system can take the supported parts, check that smaller answer again, and return it with the gaps called out. It doesn't cache that as a complete success. And if the checks fail or never finish, the answer doesn't get a verified label. Strict mode is intentionally boring. If the evidence is weak, the system should say that instead of dressing up a guess as an answer. That matters more for personal documents than it would for a toy demo. I would rather get a partial answer with clear source gaps than a polished paragraph that quietly invented the missing piece. That distinction matters. "Agentic RAG" can turn into a vague blob pretty quickly if every step is just another model call. I wanted the opposite: a deterministic retrieval engine with small agent steps where judgment actually helps. The result is slower than a single vector search, but much easier to trust because I can inspect the query plan, the retrieved sources, the verification result, and the final claim ledger. The key insight is still the same as the early version: vector search alone isn't enough. Pure vector similarity finds documents that are semantically related to your question, which is great for "find me documents about X." But it can't answer relationship questions like "who is my insurance agent and what policies do they manage?" The graph captures structural relationships that embeddings lose. The newer pipeline adds planning and verification around that hybrid retrieval core instead of replacing it. The KG routes classification, extraction, embeddings, query planning, source auditing, repair, and synthesis through a self-hosted LiteLLM proxy. The separate OCR worker is the exception; it calls the providers directly. This is one of those decisions that seemed like overkill at first but has paid for itself many times over. LiteLLM sits between my applications and the upstream API providers (Anthropic, OpenAI, Google). It provides:
  • Model abstraction: My application code calls a model alias rather than specific model versions. When a provider releases a new model, I update the mapping in one place.
  • Key management: Virtual keys for different services with separate budgets and rate limits. That lets me put limits around the KG without sharing an unrestricted provider key. The OCR worker's direct API usage needs its own controls.
  • Cost tracking: Calls through the proxy get logged with token counts and costs. I can see what extraction and queries are spending without digging through separate provider dashboards.
  • Failover: If one provider is down or rate-limited, LiteLLM can fall back to another model automatically.
The proxy runs as part of my homelab AI stack. Postgres stores the spending data and virtual key configs, and Redis handles response caching for repeated queries. This setup means I can hot-swap models without changing application code, track costs for services that actually use the proxy, and set guardrails on spending before I accidentally burn through API credits on a runaway extraction job. The practical value isn't just in answering questions. It's in finding connections between documents and getting to details buried in paperwork I scanned years ago. A policy number, a payment history, a clause in a contract. Those are small things, but they're exactly the things I don't want to spend an evening hunting down. For the more complicated questions, the useful part is having the answer and its sources together. I can see which documents it used, what it found in them, and where the evidence ran out. Deeper synthesis takes longer than a graph lookup, especially with verification, but that extra context is the point. The hard part turned out not to be getting a model to extract names. It's making sure the names refer to the right people, the facts stay attached to the right documents, and a failed job doesn't leave the system thinking it's done. None of that makes the models infallible. It does make their work much easier to inspect.
  • Paperless-ngx for document ingestion, base OCR, and tagging
  • Gemini Flash for enhanced OCR, classification, extraction, and synthesis, with separate provider routing for OCR and the KG
  • OpenAI text-embedding-3-large for 3072-dimensional embeddings through LiteLLM
  • Neo4j for entity storage and relationship traversal
  • PostgreSQL + pgvector + pg_trgm for semantic chunk search and keyword search
  • LiteLLM as a self-hosted LLM proxy for model routing, cost tracking, and failover
  • Strands Agents for bounded query planning, verification, repair, and entity-review suggestions
  • Python for the orchestration layer, extraction pipelines, and API
  • Next.js for the graph explorer, document browser, and query UI
  • Container images and Kubernetes for the current app deployment, with Docker Compose still useful for local development
If you're sitting on a pile of documents and tired of keyword search being the best you can do, a knowledge graph is worth the investment. The combination of structured relationships and semantic search covers a much wider range of questions than either approach alone. Getting the first extraction working is surprisingly approachable. Getting it to behave sensibly when the documents are messy is where the real work starts. The code is split across Paperless-ngx, Paperless OCR Enhanced, and Paperless Knowledge Graph. Enhanced OCR is optional; the KG can consume Paperless's base OCR directly.
Share this post: