Document extraction is the unglamorous bottleneck of enterprise RAG
Document extraction is not complete when a parser returns readable text. For enterprise RAG, it is complete when heterogeneous sources cross a normalized, versioned interchange contract that downstream retrieval can reprocess without reacquiring the source. If a parser loses table structure, section hierarchy, or rotated text before that boundary, a better retriever cannot recover it.
Enterprise documents arrive in three broad categories, each demanding a different extraction strategy. Text-born PDFs — the kind generated from Word or LaTeX — preserve a character-level layout model that tools like pdfplumber can reconstruct with high fidelity. The real work here is layout analysis: detecting multi-column flows, identifying table boundaries via explicit ruling lines, and handling edge cases like ligature expansion and password-protected files. Scanned PDFs require an OCR branch entirely, with heuristics to detect zero-character pages and route them to Tesseract, Azure Document Intelligence, or Textract. And DOCX files demand a different approach altogether, traversing the OOXML object model to reconstruct heading hierarchies and handle merged cells in tables.
The extraction boundary I want is a persisted artifact, not a parser return value. Format-specific adapters may take different routes, but before chunking they should emit normalized content plus the metadata needed to interpret and reproduce it, then snapshot that output. Retrieval teams can rechunk, refilter, or re-embed the snapshot without recrawling Notion, reopening a PDF, or changing acquisition and retrieval in the same experiment. Markdown plus metadata works for some prose-heavy sources; typed JSON or dual representations are safer when tables, figures, coordinates, or legally significant layout carry meaning.
Code presents its own extraction challenge, but one with a cleaner solution: tree-sitter. Instead of line-based or sliding-window chunking, you parse the concrete syntax tree and extract definition-level chunks — one chunk per function or class, with docstrings attached as metadata. This produces semantically complete units that respect the actual boundaries of the code, not arbitrary token limits. The CST includes all tokens (whitespace, comments, punctuation), and the AST view within it gives you the named nodes that matter for retrieval. Hierarchical chunking — class header plus per-method chunks prepended with the class signature — preserves context without duplication.
Benchmarking still matters, but readable text is not the whole contract. A parser can score well on characters and fail if a later run cannot identify the source version, normalization rule, or metadata that accompanied the chunks. The system may retrieve an answer while remaining unable to reproduce why that answer was available. I therefore test semantic preservation and replayability beside parser accuracy.
The economic layer is often the decision constraint. Cloud OCR services range from Azure Document Intelligence at $0.50 per 1,000 pages (best accuracy/cost ratio for general documents) to Google Cloud Vision and AWS Textract at $1.50 per 1,000 pages (better for handwriting and complex scripts). Self-hosted options like Tesseract are free but require binarization and deskew preprocessing. For high-volume enterprise pipelines, the routing decision becomes: which documents justify paid OCR, and which can stay on the free tier? The answer usually involves a hybrid approach — text-born PDFs on open-source parsers, scanned invoices and forms on Azure DI, and complex engineering drawings on Google Vision.
Here is the precise concession: an API-first source can remove the expensive parser and OCR branches. The interchange boundary remains useful when downstream teams need repeatable filtering, chunking, or embedding, because structured records still need a versioned snapshot. Skip document extraction when the source is already structured; skip the contract only when reacquisition is cheap and reproducibility is genuinely irrelevant.