Why text-only RAG breaks on real enterprise documents
Enterprise document corpora are not clean text files. A compliance library is full of scanned PDFs from contracts signed before digitization was standard. Clinical repositories hold discharge summaries with handwritten annotations and embedded lab tables. Engineering knowledge bases store specification sheets where the key data lives in a schematic, not a paragraph. When a text-only RAG pipeline ingests any of these, the parser either skips the visual content or converts it to garbled character strings the embedding model cannot work with. The retrieved context looks plausible in the index but has no connection to what the question was actually asking.
The gap shows up clearly in recall numbers. On born-digital, text-native PDFs, text-only retrieval achieves recall@5 around 0.91 in standard benchmarks. On scanned pages from the same corpus, that number drops to 0.03. That is not a degradation: it is near-total failure. For teams where a meaningful fraction of the document corpus is scanned or image-heavy, the effective knowledge base that the RAG system can actually query is a fraction of what they think it is. Engineers who tested retrieval on text-native samples during development will not see this gap until a production question hits a document the parser could not read.
Closing the gap requires decisions at five points in the pipeline: how documents are parsed, how visual content is represented in the index, which embedding model handles multimodal content, how retrieval weights different modalities, and how the pipeline is evaluated against queries that depend on visual content. Each of those decisions has a default that most teams never make explicitly, and each default tends toward the same outcome: a retrieval system that silently ignores most of what the documents say.
"Text-only RAG scores well on the documents it can read. What it cannot tell you is how many of your documents it cannot read at all."
The 5 document processing approaches compared: where each one breaks
Most teams move through these approaches in sequence, discovering what each one misses only after it produces a wrong answer in production. The table maps the failure modes before you invest in the wrong one for your document type.
| Processing approach | What it handles well | Where it fails | Typical document types affected | Retrieval risk |
|---|---|---|---|---|
| Text-only parse | Born-digital PDFs with clean text layers | Scanned pages, embedded images, charts, tables stored as images | Contracts, clinical notes, engineering specs, financial statements | Critical |
| OCR-only | Scanned text pages, printed forms | Diagrams, charts, complex table layouts, handwritten annotations | Legacy contracts, lab reports, manufacturing specs | High |
| Caption-and-index | Charts and figures that can be described in a sentence | Dense tables, multi-panel figures, diagrams requiring spatial reasoning | Annual reports, research papers, slide decks | High |
| Unified vision embeddings | Mixed text and image content; structured tables; multi-column layouts | Very low-resolution scans; domain-specific symbology without fine-tuning | Most enterprise document types at scale | Moderate |
| Page-as-image (ColPali) | Complex layouts; visual reasoning across the full page; no parsing step required | High per-query compute cost; requires GPU infrastructure for production latency | Technical drawings, scanned archives, dense financial tables | Lower |
Caption-and-index is the most common first step after teams discover their text parser is missing visual content. It works when a single sentence can accurately describe the figure. It breaks when the figure is a 40-row table, a schematic with labeled components, or a chart where the query depends on reading a specific data point rather than understanding the general trend. At that point, the caption describes the document but does not encode the retrievable content the question requires. Unified vision embeddings and page-as-image approaches close this gap but require infrastructure decisions that most teams have not made at the point where they first encounter the problem.
Not sure where your document retrieval gaps are?
10decoders audits your RAG pipeline against your actual document corpus, identifies which document types your parser is missing, and builds the multimodal indexing architecture your retrieval system needs. DocuFindr, our document AI product, handles extraction from complex PDFs, tables, and scanned documents at production scale.
Book a Free AI Assessment →Why modality weighting matters more than model choice
Once a team moves to multimodal embeddings, the next assumption is that the model does the heavy lifting and retrieval just works. It does not. Published benchmarks on enterprise document corpora show that the optimal retrieval blend weights text at 30%, image content at 15%, captions at 25%, and OCR output at 30%. That combination produces a 57.3% improvement over text-only baselines. The same documents retrieved with equal modality weights produce meaningfully worse results, and teams that default to whatever the embedding model's standard retrieval mode returns are leaving accuracy on the table without knowing it.
The reason modality weighting matters is that different document types encode their information differently. Financial statements hold key data in tables, so OCR and structured extraction should carry more weight than the prose around the table. A clinical discharge summary may put everything that matters in the free-text narrative, while the scanned signature block at the bottom is noise. Technical specifications can encode their core content in labeled diagrams where image embeddings are the only signal that matters. A single fixed weighting scheme applied uniformly across all three will underserve at least two of them. The right approach runs a calibration pass on queries against each document type in the corpus and sets weights per document class, not per pipeline.
Most teams skip this calibration step because it requires labeled evaluation data they do not have at the start. The shortcut is to start with the published research weights as a baseline, run the first 100-200 production queries through both the text-only and multimodal pipelines, and measure answer quality against a human-labeled sample. That comparison surfaces which document types the multimodal pipeline is actually improving and which ones might still need a different approach, such as a dedicated table extraction model for particularly dense financial data.
Parser-dependent retrieval
PDF text layer extracted and embedded. No handling for scanned pages, images, or tables stored as images. Recall on visual content near zero.
Supplemental extraction
OCR added for scanned pages. Figures captioned before indexing. Structured tables still partially missed. Calibration not yet in place.
Full document intelligence
Vision embeddings or page-as-image retrieval. Modality weights calibrated per document class. Evaluation tracks visual-content recall separately.
The multimodal retrieval readiness checklist
Before deploying or scaling a RAG system over a mixed document corpus, each of the following should be in place. These are the controls that determine whether your pipeline can actually answer questions about what your documents contain.
"A retrieval score computed only on text-native documents is not a retrieval score. It is an audit of the documents your pipeline already knows how to read."
What to do this week
Audit your corpus before you touch the pipeline
Pull a 500-document random sample from your production corpus and classify each one: born-digital text, scanned page, image-heavy, table-dense, or mixed. Calculate the percentage of each. If more than 15% of your corpus falls outside the born-digital text category, your text-only pipeline has a coverage problem that tuning prompts or switching embedding models will not fix. That classification takes one engineer a day and gives you a concrete number to bring to any architecture conversation.
Run a retrieval recall test on your visual content today
Take 50 queries from your production logs that you know should be answered by documents containing tables, charts, or scanned text. Run them through your current pipeline and measure how often the retrieved context actually contains the relevant information. If the recall number is below 0.5, you have a confirmed problem. If it is near 0.03, you have a pipeline that is not functioning for that document type regardless of what the overall recall metric says. That test costs less than a day of engineering time and produces the evidence needed to justify the infrastructure investment.
Choose your multimodal approach based on your hardest document type
Caption-and-index is the right starting point if your visual content is primarily charts and figures that can be described accurately in a sentence. Move to unified vision embeddings if you have table-dense documents where caption accuracy drops. Consider page-as-image approaches for corpora with complex spatial layouts or heavy scanning artifacts where parser-based approaches produce unreliable text. The choice should be driven by your hardest document type, because that type is where the worst retrieval failures occur and where the most production questions go unanswered.
Connect your retrieval evaluation to a real document intelligence layer
10decoders' DocuFindr product handles extraction from complex PDFs, scanned documents, tables, and embedded figures as structured data, with confidence scoring and exception queues that flag documents the extraction pipeline is uncertain about. If your RAG system is built on top of unverified extraction, retrieval quality is bounded by what the parser surfaces. A document intelligence layer that knows what it cannot read is the difference between a RAG pipeline that works on the documents you tested and one that works on the documents your organization actually runs on.
Let 10decoders close your multimodal retrieval gap
We audit your document corpus, identify the parsing and indexing decisions your current pipeline is missing, and build the multimodal RAG architecture your document types require. From DocuFindr-powered extraction to production retrieval evaluation, 10decoders handles the document intelligence layer most teams skip.
