One thing I have been busy confusing myself with the past week is whether natural language search of a PDF should be set up using a full RAG system with a question field or just sending the entire document each time with the person’s question.
The client I built the original multi-document RAG system for has a lot of single legal PDF documents that he wants people to be able to search using natural language queries. He wants to be able to restrict access to each document which means that they can’t be added to a document store with full AI search.
I asked Claude:
in this folder is a system to make natural language searches of PDF documents. It uses Voyage AI voyage-law-2. I have another system that works with transcriptions of videos that just sends the whole document to claude for the search. What is the difference?
This system (PDF search) — true RAG (retrieval-augmented generation):
- At index time, build-index.js chunks the PDF and embeds every chunk with Voyage AI’s voyage-law-2 model (a domain-tuned embedding model for legal text), storing vectors in data/index.json.
- At query time, the query itself is embedded (also via Voyage), then compared against every chunk via cosine similarity (server.js:66-76, 98-103), and only the top 5 (TOP_K) highest-scoring passages are pulled.
- Only those 5 passages are sent to Claude (claude-sonnet-4-5) to synthesize an answer, with citations back to passage numbers — or you can skip Claude entirely and get raw excerpts (mode: “excerpts”).
- Claude never sees the whole document — only the small slice retrieval decided was relevant.
The video-transcription system — “long-context stuffing,” not RAG:
- No embeddings, no vector index, no similarity search.
- The entire transcript is sent to Claude in the prompt on every query, and Claude reads the whole thing to answer.
Practical implications:
- Scale/cost: PDF search stays cheap and fast even for huge documents since only ~5 short passages go to Claude each time. The transcript approach re-sends (and re-processes) the full document on every single query — fine for a short transcript, expensive and slow (or impossible past the context window) for long or many documents.
- Precision vs. context: Voyage’s semantic retrieval is good at surfacing narrowly relevant passages by meaning (not just keyword match) and is specifically tuned for legal/technical language here. But retrieval can miss relevant context if it’s split awkwardly across chunks or phrased very differently from the query. Full-document stuffing guarantees Claude sees everything, so it can synthesize across the entire transcript without retrieval gaps — at the cost of scale.
- Multi-document search: The PDF system is built to scale horizontally — each document gets its own container/index, and you could in principle route a query across many indexes. Sending whole documents doesn’t scale that way; you’re bounded by how much fits in one context window.
In short: one is retrieval-then-generate over a vector index (scales to large/many documents, but retrieval quality gates the answer), the other is generate-over-everything (simpler, no retrieval failure mode, but bounded by context length and cost).
So I asked it about the current system which uses a separate container for each document and a shared ‘engine’.
this system, the PDF search, would it work better for natural language search with long-context stuffing?
These documents run ~30–35K tokens each — that’s small enough to just paste the whole thing into Claude’s context on every query, so long-context stuffing would technically work fine here. My take: not worth switching, for a few reasons specific to this setup.
- Cost per query: RAG sends ~5 short passages (maybe 1-2K tokens) to Claude; stuffing would send the full ~35K tokens every single search. At 30 requests/15min per IP allowed, that’s real repeated cost for no benefit at this document size.
- Multi-document future: your INSTRUCTIONS.txt shows this is designed as one-container-per-document with a shared engine, clearly meant to grow (8084+ next free port). RAG scales to that pattern cleanly; stuffing doesn’t save you anything now and would need rework later if you ever want cross-document search or much larger PDFs.
- voyage-law-2 is doing real work: it’s a legal-domain-tuned embedding model, which should retrieve legally relevant passages by meaning better than a generic model — that’s being wasted if you bypass retrieval.
The only edge you’d gain from stuffing is avoiding retrieval-miss errors (a relevant passage scoring just outside top-5), but at 30-35K tokens you could get most of that benefit cheaper by just raising TOP_K a bit rather than abandoning retrieval entirely.
Hmm, so maybe the full RAG system is better than just sending the entire PDF text each time… Let’s raise the TOP_K as suggested:
bump TOP_K up a bit and see what changes
I’ll bump TOP_K from 5 to 8, then run a local before/after comparison against one of the real documents so we can see the actual effect on results, not just theorize.
What changed: the score drop-off is smooth, not cliff-shaped — #5→#6 only drops 0.0019, #7→#8 drops 0.0007. There’s no natural boundary at 5; passages 6-8 are just as topically relevant as 1-5. In this query (“What interest rate applies to costs awarded in litigation?”), the 3 newly-included passages cover the 8% statutory rate, the foreign-currency discretion exception, and related case law — all genuinely useful, not noise.
Practical effect of the change:
- ~60% more context sent to synthesizeAnswer per query (5→8 passages, still only a few hundred extra tokens given these are short chunks) — negligible cost increase, still nowhere near the ~35K full-document size.
- Better chance of catching answers that span slightly more source material (e.g. the foreign-currency exception above, which a strict top-5 cut would’ve missed).
- No downside observed here since the extra passages weren’t irrelevant filler.
So, using RAG techniques for a single document is the correct way to do this? Let’s ask Gemini…
For a single legal document, sending the entire PDF directly to a long-context LLM almost always yields significantly higher accuracy than traditional chunked RAG. Legal text is structurally interdependent, and standard vector chunking frequently breaks critical relationships between clauses.
Why Full-Context Wins for Legal Text
- Preserves Global Definitions: Terms defined on page 2 dictate obligations on page 50. Chunking isolates clauses from their definitions, leading to hallucinated or overly broad interpretations.
- Maintains Cross-References: Legal contracts rely heavily on conditional logic (e.g., “Subject to the exceptions in Section 14.2…”). Vector search retrieves either the target clause or Section 14.2, but rarely connects the two seamlessly without extensive top-$k$ tuning.
- Prevents Silent Omissions: Queries like “List every scenario where the buyer can terminate” require 100% document retrieval. Semantic search only pulls the top-$k$ most similar chunks, which means scattered secondary conditions can easily be missed.
Comparison: Full-Context vs. Standard RAG
| Feature | Full-Context (Whole PDF) | Standard RAG (Chunked + Vector DB) |
| Accuracy & Recall | Highest (preserves total context) | Moderate (risks chunking boundary loss) |
| Cross-References | Seamless (LLM sees entire document) | Poor (requires parent-child chunking) |
| Pipeline Overhead | Zero (no vector DB, chunking, or embedding) | High (requires parsing, embedding, DB setup) |
| Cost / Latency | Higher per query (unless cached) | Very low per query |
| Document Limit | Bound by context window (~1M–2M tokens) | Unlimited scaling across millions of docs |
Architectural Recommendation
- For Single Documents (Contracts, Briefs, Patents): Pass the full text into a long-context LLM. Enable Prompt Caching on your provider (such as Anthropic or OpenAI)—this caches the long document context on the first query, reducing latency by up to 80% and cost by up to 90% for all subsequent questions on that same PDF.
- For Multi-Document Repositories: Use a hybrid approach. Use RAG to retrieve the top 2–5 whole documents, then pass those entire selected documents into the prompt window rather than passing individual small text chunks.
Oh. So sending the whole document is better? Let’s ask ChatGPT:
I’m building various RAG systems. If I only have one document to build a natural language search for, is it a different approach than for multiple documents?
… well it rambled on a bit about approaches and hedged it’s bets with a non-answer.
I’m going to start a new project and give Claude some more to work with and maybe get it to try both approaches and test both for the results.
UPDATE 17/09/2026 – Testing and results here. For my limited testing, Sonnet 5 with RAG gave the best results and much cheaper.

Leave a Reply