Retrieval is the secret-sauce surface: getting RAG right
Your consultant's body of work is the retrieval corpus. Embeddings, chunking, hybrid search, the prompt template that pulls in the right context, this is the surface that earns the bill.
Here's the layman version. Imagine a really good consultant. A legal pro who has reviewed ten thousand contracts and can spot a one-sided indemnification clause from across the room. A sales consultant who knows which question to ask at which stage of the deal. A medical specialist who knows the difference between two diseases that look identical on paper for the first three weeks.
What makes them good isn't that they're smarter than other people. It's that when a new situation lands on their desk, they reach into a body of memory and pull out the three or four pieces that actually matter. They don't reread every contract every time. They reach for the one that's relevant.
If you're building an AI product on top of that consultant's expertise, the most important thing you build isn't the model. It isn't the chatbot. It isn't the slick UI. It's the part that reaches into the cabinet and pulls out the right three pieces.
That part has a name. It's called RAG (retrieval-augmented generation, if you want to look it up later). The model is a smart-but-amnesiac generalist. Retrieval is the part that hands it the right page from the right binder at the right moment. Get retrieval right and a small model looks brilliant. Get it wrong and the biggest model in the world will sound confidently, fluently, ruinously wrong.
This piece is about getting it right.
Why retrieval is the highest-leverage surface in the whole stack
When I look at where my time actually pays back in an AI product. It's not the prompt or the model choice. It's the retrieval pipeline. Retrieval decides what the model sees. Everything downstream is conditioned on that. If retrieval hands the model a stale doc, the answer is wrong. If retrieval hands it three irrelevant chunks and one good one, the model dilutes the good one. If retrieval hands it the wrong precedent, you get the wrong precedent confidently cited.
The model doesn't get to vote on what's relevant. The retrieval system does. So the retrieval system is where the actual judgment lives.
This is the part most teams underbuild. They drop their docs into a vector store with default chunk sizes, slap on a top-k similarity search, and call it RAG. Then they spend three months tuning the prompt and never look back at the retrieval layer, which is where eighty percent of their failures are coming from. Don't be that team.
What "the consultant's body of work" actually is
Before we get into the mechanics, let's be honest about what we're indexing. The consultant's body of work isn't one kind of thing. It's at least four.
There's reference material, the static stuff the consultant treats as ground truth. Statutes for a legal pro. Clinical guidelines for a medical specialist. Pricing sheets and competitor intel for a sales consultant.
There's annotated examples, past situations the consultant has worked, with their judgment attached. "This contract clause looked fine but I rejected it because of clause 14.3." "This patient looked like flu but I tested for X because of travel history."
There's decision rules and rubrics, the consultant's own playbook. The triage tree. The discovery script. Short, dense, and needs to come back almost every time.
And there's voice and tone artifacts, the way the consultant talks to a client. Hedging language. Reassurance language.
These four want different treatment. Reference material gets chunked one way. Annotated examples differently. Rubrics shouldn't be chunked at all. Voice artifacts often live in the prompt template rather than in retrieval. Treating them as one undifferentiated pile is the first mistake.
Where this corpus comes from is a separate piece. The capture side (annotated examples, decision rules, voice) lives in capturing the secret sauce. This piece picks up after the consultant has handed you content; the next thirty pages are about what to do with it.
Pick the embedding model on purpose, not by default
An embedding (turning a chunk of text into a list of numbers that captures its meaning, if you want to look it up later) is what makes "find me content similar to this question" possible. The embedding model decides what "similar" means. So picking it matters.
Most teams pick by accident, whichever embedding endpoint was easiest to call. Don't do that. You'll embed every chunk in your corpus with this model. Switching later means re-embedding everything. Pick once, pick well.
What I consider:
Dimension count. Smaller embeddings (384 or 768) are cheaper to store, faster to compare, and good enough for most consultant-vertical corpora up to a few million chunks. Larger embeddings (1536, 3072) carry more nuance but cost more on storage and on every query. For an MVP I default to the smaller end and only go bigger if eval results push me there.
Domain match. If the corpus is medical, an embedding model trained on medical text will outperform a generalist by a wide margin. Same for legal, same for code. For narrow verticals, look for a domain-trained option.
Open-source vs hosted. I start with an open-source embedding model running on the local Mac Studio and pushing batched embeddings into RDS. Cheap, fast, no vendor between me and my data. The Mac Studio runs the embedding job overnight; results land in pgvector in the cloud. One pipeline, one model, predictable behavior.
Chunking is where most retrieval dies
Chunking (splitting documents into pieces small enough for the model to consume) sounds boring. It is the single biggest determinant of retrieval quality I've ever seen. The default approach (split every 500 tokens, no overlap) is wrong for almost every corpus.
The reason is that meaning lives in coherent units, not in 500-token windows. A contract clause is a meaning unit. A symptom-cluster description in a case note is a meaning unit. A discovery question, with its purpose and follow-up notes, is a meaning unit. A 500-token window will hack one of those in half and put the other half in the next chunk. Now neither chunk retrieves well, because each one is missing context.
What works for me, almost always, is semantic chunking, split on meaning boundaries first, then size from there. Concretely:
- For the legal pro's contract corpus: split on clause headers. Each clause becomes a chunk. Long clauses get further split on subsection. Short related clauses get glued together up to a max length.
- For the medical specialist's case notes: split on the sections of a SOAP note (subjective, objective, assessment, plan). Then within "assessment", split per differential diagnosis discussed. Symptoms, labs, and imaging stay attached to the assessment that referenced them.
- For the sales consultant's discovery library: each discovery question is its own chunk, with its purpose, the deal stage it belongs to, the follow-ups, and the common objections all attached as metadata. The question itself is what gets embedded. The metadata is what gets filtered on.
The chunk boundaries are the boundaries of the actual ideas. Sometimes 80 tokens, sometimes 1,200. In my own evals, semantic chunking has lifted top-3 retrieval accuracy from the high sixties to the low nineties on the same corpus, with the same model, with no other change.
Add one sentence of overlap if your boundaries are fuzzy. Don't go heavy on overlap, it bloats the index without buying much.
Hybrid search: vector plus keyword, not vector or keyword
Pure vector search has a known weakness. It's great at "find me content similar in meaning" and bad at "find me content that mentions this exact term." If a customer asks "what's the indemnification cap in our standard MSA," vector search will find a bunch of chunks about indemnification generally and may or may not find the specific clause that names the cap. Keyword search ("indemnification cap") will go directly to the chunks that contain that phrase.
The right answer is to do both, then combine. This is hybrid search. In Postgres-with-pgvector land, it's:
- Run a vector similarity query for top-N candidates (say N=20).
- Run a Postgres full-text-search query (
tsvector/ts_rank) for top-N candidates on the same query terms. - Fuse the two ranked lists. I default to reciprocal rank fusion, each chunk gets a score from each list, the scores are combined, the top-K of the fused list is what the model sees.
This adds maybe ten lines of code and a tsvector index. It pays back in retrieval quality on the queries where it matters, the ones that contain a specific term, a name, a number, a clause reference. Vector-only loses those. Hybrid catches them.
For the medical case-notes corpus, hybrid is non-negotiable. A query like "patient with elevated D-dimer" needs to retrieve cases that contain "D-dimer" specifically. Vector might find cases about coagulation generally and miss the ones that name the test. Hybrid pulls both.
Filter on metadata before you compute distances
Every chunk in the index should have metadata. Tenant, document, document type, date, deal stage, symptom cluster, clause type, whatever your domain calls for. Filter first, then run the similarity search on the remaining set.
For the sales consultant's discovery library, this is huge. The questions for "early-stage deal, technical buyer, infrastructure" are different from "late-stage deal, economic buyer, CRM." If the index has 1,200 discovery questions and the customer is mid-stage with a CFO, you don't want the top-K full of early-stage technical questions. Filter on stage and persona before you compute similarity.
In Postgres this is a WHERE clause in the same query as the vector search. The index is smaller because the filter prunes most of it. Latency goes down. Quality goes up. Go look at any "RAG tutorial" online and count how many set up metadata filters. Almost none.
The prompt template is part of retrieval
The retrieved chunks don't go straight to the model raw. They go through a prompt template. This template is part of the retrieval surface, not a separate thing. The decisions you make here matter as much as the chunking decisions.
What I keep in the template, every time:
- System framing, who the model is supposed to be in this product. The legal-pro version says "you are a contract review assistant. You cite clauses by their identifier. If you can't find a clause that supports your answer, you say so." The medical-specialist version is different but the shape is the same.
- The retrieved chunks, in a stable order, with stable delimiters. Each chunk has a small header noting its source, the document, the clause number, the case ID. This lets the model cite its sources, which makes the answer auditable, which is the difference between a useful answer and a confidence trick.
- The decision rules / rubric, when the rubric is short enough to fit. For the contract-review case the playbook is fifteen rules and they all go in the prompt every time. For larger rubrics, the relevant rule gets retrieved alongside the chunks.
- The user's question, last, framed clearly.
- A response shape, what the model should output. JSON with cited chunks, a confidence score, and the answer body. This makes the downstream parsing reliable and gives the audit log something to record.
The template lives in git. It's versioned. Every change goes through the eval suite before it ships. That's a whole separate piece, which is the next one in this series.
The template is code, treat it that way. The next piece, prompts as code, digs into the versioning, A/B, and rollback shape that makes this safe to change in production.
The hand-extension on first use
When a new consultant signs up and uploads their corpus, the retrieval pipeline can't be a black box. The MVP shape I run: I show them, in their first session, exactly what was retrieved for their first few queries. Chunks, scores, final answer. They can mark a chunk as "wrong, shouldn't have come up" or "missing, you should have pulled this." Those marks become training data for the chunking heuristics and the embedding-finetune job that runs on the Mac Studio.
The first user of a new corpus gets the wizard view. Every subsequent user gets the polished view. That's how the corpus gets calibrated to the consultant's actual judgment in the first week.
What to do this week if you're shipping one of these
Three things, in order:
- Open up the chunking code. If it's
split_every_500_tokens(text), replace it with something that splits on the meaning boundaries of your domain. Headers, sections, clauses, questions, cases. Whatever the unit is. - Add hybrid search if you're vector-only. Ten lines, one extra index, big quality lift on term-specific queries.
- Pick one query that's been quietly wrong in production and trace it. Which chunks were retrieved. Why those. What was missing. The trace tells you whether your problem is chunking, embedding, ranking, or template. You can't fix it until you know which.
Retrieval is the surface that earns the AI bill. Build it like it's the product. Because it is.