Retrieval on dirty data is a distribution problem
A team ships a chatbot on top of their warehouse. It answers well in the demo. In production it starts giving wrong numbers, and the debugging goes exactly the wrong way. Tune the chunk size. Swap the embedding model. Add a reranker. Move to a larger context window. Two weeks in, the answers are marginally more coherent and still wrong.
Retrieval is a distribution mechanism. It moves whatever is in the underlying corpus to the model, faster and more precisely. If the corpus disagrees with itself, retrieval distributes the disagreement. The fixes that get most of the attention, better embeddings, smarter models, longer context, all improve the reasoning over what was retrieved. None of them repair a source that says two different things.
What retrieval actually does
Retrieval takes a question, converts it to an embedding, and finds the chunks or tables whose embeddings sit closest to it. The top few get handed to the model as context. The model reasons over that context and returns an answer.
It has no notion of truth. It has a notion of proximity. If two chunks about the same topic contradict each other, retrieval will return both, or one, depending on how the question is phrased. The model then synthesizes an answer from whichever it got, with the same confidence either way.
For a data-facing AI the corpus is your warehouse. The chunks are marts, tables, semantic definitions, whatever the retrieval layer is pointed at. This is the semantic layer problem at a different altitude. If revenue is defined three ways across three marts, retrieval will return one of them for any given query. The model will build a paragraph on top of it. Nobody will flag the ambiguity because nobody sees the other two definitions.
Why teams debug the retrieval instead of the corpus
Retrieval is the visible layer. The corpus feels done. It was populated months or years ago, and nobody wants to open it back up. That is the same reason bad hygiene persists in warehouses generally: the transformation layer is where the problems settled in, and reopening it is a project.
There is also an incentive gap. Chunk size and embedding models are knobs. Knobs feel like progress. Deleting a mart, or rewriting a definition, is a conversation with the person who owns it, which is a project. The knob is faster to turn, so the knob gets turned first.
The last piece is that the AI ecosystem is loud about retrieval improvements and quiet about corpus hygiene. Every week there is a new reranker, a new embedding model, a new retrieval technique. Almost none of the content is about the boring, unshippable work of making the source consistent. So teams keep swapping components in a layer that was not broken.
What a dirty corpus looks like in practice
An engagement earlier this year. The client had a warehouse chatbot pointed at their marketing marts. Numbers were wrong often enough that the marketing team stopped using it. When we looked, three marts each had a version of a spend metric. One rolled up at platform grain. One at campaign grain with a filter that excluded internal test spend. One had a WHERE clause added six months ago to patch a specific report and never removed. All three were reachable by the retrieval layer. Depending on the exact phrasing of the question, it returned a different one.
The team had been tuning the reranker for a month. The fix was deleting two of the three marts and rewiring the retrieval layer to point at one mart, one definition, one number. The reranker did not need to be smarter. The corpus needed to stop contradicting itself.
None of the retrieval tuning would have fixed this. A perfect reranker returning one of three wrong numbers is still returning a wrong number. Two of the three had to go.
Why a bigger model does not fix it
The obvious next move when retrieval gives inconsistent answers is to reach for a smarter component. A larger context window, so the model sees more chunks. A stronger model, so it reasons better. A reranker, so the top chunks are more relevant.
All three help with reasoning over what was retrieved. None of them repair the underlying inconsistency. A smarter model given three contradictory chunks does not tell you the corpus is broken. It picks one, or averages, or hedges, and returns a confident paragraph. The failure is not intelligence. It is that the model is doing exactly what it was asked to do on inputs that disagree.
The trap is that each upgrade produces a small improvement, which reads as progress, which delays the corpus conversation. The corpus conversation is the one that actually fixes it, and it does not ship in a weekend.
What to do this week
The check is smaller than the tuning work you were about to do. It takes a few hours.
Pull the top twenty queries the assistant has received. You should have logs. If you do not, that is the first thing to fix. For each query, note whether a knowledgeable human would give a different answer depending on which document, mart, or table they reached first.
Grep the corpus for the concept in each query. How many places is "revenue" defined? How many marts have a version of "active user"? If the answer is more than one and the definitions do not match, retrieval was never the problem.
Delete or archive anything outdated, draft, or superseded. Old marts that got replaced but never removed. Draft docs that got scraped. Test tables. Every one of them is a source retrieval will happily reach for. If nobody has referenced it in six months, it should not be in the corpus.
Point retrieval at one canonical version per concept. For docs, one source of record with everything else linking to it. For data, the semantic layer or a small set of documented marts. If you cannot name the canonical version, retrieval cannot find it either.
Log what retrieval actually returned for the wrong answers. Not just the query and the response. The chunks or tables that got pulled. When the top three come from sources that disagree with each other, that is the bug. It is easier to see when you can look at it.
Only then start tuning. A cleaner corpus makes every retrieval improvement land harder. A dirtier corpus absorbs them. Doing the cleanup first is the difference between a reranker that helps and a reranker that gives you faster wrong answers.
Retrieval was fine. The corpus was the problem. That is the version of this story that shows up in almost every failed rollout, and it is the one nobody debugs first.
If your AI is answering questions from a warehouse that quietly says three things at once, a stack audit finds the disagreements before you spend another sprint tuning retrieval. See the stack audit →
Related reading
Why your AI chatbot keeps giving wrong answers
The chatbot isn't broken. It's working exactly as designed on bad inputs. Where the real fix lives.
dbt vs natural language: do you still need to write SQL?
The transformation layer isn't going anywhere. The question is who gets to interact with it.
What is data hygiene, and why does it matter for AI
Accurate, consistent, documented. Most stacks are one of those three things. Where the gaps hit hardest with AI in the loop.
"We're using Claude for our data" is not a plan
Four different jobs get hidden inside that sentence. Where Claude belongs in the stack and where it does not.
The semantic layer nobody owns
The place where metrics get defined once. In most stacks it does not exist, and AI on top makes the gap harder to see.
The three-day question
Why a client's quick question takes the account team three days to answer, why it is not a capacity problem, and what actually shortens the loop.
Should I use dbt or write my own SQL pipelines?
A candid, practitioner take on when dbt wins, when hand-rolled wins, and what happens when a homegrown stack outgrows the person who built it.
Case study
A regional bank cut correction time from 2.5 days to 30 minutes
What the rebuild looked like in practice, and where the four hours a week came from.