"We're using Claude for our data" is not a plan
I hear some version of this sentence in most first conversations now. "We're using Claude to help with our data and insights." It is said with the confidence of a decision already made, and when I ask what that means in practice, the answer is usually that someone on the team exports a CSV, pastes it into a chat, and asks what the numbers say.
That is not an AI strategy. It is an intern with no context and unlimited confidence, and it produces the same three problems every time: hallucinated answers, silent shortcuts, and eventually, wrong data written back into the stack. The fix is not to stop using Claude. The fix is to know your stack well enough to decide exactly where it belongs.
What "using Claude for data" usually means
The phrase covers at least four different activities, and teams rarely separate them.
Asking questions of an export. Someone drops a spreadsheet into a chat and asks for insights. The model has no idea what the columns mean, which rows are test accounts, or that the revenue field was redefined in Q2. It answers anyway. I covered why in the chatbot post: the model is correct about whatever it was given, and what it was given is a file with no lineage.
Writing SQL from a description. An analyst describes what they want and the model produces a query. This works surprisingly well when the model can see the schema and the definitions. It works badly when it cannot, because the model will guess a join key that looks plausible, pick the first table with a matching column name, and return a number that reconciles with nothing.
Cleaning or transforming data. The model is asked to dedupe, fill nulls, standardize labels, or reshape a dataset. This is where shortcuts happen. Dropping rows with nulls is a shortcut. Coalescing two date formats by assuming one of them is always US-style is a shortcut. Each one is a reasonable move for a script and a wrong move for your business, and the output looks clean.
Writing back to the stack. Someone takes the cleaned output, or the generated transformation, and loads it. Now the shortcut is in the warehouse, downstream models read from it, and the next person to ask a question is querying a table whose logic nobody wrote down. This is how a chat session becomes a data quality incident.
The first activity produces wrong answers. The last one produces wrong data. The difference matters, because wrong answers get corrected in a meeting and wrong data gets discovered in a quarterly review.
Know the stack before you introduce the model
You cannot decide where Claude belongs until you can draw the stack it is being introduced to. Most teams cannot draw it. They can name the tools, but not who owns each layer, what flows between them, or which tables are actually trusted.
Here is the version I build in the first week of an audit. It is a table, not a diagram, because a table forces you to fill in the blanks.
| Layer | Tool | Owner | Trusted? | Claude's role |
|---|---|---|---|---|
| Ingestion | Fivetran, Airbyte, custom scripts | Data eng | Mostly | None yet |
| Raw storage | BigQuery raw datasets | Data eng | Yes, but unusable | None |
| Transformation | dbt | Analytics eng | Partially | Read and propose |
| Semantic / metrics | dbt metrics, LookML, or nothing | Nobody | No | Not until it exists |
| Reporting | Power BI, Looker | Analysts | Depends on the dashboard | None |
| Ad hoc analysis | Sheets, notebooks, chat | Everyone | No | Read only, with context |
The two columns that matter are "Trusted?" and "Claude's role." If you cannot honestly fill in the first, you have no business filling in the second. A model pointed at a layer nobody trusts does not become a source of truth. It becomes a faster way to distribute the distrust.
Alongside the table, you need the data side of the same inventory. Which fields drive the questions people actually ask. Where each field is produced. Whether the definition in the YAML matches the SQL. Whether there are tests. The hygiene post walks through that check. It takes an afternoon and most teams have never done it.
Introduce Claude to specific parts, with a specific job
Once the map exists, the placement decision gets easy, because the good jobs share a shape. They are narrow, the model can be given the full context it needs, and there is a human review step between the model's output and anything that persists.
Documentation in the transformation layer. Point the model at a dbt model and its upstream sources and ask it to draft the YAML description and column-level docs. An engineer reviews and commits. This is the highest-value, lowest-risk job in the stack, and it directly fixes the documentation gap that makes every other problem worse.
Tests. Given a model and its description, the model can propose accepted_values, not_null, unique, and relationship tests, and it is good at spotting the columns that should have one and do not. Same review gate: proposed tests go in a PR.
Corrections as pull requests. An analyst describes a mapping fix in plain English. The model translates it into a change to the seed file or the CASE statement and opens a PR. The engineer reviews a diff instead of a ticket. I built this for a regional bank whose marketing team was spending half a day a week reconciling partner naming by hand. Correction time went from two and a half days to thirty minutes, and the reason it worked is that the model never touched the warehouse. It touched a branch.
SQL against a defined layer. Natural language querying is fine when the model is reading the semantic layer or a set of documented marts, with the definitions in its context, and the query is logged. It is not fine against raw tables, and it is not fine when "the definitions" are in someone's head.
Explaining what a model does. Hand it a 400-line model that the original author no longer works here to explain. Ask what it does and where the logic is fragile. This is read-only, cheap, and useful the day a number gets challenged in a meeting.
The pattern across all five: the model reads freely, proposes in a branch, and never writes to production without a person approving the diff. That one rule eliminates the "making wrong data" problem entirely. The others get smaller as the documentation and tests the model helped write start to exist.
What to stop doing
Stop pasting exports into a chat and calling the output insight. Stop letting generated transformation code land anywhere without a review. Stop pointing a chatbot at a layer you would not trust a new analyst to query unsupervised.
None of that is anti-Claude. I use it daily and it is the reason a one-person shop can take on a warehouse rebuild. But it is a tool that does exactly what the context tells it to, and if the context is a CSV with no lineage, that is the quality of the work you will get back.
Map the stack. Fix what you find. Then give the model a specific job in a specific layer, with a person between it and production. That is what "using Claude for data" should mean.
If drawing that stack table feels like more work than you have time for, that map is what week one of a stack audit produces. See the stack audit →
Related reading
Why your AI chatbot keeps giving wrong answers
The chatbot isn't broken. It's working exactly as designed on bad inputs. Where the real fix lives.
dbt vs natural language: do you still need to write SQL?
The transformation layer isn't going anywhere. The question is who gets to interact with it.
What is data hygiene, and why does it matter for AI
Accurate, consistent, documented. Most stacks are one of those three things. Where the gaps hit hardest with AI in the loop.
The semantic layer nobody owns
The place where metrics get defined once. In most stacks it does not exist, and AI on top makes the gap harder to see.
Retrieval on dirty data is a distribution problem
Retrieval is a distribution mechanism. If your corpus disagrees with itself, a smarter reranker gives you the same wrong answer with better prose.
The three-day question
Why a client's quick question takes the account team three days to answer, why it is not a capacity problem, and what actually shortens the loop.
Should I use dbt or write my own SQL pipelines?
A candid, practitioner take on when dbt wins, when hand-rolled wins, and what happens when a homegrown stack outgrows the person who built it.
Case study
A regional bank cut correction time from 2.5 days to 30 minutes
What the rebuild looked like in practice, and where the four hours a week came from.