The traceability shape that survives a model swap

When you swap GPT-4 for Claude or for a local Llama, your traceability has to survive. The shape of a model-agnostic trace, prompt-as-data, response-as-data, decision-as-data, model-as-attribution-only, and what happens to teams that couple their traces to a specific vendor's API.

The traceability shape that survives a model swap

The first time I watched a team's traceability story collapse during a model swap. I was sitting in a room where the decision had already been made. They were moving an agent off GPT-4 onto Claude, partly for cost, partly for a regulatory wedge that wanted on-soil inference. The model interface was abstracted. The prompts were templated. On the architecture diagram, it looked like a clean swap.

Then somebody asked what happens to the audit trail. Not the new one. The old one, the one the compliance team had been pointing at for two years as evidence of how well the system was traced.

The answer, after fifteen minutes of silence, was: most of it stops making sense. Half the fields in the trace were synthesized from the OpenAI response object. Some came from the function-calling shape OpenAI used at the time, which Claude did differently and the local Llama did differently again. The prompt itself wasn't logged, it was reconstructed at query time from the template plus the variables, which mostly worked unless the template had changed, which it had, twice. The model version was inferred from the deploy tag, which had been overwritten when the deploy pipeline got rewritten in March.

The team had built a trace that was, without anyone meaning to, a thin wrapper around the GPT-4 API surface. When the model went away, the shape went with it.

Here's how to build a trace that survives. The one whose shape doesn't depend on which model produced the response.

Four properties, paid for once

The shape I keep arriving at, after enough of these migrations, has four properties. They aren't surprising on their own. The discipline is treating them as load-bearing rather than as nice-to-haves.

Prompt-as-data. The exact input that went to the model is logged as a first-class field at the moment it was sent. Not the template. Not the variables that filled the template. The composed string, byte-for-byte, the way the model received it. If the prompt is a chat array, the array is logged. If it's a multimodal payload, the payload is logged with content references. The prompt is data, captured at the boundary, never reconstructed.

Response-as-data. The model's output is captured at the same boundary, in its raw form, before any of your downstream parsing has touched it. The token-by-token streaming version, the final assembled string, the structured output if there is one, the tool call if there is one, all of it. Not the parsed-and-cleaned version your application uses internally. The thing the model actually returned.

Decision-as-data. The policy outcome (what your system decided to do with the response) is logged as a separate field, at a separate emit, with its own row in the chain. If the model said "approve refund" and your policy layer decided "no, the refund exceeds the cap, escalate," both rows exist. The model's output and the system's decision are different events with different actors.

Model-as-attribution-only. Which model produced the response is metadata on the response row, provider, model name, version, parameters, temperature. It is not the trace. The trace would look the same shape if a different model had produced the response. The model is an attribute of one row, not the schema of the whole trail.

Each property is mundane on its own. The combined effect is a trace whose shape is independent of the vendor whose response it carries. A trace you can read across a swap.

Why teams couple their traces to the model API

Most teams build the other shape because the vendor SDK makes it the path of least resistance. The OpenAI Python client returns a response object with a particular structure. Logging that object verbatim is a one-liner. Projecting it into a neutral schema first is fifteen lines, plus a schema definition, plus a versioning policy, plus the discipline to update the projection when the SDK changes shape.

Every team I've watched couple to the API did it because the one-liner was Tuesday's work and the fifteen-liner was the work nobody had time to do. The SDK's response object becomes the trace row. The function-calling field names become trace field names. By month six, the trace's shape is the SDK's shape, and the SDK is the only thing the trail can describe.

This is the same failure mode I covered in why traceability dies in most platforms, wearing a model-vendor mask. Asymmetric incentive (writing logs is cheap, reading them is expensive) projected onto a vendor-shaped surface that promises convenience and delivers lock-in.

What happens when you actually swap

Until you swap, the trace looks fine. The on-call engineer's questions get answered. The compliance export goes through. The auditor signs off. The shape works because there's only ever been one model behind it.

The swap exposes the coupling. Three things break, in roughly the same order every time.

The trace queries break first. The dashboard that filtered on response.choices[0].message.content is querying a field that doesn't exist in the new SDK's shape. Engineering writes a migration script. It almost works. Now you have one dashboard for the old period and one for the new period, and the mismatch leaks into every query that tries to span the cutover.

The historical traces become unreadable second. The rows are still on disk, but the schema has lost its referent. Fields named after OpenAI concepts don't cross-walk mechanically to Claude concepts because the conceptual shapes don't align. Tool calls stream differently. The "stop reason" enum is different. You end up with a corpus only an engineer who remembers the old SDK can read.

The compliance position weakens third. The auditor's next "show me what the system was doing six months ago" lands on fields that map onto a model that no longer exists. The trail has become a historical record of a vendor relationship rather than of system behavior. The auditor doesn't always notice. When they do, the conversation gets long.

I've watched all three play out. The team pays the cost three times, dashboards, historical corpus, audit position. None of it was on anyone's roadmap.

The model-agnostic trace, walked through

Here's what the same emit looks like when the four properties are in force from day one.

A user asks the agent to process a refund. The system composes a prompt, system message, history, tool surface, the request. Before it goes to the model, a row is written: prompt.composed, with the full prompt string, the assembled chat array, the tool surface as structured data, a coordination ID linking it to the user's request, and the intended target model as an attribution attribute.

The model responds. A second row: response.received, with the response captured raw at the boundary, content blocks, tool calls, stop reason, token counts, timing. The attribution field on this row names the model. The shape of the row is the same whether GPT-4, Claude, a local Llama, or a router produced it.

The system parses the response and applies policy. A third row: decision.made, with the policy outcome (approved, escalated, denied) the rule that gated it inline with version, the inputs evaluated, the resulting action. The model said one thing, the policy did another, both are recorded.

If a side effect kicks off (refund written, customer emailed) that's a fourth row, with the IDs of the writes and the chain pointer back to the decision.

Now swap the model. Tomorrow the same agent runs on Claude, or on a local Llama, or on a routing layer. The shape of the trail does not change. The new response.received rows land in the same schema as the old. The attribution field tells you the model is different. Dashboards keep working. Historical corpus stays readable. Compliance export covers both periods with the same query. The migration is a configuration change, not a traceability project.

This is the same discipline I covered in traceability as a debugging tool, not a compliance one, design for the rich, debug-shaped trail and the compliance projection comes free. The model-agnostic property extends that argument across the vendor axis.

Where the chain shape pays back

The four properties pair naturally with the chain-shaped audit I covered in what I steal from version control when designing audit. When prompt, response, decision, and side effect each emit a content-addressed entry with parent pointers, you get something the SDK-shaped trace can never give you: a forward replay.

You can take the prompt row, send it to a different model, capture the new response in the same schema, and see what would have happened. The decision policy is a separate function from the model output, so you can replay the decision against the alternative response without changing application code. "How would Claude have handled this" stops being a hand-wave and becomes a query.

And then there's what a model swap actually wants. Not "rewrite the trace schema and migrate." But "replay historical prompts against the candidate model in a sandbox, compare outputs, validate the decision layer is stable, then cut over." The chain is the foundation. The four properties make rows comparable across models. The replay de-risks the migration.

A team I worked with last year did exactly this before swapping a customer-facing classification agent between providers. They had three months of prompt and response rows. They re-ran every prompt through the candidate, ran the same decision policy against both sets, quantified the disagreement, then shipped. The migration took an afternoon. No audit-trail project. No dashboard rewrite. The shape carried.

The cost of getting this wrong

Coupling your trace to a model API is a deferred cost that lands the day you swap. And you will swap. Models get deprecated. Vendors raise prices. Regulators push for on-soil inference. Local models catch up. The obvious choice in year one isn't the obvious choice in year three.

Every team I've seen survive a model swap had the four properties in their trace. Every team that didn't was paying down coupling. The work to project the SDK response into a neutral schema is the cheapest insurance you can buy against vendor entropy. Two weeks at design time. Pays back the first migration, twice more after that.

The shape that survives isn't clever or novel. Prompt-as-data, response-as-data, decision-as-data, model-as-attribution-only, four mundane properties, treated as load-bearing, gated through the standards library, enforced at emit. The teams that do this can swap models on a Friday. The teams that don't are rebuilding the dashboard from the last swap when the next vendor announcement hits.

The model is an attribute of a row. Your trace is the foundation. Build the foundation first.

, Sid