DaC for AI agents: the decision surface that points at a model

When the action behind a decision is 'ask an LLM,' the decision surface is what protects you from model drift. The five questions, the bounded scope, the audit trail, the ownership boundary, they all still apply. The DaC surface becomes the contract that makes AI agents safe to swap and govern.

DaC for AI agents: the decision surface that points at a model

The first time I shipped a Decisions as Code surface where the build behind one of the decisions was "call an LLM," I got nervous in a way I hadn't been about a DaC surface before. The decision was small, classify an inbound support message as billing / technical / sales / abuse, route accordingly. The shape was identical to one I'd shipped a hundred times: bounded inputs, enumerated outputs, a clear owner, an audit trail. Where I'd previously written case statements. I was now sending a prompt to a model and trusting the response.

What surprised me, six weeks in, was how not different it actually was. The DaC surface absorbed the implementation choice cleanly. The five questions still applied, the audit trail still answered them, and the ownership boundary still held. What had changed was the validation discipline at the seam between the decision and the model. You can't unit-test a model's behavior. You can pin its inputs, contract its outputs, monitor its drift, swap it on a schedule, but you can't write assert classify("...") == "billing" and have that be the test that protects you.

Agent-driven decisions are still DaC decisions. The surface, the contract, the audit, the boundary, all of it survives the swap from deterministic build to model call. What you add is a different validation discipline at the seam, and a discipline about what level of autonomy the agent is operating at when it makes the call. Get those right and the DaC surface becomes the thing that makes AI agents safe to swap and govern.

The five questions still apply

Every DaC surface I've ever shipped has had to answer the same five questions, regardless of foundation. What decision is being made? Who owns it? What inputs feed it? What outputs does it produce? What's the audit trail? Those five are the irreducible shape. If the surface can't answer them cleanly, it isn't curated; it's just configuration with optimistic naming.

When the build is an LLM call, every one of those questions is still answerable, often more answerable than with a tangle of nested if statements, because the model call forces the decision to be expressed declaratively. There's a prompt, an input contract, an output schema, a model identifier. The surface can name those things plainly.

Take the support classifier. What decision? Route an inbound message. Who owns it? The support ops lead, on the hook for routing accuracy. What inputs? Message body, customer tier, recent ticket history. What outputs? One of four enumerated buckets, plus a confidence score, plus a "needs human review" flag. What audit trail? Prompt sent, model identifier, response, routing applied, timestamp, any human reviewer. All five answerable. The fact that the build in the middle is an LLM is a footnote.

The pattern that breaks the surface (the one teams reach for when they first put an agent behind a decision) is to skip the curation. The surface becomes "the agent will figure it out," inputs are "whatever the agent can see," outputs are "whatever the agent decides." None of the five get a clean answer. The surface isn't a contract; it's a hope. The wrongness has nothing to do with the model and everything to do with abandoning the curation DaC has always demanded.

The bounded scope is doing the work

The reason DaC surfaces hold up under an LLM build is the bounded scope. The agent is not "the support agent that handles support." It's "the classifier that decides which of these four buckets a message belongs to, based on these inputs, returning this schema, with this confidence threshold." That bound is what lets the rest of the system (routing, escalation, audit, SLA) stay deterministic and auditable around the model call.

This is where the platform / product boundary shows up. That the message gets classified into one of those four buckets is a product decision. Which model performs the classification, supply-chain controls, eval pipeline, what happens during a model outage, platform decisions. Product picks "I want a classifier with these inputs and outputs at this latency budget"; platform supplies the model, prompt template, eval harness, fallback path. Same DaC contract pattern, costumed for the AI era.

If the agent's scope creeps (if the prompt evolves to "and also handle these other intents") the bound gets violated and the surface stops being a contract. The fix isn't a smarter agent. The fix is to recognize you've got a different decision (or several), each deserving its own bounded surface, owner, and eval pipeline. Keep the bound tight enough that the model does one nameable thing per call.

The audit trail is more important, not less

When the build was deterministic, the audit trail was a courtesy, produced because compliance asked, rarely used in operations because you could re-derive the decision from inputs and rules. With an LLM in the middle, you cannot re-derive the decision. The model is non-deterministic in ways that range from mild (sampling temperature) to severe (model version changes you don't control). The only way to know why the agent decided what it decided is to have captured the decision faithfully when it fired.

What I log, without exception: prompt template version, resolved prompt with variables interpolated, model identifier with version pin, inference parameters, raw response, parsed structured output, schema-validation result, confidence score, action taken, timestamp, requesting identity, any human review downstream. It's a lot. It's also the only thing that makes the decision investigable after the fact, because any of those fields can be the variable that explains a decision that looked wrong in retrospect. Without the full trail you spend a week guessing. With it you find the variable in an afternoon.

This is where the OPA Gatekeeper enforcement layer earns its keep. The admission-time policy that says "every workload calling a model must emit these audit fields against this standard schema" is what keeps the trail honest across the fleet. The DaC surface specifies what the audit trail looks like; the policy engine verifies that every build actually emits it. Specify, then verify. Same loop the approach has always demanded.

What changes is the validation discipline

Here's the part that really is different. With deterministic logic, validation happens at unit-test time, and the tests are stable across runs because the logic is stable. With an LLM, neither half is true. You can't unit-test the model's behavior because it's a probability distribution, not a function. You can't integration-test it once and trust the result, because the model can drift between deploys (and the model behind your endpoint can change without your deploy at all).

The validation discipline that replaces unit tests is an eval pipeline. A curated set of golden inputs with expected outputs or ranges, re-run on a schedule against the production model and any candidate replacement. Above a drift threshold, the agent gets demoted from whatever autonomy rung it was operating at to a lower one (back to "execute with confirmation" or "draft for review") until the drift is investigated and either explained or fixed. The eval pipeline is the validation surface; the demotion mechanism is the safety net.

The other piece is production observability. Every agent-driven decision produces telemetry: confidence distributions, schema-validation pass rates, downstream-correction rates, escalation rates. Those get watched the way error rates get watched on a deterministic service. A sudden drop in confidence or a spike in human override is a model-behavior incident, even if no error was raised.

What this buys you is the ability to swap models without changing the surface. Llama on Tuesday, an open-weights fine-tune on Wednesday, a frontier-closed model on Thursday, if all three pass the eval pipeline and produce the expected telemetry profile, the surface doesn't care. The product team doesn't care. The audit trail doesn't care, because the model identifier is a field in the record. That is the contract working as intended.

Want to go deeper on the autonomy ladder? The six rungs piece spells out the rungs and what promotes (or demotes) an agent between them. The eval-pipeline-as-demotion-trigger pattern lives at the seam between this article and that one.

The ownership boundary holds

A failure mode I keep seeing is the assumption that "AI work" needs a different ownership model. There's a temptation to spin up an "AI team" that owns everything model-shaped and let product teams hand requirements over a wall. It doesn't work, for the same reason a generic platform team doesn't work: the team that owns the decision can't be the team that doesn't see the consequences.

The boundary on agent-driven decisions is the same as everywhere else. Product owns the decision. Platform owns the foundation, which models are approved, supply-chain controls, eval pipeline, observability, fallback during outage. The interface is the DaC surface.

The conversational test from the original DaC piece still applies. Sit the product owner down in front of the surface and ask them to walk through it. They should read it in five minutes and tell you what they're choosing. "Four buckets. This confidence threshold. Escalate below threshold. This approved model from the catalog." If the surface forces them to think about prompt engineering, sampling parameters, or training data, the platform is leaking. Push it back below the line.

The contract that makes agents swappable

The reason this matters operationally is that the DaC surface is what makes agents safe to swap. Every six months, somebody on the platform team will want to retire the model behind one of these surfaces, for cost, performance, a security advisory, a vendor change, a fine-tune they think is better. The product team that consumes the surface should be insulated from that swap. They wrote against the contract; the contract should hold.

What lets the swap happen safely is the discipline above, combined. Bounded scope means the new model only has to be good at one nameable thing. The eval pipeline lets you verify it on the golden set before cutting over. The audit trail lets you compare new and old decisions on the same inputs and quantify the change. The autonomy rung lets you drop the new model to "draft for review" while it builds trust, then promote it as eval data accumulates. The DaC surface stays stable. The build rotates.

Same shape DaC has always promised. In the OneFuse era, the foundation was vRA7 vs vRA8 vs Terraform; the contract was the Property Toolkit. In the K8s era, the foundation was Helm vs Crossplane; the contract was the standards library. In the agent era, the foundation is whatever model is behind the call this week; the contract is the DaC surface, the eval pipeline, the audit schema, the autonomy bound.

Five questions. Bounded scope. Audit trail. Ownership boundary. New seam discipline at the model. Swap freely below the line. The model is just the build. The decision is still yours to own.

, Sid