The graduated autonomy pattern across products

The autonomy ladder isn't a per-agent property. It's a per-product, per-context, per-tenant property. The same agent can be Tier 4 in one place and Tier 1 in another at the same time. Here's how to build a platform that hosts the same agent at multiple autonomy levels at once.

The graduated autonomy pattern across products

I keep watching the same conversation play out in different rooms. A team has an agent. The agent has earned a Tier 3 slot in the workflow it was originally built for. Then a second team picks up the same agent for an adjacent job (same model, same code path, same infrastructure) and the conversation immediately turns into a referendum on the agent's tier. "It's a Tier 3 agent." "We can't run it at Tier 3 here." "Is it actually a Tier 3 agent then?"

That's the wrong fight. The agent isn't a Tier 3 agent. It was operating at Tier 3 for one product, in one context, on one tenant, after a specific build of calibration, audit, and rollback infrastructure was put underneath it for that surface. None of that travels into the second product. The rope it earned in one context isn't the rope it's earned in another, and pretending otherwise is how you give an agent autonomy it never paid for.

Here's how I think about it: the autonomy tier is a property of the (agent, product, context, tenant) tuple, not of the agent. The six-rung ladder lives in that tuple. Every promotion and every demotion applies inside one cell of the matrix and tells you nothing about the others.

The matrix, not the label

Pick any agent in production and try to answer "what tier is this agent at." The honest answer is a table, not a number.

For the workflow it was originally built for, in the production tenant, on the action class that's been observed long enough. Tier 3. For the same workflow, on a tenant that signed up two weeks ago and hasn't accumulated enough audit volume to calibrate confidence. Tier 2. For an adjacent product that uses the same agent for a structurally different action class. Tier 1, until that product builds the rollback infrastructure for its own action surface. For a Sunday-night experimental workflow someone in a different team wired up last sprint. Tier 0, because nobody has built any of the supporting structure yet.

Same agent. Four tiers, simultaneously. None contradict each other, because each reflects what's been built underneath the agent in that cell, not what the agent is capable of in the abstract.

The instinct teams resist most is that the agent's capability and the agent's authorized tier are not the same thing. A capable agent isn't a high-tier agent until the surrounding infrastructure has been built and the trust earned, and both live per-cell. That's been true of every production system I've operated in twenty years, a service is permitted in production because the operations around it are good, not because the service is, and AI agents are no exception. They're just the first systems where the marketing tries to make capability and authority sound like the same word.

Per-context tier registries

The implementation pattern that drops out of the matrix framing is a per-context tier registry. Not a config file with tier: 3 on the agent. A lookup table keyed by (agent_id, product_id, context_id, tenant_id) that returns the tier the platform should enforce for any specific invocation.

The lookup happens at the moment of invocation, not at deploy time. The agent doesn't know its tier. The platform looks up the tier for the cell the call belongs to and routes the output through the path that tier specifies. Tier 3 goes straight to the action layer; Tier 2 into the confirmation queue; Tier 1 into the draft folder. The agent code does not branch on this. The platform does.

Teams skip this because it sounds heavy. The first version is always the agent calling its own action layer with a tier baked in. That works for one product and one tenant; it falls over the moment the second product wants the agent on a different tier, because there's no place to vary the routing without forking the agent. By the time the team realizes they need the registry, they've forked the agent three times and the question of what tier the agent is at has become which fork you're calling. The registry is what you build before that fork happens.

The registry should carry, per cell, at minimum: the current tier, the policy version paired with it, the rollback contract for actions at that tier, the report stream the cell publishes to, and the date and reason of the last tier change. That last field is the one teams underweight; it's how the on-call engineer paged at 2am answers "why is this cell at Tier 3" without hunting for the original promotion ticket.

Per-tenant overrides

The cell that matters most in practice is the tenant cell. Two tenants of the same product, with the same agent doing the same job, can sit at different tiers without contradiction. Tenant A has been on the platform for eighteen months, has a clean audit history, has signed off on the rollback contract; the agent runs at Tier 3 for them. Tenant B onboarded last month, has no operational history, hasn't read the rollback contract; the agent runs at Tier 2 for them, with a per-action confirmation step Tenant B's human-in-the-loop absorbs.

Per-tenant overrides aren't a discount or a premium tier. They're an honest reflection that trust is built per relationship, and a fresh tenant hasn't built any yet. The autonomy the agent earned in the combine doesn't transfer; it has to be earned again, on this tenant's data, against this tenant's recovery surface.

The override has to live in the registry, not in the agent prompt and not in a per-tenant fork. A default tier on the (agent, product, context) cell and an optional override on the (agent, product, context, tenant) cell that takes precedence. A tenant graduating from Tier 2 to Tier 3 should be a registry update, not a deploy.

Treat the per-tenant override path as a first-class promotion path, with the same gates as the original product-level promotion. Calibration evidence on this tenant's data. Audit completeness on this tenant's actions. Rollback rehearsal on this tenant's environment. The cliff between Tier 2 and Tier 3 doesn't get smaller because the agent has crossed it elsewhere. It's a fresh cliff for each tenant, even though most of the underlying infrastructure can be shared.

Observability that names the cell

The hardest operational question in a graduated-autonomy world is "what tier was active for this specific action." If the answer requires an engineer to reconstruct the cell from the agent's code path and the policy file at the moment of action, the answer doesn't exist in any usable form. The audit trail has to name the cell.

Every action gets a record. The record carries the cell (the (agent, product, context, tenant) tuple) the tier active at the moment the action fired, the policy version paired with it, the cell's rollback contract version, and the report-stream destination. Five fields, no inference at read time. An on-call engineer reviewing an incident at Tier 3 on Tenant A doesn't need to ask whether some other cell was at Tier 2, the record says which cell this was, and the dashboards filter on cell, not on agent.

A view that combines "the agent's confidence distribution this week" across cells is dangerous, it averages a Tier 3 cell with confident actions against a Tier 1 cell with exploratory ones. The honest view is one chart per cell (confidence, action volume, override rate, rollback frequency) and a meta-view that ranks cells by which need attention soonest. That becomes the operating cockpit across all the agent's deployments.

The same surface is what makes the downgrade discipline work in a multi-cell world. A signal that fires on one cell should drop that cell's tier, not the agent's. Tenant B's audit deviation is not Tenant A's problem; the per-cell registry update isolates the response. Without it, the options are "downgrade the agent everywhere" or "ignore the signal," and both are wrong.

The shape of the platform

Pull the pattern up a level and the platform that hosts a graduated-autonomy agent has a recognizable shape.

A single agent runtime. The agent code lives in one place; the model is the same instance for every cell; updates to prompt or weights happen once and propagate. The cells aren't forks, they're configuration around the agent.

A tier registry, addressable by cell, that any production action looks up before it lands. Versioned, audited, runtime-mutable, observable.

A routing layer between the agent's output and the action surface that consults the registry and dispatches through the appropriate tier-specific path, direct execution, confirmation queue, draft folder, suggestion stream. The agent does not own this routing.

An audit pipeline that captures, per action, the cell, the tier, and the policy version. Queryable along all those dimensions, not just along agent identity.

A per-cell dashboard surface and a meta-view that ranks cells by the signals that should drive tier changes. The on-call engineer sees which cells are healthy, which are drifting, and which need a downgrade.

A tier-change workflow that operates per cell, with the same promotion and demotion gates the six-rung ladder prescribes, applied locally, one cell at a time.

That's the platform. It looks more complicated than the single-tier-per-agent picture, and it is. The complication is real because the underlying reality is real. The agent is doing different things, for different consumers, with different histories, on different surfaces. Pretending one tier label can describe all of that is the simplification that makes the shipping conversation feel cleaner and the operating conversation worse and worse.

Want to go deeper on the operational side? The downgrade pattern spells out the runtime move that drops a cell's tier when a signal flares, which is the only thing that makes per-cell tier registries safe to operate.

What this changes about how I talk about agents

The framing I'm trying to retire is "this is a Tier N agent." It's a category error and it's caused more confusion in agent operations than any other shorthand I can name. The framing I want in its place: "here's the agent; here are the cells it's deployed in; here are the tiers each cell currently sits at." Long, ugly, accurate.

The agents I trust most in my own operation are the ones whose registry is dense, multiple cells, multiple tiers, with downgrades and upgrades scattered through each cell's history. That density is the sign the platform underneath is paying attention to the surface the agent is acting on, not just to the agent. Bounded autonomy isn't a property the agent earns once and carries forever. It's a property the platform maintains, per cell, with the agent running through it.

The work isn't building smarter agents. It's building the platform that can host the same agent at five different tiers without losing track of which is doing what.

, Sid