The approve/deny gate: and when it goes away
Day-one, every AI suggestion goes through a human approval. Then you mine the approvals, find the safe classes, let those auto-resolve, and keep the audit trail through the whole transition.
There's a moment in every AI product's life when the founder looks at a queue full of human-approved suggestions and asks: do we still need the human? The honest answer is "for some of these, no, and for some, never." Knowing which is which is most of the job.
Layman version. An IT operations consultant has packaged her decade of triage instinct into an AI helpdesk product. A ticket lands: the office printer keeps going offline. The AI proposes: restart the print spooler on the front-desk PC, reseat the network cable. On day one, that proposal goes to a queue, not the customer. A tech looks at it, nods, approves. The customer gets the answer. Three weeks later the consultant notices this exact suggestion has been approved 47 times, denied zero, with no follow-up complaints. That pattern is the signal. This class can graduate. From now on the AI's answer goes straight to the customer; the human reviews a sample.
That arc (suggesting → acting) is the whole point of the approve/deny gate. The gate is not a permanent tax. It's a learning instrument. You build it on day one because you have to. You let pieces of it dissolve over time because the data tells you they can. And the audit trail survives the transition because you designed it to.
Why the gate exists on day one
Two reasons, both non-negotiable.
You don't yet know how the AI fails. Every AI product I've shipped has had a failure mode that wasn't in any pre-launch eval. Not because the eval was bad, because the world is bigger than your test set. The first hundred real interactions are where you learn what the model actually does in your domain. A human between the AI and the customer during that window is how you learn without burning customers.
Then the audit story. Who decided this? On what evidence? When? On day one the answer should always be "a named human." A product that day-one auto-resolves anything has nobody to point at when something goes wrong.
So: gate. Every action. Day one.
The mechanics are simple. The AI runs through its triage-diagnose-resolve loops (covered in the three loops piece) and produces a proposed action with a confidence band, a rationale, and the evidence it used. The proposal lands in a queue. A reviewer sees it, with approve/deny buttons. Both outcomes get logged with reviewer identity, timestamp, and full context.
Both outcomes. Most teams log the approves. The denies are where the gold is.
Pattern-mining the approvals (and the denies)
After a few weeks, your database knows things nobody else does. Which classes the AI handles well (high approval rates, fast reviews, no follow-up complaints), which classes it fumbles (high deny rates, slow reviews, lots of edits before approval), and the gray middle where humans approve with hesitation visible in the timing data.
Mining that table is how you decide what graduates.
For an HR consultant who's productized her interview-rubric scoring: applications scored "strong yes" with all four criteria present and no flags get approved 98% of the time, in a median of 12 seconds, with zero post-approval reversals across 200 cases. That's a class. Narrow, well-defined, and the human review is performative, every reviewer just clicks approve. Those clicks consume reviewer time that should go to harder cases.
Meanwhile "weak yes" applications with one criterion missing and a tone flag get approved only 60% of the time, take 4 minutes on average, with a 12% reversal rate. That class is not graduating anywhere. The AI's confidence isn't justified by the outcomes.
The pattern-mining isn't sophisticated. The first version I built was a Postgres view with three columns: input category, AI proposal, human action. Read the rows by hand, eye the rates, pick the obvious classes, write the rule. Later you can make it a formal classifier, but the manual phase is more honest. You see what's actually in the queue.
The graduation rule
The rule I now use for letting a class auto-resolve is deliberately conservative.
A class is eligible when all of these are true: approval rate over the last 90 days above 95%; volume of at least 100 cases (so the rate isn't a small-sample artifact); post-approval reversal rate below 1%; the deny reasons that did show up are not about safety or correctness (scope, formatting, stylistic preference); and the consultant whose secret sauce drives the product has personally signed off.
That last bit matters. Graduation isn't a system decision. It's a human decision informed by data the system collected. The system makes the decision easy to make and easy to defend.
When a class graduates: future cases skip the queue and ship the action directly; a sampled audit kicks in (one in twenty cases still gets a post-hoc human review); the audit table gets a new field recording whether this row was human-approved or auto-resolved, which rule version, signed by whom. That field is the bridge between "the AI did it" and "a named human authorized the AI to do it under these conditions."
The audit trail through the transition
This is the part most teams botch.
The audit table on day one: case ID, input, proposed action, evidence, reviewer ID, decision, timestamp. Reviewer ID always a real human.
On day 200, after several classes have graduated: same columns plus decision-maker (human or rule), rule ID and version, graduation authority (the consultant who signed), sampled review (yes/no, reviewer ID if yes). Every row is still answerable. Every row still has a chain of accountability. The chain just runs through a versioned rule signed by a named person, instead of a live reviewer.
Two things to insist on.
Graduation rules live in the same versioned, reviewed place as your prompts and eval cases. Every rule change goes through PR. I keep mine as small declarative YAML files alongside the prompts. (See prompts as code for why.)
The sampled-review path feeds back into the eval harness. When a post-hoc reviewer finds an auto-resolved case that should have been denied, that case becomes a golden example, the rule gets re-evaluated, and if the failure rate creeps above the threshold, the rule gets pulled. The eval harness is what makes graduation safe. Without it, graduation is just deletion of safety.
For the regulated-vertical reader. In medical, legal, financial domains, the graduation rule may need to stay shallow even when the data says it could go deeper. A medical specialist running a second-opinion review product might decide no class auto-resolves, ever. The gate doesn't have to dissolve, it can just get faster: better summaries, evidence presentation, keyboard shortcuts. Speed of human approval is a separate axis from removing it.
The queue UI
A bad queue is a wall of text with two buttons. The reviewer skims, gets bored, starts clicking approve to clear the backlog. The audit trail says "approved by Jane" but Jane is a rubber stamp.
A good queue is a one-line summary enough to evaluate easy cases at a glance, everything else collapsed but one click away. Approve is the default-focus button. Enter ships it. Deny opens a small box where the reason is required. Edit-and-approve is a third option, captured as "human modified the AI's proposal before shipping", those cases are gold for pattern-mining (systematic edits signal a prompt fix).
I track median time to review by class. Dropping because cases are obviously fine, graduation candidate. Rising because cases are getting harder, the AI shouldn't be touching those at all, and triage needs to route them elsewhere.
When the gate doesn't go away
Some classes never graduate, and that's the right answer.
Anything where the cost of a wrong answer is asymmetric and large, a contract going out under wrong terms, a financial recommendation a customer will act on, a medical interpretation affecting treatment, stays gated. Even if the AI is right 99% of the time, the 1% is too expensive to absorb. Mark these manually held and never auto-graduate, no matter what the numbers say.
Anything where the input distribution is unstable stays gated until it stabilizes. Strong historical approval rates with a recent spike in denials means the world moved. Graduate after the shift has settled.
Anything where the consultant's brand depends on the human touch stays gated as positioning. Some products are the human review; the AI's job is to make it faster and better-evidenced, not to replace it. That's a fine business model. The gate isn't a failure state. It's the product.
The arc, plainly
Day one: gate everything. Both outcomes logged.
Weeks one through twelve: mine the queue. Find classes where approval is reflexive and reversal is rare. Confirm with the consultant. Write the graduation rules.
Months three onward: confident classes auto-resolve, with sampled audit. Hard classes stay queued. The gate has gotten thinner, not absent. Every action (auto or human) is still answerable in the audit table.
Year two: the queue is small. The reviewer's role has shifted from "decide every case" to "decide the hard cases and supervise the rules." The audit story is stronger than on day one, because now you can show not just the decisions but the rules behind the decisions, the human who authorized each rule, and the sampled-audit data proving each one continues to behave.
The AI graduates from suggesting to acting. The human graduates from deciding to supervising. The audit trail does neither, it stays the same shape, the same "named human, named evidence, named decision" all the way through. The trail is what makes the rest safe. Build it that way on day one, and graduation is a feature you ship; build it any other way, and graduation is a story you can't tell.