Article 7: Test everything
Article seven of my personal coding constitution: tests are a property of the work, not a phase of it. The minimum viable test is 'does the documented behavior actually happen?' Anything less is hope.
The seventh article in my personal coding constitution is a rule about evidence, again, but at a different layer than article five. Article five says: every action that mutates state must leave a trail. Article seven says: every claim the code makes about its own behavior must be checkable, on demand, by something that isn't me.
Test everything.
Not "test the important stuff." Not "write tests for the gnarly bits." Everything. Every claim the code makes, every documented behavior, every "this function returns X when given Y" sentence in the README, all of it has a matching executable artifact that confirms the claim is currently true.
I had to write this one down because the failure mode it prevents is the one I'm worst at noticing. Untested behavior isn't behavior. It's a guess, written in source code, that nobody has verified since the day it was committed. The constitution treats guesses and behavior differently, and the rule is what keeps the difference legible.
Tests are a property, not a phase
The thing I had backwards for the first decade of writing software is that I treated testing as a phase of the work, the thing I did after the build, when there was time. The phase rarely arrived. When it did, the tests were thin and ceremonial, because I was testing what I'd already convinced myself was working, not the actual contract the code was supposed to satisfy.
The reframing is small and total: tests are a property of the work, like the function signature or the file the code lives in. Untested code is incomplete code in the same way uncompiled code is incomplete code. Not "needs polish" incomplete. Not done incomplete.
This isn't dogma. The shift is operational, not religious. When tests are a property, the question "is this done?" has a single answer. When tests are a phase, the question has two answers ("the code is done, the tests aren't yet") and the second half never gets done because the first half already shipped.
The reframing is the rule. Everything downstream is mechanics.
The minimum viable test
People over-engineer this rule the moment they hear it. They imagine the discipline requires property-based testing, mutation testing, ninety-percent line coverage, a CI pipeline that runs for forty minutes. None of that is the rule. Most of it is theater built around the rule.
The minimum viable test is one sentence: does the documented behavior actually happen?
If the README says the function returns the user's email when given a user ID, the minimum viable test calls the function with a user ID and checks that it returns the email. That's the whole test. It doesn't have to be elegant. It doesn't have to handle every edge case. It has to confirm the one claim the code makes about itself.
The rule isn't "test thoroughly." The rule is "every claim has at least one test." Thoroughness is a separate conversation about coverage and risk. The minimum is much smaller and much more stubborn: no claim without a matching check.
Anything less than the minimum viable test is hope. Hope is what a system runs on when nobody has bothered to confirm it works. Hope is fine for a hobby. It is not the foundation I want under the platform.
Want the deeper version of this? The agent-side discipline (handing the agent an assertion-shaped contract before it writes a single line) is the one I leaned on the most when I started working with AI on real code.
Why retroactive tests almost never happen
The thing I observed that forced me to make this an article instead of a habit: in my own work, across years of trying. I have almost never successfully written tests for code I shipped without them.
The promise ("I'll add tests later") has the same arc every time. The code ships. It works, mostly. The "add tests later" task moves down the backlog. Six weeks later, the code has been changed three times by three different people (or three different agent sessions) and the original behavior is no longer the current behavior. The question "what should the test assert?" is now unanswerable, because the contract drifted and nobody recorded the drift. The retroactive test, if I write it at all, ends up testing what the code currently does, not what it was supposed to do, and that's not a test. That's a snapshot.
The honest version of "I'll add tests later" is "I won't." Knowing this about myself is what makes the rule load-bearing. If I'm not going to add tests later, the only honest option is to add them now, before the code. Not because test-first is an approach I subscribe to. Because it's the only mode in which the test actually gets written.
Test-first as a thinking tool, not dogma
I want to be careful with the test-first framing because it carries methodology baggage I don't endorse. The TDD red-green-refactor cycle, the kata, the dogma about test-driven design as a way of life, none of that is what this rule asks for.
The rule asks for something narrower. Before I write the build, I write the assertion. Not the whole test. The single sentence in code that says "this is what done looks like."
Writing the assertion first does one specific thing: it forces me to commit to the contract before I commit to the code. When I write the code first, I conflate the two, the contract becomes whatever the code happens to do, and the test (if I write one) ends up restating the code in different words. When I write the assertion first, the contract is fixed, and the code has to satisfy it instead of inventing it.
This is why I call it a thinking tool. The act of writing the assertion is the act of deciding what the code is for. That decision is the hard part. The code, once the contract is named, is usually the easy part. It pairs almost perfectly with article one, the issue holds the intent, the assertion holds the testable form of the intent, the code satisfies the assertion. Three artifacts, one chain, no missing links.
What "everything" actually means
The most honest objection to "test everything" is that "everything" is a moving target. Test every line? Every branch? Every possible input? At some reading, the rule is impossible.
The answer is narrower than people expect. "Everything" means every claim the code makes about itself. Not every line. Not every branch. Every documented, externally-visible, contractual behavior.
A claim is anything a future reader would treat as a guarantee. The function signature is a claim, when it says it returns a User, it's claiming it returns a User. The README is a claim. The error message is a claim. Every one of those claims needs an executable check.
Internal helpers nobody outside the module ever calls don't need their own claims tested directly, the claims they support get tested through the public surface they serve. The rule isn't a coverage metric. It's a contract-enforcement principle.
"Every claim" is a much smaller set than "every line," and a much larger set than "the gnarly bits I happen to feel uncertain about." The rule sits in the middle, and the middle is where the discipline actually does work.
The agent version
When I hand an agent a task and it produces a diff, the question I have to answer before merging is: does the diff satisfy the contract? The agent will tell me yes. The agent's confidence is not evidence. The diff being well-formatted is not evidence. The code looking sensible is not evidence. The only evidence that does the job is a passing test that exercises the documented behavior.
This is why I now hand agents the assertion before the build, whenever I can. The prompt isn't "implement this function." The prompt is "implement this function such that this assertion passes", and the assertion is in the prompt, in code, in the form the runtime will evaluate. The agent's job is to satisfy the check. My job, when the agent says it's done, is to run the check and see whether it actually passes.
The assertion-first discipline closes the gap I've lived through too many times: the agent produces a confident, well-structured build that looks correct, I skim the diff, it ships, and a week later the contract turns out to be subtly different from what the agent guessed. With the assertion in the prompt, the agent doesn't get to guess the contract. The contract is executable and the test result is binary. The agent's confidence stops mattering. The check matters.
Where the rule pairs with the others
Article seven is the rule that makes the other rules verifiable. Article one is "track the intent" with no way to check whether the build matched it; the test is the check. Article four's three-strikes rule presumes you can tell when an attempt failed, which requires a test that defines what success would look like. Article five's trail records what happened; the tests record what was supposed to happen. The two together are how you tell whether the system is doing what it claimed.
It pairs with article six in a way I didn't expect. The override for article seven ("this code ships without a test because") is the override I'm most suspicious of in myself, because it's the one whose justification I'm best at inventing in the moment. The protocol is built to make me write the justification down anyway, and the justification, written down, is usually thin enough that I end up writing the test instead. About a third of the time, again. The override doing its job in the other direction.
What the rule is really about
The article looks like it's about testing. The rule underneath is about the difference between thinking it works and knowing it works. Code I think works is not the same as code I have evidence works. The first is hope. The second is engineering. The constitution is built to make me operate in the second mode by default, even when the first is faster, even when the agent has produced something that looks fine.
Tests are the cheapest evidence I know. They cost a few minutes per claim, run in seconds, and survive everybody, me, the agent, the file being rewritten. Without them. Every change is operating on a contract reconstructed from the code, which is exactly how the contract drifts and the system stops being knowable.
So: test everything. Every claim, every documented behavior. The minimum viable test is the one that confirms the claim is currently true. Anything less is hope, and hope is not a property the system can rely on.
, Sid