Build the Harness Before You Hand the Agent Real Work
Agents guess
A model produces the most plausible next chunk of code. Plausible is usually close to correct, and close is the problem. A human who writes a plausible line carries a nagging doubt about it and goes back to check. An agent carries no doubt. It has whatever signal you give it, and if the only signal is “the file was written,” then as far as the agent is concerned the work is done.
So the question for any repo is: after the agent writes a line, what tells it that line is wrong? If the answer is “the reviewer, tomorrow,” the loop is a day long and the agent will stack ten more guesses on top before anyone notices. The harness is the set of checks that turn a guess into a verified claim before the agent takes its next step.
What each layer catches
Every check in the harness catches a different class of mistake, and the cheap ones go first.
Linters catch the sloppy mistakes. Unused variables, shadowed names, an if with no body, a method that’s 80 lines long. None of these are bugs on their own, but they’re the exhaust of guessing, and an agent that gets told about them immediately stops producing them. RuboCop on this site, eslint or clippy elsewhere. Seconds to run.
Type checkers stop nonsense before it runs. Calling a method that doesn’t exist. Passing a string where the function wants a struct. Returning nil down a path the caller never guarded. An agent working in a typed codebase gets a precise, line-numbered correction for each of these. An agent working in an untyped one finds out at runtime, and only on the paths a test happens to walk. Dynamic languages defer this check on purpose, and that was a fair trade while the person writing the code held the types in their head. An agent holds nothing between guesses, and its signature mistake is a method that doesn’t exist. You don’t have to change languages, but you do have to buy the layer back: gradual types where the codebase can bear them, or tests that exercise every call site the agent touched.
The compiler proves the blueprint builds. tsc, go build, cargo build. A green build says nothing about behavior. It does prove that the thing the agent described is a machine at all, that every piece it referenced exists and fits. An agent that can’t get a green build has no business running tests yet.
Formatters keep the codebase looking like one team built it. This one gets dismissed as cosmetic. It isn’t, and it matters more with agents than it did with humans. An agent left to its own devices produces diff noise: reflowed arguments, drifting indentation, quote styles that change file to file. The formatter erases all of it, so the diff you review is only the change. It also deletes an entire category of review comment that no human should be spending judgment on.
A real test bed is what makes a test drive mean anything. Tests are where the agent learns whether its change did what it thought. But the track has to be real. A suite that hits the network and fails one run in twelve gives the agent noise instead of a signal, and the agent will start reasoning around red instead of fixing it. Deterministic, local, fast. Anything else is a test drive on a simulator you don’t trust.
Order and speed decide whether it’s a loop
Run the layers cheapest first, on the changed files, and get the whole pass under a minute. Lint, then types, then build, then the relevant tests. Each layer filters what reaches the next, so the expensive check only runs on code that’s already survived the cheap ones. I wrote about the speed targets earlier this year, and they haven’t moved. A harness that takes ten minutes to answer is a queue the agent waits in, and it will stack more guesses while it waits.
Then take the decision away from the agent. An agent asked to “run the checks when you’re done” will sometimes decide it’s done early. Wire the harness into a hook that fires after every edit, or into a single dev check command the agent is told to run before it reports back, or into a pre-commit hook it can’t bypass. The agent should never get to choose whether the harness runs. It only gets to choose what to do about the result.
The loop sets the ceiling on autonomy
The harness is what decides how much of the build an agent can own. A repo with no types, no lint, and a twenty-minute suite can hand an agent small, supervised tasks, and you’ll read every line it produces because nothing else will. A repo where lint, types, build, and tests all answer in under a minute can hand an agent a whole feature branch, and you can review the outcome instead of the output, because the output-level mistakes were already caught by machines before you ever saw the diff.
Every check you add to the harness is a class of mistake the agent can now catch and fix itself. Every check you leave out is a class of mistake you’re volunteering to catch by hand, forever, at human speed. The tripwire tests I described in the agentic test pyramid are the same idea taken one step further: encode the rules of your codebase as checks, and the agent can’t drift from them without hearing about it in milliseconds.
Build it first
The tempting sequence is to hand the agent real work, watch what goes wrong, and add checks as the damage comes in. That’s backwards, and it’s expensive, because every mistake that reaches you is one the harness would have caught for free. A good SKILL.md is how you get better guesses in. The harness is how you verify what comes out. You need both halves, and the harness is the one you can build before the agent writes a single line.
Build that harness before you hand an agent real work. The tighter the loop, the more of the build it can safely own.