How I run AI agents on production code
For six months a fleet of AI coding agents has done most of my production work. The agents were never the hard part. This is the machinery that makes it safe.
Introduction
For the last six months I have run my production work through a fleet of AI coding agents. iOS, backend, web, the occasional Android fix. Several agents at once, each on its own ticket.
People usually ask about the prompts. The prompts are the least interesting part.
An agent writing code is easy to get. An agent you can trust to merge into a production codebase, while you are in a meeting, is a different problem. And it turns out it is not a prompting problem at all. It is a systems problem.
This post is the system.
The mental model
Here is the idea that made everything else click for me.
An AI agent is a process that forgets. And a process that forgets needs an operating system around it.
The agent has four properties that cause trouble, and each one needs its own piece of machinery.
It forgets. The context window fills up, gets summarised, and whatever lived only in the conversation is gone. So the conversation is never the state. Every task has a file on disk with the plan, the investigation, a log of what happened and the next action. The context window is scratch space.
It stops. An agent only works while a turn is running. A turn that ends with “next I will do X” schedules nothing. So the things that must keep happening, merging, releasing, chasing a stuck build, run on timers outside any conversation.
It is confident. An agent will tell you it is done, with total conviction, about code it never ran. So “done” is treated as a claim, and claims need evidence.
It is parallel. Two agents editing the same file is the fastest way to lose an afternoon. So each task gets its own git worktree, and an agent has to claim a ticket before it touches it. The claim is refused if another live agent is already changing the same files.
A rule in a markdown file is a request
This is the most important lesson of the six months.
I started like everyone does: a long instructions file. “Always run the tests.” “Never push to main.” “Get a review before opening a PR.”
The agent followed them. Mostly. And “mostly” is the problem, because the time it skips the rule is exactly the time you are not watching.
A rule in a markdown file is a request. A hook is a decision.
Claude Code can run a script before every tool call, and that script can refuse the call. The model can talk its way past a sentence. It cannot talk its way past a process that exits with an error.
So every rule that really matters became code. The markdown file still explains why. The hook decides whether.
The loop
Every task, on every repo, runs the same loop.
Two steps in it carry most of the weight.
A second model reviews every change
A model reviewing its own work brings its own blind spots to the review. It agrees with itself.
So before a pull request can open, a separate AI session reviews the diff, preferably from a different model family. It reviews that exact commit. If the agent pushes anything afterwards, the review no longer counts and has to run again.
This is enforced by a hook. The command that opens a pull request is refused unless a signed passing review exists for the current commit, written by a different session than the one that wrote the code.
When the review fails, the findings go back to the agent. My teammates only ever see code that has already passed one independent review.
Proof is a merge requirement
A green test suite proves the code does what the tests say. It does not prove the feature works.
So every pull request has to carry one of three kinds of proof in its description:
- A screenshot, for anything a user can see. If the diff touches a view or a layout, text alone is refused.
- The captured output of the real run.
- A red and green pair: the same check failing before the fix and passing after it. This is the strongest, because a check you have never seen fail might not be checking anything.
That last one catches a surprising amount. An agent writes a test, the test passes, and only when you break the code on purpose do you find out the test was never looking at it.
Default-deny around production
Agents are fast. That is the point. It is also the risk.
The line is simple: inside the laptop, agents move freely. Outside it, they stop.
A hook reads every command before it runs. Anything that touches production, writes to a production database, changes cloud infrastructure, or sends a message under my name is put in front of me with a plain-language explanation. If nobody is at the keyboard, it is refused.
There is one more rule that took me a while to see. An agent must not be able to edit its own safety checks. Otherwise the cheapest way to get past a gate is to change the gate.
When something goes wrong
Things go wrong all the time. The interesting part is what happens next.
Every mistake gets the same treatment: find the class of failure, not just the instance, change the machinery so it cannot happen again, and prove the fix by showing the old version failing.
A few of the lessons that came out of it:
A detector with no hands is just a report. I had scripts that could tell when an agent was stuck or a pull request was ready. They wrote a line in a log and waited for me. The moment I stepped away, nothing moved. Now anything that can see a problem also fixes it: a ready pull request gets merged, a finished agent gets closed.
Claim first, act second, record third. A job that sends review requests checked “did I already ask today?”, then sent the message, then wrote it down. When two runs overlapped, both saw “not yet” and the team got the same request twice. The fix is to claim the action before doing it.
A saved state is a memory, and memories get checked. An agent once marked an entire account as out of quota because of one old error message. Five agents then sat idle for hours while that account was working fine. Anything cached about the outside world is now re-tested before it is believed.
The result I am proudest of
The best test of all this was a migration that touches almost every file in an app.
We moved our iOS app to Swift 6 strict concurrency in 10 days.
The agents worked in parallel, each owning one part of the codebase, split by directory so they never touched the same file. Every change was small enough to review in a few minutes. Every one was measured against a baseline, so a change that fixed warnings in one place and created new ones somewhere else was caught before review. Every one went through the second-model review.
At the end: zero concurrency warnings, no big-bang rewrite, regular feature work never stopped, and a test suite 25% larger than when we started, because pinning behaviour with a test before changing it was part of every step.
Wrapping up
If you take one thing from this post, take this: the agents were never the hard part.
What makes agents safe on real code is the same thing that makes any automated system safe. State that survives a crash. Rules enforced in code, not in prose. Proof instead of claims. A human in front of anything irreversible.
It is control-loop design, not prompt engineering.
If you are running agents on a real codebase, I would love to compare notes.
Enjoyed this post?
Subscribe to get new articles delivered to your inbox.
No spam. Unsubscribe anytime.