Code got cheap. Judgment didn't.
In August, Anthropic published a playbook for the software development lifecycle when agents do the building [1]. It's a 46-minute read, and the observation that opens it is one I recognize from my own seat. Agents write code in hours. Planning, review and deployment still run at human speed. So the bottleneck didn't disappear. It moved, from the engineer's keyboard to the reviewer's queue, and a queue that can't keep up has two outcomes: it grows, or code goes out under-reviewed.
Their answer is to rebuild the six stages as a loop. Every stage commits one file to git, the next stage reads it, and a human holds a specific gate at each turn. I turned the playbook into a two-minute walkthrough, and then into the thing I wanted more: a run of the loop you can play, where an agent does the work and you make the six calls. Both are linked at the end. This note is what I took from building them.
Every stage leaves a file. The files are the control.
Follow one request through the loop and the shape is simple. Someone with an idea explains it to the agent, which writes it up as a one-page brief, intent.md, that a product owner approves. The agent turns the brief into a spec, spec.md, while the company's rules for security, brand and accessibility are applied as it writes, so a problem surfaces in the draft rather than in a review six weeks later. In the build stage the agent writes plan.md first and can't touch a file until an engineer approves it. Then it writes code, runs the tests, fixes what fails and runs them again until they're green, with the test files locked so a failing test can't be made to pass by rewriting it. A set of real tasks runs as evals in CI and the pass rate gates the merge. The agent reviews every pull request and can never approve its own. A code owner merges. Production stays locked until a release manager signs off. And when something drifts at 3 a.m., a deterministic detector, no model involved, calls the agent to diagnose and propose, and the proposal is written as a new intent.md that lands back at the start of the loop with a human owner.
Read it as an engineering practice and it's a sensible way to work with an agent. Read it as a control design and it's something I've been arguing for since the first of these notes.
Policy as code, evidence at execution, autonomy by environment.
Risk at Runtime drew a line between a control that lives in a policy document, which is an intention, and a control expressed as code that executes on every relevant action and can't be skipped by a busy human. The playbook draws the same line with different words. Skills are rules the agent follows most of the time. Hooks are scripts that block an unsafe action every time, and the agent is told why. Skills make mistakes rare. Hooks make them impossible. That's policy as code, running in the path of the work.
The same note argued that every execution should self-evidence, so that nobody assembles an evidence package before an exam because the evidence already exists. Look at what the loop leaves behind: the intent, the spec, the plan, the diff, the test output, the review findings, each committed with an author and an approver. Together they're the audit trail, and no one wrote it as a separate deliverable. It's the git history.
Two more mappings. The evals that run in CI on every configuration change are control testing on the full population rather than a quarterly sample. And the tiers of autonomy by environment, where the agent deploys freely to dev, staging sits in between, and production opens only on a human signature, are earned autonomy in its most literal form.
I didn't expect to find my own argument laid out as a developer workflow. I take it as a good sign for the argument.
Six decisions. One owner each.
Decision Integrity has an Ownership test: does one person own this decision and its consequences? The loop passes it at every turn, and that's the part I'd want any executive to see before the tooling.
For a regulated institution these land on controls that already exist and already have owners. The author can't approve its own change: segregation of duties. Production opens on a named signature: change management. Output attached to the pull request: evidence retention. A triage queue with an owner: incident management. The playbook leaves the control regime where it is and gives it something it has rarely had, which is evidence produced by the work itself at the moment the work happens.
The scarce skill is approving well.
When build collapses from weeks to hours, the constraint becomes the capacity of a human to decide well with the evidence in front of them. The review queue that piles up in the playbook's opening is a judgment-capacity problem wearing a tooling disguise. That has two consequences I'd plan around.
The first is that every artifact has to be readable at the gate in minutes. The playbook is quietly insistent on this: a one-page intent, a plan that says which files, in what order, with what risks and what proof, a CLAUDE.md kept under a page. Those limits aren't style. They're a design for human attention, because the person at the gate is the slowest component in the loop and the only one that can't be scaled by adding sessions.
The second is where the work goes. In the run I built, the request moves from a note in fraud operations to a production incident writing its own follow-up in about three days of simulated time, against fourteen weeks of illustrative old-way durations. The agent did the typing. The human made six calls, and each one is on the record with a name. If that's the shape of the job, then the job of a product owner, an engineer and a release manager tilts toward the gate, and the organizations that get this right will be the ones that treat approving well, fast, on evidence, as a skill to hire for and train, not an interruption to the real work.
Three things before I'd run this in a bank.
The playbook is written for engineering teams, and it's good at that. The governance object is still implicit, and I'd make it explicit in three places.
The grant. Governing Digital Labor argued that the thing to govern is the delegation of authority to a non-human principal: what was delegated, to what, by whom, under what bounds, on what evidence, until when. In the loop, an engineer steers a session, and there's no record of what authority the session was given or when it lapses. Each session should run under a named delegation with a scope and an expiry, and that record should sit next to the commits it produced.
Coverage. The loop governs what's inside it. An agent session running outside the loop, writing to a shared store nobody watches, is exactly how the July incident I wrote about in Authority Provenance began. Coverage of the runtime rail is itself a control. An agent workload the loop can't see is a finding, not an unknown.
Stop authority. The playbook has gates that open. It needs a named person who can stop the loop, a default that stops on a timer if that person can't be reached, and a separate decision to restart. Gates that only open are half a control.
And one scale note. One engineer running three sessions is a team. A thousand engineers running three sessions each is a population, and the skills, hooks, evals and control bands that govern it are shared configuration. A change to a hook changes every session at once. The population has to be monitored as a population, which is the instantiation point from Governing Digital Labor arriving in practice.
Two minutes to read, five to run.
The walkthrough is at allynshaw.com/sdlc. It follows one request through the six stages in about two minutes, with the commit trail building as you go.
The run is at allynshaw.com/sdlc/run. It's set in a bank. Pick a scenario first: a blocked card payment, a mortgage application, or a deposit on hold. An agent writes the brief, the spec, the plan and the code. You hold the gates. It asks for your name, which goes on the commits, and it hands you a run report at the end: the audit trail, your decisions, and every file the loop produced. Three of the decisions have consequences later in the run. Make the wrong one and the loop will tell you, which is rather the point.
Both are prototypes, and the scenarios are illustrative. The question I'd ask anyone who finishes a run is the one the playbook ends on: which stage is slowest where you work? I suspect it's one of the gates, and I'd like to hear which.
What this borrows.
From Risk at Runtime: controls as versioned, executable code; evidence generated at execution; autonomy earned through evidence gates. The playbook is the first published workflow I've seen that does all three by default.
From Decision Integrity: the Ownership test, applied here to six gates, and the distinction between deliberate and silent debt, which is what the triage decision is really about.
From Governing Digital Labor and Authority Provenance: the grant as the object of record, coverage as a control, and a named stop authority. These are the three things the playbook leaves implicit, and the three things I'd add before running it where I work.