The AI-native SDLC loop
Anthropic published a 46-minute playbook for building software with agents. I turned it into a loop you can scroll: one request, six stages, about two minutes.
Allyn Shaw, October 2026
Or run it yourself. An agent does the work, you hold the gates.
SDLC is how software gets from idea to production.
Plan, design, build, test, deploy, maintain. The playbook keeps the six stages and rebuilds the cycle around AI agents.
Why they wrote it.
Anthropic's applied AI team works with companies using Claude Code. They kept seeing the same problem.
Agents write code in hours.
The build stage went from weeks to an afternoon.
Everything around it still runs at human speed.
Planning, review and deployment didn't get faster. The bottleneck didn't disappear. It moved.
The review queue piles up.
Agents open pull requests faster than one reviewer can read them. Either the queue grows, or code goes out under-reviewed.
So they rebuilt the process. A loop, not a line.
The chain of handoffs and sign-offs becomes a cycle that feeds itself.
Each stage saves one file to git.
That file is what starts the next stage. An intent, a spec, a plan, the code, a reviewed pull request, and a new intent when something breaks.
Follow one request.
A note from fraud operations: customers want to confirm a blocked card payment in the app instead of phoning.
Step one: plan.
The old way: backlog entry, user story, story points, refinement meeting, sign-off. Three weeks before anyone writes a line.
Whoever has the idea explains it to Claude.
It asks who's affected and what the limits are, then writes the brief.
The brief is intent.md.
Problem, outcome, constraints, open questions. One page, in plain language.
A human approves it.
The product owner signs off and the file is committed. Three weeks become four hours.
Step two: design.
Claude turns the brief into a full spec. Requirements, security, design. That's spec.md.
Company rules apply while it writes.
Security, brand and UX policies are packaged as skills. The spec follows them as it's written, not in a review weeks later.
Problems get flagged now.
The draft wants to show the fraud model's score so customers see why a payment was blocked. That also tells fraudsters how the model decides, so it's routed to the policy owner before an engineer ever sees the spec.
Step three: build.
Claude Code writes plan.md first. In plan mode it can read the whole repo and can't change a file.
An engineer approves the plan.
Files, order, risks, and how the work will be proven. Only after that approval can it edit.
Every session starts by reading CLAUDE.md.
The team's commands, conventions and past mistakes, kept under one page. Money is BigDecimal, never double. The v1 directory is frozen.
One engineer runs three sessions at once.
Each in its own worktree. The engineer steers. Two or three is the practical ceiling.
What stops it going rogue? Two layers.
Skills are rules the agent follows. Most of the time.
Hooks are scripts that block unsafe actions.
Every time. A pre-tool-use hook refuses the edit to a frozen file or the command that prints a secret, and tells Claude why.
Skills make mistakes rare. Hooks make them impossible.
Policy as instructions, then policy as code.
Step four: test.
Claude checks its own work. Run the tests, fix, run again, until it's green.
The evidence goes on the pull request.
Forty-two passed, output attached. A suite of 20 to 50 real tasks runs as evals in CI, and the pass rate gates the merge.
No cheating the tests.
A hook locks the test files while a fix is in progress. A failing test can't be made to pass by rewriting it.
Step five: deploy.
Claude reviews every pull request in three passes: bugs, security, and compliance, including whether the code matches spec.md and plan.md.
It can never approve its own code.
The author can't approve. A code owner keeps the merge button.
Production stays locked.
The agent deploys freely to dev. Staging sits in between. Prod opens only when a release manager signs off.
Step six: maintain. 3 a.m., errors spike.
Detection is deterministic: a control band on a 30-day baseline. No model is involved in deciding something's wrong.
A script calls Claude.
One sigma, log it. Two, diagnose with read-only tools. Three, propose a fix. The service owner triages in the morning: fix now, schedule, or dismiss.
Claude finds the cause and writes a new intent.md.
The fraud engine was re-scoring every confirmation, uncached, and timing out. The finding lands in the triage queue as the next request, and the incident becomes a permanent eval.
The loop restarts itself.
Weekly security scans and an on-call agent in Slack feed the same queue.
Every step left a commit.
Who asked. What the agent built. Who approved. The audit trail is the git history, nothing extra.
Who holds each gate.
A product owner on the intent. A policy owner on anything flagged. An engineer on the plan. Evals in CI on the merge, a code owner on the button, a release manager on production. A service owner on what gets fixed. Agents do the work in between.
Code isn't the bottleneck anymore. The process is.
The loop keeps running. Human judgment stays above it.
Which stage is slowest where you work?
Now run the loop yourself. Six gates, five minutes, your name on the audit trail.