The process around the code is the bottleneck now.
Agents write the diff in hours. The plan, the review and the release still run at human speed — so that is where the queue forms. This is a course on rebuilding the lifecycle as a loop of committed artifacts, and on making the stage order enforceable rather than aspirational.
This is a training site about the ai-native-sdlc skill.
It is not the skill's source repository — the skill, its gate scripts,
templates and reference docs live in
agent-skills-best-practice.
This portal is independent and not affiliated with or endorsed by Anthropic.
Why the lifecycle has to change
The traditional lifecycle was optimised for an era when writing code was the slowest, most expensive stage. Its controls also assume a human performs every step.
The bottleneck moves outward
Build collapses to hours; plan, review and deploy keep their length. The constraint is now the human-speed stages on either side of build.
Controls stop matching reality
Reading every line by hand made sense when a person wrote every line. It cannot keep pace once agents write most of the diff.
Governance cost rises
Exceptions still route through committees that meet weekly. A security team sized for human output either builds a queue or ships under-reviewed code.
The loop, and the artifact that carries it
intent.md, which re-enters Stage 1.Each stage ends by committing one machine-readable artifact, and the next stage begins by reading it. The chain of commits becomes the audit trail: who asked for what, what the agent produced, and who approved it.
intent.mdspec.mdplan.md → diff + testsverification target + evals/PR + REVIEW.md findingsbands.yaml breach → new intent.mdStage 6 writes a fresh
intent.md, which re-enters Stage 1. That is what makes it a loop rather than a
line — and why "closing the loop" is a deliverable rather than a metaphor.
Why markdown early, code later
In the early stages a product owner and an agent must both read and act on the same file, so markdown wins. From Build onward the artifact is the code and its records.
The six stages
Each stage shows the shift from traditional practice, the artifact it commits, the gate and who holds it, and how it is measured.
Choose how you work through this
Both routes read the same six canonical stages below — they add guidance, they do not replace the material. Without JavaScript both sets of guidance stay visible.
Progress stays in this browser. If storage is blocked, this visit uses session-only progress and nothing is saved.
Stage 1 — Plan
Ideas stop waiting for someone to write them up. Intent is captured once, in the originator's own words, as a version-controlled artifact the next stage can act on.
Backlog entries, user stories, story points and refinement meetings. Ownership transfers at each handoff, so what reaches engineering is several steps removed from what the originator meant.
The originator brainstorms with an agent and writes the
result as intent.md: what is wanted, why, under which constraints. Repeat
processes are encoded as skills.
- Contents. Problem, desired outcome, affected users and systems, constraints, success criteria, open questions.
- Gate. The product owner sets
Status: accepted. The accepting commit is the record — and it must not be the author's own. - Evidence. The committed file, with author, timestamp and full revision history.
- Measure. Leading: hours from first conversation to a committed intent. Lagging: the share of intents accepted into Design rather than closed.
Mentor guide · Plan ≈ 8 min
- Objective
- Explain why the originator commits intent once.
- Ask the room
- Where does meaning get lost between idea and engineering ticket?
- Demonstrate
- Show why self-accepting
intent.mdis no gate.
Recap
Plan commits intent.md; a different product owner accepts it.
Stage 2 — Design
Requirements and design collapse into one session. Policy is applied while the spec is written, not discovered in a review weeks later.
Analysts formalise the idea into requirements; designers parse those back into a design. The separation buys accountability and costs weeks, losing fidelity at each hop.
One session takes the accepted intent.md and
produces requirements plus design, constrained by the organisation's skills, with concerns
flagged.
- Flagged concerns first. These are the points an analyst would have escalated. Each names a policy owner and is resolved with them before engineering sees the spec.
- Gate. The owner sets
Status: signed-off, consulting a technical lead for anything classed as higher risk. - Evidence. The spec, the prompt that produced it, and the skill versions in force — all in version control.
- Measure. Leading: elapsed time between the intent and spec commits. Lagging: spec commits dated after the first plan commit, which is requirements rework.
Mentor guide · Design ≈ 8 min
- Objective
- Explain why policy owners resolve concerns before engineering.
- Ask the room
- Which recent concern belonged in design, not review?
- Demonstrate
- Walk a
spec.mdwhose concerns name owners and resolutions.
Recap
Design commits owner-signed spec.md, resolving concerns before build.
Stage 3 — Build
Nothing is implemented without an accepted plan. Institutional knowledge becomes files the agent reads, and guardrails run as code rather than as habits.
How the change will be made stays in the engineer's head. The first thing a reviewer sees is the finished diff, and by then rework is slow.
The agent reads the codebase and changes nothing while planning. The engineer corrects the plan before code exists, and the approved version is committed.
- The plan names the files that change, the order of work, the risks and the tests that prove it — detailed enough that an engineer who never saw the conversation could implement it.
- Record rejected approaches and why, so a later session does not re-walk them.
- Gate.
Status: acceptedbefore any file is edited. If implementation departs from the plan, update the plan in the same commit. - Binding. An acceptance is recorded for a base commit. Bind it to the merge base — the fork point the approval was granted at — not the branch tip, which moves after the branch is cut.
- Coverage. Every source file the change touches must be named in the plan; the gate refuses a diff that touches a file the active plan does not mention.
Mentor guide · Build ≈ 9 min
- Objective
- Explain pre-edit acceptance and merge-base binding.
- Ask the room
- What can branch-tip binding let a later push bypass?
- Demonstrate
- Run the gate on an unnamed file and read its refusal.
Recap
Build accepts plan.md before edits, binds it to the merge base and names every touched file.
Stage 4 — Test
Every session checks its own work before a human sees it, and the configuration that steers the agent gets regression-tested like the code it writes.
The signal that code works arrives late: CI in minutes, a tester in days, production in weeks. With an agent producing the code, a late signal makes a person the bottleneck.
The session is given a way to verify itself and iterates until the check passes, so what reaches the engineer has already passed it.
- Write the target first, watch it fail. A test that has only ever been seen green is not evidence.
- For a bug fix, commit the failing test first, then make it pass without editing the test. A test the agent could not rewrite is proof the bug is gone.
- Protect the loop. An agent fixing code must not be able to weaken the check on that code — block edits to test files during a fix, or reject such a diff in review.
- Evals are the stage gate that keeps up. Run them on any change to the agent's own configuration — instructions, skills, hooks — and on a schedule. Every production incident becomes a permanent eval case.
- Prove a test can fail. Mutate the implementation and confirm the suite goes red. A surviving mutation is an untested behaviour, not a pass.
display:none —
invisible and unreachable. A green gate means "the declared checks passed", never "the change
is correct".Mentor guide · Test ≈ 9 min
- Objective
- Explain red-first evidence and the weak-eval hole.
- Ask the room
- How do you prove a check can fail?
- Demonstrate
- Hide an eval-checked id with
display:none; watch it pass.
Recap
Test commits a red-first target and evals/; green proves only declared checks.
Stage 5 — Deploy
Review runs in both directions, and governance is enforced as the agent acts. The agent may do everything up to the production gate and nothing past it.
Review capacity was planned around human output. Quality varies with the reviewer's load, the author chases, the backlog grows.
Every pull request gets an identical set of severity-ranked passes. Human attention moves up a level: does the change do what the plan intended, and is the risk acceptable?
REVIEW.mddefines the passes — bugs, security, then compliance againstspec.mdandplan.md— plus what counts as Important versus a Nit, a nit cap, and what to skip.- Separation of duties. The agent that wrote the code has no route to approve it. A human code owner still merges.
- Close the intent on merge. Set
Status: shipped. This is not bookkeeping: a still-acceptedchain keeps authorising later unrelated edits to every file its plan named. Never reset a shipped chain to keep working under it — start a new intent. - Hooks as approval gates. A hook can allow, block, or ask. A block should explain itself and name the route to approval.
- Measure. Leading: time to first review, and comments resolved without a human touching the branch. Lagging: defects caught before merge versus those escaping to production.
Mentor guide · Deploy ≈ 8 min
- Objective
- Explain separation of duties and closing chains as shipped.
- Ask the room
- What stays authorised when a merged chain remains
accepted? - Demonstrate
- Show branch protection treating required
skippedas green.
Recap
Deploy commits PR + REVIEW.md; a human merges and marks the intent shipped.
Stage 6 — Maintain
The loop closes. A trigger invokes the agent with no
person in the invocation path, and what it finds re-enters the pipeline as
intent.md.
Maintenance is reactive. An alert fires at 3am and can be missed; a ticket sits in the backlog; post-mortem actions may never reach the codebase.
A control-band breach, ticket, channel message or
schedule invokes the agent. It diagnoses, acts only through gated routes, and writes what
it finds as a new intent.md.
- Detection stays deterministic. A version-controlled, unit-tested script watches the metric. The model is never in the detection path.
- Tier by deviation. 1σ log · 2σ diagnose read-only · 3σ act, but only by opening a pull request into the review gate or running a pre-approved runbook.
- Rollback is the most rehearsed path in the pipeline, exercised regularly in staging, because the 3σ tier may call it.
- Dismissals tune the bands and are recorded, so noise falls over time.
- Measure. Leading: time from band breach to an intent in the triage queue. Lagging: the share of findings that become merged fixes, and repeat incidents of the same class.
# bands.yaml — monitoring the CI test failure rate
metric: ci_test_failure_rate
baseline: rolling_30d
rules: western_electric
tiers:
1sigma: { action: log }
2sigma: { action: diagnose, tools: "Read,Grep,Bash(gh run view *)" }
3sigma: { action: propose, routes: [pull_request, runbook:rollback-deploy] }
Mentor guide · Maintain ≈ 8 min
- Objective
- Explain deterministic detection and how breaches re-enter the loop.
- Ask the room
- Why must detection be deterministic rather than model judgement?
- Demonstrate
- Trace a
bands.yamlbreach to a newintent.md.
Recap
Maintain turns a bands.yaml breach into a gated new intent.md.
Advisory, deterministic, or actually binding
This is the distinction most adoptions get wrong. Writing the process down does not enforce it.
| Strength | Mechanism | Bypassable |
|---|---|---|
| Advisory | A skill. Makes the agent likely to follow the process. | Yes — by not consulting it |
| Deterministic | A PreToolUse hook at write time. It
fails open. | Yes — by removing it |
| Merge gate | A CI gate on the pull request. It fails closed. | Only by an admin unsetting the required check |
Three failure modes worth memorising
A skipped check reads as passing
GitHub treats a
skipped required check as green. A summary job must therefore run
if: always() and inspect its dependencies' results explicitly, rather than
being skipped along with them.
An unbound approval outlives its change
If the acceptance records no base commit, it keeps authorising later edits. Record the base, and compare it with the pull request's real merge base.
Admins bypass repository protection
Only an organisation- or enterprise-level ruleset whose bypass list excludes repository admins is genuinely unbypassable. In a personal repository this gap cannot be closed at all.
Accepted-for must record.Hands-on lab: run one change end to end
The fastest way to learn the loop is to push one small change through it and let the gate refuse you a few times.
-
Install the two enforcement layers
Vendor the gate scripts, the write-time hook and the merge gate into the repository. The write-time hook ships as two configs for one script — Kiro's hook format and Claude Code's
.claude/settings.json— carrying an identical command, so install the one for the runtime the repo actually uses (both is harmless).SKILL=<the skill's directory> mkdir -p .kiro/hooks .claude .sdlc/scripts .github/workflows cp "$SKILL"/templates/kiro-hooks/sdlc-gate.json .kiro/hooks/ # Kiro cp "$SKILL"/templates/claude-code-hooks/settings.json .claude/settings.json # Claude Code cp "$SKILL"/scripts/sdlc_pretooluse_hook.py .sdlc/scripts/ cp "$SKILL"/scripts/sdlc_gate.py .sdlc/scripts/ cp "$SKILL"/scripts/sdlc_ci_gate.py .sdlc/scripts/ cp "$SKILL"/templates/github-workflows/sdlc-gate.yml .github/workflows/ echo "<slug>" > .sdlc/activeIf
.claude/settings.jsonalready exists, merge thehooks.PreToolUseentry into it rather than copying over it. -
Do the one step only a human can
Mark the gate a required status check in branch protection. Until you do, everything above is advisory.
-
Write the intent, then stop
Create
intent/<slug>/intent.mdfrom the template. Do not write a spec and do not touch code. Ask the owner to accept it — someone other than the author. -
Watch the gate refuse you
Run it against the draft and read the reason on stderr. Exit
2means closed. This is the lesson, not an error.python3 .sdlc/scripts/sdlc_gate.py intent/<slug> design # exit 0 = gate open, exit 2 = gate closed (reason on stderr) -
Plan before you edit anything
Produce
plan.mdnaming files, order, risks and proving tests. Bind the acceptance to the merge base, not the branch tip:git merge-base origin/main HEAD # the fork point the approval covers -
Red first, then green
Write the verification target, run it, and confirm it fails for the reason you expect — a traceback is not a red test. Commit that red target separately, then implement until it passes.
-
Open the pull request and let review run both ways
Findings inform; a human code owner still merges. Treat a required check as green only when its conclusion is literally
success— neverskipped. -
Close the chain
After merge, set
Status: shippedon all three artifacts and point.sdlc/activeat the next slug. Skipping this leaves a spent approval reading as live.
Exercises that make each trap observable
Try to edit source before the plan is accepted
The write-time hook should stop you. Now remove the hook and try again — it succeeds. That is the difference between "deterministic" and "binding", felt rather than read.
Accept a plan bound to the branch tip instead of the merge base
Push a second commit, then open the pull request. The gate refuses with a binding mismatch. This is why the guidance says merge base: the base moves after the branch is cut.
Leave a merged chain as accepted and start unrelated work
Edit any file that plan named. The gate passes — with no new approval. Now mark the chain
shipped and repeat: it refuses, and tells you to open a new intent.
Make a required check report skipped
Add a paths filter to the workflow behind a required check, then open a pull request that
touches other files. Watch branch protection accept the skip as green — or, with a summary
job that lacks if: always(), watch the pull request wait forever.
Write an eval that passes while the feature is broken
Assert that an element's id appears in the HTML, then hide the element with
display:none. Green gate, broken feature. No gate catches this — only a human
reading the requirement does.
Honest limits — read this before citing any of it
A governance tool that overstates its own governance teaches the wrong lesson. Scored against an enterprise control rubric, the reference implementation stands at 36/80 (45%), and roughly 48% is the ceiling a standalone repository can reach on its own.
| Adoption position | Status | What it requires |
|---|---|---|
| Individual or small team, production use | Usable now | Nothing beyond the repository. |
| Controlled enterprise pilot | Conditional | The enterprise supplies central policy, external evidence custody and identity governance for the pilot scope. |
| Enterprise-wide mandatory control | Not achieved | Fleet-wide policy, ownership, telemetry and drift visibility. |
| Regulated or auditable compliance control | Not achieved | All of the above, plus retention, longitudinal evidence and independent assurance. |
Three controls no repository can grant itself
An organisation policy plane
Policy authority must sit above the repository wall. Anyone with write access can delete the local control, so the tier that mandates it cannot live inside the thing being governed.
A tamper-evident audit sink
The trail is git history in the repository being governed, and the governed party can rewrite it. Self-audit is not audit.
Independent assurance
Every test in the reference implementation was written by its implementer. Self-authored tests are not independence, and no badge substitutes for a report.
Self-check
Six scenario checks on the things that actually bite. Each one now lives inside its matching stage panel above, in Self-paced mode — one check per stage rather than a single list here. Not started
- Plan. What a skipped required check looks like to branch protection.
- Design / Build. What binding acceptance to
HEADcosts you. - Test / Deploy. What a green gate certifies, and what a chain left
acceptedleaves open. - Maintain. Where the model belongs in Stage 6 detection.
Open Self-paced mode above, work each stage, answer its check, then use Mark Stage Complete when you are satisfied you understand it — completion is yours to declare, never inferred.
Sources and next steps
This portal is a study guide over two primary sources. Read them directly — they carry detail this page compresses.
The AI-Native SDLC playbook
Published by Anthropic on the Claude blog, 21 August 2026, by Louis Claxton. The conceptual framework: the six stages, the traditional-versus-AI-native shifts, and each play's prerequisites, governance considerations and metrics.
The ai-native-sdlc skill
The enforceable implementation: artifact templates, the stage gate, the
PreToolUse hook, the CI gate, and reference docs on enforcement, limitations,
the threat model and enterprise adoption.
Where to go from here
Start with any stage whose prerequisites you already meet — the stages are modular, and nothing points into Stage 1. If you only do one thing, make one check required and push a change that violates it, so you learn what your process actually enforces rather than what it says.