AI-Native SDLC
A training portal

The process around the code is the bottleneck now.

Agents write the diff in hours. The plan, the review and the release still run at human speed — so that is where the queue forms. This is a course on rebuilding the lifecycle as a loop of committed artifacts, and on making the stage order enforceable rather than aspirational.

This is a training site about the ai-native-sdlc skill. It is not the skill's source repository — the skill, its gate scripts, templates and reference docs live in agent-skills-best-practice. This portal is independent and not affiliated with or endorsed by Anthropic.

6 stages one artifact per stage hook + CI gate Usable now: individual & small team Not a compliance control

Why the lifecycle has to change

The traditional lifecycle was optimised for an era when writing code was the slowest, most expensive stage. Its controls also assume a human performs every step.

The bottleneck moves outward

Build collapses to hours; plan, review and deploy keep their length. The constraint is now the human-speed stages on either side of build.

Controls stop matching reality

Reading every line by hand made sense when a person wrote every line. It cannot keep pace once agents write most of the diff.

Governance cost rises

Exceptions still route through committees that meet weekly. A security team sized for human output either builds a queue or ships under-reviewed code.

The principle that survives. Humans stay accountable for every decision that needs judgement. What changes is where the attention sits: at the gates, reviewing what the agent flagged, instead of at the start of every stage.

The loop, and the artifact that carries it

Each stage commits one artifact, and Stage 6 re-enters Stage 1 The six stages run in order. Each ends by committing one machine-readable artifact and the next begins by reading it: Stage 1 Plan commits intent.md, Stage 2 Design commits spec.md, Stage 3 Build commits plan.md then the diff and tests, Stage 4 Test commits the verification target and evals, Stage 5 Deploy commits the pull request with REVIEW.md findings, and Stage 6 Maintain commits bands.yaml. A bands.yaml breach writes a fresh intent.md, which re-enters Stage 1. That is what makes it a loop rather than a line. Each green dot marks a commit, and the chain of commits becomes the audit trail: who asked for what, what the agent produced, and who approved it. Six stages, one loop each stage ends by committing one artifact re-enters Stage 1 Stage 1 · Plan intent.md Stage 2 · Design spec.md Stage 3 · Build plan.md → diff + tests Stage 4 · Test verification target + evals/ Stage 5 · Deploy PR + REVIEW.md findings Stage 6 · Maintain bands.yaml breach → new intent.md each dot is a commit the chain of commits becomes the audit trail
Stage 6 is not the end: a breached band writes the next intent.md, which re-enters Stage 1.

Each stage ends by committing one machine-readable artifact, and the next stage begins by reading it. The chain of commits becomes the audit trail: who asked for what, what the agent produced, and who approved it.

Stage 1 · Planintent.md
Stage 2 · Designspec.md
Stage 3 · Buildplan.md → diff + tests
Stage 4 · Testverification target + evals/
Stage 5 · DeployPR + REVIEW.md findings
Stage 6 · Maintainbands.yaml breach → new intent.md

Stage 6 writes a fresh intent.md, which re-enters Stage 1. That is what makes it a loop rather than a line — and why "closing the loop" is a deliverable rather than a metaphor.

Why markdown early, code later

In the early stages a product owner and an agent must both read and act on the same file, so markdown wins. From Build onward the artifact is the code and its records.

The six stages

Each stage shows the shift from traditional practice, the artifact it commits, the gate and who holds it, and how it is measured.

Choose how you work through this

Both routes read the same six canonical stages below — they add guidance, they do not replace the material. Without JavaScript both sets of guidance stay visible.

Stage 1 of 6
0 of 6 stages complete

Progress stays in this browser. If storage is blocked, this visit uses session-only progress and nothing is saved.

Stage 1 — Plan

Ideas stop waiting for someone to write them up. Intent is captured once, in the originator's own words, as a version-controlled artifact the next stage can act on.

Traditional

Backlog entries, user stories, story points and refinement meetings. Ownership transfers at each handoff, so what reaches engineering is several steps removed from what the originator meant.

AI-native

The originator brainstorms with an agent and writes the result as intent.md: what is wanted, why, under which constraints. Repeat processes are encoded as skills.

commits → intent/<slug>/intent.md
  • Contents. Problem, desired outcome, affected users and systems, constraints, success criteria, open questions.
  • Gate. The product owner sets Status: accepted. The accepting commit is the record — and it must not be the author's own.
  • Evidence. The committed file, with author, timestamp and full revision history.
  • Measure. Leading: hours from first conversation to a committed intent. Lagging: the share of intents accepted into Design rather than closed.
Trap. Open questions must be closed or explicitly deferred before Design. An unanswered question does not disappear; it resurfaces as rework after build starts.

Mentor guide · Plan ≈ 8 min

Objective
Explain why the originator commits intent once.
Ask the room
Where does meaning get lost between idea and engineering ticket?
Demonstrate
Show why self-accepting intent.md is no gate.

Recap

Plan commits intent.md; a different product owner accepts it.

Stage plays

Stage 2 — Design

Requirements and design collapse into one session. Policy is applied while the spec is written, not discovered in a review weeks later.

Traditional

Analysts formalise the idea into requirements; designers parse those back into a design. The separation buys accountability and costs weeks, losing fidelity at each hop.

AI-native

One session takes the accepted intent.md and produces requirements plus design, constrained by the organisation's skills, with concerns flagged.

commits → intent/<slug>/spec.md
  • Flagged concerns first. These are the points an analyst would have escalated. Each names a policy owner and is resolved with them before engineering sees the spec.
  • Gate. The owner sets Status: signed-off, consulting a technical lead for anything classed as higher risk.
  • Evidence. The spec, the prompt that produced it, and the skill versions in force — all in version control.
  • Measure. Leading: elapsed time between the intent and spec commits. Lagging: spec commits dated after the first plan commit, which is requirements rework.

Mentor guide · Design ≈ 8 min

Objective
Explain why policy owners resolve concerns before engineering.
Ask the room
Which recent concern belonged in design, not review?
Demonstrate
Walk a spec.md whose concerns name owners and resolutions.

Recap

Design commits owner-signed spec.md, resolving concerns before build.

Stage plays

Stage 3 — Build

Nothing is implemented without an accepted plan. Institutional knowledge becomes files the agent reads, and guardrails run as code rather than as habits.

Traditional

How the change will be made stays in the engineer's head. The first thing a reviewer sees is the finished diff, and by then rework is slow.

AI-native

The agent reads the codebase and changes nothing while planning. The engineer corrects the plan before code exists, and the approved version is committed.

commits → intent/<slug>/plan.md, then the diff
  • The plan names the files that change, the order of work, the risks and the tests that prove it — detailed enough that an engineer who never saw the conversation could implement it.
  • Record rejected approaches and why, so a later session does not re-walk them.
  • Gate. Status: accepted before any file is edited. If implementation departs from the plan, update the plan in the same commit.
  • Binding. An acceptance is recorded for a base commit. Bind it to the merge base — the fork point the approval was granted at — not the branch tip, which moves after the branch is cut.
  • Coverage. Every source file the change touches must be named in the plan; the gate refuses a diff that touches a file the active plan does not mention.
Why coverage matters. Without it a single historical acceptance launders every later change to any file that plan happened to name.

Mentor guide · Build ≈ 9 min

Objective
Explain pre-edit acceptance and merge-base binding.
Ask the room
What can branch-tip binding let a later push bypass?
Demonstrate
Run the gate on an unnamed file and read its refusal.

Recap

Build accepts plan.md before edits, binds it to the merge base and names every touched file.

Stage plays

Stage 4 — Test

Every session checks its own work before a human sees it, and the configuration that steers the agent gets regression-tested like the code it writes.

Traditional

The signal that code works arrives late: CI in minutes, a tester in days, production in weeks. With an agent producing the code, a late signal makes a person the bottleneck.

AI-native

The session is given a way to verify itself and iterates until the check passes, so what reaches the engineer has already passed it.

a verification target (one command) + evals/
  • Write the target first, watch it fail. A test that has only ever been seen green is not evidence.
  • For a bug fix, commit the failing test first, then make it pass without editing the test. A test the agent could not rewrite is proof the bug is gone.
  • Protect the loop. An agent fixing code must not be able to weaken the check on that code — block edits to test files during a fix, or reject such a diff in review.
  • Evals are the stage gate that keeps up. Run them on any change to the agent's own configuration — instructions, skills, hooks — and on a schedule. Every production incident becomes a permanent eval case.
  • Prove a test can fail. Mutate the implementation and confirm the suite goes red. A surviving mutation is an untested behaviour, not a pass.
The weak-eval hole. A gate can verify that an eval exists and passes. It cannot verify the eval is any good. An eval asserting that an element's id appears in the HTML passes while that element sits behind display:none — invisible and unreachable. A green gate means "the declared checks passed", never "the change is correct".

Mentor guide · Test ≈ 9 min

Objective
Explain red-first evidence and the weak-eval hole.
Ask the room
How do you prove a check can fail?
Demonstrate
Hide an eval-checked id with display:none; watch it pass.

Recap

Test commits a red-first target and evals/; green proves only declared checks.

Stage plays

Stage 5 — Deploy

Review runs in both directions, and governance is enforced as the agent acts. The agent may do everything up to the production gate and nothing past it.

Traditional

Review capacity was planned around human output. Quality varies with the reviewer's load, the author chases, the backlog grows.

AI-native

Every pull request gets an identical set of severity-ranked passes. Human attention moves up a level: does the change do what the plan intended, and is the risk acceptable?

PR + REVIEW.md findings → gated release
  • REVIEW.md defines the passes — bugs, security, then compliance against spec.md and plan.md — plus what counts as Important versus a Nit, a nit cap, and what to skip.
  • Separation of duties. The agent that wrote the code has no route to approve it. A human code owner still merges.
  • Close the intent on merge. Set Status: shipped. This is not bookkeeping: a still-accepted chain keeps authorising later unrelated edits to every file its plan named. Never reset a shipped chain to keep working under it — start a new intent.
  • Hooks as approval gates. A hook can allow, block, or ask. A block should explain itself and name the route to approval.
  • Measure. Leading: time to first review, and comments resolved without a human touching the branch. Lagging: defects caught before merge versus those escaping to production.

Mentor guide · Deploy ≈ 8 min

Objective
Explain separation of duties and closing chains as shipped.
Ask the room
What stays authorised when a merged chain remains accepted?
Demonstrate
Show branch protection treating required skipped as green.

Recap

Deploy commits PR + REVIEW.md; a human merges and marks the intent shipped.

Stage plays

Stage 6 — Maintain

The loop closes. A trigger invokes the agent with no person in the invocation path, and what it finds re-enters the pipeline as intent.md.

Traditional

Maintenance is reactive. An alert fires at 3am and can be missed; a ticket sits in the backlog; post-mortem actions may never reach the codebase.

AI-native

A control-band breach, ticket, channel message or schedule invokes the agent. It diagnoses, acts only through gated routes, and writes what it finds as a new intent.md.

bands.yaml breach → a new intent.md
  • Detection stays deterministic. A version-controlled, unit-tested script watches the metric. The model is never in the detection path.
  • Tier by deviation. 1σ log · 2σ diagnose read-only · 3σ act, but only by opening a pull request into the review gate or running a pre-approved runbook.
  • Rollback is the most rehearsed path in the pipeline, exercised regularly in staging, because the 3σ tier may call it.
  • Dismissals tune the bands and are recorded, so noise falls over time.
  • Measure. Leading: time from band breach to an intent in the triage queue. Lagging: the share of findings that become merged fixes, and repeat incidents of the same class.
# bands.yaml — monitoring the CI test failure rate
metric: ci_test_failure_rate
baseline: rolling_30d
rules: western_electric
tiers:
  1sigma: { action: log }
  2sigma: { action: diagnose, tools: "Read,Grep,Bash(gh run view *)" }
  3sigma: { action: propose, routes: [pull_request, runbook:rollback-deploy] }

Mentor guide · Maintain ≈ 8 min

Objective
Explain deterministic detection and how breaches re-enter the loop.
Ask the room
Why must detection be deterministic rather than model judgement?
Demonstrate
Trace a bands.yaml breach to a new intent.md.

Recap

Maintain turns a bands.yaml breach into a gated new intent.md.

Stage plays

Advisory, deterministic, or actually binding

This is the distinction most adoptions get wrong. Writing the process down does not enforce it.

Enforcement strength by mechanism
StrengthMechanismBypassable
AdvisoryA skill. Makes the agent likely to follow the process.Yes — by not consulting it
DeterministicA PreToolUse hook at write time. It fails open.Yes — by removing it
Merge gateA CI gate on the pull request. It fails closed.Only by an admin unsetting the required check
A green-or-red check is not a gate. Until the check is marked required in branch protection, a red check can still be merged. And a required check whose workflow is path-filtered never reports on a pull request that touches other files — which blocks that pull request forever, waiting on a check that can never arrive.
Enforcement strength by mechanism, and how each one is bypassed Three strengths of enforcement, each bypassable in a different way. Advisory is a skill: it makes the agent likely to follow the process, and is bypassable by not consulting it. Deterministic is a PreToolUse hook at write time; it fails open and is bypassable by removing it. A merge gate is a CI gate on the pull request; it fails closed and is bypassable only by an admin unsetting the required check. A green-or-red check is not a gate: until the check is marked required in branch protection, a red check can still be merged. Enforcement strength by mechanism writing the process down does not enforce it Advisory A skill. Makes the agent likely to follow the process. Yes — by not consulting it Deterministic A PreToolUse hook at write time. It fails open. Yes — by removing it Merge gate A CI gate on the pull request. It fails closed. Only by an admin unsetting the required check A green-or-red check is not a gate. Until the check is marked required in branch protection, a red check can still be merged.
Each tier removes one escape route. None removes them all — the strongest still yields to a repository admin.

Three failure modes worth memorising

A skipped check reads as passing

GitHub treats a skipped required check as green. A summary job must therefore run if: always() and inspect its dependencies' results explicitly, rather than being skipped along with them.

An unbound approval outlives its change

If the acceptance records no base commit, it keeps authorising later edits. Record the base, and compare it with the pull request's real merge base.

Admins bypass repository protection

Only an organisation- or enterprise-level ruleset whose bypass list excludes repository admins is genuinely unbypassable. In a personal repository this gap cannot be closed at all.

An unbound approval outlives its change A branch is cut from commit A. The base branch then advances to B and C while the branch adds X and Y. If the acceptance records no base commit, it keeps authorising later edits, so the base must be recorded and compared with the pull request's real merge base, which is A. Recording the base tip C instead claims review of commits B and C that nobody looked at, and recording the branch tip Y makes the approval self-certifying because it moves with every push. An unbound approval outlives its change If the acceptance records no base commit, it keeps authorising later edits. base A B C base tip yours X Y branch tip A — the real merge base Record the base, and compare it with the pull request's real merge base. C — the base tip Claims review of B and C, which nobody looked at. Y — the branch tip Moves with every push, so the approval certifies itself.
The approval was granted against the fork point, so that is the commit Accepted-for must record.

Hands-on lab: run one change end to end

The fastest way to learn the loop is to push one small change through it and let the gate refuse you a few times.

  1. Install the two enforcement layers

    Vendor the gate scripts, the write-time hook and the merge gate into the repository. The write-time hook ships as two configs for one script — Kiro's hook format and Claude Code's .claude/settings.json — carrying an identical command, so install the one for the runtime the repo actually uses (both is harmless).

    SKILL=<the skill's directory>
    mkdir -p .kiro/hooks .claude .sdlc/scripts .github/workflows
    cp "$SKILL"/templates/kiro-hooks/sdlc-gate.json        .kiro/hooks/          # Kiro
    cp "$SKILL"/templates/claude-code-hooks/settings.json  .claude/settings.json # Claude Code
    cp "$SKILL"/scripts/sdlc_pretooluse_hook.py            .sdlc/scripts/
    cp "$SKILL"/scripts/sdlc_gate.py                       .sdlc/scripts/
    cp "$SKILL"/scripts/sdlc_ci_gate.py                    .sdlc/scripts/
    cp "$SKILL"/templates/github-workflows/sdlc-gate.yml   .github/workflows/
    echo "<slug>" > .sdlc/active

    If .claude/settings.json already exists, merge the hooks.PreToolUse entry into it rather than copying over it.

  2. Do the one step only a human can

    Mark the gate a required status check in branch protection. Until you do, everything above is advisory.

  3. Write the intent, then stop

    Create intent/<slug>/intent.md from the template. Do not write a spec and do not touch code. Ask the owner to accept it — someone other than the author.

  4. Watch the gate refuse you

    Run it against the draft and read the reason on stderr. Exit 2 means closed. This is the lesson, not an error.

    python3 .sdlc/scripts/sdlc_gate.py intent/<slug> design
    # exit 0 = gate open, exit 2 = gate closed (reason on stderr)
  5. Plan before you edit anything

    Produce plan.md naming files, order, risks and proving tests. Bind the acceptance to the merge base, not the branch tip:

    git merge-base origin/main HEAD   # the fork point the approval covers
  6. Red first, then green

    Write the verification target, run it, and confirm it fails for the reason you expect — a traceback is not a red test. Commit that red target separately, then implement until it passes.

  7. Open the pull request and let review run both ways

    Findings inform; a human code owner still merges. Treat a required check as green only when its conclusion is literally success — never skipped.

  8. Close the chain

    After merge, set Status: shipped on all three artifacts and point .sdlc/active at the next slug. Skipping this leaves a spent approval reading as live.

Exercises that make each trap observable

Try to edit source before the plan is accepted

The write-time hook should stop you. Now remove the hook and try again — it succeeds. That is the difference between "deterministic" and "binding", felt rather than read.

Accept a plan bound to the branch tip instead of the merge base

Push a second commit, then open the pull request. The gate refuses with a binding mismatch. This is why the guidance says merge base: the base moves after the branch is cut.

Leave a merged chain as accepted and start unrelated work

Edit any file that plan named. The gate passes — with no new approval. Now mark the chain shipped and repeat: it refuses, and tells you to open a new intent.

Make a required check report skipped

Add a paths filter to the workflow behind a required check, then open a pull request that touches other files. Watch branch protection accept the skip as green — or, with a summary job that lacks if: always(), watch the pull request wait forever.

Write an eval that passes while the feature is broken

Assert that an element's id appears in the HTML, then hide the element with display:none. Green gate, broken feature. No gate catches this — only a human reading the requirement does.

Honest limits — read this before citing any of it

A governance tool that overstates its own governance teaches the wrong lesson. Scored against an enterprise control rubric, the reference implementation stands at 36/80 (45%), and roughly 48% is the ceiling a standalone repository can reach on its own.

Adoption positions, and what each requires
Adoption positionStatusWhat it requires
Individual or small team, production use Usable nowNothing beyond the repository.
Controlled enterprise pilot ConditionalThe enterprise supplies central policy, external evidence custody and identity governance for the pilot scope.
Enterprise-wide mandatory control Not achievedFleet-wide policy, ownership, telemetry and drift visibility.
Regulated or auditable compliance control Not achievedAll of the above, plus retention, longitudinal evidence and independent assurance.

Three controls no repository can grant itself

An organisation policy plane

Policy authority must sit above the repository wall. Anyone with write access can delete the local control, so the tier that mandates it cannot live inside the thing being governed.

A tamper-evident audit sink

The trail is git history in the repository being governed, and the governed party can rewrite it. Self-audit is not audit.

Independent assurance

Every test in the reference implementation was written by its implementer. Self-authored tests are not independence, and no badge substitutes for a report.

What documentation does not buy. Writing a contract is not deploying a control. A passing suite written by the implementer is not independent review. Synthetic CI fixtures are not representative production adoption. A roadmap is not a funded commitment.
What this deliberately does not cover. Out of scope: auto mode, CLAUDE.md, parallel sessions, closing the loop, recurring security scans, Claude on call. Cross-cutting: legacy-system onboarding, managed settings.

Self-check

Six scenario checks on the things that actually bite. Each one now lives inside its matching stage panel above, in Self-paced mode — one check per stage rather than a single list here. Not started

  • Plan. What a skipped required check looks like to branch protection.
  • Design / Build. What binding acceptance to HEAD costs you.
  • Test / Deploy. What a green gate certifies, and what a chain left accepted leaves open.
  • Maintain. Where the model belongs in Stage 6 detection.

Open Self-paced mode above, work each stage, answer its check, then use Mark Stage Complete when you are satisfied you understand it — completion is yours to declare, never inferred.

Sources and next steps

This portal is a study guide over two primary sources. Read them directly — they carry detail this page compresses.

The AI-Native SDLC playbook

Published by Anthropic on the Claude blog, 21 August 2026, by Louis Claxton. The conceptual framework: the six stages, the traditional-versus-AI-native shifts, and each play's prerequisites, governance considerations and metrics.

The ai-native-sdlc skill

The enforceable implementation: artifact templates, the stage gate, the PreToolUse hook, the CI gate, and reference docs on enforcement, limitations, the threat model and enterprise adoption.

Where to go from here

Start with any stage whose prerequisites you already meet — the stages are modular, and nothing points into Stage 1. If you only do one thing, make one check required and push a change that violates it, so you learn what your process actually enforces rather than what it says.

Reset your local progress?

This clears the completed stages and current position stored in this browser only. Nothing was ever sent anywhere, and this cannot be undone.