The portal teaches the six stages. This page lists every
play in the Playbook on its own: what it changes, what to adopt first, how it is governed, how
you would know it worked, and how far this repository has adopted it. Every entry is a
paraphrase. The source has the full steps, examples and context.
Each status is judged only against files committed here:
implemented means the evidence files exist;
partial names the missing piece;
declared out of scope gives the reason. Three plays have no
prerequisite, governance or indicator text in the source, and their fields say
"Not stated in the source" rather than inventing one.
Stage 1 · Plan
The person with the idea writes it down once, as a committed file the next stage can pick
up and act on.
The originator talks the idea through with Claude and commits the
outcome as intent.md, a short proto-spec of problem, outcome and limits. No
intermediary rewrites it into tickets first.
Prerequisites
None
Governance
Git history on the committed file shows author and time. A
product owner accepts it by merging or rejects it by closing the review, and acceptance is
what opens Design.
Leading indicator
Elapsed time between first discussing an idea and committing it,
which should drop from weeks to hours.
Lagging indicator
How many intents get accepted, and how often one is still being
edited once its spec exists.
Status here
implemented
Evidence
Each change starts as a
folder under intent/, and .sdlc/active names the one in
force.
One prompted session replaces two phases: Claude turns the
accepted intent into spec.md under the organisation's policy skills and lists
the concerns it could not resolve.
Policy is enforced during drafting instead of surfacing in a
late review; spec, prompt and skill versions are all versioned. The product owner signs
off and routes each flagged concern to its policy owner.
Leading indicator
The gap between the two commits, intent then spec, set against the
old two-phase cycle.
Lagging indicator
Spec revisions that land after planning for the change has
begun.
Status here
partial
Evidence
Every chain has a signed-off
intent/*/spec.md.
Gap
No brand, security or UX policy skills are committed here, so specs are
not skill-constrained.
Claude Code plan mode as the default starting point
What changes
Sessions open in plan mode with the approved spec. Claude reads the
code but edits nothing, the engineer challenges the plan until a stranger could follow it,
and it is committed as plan.md before any code.
Design is reviewed while changing course only means editing a
document, since plan mode cannot touch files before acceptance. Who accepted each revision
is recorded; riskier changes go to a tech lead or architect.
Leading indicator
How many changes merge from the first implementation attempt, and
how long approval-to-merge takes.
Lagging indicator
Rework rounds per change, and whether the merged diff still agrees
with its plan.
Status here
implemented
Evidence
Plans
are accepted in intent/*/plan.md before code changes, and
.github/workflows/sdlc-gate.yml checks coverage at merge. Whether a session
really ran in plan mode is a runtime fact the repository cannot show.
Claude Code on auto mode
What changes
Once the plan is approved, Claude applies edits without asking each
time. As the other guardrails mature this becomes normal for routine, well-tested work, and
review shifts from watching edits to reading what a long session produced.
Prerequisites
Not stated in the source
Governance
Not stated in the source
Leading indicator
Not stated in the source
Lagging indicator
Not stated in the source
Status here
declared out of scope
Evidence
None; see the reason.
Reason
An agent-runtime permission setting, not a repository artifact.
The CLAUDE.md
What changes
One file tells the agent what a newcomer would need on day one:
conventions, commands, architecture and recurring mistakes. The team edits it each time the
agent gets something wrong.
Prerequisites
None
Governance
Being versioned, the agent's standing instructions can be
reviewed and audited; edits appear in history and code owners approve them through
review.
Leading indicator
How often the agent repeats a mistake the file was meant to
prevent.
Lagging indicator
How long a new team member takes to merge a first change.
Status here
declared out of scope
Evidence
None; see the reason.
Reason
This repository commits no agent-context file. Adopting one is its
own intent, not a status fix.
Skills as institutional knowledge
What changes
Policy that must be applied the same way every time becomes a
versioned skill, updated in one place when the policy moves. General context belongs in
CLAUDE.md or a prompt instead.
A skill only advises: it makes compliance likely, not certain,
so a rule that must always hold needs a hook or review pass behind it. Policy owners review
skill edits as they would code.
Leading indicator
Delay between a policy owner approving a change and the updated
skill merging.
Lagging indicator
Review comments citing the policy, which should approach zero; if
they do not, the skill is not firing or has drifted.
Status here
partial
Evidence
README.md links the
separate repository that holds the lifecycle skill.
Gap
No skill is committed in this repository.
Hooks as build-time guardrails
What changes
Hooks are the deterministic layer under advisory skills. While
code is written they can refuse edits to protected paths, format and lint after each edit,
and keep secrets out of the diff, so they must be quick and scoped to the changed file.
Prerequisites
Not stated in the source
Governance
Not stated in the source
Leading indicator
Not stated in the source
Lagging indicator
Not stated in the source
Status here
implemented
Evidence
.kiro/hooks/privacy-scan.json
runs scripts/privacy_pretooluse_hook.py before the agent writes a file.
Parallel sessions and subagents
What changes
An engineer drives several Claude Code sessions at once, each in
its own git worktree, and turns recurring jobs into subagents, each scoped to a separate
context and a restricted toolset. The engineer's work becomes steering and review.
With more output, control has to come from repository
configuration: hooks and permissions bind every session, and each session's actions are
attributed to whoever started it.
Leading indicator
Sessions an engineer runs at once without review quality slipping,
and time spent steering instead of waiting.
Lagging indicator
Weekly merged changes per engineer, read together with rework.
Status here
declared out of scope
Evidence
None; see the reason.
Reason
Session orchestration is a runtime practice; no subagent definition
is committed.
Each session gets something to check itself against, such as
tests, a build or a screenshot comparison, and iterates until it passes. A verifier
subagent is one way to package the last check in a fresh context.
Prerequisites
None
Governance
Hooks can require verification before work is called done. The
proof is the toolchain's own output, kept in the transcript and the PR checks, which lets
the code owner concentrate on intent and risk.
Leading indicator
How often agent-written changes pass CI on the first run.
Lagging indicator
Review time per PR and the rate of changes that fail.
Status here
implemented
Evidence
verify.py
runs every portal check in one command, and
.github/workflows/portal-verify.yml runs it in CI.
Continuous evals in CI
What changes
An eval suite runs whenever the agent's setup changes, such as a new
model or a reworded prompt, and shows whether quality held. Cases that stop telling models
apart are retired; monitoring supplies new ones.
QA gets a gate that keeps pace: a pass-rate threshold blocks the
merge, results are kept for comparison, and the team owning the configuration approves its
changes.
Leading indicator
Pass rate across runs, and how fast an incident becomes a lasting
eval case.
Lagging indicator
Regressions stopped in CI compared with those that escape to
production.
Status here
partial
Evidence
scripts/privacy_mutation_proof.py
proves each scanner rule can fail, run by .github/workflows/portal-verify.yml.
Gap
These test the portal and scanner, not the agent's configuration; there
is no eval suite over agent tasks.
Claude reviews incoming PRs against written policy and answers
comments on its own. Every PR gets the same ranked findings, and people judge whether the
change fits the plan and whether its risk is acceptable.
Duties stay separate because the authoring agent cannot approve
its own change. REVIEW.md holds the policy, the PR records findings and
approvals, and a human approves through branch protection.
Leading indicator
Minutes to a first review, and comments settled without a person
editing the branch.
Lagging indicator
Defects and vulnerabilities stopped before merge compared with
those found later.
Status here
partial
Evidence
.github/pull_request_template.md
structures every PR, and .github/workflows/sdlc-gate.yml checks it against the
accepted chain.
Gap
No REVIEW.md and no AI review pass.
Hooks as approval gates
What changes
Beyond allowing or blocking, a hook can hold an action until a named
person approves it. Release is the obvious case, but the same gate works anywhere Claude
acts, such as migrations or test files.
Prerequisites
None
Governance
The hook is the gate: its rule applies on every run and to
every person, each verdict is logged with a time, and the gate itself defines what counts as
approval.
Leading indicator
Waiting time at each gate, read from the logged hook verdicts.
Lagging indicator
Gate breaches reaching production, before versus after the hooks.
Status here
partial
Evidence
.githooks/pre-push
refuses a push that the privacy scan flags.
Gap
No hook asks a named person to approve; human approval comes from branch
protection, which is repository configuration, not a committed file.
CI/CD integration and deployment
What changes
Claude runs headless in the pipeline for steps needing judgement,
sandboxed with scoped credentials, reaches deployment tools through MCP, and has rollback
paths rehearsed before it ever needs them.
The agent may work up to the production gate, never through it:
its output arrives as PRs, a named release manager authorises production, it runs under its
own identity, and each environment sets its tier.
Leading indicator
Pipeline failures triaged without paging anyone.
Lagging indicator
The four DORA delivery measures.
Status here
partial
Evidence
.github/workflows/
runs the portal, privacy and lifecycle gates on every PR.
Gap
CI is deterministic only; no agent step, MCP deployment or rehearsed
rollback.
A trigger starts the agent with nobody in the loop, and what it finds returns as a new
intent.md. The stage's opening part, on maintenance in general, is context rather
than a play.
A deterministic script watches production and calls Claude only
once a control band is breached. Claude diagnoses, acts within its tier, and records its
findings as a new intent.md that re-enters Plan.
Tier limits come from versioned configuration that denies
production access. Every invocation and finding is timestamped, a service owner triages,
fixes take the normal review path, and runbooks are pre-approved.
Leading indicator
Time from a breach to a new intent waiting for triage.
Lagging indicator
Findings that end as merged fixes, and whether the same class of
incident returns.
Status here
declared out of scope
Evidence
None; see the reason.
Reason
A static page has no production metric to band; no
bands.yaml is committed.
Recurring codebase scans
What changes
Security scans run on a timetable that no one has to start, since
both the code and the models that find flaws keep moving. Findings take the usual gates:
small fixes as PRs, larger work as a new intent.
Repositories, scan seats and spend are controlled centrally.
Each finding carries a validation and confidence rating and each dismissal a reason, and
fixes still reach production only through review.
Leading indicator
The share of repositories scanned on a timetable, and how long a
finding waits before its patch enters review.
Lagging indicator
Flaws found by the scans compared with those found in production or
reported from outside.
Status here
declared out of scope
Evidence
None; see the reason.
Reason
The privacy scan runs on changes, not on a schedule, and is not a
model-driven scan.
Claude on call with Claude Tag
What changes
Claude sits in incident and work channels under its own identity,
so each incident gets a first responder. It confirms recovery, files the post-mortem as a
versioned lessons file, and hands bigger work back as an intent.
Prerequisites
Not stated in the source
Governance
The channel history doubles as the audit record, keeping the
request, the diagnosis, the human sign-off and the fix in one place.
Leading indicator
Not stated in the source
Lagging indicator
Not stated in the source
Status here
declared out of scope
Evidence
None; see the reason.
Reason
No channel integration exists or is planned for this
repository.
Two passages in the source apply across stages. They are not plays.
Legacy systems and the source of truth
Most teams already keep these records in tools auditors trust, such as trackers,
requirements systems and change boards. For each artifact, pick one authoritative home: the
repository, the existing tool, or at minimum both, cross-linked by record ID and commit
SHA.
Managed settings
The source's worked example has a platform team push locked settings that deny secret
reads, confine network access with an OS sandbox, accept only managed hooks, MCP servers and
plugins, and set a minimum client version. It is a starting point to tailor; this page does
not reproduce the configuration.
What to adopt first
Edges come only from each play's prerequisites text. Required means adopt that play
first; helps means it makes the later play easier. The list is the canonical form; the
figure draws the same edges.
Capture as intent.md → Requirements and design (required)
Capture as intent.md → Closing the loop (required)
PR review → Closing the loop (required)
Approval hooks → Closing the loop (required; interpretive reading)
CI/CD → Closing the loop (required)
PR review → Recurring scans (required)
Approval hooks → Recurring scans (required)
Capture as intent.md → Recurring scans (required)
Auto mode, build-time hooks and Claude Tag have no prerequisites text in the source, so they
have no edges.
Adoption order for the thirteen plays with prerequisites. Solid lines are
required, dashed lines help. Where lines cross, the list above is authoritative.
Glossary
Intent
A short committed file, intent.md, stating a problem, the outcome wanted and
its limits, in the words of whoever raised it.
Spec
The requirements and design for one intent, spec.md, signed off by its
owner before planning.
Plan
The accepted implementation plan, plan.md: files that change, order of work
and the tests that prove it.
Play
One adoptable practice in the Playbook, belonging to a single stage.
Gate
A check that must pass before work moves on, ideally enforced by code rather than by
habit.
Control band
The normal range for a production metric; leaving it is what triggers the agent in
Maintain.
Hook
A script the agent runtime runs on a matching action, able to allow, block or ask for
approval.
Skill
A versioned instruction set the agent loads to apply a policy or procedure
consistently.
Subagent
A scoped helper inside one session that keeps a separate context and a restricted
toolset.
Worktree
An extra git working directory on its own branch, so parallel sessions do not collide.
Eval
A repeatable test of the agent's behaviour, used to catch regressions when its setup
changes.
MCP
Model Context Protocol: the standard way to expose tools and data sources to the
agent.
Managed settings
Agent configuration pushed centrally by a platform team that individual engineers cannot
override.
Merge base
The commit where a branch left main; plan acceptance in this repository is
bound to it.
Playbook terms and this repository
The Playbook names Claude Code files; this repository runs a different agent runtime. Each
row gives the local path, or says there is none.