The Standard · Adopters · GitHub ↗
How to structure agent-instruction files across a set of repos: one source of truth, an in-repo fix log, written anti-drift rules, and a self-healing session hook.
Apply this to any repo an AI agent (Claude Code, Codex, Cursor, Gemini) works in.
A repo has one canonical instruction file:
AGENTS.md. It holds everything an agent (or human) needs to
work on the repo: architecture pointer, build/test commands,
conventions, gotchas, ship rules.
CLAUDE.md is one line, an include:
@AGENTS.md
Why: AGENTS.md is the cross-harness convention (Codex,
Cursor, Gemini, Agent Skills all read it). Claude Code reads
CLAUDE.md, so the include points it at the same canonical
file. Never maintain the same content in two files.
That is the failure mode this prevents (a large monolithic instruction
file drifted because there was no single source).
A CLAUDE.md symlink →
AGENTS.md is equally compliant and in some ways
stronger: any tool that opens CLAUDE.md gets the full
canonical content with zero drift, no @import support
required (a CLAUDE.md -> AGENTS.md symlink, git mode
120000). Leave such repos as-is.
Tradeoff: symlinks don’t survive some Windows checkouts / zip
exports, where the @AGENTS.md text include is more
portable. Pick per repo; both satisfy one-source-of-truth. Do
not convert a working symlink to an include just for
uniformity; it is a lateral move.
If a repo genuinely needs Claude-only nuance, put the
@AGENTS.md line first, then the small Claude-specific
addendum below it. This should be rare.
A repo may keep CLAUDE.md and AGENTS.md as
complementary files (not duplicates) when the content
genuinely divides by audience, e.g. AGENTS.md = install +
operating protocol + routing (cross-harness onboarding),
CLAUDE.md = architecture reference / key files / test
layout. This is compliant as long as the two never hold the same
content and a ## Keep in sync rule covers any
overlap. The anti-pattern is duplication, not two
files.
The root AGENTS.md doesn’t have to carry everything.
Context that belongs to one part of the repo can live there, so it loads
only when an agent actually works there:
AGENTS.md for a folder
with its own live conventions, constraints, or locked decisions (a
package in a monorepo, a deploy dir). Harnesses that read nested
instruction files load it on top of the root; open it with “Apply the
root AGENTS.md first, then this” so precedence is written down. The
one-source rule applies at every level — if the harness wants a folder
CLAUDE.md too, it’s a symlink or include, never a
copy.paths: glob), so a guardrail loads
only when matching files are touched.Two disciplines keep scoped files from becoming the drift problem
they solve: a folder file holds only what the root doesn’t (zero
duplication), and it pins decisions and constraints, not
structure — file trees and stack lists are derivable from
ls and rot fast (§3). A folder of static reference files
needs no instruction file at all.
For the wider doc set an AGENTS.md routes to (runbooks,
design docs, system notes), make discovery mechanical: open every doc
with a dense summary in its first few lines, so an agent can
head across the folder and pick the right file instead of
loading them all. Note the convention in AGENTS.md — both
so agents rely on it when searching and so they keep the summary true
when they edit the doc.
# <Repo>: Agent Guide
One-sentence description of what this repo is.
## Architecture
Where the code lives, the 3 to 5 things you must understand before editing.
## Commands
Build / dev / test commands. The ones you actually run.
## Conventions
House style, naming, patterns to follow, orphaned code NOT to recreate.
## Gotchas
The traps. (See docs/solutions/ for the full fix log.)
## Before shipping
The ship gate, in order: static → behavior → system. A failure stops the run.
Name the exact command(s); prefer one entry point (see "What Before shipping has
to say" below).
## Keep in sync
<see section 3>
## Commit flow
One or two lines: the repo's default (commit straight to `main`, or branch + PR),
deferring to §6 for the rule. Only the repo-specific part belongs here — the
account/deploy note, or "this repo branches because X."
## Diagrams
Where the architecture/data-flow diagrams live, if the repo has them (e.g.
`docs/diagrams/` — mermaid source + rendered images). Omit the section entirely
when there are none.
Keep AGENTS.md scannable. Anything that is a specific past
incident goes in docs/solutions/, not inline. That
keeps AGENTS.md from growing without bound.
The last two sections are conditional, not
mandatory. ## Commit flow earns its place only
when the repo diverges from §6’s default or carries an account/deploy
note; a repo that just follows the default needs no section.
## Diagrams exists only when diagrams do — an empty
“Diagrams” heading is the ceremony §3 warns against, not compliance.
They are in the skeleton because nearly every repo grows them, not
because every repo must.
Write hard rules with their exceptions attached: “never commit
secrets, except .env.example” survives contact with edge
cases; a bare NEVER gets ignored the first time an edge case makes it
wrong.
Some rules are not “follow this convention” but “this value must not
change” — a flag that has to stay set, a config key that must keep its
value, a file an agent must never touch. A bare NEVER buried in prose is
the wrong shape for these: it reads as one line among many, and an agent
editing the surrounding file steps on it without noticing. Give them a
named block near the top of AGENTS.md, before the sections
an agent skims for its task:
telemetry.enabled in config.json must stay
true in every change” is checkable; “keep telemetry on” is
a wish. Name the exact key, file, and required value.An invariant that can be checked mechanically belongs in the ship gate or a sync contract (§3) as well — the block tells the agent, the check catches the agent that didn’t read it. The block is not a substitute for the gate; it’s what makes a violation legible when the gate flags it.
The ## Before shipping line in the skeleton is a gate,
and a gate is only worth writing down if it answers two questions the
agent will otherwise answer for itself: what counts as done,
and in what order do I check.
Done is runtime evidence, not written code. The most common way an agent ships a regression is not a bad edit — it’s declaring victory early. Code compiles, the diff looks right, the model is confident, so the work is announced as complete without ever running. State the bar explicitly, because the default bar is the agent’s confidence: a change is done when the gate has been run and passed on this change, and its output is quoted. Not when the code is written, not when it “should work”, not when a previous run passed. An agent’s report that the tests pass is a claim about tests; it is not the tests. If the gate wasn’t run, the honest status is “written, unverified” — which is a fine thing to say and a much cheaper thing to hear than a false green.
Order the gate in layers, and don’t skip ahead. List the ship gate as an ordered sequence, not a flat pile of commands:
Each layer is only meaningful if the one below it passed, so
a failure at layer N stops the run — it doesn’t get noted and
skipped. The ordering exists because layer 1 is the layer
agents over-trust: a clean typecheck feels like proof and proves almost
nothing. Say in AGENTS.md which layers a given change
requires — a comment fix may need only layer 1, but anything crossing a
module boundary needs layer 3, because that’s the boundary unit tests
are least likely to cover.
Repos that can express the whole ladder as one command should:
make check, one script, one entry point. A single command
is a gate an agent can’t partially run.
The strongest form is a named, seconds-fast command whose
exit code is the definition of done —
make check, npm run verify,
./ship — stated in AGENTS.md as “exit 0 or the
work is not done.” Two properties do the work: it is one command (can’t
be partially run, above), and it is fast enough that it actually gets
run every time — a gate slow enough to skip is a gate that gets skipped,
so push the slow browser/e2e/build checks into a separate
deploy gate and keep the commit gate lean. Write down
what each gate covers and when it runs. This is the single least-adopted
part of the standard in practice — most repos leave the “Before
shipping” section as a placeholder — so make it a real command, not a
paragraph describing one.
An instruction file that documents the build perfectly is still a broken harness if the agent can’t reproduce the toolchain that build assumed. Two things belong in the repo, both cheap:
package-lock.json, uv.lock,
Cargo.lock, go.sum, Gemfile.lock,
…). Without one, the agent resolves dependencies to whatever is newest
today, and a failure it hits is not reproducible by the next session —
or by you..tool-versions, .nvmrc,
.python-version, or the equivalent field in the project
manifest). “Works on my machine” is a human-scale problem; for an agent
it’s a wrong-version error it will try to fix by editing your code.Neither is scored by adopt --check — they’re
language-specific, and the six checks are deliberately
language-agnostic. They are still the first thing to look at when an
agent’s failures don’t reproduce.
Instruction-file content comes in four layers — Syntax, Service, System, Strategy. Each builds on the one below; skipping a layer produces slop regardless of model quality:
A minimal AGENTS.md covers Syntax and Service. System often lives in a shared org-level file the repo points to. Strategy is frequently private config — but each layer should exist in writing somewhere an agent can read, because this hierarchy is exactly the implicit knowledge senior engineers carry, made explicit and machine-readable.
Content drift has a twin: naming drift — one concept
accruing synonyms (“issue”, “ticket”, “task item”) until an agent builds
around the wrong one. Where a repo’s domain has terms worth defending,
give AGENTS.md (or a file it points to) a short
domain-language block:
Keep it to terms that actually get confused. A glossary of the obvious is an inventory, and inventories rot (§3).
Naming drift has a structural twin. A repo maintained by a succession of leads accumulates strata: each lead built in the idiom they understood, and because nobody had enough context to migrate the previous one, the old idiom survives underneath. The result is a codebase where two files solve the same problem in two shapes, both load-bearing, neither wrong. This is most common in small teams on complex domains — permissions, billing, identity — where the work is permanently valuable but rarely the quarter’s priority, so nothing ever forces the layers to compact.
Strata break the instruction “write code consistent with the surrounding code”, because surrounding is ambiguous by construction. An agent that copies the nearest pattern picks a layer at random, and a harness that ships all day picks a new one every session — which is how the repo gains a stratum per contributor instead of per lead.
Where a repo has strata, AGENTS.md (or a file it points
to) says so plainly:
Left unwritten, this knowledge lives only with whoever has been there longest, and it decays exactly the way §11 describes — except here the cost isn’t a lost runbook, it’s a repo that gains a layer every time someone new starts shipping.
The block above assumes someone can enumerate the layers. Often nobody can — the person who could is exactly the person who left, which is the condition that produced the strata in the first place. So the layers have to be recovered from evidence, and the repo carries plenty: strata are visible as clusters of files that solve the same problem in different shapes, and those clusters correlate with periods of history.
The recovery is a §11 discovery pass narrowed to one question — what idioms are live here? — and it is read-only:
AGENTS.md
explicitly (“two live shapes, no ruling”) — an honest ambiguity an agent
can escalate on beats a confident wrong answer it will build on.The output is the block above, plus a fix-log entry (§2) recording
how the strata arose where that is knowable — the entry is the
why, and the AGENTS.md block is the compiled rule.
Re-run the pass when a migration finishes or a new shape starts
appearing, not on a schedule; strata change on the timescale of leads,
not sprints.
Compliance is checkable (adopt --check), and the score
is deliberately shaped as a maturity level, not a raw count of
passing checks. A repo with a rich AGENTS.md, a
fix log, and a sync block but no secret hygiene is not
“almost compliant” — it is one clear rung below a repo that also keeps
secrets out of history, because the missing piece is a floor,
not one point among many. Levels gate on the shape of
the harness:
AGENTS.md; agents relearn the repo every session.AGENTS.md exists..gitignore that keeps .env out of
history). The secret floor is what separates the top level.A point total lets 80% of beautiful docs mask a dangerous gap; a
level makes the gap the headline. The gate (--check exit 0)
is reaching the top level, not collecting the most points.
The check IDs are a contract. Each check has a
stable identifier (STD-01 … STD-06); an ID
never changes meaning, and the --json output only gains
fields, never renames them — so a CI pipeline gating on the score never
silently misreads a new version.
What the check does not measure. A
deterministic scan can confirm a file exists, parses, and matches a
pattern — never that its contents are true. A stale rule scores
like a fresh one; an AGENTS.md full of wrong commands
passes the presence check; a fix log of outdated entries still counts as
a fix log. A high score means the infrastructure for reliable
agent work is in place — it is necessary, not sufficient, and that is
the honest ceiling of any automated check. Keeping the contents
true is the §2 “compile, don’t retrieve” discipline, which no
scanner can do for you.
docs/solutions/: the
fix logA committed, queryable record of past bugs, fixes, and hard-won patterns. The in-repo, shared version of per-machine agent memory: every agent and human that opens the repo sees it.
One fix per file. Filename:
docs/solutions/<area>-<short-slug>.md.
Required frontmatter:
---
module: <which part of the codebase, e.g. "booking", "auth", "build">
tags: [<keywords for search>]
problem_type: bug | gotcha | pattern | workflow
date: YYYY-MM-DD
---
## Problem
What went wrong / what's confusing.
## Cause
Why it happens (the root cause, not the symptom).
## Fix
What to do. Concrete, copy-pasteable where possible.
When a repo already grows an inline “Corrections Log” or “Things
Claude Has Learned” section, migrate those entries into
docs/solutions/ one file each and leave a one-line
pointer in AGENTS.md.
Not every entry is a bug. Recurring slop — agent output that
compiles, passes the cheap checks, and looks plausible but is subtly
wrong — gets logged the same way (problem_type: pattern),
one category per file. The categories that show up everywhere:
Capture a category once and it becomes context that prevents it forever. The slop list is the institutional memory of how agents fail on this codebase — and the review lens for §10’s contract check.
A fix log only pays off if entries get read into
AGENTS.md and each other, not just accumulated as files an
agent might grep. Retrieval re-derives an answer from raw entries on
every session and compounds nothing; compilation folds an entry’s
implication into the standing instructions once, so every future session
starts from the compiled result instead of re-discovering it.
Concretely: when a fix-log entry reveals a rule an agent should follow
by default (not just a past incident to know about), promote that rule
into AGENTS.md’s Gotchas or Conventions section — don’t
leave it as something only found by searching
docs/solutions/. The entry stays as the record of
why; the rule it produced belongs in the file agents read every
session.
Prose is the floor of compilation, not the ceiling. When a logged
pattern is mechanically checkable, promote it past
AGENTS.md into a linter or pre-commit hook — one that
fixes the problem (--fix), or at minimum
blocks it, never one that only flags it. A rule enforced by a hook
cannot be skimmed past, and it frees the instruction file’s budget for
rules that need judgment. The same goes for workflows: a multi-step
incantation agents keep re-deriving (how to kick off a review, how to
run one targeted test) gets compiled into a small script in the repo’s
bin/, pointed to from AGENTS.md.
A hook only bites the surface it watches. A PreToolUse
hook that gates the Read tool is bypassed the moment the
agent reaches the same file through
cat/head/tail/less
in a Bash call — same capability, different surface, and
the guard never fires. So guard every surface that reaches the
capability, not just the obvious one: pair the
Read-tool gate with a Bash gate that catches
the shell path, and let genuinely targeted access through (a
Read with an offset/limit, a
cat that pipes or redirects rather than dumping into
context) so the guard blocks the bypass without blocking real work.
templates/hooks/scripts/guard-large-read.sh is a copy-in
pair implementing exactly this — one script, two matchers. The rule
generalizes past this one case: any hook enforcing a limit has to be
checked against the cheapest way around it, because an agent under a
blocked path will find that way without being told to.
Add entries one at a time. Write a fix-log entry
right after the incident, while the cause is fresh. The highest-signal
trigger is a human correcting the agent on something the instructions
should have prevented: log it and promote the rule in the same session,
not in a later documentation pass. Cross-link each entry to related
entries and to the AGENTS.md rule it feeds (§3’s “Keep in
sync” is the place to declare that link if it’s easy to miss). Do not
batch-import a backlog of old incidents in one pass — a bulk import
produces isolated files with no cross-links and no promoted rules, which
is a pile, not a compiled fix log.
Promoting a rule into a hook (above) only bites if the hook actually
runs before the commit lands. The reliable shape is a review
marker: a review step (a /code-review, a reviewer
agent, a lint pass) writes a marker file on pass, and a
PreToolUse hook blocks git commit
until the marker exists. The difference from a convention is
that the agent cannot skip it — an unreviewed commit is refused, not
merely frowned upon.
Make the marker session-scoped (keyed by repo +
agent session id), so parallel agent sessions in the same checkout don’t
clear each other’s gate, and so the requirement resets per session
rather than leaking across unrelated work. Scope the strictness
to the diff: require only a light review for most changes, and
a deeper one (a data-migration reviewer, a security pass) only when
the staged paths match the sensitive set — a blanket heavy gate on
every commit gets disabled within a week. Keep a single documented
bypass for genuine exceptions (an env flag, --no-verify); a
gate with no escape hatch gets ripped out instead of bypassed.
templates/hooks/scripts/review-gate.sh is a copy-in
implementation.
Some rules can’t be enforced as a hard line without failing on day
one. “No source file over 500 lines” is unachievable in a repo that
already has fifty such files, so the rule never ships. Ratchet
instead: snapshot the current count of the thing you want less
of (files over N lines, TODOs, suppressions,
any-casts) into a committed baseline, and fail CI only when
a metric goes up. Existing debt is grandfathered; new debt is
blocked; lowering the baseline is a deliberate commit, so the number
only moves in the good direction — never silently. This turns an
aspiration a repo can’t meet today into a gate it can adopt today and
tighten over time. templates/hooks/scripts/ratchet.sh is a
language-agnostic implementation; edit its metric list for the repo.
The fix log (above) captures what broke and why — after the
fact. Its forward-looking twin is a short decision
record written before a significant change: what’s
being decided, why, and the alternatives rejected. Keep it lightweight —
a few paragraphs in docs/adr/NNNN-slug.md — and, exactly
like the fix log, bound it with an explicit skip-list
so it doesn’t become ceremony. Require it for: new features,
architectural changes, new external integrations, changes spanning many
modules, or a new cross-cutting pattern. Do not require
it for: bug fixes, single-file refactors, doc-only changes, or test
additions. The rule earns its keep precisely because it says loudly when
it does not apply — an unbounded “write an ADR for everything”
is the drift, not the discipline.
The alternatives are generated before the choice, not after it. “Alternatives rejected” is the one field a record can satisfy dishonestly: having built the thing, write down two options nobody weighed and mark them rejected. It reads identical to real deliberation and carries none of it. The tell is that the rejected options are always strawmen — no one lists a rival they might have picked. This failure is near-universal for agents, which answer an architectural question with one confident design and implement it; the search never happened, so the record documents a preference with formatting. Make the field carry weight: each alternative names a condition under which it would have won (“if we needed X, this one”), and the record is written at the point the design is still open. An alternative with no winning condition wasn’t in the running, and listing it is decoration. This is §10’s “the author is not the judge” moved one step earlier — a producer who never generated a rival cannot have chosen between them.
When two files must agree, say so in writing. Drop a
## Keep in sync block into AGENTS.md:
## Keep in sync
- Add an env var → document it in `CONFIGURATION.md` (or `.env.example`).
- Add a CLI flag / route → update the relevant section here and the README example.
- Change <file A> → update <file B> (and the test that pins them, if any).
List only the pairs that actually drift in this repo. The rule exists because prose inventories rot: a hand-kept “Inventory (legacy)” file-list drifts from the codebase until it is removed. Prefer “read the directory” over a hand-kept list; where a list is unavoidable, pin it with a sync rule or a test.
A sync contract catches drift between files. It cannot catch
AGENTS.md going stale by itself — a build command that
quietly stopped working breaks no pair, so no rule fires, and the file
keeps asserting it with full authority. Since §1 makes
AGENTS.md the file every agent reads on every task, an
unverifiable claim there is the most expensive stale instruction in the
repo. So give the executable claims a way to be re-checked, the same way
§11 asks of skills:
## Tech stack (source of truth: package.json)) and cite
the field on the individual claim where it isn’t obvious. Re-verifying
becomes a grep instead of an investigation, and a reader
who hits a contradiction knows immediately which side is
authoritative.Both rules above assume the stale claim has a source you can re-read
— a version in package.json, a command in a workflow file.
Some claims have no such source because they are a judgment:
status: complete, done, shipped,
active. Nothing drifts against them, so no sync
rule fires, and a status left behind by abandoned work keeps asserting
itself to every agent that reads the file.
Where the repo already contains the evidence for such a judgment, derive it a second way and compare:
in-review from open tickets [#13] and unsigned ACs
[none]” lets a reader adjudicate on the spot; a bare “status drift”
sends them to re-derive it. Name empty inputs explicitly, so “checked,
found none” is distinguishable from “not checked”.Reach for this only where the derivation is cheap and unambiguous. A derived status that is itself a guess is a second unreliable claim, not a check on the first — and §9’s rule that an uncalibrated checker is worse than none applies here exactly as it does to model-graded review.
Only for repos with config that fails silently (an
.env with secrets, a required output dir). A
SessionStart hook fixes the common problems before they
bite instead of after.
hooks/hooks.json:
{
"hooks": {
"SessionStart": [
{
"matcher": "",
"hooks": [
{
"type": "command",
"command": "bash \"${CLAUDE_PLUGIN_ROOT:-.}/hooks/scripts/check-config.sh\""
}
]
}
]
}
}
The script (template in
templates/hooks/scripts/check-config.sh) does two
generically-useful things, both cross-platform-guarded:
mkdir -p
the output/cache dir).chmod 600 a loose-permission
.env and warn, so a world-readable secrets file
gets locked down at session start.Adapt the env-file path and dir per repo. Skip the hook entirely where there’s no silent-failure config to heal. Do not add ceremony for its own sake.
A SessionStart hook may also inject orienting
context — repo state, open work, health — so a session starts with
its bearings instead of re-deriving them. That is a different job from
the repair above and carries its own risk: it spends prompt space on
every session, and it is a second place for repo facts to live (§1).
Three disciplines keep it honest:
Commits in any repo under this standard must be authored by one of a small, explicit set of sanctioned identities. No stray author (a work email, a machine default, a bot) should ever land in history. Pick your allowed identities and list them, e.g.:
you <you@example.com>you-alt <you-alt@example.com>Before committing, verify the local identity resolves to one of them:
git config user.name && git config user.email
If it doesn’t, set it per-repo
(git config user.email you@example.com). Do
not commit under a different identity and fix it later.
An agent committing on your behalf uses whichever sanctioned identity
the repo is already configured for; if unset, fall back to a documented
default.
Co-author trailers (Co-Authored-By:) for the agent are
fine and don’t count as the commit author.
Separate from whose identity a commit carries is whether the work was agent-generated at all — and that must never be invisible. Disclosure is continuous, not a one-time note:
Assisted-by: (or Co-Authored-By:)
trailer naming the agent and whether it acted autonomously or
under direct human supervision. A human git identity on the commit does
not exempt it — the trailer states agent involvement regardless of whose
name is on the author line.Multi-account hosting. If your repos live under more than one GitHub (or GitLab) account, remember the hosting account is separate from commit identity, and CLIs like
ghkeep only one account active at a time. Working in a repo owned by a non-default account without switching first (gh auth switch --user <account>) makes reads/pushes hit the wrong account, which returns a bare404 / repository not found, a silent “wrong active account,” not a missing repo. Note the required account at the top of that repo’sAGENTS.mdand switch before anygh/push operation.
Default: commit straight to the default branch
(main/master) and push it. For solo /
small-team repos, a feature branch + PR for routine work just adds
ceremony and leaves stale branches behind (see the anti-pattern below).
Complete the loop: commit and
git push origin <default>, so the work is actually on
the remote, not parked on a local branch waiting for a second ask.
Branch + PR only when the change is risky. Open a branch instead of committing to main when the change is any of:
For those, branch off the default, push the branch, and open a PR so main stays green. Everything else goes straight to main. The user can always override in either direction (“just commit it”, “put it on a branch”); when they do, that wins for that change.
Do not create a
feature/add-x branch for a routine, low-risk
edit (docs, a diagram, a copy tweak, a one-line fix) and then merge it
yourself moments later. The branch adds no review value on a solo repo,
and if it’s fast-forwarded or rebased into main the leftover branch
lingers on the remote showing a misleading “Compare & pull request”
banner. Commit low-risk work directly to main; reserve branches for the
risky cases above. If a redundant branch does get created, delete it
(remote + local) once its content is on main.
If you deploy across more than one account on a host (Vercel, Netlify, Fly, Cloudflare, etc.), the CLI usually keeps one account logged in at a time, and the account that owns a deployment is independent of the git remote owner. Deploying under the wrong account fails (“Could not retrieve Project Settings”) or, worse, deploys to the wrong project.
Rules for any agent about to deploy:
deploy-check wrapper that reads the required account from
the repo’s AGENTS.md and diffs it against the active login
pays for itself.> **Deploy:** line in AGENTS.md, or
cross-check the linked project’s org id against a maintained
account→repo map.find . -path '*/.vercel/project.json', adjust per host)
and run deploy commands from that directory, otherwise
the CLI silently uses whatever account is logged in.--token and avoids flipping the
global CLI session to the wrong account. The fix for a wrong-account
error is the correct account’s token, not a bare
login that mutates global state.If the account-check reports a mismatch, stop and switch accounts. Do not guess your way through auth.
Keep the concrete account↔︎repo map (emails, org ids, domains) in a private file or a secrets manager, not in this public standard. This section is the policy; your account list is config.
If more than one model or agent CLI is available, keep a small ranking table of the models you use, scored on three axes: cost, intelligence (how hard a problem it can be handed unsupervised), and taste (UI/UX, code quality, API design, copy). The table is config — keep it private and current. This section is the policy for using it.
A worked, harness-specific setup for all of this is in examples/orchestration-workflow.md.
Rules for work that spans subagents, background jobs, or hours. The theme: files are the state, context is scarce, and the user is not a polling target.
TODOS.md or a tracker reachable by
CLI, never a list that lives only in chat memory. Where the queue’s
state can be derived from the work itself, prefer the derivation over
the hand-maintained field (§3): a queue whose entries can be checked
against reality is the only kind that survives a long unattended
run.git commit -a — commit with explicit pathspecs so one
commit can’t bundle another worker’s work-in-progress. One commit per
completed chunk.cwd, or simply the human’s own uncommitted work. Treat the
working tree as on loan: your edits are yours, and everything
else in it belongs to someone still using it. So work
additively — edit your files, stage them by path,
commit them, and leave the rest of the tree as you found it; a dirty
tree is the normal resting state of a shared checkout, not a problem to
clear before starting. An agent therefore never runs, unless asked for
it right now: git stash in any form (including the
--autostash that rides along with
git pull --rebase, which is how it usually arrives), a
branch switch, a git worktree add/remove, or any reset that
discards work. Each of these silently takes a peer’s
uncommitted changes, and the peer’s only symptom is that its
edits vanished — so it rewrites them, racing a stash entry nobody will
pop. When the tree holds files you don’t recognize, leave them
and keep going: unrecognized is not the same as stray, and
tidying up is the most common way an agent destroys work it was never
asked to touch. A shared tree that genuinely blocks you is an
escalation, not a cleanup job.Rules for keeping autonomous work safe when things go wrong — and for deciding, in writing, when a human takes over.
AGENTS.md or the loop’s config):
templates/docs/EXAMPLE-acceptance-criteria.md. And never
let the model that produced the work be the sole judge of whether it met
the bar: a producer grading itself struggles to notice it went in the
wrong direction, and self-reported completion is a claim, not a result —
“done” is what the compiler, the tests, and an independent checker say.
And “the tests” means the change exercised the way a user actually hits
it — run the app, drive the flow end to end — not only the unit tests
written alongside the change, which encode the author’s own assumptions.
Route the gate through a test, an independent reviewer, or a different
model (§8, §9). Review the output against the contract — “did
it satisfy the contract, and did it add anything beyond it?” — not the
diff line by line; line-by-line reading is how slop (§2) slips through
while the reviewer feels thorough.AGENTS.md (§1) and the fix log (§2) cover a repo’s
day-to-day operating knowledge. Some repos also carry knowledge that
lives only in one person’s head — the debugging instincts, the settled
arguments, the unwritten rules nobody documented because the senior
engineer just knew them. When that knowledge needs to survive
the person, or needs to run on a cheaper model than the one that holds
it today, generalize it into a skill library
(.claude/skills/<name>/SKILL.md or the
harness-equivalent path) instead of letting it stay tacit.
A skill library is a compiled artifact in the same sense as §2’s “compile, don’t retrieve”: it is the settled output of someone’s tacit judgment, written once so a reader gets the answer directly instead of re-deriving it from raw history, Slack threads, or trial and error. A library that just links out to source material without stating the settled rule has not actually succeeded the knowledge — it has relocated the retrieval step.
Discover before you write. Read the repo like an incoming engineer first — history, docs, tests, CI, the trail of reverted or abandoned attempts — then ask a small, bounded number of questions for what the repo genuinely cannot tell you (the hardest live problem, the unwritten discipline rules, who the audience is and what they don’t know). Fold the answers into the library; don’t author from assumption.
One skill, one topic — no duplicate homes for a fact. Split a library by concern (architecture, debugging, config, domain reference, validation discipline, the hardest live problem as its own guided runbook) rather than one sprawling file. Each skill states when not to use it and which sibling to use instead, so a loader doesn’t have to guess.
The description is the routing contract. A skill’s frontmatter description is all a loader sees before deciding to read it. State what the skill does, when to use it (the trigger phrases someone would actually say), and what distinguishes it from siblings — and never summarize the workflow itself, or the loader follows the summary and skips the body. Debug accordingly: a skill that doesn’t fire has a description problem; a skill that fires and produces the wrong output has a body problem.
Invocation is a cost decision — put each skill in the cheapest tier that meets its activation need. Three tiers, and only the last one costs prompt space:
| Tier | How it fires | Resident cost |
|---|---|---|
| Referenced — read at the point of use, nothing stored | you name the path | none |
| Saved — checked into the repo, invoked by name | you name the skill | none |
| Auto-firing — description is loaded every session | unprompted, on a description match | ~50–280 tokens per skill, on every message |
A skill only needs the top tier if the agent must reach it without being asked. Everything else — content, and keeping a copy — costs nothing resident. Default to the lowest tier that works and promote deliberately; nothing should reach the top tier by accident. Note the second cost of the bottom tiers: the human becomes the index that must remember the skill exists. When by-name skills multiply past what a person can remember, add one router skill that names the others and says when to reach for each — that router, not the whole library, is what earns auto-firing.
The auto-fire budget is small, and it is a budget. Assume well under a hundred reliably-firing skills per agent, and treat that as a working ceiling rather than a measured one. The cap is not “however many descriptions fit in the window”: input length alone degrades a model’s reasoning well before the nominal context limit, a description at the top of the context competes with every turn, file read, and tool result that lands after it, and each added description dilutes the rest. So an auto-firing skill is not free even when there is room for it — count them, and know what each one is buying.
Don’t bid for attention in the description. A missed match is silent — nobody is told the skill was there and didn’t fire — so the tempting fix is to pad the one-liner with shouted trigger and skip conditions until it wins. It works for one skill and costs everyone: a padded description runs an order of magnitude more tokens than a quiet one, paid on every message, and the extra imperatives compete with the repo’s actual instructions. Keep the description to what the skill does and when to reach for it. If it still doesn’t fire, the honest fixes are a sharper trigger phrase, a narrower scope, or demoting it to by-name invocation behind a router — never volume.
Lean body, deep references. Keep the skill file short enough to load cheaply; move rarely-needed detail into reference files linked one level deep (no chains), each with an explicit “read this when …” trigger. Push anything deterministic into a script the skill invokes rather than prose the model re-derives. Don’t re-teach what the model already knows — every paragraph must carry knowledge specific to this repo or this person.
Test the trigger, not just the content. Before calling a skill done, pose the task in the words a user would use, without naming the skill, and check it fires — then pose a neighboring task and check it doesn’t. A library whose skills only load when invoked by name has failed at succession: the person who knew which file to open is exactly who’s gone.
Ground truth only. Every command, flag, path,
and claim gets verified against the repo before it’s written down — a
wrong runbook is worse than no runbook, because it’s trusted. Unproven
or open items stay explicitly labeled as such; nothing in the library
may contradict AGENTS.md or route around this standard’s
ship gates.
Provenance and re-verification. Date-stamp anything that can drift (config defaults, flag lists, tool versions) and give each skill a one-line command that re-checks it. A skill without a re-verification path decays into the exact stale-instruction problem §1 exists to prevent.
Write-scope discipline. A skill-authoring pass
writes only inside the skills directory; it doesn’t mutate the rest of
the repo. Keep the authoring and review passes separate — author, then
have an independent pass check facts, check for contradictions between
skills or with AGENTS.md, and check that a zero-context
reader could actually follow each one.
This is expensive relative to a normal AGENTS.md update,
so reserve it for knowledge that is genuinely at risk of being lost or
that must run on a materially cheaper model than the one that holds it —
not as the default way to document a repo.
The standard fails one skipped step at a time, and every skip arrives wearing a plausible excuse. These are the recurring ones, each with why it doesn’t hold. An agent about to act on an excuse from the left column should treat that as the signal to stop and follow the section on the right instead.
| Rationalization | Reality |
|---|---|
| “I’ll add this note to CLAUDE.md too, so it’s visible everywhere.” | Two copies is the exact failure mode §1 exists to prevent. The
second copy starts drifting the moment it lands; put it in
AGENTS.md once. |
| “This fix is too small to log.” | Size of fix and cost of rediscovery are unrelated — one-line fixes with invisible causes are exactly what the fix log (§2) is for. If it took real digging, log it. |
| “The fix-log entry exists; anyone can grep for it.” | Retrieval isn’t compilation (§2). If the entry implies a standing
rule, promote the rule into AGENTS.md — an entry only found
by searching protects nobody by default. |
| “I’ll update the paired file in a follow-up.” | The follow-up is the step that never happens; that’s why the pair is
listed in ## Keep in sync (§3). The sync is part of this
change, not a second task. |
| “My identity is on the commit, so the agent trailer is redundant.” | The author line says whose commit it is; the trailer says how it was made. Agent involvement must never be invisible (§5), whoever’s name is on it. |
| “Safer to put this on a branch.” | Unless it’s risky per §6’s list (migration, wide refactor, hard to revert, build-breaking), the branch is ceremony that leaves litter behind. Safe-by-default is main. |
| “Tests are slow and this change is obviously safe.” | “Obviously safe” is a self-grade, and the author is not the judge (§10). The ship gate exists precisely for changes that look safe. Run it. |
| “The output looks right, so it’s done.” | Looks-right is how slop (§2) ships. Done is what the tests and an independent check say (§10) — verify against the contract, not the vibe. |
| “I matched the file next to it, so it’s consistent.” | In a repo with strata (§1), the nearest file is a random layer, not
the current one. Match the layer AGENTS.md names as
current, and leave frozen layers unseeded. |
| “The newest pattern has the most files, so that’s the canonical one.” | File counts and recency establish which layers are live, never which one is meant to win (§1). That’s a decision a human makes; inferring it enshrines a layer nobody chose. |
| “The query ran clean, so the numbers are right.” | Clean execution proves the query ran, not that it covered the right rows (§10). State the population and reconcile it against a second source before reasoning on the result. |
| “We’re still learning the tool, so a rough result is expected here.” | Then it wasn’t production work (§10). Learning gets its own bounded space; a real deliverable clears the same gates regardless of how new the tooling is. |
| “The tree was dirty, so I stashed it first to get a clean start.” | A dirty tree is the resting state of a shared checkout, not a
problem to clear (§9). git stash takes every modified file,
including a peer’s and the human’s, and the victim’s only symptom is
that its work vanished. Work additively or escalate. |
| “These files aren’t mine, so I tidied them up.” | Unrecognized is not stray (§9). Cleaning up the tree is the most common way an agent destroys work nobody asked it to touch. Leave them and keep going. |
| “It’s one line in package.json.” | Size is not the test; reversibility is (§10, §6). A dependency buys a transitive tree, a licence, and a supply-chain surface, and gets harder to remove with every import. Propose and stop. |
| “The command is documented in AGENTS.md, so it’s covered.” | Documented is not enforced (§3). A command that quietly stopped working breaks no pair, so no sync rule fires and the file keeps asserting it. Pin the executable claim to a gate. |
CLAUDE.md
(scratchpad copy; git init + commit first if the repo isn’t
under git).AGENTS.md
(if AGENTS.md exists as a stub, merge into it; if it’s a duplicate, the
content is already there).CLAUDE.md with the single line
@AGENTS.md.docs/solutions/*.md with frontmatter; leave a
pointer in AGENTS.md.## Keep in sync block for this
repo’s drift-prone file pairs.head -2 CLAUDE.md shows the
include; diff AGENTS.md against the backup to confirm zero content loss
(relocation only); git diff is reviewable.