Cheat Sheet
Amplify Your AI Agents. Give Them a Computer.
The quick-reference version of the field guide. Every task gets its own environment, and every layer takes a check off your plate.
The eight proofs
Every change has to prove the same things before it ships:
| Tier | B | 1 | 2 | 3 | 4 | 5 | 6 | |
|---|---|---|---|---|---|---|---|---|
| 1 | It builds and follows the house style. | |||||||
| 2 | It does what it should. | |||||||
| 3 | It's secure. | |||||||
| 4 | It's fast enough. | |||||||
| 5 | It's maintainable and fits the architecture. | |||||||
| 6 | It works for a user. | |||||||
| 7 | It runs with the rest of the system. | a | b | |||||
| 8 | It's healthy in production. |
The tiers hand each proof to the system, cheapest and most certain first.
| Tier | Misses | Cheapest start |
|---|---|---|
| Baseline | Its own assumptions; whether the code runs | Free |
| 1. Deterministic checks | Anything that needs judgement | A hook that runs them after every edit |
| 2. Independent AI reviewer | Blind spots it shares with the author's model | A CI review step with your own prompts |
| 3. Eyes | Whether the flow feels right | Playwright MCP on the web, Maestro or mobile-mcp for apps, and a frontend that runs without its backend |
| 4. Run the stack, once per worktree | Data-scale bugs | Names and ports derived from the worktree |
| 5. An environment per PR | Anything an empty database hides | Create and destroy a real database per preview |
| 6. Access to the real world | Nothing, if you keep user data out | One read-only MCP server for logs |
The split between tiers 1 and 2: if a tool can catch it every time, make it a check. If it needs judgement, put it in the reviewer's prompt. When the reviewer flags the same thing twice, turn it into a check.
The verification contract
copy into CLAUDE.md or AGENTS.mdPaste this into CLAUDE.md or AGENTS.md and adjust it to your stack. It's a trimmed version of my own. The full file, and my pr-feedback-loop skill, are in my dotfiles (github.com/acrookston/dotfiles).
## Definition of ready Substantial work starts only when the plan states: 1. Intent: the problem, the outcome, and what is out of scope. 2. Acceptance criteria: observable behaviour, including error and empty states. 3. Verification plan: the check that proves each criterion. Say up front if a criterion can't be checked automatically. ## Definition of done - Run the checks that cover the change (types, lint, tests; for UI, drive it in the browser). Show the evidence: the command and its result, or a screenshot. List anything left unverified. - Bug fixes start with a test that fails without the fix. - Before opening a PR, review the diff in a fresh subagent. Fix findings that affect correctness or the acceptance criteria. ## Autonomy - Don't stop to ask what you can check yourself. Come back when the checks pass, or when you're blocked on a decision that's mine. ## Local environment - Start the stack with `<your boot command>`. - Worktrees share ports. Check `lsof -i :<port>` before starting servers. - Never kill dev servers or databases you didn't start.
The worktree test
The test: can two sessions run the full check at once without knowing about each other? If not, you have worktrees without parallelism.
Fixes, from cheap to thorough:
- A Compose project per worktree. Set
COMPOSE_PROJECT_NAMEfrom the worktree name. Derive ports, database name and Redis URL from it too. - One database per worktree. Seed a template once, then
CREATE DATABASE wt_feature TEMPLATE app_seeded. Give each worktree its own Redis index or key prefix. - A cloud session per task. Each session gets its own machine. Your boot command must run headless.
What I check by hand
- Does it do what I asked? Open the preview and click through.
- Does it look and feel right? Agents can confirm a button exists. They can't tell you the flow feels wrong.
- Are the fundamentals sound? Models, migrations, API contracts, external calls, module boundaries, deployment. Plus a quick security pass.
What the agent never gets
- User data or personal data. If it shows up in logs, fix the logging.
- Write access to production.
- Production secrets in preview environments.