Andrew Crookston. @acr
Updated September 30, 2026 · 18 essays · Stockholm RSS ↗

Cheat Sheet

Amplify Your AI Agents. Give Them a Computer.

The quick-reference version of the field guide. Every task gets its own environment, and every layer takes a check off your plate.

Andrew Crookston

The six tiers drawn as a stepped column, with the baseline at the bottom and tier 6 at the top B 1 2 3 4 5 6 Baseline Reads back own work 1. Deterministic checks Linters, types, tests 2. Independent AI reviewer Separate agent in CI 3. Eyes Drives browser or simulator 4. Run the stack, once per worktree One boot command 5. An environment per PR Short-lived preview 6. Access to the real world Read-only logs, CI, deploys
Fig. 1: The six tiers.

The eight proofs

Every change has to prove the same things before it ships:

TierB123456
1It builds and follows the house style.
2It does what it should.
3It's secure.
4It's fast enough.
5It's maintainable and fits the architecture.
6It works for a user.
7It runs with the rest of the system.ab
8It's healthy in production.
Proves The known parts of 3 a 7, locally b 7, integrated, with a URL for reviewers B Baseline proves obvious mistakes

The tiers hand each proof to the system, cheapest and most certain first.

TierMissesCheapest start
BaselineIts own assumptions; whether the code runsFree
1. Deterministic checksAnything that needs judgementA hook that runs them after every edit
2. Independent AI reviewerBlind spots it shares with the author's modelA CI review step with your own prompts
3. EyesWhether the flow feels rightPlaywright MCP on the web, Maestro or mobile-mcp for apps, and a frontend that runs without its backend
4. Run the stack, once per worktreeData-scale bugsNames and ports derived from the worktree
5. An environment per PRAnything an empty database hidesCreate and destroy a real database per preview
6. Access to the real worldNothing, if you keep user data outOne read-only MCP server for logs

The split between tiers 1 and 2: if a tool can catch it every time, make it a check. If it needs judgement, put it in the reviewer's prompt. When the reviewer flags the same thing twice, turn it into a check.

The verification contract

copy into CLAUDE.md or AGENTS.md

Paste this into CLAUDE.md or AGENTS.md and adjust it to your stack. It's a trimmed version of my own. The full file, and my pr-feedback-loop skill, are in my dotfiles (github.com/acrookston/dotfiles).

## Definition of ready
Substantial work starts only when the plan states:
1. Intent: the problem, the outcome, and what is out of scope.
2. Acceptance criteria: observable behaviour, including error and empty states.
3. Verification plan: the check that proves each criterion.
   Say up front if a criterion can't be checked automatically.

## Definition of done
- Run the checks that cover the change (types, lint, tests; for UI, drive it
  in the browser). Show the evidence: the command and its result, or a
  screenshot. List anything left unverified.
- Bug fixes start with a test that fails without the fix.
- Before opening a PR, review the diff in a fresh subagent. Fix findings
  that affect correctness or the acceptance criteria.

## Autonomy
- Don't stop to ask what you can check yourself. Come back when the checks
  pass, or when you're blocked on a decision that's mine.

## Local environment
- Start the stack with `<your boot command>`.
- Worktrees share ports. Check `lsof -i :<port>` before starting servers.
- Never kill dev servers or databases you didn't start.

The worktree test

The test: can two sessions run the full check at once without knowing about each other? If not, you have worktrees without parallelism.

Fixes, from cheap to thorough:

  1. A Compose project per worktree. Set COMPOSE_PROJECT_NAME from the worktree name. Derive ports, database name and Redis URL from it too.
  2. One database per worktree. Seed a template once, then CREATE DATABASE wt_feature TEMPLATE app_seeded. Give each worktree its own Redis index or key prefix.
  3. A cloud session per task. Each session gets its own machine. Your boot command must run headless.

What I check by hand

  1. Does it do what I asked? Open the preview and click through.
  2. Does it look and feel right? Agents can confirm a button exists. They can't tell you the flow feels wrong.
  3. Are the fundamentals sound? Models, migrations, API contracts, external calls, module boundaries, deployment. Plus a quick security pass.

What the agent never gets

  • User data or personal data. If it shows up in logs, fix the logging.
  • Write access to production.
  • Production secrets in preview environments.

Grade your own repo

Paste this into a fresh session:

Then ask the same agent to set up the first missing tier. It knows your language and stack better than any list here.

andrewcrookston.com/articles/amplify-your-ai-agents.html

Prompt

Read https://andrewcrookston.com/articles/amplify-your-ai-agents.html.
Then review this codebase against it. For each of the six tiers, tell me:
- what we already have, with file paths
- what is missing
- the smallest change that would move us up a tier
End with the one change you would make first.