All checks were successful
Check / eval (push) Successful in 4m28s
The repo is maintained by agents from multiple vendors, so agent instructions move out of vendor-specific locations into shared, git-tracked markdown: - AGENTS.md is the new entry point for any harness; CLAUDE.md becomes a symlink to it (Claude Code keeps working unchanged). - Skill bodies relocate to agent/: VERIFICATION.md (the enforcement rules, ex .claude/skills/nomarchy), DELEGATION.md (capability tiers light/standard/frontier, scout/runner role contracts, token economy, parallel fan-out — consolidates the CLAUDE.md model table, LOOP.md's economy section, and skill §6.5 into one place; vendor model names survive only in the per-harness mapping table), THEME-DESIGN.md (ex .claude/skills/theme, which previously lacked frontmatter). - .claude/ shrinks to a thin Claude Code adapter: settings, subagent defs, and skill shims that route into agent/. - Maps updated: agent/README.md (instructions vs state vs adapters), docs/README.md, README.md layout tree. Cleanup: stale old_distro .gitignore entry dropped (dir long gone); local result*/__pycache__ artifacts removed; JOURNAL.md rotated (29 recent entries kept, 120 older moved to agent/JOURNAL-ARCHIVE.md, rotation rule documented in the header). Verification: V0 (docs/meta only) — nix flake check --no-build exit 0 with the tracked symlink; grep sweep confirms no vendor model names outside DELEGATION.md's mapping table. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
113 lines
5.1 KiB
Markdown
113 lines
5.1 KiB
Markdown
# Delegation — capability tiers, roles, token economy
|
|
|
|
How to spend model capacity in this repo, for **any** agent harness.
|
|
Tasks are matched to capability *tiers*, not vendor model names; each
|
|
harness maps tiers to its own models (table at the bottom). GOALS says
|
|
what to build, CONVENTIONS how to write it — this says who does which
|
|
part.
|
|
|
|
## The dividing line: evidence vs judgment
|
|
|
|
Cheap models gather evidence; they never make verification claims or
|
|
design decisions. The following always stay on the **frontier** tier and
|
|
are never delegated downward:
|
|
|
|
- deciding what to test and the regression scope (VERIFICATION.md)
|
|
- interpreting an ambiguous or flaky failure
|
|
- viewing screenshots and judging visual quality — aesthetic judgment is
|
|
exactly what small models do badly, and it's load-bearing here
|
|
- writing non-trivial Nix (module structure, overlays, cross-cutting
|
|
refactors)
|
|
- the final diff review and report
|
|
|
|
Product calls ("finish vs quarantine", promoting PROPOSED items) sit
|
|
above even that: they belong to the human.
|
|
|
|
## Tiers
|
|
|
|
| Tier | Best for | Delegate to it? |
|
|
|------|----------|-----------------|
|
|
| **light** | bulk mechanical: search, summarize, audit sweeps, running the harness | freely — never for judgment |
|
|
| **standard** | well-specified scoped edits, routine research | when the spec is already written |
|
|
| **frontier** | design, novel code, ambiguous failures, taste | this is the loop's own tier |
|
|
|
|
Reasoning effort: default moderate; go maximum only for the hardest
|
|
calls. A frontier parent may spawn a frontier child for one hard call.
|
|
Work above your tier gets returned, not attempted.
|
|
|
|
## Standing roles
|
|
|
|
Two mechanical roles exist for any harness that supports subagents
|
|
(Claude Code implementations: `.claude/agents/`; other harnesses
|
|
implement the same contracts in their own format):
|
|
|
|
- **scout** (light, read-only): locate where things are defined, map
|
|
which files touch a subsystem, scan build logs / serial output for
|
|
error lines, docs-vs-code drift sweeps. Reports facts with paths and
|
|
line numbers, quotes the minimum snippet, says "not found" plainly —
|
|
never guesses, never recommends.
|
|
- **runner** (light, executes): builds, VM boots, screenshot capture,
|
|
the scripted `tools/` checks. Headless and unattended, every wait
|
|
bounded by a timeout. Returns commands, exit codes, wall time, and
|
|
artifact paths — never marks anything passed or verified; the caller
|
|
makes the verification claim.
|
|
|
|
Use them for pure information-gathering or pure execution instead of
|
|
pulling bulk (log files, wide scans) into the main context. When a
|
|
result surprises you, spot-check it yourself before building on it —
|
|
cheap models are allowed to be wrong about hard things, which is
|
|
precisely why they're not allowed to make claims.
|
|
|
|
## Economy rules
|
|
|
|
- **Only delegate when writing the spec is cheaper than doing the
|
|
work** — a spawned agent starts cold and must re-derive context. A
|
|
one-file read is cheaper done directly.
|
|
- **Brief every child cold; point at the spec, don't restate it.** "The
|
|
spec is in `agent/BACKLOG.md` #NN — implement it" plus only the
|
|
*constraints* (scope files, branch, no-VM, no-push).
|
|
- **Match the tier to the task, not the prestige.** An item whose spec
|
|
is already written (exact files, exact fixes) is standard-tier work;
|
|
reserve frontier children for genuine multi-step reasoning. This is
|
|
the single biggest saving.
|
|
- The strong model writes the spec, reviews the result, and owns the
|
|
commit.
|
|
|
|
## Fanning out parallel work (V0/V1, no VM)
|
|
|
|
When several NEXT items are independent and don't need the VM, spread
|
|
them across worktree-isolated subagents in parallel:
|
|
|
|
- **Disjoint file lanes.** Partition items so no two agents touch the
|
|
same file (map the touched files first). If two items must share a
|
|
file (README, flake.nix, rofi.nix), give both to one agent or keep
|
|
one for yourself.
|
|
- **Isolation + you own landing.** Each agent works in an isolated
|
|
worktree, commits to its own branch, and **never pushes or touches
|
|
`main`/`v1`**. You review each diff, cherry-pick onto `main`, and do
|
|
the bookkeeping — a single landing agent can't race itself. Clean up
|
|
worktrees and branches after landing.
|
|
- **Lean on scriptable checks as primary evidence.** Where a
|
|
deterministic `tools/` check or `checks.*` guard already proves the
|
|
property, that near-free run *is* the V0/V1 evidence.
|
|
- **Batch V2 at the end, once.** Delegated visual/behavioural items come
|
|
back "V2 pending"; collect the landed changes and do **one** VM pass
|
|
covering all of them — the VM render + screenshot review is the most
|
|
expensive step in the loop; amortise it.
|
|
- **Re-verify on `main`, but leanly.** After landing, confirm the
|
|
agent's V0/V1 on the merged tree with a targeted build, not a full
|
|
re-run. Trust-but-spot-check scales; blind re-running doesn't.
|
|
|
|
The judgment list above still holds: *you* review every diff before it
|
|
lands.
|
|
|
|
## Per-harness model mapping
|
|
|
|
| Tier | Claude Code |
|
|
|------|-------------|
|
|
| light | `haiku` |
|
|
| standard | `sonnet` |
|
|
| frontier | the session's top model (`opus` and up) |
|
|
|
|
Other harnesses: add a column when one is actually used on this repo.
|