Files
Nomarchy/agent/DELEGATION.md
Bernardo Magri c0fc16e25c
All checks were successful
Check / eval (push) Successful in 4m28s
docs(agent): vendor-neutral agent docs — AGENTS.md entry, agent/ SoT, .claude shims
The repo is maintained by agents from multiple vendors, so agent
instructions move out of vendor-specific locations into shared,
git-tracked markdown:

- AGENTS.md is the new entry point for any harness; CLAUDE.md becomes a
  symlink to it (Claude Code keeps working unchanged).
- Skill bodies relocate to agent/: VERIFICATION.md (the enforcement
  rules, ex .claude/skills/nomarchy), DELEGATION.md (capability tiers
  light/standard/frontier, scout/runner role contracts, token economy,
  parallel fan-out — consolidates the CLAUDE.md model table, LOOP.md's
  economy section, and skill §6.5 into one place; vendor model names
  survive only in the per-harness mapping table), THEME-DESIGN.md
  (ex .claude/skills/theme, which previously lacked frontmatter).
- .claude/ shrinks to a thin Claude Code adapter: settings, subagent
  defs, and skill shims that route into agent/.
- Maps updated: agent/README.md (instructions vs state vs adapters),
  docs/README.md, README.md layout tree.

Cleanup: stale old_distro .gitignore entry dropped (dir long gone);
local result*/__pycache__ artifacts removed; JOURNAL.md rotated (29
recent entries kept, 120 older moved to agent/JOURNAL-ARCHIVE.md,
rotation rule documented in the header).

Verification: V0 (docs/meta only) — nix flake check --no-build exit 0
with the tracked symlink; grep sweep confirms no vendor model names
outside DELEGATION.md's mapping table.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-11 08:58:25 +01:00

5.1 KiB

Delegation — capability tiers, roles, token economy

How to spend model capacity in this repo, for any agent harness. Tasks are matched to capability tiers, not vendor model names; each harness maps tiers to its own models (table at the bottom). GOALS says what to build, CONVENTIONS how to write it — this says who does which part.

The dividing line: evidence vs judgment

Cheap models gather evidence; they never make verification claims or design decisions. The following always stay on the frontier tier and are never delegated downward:

  • deciding what to test and the regression scope (VERIFICATION.md)
  • interpreting an ambiguous or flaky failure
  • viewing screenshots and judging visual quality — aesthetic judgment is exactly what small models do badly, and it's load-bearing here
  • writing non-trivial Nix (module structure, overlays, cross-cutting refactors)
  • the final diff review and report

Product calls ("finish vs quarantine", promoting PROPOSED items) sit above even that: they belong to the human.

Tiers

Tier Best for Delegate to it?
light bulk mechanical: search, summarize, audit sweeps, running the harness freely — never for judgment
standard well-specified scoped edits, routine research when the spec is already written
frontier design, novel code, ambiguous failures, taste this is the loop's own tier

Reasoning effort: default moderate; go maximum only for the hardest calls. A frontier parent may spawn a frontier child for one hard call. Work above your tier gets returned, not attempted.

Standing roles

Two mechanical roles exist for any harness that supports subagents (Claude Code implementations: .claude/agents/; other harnesses implement the same contracts in their own format):

  • scout (light, read-only): locate where things are defined, map which files touch a subsystem, scan build logs / serial output for error lines, docs-vs-code drift sweeps. Reports facts with paths and line numbers, quotes the minimum snippet, says "not found" plainly — never guesses, never recommends.
  • runner (light, executes): builds, VM boots, screenshot capture, the scripted tools/ checks. Headless and unattended, every wait bounded by a timeout. Returns commands, exit codes, wall time, and artifact paths — never marks anything passed or verified; the caller makes the verification claim.

Use them for pure information-gathering or pure execution instead of pulling bulk (log files, wide scans) into the main context. When a result surprises you, spot-check it yourself before building on it — cheap models are allowed to be wrong about hard things, which is precisely why they're not allowed to make claims.

Economy rules

  • Only delegate when writing the spec is cheaper than doing the work — a spawned agent starts cold and must re-derive context. A one-file read is cheaper done directly.
  • Brief every child cold; point at the spec, don't restate it. "The spec is in agent/BACKLOG.md #NN — implement it" plus only the constraints (scope files, branch, no-VM, no-push).
  • Match the tier to the task, not the prestige. An item whose spec is already written (exact files, exact fixes) is standard-tier work; reserve frontier children for genuine multi-step reasoning. This is the single biggest saving.
  • The strong model writes the spec, reviews the result, and owns the commit.

Fanning out parallel work (V0/V1, no VM)

When several NEXT items are independent and don't need the VM, spread them across worktree-isolated subagents in parallel:

  • Disjoint file lanes. Partition items so no two agents touch the same file (map the touched files first). If two items must share a file (README, flake.nix, rofi.nix), give both to one agent or keep one for yourself.
  • Isolation + you own landing. Each agent works in an isolated worktree, commits to its own branch, and never pushes or touches main/v1. You review each diff, cherry-pick onto main, and do the bookkeeping — a single landing agent can't race itself. Clean up worktrees and branches after landing.
  • Lean on scriptable checks as primary evidence. Where a deterministic tools/ check or checks.* guard already proves the property, that near-free run is the V0/V1 evidence.
  • Batch V2 at the end, once. Delegated visual/behavioural items come back "V2 pending"; collect the landed changes and do one VM pass covering all of them — the VM render + screenshot review is the most expensive step in the loop; amortise it.
  • Re-verify on main, but leanly. After landing, confirm the agent's V0/V1 on the merged tree with a targeted build, not a full re-run. Trust-but-spot-check scales; blind re-running doesn't.

The judgment list above still holds: you review every diff before it lands.

Per-harness model mapping

Tier Claude Code
light haiku
standard sonnet
frontier the session's top model (opus and up)

Other harnesses: add a column when one is actually used on this repo.