Files
Nomarchy/agent
Bernardo Magri 67400b07dc
Some checks failed
Check / eval (push) Failing after 7m59s
fix(ci): memory-bounded eval — one output per nix process
check.yml has been permanently red since the runner container was
capped at 2 GB (uncapped, it OOM'd the whole 4 GB VPS). Root cause,
measured with GNU time (eval-cache off): `nix flake check --no-build`
walks every output in ONE evaluator process and peaks at ~6.0 GB RSS
on this repo — the checks.* suite alone is ~15 runNixOSTests, each a
full NixOS eval, all accumulating in one heap.

New tools/ci-eval.sh: identical coverage (all checks.* drvPaths, both
nixosConfigurations toplevels, the HM activationPackage — enumerated
dynamically, no drift), one fresh nix process per output so the heap is
freed between outputs. Cold per-output peaks: worst check 0.99 GB
(hardware-toggles), nomarchy toplevel 0.78 GB, HM 0.60 GB; the one
outlier is nomarchy-live at 2.69 GB (live-ISO eval), which fits the
2 GB cap only via the container's swap allowance (docker's default
--memory-swap is 2x memory). check.yml and bump.yml both call the
script now; bump.yml also gains max-jobs=1/cores=2 (it had no limits,
so its V1 build gate could spawn parallel builders in the capped
container). ROADMAP's stale "no CI today" corrected: bump.yml landed
8fded63 on schedule 2026-07-06.

Verified: V1+ — ci-eval.sh green locally end-to-end (3m22s, tree peak
2.75 GB = nomarchy-live's process); per-output peaks measured
individually. The decisive proof is the next Actions run on the VPS
itself — watch the nomarchy-live step; if it OOMs there, the runner
needs swap allowed: --memory=2g --memory-swap=6g.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-12 19:39:46 +01:00
..

Agent instructions + loop state

Everything an AI agent needs to work on Nomarchy, vendor-neutral and git-tracked. Protocol: LOOP.md. Entry point for every harness: repo-root AGENTS.md.

Instructions (how to work)

File Who writes Role
LOOP.md Human One-iteration protocol (orient → pick → work → verify → commit → record) + the V0V3 ladder
VERIFICATION.md Human (agents propose) Enforcement: preflight, honesty rules, visual protocol, hardware-blocked checks, reporting
DELEGATION.md Human (agents propose) Capability tiers, scout/runner roles, token economy, parallel fan-out
GOALS.md Human (agents propose) Pillars, quality bars, non-goals
CONVENTIONS.md Human (agents propose) How to write code/menu/state while shipping
THEME-DESIGN.md Human (agents propose) Theme/visual design instructions

State (what's happening)

File Who writes Role
BACKLOG.md Both Prioritized queue — only executable work list
JOURNAL.md Agents Append-only iteration log (read last 35 entries; older → JOURNAL-ARCHIVE.md)
MEMORY.md Agents Curated durable gotchas
HARDWARE-QUEUE.md Agents append, human checks On-hardware V3 tests only Bernardo can run

Product / design docs (not a queue)

File Role
../docs/VISION.md v1.0 product themes — agents slice into BACKLOG PROPOSED
../docs/ROADMAP.md Design history + shipped log
../docs/README.md Full docs map

Harness adapters (vendor-specific, thin)

Shared content never lives in an adapter — adapters only register/route into the files above, in whatever format their harness requires.

Path Harness Role
../AGENTS.md any Entry point (CLAUDE.md is a symlink to it)
../.claude/settings.json Claude Code Tool permissions
../.claude/agents/ Claude Code nomarchy-scout / nomarchy-runner role defs (contracts in DELEGATION.md)

Do not put backlog items, vision text, or policy under an adapter directory — it is not shared with other agent runners.

Rules of thumb

  1. Execute from BACKLOG only (NOW → NEXT; never PROPOSED without human triage).
  2. Orient with GOALS + CONVENTIONS + MEMORY + last journal + BACKLOG; when the task is product-shaped, also read the relevant VISION §.
  3. Record lasting design in ROADMAP ✓ when something ships that future humans should know; delete the BACKLOG line.
  4. v1 branch is human-only — never advance from an agent session.