Files
Nomarchy/agent/BACKLOG.md
Bernardo Magri 44aac0fcde
All checks were successful
Check / eval (push) Successful in 3m49s
docs(agent): #138 root-caused — our own reprobe restarts the graph under Chromium
Bernardo's hypothesis (pipewire restarts, Chromium never reconnects) is
right, and the journal names the restarter: dock-audio.nix:153 runs
`systemctl --user restart pipewire pipewire-pulse wireplumber` on every
monitoradded (#100). Chromium's audio service does not reconnect, so it
enumerates nothing — Meet's "no mic or speakers".

Ordering across four boots is decisive, and Bernardo asked for exactly
this check: restart-after-Chromium in both broken sessions (boot -3
14:06:49 → 14:07:52; boot -2 11:10:34 → 11:31:22), restart-BEFORE in
today's working one (08:04:19 vs 08:03:55, a 24s miss). Hibernate is not
the cause, only how a Chromium lives long enough to meet a plug event.

Retiers off [human] — no repro needed, and it predicts a live one:
dock/undock now should break Meet in this session. Fix is ours, not
Chromium's: the graph restart is inherited from "the old working flake",
never shown necessary, and repair_dock_cards already handles the parked-
card case without it.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-16 09:32:28 +01:00

584 lines
33 KiB
Markdown
Raw Blame History

This file contains invisible Unicode characters
This file contains invisible Unicode characters that are indistinguishable to humans but may be processed differently by a computer. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Backlog — the prioritized task queue
**This is the only executable work list for agents.** Product themes and
v1.0 intent live in [`docs/VISION.md`](../docs/VISION.md); design history
in [`docs/ROADMAP.md`](../docs/ROADMAP.md); map in
[`docs/README.md`](../docs/README.md) and [`agent/README.md`](README.md).
**Rules:**
- Agents take the topmost actionable item (see LOOP.md). Finished items
are **deleted** here — the journal + git log are the record; durable
design notes get a ✓-entry in docs/ROADMAP.md (and/or a note in VISION)
if worth keeping.
- Item numbers are **stable IDs** — never renumbered or reused. A gap in
the sequence means shipped (or dropped) work; new items take the next
free number regardless of tier.
- Tags: `[blocked:hw]` needs real hardware (see HARDWARE-QUEUE.md) ·
`[human]` needs Bernardo · `[stuck]` two failed attempts, needs help ·
`[big]` must be split before starting.
- Agents may append to **PROPOSED** and **Decisions** freely (include
`VISION § …` or `ROADMAP § …` when relevant); only Bernardo moves items
*out* of PROPOSED into the tiers.
---
## NOW
### Live ISO / install hardware findings — Acer Aspire M5-481T + Dell XPS 9350
Bernardo, real installs 2026-07-1314 (photos of the install end screens and
post-boot sessions). Preserve separation: the installer bake failure and the
flake pin are different root causes even when they show up on the same machine.
(Terminal / Ghostty-on-Acer → shipped as Kitty-only, #95.)
### 127. `[blocked:hw]` Docked idle: black screens, no wake, undock does not restore panel
Bernardo 2026-07-15 (AMD dev box, docked): left the laptop docked; after a
few minutes of idle it appeared to "sleep". Then:
1. External keyboard/mouse did **not** wake any display.
2. Unplugging the external monitor did **not** turn the laptop panel back
on (dock mode has eDP disabled — undock *should* re-enable it).
3. The machine was still alive: Caps Lock LED toggled on the keyboard.
This also killed a long agent session (looked like a "crash" from the
outside). Separate from awake undock recovery (HARDWARE-QUEUE round 8):
the failure is **idle while docked**, then no path back to a usable panel.
**Likely shape (code-side, unconfirmed on the incident):** on AC, hypridle
does **not** call `systemctl suspend` (`onAc || suspend` — only battery
suspends at 15 min). Idle path is lock @5 min + `dpms off` @10 min
(`modules/home/idle.nix`). Caps Lock working fits **DPMS/lock blackout**
better than deep S3. Dock mode has already disabled eDP; recovery then
needs either (a) input → hypridle `on-resume` / `dpms on` on the external,
or (b) undock → `nomarchy-display-transition undock` re-enabling the
panel. Both failed. `after_sleep_cmd` only does `hyprctl dispatch dpms on`
and never re-enables a disabled internal.
**Progress 2026-07-15:**
- Recovery tooling: undock ends with `dpms on`; `nomarchy-display-wake` /
`nomarchy-display-dump` (SUPER+SHIFT+D; TTY if it works).
- **Mitigation (shipped):** do **not** DPMS-off when docked — if any
laptop internal (eDP/LVDS/DSI) exists in `monitors all` but is not
enabled, the 10 min listener skips blanking. Lock at 5 min still runs;
undocked laptops still DPMS-off as before. Bernardo: Ctrl+Alt+F3 did
not yield a usable TTY on the original brick, so prevention > dump path.
- **Revisit later (LATER):** intentional DPMS-when-docked once wake is
proven trustworthy (monitor power-save without a brick). Tracked as
**#135** (state key + the surface to delete this by) and **#136** (the
gate is clamshell, not "docked" — lid-open docked still blanks). Both
are downstream of the hardware test below.
- **Caveat the test must settle:** the mitigation is a cure for an
**unconfirmed cause**. "Blanking the only live output caused the brick"
is inference from one incident, shipped at V1. If DPMS-off is not the
trigger, we are paying a permanent behaviour cost (docked externals never
sleep) and the brick is still live — so confirming the cause outranks
#135/#136.
**Pass (current mitigation):** docked on AC, idle past 10+ min → external
stays lit under hyprlock (journal: `dpms-off skipped: docked`); undocked
still blanks at 10 min and wakes on input. Full original pass (wake after
DPMS-while-docked) deferred to the LATER revisit.
**Fallback if the cause resists (Bernardo 2026-07-16):** diagnosing stays the
preference — DPMS when docked is the outcome we want. But the mitigation as
shipped holds a near-static lock screen on the externals indefinitely, so if
the cause cannot be pinned down, a **screensaver while docked** is the
consolation prize: it protects the panels (burn-in — the same cost #135
names) without re-entering the blanking path that bricked the session. Not a
new item: it is the "if diagnosis fails" branch of this one, and #135's state
key is where it would land.
## NEXT
### 135. Give the #127 DPMS mitigation a state key (and a way back)
The shipped mitigation is hardcoded policy: docked (clamshell) externals now
**never** blank, and hyprlock has no timeout of its own, so they stay lit
indefinitely. Two real costs land on the user — a couple of 27" panels at
~60100W all night, and **OLED burn-in** from a near-static lock screen held
for hours. Today there is no way to opt back in, which also breaks the
in-flake-state rule that settings are menu-writable.
Add `settings.idle.dpmsWhenDocked` (default **false** = today's safe
behaviour), read by `modules/home/idle.nix`'s `dpmsOff` gate, with a
Preferences row. The real point is not the option: it gives the mitigation a
**named surface to delete** once wake is proven, instead of an unowned LATER
note that quietly becomes the product. Whoever removes the mitigation removes
the key in the same commit.
**Order:** after the #127 hardware test. If DPMS-off turns out not to be the
trigger, the mitigation is reverted wholesale and this key never exists —
building it first risks shipping an option for a policy we withdraw.
### 136. `[human]` #127's gate is clamshell, not "docked"
The code skips DPMS-off when a laptop internal is present in `monitors all`
but absent from the enabled list — i.e. **clamshell**. Commit 060bf52, the
#127 entry and the journal all say "docked", which is broader: **docked with
the lid open** (internal + externals all enabled) falls straight through to
`hyprctl dispatch dpms off` and still blanks. So either
- the narrow scope is deliberate (the brick needs the internal disabled, and
lid-open docked is genuinely safe) → fix the wording in 060bf52's
descendants so the docs stop claiming more than the code does; or
- it is an accidental gap → the mitigation misses half the docked cases and
the brick can still fire with the lid open.
Bernardo's call, and the #127 hardware test is what settles it: if the brick
reproduces lid-open, this is a gap. Cheap either way — one `jq` predicate or
one docs sweep. Filed because a doc that overclaims its own fix is how the
next session mis-diagnoses the recurrence.
### 129. Permanent guard for the GTK button-colour trap (#98 follow-up)
The #98 regression (destructive-action at 1.03:1, invisible) was caught only
by rendering a dialog — every scripted check was green, because the bug lives
in adw-gtk3's `mix(@destructive_color, alpha(currentColor,…))` background, not
in our CSS. `tools/theme-shot.nix` is deliberately not a CI gate (full-desktop
VM), so add the cheap static half: assert that for every theme, any rule that
pins a button label colour also pins that button's `background-color` — i.e.
no pinned label may sit on a `currentColor`-derived background. Guards the
class, not the instance. The scratch render harness that found it is worth
promoting to `tools/dialog-shot.nix` as a maintainer tool alongside theme-shot.
### 130. GTK accent buttons sit at ~2.7:1 in light themes
Fallout noted while fixing #98: in light palettes the suggested/destructive
labels are base00 (cream) on saturated accent/bad — summer-day measures
**2.72:1** (suggested) and **2.74:1** (destructive), under AA 4.5 and even
under the 3:1 large-text floor. Not a regression: upstream adw-gtk3 does the
same (`color: white` on the accent), so we currently match GNOME. Decide
whether Nomarchy holds itself higher — darker label, or a darker accent mix
for light palettes — and note `tools/check-theme-contrast.py` covers palette
pairings only, never GTK widget surfaces. `[human]` for the aesthetic call.
### 131. Recovery cost labels truncate on 1366-wide panels
#111 gave Recovery scope-first labels that carry the cost inline — "Desktop
generation — themes & home config (instant, reversible)" (462px in Inter 11)
and "System boot generation — older NixOS (reboot, then pick in boot menu)"
(509px). The picker is `width: 40%`, so text room is ~668px at 1920 (fits)
but only ~446px at 1366 — **both long labels ellipsize, and the cost hint is
exactly the part that gets cut**. The Acer Aspire M5-481T (1366×768) is an
active QA machine, so this is live, not theoretical. Options: shorten to
"Desktop generation (instant)" and put cost in a rofi `-mesg` line, widen the
picker for narrow screens, or let rofi wrap. Measured, not observed — the VM
renders menu geometry unfaithfully (see #132); confirm on the Acer.
### 132. VM cannot judge rofi menu geometry
The guest renders the picker with icons at roughly a quarter of `ui.iconSize`
(44 → ~10px) and rows short enough that a 6-row root clips its last entry
behind a scrollbar, though `lines = 8` and `dynamic = true` should fit it.
Root is unchanged since before #105, so this is the guest's font/icon
environment, not a regression — but it means `theme-shot.nix`/`menu-shot`
screenshots are trustworthy for **menu content** and not for spacing, icon
size, or truncation. Either fix the guest (fontconfig + icon cache in the
test node) or document the gap in docs/TESTING.md §5 so the next agent does
not chase it — right now nothing warns them.
### 133. No check pins the #107 legacy-name compat shim
`lib.mkFlake`, doctor and lifecycle all still accept `theme-state.json`, and
`nomarchy-state-sync` migrates it on write — but nothing in `checks.*`
evaluates a checkout that has *only* the legacy name, so the shim can rot
silently while every check stays green. It is load-bearing until every
existing machine has taken one menu write. Add a cheap eval check: a fixture
dir with `theme-state.json` only → `mkFlake` resolves it; plus the write-side
migration (state.json created + git-staged, legacy `git rm`'d). Proven by
hand 2026-07-15 (drvPaths identical across the rename) — this just keeps it
proven. Delete together with the shim + the `nomarchy-theme-sync` alias.
### 137. Plymouth splash draws on head 0's geometry, so it is off-centre when docked
Bernardo 2026-07-16 (T14s, docked): the boot **and** shutdown logo is not
centred on the external monitor. Cause is in the theme script, not the
hardware: `modules/nixos/plymouth/nomarchy.script` calls `Window.GetWidth()`
/ `Window.GetHeight()` with **no head index**, which returns head 0, while
sprite coordinates live in a device-wide space spanning every head. So every
element is placed at "head 0's centre" measured from the union origin — right
on a single monitor, wrong the moment a second one exists, and wrong in a
different direction depending on which head is 0 and how they are arranged.
It is the whole splash, not just the logo: the same unindexed call places the
logo (`:9`, `:1516` — note the 15%-of-height *scale* also tracks the wrong
head), the password entry (`:129`), the lock, the bullets, and the progress
box/bar (`:210`, `:219`). Fix shape is the standard Plymouth one: loop
`Window.GetX(i)` / `GetY(i)` / `GetWidth(i)` / `GetHeight(i)` over the heads
and give each its own sprite set, so each monitor gets a centred, correctly
scaled splash — the LUKS prompt especially, since a passphrase box half off
the panel is the failure that matters. Pass = docked boot + shutdown centred
on both heads; undocked unchanged.
### 138. Our dock-audio reprobe restarts PipeWire under running apps (Chromium loses its devices)
Bernardo 2026-07-16, refined by his own hypothesis (pipewire restarts →
Chromium never reconnects) — which the journal confirms, and names **our own
code as the thing doing the restarting**. Symptom: Google Meet reports no mic
or speakers in native Chromium; Zoom is unaffected; it did **not** reproduce on
a fresh boot, only after days of uptime + a hibernate.
**Mechanism.** `modules/home/dock-audio.nix:153` — on every `monitoradded`, the
`reprobe` path runs `systemctl --user restart pipewire.service
pipewire-pulse.service wireplumber.service`. A full graph restart, on every
physical monitor plug (#100). Chromium's audio service does not re-establish
its PulseAudio connection when the server disappears under it, so it keeps
running with a dead one and enumerates nothing — hence Meet's wording, hence
"restarting Chromium fixes it". Zoom survives because it is a Flatpak that
reconnects, and because it is generally started fresh into an already-restarted
graph. **Hibernate is not the cause** — it is only how a Chromium survives long
enough to still be running when a plug event arrives; resume-while-docked then
generates the `monitoradded` that fires the reprobe.
**Evidence — ordering is the whole tell** (`Started app-org.chromium.Chromium-*.scope`
vs `reprobe-start`, four boots, no exceptions):
| Session | Chromium | Graph restart | Order | Meet |
|---|---|---|---|---|
| boot -3 (Jul 14→15) | 14:06:49 | 14:07:52 (+2 more; then 10:52 next day) | **after** | broken |
| boot -2 (Jul 15 daytime) | 11:10:34 | 11:31:22 | **after**, +21 min | broken |
| boot 0 (Jul 16, fresh) | 08:04:19 | 08:03:55 | **before**, 24 s | works |
Today is the negative control, and it is luck, not a fix: the restart missed
Chromium by 24 seconds. **Prediction that confirms this outright — dock/undock
now, then reload Meet: it should break in the current session**, with no reboot
and no hibernate. That is the cheap V3 for the diagnosis, and it is also the
regression test.
**What to fix.** Not Chromium — the graph restart is a sledgehammer we
inherited: the header comment justifies it only as "the recovery used by the
old working flake", i.e. it was never shown to be *needed*, while its collateral
(every long-lived audio client dropped) was never priced. It is also not rare —
boot -3 took three restarts in 82 seconds (14:07:52 / 14:08:58 / 14:09:09;
the `mkdir` debounce only collapses events within one run, not a plug storm),
and each one drops audio for every app on the machine. Shape: **try selection
first, restart only if nothing routable appears** — `select_first_dock` already
falls back to `repair_dock_cards`, which fixes the parked-card case
(`set-card-profile`) *without* touching the graph, so the restart may have no
remaining job at all. If some HDMI codec really does need it (the other claim
in that comment), prove it on this hardware and scope the restart to that case.
Pass = plug/unplug the dock with Chromium open → output still follows to the
monitor **and** Meet still sees its devices; `journalctl -t nomarchy-dock-audio`
shows no graph restart on the ordinary path.
### 139. Terminal floats open full-screen — % size rules are ignored, and Kitty remembers
Bernardo 2026-07-16: calcurse and nomarchy-doctor float but fill the screen —
the reduced size we had under Ghostty is gone. Reproduced and root-caused live
on the T14s (Hyprland 0.55.4, single 2560×1440 head), and it is **two** faults
stacked, which is why it looks like the float rules are dead when they are not:
1. **Percentage sizes silently no-op.** `com.nomarchy.calendar` came up
`floating=true` but `size=[2526,1360]` — the full work area, not the
`size 60% 65%` in `modules/home/hyprland.nix:1064`. A runtime rule with
**absolute** pixels (`size 1400 900`) applied perfectly and even centred
itself, so `size` works and the matcher works: it is the `%` form that this
Hyprland drops. No `hyprctl configerrors` — it fails silently, which is
how it survived the rule-syntax rewrite.
2. **Kitty remembers its last window size.** With the rule not applying,
nothing overrides Kitty's own request, and Kitty defaults to
`remember_window_size yes``~/.cache/kitty/main.json` literally holds
`{"window-size": [...]}` and is replayed into every new OS window. So the
float inherits whatever the last Kitty was, which after any tiled window is
~full-screen. **This is the Ghostty regression** (#95, Kitty-only): Ghostty
kept no such memory, so fault 1 was invisible — fault 2 exposed it.
Both need fixing, and fixing either alone is a trap: absolute pixels alone
stop being right on the 1366-wide Acer (#131's machine) or any scaled head,
and pinning Kitty alone leaves the `%` rules rotting for the next window we
float. Likely shape: `-o remember_window_size=no` (plus
`initial_window_width/height`) in the `nomarchy-calendar` / doctor launchers
so Kitty stops fighting us, and either the working syntax for `%` or a
documented note that this Hyprland wants absolute px. Affects every rule using
`%` — today calendar (`:1064`) and doctor (`:1070`). Pass = both sheets open
centred at their intended fraction on the dev box **and** on the Acer.
_(My probes wrote `{"window-size": [1400, 900]}` into your Kitty cache — it is
a self-updating cache, harmless, and any resize overwrites it.)_
### 141. Waybar's update module is dead in the whole-swap themes (`$TERMINAL` is unset)
Bernardo 2026-07-16: clicking the Waybar updates module does nothing.
Confirmed live and it is a **Waybar parity break** (CONVENTIONS § whole-swaps),
not an updates bug. The generated module bakes the terminal in —
`"${config.nomarchy.terminal} -e nomarchy-updates upgrade"`
(`modules/home/waybar.nix:356`) — but the four whole-swap themes hand-write
`"sh -c '$TERMINAL -e nomarchy-updates upgrade'"`, and **Waybar's environment
has no `TERMINAL`**: its 55 vars contain zero matches (pid 1688, `/proc/*/environ`).
`TERMINAL` is a `home.sessionVariables` entry (`modules/home/default.nix:104`),
so login shells get it and the Hyprland-spawned bar does not. The click expands
to `sh -c ' -e nomarchy-updates upgrade'` → nothing, no error, no toast. You
are on **boreal**, a whole-swap, which is why you see it.
Affected, one line each: `themes/{boreal,summer-day,summer-night,executive-slate}/waybar.jsonc`
(`custom/updates.on-click` — grep says this is the *only* `$TERMINAL` use left
in any whole-swap, so the blast radius is exactly these 4 lines). Fix = drop
the env dependency the way calendar/doctor already do — name `kitty` outright,
Kitty being sole supported since #95 — or, better and one place instead of
five, give `nomarchy-updates` a subcommand that opens its own terminal so
every caller stops needing `$TERMINAL`. Worth a check against the class of
bug, not the instance: nothing today asserts that a whole-swap's on-click
survives Waybar's actual environment, so the next hand-written jsonc can
reintroduce it silently.
### 120. A netinstall ISO, next to the fat offline one
Bernardo 2026-07-14, after seeing the measured size: **keep the current ISO
exactly as it is** — the guaranteed offline install is the feature it buys —
and ship a **much lighter netinstall variant alongside it**. Two products, one
distro: "works on a plane" and "8 GiB is absurd to download" are both true, and
a second target settles them without compromising either.
**Measured facts (2026-07-14), so this starts from numbers, not vibes.**
*(These stand as measured: #121 would have cut ~0.67 GiB of duplicate chromium
from them, but it was **reverted** — decided against, ROADMAP § one chromium,
not two. If a netinstall ships, revisit it: the duplicate is worth ~195 MiB of
**download**, which is this item's whole currency, even though it is worth
almost nothing on disk or on the ISO.)*
> **Read this before using the numbers below.** They are **closure arithmetic**,
> and #121 proved the hard way that closure size is neither disk size nor image
> size: removing a 687 MiB path shrank the ISO by **8 KiB**, because
> **mksquashfs dedupes duplicate files** and **`auto-optimise-store` hardlinks**
> them on disk. So a change that looks like it sheds gigabytes of closure can
> shed nothing off the actual image. **Measure the artifact — build the ISO and
> `stat` it.** The corollary cuts the other way and is the good news for this
> item: what dedupe cannot help is the **wire**, so a netinstall's download is
> the one figure closure/NAR size predicts honestly (`nix path-info --store
> https://cache.nixos.org --json` gives the real `downloadSize`).
- Current ISO **8.078 GiB** compressed; **18.03 GiB** of store uncompressed
(`zstd -19`, 2.23:1 — compression is already near-max, not the lever).
- The offline pin (`system.extraDependencies`, 60 roots: a representative
installed system + the template HM closure + all flake inputs) is **4.02 GiB
uncompressed of that — only ~22%**. Dropping it entirely still leaves a
**~13.3 GiB** desktop → roughly **6 GiB** compressed at the same ratio.
**So "no pin" alone is NOT the lighter ISO** — this is the trap to avoid.
- The desktop's own top weights: libreoffice 1457 MiB, initrd 1369,
linux-firmware 770, chromium 1391 (two builds — #121, left in), llvm-lib 540,
bibata-cursors 322, mesa 264, mbrola-voices 259, nerd-fonts ~420 combined.
Note what that list implies: no single lever gets a desktop ISO under ~4 GB —
which is the case for (b) below.
**So the real decision is what a netinstall ISO IS**, and it should be settled
first (`[human]`): (a) the full try-before-install desktop minus the pin
(~6.3 GiB — barely lighter, probably not worth a second target); (b) a **TUI
installer only, no desktop** (~1 GiB, the actual "netinstall" in the Debian
sense) which drops "try before install" from that medium — the fat ISO still
offers it; (c) a middle desktop (no libreoffice/chromium — but note #103 just
put those there deliberately, and a *demo* desktop that can't browse is the
bug #103 fixed).
**The gotcha that decides feasibility:** without the pin, a netinstall target
fetches from `cache.nixos.org` for stock nixpkgs paths — but **Nomarchy's own
derivations are in no binary cache**, so they would build *from source on the
user's machine* during install. That is the same failure `tools/vm/gap-analysis.py`
exists to diagnose (and #113 is a live instance of). So this item probably
depends on a public binary cache (cachix) for the flake's own outputs, or it
trades an 8 GiB download for a 40-minute install. Establish that before
building the target.
Pass = a second, documented ISO target that is *substantially* smaller (state
the measured number, both ISOs built from one tree), installs successfully with
a network in a QEMU run, says clearly at boot that it needs one, and leaves the
offline ISO's behaviour untouched (`checks.*` for the offline path stay green).
## LATER
- **DPMS when docked (revisit #127):** today we **skip** display blanking
while any laptop internal is present-but-disabled (dock/clamshell), so
idle only locks. That avoids blanking the sole external and the brick
where wake/TTY failed (2026-07-15). Revisit when we want monitor
power-save on dock again: re-enable DPMS-off for that mode only if
`nomarchy-display-wake` (or better) reliably restores the external and
undock-while-black restores eDP on the AMD dock setup — without
depending on SSH. Until then, leave the skip in `modules/home/idle.nix`.
- **Wallpapers artifact split** (ROADMAP § Faster switches — decided,
deferred): pinned `Nomarchy-wallpapers` input so a state write stops
re-copying 86 MB. Follow-on: pre-built theme variants if switches are
still slow after.
- **Installer round 2** (ROADMAP § Installer): multi-disk BTRFS RAID,
impermanence, BIOS/legacy boot.
- **Boot-from-snapshot**: a systemd-boot equivalent of grub-btrfs.
- **MIPI/IPU software-ISP camera** support (no-UVC machines).
- **NixOS release bump → v2** `[human]`: deliberate, hand-edited, never
automated; the previous attempt was discarded (2026-06-22) over a
Hyprland OOM blocker — see MEMORY.md before retrying.
## FUTURE (decided deferred — not the agent queue head)
Work we **intend** someday but explicitly **not** NEXT. Agents do not
pick these unless Bernardo promotes one into NEXT/NOW.
### 20. KVM runner → VM suite in CI `[human]`
**Status (2026-07-10):** keep **eval-only** CI on the current Gitea
stack (act_runner in docker-compose on the 4c/4GB IONOS VPS). Nested
KVM + RAM headroom on that host are a poor fit next to Gitea; full
`checks.*` VMs stay local / promotion-time until a **separate**
KVM-capable machine exists.
**When ready:** register a second runner (host-mode nix + `/dev/kvm`,
label `nix-kvm` — not the existing docker eval runner), then uncomment
the `vm-checks` job in `.gitea/workflows/check.yml` (`runs-on: nix-kvm`,
`nix flake check` + toplevel/HM builds). Do not enable the job until
that label is online (Gitea queues forever otherwise).
### Formatter — adopt later `[human]`
**Intent:** add a Nix formatter (likely `nixfmt-rfc-style`) in a dedicated
pass: reformat the tree once, document in CONVENTIONS, optional CI
check. **Not** the queue head — no drive-by reformats until that pass.
## PROPOSED (agent suggestions — await human triage)
*Agents: append here with a one-paragraph pitch (what/why/cost). Do not
implement. Bernardo moves accepted items into a tier.*
*Open work only. Shipped exam/AC items (#47#63, #14, #52 theme
high-ROI, etc.) live in the journal + ROADMAP — not here.*
### Product / day-2
### 140. Ask Claude should hand off to a web chat, not open a coding agent
Bernardo 2026-07-16: "opening claude-code is too disruptive" — he would prefer
the query go to a **web interface** (Gemini or Claude) instead. Today
SUPER+CTRL+A → `nomarchy-menu ask` (`modules/home/rofi.nix:899905`) takes the
free-text line and does `${cfg.terminal} -e npx --yes @anthropic-ai/claude-code@latest "$q"`:
a terminal window, an npm fetch on first run, and a REPL that stays open — a
coding agent answering "what's the syntax for…". The prompt surface is right;
the destination is wrong for the question the surface invites.
Cheap version: URL-encode `$q` and `xdg-open` a prefill URL
(`https://claude.ai/new?q=…`, `https://gemini.google.com/app?q=…`), landing in
the default browser (Chromium, per Decisions). That is a handful of lines and
deletes the `npx`/network/REPL cost — plus `pkgs.nodejs`, carried in
`home.packages` (`:1685`) **only** for this feature, so the closure shrinks
with it. To settle before building:
- **Which destination, and is it a choice?** Bernardo named Claude *or*
Gemini. A `settings.ask.provider` key keeps it in-flake and menu-writable
(the convention) and stops us hardcoding a vendor into a vendor-neutral
distro; a fixed default is cheaper. Prefill-URL support is per-vendor and
can be changed by them unilaterally — worth verifying each still prefills,
since a silently-ignored `?q=` degrades to "opens a blank chat", which is
survivable but should be a known cost, not a surprise.
- **Keep the CLI at all?** A second row (Ask ▸ web / Ask ▸ agent) preserves
today's behaviour for whoever wants it; deleting it is the simpler surface
and the one the complaint argues for. Bernardo's call.
- Renaming: "Ask Claude" is the keybind desc (`keybinds.nix:57`) and shows in
the generated cheatsheet — if the destination becomes configurable, the
label should stop naming one vendor.
### 134. `unstable.<pkg>` in the downstream — a newer app without a second flake
Bernardo 2026-07-15, wanting a newer LM Studio (pinned 0.4.15-2 vs unstable
0.4.19-2) and disliking the only path that works today. There is **no seam**:
`mkFlake` takes `src`/`username`/`hardwareProfile`/`system` only, and
`homeConfigurations` is built from mkFlake's own `pkgs` (`inherit pkgs`), so a
`nixpkgs.overlays` in system.nix cannot reach a home package. The user is left
hand-pinning a `fetchTarball` rev+sha256 in home.nix (verified working) — a
second pin with no lock, which is exactly the Nix expertise the distro exists
to spare them. Sketch: `home.packages = [ unstable.lmstudio ];`, and maybe
`stable.` as an explicit synonym for today's bare `pkgs`.
The shape that keeps the promises: Nomarchy carries the `nixpkgs-unstable`
input **itself** and exposes it through `overlays.default` as an `unstable`
attrset, so downstream stays **one input**, stays locked, and `nomarchy-pull`
still moves it. Cost/questions to settle before building:
- Lock churn: every lock bump now moves two channels, and the full-checklist
rule (VERIFICATION §5) applies to both.
- Duplication: measured — pinned vs unstable lmstudio build closures share
only 690 of ~3.74k paths (separate stdenv bootstrap). Opt-in per package,
but a user who reaches for it pays a second glibc/stack in the closure.
- "Tested together upstream" weakens by construction: nothing validates
pinned-Nomarchy + unstable-app. Needs an explicit, documented "you own this
combination" line, not silence.
- Does `unstable` belong in the eval by default, or only when referenced?
(An unused input still locks + fetches.)
- Scope: home packages only, or system.nix too?
Alternative if that is too much surface: give `mkFlake` an `overlays ? []`
param (one-line seam, user still owns their own input/pin — cheaper for us,
still expert-only for them). Related: #133.
### 114. Greeter ignores per-device keyboard layouts
Found by Bernardo 2026-07-14: logging out while docked lands on tuigreet,
where his external keyboard (remembered as `us` via
`settings.keyboard.devices`) types the session layout `gb` — so the password
prompt fights him. Not a regression and not docking-related: per-device
layouts are applied with `hyprctl keyword device[<name>]:kb_layout`, which
only exists inside a running Hyprland session, while tuigreet draws on a
kernel VT whose single keymap comes from `console.useXkbConfig`
`services.xserver.xkb.layout`. A VT structurally cannot do per-device
layouts, so this is a design gap, not a bug to patch. Options, cheapest
first: (a) document it and stop there; (b) a greeter layout-cycle key
(tuigreet has no such feature — would need the keymap swapped under it);
(c) host tuigreet inside a small Wayland compositor (cage/labwc), which
*does* get per-device XKB and would let the greeter honour the same
in-flake state the session uses — real work, and it changes the greeter's
whole rendering path (`modules/nixos/greeter.nix` themes tuigreet through
the 16 ANSI console slots, so (c) is not a drop-in). Worth deciding whether
the greeter is meant to be layout-aware at all before costing it.
- **NVIDIA first-class options** — **deferred past v1** (Bernardo
2026-07-10). Keep #59 commented install guidance; no
`nomarchy.hardware.nvidia.*` until a hybrid maintainer + queue.
- **Post-install hardware hints** (`VISION § B`) — After the general
“you're set” card (#81), optionally fire **one** additional
self-gated notify when the machine actually has the hardware:
(a) `fwupdmgr` on PATH → “System Firmware to check LVFS updates”;
(b) `fprintd-list` on PATH → “System Fingerprint to enroll”.
One-shot markers in `settings.*` (same in-checkout discipline as
`firstBootShown`); never a permanent MOTD nag. Cost: small — extend
`nomarchy-first-boot` or a sibling oneshot + `checks.first-boot`
fixture. Control-center / MOTD already mention these; the gap is the
silent first *graphical* session for people who never open those.
_(#80#83 + #85#88 shipped 2026-07-11. Theme A day-2 + neon-glass finish
shipped — VISION ✓. Dock/hibernate V3 → HARDWARE-QUEUE. Parallel
fingerprint-or-password shipped 2026-07-12 (Bernardo promoted it live;
`fingerprint.parallel`, pam-fprint-grosshack) — reader V3 →
HARDWARE-QUEUE.)_
### v1.0 pointer
See **VISION**. Open PROPOSED: post-install hardware hints; NVIDIA
deferred past v1; IR portal (b)/(c) need T14s (HARDWARE-QUEUE § T14s).
Standing calls: browser = Chromium; power = PPD.
## Decisions `[human]`
Open calls only Bernardo can make; agents add options/evidence but never
decide. **Resolved** entries stay for history; agents treat them as closed.
### Resolved (2026-07-10)
- **Docs site vs Markdown-in-repo** — **markdown in-repo for now**
(`docs/`, README). A rendered docs site is FUTURE if wanted.
- **Default browser** — **ship Chromium** in
`templates/downstream/home.nix`; mime → `chromium-browser.desktop`.
Opt out: delete the line / override mime.
- **Default power backend** — **keep PPD** (`nomarchy.system.power.backend`
default). TLP remains the one-line opt-in. Rationale: stability + live
profile API for menu/Waybar; Omarchys TLP experiment reverted.
### Resolved (2026-07-10, more)
- **Formatter adoption** — **yes, but not now.** Tracked as FUTURE
(below). Nix-source style only (`nixfmt-rfc-style` or similar); one
bulk reformat + CI/check when promoted. Until then: hand-aligned
style per CONVENTIONS.
- **Hibernation** — **want by default** (product intent). Needs a
disk-backed swap (file or partition) sized for resume; not zram alone.
**Shipped as #76**; V3 power-cycle PASSED on TuringMachine 2026-07-12
(ROADMAP § Hibernation + zram by default).
### Resolved (2026-07-10, #76 design)
- **Swap sizing** — **exactly RAM** (installer default, unchanged). Hibernate
image ≤ RAM; zram takes day-to-day paging. **`swapSize=0`** stays no-swap.
- **Migration** — **docs runbook** (`docs/MIGRATION.md`), not a tool.
- **No-swap Hibernate** — keep the menu row; **notify on failure**.