fix(nixos): smartd self-gates on the hardware — the doctor was right, smartd wasn't
Some checks failed
Check / eval (push) Has been cancelled
Some checks failed
Check / eval (push) Has been cancelled
BACKLOG #118. Bernardo booted the live ISO and the Waybar health icon was red, reporting smartd. Everything downstream turned out to be working correctly, which is the part worth recording: smartd's config is DEVICESCAN, and where no drive answers SMART it exits 17 ("Unable to monitor any SMART enabled devices"), systemd marks the unit failed, nomarchy-doctor faithfully reports a failed system unit, and Waybar paints @bad. The doctor was telling the truth. smartd was the bug. Scope was never live-only, which is why this sat in NOW rather than as a live nit: services.smartd.enable mkDefaults true on every machine and QEMU virtio exposes no SMART, so every VM install has been booting to a health warning about a daemon with nothing to do — and every V2 run had been showing it as noise. Fixed with the distro's own self-gate convention: an ExecCondition running `smartctl --scan`, which prints nothing exactly when smartd would find nothing. A failed condition leaves the unit inactive rather than failed. Deliberately NOT SuccessExitStatus = 17: that would also swallow exit 17 from a machine that does have drives, which is the entire reason the daemon ships. V2. checks.smartd-gate boots the REAL distro module rather than a restatement of it (its nixpkgs.config needs mkForce to yield to the test's pkgs) and asserts both halves, because they pull in opposite directions: a gate that never skips leaves the red icon, and a gate that always skips silently disables drive-health monitoring on real hardware — the failure nobody notices until a disk dies quietly. So the no-SMART node must go ActiveState=inactive, unfailed, and absent from `systemctl --failed` (what the doctor actually reads); and the gate's logic is driven against a scan that DOES find a device, since QEMU cannot answer SMART honestly and pretending otherwise would test nothing. The check was proved to fail by unwiring the condition. flake check, doctor, hardware-toggles, live-baseline-apps, option-docs and state-bridges pass. V3 pending: that smartd still RUNS where drives have SMART (dev box, real NVMe). Queued with an explicit fail-condition — if it skips there, revert the gate rather than tune it. Also swept: #120's size table said the duplicate chromium was gone; #121 was reverted, so it is back and the table says so. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
@@ -429,6 +429,18 @@ the **T14s** (webcam case).
|
||||
double-clicking a text file is #119's problem, so don't file it twice.
|
||||
|
||||
## AMD dev box only
|
||||
- [ ] **#118 smartd still runs where drives DO have SMART** (this commit) — the
|
||||
half a VM cannot answer: QEMU exposes no SMART, so `checks.smartd-gate`
|
||||
proves the skip but has to drive the *with-device* path through a stub.
|
||||
On the dev box (real NVMe), after `nomarchy-rebuild` + reboot:
|
||||
`systemctl status smartd` is **active/running**, and
|
||||
`systemctl show -p ExecCondition --value smartd` names
|
||||
`smartd-any-smart-device`. **Pass** = smartd is running, exactly as
|
||||
before this commit — i.e. the gate skips nothing on real hardware.
|
||||
**Fail** = inactive/skipped, which would mean the gate is silently
|
||||
disabling drive-health monitoring: revert it, don't tune it.
|
||||
Cheap bonus while you are there: `nomarchy-doctor` reports no failed
|
||||
units (the red-icon symptom that started #118).
|
||||
- [ ] **AMD runtime bits** — VA-API (`vainfo` → radeonsi), amd-pstate EPP
|
||||
active and PPD switching governors; opt-ins: ROCm (`rocminfo`, a GPU
|
||||
PyTorch/Ollama smoke) and the XDNA NPU driver loading.
|
||||
|
||||
Reference in New Issue
Block a user