fix(nixos): smartd self-gates on the hardware — the doctor was right, smartd wasn't
Some checks failed
Check / eval (push) Has been cancelled
Some checks failed
Check / eval (push) Has been cancelled
BACKLOG #118. Bernardo booted the live ISO and the Waybar health icon was red, reporting smartd. Everything downstream turned out to be working correctly, which is the part worth recording: smartd's config is DEVICESCAN, and where no drive answers SMART it exits 17 ("Unable to monitor any SMART enabled devices"), systemd marks the unit failed, nomarchy-doctor faithfully reports a failed system unit, and Waybar paints @bad. The doctor was telling the truth. smartd was the bug. Scope was never live-only, which is why this sat in NOW rather than as a live nit: services.smartd.enable mkDefaults true on every machine and QEMU virtio exposes no SMART, so every VM install has been booting to a health warning about a daemon with nothing to do — and every V2 run had been showing it as noise. Fixed with the distro's own self-gate convention: an ExecCondition running `smartctl --scan`, which prints nothing exactly when smartd would find nothing. A failed condition leaves the unit inactive rather than failed. Deliberately NOT SuccessExitStatus = 17: that would also swallow exit 17 from a machine that does have drives, which is the entire reason the daemon ships. V2. checks.smartd-gate boots the REAL distro module rather than a restatement of it (its nixpkgs.config needs mkForce to yield to the test's pkgs) and asserts both halves, because they pull in opposite directions: a gate that never skips leaves the red icon, and a gate that always skips silently disables drive-health monitoring on real hardware — the failure nobody notices until a disk dies quietly. So the no-SMART node must go ActiveState=inactive, unfailed, and absent from `systemctl --failed` (what the doctor actually reads); and the gate's logic is driven against a scan that DOES find a device, since QEMU cannot answer SMART honestly and pretending otherwise would test nothing. The check was proved to fail by unwiring the condition. flake check, doctor, hardware-toggles, live-baseline-apps, option-docs and state-bridges pass. V3 pending: that smartd still RUNS where drives have SMART (dev box, real NVMe). Queued with an explicit fail-condition — if it skips there, revert the gate rather than tune it. Also swept: #120's size table said the duplicate chromium was gone; #121 was reverted, so it is back and the table says so. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
@@ -404,6 +404,30 @@ Design/decision records and a running log of shipped work (items marked
|
||||
decision rather than a drive-by. Proved to fail: dropping snapshot makes it
|
||||
name the missing entry. **V3 pending** — that the apps *launch* needs real
|
||||
hardware (HARDWARE-QUEUE, Acer M5-481T).
|
||||
- ✓ **smartd self-gates on the hardware (#118):** Bernardo booted the live ISO
|
||||
and the Waybar health icon was red. Everything downstream was working
|
||||
correctly, which is what made it worth writing down: smartd's config is
|
||||
DEVICESCAN, and where no drive answers SMART it exits **17** ("Unable to
|
||||
monitor any SMART enabled devices"), systemd marks the unit **failed**,
|
||||
nomarchy-doctor honestly reports a failed system unit, and Waybar paints
|
||||
`@bad`. **The doctor was right; smartd was the bug.** Scope was never
|
||||
live-only — `services.smartd.enable` mkDefaults true on every machine and
|
||||
QEMU virtio exposes no SMART, so *every VM install booted to a health
|
||||
warning* and every V2 run had been showing it as noise.
|
||||
Fixed by the distro's own self-gate convention: an `ExecCondition` running
|
||||
`smartctl --scan` (which prints nothing exactly when smartd would find
|
||||
nothing) leaves the unit **inactive** instead of failed. **Deliberately not
|
||||
`SuccessExitStatus = 17`** — that would also swallow exit 17 from a machine
|
||||
that *does* have drives, which is the entire reason the daemon ships.
|
||||
**V2:** `checks.smartd-gate` boots the real distro module and asserts both
|
||||
halves, because they pull in opposite directions — a gate that never skips
|
||||
leaves the red icon; a gate that always skips silently disables drive-health
|
||||
monitoring on real hardware, the failure nobody notices until a disk dies
|
||||
quietly. So: no-SMART node → `ActiveState=inactive`, not failed, absent from
|
||||
`systemctl --failed` (what the doctor reads); and the gate's logic driven
|
||||
against a scan that *finds* a device → exit 0, since QEMU cannot answer SMART
|
||||
honestly. Proved to fail by unwiring the condition. **V3 pending:** on real
|
||||
drives smartd still runs (HARDWARE-QUEUE, dev box).
|
||||
- ✗ **One chromium, not two (#121) — investigated, measured, DECIDED AGAINST
|
||||
(Bernardo, 2026-07-14): "too much work for a negligible gain".** The bug is
|
||||
real and the fix worked; the *gain* is what failed to survive measurement.
|
||||
|
||||
Reference in New Issue
Block a user