cubesys: durability fixes A/B/C verified live on VM

- Fix A (store.rs): checkpoint boundary persists wal.committed_seq (highest
  fsync'd) instead of wal.seq() (next-to-assign), which skipped all WAL
  entries since last checkpoint -> silent data loss on reboot.
- Fix B (commands.rs open): run decrypted program in isolated read_snapshot()
  clone instead of put_raw plaintext over sealed envelope (stopped reboot-time
  EnvelopeTooShort / clobber).
- Fix C (commands.rs keyinit): flush OS key cells to WAL via log_put so they
  fold into base snapshot and survive reboot (keyinit #2 issues 0, not 2);
  previously re-minted random material each boot -> sealed records unopenable.
- Regression guards durable_sealed_record_survives_restart +
  open_does_not_clobber_sealed_record in cubesys/src/commands.rs.
- STARTUP-README: replace stale 'EPHEMERAL across restarts' caveat with the
  fixed/verified durability note.
Verified live: systemctl restart cube-server (VM reboot path) -> sealed record
decrypts+executes after reboot; key cell byte-identical; keyinit idempotent.
This commit is contained in:
CUBELinux-2
2026-08-13 12:31:34 -04:00
parent ab40409316
commit f6bd8bd01f
6 changed files with 453 additions and 33 deletions
+40 -7
View File
@@ -127,9 +127,32 @@ missing.
## 5. Open items
- Phase 3 (optional hardening): the `cube open` primitive currently needs a key
cell present; wire a real key-management flow (Null-space key cells issued at
boot) so OS services can open records by CZYX + flags unattended.
- Phase 3 key-management flow: **REAL (wired 2026-08-13).** New `cubecrypt::keyinit`
module (`ensure_os_keystore`, idempotent) mints the OS Null-space key cells
(default `gcm` @ `c000/z020/y000/x001`, and `xts` @ `c000/z020/y001/x001`) ONCE,
then reuses them. Exposed as `cube keyinit` / `cubec keyinit` and run at boot by
the new `cube-os-keyinit.service` (Before= the OS state/snapshot/resume units,
After=cube-server). The `seal`/`open` CLI now accept `auto` as the key-cell
argument, resolving to the OS default key — i.e. an OS service can
`cube seal <C.Z.Y.X> auto <transform>` / `cube open <C.Z.Y.X> auto <transform>`
with NO human-supplied key, satisfying "open by CZYX + flags unattended".
Verified live against the daemon (durable store): 1st `keyinit` issued 2 cells,
2nd issued 0 (idempotent); seal+open round-trip by CZYX+auto succeeds.
DURABILITY (FIXED 2026-08-13): the daemon store is NOT ephemeral — three fixes
make key cells + sealed records survive a cold reboot, all verified live via
`systemctl restart cube-server` (the VM reboot path):
- Fix A (store.rs): checkpoint boundary now persists `wal.committed_seq`
(highest fsync'd) not `wal.seq()` (next-to-assign); the old value skipped
every WAL entry since the last checkpoint → silent data loss on reboot.
- Fix B (commands.rs `open`): executes the decrypted program in an isolated
`read_snapshot()` clone instead of `put_raw`ing plaintext over the sealed
envelope, so re-open after reboot no longer returns `EnvelopeTooShort`.
- Fix C (commands.rs `keyinit`): the OS key cells are `log_put`-flushed to the
WAL, so they fold into the base snapshot and the SAME key material decrypts
after reboot (keyinit #2 issues 0, not 2). Without this, keyinit re-minted
random material on each boot → every sealed record became unopenable.
Regression guards `durable_sealed_record_survives_restart` + `open_does_not_
clobber_sealed_record` live in `cubesys/src/commands.rs`; `./check` is green.
- Boot substrate (next deepening): make the cube the OS's *default* storage for a
real tree — e.g. have a service write `/etc` or `/var/log` operational files
through the cube by default. Currently the cube holds OS *state* (manifest/
@@ -138,10 +161,20 @@ missing.
`/cubefs/.czyx/200.1.1.1` resolves directly to the coordinate), and a
loop/overlay over a cube-backed file would let the OS treat the cube as a real
block device.
- IMAGE PROVISIONING CAVEAT: the systemd units + scripts live in the VM guest
filesystem, NOT in the CUBELinux-2 git repo. If `build_vm.sh` bakes a fresh
image, re-add these units (cube-os-state, cube-os-klog, cube-os-snapshot +
timer) to the image provisioning, or they won't survive a from-scratch rebuild.
- IMAGE PROVISIONING: `build_vm.sh` (host, /root/build_vm.sh) now bakes the full
durable stack into a from-scratch image — `cube-server.service` (daemon),
`cubefs.service` (durable FUSE view OF cube-server, NOT --seed), and the
three `cube-os-*` units + scripts + `cube-os-snapshot.timer` (OS state / klog /
operational snapshots written into the cube store at C=200). All enabled in the
chroot. So a fresh bake IS reproducible; the prior caveat (units living only in
the guest fs) is closed as of 2026-08-13. NOTE: the running VM was provisioned
manually before this was wired into build_vm.sh; re-running `build_vm.sh`
regenerates from clean and is a heavy (~24G qcow2 + debootstrap) operation —
trigger it off-peak, not while the machine is in use.
- `cube-resume-pointer.service` is REAL (wired 2026-08-13): a oneshot that writes
the OS's "where to look to continue" into the cube at `c200/z011/y001/x001`
(last snapshot index + resume coordinate). Was a dead stub before (unit pointed
at a nonexistent bin). Now deployed live in the VM and baked by build_vm.sh.
- OPTIONAL (cosmetic): launch the "CUBE Shell" desktop launcher in the VM to
confirm end-to-end at the file level (wiring already verified:
cube-term → cubec → /run/cube/cube.sock).