cubesys: durability fixes A/B/C verified live on VM
- Fix A (store.rs): checkpoint boundary persists wal.committed_seq (highest fsync'd) instead of wal.seq() (next-to-assign), which skipped all WAL entries since last checkpoint -> silent data loss on reboot. - Fix B (commands.rs open): run decrypted program in isolated read_snapshot() clone instead of put_raw plaintext over sealed envelope (stopped reboot-time EnvelopeTooShort / clobber). - Fix C (commands.rs keyinit): flush OS key cells to WAL via log_put so they fold into base snapshot and survive reboot (keyinit #2 issues 0, not 2); previously re-minted random material each boot -> sealed records unopenable. - Regression guards durable_sealed_record_survives_restart + open_does_not_clobber_sealed_record in cubesys/src/commands.rs. - STARTUP-README: replace stale 'EPHEMERAL across restarts' caveat with the fixed/verified durability note. Verified live: systemctl restart cube-server (VM reboot path) -> sealed record decrypts+executes after reboot; key cell byte-identical; keyinit idempotent.
This commit is contained in:
+40
-7
@@ -127,9 +127,32 @@ missing.
|
||||
|
||||
## 5. Open items
|
||||
|
||||
- Phase 3 (optional hardening): the `cube open` primitive currently needs a key
|
||||
cell present; wire a real key-management flow (Null-space key cells issued at
|
||||
boot) so OS services can open records by CZYX + flags unattended.
|
||||
- Phase 3 key-management flow: **REAL (wired 2026-08-13).** New `cubecrypt::keyinit`
|
||||
module (`ensure_os_keystore`, idempotent) mints the OS Null-space key cells
|
||||
(default `gcm` @ `c000/z020/y000/x001`, and `xts` @ `c000/z020/y001/x001`) ONCE,
|
||||
then reuses them. Exposed as `cube keyinit` / `cubec keyinit` and run at boot by
|
||||
the new `cube-os-keyinit.service` (Before= the OS state/snapshot/resume units,
|
||||
After=cube-server). The `seal`/`open` CLI now accept `auto` as the key-cell
|
||||
argument, resolving to the OS default key — i.e. an OS service can
|
||||
`cube seal <C.Z.Y.X> auto <transform>` / `cube open <C.Z.Y.X> auto <transform>`
|
||||
with NO human-supplied key, satisfying "open by CZYX + flags unattended".
|
||||
Verified live against the daemon (durable store): 1st `keyinit` issued 2 cells,
|
||||
2nd issued 0 (idempotent); seal+open round-trip by CZYX+auto succeeds.
|
||||
DURABILITY (FIXED 2026-08-13): the daemon store is NOT ephemeral — three fixes
|
||||
make key cells + sealed records survive a cold reboot, all verified live via
|
||||
`systemctl restart cube-server` (the VM reboot path):
|
||||
- Fix A (store.rs): checkpoint boundary now persists `wal.committed_seq`
|
||||
(highest fsync'd) not `wal.seq()` (next-to-assign); the old value skipped
|
||||
every WAL entry since the last checkpoint → silent data loss on reboot.
|
||||
- Fix B (commands.rs `open`): executes the decrypted program in an isolated
|
||||
`read_snapshot()` clone instead of `put_raw`ing plaintext over the sealed
|
||||
envelope, so re-open after reboot no longer returns `EnvelopeTooShort`.
|
||||
- Fix C (commands.rs `keyinit`): the OS key cells are `log_put`-flushed to the
|
||||
WAL, so they fold into the base snapshot and the SAME key material decrypts
|
||||
after reboot (keyinit #2 issues 0, not 2). Without this, keyinit re-minted
|
||||
random material on each boot → every sealed record became unopenable.
|
||||
Regression guards `durable_sealed_record_survives_restart` + `open_does_not_
|
||||
clobber_sealed_record` live in `cubesys/src/commands.rs`; `./check` is green.
|
||||
- Boot substrate (next deepening): make the cube the OS's *default* storage for a
|
||||
real tree — e.g. have a service write `/etc` or `/var/log` operational files
|
||||
through the cube by default. Currently the cube holds OS *state* (manifest/
|
||||
@@ -138,10 +161,20 @@ missing.
|
||||
`/cubefs/.czyx/200.1.1.1` resolves directly to the coordinate), and a
|
||||
loop/overlay over a cube-backed file would let the OS treat the cube as a real
|
||||
block device.
|
||||
- IMAGE PROVISIONING CAVEAT: the systemd units + scripts live in the VM guest
|
||||
filesystem, NOT in the CUBELinux-2 git repo. If `build_vm.sh` bakes a fresh
|
||||
image, re-add these units (cube-os-state, cube-os-klog, cube-os-snapshot +
|
||||
timer) to the image provisioning, or they won't survive a from-scratch rebuild.
|
||||
- IMAGE PROVISIONING: `build_vm.sh` (host, /root/build_vm.sh) now bakes the full
|
||||
durable stack into a from-scratch image — `cube-server.service` (daemon),
|
||||
`cubefs.service` (durable FUSE view OF cube-server, NOT --seed), and the
|
||||
three `cube-os-*` units + scripts + `cube-os-snapshot.timer` (OS state / klog /
|
||||
operational snapshots written into the cube store at C=200). All enabled in the
|
||||
chroot. So a fresh bake IS reproducible; the prior caveat (units living only in
|
||||
the guest fs) is closed as of 2026-08-13. NOTE: the running VM was provisioned
|
||||
manually before this was wired into build_vm.sh; re-running `build_vm.sh`
|
||||
regenerates from clean and is a heavy (~24G qcow2 + debootstrap) operation —
|
||||
trigger it off-peak, not while the machine is in use.
|
||||
- `cube-resume-pointer.service` is REAL (wired 2026-08-13): a oneshot that writes
|
||||
the OS's "where to look to continue" into the cube at `c200/z011/y001/x001`
|
||||
(last snapshot index + resume coordinate). Was a dead stub before (unit pointed
|
||||
at a nonexistent bin). Now deployed live in the VM and baked by build_vm.sh.
|
||||
- OPTIONAL (cosmetic): launch the "CUBE Shell" desktop launcher in the VM to
|
||||
confirm end-to-end at the file level (wiring already verified:
|
||||
cube-term → cubec → /run/cube/cube.sock).
|
||||
|
||||
Reference in New Issue
Block a user