A fold turns the log's effects into the image's by keeping, per coordinate, the *last* entry in the
merged order — that is how "the later write wins" is implemented. The order came from a derive on
`Entry`, which compares fields in declaration order, and `flags` was declared before `seq`. So the
tie-break was not arrival at all: it was the class mask.
A `del` is written with no class (0). A **sealed** record is written with `0x0800`. For the same
coordinate that put the removal *before* the record it removed, so the record won and the fold wrote
it back — a deleted record returning from the fold, with its old value, on any store whose writes are
sealed. Found on the box on 2026-09-25: put → present, del → "(not found)", `cube-fold` → "folded the
log into the image" (generation 2647 → 2648, log emptied), get → the value, back. It reproduced
through a sealed front-end and an unsealed one, and the live store still holds the test records it
resurrected.
Why it hid: an unsealed store writes both entries with flags 0, the tie falls through to `seq`, and
the order is right — which is exactly what the guest gate did before this, so the gate agreed with a
bug it could not see. The shape that fails is the shape the machine uses.
`Entry` now orders by `(space, key, seq, …)` explicitly, with the reason written beside it, because
"the fields happen to be declared in this order" is what went wrong. The userspace store cannot have
this bug and does not: it applies the log into a `HashMap`, where a removal *is* the removal of the
key, so no ordering decides anything.
Every part of this was in place except the point of it: a record can carry a 16-bit class,
CUBE_OP_PUT stamps one and CUBE_OP_FLAG_SCAN retrieves by class, the vocabulary is a userspace
convention, and the kernel already stamped its own boot record with the boot class. What was
missing was a kernel that captures *events* — one that writes about what happened to it rather
than only what a caller asked for. Until now the kernel's account of saying no was a line in
dmesg, which is not somewhere a later reader can ask.
Three events, in the events space (0xFB), each classed by the vocabulary it belongs to:
bad-op ERROR a request the interface does not offer, refused and recorded
clock-set BOOT the epoch was provisional, and the record was rewritten
boot-record-late ERROR|BOOT the record only landed on a retry — the one that matters most,
because that defect's effect was invisible in the store
The write is bounded twice (EVENT_BUDGET, 32 a boot; REFUSAL_BUDGET, 8 among them) and that is
not tidiness: a kernel that appends a record per event can turn a storm of refused requests into
a storm of writes, which this store has already met from the other side when a walk with no store
behind it served ~200,000 invented records a minute. Past the ceiling the kernel logs and stops.
The refusal capture is exported for the syscall layer to call (`cubelinux_kernel_capture_refusal`)
because that is where the refusals that leave the store usable happen — a bad size, an op that
does not exist, a value past the maximum. A store that cannot be read at all is the one refusal
the kernel cannot record into itself, and that is stated rather than papered over.
The record the kernel writes about its own boot carried `boot=<epoch>` and nothing else — a
reading taken at the first write of the boot, which is exactly when this machine's clock is
least likely to be right: it arrives from a night powered off tens of seconds out, and on the
boot before it, 10h 27m behind. Nothing in the record said the time was provisional.
Two fields carry the account now. `uptime` comes from the monotonic clock, so `boot - uptime`
is the instant the boot began whatever the wall clock was doing. `clock=raw|set` says whether
the reading predates a correction. And the kernel acts on the difference: the wall clock can be
set from anywhere and the monotonic clock cannot be set at all, so a wall clock that moves
without it is a correction and nothing else — when that is observed, the next write rewrites
the record at the same coordinate with the corrected time, bounded like the first write.
Measured in the guest (verify-boot-record, which now moves the clock forward an hour mid-boot):
boot=1790308025 uptime=3 clock=raw, then the same coordinate at boot=1790311625 uptime=3
clock=set — exactly the hour that was moved. The gate also refuses a record whose epoch
precedes its own uptime, and one that claims to have been written long after the boot began.
The retry half of this was already in the tree (a count rather than a swap that cannot
un-claim itself, from e968b3964); this is the epoch half, and the record's own "Next" list
named both.
`verify-kernel-append` asks the question a log exists to answer: four writes are acknowledged, the
VM is killed with `SIGKILL` with no shutdown at all, and the log must still hold them. It is red on
`#95` and on `#87` — the kernel this box runs — and green on the kernel built at 18:41, and the one
commit between them is e0218ec96: the collapse to a single `fsync` per append, where the entry goes
down unsynced and the control write that counts it carries the flush for both.
That collapse is right where its sentence is true, and the sentence assumes there is a control block
to write. *A store with no control block has no count write*, so nothing on the append path was
flushed at all — and the caller was told a write it may never get back. That is not an exotic shape:
it is what a store the CLI builds looks like, and it is the shape that gate's own device has.
Measured, same device, same four writes, the host's copy of the device read while the guest is still
hung:
no control block the log region holds its header and zeros where the entry should be, and
the fold returns the base store: records=2
a formatted device the entry is on the device, and the fold reproduces the userspace store
exactly: records=6, fnv1a64=6b679d39597a62b3
and a pre-collapse kernel keeps the entry in *both* shapes, because the entry's own `fsync` is what
made it durable there.
So the fix is not the revert. `append` takes the entry's own `fsync` exactly when no later write in
the same operation will carry one — `layout.control.is_some()`, one boolean, both branches gated by
the two shapes above. The box's store is a formatted file with a control block, so it keeps the
measured 4.96 ms win; a bare store pays its entry's flush, which is what it always did before the
collapse. The count still goes second, so a crash can still lose an unacknowledged mutation rather
than count one that is not there.
The gate was wrong in the same way the first diagnosis was, and is fixed with the kernel: it ran its
crash check over one device shape, which is the shape that made the collapse look safe everywhere. It
now runs it over both — a store with no control block and a formatted device — folds each the way
userspace reads it, and compares both against the same userspace store, so "it survived" means the
same thing for each. It also stops SIGKILLing the `timeout` wrapper rather than QEMU: SIGKILL cannot
be forwarded, so every run of this gate left a 512 MiB VM spinning in `pause()` for the rest of the
wall, twice found beside a gate that measures latency.
Built as #96. The wall, whole: 26 of 26.
The walk with no store behind it served ~200,000 invented records a minute and never ended.
Half of that was fixed and proven in fc057820f: the driver asks whether the store can be read
before either walk op answers anything, and it refuses. The client was still handed
spaces returned 0, len=0, cursor=1
— success, no space, a cursor one further on — so it copied the space it was handed out of a
buffer the kernel never wrote, and with the cursor moving by itself the rule that ends every
other walk here, *no progress is the only end signal*, had nothing to fire on.
The remaining half was not in the C arm, and both candidate explanations recorded there are
wrong: the `ret < 0` test IS on the path, and nothing overwrote the answer. One boot at
loglevel=7 says what crosses the boundary instead:
cubelinux: the store is not readable; refusing to answer (x138 in 40 s)
walk: spaces returned 0, len=0, cursor=1 (the client, told success)
The guard fires and the client is told success, so the value is not negative. On that path the
driver's only return is `-(e.to_errno() as i32)`, and `kernel::error::Error` IS the kernel error
code: `from_errno(-2) == ENOENT`, `to_errno()` is documented as "the kernel error code", and the
API's own conversion of a `Result` to a C result — `kernel::from_result` — writes
`T::from(e.to_errno() as i16)` with no negation, as every other driver in this tree does. The
second minus sign made every refusal in this file a *positive* number, and a positive return is
what the C arm's `ret < 0`, the client's own wrapper, and the walk arms' `ret == 0` all read as
success.
38 sites in cubelinux_store.rs were shaped `-(e.to_errno() as i32)` / `as isize`. All 38 now
return `e.to_errno()`, which is what makes them refusals. Nothing else about them changed.
Why this became a *loop* in the space walk and nowhere else: CUBE_OP_SPACES is the one arm that
advances the cursor itself. Every other walk arm leaves the cursor where the caller put it, so a
bogus return there ends the walk on the client's no-progress rule — which is why one defect was
invisible at every other verb for as long as it existed. That arm now refuses a return that is
neither 0 ("here is a space") nor negative (a refusal), so a defect of this shape cannot be read
as a space again.
One more correction in the same class, found while proving the above. A store that cannot be
*opened* answered -ENOENT, and -ENOENT is this interface's own end-of-walk signal — "no such
space; the walk is finished" — so a walk over a store on a disk whose driver had not loaded read
exactly like a walk over an empty store, which is what the first benchmark boot was.
`store_file()` now answers ENODEV when the store device is not there. A store that is not there
is not an empty store.
Proven against the reproduction, in the guest, on the bench initramfs:
cube_store=/dev/null -> "walk: spaces failed: Invalid argument", 0 records, 5 s, boot finishes
cube_store=/nowhere/x.img -> "walk: spaces failed: No such device", 0 records, 5 s, boot finishes
before: 364,994 invented record lines in the 90 s the instrument allowed
The three checks are `kernel/verify-no-store.sh`, a gate on the wall now, and the same three
inside the benchmark rehearsal — which is where this defect was found, and where they were
warnings while it was open.
The diagnostic prints this was hunted with come off in the same commit: the `cube_store=
resolved to` line in cube_syscall.c, and the per-call "not readable" warning in both walk ops.
The refusal is the return value, and the walk's own transcript is where a reader learns what
happened; a message per call is a diagnostic, not the interface. The guard itself stays, and
where to find it is written down at its definition rather than implied.
Built as #95, which is what the box now has installed: the machine boots it on its next reboot.
The guard is `store_readable()`: four bytes at offset zero, asked before either walk
op answers anything. Proven against the reproduction, which is the reproduction
from the record — `cube_store=/dev/null cubelinux.enum=1`:
cubelinux: cube_store= resolved to /dev/null (the token was honoured)
cubelinux: the store is not readable; refusing to answer (the kernel refuses)
and the client is *still* handed `spaces returned 0, len=0, cursor=1`, so it copies
an unfilled space and walks 199,000 invented records. That is the whole defect in
one transcript: the driver refuses and the syscall boundary reports success anyway.
So this commit fixes and proves half of it. What remains is `cube_syscall.c`'s
CUBE_OP_SPACES arm, where the driver's -EINVAL becomes a 0 with a cursor advanced by
one — either its `ret < 0` test is not on the path the client takes, or the answer is
overwritten before the copy-out. Nothing above or below that needs touching.
Three earlier guesses are recorded as wrong rather than deleted: the packed fallback
and the magic guard written for it are never reached (probed, proven), and
`read_exact_at` already rejects a short read. The guard here is the first one that
fires.
The device-path print in cube_syscall.c stays on purpose: it is what turned four
builds of inference into one line of fact.
Built as #93, deliberately NOT installed — installing it alone would leave the walk
still looping. The machine runs #87.
This is the guard for one of the two doors on the walk-with-no-store defect:
"neither control block decodes" is not evidence of a packed image, so the packed
reader is allowed only for a device that begins with the image header's own magic —
the question the userspace raw reader has always asked and the kernel did not.
It builds clean and it does NOT stop the reproduction: `cube_store=/dev/null
cubelinux.enum=1` still serves invented records at the same rate. So the guard is
not on the path that answers, and it is committed for that reason stated plainly
rather than installed: this machine stays on #87, which is measured, instead of
moving to a kernel whose only change is one that has not been shown to do anything.
Kept rather than reverted because the check is right on its own terms — an
addressed store begins with CUBS, a bare packed image with CUBE — and because the
next attempt should begin by asking whether this guard is even reached before
writing another one.
The arrangement vocabulary's second half — the producer and the userspace
reader have existed since 0244-ish, and this is the kernel joining them for a
caller who does not know it holds a chain. Bounded by the same 65,536 slices the
userspace reader uses and by MAX_CHAIN_BYTES (CUBE_MAX_VALUE, 16 MiB: a longer
chain could not be handed over in one call, so reading on would be work whose
answer nobody can receive). A hole, a slice that claims to continue and does not,
and a second beginning are each -EIO rather than a guessed end.
A joined read reports START|END, because the flags describe the bytes handed over
rather than the slice they came from — which is also why the userspace reader
needed no change: it already stops at END_RECORD.
It does NOT join a sealed chain, and that is a fact about the format: a sealed
slice is an envelope with its own header and nonce, so N of them concatenated are
not one openable value, and joining them belongs to whoever holds the key. When
it cannot join (sealed, or a packed image with no index) it returns the slice with
the record's own flags — CONTINUATION without END_RECORD says "this is a slice,
not the whole value" — so the caller is told rather than misled.
Reading those bits at all was safe because nothing uses them, and that was
measured before the code was written: over the live store's 69,749 records,
flag-scan finds 1 record in 0x0080, 137 in 0x0800, and zero in every one of
0x0001, 0x0002, 0x0003, 0x0004, 0x0008, 0x0010, 0x0020, 0x0040, 0x0100, 0x0200.
Gated by kernel/verify-chain.sh: PASS, seven checks.
2026-09-23 19:59:57 -04:00
2 changed files with 674 additions and 68 deletions
File diff suppressed because it is too large
Load Diff
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.