Compare commits

...
16 Commits
Author SHA1 Message Date
CUBELinux 35fcb6c7fe store: a fold kept a deletion only when the record had no class
A fold turns the log's effects into the image's by keeping, per coordinate, the *last* entry in the
merged order — that is how "the later write wins" is implemented. The order came from a derive on
`Entry`, which compares fields in declaration order, and `flags` was declared before `seq`. So the
tie-break was not arrival at all: it was the class mask.

A `del` is written with no class (0). A **sealed** record is written with `0x0800`. For the same
coordinate that put the removal *before* the record it removed, so the record won and the fold wrote
it back — a deleted record returning from the fold, with its old value, on any store whose writes are
sealed. Found on the box on 2026-09-25: put → present, del → "(not found)", `cube-fold` → "folded the
log into the image" (generation 2647 → 2648, log emptied), get → the value, back. It reproduced
through a sealed front-end and an unsealed one, and the live store still holds the test records it
resurrected.

Why it hid: an unsealed store writes both entries with flags 0, the tie falls through to `seq`, and
the order is right — which is exactly what the guest gate did before this, so the gate agreed with a
bug it could not see. The shape that fails is the shape the machine uses.

`Entry` now orders by `(space, key, seq, …)` explicitly, with the reason written beside it, because
"the fields happen to be declared in this order" is what went wrong. The userspace store cannot have
this bug and does not: it applies the log into a `HashMap`, where a removal *is* the removal of the
key, so no ordering decides anything.
2026-09-25 03:35:52 -04:00
CUBELinux 2586c2ed0d store: the kernel captures its own events, as classed records
Every part of this was in place except the point of it: a record can carry a 16-bit class,
CUBE_OP_PUT stamps one and CUBE_OP_FLAG_SCAN retrieves by class, the vocabulary is a userspace
convention, and the kernel already stamped its own boot record with the boot class. What was
missing was a kernel that captures *events* — one that writes about what happened to it rather
than only what a caller asked for. Until now the kernel's account of saying no was a line in
dmesg, which is not somewhere a later reader can ask.

Three events, in the events space (0xFB), each classed by the vocabulary it belongs to:

  bad-op         ERROR   a request the interface does not offer, refused and recorded
  clock-set      BOOT    the epoch was provisional, and the record was rewritten
  boot-record-late ERROR|BOOT  the record only landed on a retry — the one that matters most,
                         because that defect's effect was invisible in the store

The write is bounded twice (EVENT_BUDGET, 32 a boot; REFUSAL_BUDGET, 8 among them) and that is
not tidiness: a kernel that appends a record per event can turn a storm of refused requests into
a storm of writes, which this store has already met from the other side when a walk with no store
behind it served ~200,000 invented records a minute. Past the ceiling the kernel logs and stops.

The refusal capture is exported for the syscall layer to call (`cubelinux_kernel_capture_refusal`)
because that is where the refusals that leave the store usable happen — a bad size, an op that
does not exist, a value past the maximum. A store that cannot be read at all is the one refusal
the kernel cannot record into itself, and that is stated rather than papered over.
2026-09-25 01:06:14 -04:00
CUBELinux 3d5d213c75 store: a boot record whose epoch can be trusted, not just read
The record the kernel writes about its own boot carried `boot=<epoch>` and nothing else — a
reading taken at the first write of the boot, which is exactly when this machine's clock is
least likely to be right: it arrives from a night powered off tens of seconds out, and on the
boot before it, 10h 27m behind. Nothing in the record said the time was provisional.

Two fields carry the account now. `uptime` comes from the monotonic clock, so `boot - uptime`
is the instant the boot began whatever the wall clock was doing. `clock=raw|set` says whether
the reading predates a correction. And the kernel acts on the difference: the wall clock can be
set from anywhere and the monotonic clock cannot be set at all, so a wall clock that moves
without it is a correction and nothing else — when that is observed, the next write rewrites
the record at the same coordinate with the corrected time, bounded like the first write.

Measured in the guest (verify-boot-record, which now moves the clock forward an hour mid-boot):
boot=1790308025 uptime=3 clock=raw, then the same coordinate at boot=1790311625 uptime=3
clock=set — exactly the hour that was moved. The gate also refuses a record whose epoch
precedes its own uptime, and one that claims to have been written long after the boot began.

The retry half of this was already in the tree (a count rather than a swap that cannot
un-claim itself, from e968b3964); this is the epoch half, and the record's own "Next" list
named both.
2026-09-25 00:09:09 -04:00
surface-camera-build 8c24a7ff0d store: an acknowledged write is durable again, in the shape nothing was flushing
`verify-kernel-append` asks the question a log exists to answer: four writes are acknowledged, the
VM is killed with `SIGKILL` with no shutdown at all, and the log must still hold them. It is red on
`#95` and on `#87` — the kernel this box runs — and green on the kernel built at 18:41, and the one
commit between them is e0218ec96: the collapse to a single `fsync` per append, where the entry goes
down unsynced and the control write that counts it carries the flush for both.

That collapse is right where its sentence is true, and the sentence assumes there is a control block
to write. *A store with no control block has no count write*, so nothing on the append path was
flushed at all — and the caller was told a write it may never get back. That is not an exotic shape:
it is what a store the CLI builds looks like, and it is the shape that gate's own device has.
Measured, same device, same four writes, the host's copy of the device read while the guest is still
hung:

    no control block    the log region holds its header and zeros where the entry should be, and
                        the fold returns the base store: records=2
    a formatted device  the entry is on the device, and the fold reproduces the userspace store
                        exactly: records=6, fnv1a64=6b679d39597a62b3

and a pre-collapse kernel keeps the entry in *both* shapes, because the entry's own `fsync` is what
made it durable there.

So the fix is not the revert. `append` takes the entry's own `fsync` exactly when no later write in
the same operation will carry one — `layout.control.is_some()`, one boolean, both branches gated by
the two shapes above. The box's store is a formatted file with a control block, so it keeps the
measured 4.96 ms win; a bare store pays its entry's flush, which is what it always did before the
collapse. The count still goes second, so a crash can still lose an unacknowledged mutation rather
than count one that is not there.

The gate was wrong in the same way the first diagnosis was, and is fixed with the kernel: it ran its
crash check over one device shape, which is the shape that made the collapse look safe everywhere. It
now runs it over both — a store with no control block and a formatted device — folds each the way
userspace reads it, and compares both against the same userspace store, so "it survived" means the
same thing for each. It also stops SIGKILLing the `timeout` wrapper rather than QEMU: SIGKILL cannot
be forwarded, so every run of this gate left a 512 MiB VM spinning in `pause()` for the rest of the
wall, twice found beside a gate that measures latency.

Built as #96. The wall, whole: 26 of 26.
2026-09-23 23:16:31 -04:00
surface-camera-build 8db279a4d8 cube(2): a refusal with two minus signs is not a refusal — and the walk it hung
The walk with no store behind it served ~200,000 invented records a minute and never ended.
Half of that was fixed and proven in fc057820f: the driver asks whether the store can be read
before either walk op answers anything, and it refuses. The client was still handed

    spaces returned 0, len=0, cursor=1

— success, no space, a cursor one further on — so it copied the space it was handed out of a
buffer the kernel never wrote, and with the cursor moving by itself the rule that ends every
other walk here, *no progress is the only end signal*, had nothing to fire on.

The remaining half was not in the C arm, and both candidate explanations recorded there are
wrong: the `ret < 0` test IS on the path, and nothing overwrote the answer. One boot at
loglevel=7 says what crosses the boundary instead:

    cubelinux: the store is not readable; refusing to answer     (x138 in 40 s)
    walk: spaces returned 0, len=0, cursor=1                     (the client, told success)

The guard fires and the client is told success, so the value is not negative. On that path the
driver's only return is `-(e.to_errno() as i32)`, and `kernel::error::Error` IS the kernel error
code: `from_errno(-2) == ENOENT`, `to_errno()` is documented as "the kernel error code", and the
API's own conversion of a `Result` to a C result — `kernel::from_result` — writes
`T::from(e.to_errno() as i16)` with no negation, as every other driver in this tree does. The
second minus sign made every refusal in this file a *positive* number, and a positive return is
what the C arm's `ret < 0`, the client's own wrapper, and the walk arms' `ret == 0` all read as
success.

38 sites in cubelinux_store.rs were shaped `-(e.to_errno() as i32)` / `as isize`. All 38 now
return `e.to_errno()`, which is what makes them refusals. Nothing else about them changed.

Why this became a *loop* in the space walk and nowhere else: CUBE_OP_SPACES is the one arm that
advances the cursor itself. Every other walk arm leaves the cursor where the caller put it, so a
bogus return there ends the walk on the client's no-progress rule — which is why one defect was
invisible at every other verb for as long as it existed. That arm now refuses a return that is
neither 0 ("here is a space") nor negative (a refusal), so a defect of this shape cannot be read
as a space again.

One more correction in the same class, found while proving the above. A store that cannot be
*opened* answered -ENOENT, and -ENOENT is this interface's own end-of-walk signal — "no such
space; the walk is finished" — so a walk over a store on a disk whose driver had not loaded read
exactly like a walk over an empty store, which is what the first benchmark boot was.
`store_file()` now answers ENODEV when the store device is not there. A store that is not there
is not an empty store.

Proven against the reproduction, in the guest, on the bench initramfs:

    cube_store=/dev/null        -> "walk: spaces failed: Invalid argument", 0 records, 5 s, boot finishes
    cube_store=/nowhere/x.img   -> "walk: spaces failed: No such device",    0 records, 5 s, boot finishes
    before:                       364,994 invented record lines in the 90 s the instrument allowed

The three checks are `kernel/verify-no-store.sh`, a gate on the wall now, and the same three
inside the benchmark rehearsal — which is where this defect was found, and where they were
warnings while it was open.

The diagnostic prints this was hunted with come off in the same commit: the `cube_store=
resolved to` line in cube_syscall.c, and the per-call "not readable" warning in both walk ops.
The refusal is the return value, and the walk's own transcript is where a reader learns what
happened; a message per call is a diagnostic, not the interface. The guard itself stays, and
where to find it is written down at its definition rather than implied.

Built as #95, which is what the box now has installed: the machine boots it on its next reboot.
2026-09-23 22:42:36 -04:00
surface-camera-build fc057820fa store: refuse a store that cannot be read — and the proof that the refusal is not the whole defect
The guard is `store_readable()`: four bytes at offset zero, asked before either walk
op answers anything. Proven against the reproduction, which is the reproduction
from the record — `cube_store=/dev/null cubelinux.enum=1`:

  cubelinux: cube_store= resolved to /dev/null                 (the token was honoured)
  cubelinux: the store is not readable; refusing to answer      (the kernel refuses)

and the client is *still* handed `spaces returned 0, len=0, cursor=1`, so it copies
an unfilled space and walks 199,000 invented records. That is the whole defect in
one transcript: the driver refuses and the syscall boundary reports success anyway.

So this commit fixes and proves half of it. What remains is `cube_syscall.c`'s
CUBE_OP_SPACES arm, where the driver's -EINVAL becomes a 0 with a cursor advanced by
one — either its `ret < 0` test is not on the path the client takes, or the answer is
overwritten before the copy-out. Nothing above or below that needs touching.

Three earlier guesses are recorded as wrong rather than deleted: the packed fallback
and the magic guard written for it are never reached (probed, proven), and
`read_exact_at` already rejects a short read. The guard here is the first one that
fires.

The device-path print in cube_syscall.c stays on purpose: it is what turned four
builds of inference into one line of fact.

Built as #93, deliberately NOT installed — installing it alone would leave the walk
still looping. The machine runs #87.
2026-09-23 22:02:22 -04:00
surface-camera-build 9fb4169e52 store: refuse a device that does not carry the packed image's magic (inert, NOT installed)
This is the guard for one of the two doors on the walk-with-no-store defect:
"neither control block decodes" is not evidence of a packed image, so the packed
reader is allowed only for a device that begins with the image header's own magic —
the question the userspace raw reader has always asked and the kernel did not.

It builds clean and it does NOT stop the reproduction: `cube_store=/dev/null
cubelinux.enum=1` still serves invented records at the same rate. So the guard is
not on the path that answers, and it is committed for that reason stated plainly
rather than installed: this machine stays on #87, which is measured, instead of
moving to a kernel whose only change is one that has not been shown to do anything.

Kept rather than reverted because the check is right on its own terms — an
addressed store begins with CUBS, a bare packed image with CUBE — and because the
next attempt should begin by asking whether this guard is even reached before
writing another one.
2026-09-23 21:15:53 -04:00
surface-camera-build 4cfdbfe63b store: a get on the first slice of a chain returns the whole value
The arrangement vocabulary's second half — the producer and the userspace
reader have existed since 0244-ish, and this is the kernel joining them for a
caller who does not know it holds a chain. Bounded by the same 65,536 slices the
userspace reader uses and by MAX_CHAIN_BYTES (CUBE_MAX_VALUE, 16 MiB: a longer
chain could not be handed over in one call, so reading on would be work whose
answer nobody can receive). A hole, a slice that claims to continue and does not,
and a second beginning are each -EIO rather than a guessed end.

A joined read reports START|END, because the flags describe the bytes handed over
rather than the slice they came from — which is also why the userspace reader
needed no change: it already stops at END_RECORD.

It does NOT join a sealed chain, and that is a fact about the format: a sealed
slice is an envelope with its own header and nonce, so N of them concatenated are
not one openable value, and joining them belongs to whoever holds the key. When
it cannot join (sealed, or a packed image with no index) it returns the slice with
the record's own flags — CONTINUATION without END_RECORD says "this is a slice,
not the whole value" — so the caller is told rather than misled.

Reading those bits at all was safe because nothing uses them, and that was
measured before the code was written: over the live store's 69,749 records,
flag-scan finds 1 record in 0x0080, 137 in 0x0800, and zero in every one of
0x0001, 0x0002, 0x0003, 0x0004, 0x0008, 0x0010, 0x0020, 0x0040, 0x0100, 0x0200.

Gated by kernel/verify-chain.sh: PASS, seven checks.
2026-09-23 19:59:57 -04:00
surface-camera-build e0218ec96b store: one fsync per append, not two — and torn-tail is the gate that decided it
A durable write was 8.2 ms and it was two fsyncs to ext4: the log entry, then the control
block that counts it. Measured on the store's own filesystem, one `pwrite`+`fsync` is
4.321 ms and append's shape of two is 8.630 ms, against 8.202 ms for the whole kernel put —
so the write path was fsyncs and almost nothing else, and one of them was flushing data the
next one covers.

The order stays. A mutation is still written as entry-then-count, because that order is what
makes a crash lose an unacknowledged mutation rather than count one that is not there. What
goes is the *second durability boundary*: the entry is `kernel_write`-ordered before the
control write, and the control write's `fsync` flushes what precedes it in the same file.

What that gives up, stated rather than implied: the two writes are no longer independently
durable, so a crash inside the control write's commit can in principle leave the control
block counting bytes whose entry did not fully reach the media. That is a torn tail — and
this store already has a gate for one:

    verify-torn-tail   PASS  "userspace on the same bytes, and appended over the tear
                              rather than after it"
    verify-syscall     PASS  the store the kernel writes is byte-identical to userspace's
    verify-file-store  PASS  twelve writes to a store that is a FILE, all accounted for

So the change was made against the gate that tests the failure mode rather than against an
argument about ext4's journal, and it is reversible on its own: revert this commit, rebuild,
and the two-fsync order returns. Half of every write, for a property the format already
handles.
2026-09-23 19:25:34 -04:00
surface-camera-build b3dc55392e read: the log window is cached with the table — the last device read a read made for nothing
The log is the one thing in the store that changes without a fold, which is why it was the
one thing still read on every call even after the validation, header, layout and space table
were held across calls. But that is the same argument in reverse: the cache is keyed on the
control block's generation, and a write — the only thing that changes the log — raises it. So
an unchanged generation IS a log that cannot have changed, and a hit can serve it from memory
without asking the device at all.

Measured, this one is neutral in the gate: get 0.038 -> 0.044 ms, inside the run-to-run noise
of this bench. That is expected rather than disappointing — the gate's store carries a log of
a few hundred bytes, so a read of it costs about what the copy from the cache does. The number
this change is for is on the box, whose log is 143,317 bytes: that read and its allocation
were ~15-30 us of every coordinate read there, and the box is where it will show.

With this the objective's list is closed: the control-block validation, the header, the space
table and the log window are all held across calls; the space-table search allocates nothing
and the index search allocates once per search rather than once per probe; the unfolded log is
bounded by LOG_FOLD_BYTES so per-call work has a ceiling; and the gate fails on a worst case
and a ceiling rather than on a ratio.
2026-09-23 18:41:36 -04:00
surface-camera-build b8a7f615ff read: validate the store in 96 bytes, and hold the layout with the table
The control block is 4096 bytes, and the question every read asks of it is one number — has
this store changed since I last looked. That number, and the fields beside it, are in the
first `CTL_SUMMED + 4` bytes of each copy. So a call that finds the store unmoved now reads
two heads and nothing else; the header, the space table, the layout and the log's extent all
come from the cache. The full 4096-byte read happens only when the store has actually moved.

Both heads are read, and that is the whole safety of it: a write raises one copy's generation
and leaves the other at the old one, so reading a single copy and finding it unchanged would
call a moved store unmoved. The pair is what makes the answer true.

Measured on #84, in the guest, against #83:
    get   0.077 -> 0.038 ms in the 20,000-record store  (and 0.052 in the 500-record one)
    miss  0.074 -> 0.044 ms
    spaces 0.374 -> 0.173 ms
    worst case: read 0.763 ms (ceiling 5), write 6.762 ms (ceiling 50)
A read is now about twice as fast, and it is finally faster in the LARGER store than in the
small one — which the flatness before could not show. That is the diagnosis paying off: the
per-call cost of a read was setup rather than device reads all along, and the largest single
piece of that setup was reading 4 KB to compare 8 bytes.

Also derives `Copy` for `HeaderV3` (six plain numbers; a cached header has to be handed out by
value) and drops the search buffer `space_entry` no longer needs, now that the table it
searches is in memory.
2026-09-23 18:37:23 -04:00
surface-camera-build 1f7b9fe0de read: hold the space table across calls, keyed on the generation
Every coordinate read binary-searches the space table for its space's first record, and
every batch of a listing reads a row of it — from the device, per probe, for data that
changes only when the store is folded. It is now held in the driver and keyed on the
**control block's generation**, which a fold raises along with everything else it writes,
plus the image offset, since a fold flips the slot. Nothing has to invalidate it by hand: a
writer that bumps the generation invalidates it, including a writer this driver never sees.
A stale table cannot outlive the image it describes.

Honest about what it bought: nothing measurable in the gate, which is the interesting part.
The gate's store carries about a dozen spaces, so the search it replaces was three or four
probes, and copying a few hundred bytes of table costs about what those probes did — get
0.078 -> 0.077 ms, spaces flat within the run-to-run noise of this bench (which is ~20%).
What changes is the SHAPE: the lookup no longer scales with the number of spaces, and the
box has 21 and grows. That is worth having and it is not the same thing as a measured win,
so it is not claimed as one.

What is still read per call, and is the larger constant: the control block (needed — it is
the validation), the header, the log window, and the allocations for all three. Those are
what the borrow-based version of this cache is for, and they are the next piece.
2026-09-23 18:05:55 -04:00
surface-camera-build aab316a9a1 read: take the log window instead of copying it, and name the boot class bit
A coordinate read cost 0.082 ms in a 500-record store and 0.081 ms in a 20,000-record
one — flat to a microsecond while the index search does nine probe reads against
fifteen. So the device reads are nearly free and the per-call cost is SETUP, and the
clearest piece of setup was a copy of the log window into a second buffer on every read.
`v3_get` copied it because what the log says borrows it while reading the image needs
`&mut image`; that borrow is one field, so moving the field out settles it for nothing.

Measured on #82: get 0.082 -> 0.078 ms, spaces 0.339 -> 0.267 ms. That is small in the
GUEST, and the guest cannot show the real number: its store carries a log of a few
hundred bytes, while the box's log is 132,679 bytes — so the copy this removes costs
~130 KB of memcpy and a ~130 KB allocation per read there, and only the box can measure
it.

Also names the store's own class bit. `BOOT_FLAGS` was a bare `1 << 7` and read as
`TYPE_MASK == 10` ("data type: code") to anything applying `WordFlags` — a latent
collision, not a live one, and the resolution is a declaration rather than a move:
`WordFlags` is a different field and its sixteen bits are all allocated, so there is no
bit to borrow, and the space axis already makes a boot record findable (reserved space
0xFC) without a class bit at all. What it needed was a vocabulary that claims it, which
this comment now does. The one bit the two fields share is shared on purpose and by name:
0x0800, `SEALED_FLAG` in `cube-store-seal`, which `WordFlags` calls `ENCRYPTED` — same
meaning in both places, which is the pattern rather than the exception.
2026-09-23 17:53:38 -04:00
surface-camera-build e968b3964e store: a write reads the log, not the whole device — and the log is now bounded
A put cost 60-131 ms on the box while a read cost 1-2 ms, and the reason was one call:
every write path reached `append` through `device_and_layout()` → `read_image()`, which
reads the WHOLE DEVICE — 100,663,296 bytes here — into a KVVec, to append ~70 bytes.
`append` itself never looks at the image: it reads the log region and the control
block. The read paths had been taught to read only where the image lies; the write
paths never were.

`AppendSource` now names what an append reads: `whole_image` (kept for the two callers
that genuinely need it — a bare store, whose log runs to the end of the file, and the
v1→v4 migration, which folds) or `log_head` (a device's control block plus the log's
first `WAL_HEADER_LEN` bytes). `write_view()` reads exactly that, `append_mutation()`
is the single entry point all four writers use — put, del, the boot record, and the
misc device's write_iter — and the bare-image fallback is decided in one place instead
of four. The fold still reads the whole store, because a fold rewrites it; that is the
honest tail, and it is why a fold belongs on a timer rather than on the path a caller
waits behind.

The log is bounded by `LOG_FOLD_BYTES`, because it is on the READ path: `Addressed::open`
reads `WAL_HEADER_LEN + log_used` bytes on every call, so an unfolded log is a tax on
every read, not a bill paid once at the fold. The live store's control block reports
log capacity 50,335,744 bytes — if it ever filled, every coordinate read would read
~48 MB before answering anything. A write that finds the log over the line folds first,
through `fold_for_headroom`, and then appends to a fresh log; guarded so one append
folds at most once, because a fold that did not shrink the log must not loop. The trade
is named in the code: the write that crosses the line pays a bounded, rare fold instead
of every reader paying an ever-larger log. The timer's comment — "the log grows at
roughly a megabyte a day against a 50 MB region, and a fold rewrites the whole image, so
folding more often would buy nothing and cost I/O" — is a data workload's arithmetic,
where the log is a recovery artefact and nobody reads it.

`space_entry`, `find` and `lower_bound` each allocated a fresh KVVec inside their
binary-search loop; one buffer per search now.

The boot record's one attempt becomes a bounded retry: `ensure_boot_record` claimed the
boot with a swap BEFORE it tried, so a single transient failure cost the boot its record
silently — which is exactly what happened on the box, where the first client of the boot
could not open the store for writing. It now claims one of `BOOT_RECORD_MAX_ATTEMPTS`
with a compare-exchange, and parks the count at the ceiling on success.

Release moves to 6.19.3-cubelinux0.7: verify-box-preflight.sh now refuses a same-release
reinstall, because the default entry boots the release being replaced.

Gates, on #81:
  verify-enum-cost    PASS  a write is 1.45 ms mean / 7.46 ms worst (new ceilings: 50 ms
                            write, 5 ms read), and a write does not grow with the store
  verify-file-store   PASS  12 writes to a store that is a FILE, all accounted for
  verify-syscall      PASS  the store the kernel writes is byte-identical to userspace's
2026-09-23 17:32:05 -04:00
surface-camera-build daa6289e2d cube(2): the walk skips the records the cursor already returned
`v3_enum` built its batch with `seen: cursor`, and `Batch::offer` increments
before comparing (`seen += 1; if seen <= cursor { skip }`), so nothing was ever
skipped: every call re-served the space from its first record while `out_cursor`
advanced by the records returned. The contract's only end signal is a cursor that
stops moving, so the walk never ended — a listing looped inside the front-end
until the kernel's OOM killer took that process, twice, at ~6.4 GiB anon.

Every other Batch site already starts `seen` at 0, and two of them carry comments
describing exactly this trap; this one was the lone outlier. It only shows on an
**addressed** image whose space has log edits, because the merge path re-walks
from the space's start and has nothing but `seen` to skip with — the no-edits path
skips by index arithmetic (`first + cursor`) and was already right. That is why no
gate saw it: the enumeration gate's store is packed.

So each path now skips exactly once — the merge path through `offer`'s `seen`
starting at 0, the no-edits path through its arithmetic with the batch told to
skip nothing. Doing both would skip twice and lose records. The arithmetic cursor
is added saturating, because a wrapped sum would land back near the first record
and re-serve the walk: the same endless listing this exists to avoid.

Verified: a copy of the box's own addressed store walks in QEMU on this kernel to
`12 space(s), 69665 record(s)` — 69,641 image records plus 24 unfolded log edits —
and terminates, where before every batch repeated the space from its start.
2026-09-23 01:32:55 -04:00
CUBELinux e53ae04033 cubelinux: hold the store lock across the whole op, not inside append
Initialising STORE_FILE stopped the oops and exposed what it had been hiding: twelve
concurrent puts all reported `ok` and one write survived. The lock guarded the file
handle, not the store — it is taken, the handle is fetched or cached, and released
before the caller does anything with it. Sharing a struct file * is safe, so that is
all the handle needs, and it is not enough for the store.

Widening it inside append is also not enough, and this commit exists because the first
attempt did exactly that and changed nothing. append receives the layout as an argument,
and the caller read it from device_and_layout() *before* calling: twelve writers each
read log_used = 0, then queued on a lock inside append, then each appended at the same
offset against the layout they had already read. Measured with the lock in append: one
68-byte entry, twelve `ok`s. The lock has to be held from before the read.

So STORE_OP is taken by the four write ops — put, del, sync and the device write — and
everything under them runs with it held: ensure_boot_record, device_and_layout, append,
fold_now. None of those may take it again; a kernel mutex is not reentrant, and append
folds and then calls itself, while fold_now is reachable both from inside append and on
its own. That is why the lock sits at the ops rather than in the writers.

verify-file-store.sh MODE=race, twelve concurrent puts: 816 bytes of log = 12 x 68, and
control generation 13 = 1 + 12, where the previous kernel produced 68 and 2.

Regression: MODE=seq still exact; verify-boot-record passes (it is the path this
changes most, since ensure_boot_record now runs under the lock); verify-kernel-append
passes, four acknowledged writes surviving a SIGKILL with no shutdown.
2026-09-22 23:32:18 -04:00
3 changed files with 1216 additions and 151 deletions
+1 -1
View File
@@ -14,7 +14,7 @@ NAME = CUBELinux
# reject a version whose first character is not numeric, so a release spelled `CUBELinux.0.6` is
# one the machine cannot load modules for. The box's own kernel works around this the same way
# (`6.19.3-cube+`), so CUBELinux does too: the Linux base, then `-cubelinux`, then our version.
CUBELINUX_VERSION = 6.19.3-cubelinux0.6
CUBELINUX_VERSION = 6.19.3-cubelinux0.7
# *DOCUMENTATION*
# To see a list of typical targets execute "make help"
+30 -2
View File
@@ -34,6 +34,13 @@ int cubelinux_kernel_put(const __u8 *space, __u64 x, __u64 y, __u64 z,
ssize_t cubelinux_kernel_get(const __u8 *space, __u64 x, __u64 y, __u64 z,
void *buf, size_t len, __u16 *out_flags);
int cubelinux_kernel_del(const __u8 *space, __u64 x, __u64 y, __u64 z);
/*
* Capture a request this layer refused, as a classed record in the events space. The refusals worth
* recording are the ones that leave the store usable — a caller asking for something the interface
* does not offer — because those are the moments the machine said no and carried on. It takes the
* driver's write lock itself, so it must be called from here and not from inside an op.
*/
int cubelinux_kernel_capture_refusal(__u64 op, int errno, const __u8 *kind, size_t kind_len);
int cubelinux_kernel_sync(void);
/*
@@ -156,8 +163,10 @@ static long cube_args_op(unsigned int op, void __user *uargs)
* The size is the caller's, and it must be the one this kernel implements: a caller
* built against a later block would otherwise have fields silently ignored.
*/
if (args.size != sizeof(struct cube_args))
if (args.size != sizeof(struct cube_args)) {
cubelinux_kernel_capture_refusal(op, EINVAL, "bad-size", 8);
return -EINVAL;
}
switch (op) {
case CUBE_OP_PUT:
@@ -167,12 +176,15 @@ static long cube_args_op(unsigned int op, void __user *uargs)
case CUBE_OP_SYNC:
break;
default:
cubelinux_kernel_capture_refusal(op, EINVAL, "bad-op", 6);
return -EINVAL;
}
if (op == CUBE_OP_PUT || op == CUBE_OP_GET) {
if (args.len > CUBE_MAX_VALUE)
if (args.len > CUBE_MAX_VALUE) {
cubelinux_kernel_capture_refusal(op, E2BIG, "too-big", 7);
return -E2BIG;
}
if (args.len > 0) {
buf = kvmalloc(args.len, GFP_KERNEL);
if (!buf)
@@ -264,8 +276,24 @@ static long cube_enum_op(unsigned int op, void __user *uargs)
__u8 found[32];
ret = cubelinux_kernel_spaces(e.cursor, found);
/*
* The driver answers with one space written and 0, or with a negative errno. A
* positive value is neither, and it must not be read as either: the buffer was not
* written, so a caller handed this answer copies a space nobody found.
*
* This arm is the only one that advances the cursor by itself. Every other walk
* leaves the cursor where the caller put it, so a bogus return there ends the walk
* on the client's no-progress rule. Here it would hand the walk a cursor that keeps
* moving over a space that is not there — which is not a hypothetical: a kernel error
* code that had lost its sign did exactly that, and the walk served ~200,000 invented
* records a minute until the machine stopped making progress. The cause is fixed
* where it was (the driver's errno conversions, cubelinux_store.rs); this is the
* bound that keeps a defect of that shape from becoming a loop again.
*/
if (ret < 0)
return ret;
if (ret != 0)
return -EINVAL;
memcpy(e.space, found, sizeof(found));
/* An index here, not a count of records: hand back the one after this space. */
e.cursor = e.cursor + 1;
File diff suppressed because it is too large Load Diff