Compare commits

...
26 Commits
Author SHA1 Message Date
CUBELinux 35fcb6c7fe store: a fold kept a deletion only when the record had no class
A fold turns the log's effects into the image's by keeping, per coordinate, the *last* entry in the
merged order — that is how "the later write wins" is implemented. The order came from a derive on
`Entry`, which compares fields in declaration order, and `flags` was declared before `seq`. So the
tie-break was not arrival at all: it was the class mask.

A `del` is written with no class (0). A **sealed** record is written with `0x0800`. For the same
coordinate that put the removal *before* the record it removed, so the record won and the fold wrote
it back — a deleted record returning from the fold, with its old value, on any store whose writes are
sealed. Found on the box on 2026-09-25: put → present, del → "(not found)", `cube-fold` → "folded the
log into the image" (generation 2647 → 2648, log emptied), get → the value, back. It reproduced
through a sealed front-end and an unsealed one, and the live store still holds the test records it
resurrected.

Why it hid: an unsealed store writes both entries with flags 0, the tie falls through to `seq`, and
the order is right — which is exactly what the guest gate did before this, so the gate agreed with a
bug it could not see. The shape that fails is the shape the machine uses.

`Entry` now orders by `(space, key, seq, …)` explicitly, with the reason written beside it, because
"the fields happen to be declared in this order" is what went wrong. The userspace store cannot have
this bug and does not: it applies the log into a `HashMap`, where a removal *is* the removal of the
key, so no ordering decides anything.
2026-09-25 03:35:52 -04:00
CUBELinux 2586c2ed0d store: the kernel captures its own events, as classed records
Every part of this was in place except the point of it: a record can carry a 16-bit class,
CUBE_OP_PUT stamps one and CUBE_OP_FLAG_SCAN retrieves by class, the vocabulary is a userspace
convention, and the kernel already stamped its own boot record with the boot class. What was
missing was a kernel that captures *events* — one that writes about what happened to it rather
than only what a caller asked for. Until now the kernel's account of saying no was a line in
dmesg, which is not somewhere a later reader can ask.

Three events, in the events space (0xFB), each classed by the vocabulary it belongs to:

  bad-op         ERROR   a request the interface does not offer, refused and recorded
  clock-set      BOOT    the epoch was provisional, and the record was rewritten
  boot-record-late ERROR|BOOT  the record only landed on a retry — the one that matters most,
                         because that defect's effect was invisible in the store

The write is bounded twice (EVENT_BUDGET, 32 a boot; REFUSAL_BUDGET, 8 among them) and that is
not tidiness: a kernel that appends a record per event can turn a storm of refused requests into
a storm of writes, which this store has already met from the other side when a walk with no store
behind it served ~200,000 invented records a minute. Past the ceiling the kernel logs and stops.

The refusal capture is exported for the syscall layer to call (`cubelinux_kernel_capture_refusal`)
because that is where the refusals that leave the store usable happen — a bad size, an op that
does not exist, a value past the maximum. A store that cannot be read at all is the one refusal
the kernel cannot record into itself, and that is stated rather than papered over.
2026-09-25 01:06:14 -04:00
CUBELinux 3d5d213c75 store: a boot record whose epoch can be trusted, not just read
The record the kernel writes about its own boot carried `boot=<epoch>` and nothing else — a
reading taken at the first write of the boot, which is exactly when this machine's clock is
least likely to be right: it arrives from a night powered off tens of seconds out, and on the
boot before it, 10h 27m behind. Nothing in the record said the time was provisional.

Two fields carry the account now. `uptime` comes from the monotonic clock, so `boot - uptime`
is the instant the boot began whatever the wall clock was doing. `clock=raw|set` says whether
the reading predates a correction. And the kernel acts on the difference: the wall clock can be
set from anywhere and the monotonic clock cannot be set at all, so a wall clock that moves
without it is a correction and nothing else — when that is observed, the next write rewrites
the record at the same coordinate with the corrected time, bounded like the first write.

Measured in the guest (verify-boot-record, which now moves the clock forward an hour mid-boot):
boot=1790308025 uptime=3 clock=raw, then the same coordinate at boot=1790311625 uptime=3
clock=set — exactly the hour that was moved. The gate also refuses a record whose epoch
precedes its own uptime, and one that claims to have been written long after the boot began.

The retry half of this was already in the tree (a count rather than a swap that cannot
un-claim itself, from e968b3964); this is the epoch half, and the record's own "Next" list
named both.
2026-09-25 00:09:09 -04:00
surface-camera-build 8c24a7ff0d store: an acknowledged write is durable again, in the shape nothing was flushing
`verify-kernel-append` asks the question a log exists to answer: four writes are acknowledged, the
VM is killed with `SIGKILL` with no shutdown at all, and the log must still hold them. It is red on
`#95` and on `#87` — the kernel this box runs — and green on the kernel built at 18:41, and the one
commit between them is e0218ec96: the collapse to a single `fsync` per append, where the entry goes
down unsynced and the control write that counts it carries the flush for both.

That collapse is right where its sentence is true, and the sentence assumes there is a control block
to write. *A store with no control block has no count write*, so nothing on the append path was
flushed at all — and the caller was told a write it may never get back. That is not an exotic shape:
it is what a store the CLI builds looks like, and it is the shape that gate's own device has.
Measured, same device, same four writes, the host's copy of the device read while the guest is still
hung:

    no control block    the log region holds its header and zeros where the entry should be, and
                        the fold returns the base store: records=2
    a formatted device  the entry is on the device, and the fold reproduces the userspace store
                        exactly: records=6, fnv1a64=6b679d39597a62b3

and a pre-collapse kernel keeps the entry in *both* shapes, because the entry's own `fsync` is what
made it durable there.

So the fix is not the revert. `append` takes the entry's own `fsync` exactly when no later write in
the same operation will carry one — `layout.control.is_some()`, one boolean, both branches gated by
the two shapes above. The box's store is a formatted file with a control block, so it keeps the
measured 4.96 ms win; a bare store pays its entry's flush, which is what it always did before the
collapse. The count still goes second, so a crash can still lose an unacknowledged mutation rather
than count one that is not there.

The gate was wrong in the same way the first diagnosis was, and is fixed with the kernel: it ran its
crash check over one device shape, which is the shape that made the collapse look safe everywhere. It
now runs it over both — a store with no control block and a formatted device — folds each the way
userspace reads it, and compares both against the same userspace store, so "it survived" means the
same thing for each. It also stops SIGKILLing the `timeout` wrapper rather than QEMU: SIGKILL cannot
be forwarded, so every run of this gate left a 512 MiB VM spinning in `pause()` for the rest of the
wall, twice found beside a gate that measures latency.

Built as #96. The wall, whole: 26 of 26.
2026-09-23 23:16:31 -04:00
surface-camera-build 8db279a4d8 cube(2): a refusal with two minus signs is not a refusal — and the walk it hung
The walk with no store behind it served ~200,000 invented records a minute and never ended.
Half of that was fixed and proven in fc057820f: the driver asks whether the store can be read
before either walk op answers anything, and it refuses. The client was still handed

    spaces returned 0, len=0, cursor=1

— success, no space, a cursor one further on — so it copied the space it was handed out of a
buffer the kernel never wrote, and with the cursor moving by itself the rule that ends every
other walk here, *no progress is the only end signal*, had nothing to fire on.

The remaining half was not in the C arm, and both candidate explanations recorded there are
wrong: the `ret < 0` test IS on the path, and nothing overwrote the answer. One boot at
loglevel=7 says what crosses the boundary instead:

    cubelinux: the store is not readable; refusing to answer     (x138 in 40 s)
    walk: spaces returned 0, len=0, cursor=1                     (the client, told success)

The guard fires and the client is told success, so the value is not negative. On that path the
driver's only return is `-(e.to_errno() as i32)`, and `kernel::error::Error` IS the kernel error
code: `from_errno(-2) == ENOENT`, `to_errno()` is documented as "the kernel error code", and the
API's own conversion of a `Result` to a C result — `kernel::from_result` — writes
`T::from(e.to_errno() as i16)` with no negation, as every other driver in this tree does. The
second minus sign made every refusal in this file a *positive* number, and a positive return is
what the C arm's `ret < 0`, the client's own wrapper, and the walk arms' `ret == 0` all read as
success.

38 sites in cubelinux_store.rs were shaped `-(e.to_errno() as i32)` / `as isize`. All 38 now
return `e.to_errno()`, which is what makes them refusals. Nothing else about them changed.

Why this became a *loop* in the space walk and nowhere else: CUBE_OP_SPACES is the one arm that
advances the cursor itself. Every other walk arm leaves the cursor where the caller put it, so a
bogus return there ends the walk on the client's no-progress rule — which is why one defect was
invisible at every other verb for as long as it existed. That arm now refuses a return that is
neither 0 ("here is a space") nor negative (a refusal), so a defect of this shape cannot be read
as a space again.

One more correction in the same class, found while proving the above. A store that cannot be
*opened* answered -ENOENT, and -ENOENT is this interface's own end-of-walk signal — "no such
space; the walk is finished" — so a walk over a store on a disk whose driver had not loaded read
exactly like a walk over an empty store, which is what the first benchmark boot was.
`store_file()` now answers ENODEV when the store device is not there. A store that is not there
is not an empty store.

Proven against the reproduction, in the guest, on the bench initramfs:

    cube_store=/dev/null        -> "walk: spaces failed: Invalid argument", 0 records, 5 s, boot finishes
    cube_store=/nowhere/x.img   -> "walk: spaces failed: No such device",    0 records, 5 s, boot finishes
    before:                       364,994 invented record lines in the 90 s the instrument allowed

The three checks are `kernel/verify-no-store.sh`, a gate on the wall now, and the same three
inside the benchmark rehearsal — which is where this defect was found, and where they were
warnings while it was open.

The diagnostic prints this was hunted with come off in the same commit: the `cube_store=
resolved to` line in cube_syscall.c, and the per-call "not readable" warning in both walk ops.
The refusal is the return value, and the walk's own transcript is where a reader learns what
happened; a message per call is a diagnostic, not the interface. The guard itself stays, and
where to find it is written down at its definition rather than implied.

Built as #95, which is what the box now has installed: the machine boots it on its next reboot.
2026-09-23 22:42:36 -04:00
surface-camera-build fc057820fa store: refuse a store that cannot be read — and the proof that the refusal is not the whole defect
The guard is `store_readable()`: four bytes at offset zero, asked before either walk
op answers anything. Proven against the reproduction, which is the reproduction
from the record — `cube_store=/dev/null cubelinux.enum=1`:

  cubelinux: cube_store= resolved to /dev/null                 (the token was honoured)
  cubelinux: the store is not readable; refusing to answer      (the kernel refuses)

and the client is *still* handed `spaces returned 0, len=0, cursor=1`, so it copies
an unfilled space and walks 199,000 invented records. That is the whole defect in
one transcript: the driver refuses and the syscall boundary reports success anyway.

So this commit fixes and proves half of it. What remains is `cube_syscall.c`'s
CUBE_OP_SPACES arm, where the driver's -EINVAL becomes a 0 with a cursor advanced by
one — either its `ret < 0` test is not on the path the client takes, or the answer is
overwritten before the copy-out. Nothing above or below that needs touching.

Three earlier guesses are recorded as wrong rather than deleted: the packed fallback
and the magic guard written for it are never reached (probed, proven), and
`read_exact_at` already rejects a short read. The guard here is the first one that
fires.

The device-path print in cube_syscall.c stays on purpose: it is what turned four
builds of inference into one line of fact.

Built as #93, deliberately NOT installed — installing it alone would leave the walk
still looping. The machine runs #87.
2026-09-23 22:02:22 -04:00
surface-camera-build 9fb4169e52 store: refuse a device that does not carry the packed image's magic (inert, NOT installed)
This is the guard for one of the two doors on the walk-with-no-store defect:
"neither control block decodes" is not evidence of a packed image, so the packed
reader is allowed only for a device that begins with the image header's own magic —
the question the userspace raw reader has always asked and the kernel did not.

It builds clean and it does NOT stop the reproduction: `cube_store=/dev/null
cubelinux.enum=1` still serves invented records at the same rate. So the guard is
not on the path that answers, and it is committed for that reason stated plainly
rather than installed: this machine stays on #87, which is measured, instead of
moving to a kernel whose only change is one that has not been shown to do anything.

Kept rather than reverted because the check is right on its own terms — an
addressed store begins with CUBS, a bare packed image with CUBE — and because the
next attempt should begin by asking whether this guard is even reached before
writing another one.
2026-09-23 21:15:53 -04:00
surface-camera-build 4cfdbfe63b store: a get on the first slice of a chain returns the whole value
The arrangement vocabulary's second half — the producer and the userspace
reader have existed since 0244-ish, and this is the kernel joining them for a
caller who does not know it holds a chain. Bounded by the same 65,536 slices the
userspace reader uses and by MAX_CHAIN_BYTES (CUBE_MAX_VALUE, 16 MiB: a longer
chain could not be handed over in one call, so reading on would be work whose
answer nobody can receive). A hole, a slice that claims to continue and does not,
and a second beginning are each -EIO rather than a guessed end.

A joined read reports START|END, because the flags describe the bytes handed over
rather than the slice they came from — which is also why the userspace reader
needed no change: it already stops at END_RECORD.

It does NOT join a sealed chain, and that is a fact about the format: a sealed
slice is an envelope with its own header and nonce, so N of them concatenated are
not one openable value, and joining them belongs to whoever holds the key. When
it cannot join (sealed, or a packed image with no index) it returns the slice with
the record's own flags — CONTINUATION without END_RECORD says "this is a slice,
not the whole value" — so the caller is told rather than misled.

Reading those bits at all was safe because nothing uses them, and that was
measured before the code was written: over the live store's 69,749 records,
flag-scan finds 1 record in 0x0080, 137 in 0x0800, and zero in every one of
0x0001, 0x0002, 0x0003, 0x0004, 0x0008, 0x0010, 0x0020, 0x0040, 0x0100, 0x0200.

Gated by kernel/verify-chain.sh: PASS, seven checks.
2026-09-23 19:59:57 -04:00
surface-camera-build e0218ec96b store: one fsync per append, not two — and torn-tail is the gate that decided it
A durable write was 8.2 ms and it was two fsyncs to ext4: the log entry, then the control
block that counts it. Measured on the store's own filesystem, one `pwrite`+`fsync` is
4.321 ms and append's shape of two is 8.630 ms, against 8.202 ms for the whole kernel put —
so the write path was fsyncs and almost nothing else, and one of them was flushing data the
next one covers.

The order stays. A mutation is still written as entry-then-count, because that order is what
makes a crash lose an unacknowledged mutation rather than count one that is not there. What
goes is the *second durability boundary*: the entry is `kernel_write`-ordered before the
control write, and the control write's `fsync` flushes what precedes it in the same file.

What that gives up, stated rather than implied: the two writes are no longer independently
durable, so a crash inside the control write's commit can in principle leave the control
block counting bytes whose entry did not fully reach the media. That is a torn tail — and
this store already has a gate for one:

    verify-torn-tail   PASS  "userspace on the same bytes, and appended over the tear
                              rather than after it"
    verify-syscall     PASS  the store the kernel writes is byte-identical to userspace's
    verify-file-store  PASS  twelve writes to a store that is a FILE, all accounted for

So the change was made against the gate that tests the failure mode rather than against an
argument about ext4's journal, and it is reversible on its own: revert this commit, rebuild,
and the two-fsync order returns. Half of every write, for a property the format already
handles.
2026-09-23 19:25:34 -04:00
surface-camera-build b3dc55392e read: the log window is cached with the table — the last device read a read made for nothing
The log is the one thing in the store that changes without a fold, which is why it was the
one thing still read on every call even after the validation, header, layout and space table
were held across calls. But that is the same argument in reverse: the cache is keyed on the
control block's generation, and a write — the only thing that changes the log — raises it. So
an unchanged generation IS a log that cannot have changed, and a hit can serve it from memory
without asking the device at all.

Measured, this one is neutral in the gate: get 0.038 -> 0.044 ms, inside the run-to-run noise
of this bench. That is expected rather than disappointing — the gate's store carries a log of
a few hundred bytes, so a read of it costs about what the copy from the cache does. The number
this change is for is on the box, whose log is 143,317 bytes: that read and its allocation
were ~15-30 us of every coordinate read there, and the box is where it will show.

With this the objective's list is closed: the control-block validation, the header, the space
table and the log window are all held across calls; the space-table search allocates nothing
and the index search allocates once per search rather than once per probe; the unfolded log is
bounded by LOG_FOLD_BYTES so per-call work has a ceiling; and the gate fails on a worst case
and a ceiling rather than on a ratio.
2026-09-23 18:41:36 -04:00
surface-camera-build b8a7f615ff read: validate the store in 96 bytes, and hold the layout with the table
The control block is 4096 bytes, and the question every read asks of it is one number — has
this store changed since I last looked. That number, and the fields beside it, are in the
first `CTL_SUMMED + 4` bytes of each copy. So a call that finds the store unmoved now reads
two heads and nothing else; the header, the space table, the layout and the log's extent all
come from the cache. The full 4096-byte read happens only when the store has actually moved.

Both heads are read, and that is the whole safety of it: a write raises one copy's generation
and leaves the other at the old one, so reading a single copy and finding it unchanged would
call a moved store unmoved. The pair is what makes the answer true.

Measured on #84, in the guest, against #83:
    get   0.077 -> 0.038 ms in the 20,000-record store  (and 0.052 in the 500-record one)
    miss  0.074 -> 0.044 ms
    spaces 0.374 -> 0.173 ms
    worst case: read 0.763 ms (ceiling 5), write 6.762 ms (ceiling 50)
A read is now about twice as fast, and it is finally faster in the LARGER store than in the
small one — which the flatness before could not show. That is the diagnosis paying off: the
per-call cost of a read was setup rather than device reads all along, and the largest single
piece of that setup was reading 4 KB to compare 8 bytes.

Also derives `Copy` for `HeaderV3` (six plain numbers; a cached header has to be handed out by
value) and drops the search buffer `space_entry` no longer needs, now that the table it
searches is in memory.
2026-09-23 18:37:23 -04:00
surface-camera-build 1f7b9fe0de read: hold the space table across calls, keyed on the generation
Every coordinate read binary-searches the space table for its space's first record, and
every batch of a listing reads a row of it — from the device, per probe, for data that
changes only when the store is folded. It is now held in the driver and keyed on the
**control block's generation**, which a fold raises along with everything else it writes,
plus the image offset, since a fold flips the slot. Nothing has to invalidate it by hand: a
writer that bumps the generation invalidates it, including a writer this driver never sees.
A stale table cannot outlive the image it describes.

Honest about what it bought: nothing measurable in the gate, which is the interesting part.
The gate's store carries about a dozen spaces, so the search it replaces was three or four
probes, and copying a few hundred bytes of table costs about what those probes did — get
0.078 -> 0.077 ms, spaces flat within the run-to-run noise of this bench (which is ~20%).
What changes is the SHAPE: the lookup no longer scales with the number of spaces, and the
box has 21 and grows. That is worth having and it is not the same thing as a measured win,
so it is not claimed as one.

What is still read per call, and is the larger constant: the control block (needed — it is
the validation), the header, the log window, and the allocations for all three. Those are
what the borrow-based version of this cache is for, and they are the next piece.
2026-09-23 18:05:55 -04:00
surface-camera-build aab316a9a1 read: take the log window instead of copying it, and name the boot class bit
A coordinate read cost 0.082 ms in a 500-record store and 0.081 ms in a 20,000-record
one — flat to a microsecond while the index search does nine probe reads against
fifteen. So the device reads are nearly free and the per-call cost is SETUP, and the
clearest piece of setup was a copy of the log window into a second buffer on every read.
`v3_get` copied it because what the log says borrows it while reading the image needs
`&mut image`; that borrow is one field, so moving the field out settles it for nothing.

Measured on #82: get 0.082 -> 0.078 ms, spaces 0.339 -> 0.267 ms. That is small in the
GUEST, and the guest cannot show the real number: its store carries a log of a few
hundred bytes, while the box's log is 132,679 bytes — so the copy this removes costs
~130 KB of memcpy and a ~130 KB allocation per read there, and only the box can measure
it.

Also names the store's own class bit. `BOOT_FLAGS` was a bare `1 << 7` and read as
`TYPE_MASK == 10` ("data type: code") to anything applying `WordFlags` — a latent
collision, not a live one, and the resolution is a declaration rather than a move:
`WordFlags` is a different field and its sixteen bits are all allocated, so there is no
bit to borrow, and the space axis already makes a boot record findable (reserved space
0xFC) without a class bit at all. What it needed was a vocabulary that claims it, which
this comment now does. The one bit the two fields share is shared on purpose and by name:
0x0800, `SEALED_FLAG` in `cube-store-seal`, which `WordFlags` calls `ENCRYPTED` — same
meaning in both places, which is the pattern rather than the exception.
2026-09-23 17:53:38 -04:00
surface-camera-build e968b3964e store: a write reads the log, not the whole device — and the log is now bounded
A put cost 60-131 ms on the box while a read cost 1-2 ms, and the reason was one call:
every write path reached `append` through `device_and_layout()` → `read_image()`, which
reads the WHOLE DEVICE — 100,663,296 bytes here — into a KVVec, to append ~70 bytes.
`append` itself never looks at the image: it reads the log region and the control
block. The read paths had been taught to read only where the image lies; the write
paths never were.

`AppendSource` now names what an append reads: `whole_image` (kept for the two callers
that genuinely need it — a bare store, whose log runs to the end of the file, and the
v1→v4 migration, which folds) or `log_head` (a device's control block plus the log's
first `WAL_HEADER_LEN` bytes). `write_view()` reads exactly that, `append_mutation()`
is the single entry point all four writers use — put, del, the boot record, and the
misc device's write_iter — and the bare-image fallback is decided in one place instead
of four. The fold still reads the whole store, because a fold rewrites it; that is the
honest tail, and it is why a fold belongs on a timer rather than on the path a caller
waits behind.

The log is bounded by `LOG_FOLD_BYTES`, because it is on the READ path: `Addressed::open`
reads `WAL_HEADER_LEN + log_used` bytes on every call, so an unfolded log is a tax on
every read, not a bill paid once at the fold. The live store's control block reports
log capacity 50,335,744 bytes — if it ever filled, every coordinate read would read
~48 MB before answering anything. A write that finds the log over the line folds first,
through `fold_for_headroom`, and then appends to a fresh log; guarded so one append
folds at most once, because a fold that did not shrink the log must not loop. The trade
is named in the code: the write that crosses the line pays a bounded, rare fold instead
of every reader paying an ever-larger log. The timer's comment — "the log grows at
roughly a megabyte a day against a 50 MB region, and a fold rewrites the whole image, so
folding more often would buy nothing and cost I/O" — is a data workload's arithmetic,
where the log is a recovery artefact and nobody reads it.

`space_entry`, `find` and `lower_bound` each allocated a fresh KVVec inside their
binary-search loop; one buffer per search now.

The boot record's one attempt becomes a bounded retry: `ensure_boot_record` claimed the
boot with a swap BEFORE it tried, so a single transient failure cost the boot its record
silently — which is exactly what happened on the box, where the first client of the boot
could not open the store for writing. It now claims one of `BOOT_RECORD_MAX_ATTEMPTS`
with a compare-exchange, and parks the count at the ceiling on success.

Release moves to 6.19.3-cubelinux0.7: verify-box-preflight.sh now refuses a same-release
reinstall, because the default entry boots the release being replaced.

Gates, on #81:
  verify-enum-cost    PASS  a write is 1.45 ms mean / 7.46 ms worst (new ceilings: 50 ms
                            write, 5 ms read), and a write does not grow with the store
  verify-file-store   PASS  12 writes to a store that is a FILE, all accounted for
  verify-syscall      PASS  the store the kernel writes is byte-identical to userspace's
2026-09-23 17:32:05 -04:00
surface-camera-build daa6289e2d cube(2): the walk skips the records the cursor already returned
`v3_enum` built its batch with `seen: cursor`, and `Batch::offer` increments
before comparing (`seen += 1; if seen <= cursor { skip }`), so nothing was ever
skipped: every call re-served the space from its first record while `out_cursor`
advanced by the records returned. The contract's only end signal is a cursor that
stops moving, so the walk never ended — a listing looped inside the front-end
until the kernel's OOM killer took that process, twice, at ~6.4 GiB anon.

Every other Batch site already starts `seen` at 0, and two of them carry comments
describing exactly this trap; this one was the lone outlier. It only shows on an
**addressed** image whose space has log edits, because the merge path re-walks
from the space's start and has nothing but `seen` to skip with — the no-edits path
skips by index arithmetic (`first + cursor`) and was already right. That is why no
gate saw it: the enumeration gate's store is packed.

So each path now skips exactly once — the merge path through `offer`'s `seen`
starting at 0, the no-edits path through its arithmetic with the batch told to
skip nothing. Doing both would skip twice and lose records. The arithmetic cursor
is added saturating, because a wrapped sum would land back near the first record
and re-serve the walk: the same endless listing this exists to avoid.

Verified: a copy of the box's own addressed store walks in QEMU on this kernel to
`12 space(s), 69665 record(s)` — 69,641 image records plus 24 unfolded log edits —
and terminates, where before every batch repeated the space from its start.
2026-09-23 01:32:55 -04:00
CUBELinux e53ae04033 cubelinux: hold the store lock across the whole op, not inside append
Initialising STORE_FILE stopped the oops and exposed what it had been hiding: twelve
concurrent puts all reported `ok` and one write survived. The lock guarded the file
handle, not the store — it is taken, the handle is fetched or cached, and released
before the caller does anything with it. Sharing a struct file * is safe, so that is
all the handle needs, and it is not enough for the store.

Widening it inside append is also not enough, and this commit exists because the first
attempt did exactly that and changed nothing. append receives the layout as an argument,
and the caller read it from device_and_layout() *before* calling: twelve writers each
read log_used = 0, then queued on a lock inside append, then each appended at the same
offset against the layout they had already read. Measured with the lock in append: one
68-byte entry, twelve `ok`s. The lock has to be held from before the read.

So STORE_OP is taken by the four write ops — put, del, sync and the device write — and
everything under them runs with it held: ensure_boot_record, device_and_layout, append,
fold_now. None of those may take it again; a kernel mutex is not reentrant, and append
folds and then calls itself, while fold_now is reachable both from inside append and on
its own. That is why the lock sits at the ops rather than in the writers.

verify-file-store.sh MODE=race, twelve concurrent puts: 816 bytes of log = 12 x 68, and
control generation 13 = 1 + 12, where the previous kernel produced 68 and 2.

Regression: MODE=seq still exact; verify-boot-record passes (it is the path this
changes most, since ensure_boot_record now runs under the lock); verify-kernel-append
passes, four acknowledged writes surviving a SIGKILL with no shutdown.
2026-09-22 23:32:18 -04:00
CUBELinux 7938966d6a cubelinux: initialise STORE_FILE, the mutex that was never initialised
STORE_FILE is declared `unsafe(uninit) static ... Mutex<Option<StoreFile>> = None`
and nothing ever called STORE_FILE.init(). Its sibling SPACE_STARTS is initialised
explicitly in CubeStoreModule::init; this one was simply missed.

The failure mode is why it hid. A zeroed mutex satisfies the uncontended fast path
— the count reads 0, which means "unlocked" — so one writer at a time works and
nothing looks wrong. The first CONTENDED lock takes __mutex_lock_slowpath, which
splices the task into the mutex's wait list; that list is uninitialised, so its head
is NULL and the splice stores through it. A write to address 0 in kernel mode, after
which the task returns with interrupts disabled and preemption held — a machine that
cannot panic, log, or recover.

Found by verify-file-store.sh MODE=race: twelve concurrent puts to a store that is a
file on the root filesystem oopsed in cubelinux_store::store_file, while twelve
sequential ones passed with exact log accounting (816 bytes of log, generation 13).
The same signature is on the box, where the store had several writers and the box
froze with no panic despite panic=30.
2026-09-22 21:10:12 -04:00
surface-camera-build da0a19786e cube(2): one frame for every walk — the class mask is in it
CUBE_OP_ENUM and CUBE_OP_RANGE returned `key | value_len | value`, so a caller
could walk a store and not learn what any record it walked past *was*. A store
could not answer `entries_flagged` from a kernel store at all, and the only way
to find out was to open the device and read the format directly — which is the
second reader of the format this interface exists to make unnecessary.

Both walks now return the flag scan's frame: `key(24) | flags(2) | value_len |
value`, with the space ahead of it when the scope is every space. One frame,
one packer, one `Batch::offer`; the separate `offer_flagged` and `pack_record`
are gone, and so is the reason for them to disagree.

The class comes from wherever the value did: an addressed image's index entry
(the binary search already read it and used to throw it away), a log entry's
mask, or zero for a packed v1/v2 record, which has no field to carry one. The
merge in `SpaceWalker` hands it back alongside the key and the value for the
same reason — a record the log supplied carries the class its writer stamped,
and dropping it there is what left a listing unable to say what it was listing.
2026-09-22 02:20:31 -04:00
surface-camera-build 1081b5f1e2 cube(2): CUBE_OP_GET answers with the record's class
The last hole in the substrate: a caller could write a class through the syscall
but not read one back. `CUBE_OP_GET` now fills `args.flags` from the same index
entry the address came from, so learning what a record *is* costs nothing beyond
a read that was going to happen.

One field for both directions, because it is one thing — the class of this
record. A write states it, a read learns it, and neither is a special case of the
other. `find` hands the mask back with the address for the same reason: it is in
the same stride the binary search already read, so wanting both does not mean
searching twice. A read that finds nothing leaves 0 rather than a stale class for
the caller to believe.

The size of `cube_args` does not change, which matters because `size` is what
says which argument block arrived.
2026-09-22 01:34:23 -04:00
surface-camera-build 72e72fe7a8 cubelinux: an addressed image is readable without a log
merged_digest answered "no log" before it looked at the layout, so a bare
v3/v4 image fell through to the packed walk — which starts at the v2 header
length and reads index entries as record frames. A bare addressed image with
no log beside it therefore read as a one-record store with an empty value.

An addressed image is a complete store on its own: its index holds every
record and says where each one is, so a log is an addition rather than a
requirement. A packed image is not, which is why "no log" still means what it
meant for v1/v2.

Found by verify-image-read the moment userspace started writing v4: while
every userspace image was v2 the fallback happened to be the right walk. The
kernel's own v4 images always had a log header beside them, so no other gate
could have caught it.
2026-09-22 01:02:39 -04:00
surface-camera-build 29ae3b53f1 cube_format: a geometry encodes the version it is, not a constant
V3::encode wrote a hardcoded VERSION_V3. It had no callers, so the mistake
cost nothing — and then userspace's v4 writer became the first caller, and a
v4 geometry would have been published under a v3 header: every reader walking
a 42-byte index 40 bytes at a time, finding a store that is silently wrong.
One wrong byte, found before the first v4 image was written rather than after.
2026-09-22 01:02:29 -04:00
surface-camera-build 9586e114e7 cube(2): CUBE_OP_PUT takes a class mask
The mask is the writer's and is stamped once, at the moment the record's
class is known for certain; every later reader is spared re-deriving it. It
means nothing to this side — which bits are which class is a vocabulary's
business, and a kernel that interpreted one would be inventing a vocabulary.

0 is "no class", which is what every record written before the field existed
reads as, so a caller that does not classify is not writing a special value.
A store whose image is the legacy packed layout has no field to put a mask in
and drops it: that layout cannot carry a class, and saying otherwise would be
a lie about the bytes on disk.

The field is appended, so sizeof(cube_args) grows from 80 to 88 — still
distinct from the other three argument blocks, which is what the size-first
dispatch depends on.
2026-09-22 00:37:09 -04:00
surface-camera-build 71552fc157 cube(2): CUBE_OP_FLAG_SCAN takes a scope — one space, or every space
"Every error anywhere" and "every error here" are different questions, and
a space is a hard partition, so the scope is a field rather than a widening.
It is not a reserved space id because there is no such id to reserve: every
32-byte value is a legitimate space, root 0x00 and edge 0xFF…FF among them,
so a sentinel would be a space somebody could name. The field occupies what
was padding, which keeps sizeof unchanged — and that matters, because size is
what says which argument block arrived and cube_args is exactly eight bytes
wider.

An every-space answer carries each frame's space, and that is not decoration:
a walk's frame omits the space on the grounds that the caller named it, and
this caller named none. A coordinate is meaningless without its space, so an
answer that left it out would be unusable rather than merely terse. The space
leads because the (space, key) pair it forms is the order records are stored
in and returned in — so an every-space scan answers in exactly the order a
checkpoint writes.

The walk visits the space table's order and puts a space only the log writes
into in its place in that same order, which is the one thing a plain walk of
the table would miss entirely. A scope or mode this build does not know is
refused rather than defaulted: silently answering a narrower question than the
one asked is as quiet a way to be wrong as answering a wider one.
2026-09-22 00:31:56 -04:00
surface-camera-build 41e437c07a cube(2): CUBE_OP_FLAG_SCAN — classify at write, retrieve by class
The store's flag field is a substrate, and this is its read half: a v4
index entry carries a 16-bit class mask, and a scan by class is a walk that
reads the mask and tests it. Nothing about the merge changes — the same log
overlay, the same order — so a caller pays for its own class rather than for
the store.

The op takes a space, a mask, a mode (any/all), a cursor and a buffer, in
its own size-versioned block. A mask of zero matches nothing, because naming
no class is asking no question. Records come back as
key | flags(2) | value_len | value: the mask travels, since a record can
carry bits the scan did not name and no other op returns a mask.

Both layouts the mask can be in are read. With the record still in the log
it comes from the v2 log entry; after a fold it comes from the v4 index
entry. The non-indexed path is not a corner — it is the state of every store
between the write that classified something and the fold, so answering it
with 'nothing' would make the substrate work only after a checkpoint.

Also settles what an append writes into which log, since the entry's frame
has to match the header a reader frames it by: a log that already holds v1
entries keeps taking v1 entries (a device folds first — that is the v4
migration — and a bare image keeps the log it has), an empty log is framed
v2 on a device and left alone on a bare image, and a log with no header is
framed v2 on a device and v1 on a bare image. The bare layout is the legacy
one: it has no index for a mask to be folded into, so a mask written there
could only be scanned and never checkpointed — half a feature, bought by
making every existing reader of that layout grow a version it cannot use.
2026-09-22 00:08:19 -04:00
surface-camera-build 0267831b18 cube: fix v2 log framing — never write a v2 entry under a v1 header
The v4 migration made the kernel write v2 log entries (flags field) but
append only wrote a log header when the region had no magic. A store
formatted by userspace already carries a v1 header, so the kernel appended
v2 entries under it and every reader framed them as v1: the length landed
on the flags field, the walk stopped at the first entry, and get/fold/boot
saw an empty log.

Now append insists the header matches what it writes: a v1 log that still
holds entries is folded into the image first (the actual v4 migration),
then the entry lands in a fresh v2 log; an empty v1 log is upgraded in
place. The fold is factored into fold_now, shared by append and sync.
2026-09-21 23:41:02 -04:00
surface-camera-build a111d7e8b2 cubelinux: v4 — a 16-bit class mask in the index and the log entry
The first half of the flag substrate (DESIGN-flag-vocabularies.md): the shared
format file now defines VERSION_V4, whose index entry is `key | flags(u16) |
value_off | value_len`, and WAL version 2, whose entry carries the same mask
before its length. The mask is a raw u16 — its bits are a vocabulary's business,
never the format's.

Backward compatible, and pinned as such: a v3 index entry and a v1 log entry read
as a zero mask ("no class"), which a scan treats as matching nothing, so a store
folded before the flag existed degrades to "unclassified" rather than "matches
everything". `V3::decode` accepts both versions and `index_stride()` names the
one that differs; `wal_entry` keys its stride off the log's own version byte.

The readers and writers that actually move bytes (the driver's serialize and
append, and `cube-store-raw`) are separate and are the next commit; this is the
shared definition and the arithmetic a reader derives from it.
2026-09-21 22:38:43 -04:00
5 changed files with 2421 additions and 334 deletions
+1 -1
View File
@@ -14,7 +14,7 @@ NAME = CUBELinux
# reject a version whose first character is not numeric, so a release spelled `CUBELinux.0.6` is
# one the machine cannot load modules for. The box's own kernel works around this the same way
# (`6.19.3-cube+`), so CUBELinux does too: the Linux base, then `-cubelinux`, then our version.
CUBELINUX_VERSION = 6.19.3-cubelinux0.6
CUBELINUX_VERSION = 6.19.3-cubelinux0.7
# *DOCUMENTATION*
# To see a list of typical targets execute "make help"
+87 -18
View File
@@ -30,6 +30,10 @@ pub const VERSION_V1: u8 = 1;
pub const VERSION_V2: u8 = 2;
/// The addressed format: the same, plus a space table, a fixed-size index, and packed values.
pub const VERSION_V3: u8 = 3;
/// v3, plus a 16-bit class mask in each index entry, so a scan by flag is a seek rather than a
/// walk. The mask is the flag substrate (DESIGN-flag-vocabularies.md): a raw `u16` whose bits are a
/// vocabulary's business, written at put time and read by `CUBE_OP_FLAG_SCAN`.
pub const VERSION_V4: u8 = 4;
pub const HEADER_LEN_V1: usize = 6;
pub const HEADER_LEN_V2: usize = 6 + 8 + 8;
@@ -39,19 +43,28 @@ pub const HEADER_LEN_V3: usize = 4 + 1 + 1 + 8 + 8 + 8 + 8 + 8;
pub const SPACE_ID_LEN: usize = 32;
pub const RAW_KEY_LEN: usize = 24;
/// The class mask's width: one `u16` per record.
pub const FLAGS_LEN: usize = 2;
/// A packed record's fixed part: `space | key | value_len(u64)`, with the value behind it.
pub const RECORD_FIXED: usize = SPACE_ID_LEN + RAW_KEY_LEN + 8;
/// A v3 index entry: `key | value_off(u64) | value_len(u64)`.
pub const INDEX_ENTRY: usize = RAW_KEY_LEN + 8 + 8;
/// A v4 index entry: `key | flags(u16) | value_off(u64) | value_len(u64)`.
pub const INDEX_ENTRY_V4: usize = RAW_KEY_LEN + FLAGS_LEN + 8 + 8;
/// A v3 space-table row: `space | first index(u64) | records(u64)`.
pub const SPACE_ENTRY: usize = SPACE_ID_LEN + 8 + 8;
/// The log's framing, which the fold and the readers both parse.
pub const WAL_MAGIC: &[u8; 4] = b"CUBW";
/// The original log entry: no class mask.
pub const WAL_VERSION: u8 = 1;
/// The flagged log entry: the entry carries a `u16` class mask before its length.
pub const WAL_VERSION_V2: u8 = 2;
pub const WAL_HEADER_LEN: usize = 6;
/// `op(1) | crc(4) | space(32) | key(24) | len(4)`, with the value behind it.
pub const ENTRY_FIXED: usize = 1 + 4 + SPACE_ID_LEN + RAW_KEY_LEN + 4;
/// v2's entry: the same, with a `flags(2)` field before the length.
pub const ENTRY_FIXED_V2: usize = 1 + 4 + SPACE_ID_LEN + RAW_KEY_LEN + FLAGS_LEN + 4;
/// `CUBE_OP_PUT`: an entry that stores a value.
pub const WAL_OP_WRITE: u8 = 1;
@@ -74,7 +87,7 @@ impl Header {
pub fn records_off(&self) -> usize {
if self.version == VERSION_V1 {
HEADER_LEN_V1
} else if self.version == VERSION_V3 {
} else if self.version == VERSION_V3 || self.version == VERSION_V4 {
HEADER_LEN_V3
} else {
HEADER_LEN_V2
@@ -134,7 +147,7 @@ pub fn parse_header(bytes: &[u8]) -> Result<Header, Bad> {
record_count: Some(record_count),
})
}
VERSION_V3 => {
VERSION_V3 | VERSION_V4 => {
if bytes.len() < HEADER_LEN_V3 {
return Err(Bad::Magic);
}
@@ -153,6 +166,9 @@ pub fn parse_header(bytes: &[u8]) -> Result<Header, Bad> {
/// The v3 tables' geometry: where the index is, where the values start, and how many of each.
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub struct V3 {
/// The version this geometry belongs to: `VERSION_V3` or `VERSION_V4`, which differ only in
/// the index entry's stride (the class mask adds two bytes).
pub version: u8,
pub record_count: u64,
pub space_count: u64,
pub index_off: u64,
@@ -160,16 +176,27 @@ pub struct V3 {
}
impl V3 {
/// The width of one index entry for this geometry's version.
pub fn index_stride(&self) -> usize {
if self.version == VERSION_V4 {
INDEX_ENTRY_V4
} else {
INDEX_ENTRY
}
}
/// Read the tables' geometry, refusing anything that does not add up.
///
/// The three equalities here are the whole point of the format: a reader computes a record's
/// place as `index_off + i * INDEX_ENTRY`, and that is only a place if the index really starts
/// place as `index_off + i * stride`, and that is only a place if the index really starts
/// there and really is that wide.
pub fn decode(bytes: &[u8]) -> Result<Self, Bad> {
if bytes.len() < HEADER_LEN_V3 || &bytes[0..4] != MAGIC || bytes[4] != VERSION_V3 {
let version = bytes.get(4).copied().ok_or(Bad::Magic)?;
if bytes.len() < HEADER_LEN_V3 || &bytes[0..4] != MAGIC || (version != VERSION_V3 && version != VERSION_V4) {
return Err(Bad::Magic);
}
let geometry = V3 {
version,
record_count: le_u64(bytes, 14),
space_count: le_u64(bytes, 22),
index_off: le_u64(bytes, 30),
@@ -180,7 +207,7 @@ impl V3 {
return Err(Bad::Extent);
}
let table_end = HEADER_LEN_V3 as u64 + geometry.space_count * SPACE_ENTRY as u64;
let index_end = geometry.index_off + geometry.record_count * INDEX_ENTRY as u64;
let index_end = geometry.index_off + geometry.record_count * geometry.index_stride() as u64;
if geometry.index_off != table_end
|| geometry.values_off != index_end
|| geometry.values_off > image_bytes
@@ -192,12 +219,18 @@ impl V3 {
/// Write the header this geometry describes. `image_bytes` is filled in by the caller's
/// arithmetic, since only it knows how long the values are.
///
/// The version written is the geometry's own rather than a constant. A v4 geometry has a
/// 42-byte index, and a header saying v3 would tell every reader to walk it 40 bytes at a time
/// — one wrong byte that silently mis-addresses the whole store. This had no callers when it
/// was written, so the mistake cost nothing; it has one now, and finding it before the first
/// v4 image is written is the whole value of fixing it here.
pub fn encode(&self, out: &mut [u8], curve: u8, image_bytes: u64) -> Result<usize, Bad> {
if out.len() < HEADER_LEN_V3 {
return Err(Bad::Extent);
}
out[0..4].copy_from_slice(MAGIC);
out[4] = VERSION_V3;
out[4] = self.version;
out[5] = curve;
out[6..14].copy_from_slice(&image_bytes.to_le_bytes());
out[14..22].copy_from_slice(&self.record_count.to_le_bytes());
@@ -229,14 +262,24 @@ impl V3 {
if at >= self.record_count {
return None;
}
let off = self.index_off as usize + at as usize * INDEX_ENTRY;
if off + INDEX_ENTRY > bytes.len() {
let stride = self.index_stride();
let off = self.index_off as usize + at as usize * stride;
if off + stride > bytes.len() {
return None;
}
// The flag field exists only in v4; a v3 entry reads as a zero mask, which is honest —
// "no class" — and matches nothing in a scan.
let flags = if self.version == VERSION_V4 {
le_u16(bytes, off + RAW_KEY_LEN)
} else {
0
};
let vo = off + RAW_KEY_LEN + if self.version == VERSION_V4 { FLAGS_LEN } else { 0 };
let entry = IndexEntry {
key: bytes[off..off + RAW_KEY_LEN].try_into().ok()?,
value_off: le_u64(bytes, off + RAW_KEY_LEN),
value_len: le_u64(bytes, off + RAW_KEY_LEN + 8),
flags,
value_off: le_u64(bytes, vo),
value_len: le_u64(bytes, vo + 8),
};
// A value that is not inside the image is a truncated image, not an empty one.
if entry.value_off < self.values_off
@@ -258,10 +301,12 @@ impl V3 {
}
}
/// A v3 index entry: the key, and where its value lies.
/// An addressed index entry: the key, its class mask, and where its value lies.
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub struct IndexEntry<'a> {
pub key: &'a [u8; RAW_KEY_LEN],
/// The class mask (v4), or zero (v3, where no mask was written).
pub flags: u16,
pub value_off: u64,
pub value_len: u64,
}
@@ -304,16 +349,25 @@ pub struct WalEntry<'a> {
pub space: &'a [u8; SPACE_ID_LEN],
pub key: &'a [u8; RAW_KEY_LEN],
pub op: u8,
/// The class mask (WAL version 2), or zero (version 1, where no mask was written).
pub flags: u16,
pub value: &'a [u8],
pub next: usize,
}
/// Read one log entry at `off`.
/// Read one log entry at `off`. The log slice includes its header, so the version at `log[4]`
/// decides the entry's stride: version 2 carries a two-byte class mask before the length.
///
/// The checksum covers space, key, length and value, so a torn tail is stopped at rather than
/// applied — the rule the fold and every reader share.
/// The checksum covers space, key, the mask, length and value, so a torn tail is stopped at rather
/// than applied — the rule the fold and every reader share.
pub fn wal_entry(log: &[u8], off: usize) -> Option<WalEntry<'_>> {
if off + ENTRY_FIXED > log.len() {
let version = log.get(4).copied().unwrap_or(WAL_VERSION);
let (fixed, len_at) = if version == WAL_VERSION_V2 {
(ENTRY_FIXED_V2, RAW_KEY_LEN + FLAGS_LEN + SPACE_ID_LEN + 5)
} else {
(ENTRY_FIXED, 61)
};
if off + fixed > log.len() {
return None;
}
let op = log[off];
@@ -321,19 +375,25 @@ pub fn wal_entry(log: &[u8], off: usize) -> Option<WalEntry<'_>> {
return None;
}
let crc = le_u32(log, off + 1);
let len = le_u32(log, off + 61) as usize;
let frame_end = off + ENTRY_FIXED + len;
let len = le_u32(log, off + len_at) as usize;
let frame_end = off + fixed + len;
if frame_end > log.len() {
return None;
}
if crc32(&log[off + 5..frame_end]) != crc {
return None;
}
let flags = if version == WAL_VERSION_V2 {
le_u16(log, off + SPACE_ID_LEN + RAW_KEY_LEN + 5)
} else {
0
};
Some(WalEntry {
space: log[off + 5..off + 37].try_into().ok()?,
key: log[off + 37..off + 61].try_into().ok()?,
op,
value: &log[off + ENTRY_FIXED..frame_end],
flags,
value: &log[off + fixed..frame_end],
next: frame_end,
})
}
@@ -357,6 +417,15 @@ pub fn le_u32(bytes: &[u8], at: usize) -> u32 {
u32::from_le_bytes(w)
}
/// A little-endian `u16` at `at`, with the same rule.
pub fn le_u16(bytes: &[u8], at: usize) -> u16 {
let mut w = [0u8; 2];
if at + 2 <= bytes.len() {
w.copy_from_slice(&bytes[at..at + 2]);
}
u16::from_le_bytes(w)
}
/// CRC-32 (IEEE 802.3), bitwise.
///
/// A corruption check, not a security check: it catches a torn write or a flipped bit, and says
+103 -7
View File
@@ -30,10 +30,17 @@
* encoding, which must produce exactly the key a userspace reader decodes.
*/
int cubelinux_kernel_put(const __u8 *space, __u64 x, __u64 y, __u64 z,
const void *value, size_t len);
const void *value, size_t len, __u16 flags);
ssize_t cubelinux_kernel_get(const __u8 *space, __u64 x, __u64 y, __u64 z,
void *buf, size_t len);
void *buf, size_t len, __u16 *out_flags);
int cubelinux_kernel_del(const __u8 *space, __u64 x, __u64 y, __u64 z);
/*
* Capture a request this layer refused, as a classed record in the events space. The refusals worth
* recording are the ones that leave the store usable — a caller asking for something the interface
* does not offer — because those are the moments the machine said no and carried on. It takes the
* driver's write lock itself, so it must be called from here and not from inside an op.
*/
int cubelinux_kernel_capture_refusal(__u64 op, int errno, const __u8 *kind, size_t kind_len);
int cubelinux_kernel_sync(void);
/*
@@ -56,6 +63,15 @@ int cubelinux_kernel_range(const __u8 *space,
__u64 cursor, void *buf, size_t cap,
__u64 *out_len, __u64 *out_cursor);
/*
* The flag scan (CUBE_OP_FLAG_SCAN). The mask and its mode travel as plain numbers: which bits mean
* what is a vocabulary's business, and the kernel never interprets one — it compares masks, which is
* what lets a new vocabulary attach without a format change.
*/
int cubelinux_kernel_flag_scan(const __u8 *space, __u16 mask, __u16 mode,
__u32 every_space, __u64 cursor, void *buf, size_t cap,
__u64 *out_len, __u64 *out_cursor);
/* The store device path, resolved from the `cube_store=` boot parameter at boot. */
const char *cubelinux_store_device(void);
@@ -147,8 +163,10 @@ static long cube_args_op(unsigned int op, void __user *uargs)
* The size is the caller's, and it must be the one this kernel implements: a caller
* built against a later block would otherwise have fields silently ignored.
*/
if (args.size != sizeof(struct cube_args))
if (args.size != sizeof(struct cube_args)) {
cubelinux_kernel_capture_refusal(op, EINVAL, "bad-size", 8);
return -EINVAL;
}
switch (op) {
case CUBE_OP_PUT:
@@ -158,12 +176,15 @@ static long cube_args_op(unsigned int op, void __user *uargs)
case CUBE_OP_SYNC:
break;
default:
cubelinux_kernel_capture_refusal(op, EINVAL, "bad-op", 6);
return -EINVAL;
}
if (op == CUBE_OP_PUT || op == CUBE_OP_GET) {
if (args.len > CUBE_MAX_VALUE)
if (args.len > CUBE_MAX_VALUE) {
cubelinux_kernel_capture_refusal(op, E2BIG, "too-big", 7);
return -E2BIG;
}
if (args.len > 0) {
buf = kvmalloc(args.len, GFP_KERNEL);
if (!buf)
@@ -179,14 +200,16 @@ static long cube_args_op(unsigned int op, void __user *uargs)
break;
}
ret = cubelinux_kernel_put(args.coord.space, args.coord.x,
args.coord.y, args.coord.z, buf, args.len);
args.coord.y, args.coord.z, buf, args.len,
args.flags);
break;
case CUBE_OP_GET: {
ssize_t got;
got = cubelinux_kernel_get(args.coord.space, args.coord.x,
args.coord.y, args.coord.z, buf, args.len);
args.coord.y, args.coord.z, buf, args.len,
&args.flags);
if (got < 0) {
ret = got;
break;
@@ -253,8 +276,24 @@ static long cube_enum_op(unsigned int op, void __user *uargs)
__u8 found[32];
ret = cubelinux_kernel_spaces(e.cursor, found);
/*
* The driver answers with one space written and 0, or with a negative errno. A
* positive value is neither, and it must not be read as either: the buffer was not
* written, so a caller handed this answer copies a space nobody found.
*
* This arm is the only one that advances the cursor by itself. Every other walk
* leaves the cursor where the caller put it, so a bogus return there ends the walk
* on the client's no-progress rule. Here it would hand the walk a cursor that keeps
* moving over a space that is not there — which is not a hypothetical: a kernel error
* code that had lost its sign did exactly that, and the walk served ~200,000 invented
* records a minute until the machine stopped making progress. The cause is fixed
* where it was (the driver's errno conversions, cubelinux_store.rs); this is the
* bound that keeps a defect of that shape from becoming a loop again.
*/
if (ret < 0)
return ret;
if (ret != 0)
return -EINVAL;
memcpy(e.space, found, sizeof(found));
/* An index here, not a count of records: hand back the one after this space. */
e.cursor = e.cursor + 1;
@@ -361,7 +400,62 @@ static long cube_range_op(void __user *uargs)
}
/*
* One syscall, three argument blocks. They share a prefix — `size`, then `op` — so the size the
* The flag scan — CUBE_OP_FLAG_SCAN — which travels in `struct cube_flag_scan_args`.
*
* The walks' shape again, because it is the walks' contract: a batch fills the caller's buffer and
* returns how much was used plus the cursor to pass next; a record that does not fit ends the
* batch; a record that cannot fit in any buffer the caller offered comes back as -ERANGE with `len`
* saying what it would need.
*
* The one thing a caller must know beyond the walk's rules: the cursor counts the records that
* **matched**, not the records examined, because the records the mask rejected are not answers.
*/
static long cube_flag_scan_op(void __user *uargs)
{
struct cube_flag_scan_args f;
void *buf = NULL;
long ret = 0;
u64 out_len = 0, out_cursor = 0;
if (copy_from_user(&f, uargs, sizeof(f)))
return -EFAULT;
if (f.size != sizeof(struct cube_flag_scan_args) || f.op != CUBE_OP_FLAG_SCAN)
return -EINVAL;
if (f.mode != CUBE_FLAG_ANY && f.mode != CUBE_FLAG_ALL)
return -EINVAL;
if (f.every_space != CUBE_SPACE_ONE && f.every_space != CUBE_SPACE_EVERY)
return -EINVAL;
if (f.len > CUBE_MAX_WALK)
return -E2BIG;
if (f.len > 0) {
buf = kvmalloc(f.len, GFP_KERNEL);
if (!buf)
return -ENOMEM;
}
ret = cubelinux_kernel_flag_scan(f.space, f.mask, f.mode, f.every_space,
f.cursor, buf, f.len, &out_len, &out_cursor);
if (ret == 0) {
if (out_len > 0 && copy_to_user((void __user *)f.value, buf, out_len))
ret = -EFAULT;
f.len = out_len;
f.cursor = out_cursor;
if (copy_to_user(uargs, &f, sizeof(f)))
ret = -EFAULT;
} else if (ret == -ERANGE) {
/* Nothing was written; `len` now says how much one record needs. */
f.len = out_len;
if (copy_to_user(uargs, &f, sizeof(f)))
ret = -EFAULT;
}
kvfree(buf);
return ret;
}
/*
* One syscall, four argument blocks. They share a prefix — `size`, then `op` — so the size the
* caller declares is what says which one arrived. That is the whole point of putting `size`
* first: an interface that cannot grow has to be replaced, and this one grows by being given a
* new block with a new size.
@@ -378,5 +472,7 @@ SYSCALL_DEFINE2(cube, unsigned int, op, void __user *, uargs)
return cube_enum_op(op, uargs);
if (size == sizeof(struct cube_range_args))
return cube_range_op(uargs);
if (size == sizeof(struct cube_flag_scan_args))
return cube_flag_scan_op(uargs);
return -EINVAL;
}
File diff suppressed because it is too large Load Diff
+103 -6
View File
@@ -34,6 +34,23 @@ struct cube_args {
* out: on -ERANGE, the bytes that would be needed;
* on success for CUBE_OP_GET, the bytes read.
*/
__u16 flags; /* in: CUBE_OP_PUT — the class mask to stamp on the record;
* out: CUBE_OP_GET — the mask the record carries.
*
* One field for both directions because it is one thing: the
* class of this record. A write states it, a read learns it,
* and neither is a special case of the other. 0 is "no class",
* which is what every record written before the field existed
* reads as — so not classifying is not writing a special value,
* and a read that found nothing leaves 0 rather than a stale
* class for the caller to believe.
*
* A store whose image is the legacy packed layout has no field
* to put a mask in, and drops it: that layout cannot carry a
* class, and saying otherwise would be a lie about the bytes on
* disk. A read from such a store answers 0 for the same reason.
*/
__u16 reserved; /* must be 0 */
};
#define CUBE_OP_PUT 1 /* store bytes at a coordinate */
@@ -43,15 +60,21 @@ struct cube_args {
#define CUBE_OP_ENUM 5 /* walk the records of a space, in batches */
#define CUBE_OP_SPACES 6 /* walk the spaces that hold records */
#define CUBE_OP_RANGE 7 /* walk the records of a space that lie in a box */
#define CUBE_OP_FLAG_SCAN 8 /* walk the records of a space whose class mask matches */
/*
* The walk's argument block: its own block rather than a wider `cube_args`, because it needs a
* cursor and a buffer, and the coordinate would otherwise be both an input and an output.
*
* Records are packed as `key(24) | value_len(u32, little-endian) | value`, in the store's own
* order — space first, then key — which is the order a checkpoint writes them and the order the
* userspace store returns them, so a kernel listing and a userspace listing can be compared
* directly. The space is not repeated per record: the caller named it.
* Records are packed as `key(24) | flags(2, little-endian) | value_len(u32, little-endian) |
* value`, in the store's own order — space first, then key — which is the order a checkpoint writes
* them and the order the userspace store returns them, so a kernel listing and a userspace listing
* can be compared directly. The space is not repeated per record: the caller named it.
*
* The class mask is in the frame because a listing has to be able to say what a record *is*, and a
* caller that cannot see it has to open the store itself to find out — which is the second reader
* of the format this interface exists to make unnecessary. A record whose layout carries no mask (a
* packed v1/v2 image) reads as zero: "no class", the same answer every other reader gives.
*
* A walk ends when the cursor stops moving, and that is the only end signal: a batch holds as
* many whole records as fit, so most batches come back short, and reading a short batch as the
@@ -80,8 +103,9 @@ struct cube_enum_args {
* replaced.
*
* A region is a box, inclusive on both corners. Records come back packed exactly as CUBE_OP_ENUM
* packs them — `key(24) | value_len(u32, little-endian) | value`, in the store's own order — so a
* kernel region answer and a userspace one can be compared byte for byte.
* packs them — `key(24) | flags(2, little-endian) | value_len(u32, little-endian) | value`, in the
* store's own order — so a kernel region answer and a userspace one can be compared byte for
* byte.
*
* The cursor is a COUNT OF RECORDS ALREADY RETURNED, and it is the end-of-walk signal for the same
* reason as the walk's: a batch holds as many whole records as fit, so most batches come back
@@ -89,6 +113,9 @@ struct cube_enum_args {
* counts — a walk counts the records of the space, a region walk counts the records *in the box*,
* because those are the records it returns.
*
* Records come back in the same frame the walk uses, class mask included:
* `key(24) | flags(2) | value_len(u32) | value`.
*
* Implementation, because it is what makes this cheap: **the kernel seeks.** The box's two corner
* keys bound every key inside it (`cube_format::key_span`), so a v3 image's sorted index is
* binary-searched for the foot of that span and read forward to its head. The span is a BOUND, not
@@ -111,4 +138,74 @@ struct cube_range_args {
*/
};
/*
* The flag scan's argument block: `CUBE_OP_FLAG_SCAN`, the class-mask half of the store's
* classification substrate (DESIGN-flag-vocabularies.md). Its own block for the walks' reason — it
* needs a cursor and a buffer — and versioned by `size` like the other three.
*
* `mask` is a raw 16-bit class mask and this interface does not interpret a bit of it: the
* vocabulary that owns the bits (events today, sealing / lineage / lifecycle later) is the only
* thing that knows what they mean. `mode` says how to read the mask:
*
* CUBE_FLAG_ANY the record shares at least one bit with `mask` — "every error"
* CUBE_FLAG_ALL the record carries every bit of `mask` — "every Wi-Fi error"
*
* A `mask` of zero matches *nothing*, not everything: naming no class is asking no question, and a
* scan that answered a walk's worth of records to an empty question would be a walk wearing a
* scan's hat.
*
* The cursor counts **matches already returned**, as the region walk's counts records in its box,
* and for the same reason: the records examined before a match are not matches, so the index
* position of the cursor-th match is not arithmetic. A batch holds as many whole records as fit, so
* most batches come back short; reading a short batch as the end truncates the answer. A finished
* scan answers with no records and the cursor unchanged.
*
* Records come back as `key(24) | flags(2, little-endian) | value_len(u32, little-endian) | value`
* — the walk's frame with the class mask in it. The mask travels because a scan's answer has to say
* what class each record answered with: a record can carry bits beyond the one asked for, and no
* other operation returns a mask.
*
* Under CUBE_SPACE_EVERY the frame gains a leading `space(32)`, because it has to: a walk's frame
* leaves the space out on the grounds that the caller named it, and a caller who named no space
* cannot be told which one a record came from any other way. A coordinate is meaningless without
* its space, so an every-space answer that omitted it would be unusable rather than merely terse.
* The space leads because the (space, key) pair it forms is the order records are stored in and
* returned in.
*/
struct cube_flag_scan_args {
__u32 size; /* sizeof(struct cube_flag_scan_args) as the caller built it */
__u32 op; /* CUBE_OP_FLAG_SCAN */
__u8 space[32]; /* in: the space to scan, or ignored with CUBE_SPACE_EVERY */
__u16 mask; /* in: the class mask to match; 0 matches nothing */
__u16 mode; /* in: CUBE_FLAG_ANY or CUBE_FLAG_ALL */
__u32 every_space; /* in: CUBE_SPACE_ONE or CUBE_SPACE_EVERY — see below */
__u64 cursor; /* in: 0 to start, or what the last call returned;
* out: what to pass next — see the end-of-scan rule above
*/
__u64 value; /* user pointer: where to put the records */
__u64 len; /* in: the buffer's capacity;
* out: bytes written, or on -ERANGE what would be needed
*/
};
#define CUBE_FLAG_ANY 0 /* the record shares at least one bit with the mask */
#define CUBE_FLAG_ALL 1 /* the record carries every bit of the mask */
/*
* The scan's scope. A space is a hard partition, so this is a choice between two different
* questions and never a filter that can be widened by accident:
*
* CUBE_SPACE_ONE the records of `space` — "this class here"
* CUBE_SPACE_EVERY the records of every space — "this class anywhere"
*
* It is a field and not a reserved space id because there is no such id to reserve: every 32-byte
* value is a legitimate space (root `0x00` and edge `0xFF…FF` are both in use), so a sentinel would
* be a space somebody could name. The field occupies what was padding, so `sizeof` is unchanged —
* which matters, because `size` is what says which argument block arrived and `cube_args` is
* exactly eight bytes wider. A caller that zeroes its block (and every caller does) gets
* CUBE_SPACE_ONE, which is the narrower question.
*/
#define CUBE_SPACE_ONE 0 /* scan only `space` */
#define CUBE_SPACE_EVERY 1 /* scan every space */
#endif /* _UAPI_LINUX_CUBE_H */