b8a7f615ff806bc0af618d3d8fb466e98589e268
24
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
b8a7f615ff |
read: validate the store in 96 bytes, and hold the layout with the table
The control block is 4096 bytes, and the question every read asks of it is one number — has this store changed since I last looked. That number, and the fields beside it, are in the first `CTL_SUMMED + 4` bytes of each copy. So a call that finds the store unmoved now reads two heads and nothing else; the header, the space table, the layout and the log's extent all come from the cache. The full 4096-byte read happens only when the store has actually moved. Both heads are read, and that is the whole safety of it: a write raises one copy's generation and leaves the other at the old one, so reading a single copy and finding it unchanged would call a moved store unmoved. The pair is what makes the answer true. Measured on #84, in the guest, against #83: get 0.077 -> 0.038 ms in the 20,000-record store (and 0.052 in the 500-record one) miss 0.074 -> 0.044 ms spaces 0.374 -> 0.173 ms worst case: read 0.763 ms (ceiling 5), write 6.762 ms (ceiling 50) A read is now about twice as fast, and it is finally faster in the LARGER store than in the small one — which the flatness before could not show. That is the diagnosis paying off: the per-call cost of a read was setup rather than device reads all along, and the largest single piece of that setup was reading 4 KB to compare 8 bytes. Also derives `Copy` for `HeaderV3` (six plain numbers; a cached header has to be handed out by value) and drops the search buffer `space_entry` no longer needs, now that the table it searches is in memory. |
||
|
|
1f7b9fe0de |
read: hold the space table across calls, keyed on the generation
Every coordinate read binary-searches the space table for its space's first record, and every batch of a listing reads a row of it — from the device, per probe, for data that changes only when the store is folded. It is now held in the driver and keyed on the **control block's generation**, which a fold raises along with everything else it writes, plus the image offset, since a fold flips the slot. Nothing has to invalidate it by hand: a writer that bumps the generation invalidates it, including a writer this driver never sees. A stale table cannot outlive the image it describes. Honest about what it bought: nothing measurable in the gate, which is the interesting part. The gate's store carries about a dozen spaces, so the search it replaces was three or four probes, and copying a few hundred bytes of table costs about what those probes did — get 0.078 -> 0.077 ms, spaces flat within the run-to-run noise of this bench (which is ~20%). What changes is the SHAPE: the lookup no longer scales with the number of spaces, and the box has 21 and grows. That is worth having and it is not the same thing as a measured win, so it is not claimed as one. What is still read per call, and is the larger constant: the control block (needed — it is the validation), the header, the log window, and the allocations for all three. Those are what the borrow-based version of this cache is for, and they are the next piece. |
||
|
|
aab316a9a1 |
read: take the log window instead of copying it, and name the boot class bit
A coordinate read cost 0.082 ms in a 500-record store and 0.081 ms in a 20,000-record one — flat to a microsecond while the index search does nine probe reads against fifteen. So the device reads are nearly free and the per-call cost is SETUP, and the clearest piece of setup was a copy of the log window into a second buffer on every read. `v3_get` copied it because what the log says borrows it while reading the image needs `&mut image`; that borrow is one field, so moving the field out settles it for nothing. Measured on #82: get 0.082 -> 0.078 ms, spaces 0.339 -> 0.267 ms. That is small in the GUEST, and the guest cannot show the real number: its store carries a log of a few hundred bytes, while the box's log is 132,679 bytes — so the copy this removes costs ~130 KB of memcpy and a ~130 KB allocation per read there, and only the box can measure it. Also names the store's own class bit. `BOOT_FLAGS` was a bare `1 << 7` and read as `TYPE_MASK == 10` ("data type: code") to anything applying `WordFlags` — a latent collision, not a live one, and the resolution is a declaration rather than a move: `WordFlags` is a different field and its sixteen bits are all allocated, so there is no bit to borrow, and the space axis already makes a boot record findable (reserved space 0xFC) without a class bit at all. What it needed was a vocabulary that claims it, which this comment now does. The one bit the two fields share is shared on purpose and by name: 0x0800, `SEALED_FLAG` in `cube-store-seal`, which `WordFlags` calls `ENCRYPTED` — same meaning in both places, which is the pattern rather than the exception. |
||
|
|
e968b3964e |
store: a write reads the log, not the whole device — and the log is now bounded
A put cost 60-131 ms on the box while a read cost 1-2 ms, and the reason was one call: every write path reached `append` through `device_and_layout()` → `read_image()`, which reads the WHOLE DEVICE — 100,663,296 bytes here — into a KVVec, to append ~70 bytes. `append` itself never looks at the image: it reads the log region and the control block. The read paths had been taught to read only where the image lies; the write paths never were. `AppendSource` now names what an append reads: `whole_image` (kept for the two callers that genuinely need it — a bare store, whose log runs to the end of the file, and the v1→v4 migration, which folds) or `log_head` (a device's control block plus the log's first `WAL_HEADER_LEN` bytes). `write_view()` reads exactly that, `append_mutation()` is the single entry point all four writers use — put, del, the boot record, and the misc device's write_iter — and the bare-image fallback is decided in one place instead of four. The fold still reads the whole store, because a fold rewrites it; that is the honest tail, and it is why a fold belongs on a timer rather than on the path a caller waits behind. The log is bounded by `LOG_FOLD_BYTES`, because it is on the READ path: `Addressed::open` reads `WAL_HEADER_LEN + log_used` bytes on every call, so an unfolded log is a tax on every read, not a bill paid once at the fold. The live store's control block reports log capacity 50,335,744 bytes — if it ever filled, every coordinate read would read ~48 MB before answering anything. A write that finds the log over the line folds first, through `fold_for_headroom`, and then appends to a fresh log; guarded so one append folds at most once, because a fold that did not shrink the log must not loop. The trade is named in the code: the write that crosses the line pays a bounded, rare fold instead of every reader paying an ever-larger log. The timer's comment — "the log grows at roughly a megabyte a day against a 50 MB region, and a fold rewrites the whole image, so folding more often would buy nothing and cost I/O" — is a data workload's arithmetic, where the log is a recovery artefact and nobody reads it. `space_entry`, `find` and `lower_bound` each allocated a fresh KVVec inside their binary-search loop; one buffer per search now. The boot record's one attempt becomes a bounded retry: `ensure_boot_record` claimed the boot with a swap BEFORE it tried, so a single transient failure cost the boot its record silently — which is exactly what happened on the box, where the first client of the boot could not open the store for writing. It now claims one of `BOOT_RECORD_MAX_ATTEMPTS` with a compare-exchange, and parks the count at the ceiling on success. Release moves to 6.19.3-cubelinux0.7: verify-box-preflight.sh now refuses a same-release reinstall, because the default entry boots the release being replaced. Gates, on #81: verify-enum-cost PASS a write is 1.45 ms mean / 7.46 ms worst (new ceilings: 50 ms write, 5 ms read), and a write does not grow with the store verify-file-store PASS 12 writes to a store that is a FILE, all accounted for verify-syscall PASS the store the kernel writes is byte-identical to userspace's |
||
|
|
daa6289e2d |
cube(2): the walk skips the records the cursor already returned
`v3_enum` built its batch with `seen: cursor`, and `Batch::offer` increments
before comparing (`seen += 1; if seen <= cursor { skip }`), so nothing was ever
skipped: every call re-served the space from its first record while `out_cursor`
advanced by the records returned. The contract's only end signal is a cursor that
stops moving, so the walk never ended — a listing looped inside the front-end
until the kernel's OOM killer took that process, twice, at ~6.4 GiB anon.
Every other Batch site already starts `seen` at 0, and two of them carry comments
describing exactly this trap; this one was the lone outlier. It only shows on an
**addressed** image whose space has log edits, because the merge path re-walks
from the space's start and has nothing but `seen` to skip with — the no-edits path
skips by index arithmetic (`first + cursor`) and was already right. That is why no
gate saw it: the enumeration gate's store is packed.
So each path now skips exactly once — the merge path through `offer`'s `seen`
starting at 0, the no-edits path through its arithmetic with the batch told to
skip nothing. Doing both would skip twice and lose records. The arithmetic cursor
is added saturating, because a wrapped sum would land back near the first record
and re-serve the walk: the same endless listing this exists to avoid.
Verified: a copy of the box's own addressed store walks in QEMU on this kernel to
`12 space(s), 69665 record(s)` — 69,641 image records plus 24 unfolded log edits —
and terminates, where before every batch repeated the space from its start.
|
||
|
|
da0a19786e |
cube(2): one frame for every walk — the class mask is in it
CUBE_OP_ENUM and CUBE_OP_RANGE returned `key | value_len | value`, so a caller could walk a store and not learn what any record it walked past *was*. A store could not answer `entries_flagged` from a kernel store at all, and the only way to find out was to open the device and read the format directly — which is the second reader of the format this interface exists to make unnecessary. Both walks now return the flag scan's frame: `key(24) | flags(2) | value_len | value`, with the space ahead of it when the scope is every space. One frame, one packer, one `Batch::offer`; the separate `offer_flagged` and `pack_record` are gone, and so is the reason for them to disagree. The class comes from wherever the value did: an addressed image's index entry (the binary search already read it and used to throw it away), a log entry's mask, or zero for a packed v1/v2 record, which has no field to carry one. The merge in `SpaceWalker` hands it back alongside the key and the value for the same reason — a record the log supplied carries the class its writer stamped, and dropping it there is what left a listing unable to say what it was listing. |
||
|
|
1081b5f1e2 |
cube(2): CUBE_OP_GET answers with the record's class
The last hole in the substrate: a caller could write a class through the syscall but not read one back. `CUBE_OP_GET` now fills `args.flags` from the same index entry the address came from, so learning what a record *is* costs nothing beyond a read that was going to happen. One field for both directions, because it is one thing — the class of this record. A write states it, a read learns it, and neither is a special case of the other. `find` hands the mask back with the address for the same reason: it is in the same stride the binary search already read, so wanting both does not mean searching twice. A read that finds nothing leaves 0 rather than a stale class for the caller to believe. The size of `cube_args` does not change, which matters because `size` is what says which argument block arrived. |
||
|
|
72e72fe7a8 |
cubelinux: an addressed image is readable without a log
merged_digest answered "no log" before it looked at the layout, so a bare v3/v4 image fell through to the packed walk — which starts at the v2 header length and reads index entries as record frames. A bare addressed image with no log beside it therefore read as a one-record store with an empty value. An addressed image is a complete store on its own: its index holds every record and says where each one is, so a log is an addition rather than a requirement. A packed image is not, which is why "no log" still means what it meant for v1/v2. Found by verify-image-read the moment userspace started writing v4: while every userspace image was v2 the fallback happened to be the right walk. The kernel's own v4 images always had a log header beside them, so no other gate could have caught it. |
||
|
|
29ae3b53f1 |
cube_format: a geometry encodes the version it is, not a constant
V3::encode wrote a hardcoded VERSION_V3. It had no callers, so the mistake cost nothing — and then userspace's v4 writer became the first caller, and a v4 geometry would have been published under a v3 header: every reader walking a 42-byte index 40 bytes at a time, finding a store that is silently wrong. One wrong byte, found before the first v4 image was written rather than after. |
||
|
|
9586e114e7 |
cube(2): CUBE_OP_PUT takes a class mask
The mask is the writer's and is stamped once, at the moment the record's class is known for certain; every later reader is spared re-deriving it. It means nothing to this side — which bits are which class is a vocabulary's business, and a kernel that interpreted one would be inventing a vocabulary. 0 is "no class", which is what every record written before the field existed reads as, so a caller that does not classify is not writing a special value. A store whose image is the legacy packed layout has no field to put a mask in and drops it: that layout cannot carry a class, and saying otherwise would be a lie about the bytes on disk. The field is appended, so sizeof(cube_args) grows from 80 to 88 — still distinct from the other three argument blocks, which is what the size-first dispatch depends on. |
||
|
|
71552fc157 |
cube(2): CUBE_OP_FLAG_SCAN takes a scope — one space, or every space
"Every error anywhere" and "every error here" are different questions, and a space is a hard partition, so the scope is a field rather than a widening. It is not a reserved space id because there is no such id to reserve: every 32-byte value is a legitimate space, root 0x00 and edge 0xFF…FF among them, so a sentinel would be a space somebody could name. The field occupies what was padding, which keeps sizeof unchanged — and that matters, because size is what says which argument block arrived and cube_args is exactly eight bytes wider. An every-space answer carries each frame's space, and that is not decoration: a walk's frame omits the space on the grounds that the caller named it, and this caller named none. A coordinate is meaningless without its space, so an answer that left it out would be unusable rather than merely terse. The space leads because the (space, key) pair it forms is the order records are stored in and returned in — so an every-space scan answers in exactly the order a checkpoint writes. The walk visits the space table's order and puts a space only the log writes into in its place in that same order, which is the one thing a plain walk of the table would miss entirely. A scope or mode this build does not know is refused rather than defaulted: silently answering a narrower question than the one asked is as quiet a way to be wrong as answering a wider one. |
||
|
|
41e437c07a |
cube(2): CUBE_OP_FLAG_SCAN — classify at write, retrieve by class
The store's flag field is a substrate, and this is its read half: a v4 index entry carries a 16-bit class mask, and a scan by class is a walk that reads the mask and tests it. Nothing about the merge changes — the same log overlay, the same order — so a caller pays for its own class rather than for the store. The op takes a space, a mask, a mode (any/all), a cursor and a buffer, in its own size-versioned block. A mask of zero matches nothing, because naming no class is asking no question. Records come back as key | flags(2) | value_len | value: the mask travels, since a record can carry bits the scan did not name and no other op returns a mask. Both layouts the mask can be in are read. With the record still in the log it comes from the v2 log entry; after a fold it comes from the v4 index entry. The non-indexed path is not a corner — it is the state of every store between the write that classified something and the fold, so answering it with 'nothing' would make the substrate work only after a checkpoint. Also settles what an append writes into which log, since the entry's frame has to match the header a reader frames it by: a log that already holds v1 entries keeps taking v1 entries (a device folds first — that is the v4 migration — and a bare image keeps the log it has), an empty log is framed v2 on a device and left alone on a bare image, and a log with no header is framed v2 on a device and v1 on a bare image. The bare layout is the legacy one: it has no index for a mask to be folded into, so a mask written there could only be scanned and never checkpointed — half a feature, bought by making every existing reader of that layout grow a version it cannot use. |
||
|
|
0267831b18 |
cube: fix v2 log framing — never write a v2 entry under a v1 header
The v4 migration made the kernel write v2 log entries (flags field) but append only wrote a log header when the region had no magic. A store formatted by userspace already carries a v1 header, so the kernel appended v2 entries under it and every reader framed them as v1: the length landed on the flags field, the walk stopped at the first entry, and get/fold/boot saw an empty log. Now append insists the header matches what it writes: a v1 log that still holds entries is folded into the image first (the actual v4 migration), then the entry lands in a fresh v2 log; an empty v1 log is upgraded in place. The fold is factored into fold_now, shared by append and sync. |
||
|
|
a111d7e8b2 |
cubelinux: v4 — a 16-bit class mask in the index and the log entry
The first half of the flag substrate (DESIGN-flag-vocabularies.md): the shared
format file now defines VERSION_V4, whose index entry is `key | flags(u16) |
value_off | value_len`, and WAL version 2, whose entry carries the same mask
before its length. The mask is a raw u16 — its bits are a vocabulary's business,
never the format's.
Backward compatible, and pinned as such: a v3 index entry and a v1 log entry read
as a zero mask ("no class"), which a scan treats as matching nothing, so a store
folded before the flag existed degrades to "unclassified" rather than "matches
everything". `V3::decode` accepts both versions and `index_stride()` names the
one that differs; `wal_entry` keys its stride off the log's own version byte.
The readers and writers that actually move bytes (the driver's serialize and
append, and `cube-store-raw`) are separate and are the next commit; this is the
shared definition and the arithmetic a reader derives from it.
|
||
|
|
2e5021dd83 |
cubelinux: the region walk's cursor actually advances
The gate caught it on its first run: the region walk answered the box correctly five times over — the seek and the membership test were right — but the batch re-served the same five records forever, because the cursor never skipped. The cause is a one-word difference between the two existing walk paths, and it was mine. `v3_enum` positions itself with `at = first + cursor` and so does not need `Batch.seen` to skip; the packed walk has no index to reposition and relies on `seen` instead, which is why it initialises `seen: 0`. My two region paths re-seek to the span's foot every batch — there is no `first + cursor` for a box, because a span holds records that are *not* returned, so the index position of the cursor-th match is not arithmetic — and I had set `seen: cursor`, which disables the skip and re-serves the span from its foot forever. The fix is `seen: 0`, and the gate now shows why the trap matters: an unaligned box [3,3,3]-[10,10,10] answers five records, and the record at (2,7,7) — outside the box but inside the span — is examined and rejected, not returned. |
||
|
|
88ce1bf2bf |
cubelinux: CUBE_OP_RANGE — a region walk that seeks on the box's key span
The seventh operation: `cube(2)` gains CUBE_OP_RANGE, a bounded walk of one space's records that lie in a box. Its argument block is its own (`cube_range_args`, versioned by `size` like the walk's), and the box travels as its six numbers for the reason a coordinate does — the key it has to become is the driver's business. The operation rests on the span that the shared file already owns. `key_span(lo, hi)` bounds every key in the box because the interleave is monotone on each axis, so a v3 image is **sought**: the fixed-stride index is binary-searched for the span's foot (`Addressed::lower_bound`) and read forward to its head, merging the log's edits exactly as a walk does. A packed v1/v2 image has no index to search, so its space is walked with the same span used only to stop early — and the contract is the same either way, so a caller is not told which path it got. The trap that shaped the code, and the reason it is written the way it is: **the span is a bound, not the set.** Keys of points outside the box fall inside it (Z-order amplification), so every candidate is decoded and tested against the box before it is returned — which is what `morton_decode`, the interleave's inverse, is for, now in the shared file with the same kind of hand-pinned tests the interleave has. And the cursor counts the records *in the box*, not the records of the space, because those are the records the walk returns. The over-coverage — records examined versus records returned — is the number this operation is meant to publish, and it is not wired to the caller yet: the cost gate measures it by comparing against the userspace store, which counts the same thing. |
||
|
|
72a6bbe173 |
cubelinux: the key and the span move into the shared file, and the driver stops having its own
`morton_encode` existed twice — once in the driver, once (as the curve) in cube-core — and the two
had to agree byte for byte, because one side computes a key to look a record up and the other
computes it to lay an image out. Nothing structural held that agreement: only the gates, which
would notice afterwards. That is the same shape as the cube_format.rs finding recorded in
DESIGN-coordinate-surface.md §7.4 — a shared *file* whose shared half has no callers — so this
starts paying it off where it is cheap and exact.
The interleave now lives in cube_format.rs, which both builds compile, and the driver's copy is a
call to it. It is written as plain `while` loops rather than iterator chains because the kernel
build of this file has no `std` and no `alloc`, which is the constraint the whole arrangement
exists under.
Alongside it, two things the range work needs and neither side had:
key_cmp compare two keys as the 192-bit numbers they are. The key is big-endian, so this
is the comparison a sorted index performs, and it is stated once rather than
assumed at each site.
key_span the key span a box covers, (key(lo), key(hi)) inclusive. This is the bound a seek
needs, and — the point — it needs no decomposition at all: the interleave is
monotone in every axis, so every point in the box has a key between its corners'
keys. A sorted index can be binary-searched for the foot and scanned forward to the
head. The span is a BOUND and not the set: records outside the box can have keys
inside it, which is the classic Z-order amplification and is why a caller tests
membership per candidate. Confusing the bound for the set is close to the mistake
that put "an aligned box is one run" into two doc comments.
aligned_box_is_one_run
the condition, proved in crates/cube-format's tests and pinned there and in
cube-store: a power-of-two-aligned box is exactly one contiguous run of the key iff
max(k) - min(k) <= 1 AND the axes at the top level form a prefix of (x, y, z). A
size that is not a power of two is answered `false` rather than rounded, because a
caller passing one has a bug and "no" is the useful answer.
Verified on build #53: verify-enum (the kernel's walk against userspace's, which is the check that
would catch any change in the key), verify-boot-record, verify-syscall all pass.
|
||
|
|
087c6b0e60 |
cubelinux: the kernel records its own boot
The second half of "the OS stores itself". The store was already the kernel's; what was missing was the kernel *saying* something of its own rather than a client doing it. At its first write of a boot the driver now appends one record describing the boot it is having — its own version banner, the wall-clock time, and the store device it resolved — to a reserved space, through the same append path every other mutation uses. Three things had to be decided, and each had a wrong answer that looked right: WHERE IT HOOKS. "At store init" does not exist and must not be invented: there is no init-time open, deliberately, so there is no ordering to get wrong against the block driver that provides the device. An __initcall appending a record would reintroduce exactly that ordering problem. The moment is the FIRST WRITE, which needs no ordering at all and is the semantically right one — the kernel records itself when it becomes the writer. A boot in which the kernel only reads writes no record, which is honest rather than a gap. WHICH SPACE. 0xFD was the first choice, following the convention that a reserved space is one repeated byte. 0xFD is the OS KEYSTORE — the space the kill switch exists to destroy — and writing boot records into it would have been a serious bug. Only the userspace name table catches this (format_space in crates/cube-command) because the kernel keeps no table of space names, so the table is now written down where the constant is: 0x00 root, 0xFF edges, 0xFE portal, 0xFD keystore, 0xFC boot. The record lives at (0,0,0) in 0xFC: one record, the current boot. WHAT IT SAYS. boot=<epoch seconds> device=<resolved path> kernel=<the version banner>, banner last and unquoted so everything after the final = is the kernel's own words rather than a field this code parsed. Raw epoch seconds rather than a date: rendering a calendar date in the kernel is date arithmetic, and a caller with a clock can do it without a kernel bug being the reason a timestamp is wrong. The banner comes from linux_banner and the time from ktime_get_real_ts64. WHY IT IS OFF BY DEFAULT. Not caution, but an invariant. The gates' method is that the store the kernel produces is comparable, byte for byte, with the store userspace produces from the same mutations; a record the kernel injects that the caller never asked for would turn those comparisons into non-comparisons. So it is cube_boot_record=1 on the kernel command line, parsed in C beside cube_store= for the reason that parameter is in C (this kernel's Rust cannot express a string parameter), and kernel/verify-boot-record.sh is the gate that turns it on. A failure to record is logged and never propagated: the record is worth having and is not a precondition for the caller's write. It is attempted once per boot rather than retried per write, because a store that will not take it will not take it later, and one warning is information where a stream of them is noise. |
||
|
|
78540b5687 |
cubelinux: the fold reads the layout it writes
Found on the box, minutes after the live store was folded for the first time: the second fold answered -EINVAL. `build_merged` parsed the packed layout — `space | key | len | value`, repeated — so it could read a store that had never been folded and nothing else. The first fold reads packed and writes addressed; every fold after that reads addressed, which is what a store does for the rest of its life. As written, a store could be folded exactly once and then never again — the log would grow until it filled. No gate folded twice, which is why it got this far: verify-kernel-checkpoint.sh, verify-frontend.sh and verify-enum-cost.sh each folded a packed store once. The guest's bench folds twice now, and verify-enum-cost.sh insists on the second one — "a store can be folded once and then never again, which is not a layout" — so the case is covered rather than remembered. The v3 branch reads through the shared format's own arithmetic: the space table gives each space's index range, the index gives each key and the value's place, and the log is applied over the result exactly as before. The packed path is untouched. Verified: verify-enum-cost.sh (which now folds twice and still measures 1.0x for a 40x store) and verify-kernel-checkpoint.sh (the image a fold writes holds exactly what userspace holds). |
||
|
|
d569df710d |
cubelinux: the store's format is one file, and the kernel and userspace both include it
Two build systems cannot share a crate: the kernel's Rust build compiles what is in its own module
tree, and a cargo crate is not that. So they share a *file* — `drivers/cube/cube_format.rs`, which
the driver declares with `mod` and `crates/cube-format` includes by path. There is no copy to drift
from, which is the only arrangement that cannot go stale. What happened when v3 landed and userspace
did not is the argument: an hour of `unsupported-version`, one gate red, and two implementations of
one format each believing itself.
The file holds what both sides must agree about, and nothing else — pure functions over slices, no
allocation, no I/O, no logging:
* the constants every reader and writer derives its arithmetic from,
* the header, in all three versions, with refusal rather than guessing: a reader that invents an
extent can read somebody else's bytes,
* the v3 space table and index, and the three equalities a reader computes its addresses from,
* the packed record and the log entry, the two framings a fold and a walk have to parse alike,
* CRC-32, FNV-1a, and the digest line both sides print.
The driver's copies of the constants are now aliases of the shared ones, and the two functions that
*validate* a header — `parse_header` and the v3 geometry — delegate to it, because validation is
where a second description gets believed. The readers stay where they are: this driver streams from
a file with its own buffers while userspace already holds the whole image, and that difference is
real rather than duplicated.
Nothing changed in behaviour, which is the point: verify-enum.sh, verify-enum-cost.sh,
verify-frontend.sh, verify-kernel-checkpoint.sh and verify-kernel-append.sh all pass, and the shared
crate's own tests include a kernel-written image — ten records, three spaces, 110 value bytes — read
through its arithmetic.
|
||
|
|
1db3abce93 |
cubelinux: v3 — the image carries the addresses, so a coordinate is a place
A packed list (`space | key | len | value`, repeated) cannot answer "where is this coordinate":
record N's offset is the sum of every record before it. The key gives ORDER, and order is not an
address, so no coordinate could be turned into a place and not even a binary search was possible —
the middle record's offset is just as unknowable. Every lookup walked, every batch of a listing
walked again, and a read of one record cost as much as the store (measured: 110 ms per get, 119 ms
per `spaces`, 1,206 ms for a 6,127-record listing).
v3 puts the addresses in the image:
[header 46] magic, version, curve, extents, counts
[space table: space_count x 48] space | first index | records
[index: record_count x 40] key | value offset | value length
[values: packed, in index order]
A lookup is now a binary search over arithmetic addresses (`index_off + i * 40`); a listing starts
at its space's first index entry and streams, reading a span of values per batch; asking which
spaces exist is the space table; a spatial range is a contiguous run of index entries. An index
entry is 40 bytes against the packed frame's 64, so the image also gets smaller.
Measured by `verify-enum-cost.sh` between a 500-record store and a 20,000-record one:
listing the same 10-record space: 3.22 ms -> 3.27 ms (1.0x, store grew 40x)
asking which spaces exist: 3.25 ms -> 3.08 ms (0.9x)
reading a coordinate that is there: 1.07 ms -> 1.25 ms (1.2x)
reading one that is not: 1.19 ms -> 1.22 ms (1.0x)
Nothing scales with the store any more, which is the premise: knowing a coordinate is what lets you
reach it.
Two bugs found on the way, both of the kind only a real image shows. The index was written with a
4-byte value length (`Entry` carries it in a u32) while the header's arithmetic and both readers
assumed 8 — every entry four bytes short, the last of them overlapping the values. And the digest's
`bytes` figure added the layout's own header length, so two layouts of one store digested
differently; it is now the records' logical size, which is the same number either way.
Verified: verify-kernel-append.sh, verify-kernel-checkpoint.sh, verify-enum.sh, verify-enum-cost.sh
and verify-frontend.sh all pass, and the workspace tests pass — including a new fixture test in
cube-store-raw that reads a 700-byte image the kernel itself wrote, because a hand-built image
cannot catch a misunderstanding shared by the builder.
|
||
|
|
56665a303e |
cubelinux: read the store where it lies, instead of rebuilding it per call
Every read began by reading the whole device into kernel memory, parsing every
record of every space into a merged pool, and heapsorting the lot — then throwing
all of it away. So a `get` of one record cost as much as listing the store, every
batch of a listing paid it again, and the coordinate bought nothing mechanically:
knowing where a record is did not make reaching it any cheaper. Measured on the
workhorse: one `get` 110 ms, one `spaces` call 119 ms, one 6,127-record listing
1,206 ms.
The records are already in the order the coordinate computes — a checkpoint writes
them by (space, key) — so the kernel now walks them where they lie: zero-copy
values, no pool, no sort, and the log (small, and the only thing that can override
the image) consulted as an overlay. It reads the image and the log window rather
than the whole device, which is four images wide.
* `get` — the log's newest word on the coordinate, else the image walked to it.
* `enum` — one merge pass over two sorted sequences: the space's image records
and its log edits.
* `spaces` — from a table of where each space starts, cached by the control
block's generation. That table is one entry per space, not per record, and a
checkpoint is the only thing that invalidates it.
* `put`/`del`/`sync` — unchanged: a write is a log append, and the fold still
writes the whole sorted store.
Verified: verify-enum.sh (a 3,006-record listing identical to userspace's, record
for record) and verify-frontend.sh (socket -> front-end -> cube(2), including a
listing larger than one batch and the store the front-end leaves being
digest-identical to userspace's) both pass.
Not yet what it should be: listing a 10-record space in a 20,000-record store
still costs about six times what it costs in a 500-record one, so the space table
is not taking effect as intended and the lookup is not yet independent of store
size. That belongs in the image's layout — a coordinate cannot be turned into a
byte offset in a variable-length packed list, because a record's position is the
sum of every value before it — and the fix is a format one, not a caching one.
|
||
|
|
b9a7b8d8f3 |
cubelinux: cube(2) walks the store — CUBE_OP_ENUM and CUBE_OP_SPACES
Enumeration was the one operation the interface did not have, and the one a capability interface owes an explanation for: a listing is the opposite of "knowing a coordinate is the authorisation to use it". So the walk is bounded and explicit. A caller names the space, holds a cursor, and gets as many whole records as fit in the buffer it offered, packed `key(24) | value_len(u32) | value` in the store's own order. SPACES walks the distinct spaces that hold a record, one per call; range is that walk with the region as a filter, applied by the caller rather than by a second operation in the kernel. Its own argument block, versioned by its own size: SYSCALL_DEFINE2 peeks `size` and routes — 80 bytes is the coordinate block, 64 is this one. An interface that cannot grow has to be replaced, and this one grows by being given a new block. The end of a walk is the cursor alone. A batch holds as many whole records as fit, so it is FULL only when a record lands on the boundary and most batches come back short; a caller that reads "the buffer was not filled" as "the space is exhausted" truncates its listing to the first batch and cannot tell. ENUM answers a finished walk with no records and the cursor unmoved; SPACES answers it with -ENOENT, because the space after the last one is not there. That rule is in the uapi header because it is a contract detail, not an implementation one. The kernel caches no merged index: each call walks the records in order, asks the log — small, and the only thing that can override one — what the winner is, and skips what the cursor has already covered. O(records) a call, tens of milliseconds for the live store, which is worth more than a cache every write would have to invalidate. |
||
|
|
ed5ff764ba |
cubelinux: the release starts with a digit, because module tooling demands it
The store's sort no longer lives on the kernel stack (heapsort, see the parent commit's gate), and the release is now 6.19.3-cubelinux0.6 instead of CUBELinux.0.6 — depmod/mkinitramfs reject a version whose first character is not numeric, so the literal name cannot be uname -r. The box's own kernel uses the same shape (6.19.3-cube+); this is the same concession. |