8c24a7ff0dbeaa0038e5b43083396e001ab618fd
9
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
e968b3964e |
store: a write reads the log, not the whole device — and the log is now bounded
A put cost 60-131 ms on the box while a read cost 1-2 ms, and the reason was one call: every write path reached `append` through `device_and_layout()` → `read_image()`, which reads the WHOLE DEVICE — 100,663,296 bytes here — into a KVVec, to append ~70 bytes. `append` itself never looks at the image: it reads the log region and the control block. The read paths had been taught to read only where the image lies; the write paths never were. `AppendSource` now names what an append reads: `whole_image` (kept for the two callers that genuinely need it — a bare store, whose log runs to the end of the file, and the v1→v4 migration, which folds) or `log_head` (a device's control block plus the log's first `WAL_HEADER_LEN` bytes). `write_view()` reads exactly that, `append_mutation()` is the single entry point all four writers use — put, del, the boot record, and the misc device's write_iter — and the bare-image fallback is decided in one place instead of four. The fold still reads the whole store, because a fold rewrites it; that is the honest tail, and it is why a fold belongs on a timer rather than on the path a caller waits behind. The log is bounded by `LOG_FOLD_BYTES`, because it is on the READ path: `Addressed::open` reads `WAL_HEADER_LEN + log_used` bytes on every call, so an unfolded log is a tax on every read, not a bill paid once at the fold. The live store's control block reports log capacity 50,335,744 bytes — if it ever filled, every coordinate read would read ~48 MB before answering anything. A write that finds the log over the line folds first, through `fold_for_headroom`, and then appends to a fresh log; guarded so one append folds at most once, because a fold that did not shrink the log must not loop. The trade is named in the code: the write that crosses the line pays a bounded, rare fold instead of every reader paying an ever-larger log. The timer's comment — "the log grows at roughly a megabyte a day against a 50 MB region, and a fold rewrites the whole image, so folding more often would buy nothing and cost I/O" — is a data workload's arithmetic, where the log is a recovery artefact and nobody reads it. `space_entry`, `find` and `lower_bound` each allocated a fresh KVVec inside their binary-search loop; one buffer per search now. The boot record's one attempt becomes a bounded retry: `ensure_boot_record` claimed the boot with a swap BEFORE it tried, so a single transient failure cost the boot its record silently — which is exactly what happened on the box, where the first client of the boot could not open the store for writing. It now claims one of `BOOT_RECORD_MAX_ATTEMPTS` with a compare-exchange, and parks the count at the ceiling on success. Release moves to 6.19.3-cubelinux0.7: verify-box-preflight.sh now refuses a same-release reinstall, because the default entry boots the release being replaced. Gates, on #81: verify-enum-cost PASS a write is 1.45 ms mean / 7.46 ms worst (new ceilings: 50 ms write, 5 ms read), and a write does not grow with the store verify-file-store PASS 12 writes to a store that is a FILE, all accounted for verify-syscall PASS the store the kernel writes is byte-identical to userspace's |
||
|
|
ed5ff764ba |
cubelinux: the release starts with a digit, because module tooling demands it
The store's sort no longer lives on the kernel stack (heapsort, see the parent commit's gate), and the release is now 6.19.3-cubelinux0.6 instead of CUBELinux.0.6 — depmod/mkinitramfs reject a version whose first character is not numeric, so the literal name cannot be uname -r. The box's own kernel uses the same shape (6.19.3-cube+); this is the same concession. |
||
|
|
3ee59dafff |
CUBELinux.0.6: cube(2) — the coordinate interface
The write path was proven but unreachable: its operations lived behind a device node. This is the interface the decision chose (DESIGN-cube-interface.md) — one syscall number, an opcode, and a versioned argument block, with operations that are the verbs the command language already defines: put, get, del, sync. long cube(unsigned int op, struct cube_args __user *args) `size` comes first and is checked, because syscall numbers are permanent and an interface that cannot grow would have to be replaced. A coordinate is the space and its three axes; nothing here resolves a name and nothing enumerates. Split deliberately: the entry point, the user copies and the argument validation are in C (cube_syscall.c) because `SYSCALL_DEFINE*` is a C macro this kernel has no Rust equivalent for; everything that touches the store's bytes is in Rust, which passes the coordinate to the format code as its parts so that Morton encoding stays in the one module that must get it exactly right. A read that does not fit returns the size it needs rather than truncating — a short read would be worse than an error. Number 548: the x86_64 table says numbers 548 and above are available for non-x32 use. Gate (kernel/verify-syscall.sh): a static client in the initramfs does four writes (including an empty value and a second space), reads one back *through the same interface*, and folds with `sync`. Three different failures are separated — the calls failing (a broken ABI), a read not returning what a write stored (a wrong key encoding or index), and the folded image differing (a wrong format, order or merge). put 7,0,0 ok (21 bytes) put 8,0,0 ok (0 bytes) put 9,0,0 ok (28 bytes) put 1,2,3 ok (13 bytes) get 7,0,0 21 bytes: the kernel wrote this sync ok and the image left on the device is byte-identical to the one userspace writes from the same mutations, with the log empty. The interface's store is the same store. All nine gates pass on 0.6. |
||
|
|
2d4ed0fd9f |
CUBELinux.0.5: a torn log entry is overwritten, not buried
The write path's last unproven claim was the one everything else rests on: replay is prefix-trusting. Building the gate for it found that the claim was false as implemented — in *both* implementations, and in the same way. The log walkers advanced their offset past an entry's frame and only then checked the checksum. A corrupted entry therefore never moved the *replay* point (it was discarded, so the store looked right) but it did move the *append* point. The consequences, in order of how bad they are: - the kernel appended **after** the tear, burying the corruption inside a log that then looked well-formed — the exact opposite of the documented rule; - the userspace log truncated to a prefix that still contained the corrupt frame, so every later append started past it and the corruption stayed in the log forever; - and because each append then read back the same wrong prefix, every append wrote to the same offset and overwrote its predecessor — four mutations in, one on disk. The fix is one line of ordering in three places: compute the frame's end, check the length and the checksum, and only then move the offset. An entry that did not validate may not move the point that says where the log ends. Gate (kernel/verify-torn-tail.sh), on a store with an overwrite, an insertion and a deletion whose last entry has been torn by zeroing its tail in place: whole : 272 bytes, 3 records (the delete applied — the contrast) reference : 349 bytes, 4 records (userspace, torn bytes: the delete discarded) kernel : 349 bytes, 4 records appended : 667 bytes, 8 records (the kernel appended over the tear) expected : 667 bytes, 8 records (userspace: torn tail dropped, then the same four) Building the gate also corrected the gate itself: tearing a log by *truncating the device* is not a torn log, it is a smaller device — the capacity shrinks, writes past the end vanish into the page cache with no error anywhere, and the test measures an artifact. A torn write leaves the device the same size and corrupts bytes in place. All seven gates pass on 0.5. |
||
|
|
451754c7ac |
CUBELinux.0.4: the kernel folds its own log
Until now the kernel could append to a log but needed a *userspace* checkpoint to reclaim it — the dependency the write-authority decision was meant to remove. This is the fold. A store device is no longer an image at offset 0. It is: [control copy A][control copy B][slot 0][slot 1][log] with the control block saying which slot is live and how much log is in use, and two copies each carrying a generation and a checksum. Two slots buy atomicity: the folded image is written to the slot nothing is reading and flushed, and only then does the control change — one small write to the copy that is *not* current. At every instant there is either the old image with a log that still describes the mutations since it, or the new image with an empty log. A crash in the middle of the copy leaves the previous store intact, which is the entire reason for two slots. Folding into the only copy of an image is not atomic, so `sync` refuses on a bare image rather than pretending. `sync` is a real operation, not a test hook: a caller that wants the log reclaimed — a shutdown, a snapshot, a handover — is entitled to ask. Gate (kernel/verify-kernel-checkpoint.sh): the kernel writes four mutations and then folds its own log. The image it writes is compared with the one userspace writes from the same mutations byte for byte, not by digest — and it is identical, with the log empty afterwards. Three bugs came out of building this, all caught by the gates rather than by inspection: - the refactor that extracted the merge left the image walk *duplicated* inside the digest, re-adding image records after the log's entries — which gave them a newer sequence number and quietly resurrected records the log had deleted; - a v1 image has no extent, so it has no log either, and the new layout code was demanding both; - an unwritten log region is empty, not broken, and the first append was refusing it. The append gate's harness was also counting a *refused* write as accepted, which is how the second one hid. All six gates pass: three image reads (v1 curated, v1 snapshot at 35,318 records, v2 store), the log replay, the append with its SIGKILL durability test, and the checkpoint compared byte for byte. |
||
|
|
93e2c8c39b |
CUBELinux.0.3: the kernel writes, and is compared while writing
0.2 read the store. This appends to it: a mutation through /dev/cubelinux is written to the log and fsynced *before* the write is accepted, which is the kernel's version of the contract cube_duratest.py proves for the daemon. - the mutation comes in as the argument block the coordinate interface will pass: op(1) | space(32) | x(8) | y(8) | z(8) | len(4) | value[len]. The kernel Morton-encodes the point itself — a key written here has to be the key a userspace reader decodes, and that is a thing the gate would catch if it were merely similar; - the append lands after the log's valid prefix, so a torn tail is overwritten rather than appended to, exactly as the userspace log does it; - a log region that has never been written is zeros, not a log: the header is written first, the way the userspace log creates its file; - the entry's CRC covers space, key, length and value, computed the same way. The write is the byte plane — bytes at a coordinate, no header written beside them — because that is what a differential comparison against userspace's `cell put` can be exact about. The header tier sits above this. Gate (kernel/verify-kernel-append.sh), four mutations including an empty value and a record in a second space: reference : bytes=500 records=6 value_bytes=94 fnv1a64=6b679d39597a62b3 kernel : bytes=500 records=6 value_bytes=94 fnv1a64=6b679d39597a62b3 folded : bytes=500 records=6 value_bytes=94 fnv1a64=6b679d39597a62b3 survived : bytes=500 records=6 value_bytes=94 fnv1a64=6b679d39597a62b3 The third line is the one that matters: userspace *reads the log the kernel wrote*, folds it, and lands on the same store. Without it the first two would only show the kernel agreeing with itself. The fourth is a SIGKILL of the VM with no shutdown and therefore no flush for us. Not yet: the kernel folds nothing itself, so it still depends on a userspace checkpoint to reclaim its log. That is the next step, and it has its own gate. |
||
|
|
cee3e554d9 |
CUBELinux.0.2: the kernel reads the CUBE store off a block device
The first CUBE code in the kernel, and deliberately only a reader: the write
authority has not moved yet, and PLAN-kernel-cubelinux.md records both that
decision and the hazard that makes the order matter — a kernel writing while a
userspace daemon still holds the same image loses one of the two writers' work,
silently. A reader cannot do that.
- drivers/cube/: a Rust module exposing /dev/cubelinux. Reading it reads the
pinned image from the block device through the kernel's own file layer
(filp_open + kernel_read — the path this kernel version binds for Rust, and
the reason no C helper was needed), parses the records, and returns one line:
digest curve=0 bytes=8400896 records=35318 value_bytes=6139148 fnv1a64=5e20f98455387b08 errors=0
The work happens on read, not at init, so there is no initcall ordering to get
wrong against the block driver that provides the device.
- The format is restated in the kernel (32-byte space, 24-byte key, 8-byte LE
length, value), including the two rules the userspace parser documents: a value
that runs past the buffer is a truncated record, and an all-zero frame ends the
records only when every remaining byte is zero — the rule that keeps a real
record at the origin from being read as padding.
- The digest is the point. A record count alone lets two different images agree;
folding the bytes in means the kernel and userspace are *compared* rather than
assumed to agree. `cube-image digest` prints the same line in the same field
order, and the QEMU gate fails if they differ by a byte.
Gate, on both images:
curated 11 records, 4,096 bytes, fnv1a64=161113085b1573b2 — match
snapshot 35,318 records, 8,400,896 bytes, fnv1a64=5e20f98455387b08 — match
The tree carries CONFIG_CUBELINUX_STORE=y on top of defconfig + RUST; a tree
without it boots and simply has no /dev/cubelinux.
|
||
|
|
e4ae67ea7f |
CUBELinux.0.1: name the release, and build with this box's rustc
The tree is Linux 6.19.3 with the changes this machine's toolchain needs, and nothing else. No CUBE code yet — this is the base the coordinate interface will be built on, so it starts from a known-good bootable kernel. Naming: - VERSION/PATCHLEVEL/SUBLEVEL stay 6.19.3 (visible in `make kernelversion`), while the release string setlocalversion composes is CUBELinux.0.1, so `uname -r` and /lib/modules report this product rather than a Linux point release. The SCM suffix still applies: a dirty tree says so. rustc compatibility (rustc 1.100.0-nightly, clang 19.1.7): - scripts/generate_rust_target.rs emitted `"rustc-abi": "x86-softfloat"`, which this rustc rejects; it is `softfloat` now. - rust/Makefile's cmd_rustc_library did not pass -Zunstable-options, so the custom target spec would not load at all. - the generated bindings declare `strlen` with the kernel target's `c_char` (u8, from -funsigned-char) while rustc expects `*const i8`; the newer suspicious_runtime_symbol_definitions lint fires on that and -D warnings makes it fatal. Scoped to bindings.o and uapi.o, not to handwritten code. - three `'static` bounds the abstractions now need (irq handlers). - str.rs imported alloc::flags::* which prelude::* already provides, and `#![feature(used_with_arg)]` is stale now that the feature is stable; both are unused-feature/unused-import errors under -D warnings. Gates: bzImage builds (14,697,472 bytes) with CONFIG_RUST=y, virtio-blk and a serial console built in. Boot test in QEMU is next. |
||
|
|
598cf27219 |
Linux 6.19.3
Link: https://lore.kernel.org/r/20260217200002.683975158@linuxfoundation.org Tested-by: Florian Fainelli <florian.fainelli@broadcom.com> Tested-by: Takeshi Ogasawara <takeshi.ogasawara@futuring-girl.com> Tested-by: Peter Schneider <pschneider1968@googlemail.com> Tested-by: Jon Hunter <jonathanh@nvidia.com> Tested-by: Salvatore Bonaccorso <carnil@debian.org> Tested-by: Brett A C Sheffield <bacs@librecast.net> Tested-by: Mark Brown <broonie@kernel.org> Tested-by: Luna Jernberg <droidbittin@gmail.com> Tested-by: Ronald Warsow <rwarsow@gmx.de> Tested-by: Justin M. Forbes <jforbes@fedoraproject.org> Tested-by: Ron Economos <re@w6rz.net> Tested-by: Miguel Ojeda <ojeda@kernel.org> Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org> |