Files
cubelinux-kernel/drivers/cube/cubelinux_store.rs
T
surface-camera-build e968b3964e store: a write reads the log, not the whole device — and the log is now bounded
A put cost 60-131 ms on the box while a read cost 1-2 ms, and the reason was one call:
every write path reached `append` through `device_and_layout()` → `read_image()`, which
reads the WHOLE DEVICE — 100,663,296 bytes here — into a KVVec, to append ~70 bytes.
`append` itself never looks at the image: it reads the log region and the control
block. The read paths had been taught to read only where the image lies; the write
paths never were.

`AppendSource` now names what an append reads: `whole_image` (kept for the two callers
that genuinely need it — a bare store, whose log runs to the end of the file, and the
v1→v4 migration, which folds) or `log_head` (a device's control block plus the log's
first `WAL_HEADER_LEN` bytes). `write_view()` reads exactly that, `append_mutation()`
is the single entry point all four writers use — put, del, the boot record, and the
misc device's write_iter — and the bare-image fallback is decided in one place instead
of four. The fold still reads the whole store, because a fold rewrites it; that is the
honest tail, and it is why a fold belongs on a timer rather than on the path a caller
waits behind.

The log is bounded by `LOG_FOLD_BYTES`, because it is on the READ path: `Addressed::open`
reads `WAL_HEADER_LEN + log_used` bytes on every call, so an unfolded log is a tax on
every read, not a bill paid once at the fold. The live store's control block reports
log capacity 50,335,744 bytes — if it ever filled, every coordinate read would read
~48 MB before answering anything. A write that finds the log over the line folds first,
through `fold_for_headroom`, and then appends to a fresh log; guarded so one append
folds at most once, because a fold that did not shrink the log must not loop. The trade
is named in the code: the write that crosses the line pays a bounded, rare fold instead
of every reader paying an ever-larger log. The timer's comment — "the log grows at
roughly a megabyte a day against a 50 MB region, and a fold rewrites the whole image, so
folding more often would buy nothing and cost I/O" — is a data workload's arithmetic,
where the log is a recovery artefact and nobody reads it.

`space_entry`, `find` and `lower_bound` each allocated a fresh KVVec inside their
binary-search loop; one buffer per search now.

The boot record's one attempt becomes a bounded retry: `ensure_boot_record` claimed the
boot with a swap BEFORE it tried, so a single transient failure cost the boot its record
silently — which is exactly what happened on the box, where the first client of the boot
could not open the store for writing. It now claims one of `BOOT_RECORD_MAX_ATTEMPTS`
with a compare-exchange, and parks the count at the ceiling on success.

Release moves to 6.19.3-cubelinux0.7: verify-box-preflight.sh now refuses a same-release
reinstall, because the default entry boots the release being replaced.

Gates, on #81:
  verify-enum-cost    PASS  a write is 1.45 ms mean / 7.46 ms worst (new ceilings: 50 ms
                            write, 5 ms read), and a write does not grow with the store
  verify-file-store   PASS  12 writes to a store that is a FILE, all accounted for
  verify-syscall      PASS  the store the kernel writes is byte-identical to userspace's
2026-09-23 17:32:05 -04:00

4647 lines
188 KiB
Rust
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
// SPDX-License-Identifier: GPL-2.0
//! CUBELinux store reader — the kernel reading CUBE coordinates off a block device.
//!
//! The store is addressed by coordinate, and the pinned image's layout is:
//!
//! ```text
//! [MAGIC "CUBE" 4][version 1][curve tag]
//! records: [SpaceId 32][Key 24][value length u64 LE][value]
//! ```
//!
//! A store on a device is that image followed by a **write-ahead log** — the mutations since
//! the last checkpoint, in the same `op | crc32 | space | key | len | value` framing
//! `cube-store/src/wal.rs` writes. The log is a delta: the image is authoritative and
//! self-contained, and a reader that ignores the log loses only the mutations recorded after
//! the last checkpoint. `DESIGN-cubelinux-write-path.md` is why the write path looks like
//! this at all.
//!
//! This module reads both and reports the store they *describe* — the image with the log
//! applied: the record count, the total value bytes, and an FNV-1a digest over every
//! `(space, key, length, value)` in the order a checkpoint would write them. The digest exists so the kernel and userspace can be *compared*
//! rather than assumed to agree — `cube-image digest <image>` prints the same line in the
//! same field order, and the QEMU gate fails if they differ by a byte.
//!
//! # Why a reader, and only a reader
//!
//! The store's write authority has not moved yet. `PLAN-kernel-cubelinux.md` records the
//! decision (the kernel owns the store) and the hazard that makes the order matter: a
//! kernel that writes while a userspace daemon still holds the same image in memory loses
//! one of the two writers' work, silently. Until that is settled, a reader is the correct
//! amount of authority for the kernel to hold.
//!
//! # Interface
//!
//! `/dev/cubelinux` — reading it performs the read (through the kernel's own file layer)
//! and returns one line. Nothing happens at module init, so there is no ordering to get
//! wrong against the block driver that provides the device.
//!
//! ```text
//! digest curve=0 bytes=1141 records=11 value_bytes=431 fnv1a64=161113085b1573b2 errors=0
//! ```
use core::fmt::{self, Write};
use core::sync::atomic::{AtomicU32, Ordering};
// The store's format, in one file, shared with userspace.
//
// `crates/cube-format` includes this same path, so the layout, its constants and the arithmetic a
// reader derives from them have one description rather than one per build system. What is not here
// is I/O or allocation: this driver reads with its own buffers, and userspace already has the whole
// image. The constants below are aliases of the shared ones, and the two functions that *validate*
// a header delegate to it — validation is where a second description would be read as truth.
#[path = "cube_format.rs"]
#[allow(dead_code)] // the two builds use different subsets of one API
mod cube_format;
use kernel::{
alloc::AllocError,
bindings, c_str,
device::Device,
fs::{File, Kiocb},
iov::{IovIterDest, IovIterSource},
miscdevice::{MiscDevice, MiscDeviceOptions, MiscDeviceRegistration},
prelude::*,
sync::aref::ARef,
};
module! {
type: CubeStoreModule,
name: "cubelinux_store",
authors: ["CUBE OS"],
description: "CUBELinux store reader (coordinates in the kernel)",
license: "GPL",
}
/// The layout, restated here because the kernel cannot depend on the userspace crates.
/// `crates/cube-store-raw` is the source of truth; the digest comparison is what keeps
/// this copy honest.
const MAGIC: &[u8; 4] = cube_format::MAGIC;
/// The original format: magic, version, curve, then records until zero padding.
const VERSION_V1: u8 = cube_format::VERSION_V1;
/// The packed format: the same, plus the image's byte extent and record count.
/// The addressed format: the same, plus a fixed-size index, a space table, and packed values.
///
/// v2 is a packed list — `space | key | len | value`, repeated — so record N's offset is the sum of
/// every record before it. A coordinate gives the key, and the key gives ORDER, and order is not an
/// address: no coordinate can be turned into a place, and not even a binary search is possible,
/// because the middle record's offset cannot be computed either. Every lookup, therefore, walks.
///
/// v3 puts the addresses in the image:
///
/// ```text
/// [header 46] magic, version, curve, extents, counts
/// [space table: space_count x 48] space | first index | records
/// [index: record_count x 40] key | value offset | value length
/// [values: packed, in index order]
/// ```
///
/// The index is fixed-size and sorted by key, so a lookup is a binary search over arithmetic
/// addresses (`index_off + i * 40`), a listing starts at its space's first index entry and streams,
/// asking which spaces exist is the space table, and a spatial range is a contiguous run of index
/// entries. An index entry is 40 bytes against v2's 64-byte frame, so the image also gets smaller.
const VERSION_V3: u8 = cube_format::VERSION_V3;
/// v3, plus a 16-bit class mask in each index entry.
const VERSION_V4: u8 = cube_format::VERSION_V4;
const HEADER_LEN_V1: usize = cube_format::HEADER_LEN_V1;
const HEADER_LEN_V2: usize = cube_format::HEADER_LEN_V2;
/// magic(4) version(1) curve(1) image_bytes(8) record_count(8) space_count(8) index_off(8)
/// values_off(8).
const HEADER_LEN_V3: usize = cube_format::HEADER_LEN_V3;
/// `key(24) | value_off(8) | value_len(8)`.
const INDEX_ENTRY: usize = cube_format::INDEX_ENTRY;
/// `key(24) | flags(2) | value_off(8) | value_len(8)`.
const INDEX_ENTRY_V4: usize = cube_format::INDEX_ENTRY_V4;
const FLAGS_LEN: usize = cube_format::FLAGS_LEN;
/// `space(32) | first index(8) | records(8)`.
const SPACE_ENTRY: usize = cube_format::SPACE_ENTRY;
const SPACE_ID_LEN: usize = cube_format::SPACE_ID_LEN;
const RAW_KEY_LEN: usize = cube_format::RAW_KEY_LEN;
const RECORD_FIXED: usize = cube_format::RECORD_FIXED;
// The device the store lives on, as the `cube_store=` boot parameter resolved it.
//
// This was a constant — *"for the boot gate one fixed name is honest and has one less way to
// be wrong"* — and that was true while a virtual machine was the only place this driver ran.
// It runs on the box now, the box's device is not `/dev/vda`, and a constant cannot be both.
// So the box names its device on the kernel command line:
//
// cube_store=/dev/nvme0n1p2
//
// The parameter is declared in C (`cube_syscall.c`) because this kernel's Rust can express only
// *integer* module parameters — `rust/kernel/module_param.rs` implements `ModuleParam` through
// `ParseInt` and nothing else — and a path is not an integer. It is a `__setup` parameter rather
// than a module parameter for the naming reason recorded in that file.
//
// (A `///` here would be an unused doc comment: it would attach to the `extern` block, which
// documents nothing.)
extern "C" {
/// Returns a pointer to a static, NUL-terminated buffer holding the path.
fn cubelinux_store_device() -> *const core::ffi::c_char;
/// Whether the kernel was asked to record its own boots (`cube_boot_record=1`).
///
/// Parsed in C for the same reason the device path is: this kernel's Rust can express only
/// integer module parameters, and a command-line flag beside `cube_store=` belongs where that
/// parameter lives.
fn cubelinux_boot_record_enabled() -> bool;
}
/// The store's path, NUL-terminated, owned by the C side for the life of the kernel.
fn store_device() -> *const u8 {
// SAFETY: `cubelinux_store_device` returns a pointer to a static buffer that the module
// parameter filled during boot and that nothing writes afterwards, so it stays valid and
// NUL-terminated for as long as this runs.
unsafe { cubelinux_store_device().cast::<u8>() }
}
/// Refuse to pull an unbounded device into memory: the guard against a wrong device name
/// turning a read into an allocation storm. It is sized for a *store device*, not for an image.
///
/// `Control::layout` gives each of the two slots a quarter of the device, so a device must be at
/// least four times the image it holds — and a *sealed* image is larger than the plaintext one,
/// because every record gains an envelope. A sealed 18.8 MiB store therefore needs a ~76 MiB
/// device, which the original 64 MiB cap refused. 128 MiB leaves room for both to grow while
/// still being a bound: pointing this at a whole 119 GiB disk is still refused.
const MAX_BYTES: usize = 128 * 1024 * 1024;
/// The log size at which a write folds the log before appending to it.
///
/// This is a ceiling on a READER's per-call cost, which is why it lives in the write path at all.
/// Every read reads the whole log window before it can answer anything — `Addressed::open` reads
/// `WAL_HEADER_LEN + log_used` bytes — so an unfolded log is not a bill paid once, at the fold. It is
/// a tax on every call, and it grows all day.
///
/// The timer's policy said the opposite, and it was right for what it was written about: "the log
/// grows at roughly a megabyte a day against a 50 MB region, and a fold rewrites the whole image, so
/// folding more often would buy nothing and cost I/O". That is a data workload's arithmetic, where
/// the log is a recovery artefact and nobody reads it. Here the log is on the read path, and a 50 MB
/// log turns every coordinate read into a 50 MB read — a cliff that arrives at the end of the
/// region's life, not a slope. So the log is kept under this line instead.
///
/// The trade, named rather than hidden: the WRITE that crosses the line pays a fold — once, bounded,
/// and rare — instead of every reader paying an ever-larger log. It is not free, and a fold is still
/// O(store) and still serialised against other operations by [`STORE_OP`]; making the fold itself
/// cheap is a separate piece of work.
const LOG_FOLD_BYTES: u64 = 256 * 1024;
/// The log's own magic, distinct from the image's so a reader that opens the wrong one
/// cannot mistake it for the other.
const WAL_MAGIC: [u8; 4] = *cube_format::WAL_MAGIC;
const WAL_VERSION: u8 = cube_format::WAL_VERSION;
/// The flagged log entry: the same, with a two-byte class mask before the length.
const WAL_VERSION_V2: u8 = cube_format::WAL_VERSION_V2;
/// Log header: magic(4) + version(1) + curve tag(1).
const WAL_HEADER_LEN: usize = cube_format::WAL_HEADER_LEN;
/// Log entry, before the value: op(1) + crc32(4) + space(32) + key(24) + len(4).
const ENTRY_FIXED: usize = cube_format::ENTRY_FIXED;
/// The flagged entry: the same, with `flags(2)` before the length.
const ENTRY_FIXED_V2: usize = cube_format::ENTRY_FIXED_V2;
/// One log entry's fixed size, and the offset of its length, for a log version. Version 1 has no
/// class mask; version 2 carries a two-byte one before the length.
fn wal_frame(version: u8) -> (usize, usize) {
if version == WAL_VERSION_V2 {
(ENTRY_FIXED_V2, RAW_KEY_LEN + FLAGS_LEN + SPACE_ID_LEN + 5)
} else {
(ENTRY_FIXED, 61)
}
}
/// The log begins at the first 4 KiB boundary at or after the image. Fixed by geometry so
/// no superblock is needed to find it, and stated by the v2 header's extent.
///
/// That is the *bare* layout: an image on a device with a log after it. A device that is
/// meant to be checkpointed has a control block instead (see `Control`), because folding a
/// log into the only copy of an image is not atomic and a crash in the middle of it loses the
/// store.
const LOG_ALIGN: usize = 4096;
/// Control block magic, distinct from the image's `CUBE` and the log's `CUBW`.
const CTL_MAGIC: [u8; 4] = *b"CUBS";
const CTL_VERSION: u8 = 1;
/// One control copy, and where the two live.
const CTL_COPY_LEN: usize = 2048;
const CTL_COPY_A: usize = 0;
const CTL_COPY_B: usize = 2048;
/// Where the two image slots begin.
const CTL_DATA_OFF: u64 = 4096;
/// Bytes of a control copy covered by its checksum.
const CTL_SUMMED: usize = 44;
/// The one mutation this device accepts so far. The coordinate interface's opcodes will be
/// the language's verbs; this is the first of them.
const OP_PUT: u8 = 1;
/// Fold the log into the image.
const OP_SYNC: u8 = 3;
// ── The kernel's record of its own boot ────────────────────────────────────────────────
//
// "The OS stores itself" has two halves. The store is the kernel's, which is the first half; the
// second is the kernel *recording* something of its own rather than a client doing it. This is that
// record: at the first write of a boot the kernel appends one entry describing the boot it is
// having — its own version banner, the wall-clock time, and the device it resolved — to a reserved
// space, through the same append path every other mutation uses.
//
// WHERE IT HOOKS, AND WHY NOT AT INIT. There is no store open at module init, deliberately: the
// driver does nothing until it is asked, so there is no ordering to get wrong against the block
// driver that provides the device. An `__initcall` appending a record would reintroduce exactly that
// ordering problem, so the hook is not init — it is the first *write*. That is also the moment the
// claim is about: the kernel records itself when it becomes the writer. A boot in which the kernel
// only ever reads writes no record, which is honest rather than a gap.
//
// WHY IT IS OFF BY DEFAULT. Recording unconditionally would put a record into every store the
// kernel writes, and the gates' method is that the store the kernel produces is comparable, byte for
// byte, with the store userspace produces from the same mutations. A kernel-injected record the
// caller never asked for would make two of those comparisons stop being comparisons — so the feature
// is enabled by `cube_boot_record=1` on the kernel command line (`cube_syscall.c`), the gate turns it
// on, and the default keeps the invariant the gates rest on.
/// The space the kernel records its own boots in. Reserved, and reserved without a name: the kernel
/// keeps no table of space names, because a name is a userspace convention.
///
/// `0xFC` because the repeated-byte tags below it are taken, and taken by things this must not
/// collide with — which is worth spelling out, because the first choice here was `0xFD` and that is
/// the OS keystore, the space holding key material that the kill switch exists to destroy. Writing
/// boot records into it would have been a serious bug, and only a check of the userspace table
/// (`format_space` in `crates/cube-command`) catches it: the kernel has no name table to consult.
///
/// ```text
/// 0x00 root the data space, what an unqualified coordinate means
/// 0xFF edges association edges (`cube-header::NULL_SPACE`)
/// 0xFE portal portal descriptors (`cube-index::PORTAL_SPACE`)
/// 0xFD keystore the OS keystore (`cube-crypt::KEYSTORE_SPACE`)
/// 0xFC boot this — the kernel's record of its own boots
/// ```
const BOOT_SPACE: [u8; SPACE_ID_LEN] = [0xFC; SPACE_ID_LEN];
/// The class mask the boot record stamps: `EventFlags::BOOT`, bit 7. The kernel cannot link
/// `cube-core`, so the bit is spelled out here and on the userspace side, and the agreement is held
/// by `verify-boot-record` — the same shape as the `0xFC` space id, for the same reason.
const BOOT_FLAGS: u16 = 1 << 7;
/// The coordinate the record lives at. One record, the current boot: each boot overwrites the last,
/// which is the same shape as the userspace boot marker it replaces and is what "which boot is this"
/// needs. History would be a second record per boot, and nobody has asked for one.
const BOOT_POINT: (u64, u64, u64) = (0, 0, 0);
/// Whether this boot has already been recorded. One attempt, claimed with a swap so two callers
/// racing into their first write cannot write two records.
/// How many attempts this boot has made at recording itself, and the ceiling on them.
///
/// A count rather than a flag, because a flag that is claimed *before* the write cannot un-claim
/// itself, and the first attempt of a boot can fail for reasons that do not last. Measured on the
/// box: with `cube_boot_record=1` on the command line and a store image that the first client of the
/// boot could not open for writing, the kernel logged `could not read the store to record this boot`
/// and the store held no record at all — the switch was on and the effect was off, for that whole
/// boot, because the claim was made first and never retried. Bounded, so a store that is genuinely
/// broken logs a few warnings rather than one per write.
static BOOT_RECORD_ATTEMPTS: AtomicU32 = AtomicU32::new(0);
const BOOT_RECORD_MAX_ATTEMPTS: u32 = 4;
/// FNV-1a offset basis.
const FNV_OFFSET: u64 = cube_format::FNV_OFFSET;
/// A line of output, built in place. No allocation: the read path formats into a fixed
/// buffer, so it cannot fail for want of memory while holding a file open.
struct Line {
buf: [u8; 256],
len: usize,
}
impl Line {
fn new() -> Self {
Line {
buf: [0; 256],
len: 0,
}
}
fn as_bytes(&self) -> &[u8] {
&self.buf[..self.len]
}
}
impl Write for Line {
fn write_str(&mut self, s: &str) -> fmt::Result {
let bytes = s.as_bytes();
if self.len + bytes.len() > self.buf.len() {
return Err(fmt::Error);
}
self.buf[self.len..self.len + bytes.len()].copy_from_slice(bytes);
self.len += bytes.len();
Ok(())
}
}
/// FNV-1a 64 over a run of bytes, continuing from `h`.
fn fnv1a64(bytes: &[u8], mut h: u64) -> u64 {
for b in bytes {
h ^= *b as u64;
h = h.wrapping_mul(0x0000_0100_0000_01b3);
}
h
}
/// Read the whole store image through the kernel's file layer.
///
/// `kernel_read` on the block device rather than a raw bio: it is the path this kernel
/// version exposes to Rust, it goes through the page cache the way any other read does, and
/// it needs no C helper. If the store ever has to be read before the VFS is up (a root
/// filesystem, say), that is the moment to reach for the block layer directly.
/// The store device, opened once and kept for the whole boot.
///
/// Every operation used to open the device, read it, and close it again — the open is the dominant
/// cost of a `cube(2)` call, and it was paid on every single one. The path is fixed at boot
/// (`cube_store=` is parsed once), so there is nothing to invalidate: the handle is opened lazily
/// on first use and reused.
///
/// Lazy rather than at module init, for the reason the driver does nothing at init: the block
/// driver that provides the device may not be up yet, and an open at init would reintroduce the
/// ordering problem the driver's own doc says it avoided. A caller whose filesystem view cannot
/// reach the path (a chroot before its `/dev` is mounted) gets the error that open returned; every
/// later caller gets the cached handle regardless of its own root.
/// A cached store-device file. The raw pointer is kept for the whole boot — the path is fixed at
/// boot (`cube_store=` is parsed once), so there is nothing to invalidate, and the device is meant
/// to stay open. The wrapper carries the `Send`/`Sync` the raw pointer lacks, and the global lock
/// serializes access, so no two threads touch the pointer unsynchronized.
struct StoreFile(*mut bindings::file);
// SAFETY: the pointer is valid for the whole boot (the path is fixed and the block driver does not
// unbind), and every use is through the global lock below.
unsafe impl Send for StoreFile {}
unsafe impl Sync for StoreFile {}
kernel::sync::global_lock! {
/// SAFETY: Initialized (to None) before first use.
unsafe(uninit) static STORE_FILE: Mutex<Option<StoreFile>> = None;
}
kernel::sync::global_lock! {
/// Serialises the store's read-modify-write operations.
///
/// [`STORE_FILE`]'s mutex guards the file *handle*, not the store: it is taken, the handle is
/// fetched or cached, and it is released before the caller does anything with it. Sharing a
/// `struct file *` is safe, so that is all the handle needs — and it is not enough for the
/// store. An append is a read-modify-write of the control block: read `log_used`, write the
/// entry at that offset, then write a control block counting it. Writers that both read
/// `log_used = 0` write their entries to the same offset and commit the same count, so one
/// entry survives, the rest are discarded, and every caller is told `ok`.
///
/// Measured with this lock absent: `verify-file-store.sh MODE=race`, twelve concurrent puts,
/// 12 × 68 bytes of log expected — a single 68-byte entry landed and all twelve reported
/// success. The same shape is on the box, where the store has several writers.
///
/// It is taken by the *entry points* only — [`append`] and [`fold_now`] — and their inner
/// bodies assume it is already held. `append` folds and then calls itself, and `fold_now` is
/// reachable both from inside `append` and on its own, so a lock taken at the top of either
/// body would re-enter itself and deadlock: a kernel mutex is not reentrant.
/// SAFETY: Initialized (to None) before first use.
unsafe(uninit) static STORE_OP: Mutex<()> = ();
}
/// The store device's file, opened on first use and cached for the boot. O_RDWR, because the one
/// handle serves both the reads and the log-append / checkpoint writes.
fn store_file() -> Result<*mut bindings::file> {
let mut guard = STORE_FILE.lock();
if let Some(file) = guard.as_ref() {
return Ok(file.0);
}
// O_RDWR is 2. SAFETY: store_device() is a NUL-terminated C string the module parameter filled
// at boot; filp_open returns a valid `struct file *` or an error pointer, checked below.
let filp = unsafe { bindings::filp_open(store_device(), 2, 0) };
let filp = kernel::error::from_err_ptr(filp)?;
if filp.is_null() {
return Err(EINVAL);
}
*guard = Some(StoreFile(filp));
Ok(filp)
}
fn read_image(image: &mut KVVec<u8>) -> Result<()> {
// O_RDONLY is 0 in Linux; filp_open takes the raw flags word.
// SAFETY: store_device() is a NUL-terminated C string that the module parameter filled at
// boot, and filp_open either returns a valid `struct file *` or an error pointer, which is
// checked below.
let file = store_file()?;
let mut pos: bindings::loff_t = 0;
let mut chunk = KVVec::<u8>::with_capacity(4096, GFP_KERNEL)?;
chunk.extend_from_slice(&[0u8; 4096][..], GFP_KERNEL)?;
let mut result = Ok(());
loop {
// SAFETY: `file` is a live `struct file *` from filp_open; `chunk` is a 4096-byte
// kernel buffer we own; `pos` is a valid loff_t. kernel_read reads at most
// `chunk.len()` bytes into the buffer and does not retain either pointer.
let n = unsafe {
bindings::kernel_read(
file,
chunk.as_mut_ptr().cast::<core::ffi::c_void>(),
chunk.len(),
&mut pos,
)
};
if n < 0 {
result = Err(Error::from_errno(n as i32));
break;
}
if n == 0 {
break; // end of device
}
let n = n as usize;
if let Err(e) = image.extend_from_slice(&chunk[..n], GFP_KERNEL) {
result = Err(e.into());
break;
}
if image.len() >= MAX_BYTES {
break;
}
}
// Less than the shortest header is not an image at all; `digest` reports the rest.
if result.is_ok() && image.len() < HEADER_LEN_V1 {
return Err(EINVAL);
}
result
}
/// Walk the records of a store image, digesting each one, exactly as
/// `cube-store-raw::iter_records` + `cube-image digest` do.
///
/// The two rules that are easy to get wrong, both from the userspace parser: a record whose
/// value runs past the buffer is a truncated record and an error, and an all-zero frame is
/// *end of records* only when every remaining byte is also zero — otherwise a real record
/// at the origin, followed by padding, would be read as an empty one.
// What a header says, in every format, is the shared file's: `version`, `curve`, the image's
// extent, and how many records follow. Same fields, one description.
use cube_format::Header;
/// Read the header, or say what is wrong with it.
///
/// The shared file decides what a header means — that is the point of it being shared — and this
/// only translates a refusal into the words this driver logs.
fn parse_header(image: &[u8]) -> Result<Header, &'static str> {
cube_format::parse_header(image).map_err(|bad| match bad {
cube_format::Bad::Magic => "not-a-store",
cube_format::Bad::Version(_) => "unsupported-version",
cube_format::Bad::Extent => "bad-extent",
cube_format::Bad::Tables => "bad-tables",
})
}
/// The v3 header: what a reader needs to address the image by arithmetic.
struct HeaderV3 {
/// `VERSION_V3` or `VERSION_V4` — they share the geometry and differ only in the index
/// entry's stride.
version: u8,
image_bytes: u64,
record_count: u64,
space_count: u64,
index_off: u64,
values_off: u64,
}
impl HeaderV3 {
/// The width of one index entry for this header's version.
fn stride(&self) -> usize {
if self.version == VERSION_V4 {
INDEX_ENTRY_V4
} else {
INDEX_ENTRY
}
}
fn encode(&self, out: &mut KVVec<u8>, curve: u8) -> Result<(), AllocError> {
out.extend_from_slice(MAGIC, GFP_KERNEL)?;
out.extend_from_slice(&[self.version, curve], GFP_KERNEL)?;
out.extend_from_slice(&self.image_bytes.to_le_bytes(), GFP_KERNEL)?;
out.extend_from_slice(&self.record_count.to_le_bytes(), GFP_KERNEL)?;
out.extend_from_slice(&self.space_count.to_le_bytes(), GFP_KERNEL)?;
out.extend_from_slice(&self.index_off.to_le_bytes(), GFP_KERNEL)?;
out.extend_from_slice(&self.values_off.to_le_bytes(), GFP_KERNEL)?;
Ok(())
}
fn decode(image: &[u8]) -> Result<Self, &'static str> {
// The shared file owns this: the offsets a reader computes (`index_off + i * stride`)
// are only places if the index really starts there and really is that wide, and that is a
// statement about the format rather than about this driver.
let geometry = cube_format::V3::decode(image).map_err(|bad| match bad {
cube_format::Bad::Extent => "bad-extent",
_ => "bad-tables",
})?;
let header = parse_header(image)?;
Ok(HeaderV3 {
version: geometry.version,
image_bytes: header.image_bytes.unwrap_or(0),
record_count: geometry.record_count,
space_count: geometry.space_count,
index_off: geometry.index_off,
values_off: geometry.values_off,
})
}
}
fn digest(image: &[u8]) -> Line {
let mut line = Line::new();
// Where the image is, and which log belongs to it: a store device says so in its control
// block, a bare image says so in its own header.
let layout = match resolve_layout(image) {
Ok(l) => l,
Err(what) => {
let _ = write!(line, "error={what}");
return line;
}
};
let live = &image[layout.image_off..];
let header = match parse_header(live) {
Ok(h) => h,
Err(what) => {
let _ = write!(line, "error={what}");
return line;
}
};
let end = match header.image_bytes {
Some(n) => core::cmp::min(n as usize, live.len()),
None => live.len(),
};
// With a v2 image, the store is the image *plus* whatever log follows it. The order the
// digest is taken in is then the order a checkpoint would write: sorted by coordinate,
// later writes winning. Without a log there is nothing to merge, and the image's own
// order is the answer — which is also what keeps v1 images readable and unchanged.
if header.image_bytes.is_some() {
let log = log_window(image, &layout);
match merged_digest(&live[..end], log, &header) {
Ok(line) => return line,
// A log that exists but cannot be read is worth saying out loud: the digest
// alone would look like a shorter store.
Err(what) if what != "no-log" => {
let mut line = Line::new();
let _ = write!(line, "error={what}");
return line;
}
Err(_) => {}
}
}
let mut h = FNV_OFFSET;
let mut count: u64 = 0;
let mut value_bytes: u64 = 0;
let mut errors: u64 = 0;
let header_len = if header.version == VERSION_V1 {
HEADER_LEN_V1
} else {
HEADER_LEN_V2
} as u64;
let mut off = layout.image_off + header_len as usize;
let mut remaining = header.record_count;
let end = layout.image_off + end;
while off + RECORD_FIXED <= end {
if remaining == Some(0) {
break;
}
let frame = &image[off..off + RECORD_FIXED];
// Padding is not a record — but only if it is padding all the way down, and only
// in v1, where nothing else says where the records stop. A v2 image states its
// count, so a zero frame inside that count is a record like any other.
if remaining.is_none()
&& frame.iter().all(|b| *b == 0)
&& image[off..].iter().all(|b| *b == 0)
{
break;
}
let space = &image[off..off + SPACE_ID_LEN];
let key = &image[off + SPACE_ID_LEN..off + SPACE_ID_LEN + RAW_KEY_LEN];
let mut len_bytes = [0u8; 8];
len_bytes.copy_from_slice(&image[off + SPACE_ID_LEN + RAW_KEY_LEN..off + RECORD_FIXED]);
let value_len = u64::from_le_bytes(len_bytes) as usize;
let value_at = off + RECORD_FIXED;
// A promised record that is not there is a truncated image, and saying so beats
// returning a shorter list that looks complete.
if value_at + value_len > end {
errors += 1;
break;
}
let value = &image[value_at..value_at + value_len];
h = fnv1a64(space, h);
h = fnv1a64(key, h);
h = fnv1a64(&(value_len as u64).to_le_bytes(), h);
h = fnv1a64(value, h);
count += 1;
value_bytes += value_len as u64;
off = value_at + value_len;
if let Some(n) = remaining.as_mut() {
*n -= 1;
}
}
// The count was a promise; not meeting it is an error the caller must see.
if let Some(n) = remaining {
if n > 0 {
errors += 1;
}
}
// Field order matches cube-image's `digest` line so the two diff directly.
let _ = write!(
line,
"digest version={} curve={} bytes={} records={} value_bytes={} fnv1a64={:016x} errors={}",
header.version,
header.curve,
// `bytes` is a property of the store, not of how it was read: the header plus the
// records, whatever padding the reader happened to see. Reporting what was read made
// the same records disagree between a file and a device — and the gates compare this
// line, so it has to mean the same thing on both sides.
// The records' logical size, and not the layout's: a v2 image and a v3 image of one
// store hold the same records in different shapes, and this line is how the two are
// compared — a figure that changes with the layout would make them look different.
count * RECORD_FIXED as u64 + value_bytes,
count,
value_bytes,
h,
errors
);
line
}
/// What a store device's control block says: where the two image slots and the log are,
/// which slot is live, and how much of the log is in use.
#[derive(Clone, Copy)]
struct Control {
curve: u8,
slot_bytes: u64,
log_capacity: u64,
log_used: u64,
generation: u64,
active_slot: u8,
}
impl Control {
/// Parse a control copy. `None` when it is not one, or its checksum does not hold — a
/// torn update in the middle of a checkpoint is exactly what that checksum is for.
fn decode(b: &[u8]) -> Option<Control> {
if b.len() < CTL_SUMMED + 4 || &b[0..4] != CTL_MAGIC {
return None;
}
let word = |at: usize| -> u64 {
let mut w = [0u8; 8];
w.copy_from_slice(&b[at..at + 8]);
u64::from_le_bytes(w)
};
let mut c = [0u8; 4];
c.copy_from_slice(&b[44..48]);
if u32::from_le_bytes(c) != crc32(&b[..CTL_SUMMED]) {
return None;
}
let control = Control {
curve: b[5],
slot_bytes: word(8),
log_capacity: word(16),
log_used: word(24),
generation: word(32),
active_slot: b[40],
};
if b[4] != CTL_VERSION || control.active_slot > 1 || control.slot_bytes == 0 {
return None;
}
Some(control)
}
/// Encode into a control copy, checksum included.
fn encode(&self, b: &mut [u8]) {
for byte in b.iter_mut() {
*byte = 0;
}
b[0..4].copy_from_slice(&CTL_MAGIC);
b[4] = CTL_VERSION;
b[5] = self.curve;
b[8..16].copy_from_slice(&self.slot_bytes.to_le_bytes());
b[16..24].copy_from_slice(&self.log_capacity.to_le_bytes());
b[24..32].copy_from_slice(&self.log_used.to_le_bytes());
b[32..40].copy_from_slice(&self.generation.to_le_bytes());
b[40] = self.active_slot;
let crc = crc32(&b[..CTL_SUMMED]);
b[44..48].copy_from_slice(&crc.to_le_bytes());
}
fn image_off(&self) -> u64 {
CTL_DATA_OFF + self.active_slot as u64 * self.slot_bytes
}
fn spare_off(&self) -> u64 {
CTL_DATA_OFF + (1 - self.active_slot) as u64 * self.slot_bytes
}
fn log_off(&self) -> u64 {
CTL_DATA_OFF + 2 * self.slot_bytes
}
}
/// The valid control block: the copy with the higher generation. A reader that finds one
/// copy torn uses the other, which is the entire reason there are two.
fn read_control(b: &[u8]) -> Option<Control> {
if b.len() < CTL_COPY_B + CTL_COPY_LEN {
return None;
}
let a = Control::decode(&b[CTL_COPY_A..CTL_COPY_A + CTL_COPY_LEN]);
let b_copy = Control::decode(&b[CTL_COPY_B..CTL_COPY_B + CTL_COPY_LEN]);
match (a, b_copy) {
(Some(x), Some(y)) => Some(if x.generation >= y.generation { x } else { y }),
(Some(x), None) => Some(x),
(None, Some(y)) => Some(y),
(None, None) => None,
}
}
/// Where the image is and what log belongs to it.
struct Layout {
/// Byte offset of the live image.
image_off: usize,
/// Byte offset of its log.
log_off: usize,
/// Bytes of log in use.
log_used: usize,
/// The control block, when this is a store device rather than a bare image.
control: Option<Control>,
}
/// Work out the layout: a store device when the control block is there, a bare image when it
/// is not. Both are readable; only the store layout can be checkpointed atomically, which is
/// why `sync` refuses on the other.
fn resolve_layout(b: &[u8]) -> Result<Layout, &'static str> {
if b.len() >= CTL_COPY_B + CTL_COPY_LEN && &b[0..4] == CTL_MAGIC {
let control = read_control(b).ok_or("control-block-corrupt")?;
let log_off = control.log_off() as usize;
let used = control.log_used as usize;
if log_off + used > b.len() || used > control.log_capacity as usize {
return Err("log-outside-device");
}
return Ok(Layout {
image_off: control.image_off() as usize,
log_off,
log_used: used,
control: Some(control),
});
}
// A bare image. A v1 image has no extent and therefore no log — it is the whole device
// as far as this module is concerned, and it is read the way it always was.
let header = parse_header(b)?;
let log_off = match header.image_bytes {
Some(n) => log_offset(n),
None => b.len(),
};
Ok(Layout {
image_off: 0,
log_off,
// A bare device records no log length, so the log is bounded by the buffer and by
// the entries that parse: `apply_log` stops at the first one that does not.
log_used: b.len().saturating_sub(log_off),
control: None,
})
}
/// One record on its way into the merged store. Fixed size, so a `KVVec` of these sorts
/// in place; values live in a pool beside them and are referenced by offset.
#[derive(Clone, Copy, PartialEq, Eq, PartialOrd, Ord)]
struct Entry {
space: [u8; 32],
key: [u8; 24],
/// The class mask, carried from the log or the image so a fold does not lose it.
flags: u16,
/// Order in which this entry arrived. Breaks ties between the same coordinate, so the
/// later write wins after sorting.
seq: u32,
value_off: u32,
value_len: u32,
deleted: bool,
}
/// The store as a sorted set of records, which is what a checkpoint writes and what the
/// digest is taken over.
struct Merged {
pool: KVVec<u8>,
entries: KVVec<Entry>,
seq: u32,
/// Log entries that were seen but not applied, because they were torn or malformed.
dropped: u64,
}
impl Merged {
fn new() -> Result<Self, AllocError> {
Ok(Merged {
pool: KVVec::new(),
entries: KVVec::new(),
seq: 0,
dropped: 0,
})
}
/// Add a record. `deleted` marks a removal, which is kept in the list so it can
/// override an older value for the same coordinate.
fn add(&mut self, space: &[u8], key: &[u8], flags: u16, value: &[u8], deleted: bool) -> Result<(), AllocError> {
let mut sp = [0u8; 32];
sp.copy_from_slice(&space[..32]);
let mut k = [0u8; 24];
k.copy_from_slice(&key[..24]);
let value_off = self.pool.len() as u32;
self.pool.extend_from_slice(value, GFP_KERNEL)?;
let seq = self.seq;
self.seq += 1;
self.entries.push(
Entry {
space: sp,
key: k,
flags,
seq,
value_off,
value_len: value.len() as u32,
deleted,
},
GFP_KERNEL,
)?;
Ok(())
}
/// Put the entries in the order a checkpoint would write them: by space, then by key,
/// with the later write of a coordinate last — so a group's final entry decides it.
///
/// This is a heapsort, and the reason is a panic rather than a preference.
/// `slice::sort_unstable` allocates about 3.5 KiB of *kernel stack* per recursion level in
/// this kernel's Rust, so a store with a couple of hundred records overflowed the 16 KiB
/// kernel stack and took the machine down — `BUG: TASK stack guard page was hit`, with the
/// instruction pointer inside
/// `core::slice::sort::unstable::quicksort::<cubelinux_store::Entry>` and `sub rsp, 0xdd8` in
/// the function's first bytes. A heapsort has a constant frame and allocates nothing, so the
/// size of a store cannot decide whether reading it is safe, and the sort cannot fail for want
/// of memory in the middle of a read.
///
/// It is not stable, which costs nothing here: `seq` is part of the ordering key, so the order
/// is total and two entries never compare equal.
fn sort_entries(&mut self) {
let entries = self.entries.as_mut_slice();
let len = entries.len();
if len < 2 {
return;
}
// Build a max-heap, then repeatedly move the maximum to the end.
let mut start = len / 2;
while start > 0 {
start -= 1;
sift_down(entries, start, len);
}
let mut end = len;
while end > 1 {
end -= 1;
entries.swap(0, end);
sift_down(entries, 0, end);
}
}
fn value(&self, e: &Entry) -> &[u8] {
let off = e.value_off as usize;
&self.pool.as_slice()[off..off + e.value_len as usize]
}
}
/// Restore the heap property below `root` over `entries[..end]`.
///
/// Iterative on purpose: the whole point of the heapsort above is that nothing here grows with the
/// size of the store.
fn sift_down(entries: &mut [Entry], mut root: usize, end: usize) {
loop {
let left = 2 * root + 1;
if left >= end {
return;
}
let mut largest = left;
if left + 1 < end && entries[left] < entries[left + 1] {
largest = left + 1;
}
if entries[root] < entries[largest] {
entries.swap(root, largest);
root = largest;
} else {
return;
}
}
}
/// Read the log that follows the image, if there is one, and apply its entries.
///
/// The log is a delta on the image, so its entries override image records for the same
/// coordinate. Recovery is prefix-trusting, exactly as userspace does it: entries are
/// replayed from the start and the first one that is short, mis-framed or fails its
/// checksum ends the log. A fault in the middle cannot be told from a torn tail without a
/// second copy, so the log stops there and reports what it dropped instead of guessing.
fn apply_log(log: &[u8], merged: &mut Merged) -> Result<Option<u64>, &'static str> {
if log.len() < WAL_HEADER_LEN || &log[0..4] != WAL_MAGIC {
return Ok(None);
}
if log[4] != WAL_VERSION && log[4] != WAL_VERSION_V2 {
return Err("unsupported-log-version");
}
let (fixed, len_at) = wal_frame(log[4]);
let mut applied: u64 = 0;
let mut off = WAL_HEADER_LEN;
while off + fixed <= log.len() {
let start = off;
let op = log[off];
if op != 1 && op != 2 {
break;
}
let mut word = [0u8; 4];
word.copy_from_slice(&log[off + 1..off + 5]);
let crc = u32::from_le_bytes(word);
let space = &log[off + 5..off + 37];
let key = &log[off + 37..off + 61];
let flags = if log[4] == WAL_VERSION_V2 {
cube_format::le_u16(log, off + 61)
} else {
0
};
word.copy_from_slice(&log[off + len_at..off + len_at + 4]);
let len = u32::from_le_bytes(word) as usize;
let frame_end = start + fixed + len;
if frame_end > log.len() {
break;
}
// The checksum covers space, key, the mask, length and value, so a corrupted entry is
// stopped at rather than applied.
if crc32(&log[start + 5..frame_end]) != crc {
break;
}
let value = &log[start + fixed..frame_end];
off = frame_end;
merged
.add(space, key, flags, value, op == 2)
.map_err(|_| "out-of-memory")?;
applied += 1;
}
merged.dropped = (log.len() - off) as u64;
Ok(Some(applied))
}
/// CRC-32 (IEEE 802.3), bitwise — the same polynomial and the same coverage as
/// `cube-store/src/wal.rs`, because a log written on either side must validate on both.
fn crc32(bytes: &[u8]) -> u32 {
let mut crc = 0xFFFF_FFFFu32;
for byte in bytes {
crc ^= *byte as u32;
for _ in 0..8 {
let mask = (crc & 1).wrapping_neg();
crc = (crc >> 1) ^ (0xEDB8_8320 & mask);
}
}
!crc
}
/// The log window that belongs to the live image.
fn log_window<'a>(device: &'a [u8], layout: &Layout) -> &'a [u8] {
let end = core::cmp::min(
layout.log_off + WAL_HEADER_LEN + layout.log_used,
device.len(),
);
if layout.log_off <= end {
&device[layout.log_off..end]
} else {
&[]
}
}
/// The store the image and its log together describe: the image's records, then the log's
/// entries applied over them. This is what a checkpoint writes and what the digest is taken
/// over, so both go through here rather than each walking the bytes its own way.
fn build_merged(image: &[u8], log: &[u8], header: &Header) -> Result<(Merged, u64), &'static str> {
let mut merged = Merged::new().map_err(|_| "out-of-memory")?;
// A v3 image holds its records behind a space table and an index rather than packed one after
// another, so the store it describes is read through that arithmetic. This is the path a fold
// takes once a store HAS been folded — the first fold reads a packed image and writes an
// addressed one, and every fold after that reads an addressed one — so a reader that only knew
// the packed layout could fold a store exactly once. Found on the box, where that is not a
// hypothesis: the live store's first fold succeeded and its second answered -EINVAL.
if header.version == VERSION_V3 || header.version == VERSION_V4 {
let geometry = cube_format::V3::decode(image).map_err(|_| "bad-tables")?;
let mut row = 0u64;
while row < geometry.space_count {
let (space, first, records) = geometry.space_row(image, row).ok_or("truncated-table")?;
let mut i = 0u64;
while i < records {
let entry = geometry
.index_entry(image, first + i)
.ok_or("truncated-index")?;
let value = geometry
.value(image, &entry)
.ok_or("truncated-image")?;
merged
.add(space, entry.key, entry.flags, value, false)
.map_err(|_| "out-of-memory")?;
i += 1;
}
row += 1;
}
let applied = match apply_log(log, &mut merged)? {
Some(n) => n,
None => 0,
};
return Ok((merged, applied));
}
let extent = header.image_bytes.ok_or("v1-has-no-extent")? as usize;
let extent = core::cmp::min(extent, image.len());
let mut off = if header.version == VERSION_V1 {
HEADER_LEN_V1
} else {
HEADER_LEN_V2
};
let mut seen: u64 = 0;
while off + RECORD_FIXED <= extent {
if let Some(c) = header.record_count {
if seen >= c {
break;
}
}
let space = &image[off..off + SPACE_ID_LEN];
let key = &image[off + SPACE_ID_LEN..off + SPACE_ID_LEN + RAW_KEY_LEN];
let mut word = [0u8; 8];
word.copy_from_slice(&image[off + SPACE_ID_LEN + RAW_KEY_LEN..off + RECORD_FIXED]);
let value_len = u64::from_le_bytes(word) as usize;
let value_at = off + RECORD_FIXED;
if value_at + value_len > extent {
return Err("truncated-image");
}
merged
.add(space, key, 0, &image[value_at..value_at + value_len], false)
.map_err(|_| "out-of-memory")?;
seen += 1;
off = value_at + value_len;
}
let applied = match apply_log(log, &mut merged)? {
Some(n) => n,
None => 0,
};
Ok((merged, applied))
}
/// Digest the store the image and its log together describe.
///
/// `image` is the live image's bytes alone; `log` is exactly the window of log that belongs
/// to it. Returns `Err("no-log")` when there is none, so the caller can fall back to reading
/// the image by itself.
fn merged_digest(image: &[u8], log: &[u8], header: &Header) -> Result<Line, &'static str> {
// An addressed image is a complete store on its own: its index holds every record and says
// where each one is, so a log is an *addition* rather than a requirement. A packed image is
// not — without a log there is nothing this path adds over the image's own order, which is why
// it answers "no log" and lets the caller walk the records it can see.
//
// Requiring a log here regardless of layout is what made a bare v3/v4 image read as a *packed*
// one: the caller fell through to a walk that starts at the v2 header length and reads index
// entries as record frames. It went unnoticed while every userspace-written image was v2, which
// is exactly the kind of bug that a migration finds and a unit test does not.
let addressed = header.version == VERSION_V3 || header.version == VERSION_V4;
if !addressed && (log.len() < WAL_HEADER_LEN || &log[0..4] != WAL_MAGIC) {
return Err("no-log");
}
let (mut merged, applied) = build_merged(image, log, header)?;
merged.sort_entries();
let entries = merged.entries.as_slice();
let mut h = FNV_OFFSET;
let mut count: u64 = 0;
let mut value_bytes: u64 = 0;
let mut i = 0;
while i < entries.len() {
let head = &entries[i];
let mut last = i;
while last + 1 < entries.len()
&& entries[last + 1].space == head.space
&& entries[last + 1].key == head.key
{
last += 1;
}
let winner = &entries[last];
i = last + 1;
if winner.deleted {
continue;
}
let value = merged.value(winner);
h = fnv1a64(&winner.space, h);
h = fnv1a64(&winner.key, h);
h = fnv1a64(&(value.len() as u64).to_le_bytes(), h);
h = fnv1a64(value, h);
count += 1;
value_bytes += value.len() as u64;
}
// `bytes` is the records' logical size, not the size of a particular layout: a v2 image and a
// v3 image of one store hold the same records in different shapes, and this line is what the
// two are compared by, so the figure must not depend on which shape is on disk.
let folded = count * RECORD_FIXED as u64 + value_bytes;
let mut line = Line::new();
// The digest line is field-for-field what `cube-image digest` prints, so the two are
// compared by diff. What the log contributed goes on its own line: a padded log region
// makes a byte count meaningless, an entry count does not.
let _ = write!(
line,
"digest version={} curve={} bytes={} records={} value_bytes={} fnv1a64={:016x} errors=0\n",
header.version, header.curve, folded, count, value_bytes, h
);
let _ = write!(line, "log entries={} applied over the image\n", applied);
Ok(line)
}
/// Where a bare image's log starts, given the image's declared extent.
fn log_offset(image_bytes: u64) -> usize {
let n = image_bytes as usize;
(n + LOG_ALIGN - 1) & !(LOG_ALIGN - 1)
}
/// Length of a bare log's valid prefix: the header plus every entry that parses and passes
/// its checksum. This is the append position, and the reason a torn tail is overwritten
/// rather than appended to.
fn log_valid_len(log: &[u8]) -> Result<usize, &'static str> {
if log.len() < WAL_HEADER_LEN {
return Err("short-log");
}
if &log[0..4] != WAL_MAGIC {
return Err("no-log");
}
if log[4] != WAL_VERSION && log[4] != WAL_VERSION_V2 {
return Err("unsupported-log-version");
}
let (fixed, len_at) = wal_frame(log[4]);
let mut off = WAL_HEADER_LEN;
while off + fixed <= log.len() {
let start = off;
let op = log[off];
if op != 1 && op != 2 {
break;
}
let mut word = [0u8; 4];
word.copy_from_slice(&log[off + 1..off + 5]);
let crc = u32::from_le_bytes(word);
word.copy_from_slice(&log[off + len_at..off + len_at + 4]);
let len = u32::from_le_bytes(word) as usize;
// Only a validated entry moves the append point. Advancing first and checking after
// counts a torn entry as part of the prefix, so the next append lands *after* the
// corruption and buries it — the opposite of the rule that a torn tail is overwritten.
let frame_end = start + fixed + len;
if frame_end > log.len() {
break;
}
if crc32(&log[start + 5..frame_end]) != crc {
break;
}
off = frame_end;
}
Ok(off)
}
/// Morton-encode a point into the store's 24-byte key — the same interleaving
/// `cube-core`'s `Curve for Morton` does, because a key written here has to be the key a
/// userspace reader decodes. Bit `i` of each axis lands at bit `3i` of the 192-bit key,
/// counting from the least significant bit, which lives in the *last* byte.
fn morton_encode(x: u64, y: u64, z: u64) -> [u8; 24] {
// The interleave itself lives in the shared file, and this is now one of its callers rather
// than a second copy of it. That matters more than it looks: the driver and userspace have to
// agree on this key byte for byte, and they were two implementations of one rule whose
// agreement was held only by the gates. `crates/cube-format`'s tests pin the axis-to-bit
// correspondence by hand so that neither side can drift without a test failing.
cube_format::morton_key(x, y, z)
}
/// A mutation to append: the byte-plane write the store contract calls `put`, plus the class
/// mask the caller stamps at write time — the moment the event's class is known for certain.
struct Mutation {
space: [u8; 32],
key: [u8; 24],
flags: u16,
value: KVVec<u8>,
}
/// Build a log entry: `op | crc32 | space | key | [flags] | len | value`.
///
/// `version` decides whether the class mask is in the frame, and it must match the log's header:
/// a reader frames every entry by the header's version, so an entry is only well-formed against the
/// log it is written into. Version 1 has no mask field — the legacy log simply cannot carry one —
/// and an entry going into such a log is written without it rather than not written at all.
fn encode_entry(
op: u8,
space: &[u8; 32],
key: &[u8; 24],
flags: u16,
value: &[u8],
version: u8,
) -> Result<KVVec<u8>, AllocError> {
let flagged = version == WAL_VERSION_V2;
let fixed = if flagged { ENTRY_FIXED_V2 } else { ENTRY_FIXED };
let mut entry = KVVec::<u8>::with_capacity(fixed + value.len(), GFP_KERNEL)?;
entry.extend_from_slice(&[op], GFP_KERNEL)?;
entry.extend_from_slice(&[0u8; 4][..], GFP_KERNEL)?;
entry.extend_from_slice(space, GFP_KERNEL)?;
entry.extend_from_slice(key, GFP_KERNEL)?;
if flagged {
entry.extend_from_slice(&flags.to_le_bytes(), GFP_KERNEL)?;
}
entry.extend_from_slice(&(value.len() as u32).to_le_bytes(), GFP_KERNEL)?;
entry.extend_from_slice(value, GFP_KERNEL)?;
// The checksum covers everything after the crc field, so a v1 entry's checksum covers what a v1
// reader reads — the two versions are different framings of the same rule, not two rules.
let crc = crc32(&entry.as_slice()[5..]);
entry.as_mut_slice()[1..5].copy_from_slice(&crc.to_le_bytes());
Ok(entry)
}
/// Write `bytes` at `off` and flush, or say why not. The one place the module writes.
fn write_and_sync(file: *mut bindings::file, off: u64, bytes: &[u8]) -> Result<(), Error> {
let mut at: bindings::loff_t = off as bindings::loff_t;
// SAFETY: the caller passes a live `struct file *`; `bytes` outlives the call; `at` is a
// valid loff_t.
let wrote = unsafe {
bindings::kernel_write(
file,
bytes.as_ptr().cast::<core::ffi::c_void>(),
bytes.len(),
&mut at,
)
};
if wrote < 0 {
return Err(Error::from_errno(wrote as i32));
}
if wrote as usize != bytes.len() {
return Err(EIO);
}
// Durability before acknowledgement: a short write that was never flushed is a write that
// did not happen.
// SAFETY: the caller passes a live, writable file; datasync is 0 (a full sync).
let rc = unsafe { bindings::vfs_fsync(file, 0) };
if rc < 0 {
return Err(Error::from_errno(rc));
}
Ok(())
}
/// The control copy an update should write: the older of the two, so the current one stays
/// readable throughout.
fn older_copy(generation: u64) -> usize {
if generation % 2 == 0 {
CTL_COPY_A
} else {
CTL_COPY_B
}
}
/// Append a mutation to the log and make it durable, returning the updated control block
/// when the layout has one.
///
/// The order is the durability contract: the entry is written and flushed to the device
/// *before* the caller is told it happened. A mutation this function has returned `Ok` for
/// survives a power cut; one it has not may or may not, which is exactly the boundary an
/// acknowledgement is supposed to mark.
///
/// Append a mutation to the log. **Requires [`STORE_OP`] to be held by the caller.**
///
/// On a store device the entry goes after the log bytes already in use, and the control
/// block's count is raised afterwards — entry first, count second, so a crash between them
/// loses an *unacknowledged* mutation rather than counting one that is not there. A count
/// that ran ahead would make replay read past valid data; a count that lags only forgets.
///
/// That read-then-write pair is what must not interleave — and the lock has to be taken by the
/// caller, *before* it reads the layout this writes against. Taking it here is too late: writers
/// that each read `log_used = 0` before queueing on a lock inside this function still append at
/// the same offset. Measured with the lock here: one 68-byte entry and twelve `ok`s.
/// The bytes an append reads against, and where they sit in the device.
///
/// A store device's write path needs the control block and the log's head, and nothing else: the
/// image is tens of megabytes of records an append never looks at. It used to be handed the whole
/// device, so every put read all of it — 100 MB here — to append a few dozen bytes, and a write cost
/// **60131 ms** on the box while a read cost 12 ms. Measured before this change, ten puts through
/// `cube(2)`: 71, 131, 66, 120, 60, 94, 117, 104, 116, 63 ms.
///
/// The whole-image form stays, because two callers genuinely read it: a *bare* store, whose log runs
/// to the end of the file with nothing recording its length, and the v1→v4 migration, which folds
/// and therefore rewrites the image.
struct AppendSource<'a> {
/// The absolute offset `region` starts at — where the writes go.
at: usize,
/// The log region's bytes, from `at`.
region: &'a [u8],
/// The region's capacity. Comes from the control block when there is one; for a bare store the
/// region is everything to the end of the file, so its own length is the capacity.
capacity: usize,
/// The whole image, for the two callers that have it.
image: Option<&'a [u8]>,
}
impl<'a> AppendSource<'a> {
/// The whole device or image, which is what a bare store's append reads and what the v1
/// migration folds.
fn whole_image(image: &'a [u8], layout: &Layout) -> Self {
// The log starts after the bytes in use on a store device, after the image on a bare one.
let at = core::cmp::min(layout.log_off, image.len());
let region = &image[at..];
AppendSource {
at,
region,
capacity: region.len(),
image: Some(image),
}
}
/// A store device's log head: the control block already gave the offsets and the capacity, so
/// these few bytes are all the append needs to see.
fn log_head(at: usize, region: &'a [u8], capacity: usize) -> Self {
AppendSource {
at,
region,
capacity,
image: None,
}
}
}
/// Everything a device append consults, read without the image: the control block, and the log's
/// head. Both are small, and neither depends on how much the store holds.
struct WriteView {
layout: Layout,
/// The log region's first `WAL_HEADER_LEN` bytes — the header that decides how the entry is
/// framed. A region that has never been written reads as zeros, which is a log with no header
/// rather than a broken one.
head: KVVec<u8>,
capacity: usize,
}
impl WriteView {
fn source(&self) -> AppendSource<'_> {
AppendSource::log_head(self.layout.log_off, self.head.as_slice(), self.capacity)
}
}
/// Read what a write needs: the control block and the log's head, not the image.
///
/// Returns `Ok(None)` for a store with no control block — a bare image — whose append must read the
/// whole thing because nothing there records where its log ends.
fn write_view() -> Result<Option<WriteView>, Error> {
let file = store_file()?;
let mut scratch = KVVec::<u8>::with_capacity(4096, GFP_KERNEL)?;
scratch.resize(4096, 0, GFP_KERNEL)?;
let mut head = KVVec::<u8>::new();
read_exact_at(file, 0, CTL_COPY_B + CTL_COPY_LEN, &mut head, &mut scratch)?;
if head.as_slice().len() < 4 || &head.as_slice()[0..4] != CTL_MAGIC.as_slice() {
return Ok(None);
}
let control = read_control(head.as_slice()).ok_or(EINVAL)?;
let layout = Layout {
image_off: control.image_off() as usize,
log_off: control.log_off() as usize,
log_used: control.log_used as usize,
control: Some(control),
};
let mut log_head = KVVec::<u8>::new();
read_exact_at(file, control.log_off(), WAL_HEADER_LEN, &mut log_head, &mut scratch)?;
Ok(Some(WriteView {
layout,
head: log_head,
capacity: control.log_capacity as usize,
}))
}
/// Fold the log because it has grown past [`LOG_FOLD_BYTES`].
///
/// A fold rewrites the image, so unlike every other write this one reads the whole store. That is
/// the point of doing it here and rarely, rather than letting the cost sit on the read path.
fn fold_for_headroom(log_used: u64) -> Result<(), Error> {
let mut image = KVVec::<u8>::new();
read_image(&mut image)?;
let layout = resolve_layout(image.as_slice()).map_err(|what| {
pr_err!("cubelinux: {}\n", what);
EINVAL
})?;
fold_now(image.as_slice(), &layout)?;
pr_info!(
"cubelinux: folded the log at {} bytes, because every read pays for it\n",
log_used
);
Ok(())
}
/// Append one mutation, reading only what the store's shape requires.
///
/// This is the entry point every write uses. It exists so that the bounded read and the whole-image
/// fallback are one decision in one place, rather than four callers each choosing.
fn append_mutation(m: &Mutation, op: u8) -> Result<Option<Control>, Error> {
append_mutation_inner(m, op, true)
}
/// `may_fold` is false on the way back from a fold, so a log that did not shrink cannot make this
/// recurse for ever: one fold per append, whether or not it helped.
fn append_mutation_inner(m: &Mutation, op: u8, may_fold: bool) -> Result<Option<Control>, Error> {
match write_view()? {
Some(view) => {
if may_fold && view.layout.log_used as u64 > LOG_FOLD_BYTES {
let used = view.layout.log_used as u64;
fold_for_headroom(used)?;
// The fold emptied the log, so the write re-reads what it needs and appends to a
// fresh one — once only, so a fold that did not shrink anything cannot loop.
return append_mutation_inner(m, op, false);
}
let source = view.source();
append(&source, &view.layout, m, op)
}
// A bare store: no control block, and a log whose end is the file's end.
None => {
let (image, layout) = device_and_layout()?;
append(&AppendSource::whole_image(image.as_slice(), &layout), &layout, m, op)
}
}
}
fn append(
source: &AppendSource<'_>,
layout: &Layout,
m: &Mutation,
op: u8,
) -> Result<Option<Control>, Error> {
// Where it goes: after the bytes in use on a store device, after the valid prefix on a
// bare image, where nothing records the length.
let region_at = source.at;
let region = source.region;
let (used, capacity) = match layout.control {
// `log_used` counts entry bytes *after* the log's header, so the header is not
// something an append has to account for in the count.
Some(c) => (c.log_used as usize, c.log_capacity as usize),
None => {
// A region with no header is a log that has never been written; it is empty, not
// broken, and the header goes down before the first entry.
let used = match log_valid_len(region) {
Ok(valid) => valid - WAL_HEADER_LEN,
Err("no-log") | Err("short-log") => 0,
Err(what) => {
pr_err!("cubelinux: log region unreadable: {}\n", what);
return Err(EINVAL);
}
};
(used, source.capacity)
}
};
// The entry's frame has to match the log's header, because the header is what a reader frames
// it by. So the header decides the version, and there are exactly three cases:
//
// * a log that already holds v1 entries must keep taking v1 entries — rewriting the header
// would mis-frame every entry already there. A store *device* folds first instead (that is
// the v4 migration, and it has a spare slot to fold into); a bare image has nowhere to fold
// to, so it keeps writing the log it has. The mask is lost for that append, because a v1
// entry has no field to carry it — a limitation of the legacy layout stated plainly, and
// better than refusing to write to a log the store already holds data in;
// * no header, a v2 header, or a v1 header over an *empty* log: the frame is v2, and the
// header is written (or upgraded) to say so.
let header_version = if region.len() >= WAL_HEADER_LEN && &region[0..4] == WAL_MAGIC {
Some(region[4])
} else {
None
};
let entry_version = match header_version {
// A log that already holds entries and declares v1 keeps taking v1 — rewriting the header
// would mis-frame everything already there. A store *device* first folds instead (that is
// the v4 migration, and it has a spare slot to fold into); a bare image has nowhere to fold
// to, so it keeps writing the log it has.
Some(WAL_VERSION) if used > 0 => {
if layout.control.is_some() {
// [`STORE_OP`] is held by the op that reached here, so these are plain calls.
//
// The v1→v4 migration is the one device append that needs the image, because a fold
// rewrites it. Pay for the whole read here — once per store, at the migration — and
// keep it off the path every other write takes.
let mut read_here = KVVec::<u8>::new();
match source.image {
Some(image) => fold_now(image, layout)?,
None => {
read_image(&mut read_here)?;
let layout = resolve_layout(read_here.as_slice()).map_err(|what| {
pr_err!("cubelinux: {}\n", what);
EINVAL
})?;
fold_now(read_here.as_slice(), &layout)?;
}
}
let (device, layout) = device_and_layout()?;
return append(
&AppendSource::whole_image(device.as_slice(), &layout),
&layout,
m,
op,
);
}
WAL_VERSION
}
// An empty log that already declares a version: a device's is upgraded to v2, since v2 is
// the layout that carries the class mask and the one a fold turns into a v4 index. A bare
// image keeps what its log says.
Some(v) => {
if layout.control.is_some() {
WAL_VERSION_V2
} else {
v
}
}
// No header at all: this append writes the first one, and so chooses the framing for this
// log's life. A store device gets v2. A bare image gets v1, because the bare layout is the
// legacy one: it has no index, so a mask written into its log could only ever be read back
// by a scan and never be checkpointed — half a feature, bought by making every existing
// reader of that layout grow a version it has no use for. The format that carries a class
// mask is the checkpointable one, and that is the one that gets it.
None => {
if layout.control.is_some() {
WAL_VERSION_V2
} else {
WAL_VERSION
}
}
};
let entry = encode_entry(op, &m.space, &m.key, m.flags, m.value.as_slice(), entry_version)?;
if WAL_HEADER_LEN + used + entry.len() > capacity {
pr_err!(
"cubelinux: the log is full ({} of {} bytes); checkpoint before appending\n",
used,
capacity
);
return Err(ENOSPC);
}
// SAFETY: store_device() is a NUL-terminated C string filled at boot; filp_open returns a
// valid file or an error pointer, which is checked. O_RDWR is 2.
let file = store_file()?;
let mut result: Result<(), Error> = Ok(());
// A log region that has never been written is zeros, not a log: give it a header, the same
// thing the userspace log does when its file does not exist. A header that does not already
// say the version this entry is framed by is (re)written to say it, because the header is what
// a reader frames the entry by — and an entry under a header that disagrees with it is not a
// record, it is a mis-framing.
if region.len() < WAL_HEADER_LEN || &region[0..4] != WAL_MAGIC || region[4] != entry_version {
let mut hdr = [0u8; WAL_HEADER_LEN];
hdr[0..4].copy_from_slice(&WAL_MAGIC);
hdr[4] = entry_version;
hdr[5] = 0; // morton
result = write_and_sync(file, layout.log_off as u64, &hdr);
}
if result.is_ok() {
result = write_and_sync(
file,
(region_at + WAL_HEADER_LEN + used) as u64,
entry.as_slice(),
);
}
// The count second, so replay can never be told about an entry that is not there.
let mut updated = None;
if result.is_ok() {
if let Some(c) = layout.control {
let next = Control {
log_used: c.log_used + entry.len() as u64,
generation: c.generation + 1,
..c
};
let mut buf = [0u8; CTL_COPY_LEN];
next.encode(&mut buf);
result = write_and_sync(file, older_copy(c.generation) as u64, &buf);
if result.is_ok() {
updated = Some(next);
}
}
}
result.map(|_| updated)
}
/// Serialize the merged store into the pinned v2 image.
///
/// The winners are collected first because the header states the image's extent and its
/// record count, and those are only known once the coalescing is done.
fn serialize_image(merged: &mut Merged, curve: u8) -> Result<KVVec<u8>, AllocError> {
merged.sort_entries();
let entries = merged.entries.as_slice();
// The winners first: one live record per coordinate, in the store's own (space, key) order.
let mut winners = KVVec::<u32>::new();
let mut values_len: u64 = 0;
let mut i = 0;
while i < entries.len() {
let head = &entries[i];
let mut last = i;
while last + 1 < entries.len()
&& entries[last + 1].space == head.space
&& entries[last + 1].key == head.key
{
last += 1;
}
if !entries[last].deleted {
winners.push(last as u32, GFP_KERNEL)?;
values_len += entries[last].value_len as u64;
}
i = last + 1;
}
// The space table: where each space's records begin in the index, and how many. It comes from
// the same order the index is written in, so a walk of the store is a walk of this table.
let mut spaces = KVVec::<([u8; SPACE_ID_LEN], u64, u64)>::new();
for (at, w) in winners.as_slice().iter().enumerate() {
let space = merged.entries.as_slice()[*w as usize].space;
match spaces.as_mut_slice().last_mut() {
Some((last_space, _, records)) if *last_space == space => *records += 1,
_ => spaces.push((space, at as u64, 1), GFP_KERNEL)?,
}
}
let header = HeaderV3 {
version: VERSION_V4,
record_count: winners.len() as u64,
space_count: spaces.len() as u64,
index_off: (HEADER_LEN_V3 + spaces.len() * SPACE_ENTRY) as u64,
values_off: (HEADER_LEN_V3 + spaces.len() * SPACE_ENTRY + winners.len() * INDEX_ENTRY_V4) as u64,
image_bytes: 0,
};
let image_bytes = header.values_off + values_len;
let mut out = KVVec::<u8>::with_capacity(image_bytes as usize, GFP_KERNEL)?;
let mut header = header;
header.image_bytes = image_bytes;
header.encode(&mut out, curve)?;
for (space, first, records) in spaces.as_slice() {
out.extend_from_slice(space, GFP_KERNEL)?;
out.extend_from_slice(&first.to_le_bytes(), GFP_KERNEL)?;
out.extend_from_slice(&records.to_le_bytes(), GFP_KERNEL)?;
}
// The index, whose entries are the addresses: a fixed stride is what makes `index_off +
// i * stride` a place a reader can go to without reading anything before it. v4's stride is
// two wider, carrying the class mask beside the address — never inside the value.
let mut value_at = header.values_off;
for w in winners.as_slice() {
let e = &merged.entries.as_slice()[*w as usize];
out.extend_from_slice(&e.key, GFP_KERNEL)?;
out.extend_from_slice(&e.flags.to_le_bytes(), GFP_KERNEL)?;
out.extend_from_slice(&value_at.to_le_bytes(), GFP_KERNEL)?;
// The length is written as 8 bytes even though `Entry` carries it in 4: an index entry
// has a fixed stride and the header's `values_off` is computed from that, so a shorter
// field here would leave the last entries overlapping the values. Both readers assume the
// stride the header declares, which is the point of a fixed stride.
out.extend_from_slice(&(e.value_len as u64).to_le_bytes(), GFP_KERNEL)?;
value_at += e.value_len as u64;
}
// The values, packed in index order: a batch of records is one span of this, which is why a
// listing can read what it returns in one go rather than record by record.
for w in winners.as_slice() {
let e = &merged.entries.as_slice()[*w as usize];
out.extend_from_slice(merged.value(e), GFP_KERNEL)?;
}
Ok(out)
}
/// Fold the log into the image: write the merged store into the spare slot, then make it
/// live.
///
/// The order is what makes this atomic. The new image goes to the slot nothing is reading and
/// is flushed; only then does the control block change, and that change is one small write to
/// the control copy that is *not* current. At every instant there is either the old image with
/// a log that still describes the mutations since it, or the new image with an empty log. A
/// crash in the middle of the copy leaves the previous store intact, which is the entire
/// reason for two slots.
fn checkpoint(ctl: &Control, merged: &mut Merged) -> Result<u64, Error> {
let new_image = serialize_image(merged, ctl.curve)?;
if new_image.len() as u64 > ctl.slot_bytes {
pr_err!(
"cubelinux: checkpoint needs {} bytes but a slot holds {}\n",
new_image.len(),
ctl.slot_bytes
);
return Err(ENOSPC);
}
// SAFETY: as in `append`.
let file = store_file()?;
let mut result = write_and_sync(file, ctl.spare_off(), new_image.as_slice());
if result.is_ok() {
let next = Control {
log_used: 0,
generation: ctl.generation + 1,
active_slot: 1 - ctl.active_slot,
..*ctl
};
let mut buf = [0u8; CTL_COPY_LEN];
next.encode(&mut buf);
result = write_and_sync(file, older_copy(ctl.generation) as u64, &buf);
}
result.map(|_| new_image.len() as u64)
}
/// Fold the current log into the image — the shared body of `sync` and the append-time
/// migration. After it, the image is the pinned v4 shape and the log is empty, so the next
/// append starts a fresh v2 log.
///
/// **Requires [`STORE_OP`] to be held by the caller.**
fn fold_now(image: &[u8], layout: &Layout) -> Result<(), Error> {
let ctl = match layout.control {
Some(c) => c,
None => return Err(EINVAL), // a bare image has no spare slot to fold into
};
let live = &image[layout.image_off..];
let header = match parse_header(live) {
Ok(h) => h,
Err(_) => return Err(EINVAL),
};
let window = log_window(image, layout);
let extent = core::cmp::min(
header.image_bytes.unwrap_or(live.len() as u64) as usize,
live.len(),
);
let mut merged = match build_merged(&live[..extent], window, &header) {
Ok((m, _)) => m,
Err(_) => return Err(EINVAL),
};
checkpoint(&ctl, &mut merged)?;
Ok(())
}
/// Everything a syscall needs to reach the store: read the device, resolve its layout, and
/// hand back what was asked for.
///
/// These are the `cube(2)` entry points as the Rust side exposes them. The C shim
/// (cube_syscall.c) owns the user copies and the argument validation; nothing here sees a
/// userspace pointer.
fn device_and_layout() -> Result<(KVVec<u8>, Layout), Error> {
let mut device = KVVec::<u8>::new();
read_image(&mut device)?;
let layout = resolve_layout(&device).map_err(|what| {
pr_err!("cubelinux: cannot resolve the store layout: {}\n", what);
EINVAL
})?;
Ok((device, layout))
}
/// The coordinate's key, and the space, as the store addresses them.
unsafe fn coord_key(space: *const u8, x: u64, y: u64, z: u64) -> ([u8; 32], [u8; 24]) {
let mut sp = [0u8; 32];
// SAFETY: the caller (the syscall shim) passes a pointer to 32 bytes it has already
// copied from userspace into kernel memory.
unsafe { core::ptr::copy_nonoverlapping(space, sp.as_mut_ptr(), 32) };
(sp, morton_encode(x, y, z))
}
// ── The kernel's record of this boot ───────────────────────────────────────────────────
/// Push bytes, or report that the buffer would not grow.
fn push(out: &mut KVVec<u8>, bytes: &[u8]) -> Option<()> {
out.extend_from_slice(bytes, GFP_KERNEL).ok()
}
/// Push a decimal integer. The value is built in a `KVVec`, which is not a `fmt::Write`, and giving
/// the boot record its own output type to format one number into would be a second mechanism for
/// something the digest line already solved differently.
fn push_dec(out: &mut KVVec<u8>, mut n: u64) -> Option<()> {
let mut tmp = [0u8; 20];
let mut at = tmp.len();
loop {
at -= 1;
tmp[at] = b'0' + (n % 10) as u8;
n /= 10;
if n == 0 {
break;
}
}
push(out, &tmp[at..])
}
/// How long a NUL-terminated byte string is, up to `max`.
fn c_len(p: *const u8, max: usize) -> usize {
let mut n = 0;
while n < max {
// SAFETY: the caller guarantees `p` points at `max` readable bytes.
if unsafe { *p.add(n) } == 0 {
break;
}
n += 1;
}
n
}
/// The kernel's own account of the boot it is having, as one line:
/// `boot=<seconds since the epoch> device=<the store it resolved> kernel=<its version banner>`.
///
/// The banner is last and unquoted because it contains spaces, so everything after the final `=`
/// is the kernel's own words rather than a field this code parsed. The time is raw epoch seconds:
/// rendering a calendar date in the kernel is date arithmetic, and a caller with a clock can do it
/// without a kernel bug being the reason a timestamp is wrong.
fn boot_record_value() -> Option<KVVec<u8>> {
let mut out = KVVec::<u8>::new();
let mut ts = bindings::timespec64 {
tv_sec: 0,
tv_nsec: 0,
};
// SAFETY: `ts` is a live, writable `timespec64`, and the call retains no pointer to it.
unsafe { bindings::ktime_get_real_ts64(&mut ts) };
push(&mut out, b"boot=")?;
push_dec(&mut out, ts.tv_sec as u64)?;
push(&mut out, b" device=")?;
// SAFETY: `store_device` returns a static NUL-terminated buffer, valid for the life of the
// kernel; 256 is the length of the buffer the C side fills.
let dev = store_device();
push(&mut out, unsafe {
core::slice::from_raw_parts(dev, c_len(dev, 256))
})?;
push(&mut out, b" kernel=")?;
// SAFETY: `linux_banner` is a static NUL-terminated string the kernel defines. `addr_of!` takes
// its address without creating a reference to a `[c_char; 0]`, which is how bindgen declares an
// array whose size is not in the header. The bound is generous: the banner is one line.
let banner = unsafe {
let p = core::ptr::addr_of!(bindings::linux_banner).cast::<u8>();
core::slice::from_raw_parts(p, c_len(p, 512))
};
// The banner ends in a newline. A record that carries it would make every reader strip one, so
// it is dropped here, where it is known to be a terminator rather than content.
let banner = match banner.last() {
Some(b'\n') => &banner[..banner.len() - 1],
_ => banner,
};
push(&mut out, banner)?;
Some(out)
}
/// Record this boot, at a write, and try again if the store would not take it.
///
/// Called from every entry point that writes. It is not called from the read paths, so a kernel that
/// only reads leaves no record — see the note above `BOOT_SPACE`. A failure here is logged and never
/// propagated: a record of the boot is worth having, and is not a precondition for the caller's
/// write.
///
/// It retries, up to [`BOOT_RECORD_MAX_ATTEMPTS`], because "a store that will not take it will not
/// take it later either" is not true of the failure it actually had: the first write of a boot can
/// arrive before the store is reachable by the credentials of the process that made it. That is not
/// a property of the store, and treating it as one cost a boot its record silently.
///
/// What it costs: the control block and the log's head, through the same bounded read every other
/// write uses. It used to resolve the whole device — tens of megabytes — for the record and then
/// the caller resolved it again for its own mutation, which is why a retry was not affordable then
/// and is now.
fn ensure_boot_record() {
// SAFETY: a flag the C side set from the kernel command line during boot.
if !unsafe { cubelinux_boot_record_enabled() } {
return;
}
// Claim one attempt with a compare-exchange, so two callers racing into their first write
// cannot both write a record: the loser returns, exactly as the swap did. The winner's attempt
// is counted whether it succeeds or fails, which is what makes a failure retryable.
let made = BOOT_RECORD_ATTEMPTS.load(Ordering::Acquire);
if made >= BOOT_RECORD_MAX_ATTEMPTS
|| BOOT_RECORD_ATTEMPTS
.compare_exchange(made, made + 1, Ordering::AcqRel, Ordering::Acquire)
.is_err()
{
return;
}
let value = match boot_record_value() {
Some(v) => v,
None => {
pr_warn!("cubelinux: not enough memory to build this boot's record\n");
return;
}
};
let mutation = Mutation {
space: BOOT_SPACE,
key: morton_encode(BOOT_POINT.0, BOOT_POINT.1, BOOT_POINT.2),
flags: BOOT_FLAGS,
value,
};
// The record goes down the same bounded path every other write takes: the control block and the
// log's head, not the image.
match append_mutation(&mutation, 1) {
Ok(_) => {
// Recorded. Park the count at the ceiling so no later write asks again.
BOOT_RECORD_ATTEMPTS.store(BOOT_RECORD_MAX_ATTEMPTS, Ordering::Release);
pr_info!("cubelinux: recorded this boot in the store\n");
}
Err(_) => pr_warn!("cubelinux: the store would not take this boot's record\n"),
}
}
/// `CUBE_OP_PUT`: store bytes at a coordinate, under a class mask.
///
/// The mask is the writer's and is stamped here, once — this is the moment the record's class is
/// known for certain, and every later reader is spared re-deriving it. It means nothing to this
/// side: which bits are which class is a vocabulary's business, and a kernel that interpreted one
/// would be inventing a vocabulary.
///
/// # Safety
/// `space` must point to 32 readable bytes; `value` to `len` readable bytes when `len` is
/// non-zero.
#[unsafe(no_mangle)]
pub unsafe extern "C" fn cubelinux_kernel_put(
space: *const u8,
x: u64,
y: u64,
z: u64,
value: *const u8,
len: usize,
flags: u16,
) -> i32 {
// The lock spans the whole op — reading the layout and committing the count are one
// read-modify-write, and `ensure_boot_record` below writes too.
let _op = STORE_OP.lock();
ensure_boot_record();
let (sp, key) = unsafe { coord_key(space, x, y, z) };
let mut bytes = KVVec::<u8>::new();
if len > 0 {
if let Err(_) = bytes.extend_from_slice(
// SAFETY: the shim guarantees `value` holds `len` bytes.
unsafe { core::slice::from_raw_parts(value, len) },
GFP_KERNEL,
) {
return -12; // -ENOMEM
}
}
let mutation = Mutation {
space: sp,
key,
flags,
value: bytes,
};
match append_mutation(&mutation, 1) {
Ok(_) => 0,
Err(e) => -(e.to_errno() as i32),
}
}
// ── Reading the store where it lies ─────────────────────────────────────────────────────
//
// A read used to start by reading the whole device into kernel memory, parsing every record into
// a merged pool, and heapsorting the lot — then throwing all of it away. That is why a `get` of
// one record cost as much as listing the store, and why the coordinate bought nothing
// mechanically: knowing where a record is did not make reaching it any cheaper.
//
// The records are already the index. A checkpoint writes them in `(space, key)` order, which is
// the order the coordinate itself computes, and the log that can override them is small. So the
// kernel walks the records where they lie — zero-copy, no pool, no sort — and consults the log as
// an overlay. What an operation costs is proportional to the records it actually touches.
/// How much is read at a time when a length is not known ahead of it.
const VIEW_CHUNK: usize = 256 * 1024;
/// Read exactly `len` bytes at `off`. A short read is an error, not a smaller view: the store says
/// how long its image is, and quietly reading less would turn a truncated store into a shorter one
/// that looks complete.
fn read_exact_at(
file: *mut bindings::file,
off: u64,
len: usize,
out: &mut KVVec<u8>,
scratch: &mut KVVec<u8>,
) -> Result<()> {
out.clear();
out.reserve(len, GFP_KERNEL)?;
let mut pos: bindings::loff_t = off as bindings::loff_t;
let mut left = len;
while left > 0 {
let want = core::cmp::min(left, scratch.len());
// SAFETY: `file` is a live `struct file *`; `scratch` is a kernel buffer we own; `pos` is
// a valid loff_t. kernel_read copies at most `want` bytes and retains neither pointer.
let n = unsafe {
bindings::kernel_read(file, scratch.as_mut_ptr().cast::<core::ffi::c_void>(), want, &mut pos)
};
if n < 0 {
return Err(Error::from_errno(n as i32));
}
if n == 0 {
return Err(EINVAL);
}
let n = n as usize;
out.extend_from_slice(&scratch.as_slice()[..n], GFP_KERNEL)?;
left -= n;
}
Ok(())
}
/// Read from `off` to the end, or to `cap` bytes, whichever comes first.
///
/// For a bare image, whose extent nothing states, and for a log whose length is implied.
fn read_to_end_at(
file: *mut bindings::file,
off: u64,
out: &mut KVVec<u8>,
scratch: &mut KVVec<u8>,
cap: usize,
) -> Result<()> {
out.clear();
let mut pos: bindings::loff_t = off as bindings::loff_t;
while out.len() < cap {
let want = core::cmp::min(scratch.len(), cap - out.len());
// SAFETY: as in `read_exact_at`.
let n = unsafe {
bindings::kernel_read(file, scratch.as_mut_ptr().cast::<core::ffi::c_void>(), want, &mut pos)
};
if n < 0 {
return Err(Error::from_errno(n as i32));
}
if n == 0 {
break;
}
let n = n as usize;
out.extend_from_slice(&scratch.as_slice()[..n], GFP_KERNEL)?;
}
Ok(())
}
/// The bytes one operation needs: the image, and the log beside it.
///
/// The device is four images wide plus its log; reading the device to read the image copies
/// several times what is used, on every call.
struct View {
image: KVVec<u8>,
log: KVVec<u8>,
header: Header,
/// The control block's generation, which a checkpoint changes — the key to what is cached
/// about this image. Zero for a bare image, which has no control block and does not change
/// underneath a reader.
generation: u64,
}
kernel::sync::global_lock! {
/// Where each space's records begin.
///
/// A walk has to start at its space. Records are variable-length, so "where does space S begin"
/// cannot be computed from the key — it can only be found by reading the image once. Finding it
/// once per checkpoint is what makes a listing cost what it returns, and the table is one entry
/// per space (tens of bytes), not one per record. It goes stale on exactly one event: a
/// checkpoint rewriting the image, which changes the control block's generation.
///
/// SAFETY: Initialized in the module initializer before first use.
unsafe(uninit) static SPACE_STARTS: Mutex<Option<SpaceStarts>> = None;
}
struct SpaceStarts {
generation: u64,
image_len: usize,
/// Each space's first record, in the image's order.
starts: KVVec<([u8; SPACE_ID_LEN], usize)>,
}
/// Where the live image's spaces begin, from the cache when it is current and by reading the image
/// when it is not.
fn space_starts(view: &View) -> Result<KVVec<([u8; SPACE_ID_LEN], usize)>> {
let mut guard = SPACE_STARTS.lock();
if let Some(cached) = guard.as_ref() {
if cached.generation == view.generation && cached.image_len == view.image().len() {
let mut out = KVVec::new();
out.extend_from_slice(cached.starts.as_slice(), GFP_KERNEL)?;
return Ok(out);
}
}
let mut starts = KVVec::new();
let mut records = Records::new(view.image(), &view.header);
let mut previous: Option<[u8; SPACE_ID_LEN]> = None;
loop {
let at = records.off;
match records.next() {
Some((space, _, _)) => {
if previous != Some(*space) {
starts.push((*space, at), GFP_KERNEL)?;
previous = Some(*space);
}
}
None => break,
}
}
let mut out = KVVec::new();
out.extend_from_slice(starts.as_slice(), GFP_KERNEL)?;
*guard = Some(SpaceStarts {
generation: view.generation,
image_len: view.image().len(),
starts,
});
Ok(out)
}
impl View {
fn image(&self) -> &[u8] {
self.image.as_slice()
}
fn log(&self) -> &[u8] {
self.log.as_slice()
}
/// Where one space's records begin, if the image holds it.
fn space_start(&self, space: &[u8; SPACE_ID_LEN]) -> Result<Option<usize>> {
let starts = space_starts(self)?;
for (candidate, at) in starts.as_slice() {
if candidate == space {
return Ok(Some(*at));
}
}
Ok(None)
}
}
/// Read the image and its log window — not the whole device.
fn read_view() -> Result<View> {
// O_RDONLY is 0 in Linux; filp_open takes the raw flags word.
// SAFETY: `store_device()` is a NUL-terminated C string the module parameter filled at boot,
// and filp_open returns a valid `struct file *` or an error pointer, checked here.
let file = store_file()?;
let view = read_view_from(file);
let view = view?;
if view.image.len() < HEADER_LEN_V1 {
return Err(EINVAL);
}
Ok(view)
}
fn read_view_from(file: *mut bindings::file) -> Result<View> {
let mut scratch = KVVec::<u8>::with_capacity(VIEW_CHUNK, GFP_KERNEL)?;
scratch.resize(VIEW_CHUNK, 0, GFP_KERNEL)?;
// The control block, or the image's own header when the device is a bare image.
let mut head = KVVec::<u8>::with_capacity(CTL_COPY_B + CTL_COPY_LEN, GFP_KERNEL)?;
read_exact_at(file, 0, CTL_COPY_B + CTL_COPY_LEN, &mut head, &mut scratch)?;
let is_device = &head.as_slice()[0..4] == CTL_MAGIC.as_slice();
let (image_off, log_off, log_used, generation) = if is_device {
let control = read_control(head.as_slice()).ok_or(EINVAL)?;
(
control.image_off() as usize,
control.log_off() as usize,
control.log_used as usize,
control.generation,
)
} else {
(0, 0, 0, 0)
};
// What the image says about itself. A device keeps its image at an offset, so its header is
// read from there; a bare image starts with it.
let mut header = if is_device {
let mut raw = KVVec::<u8>::with_capacity(HEADER_LEN_V2, GFP_KERNEL)?;
match read_exact_at(file, image_off as u64, HEADER_LEN_V2, &mut raw, &mut scratch) {
Ok(()) => parse_header(raw.as_slice()),
// A v1 image is shorter than a v2 header, and saying "not a store" about it would be
// wrong: read the six bytes it does have.
Err(_) => {
read_exact_at(file, image_off as u64, HEADER_LEN_V1, &mut raw, &mut scratch)?;
parse_header(raw.as_slice())
}
}
} else {
parse_header(head.as_slice())
}
.map_err(|what| {
pr_err!("cubelinux: {}\n", what);
EINVAL
})?;
let mut image = KVVec::<u8>::new();
let mut log = KVVec::<u8>::new();
match (is_device, header.image_bytes) {
(true, Some(bytes)) => {
let bytes = bytes as usize;
if bytes > MAX_BYTES {
return Err(EINVAL);
}
read_exact_at(file, image_off as u64, bytes, &mut image, &mut scratch)?;
}
// A v1 image states no extent, so it is read to the end — bounded, because a device that
// never ends must not become a read that never ends.
(true, None) => {
read_to_end_at(file, image_off as u64, &mut image, &mut scratch, MAX_BYTES)?;
}
(false, _) => {
read_to_end_at(file, 0, &mut image, &mut scratch, MAX_BYTES)?;
if header.image_bytes.is_none() {
header.image_bytes = Some(image.len() as u64);
}
}
}
if is_device {
if log_used > 0 {
let len = core::cmp::min(WAL_HEADER_LEN + log_used, MAX_BYTES);
read_exact_at(file, log_off as u64, len, &mut log, &mut scratch)?;
}
} else {
// A bare image's log starts after the image, aligned, and runs to the end.
let at = log_offset(image.len() as u64);
read_to_end_at(file, at as u64, &mut log, &mut scratch, MAX_BYTES)?;
}
Ok(View {
image,
log,
header,
generation,
})
}
/// Walk a store image's records, in the order the image holds them.
///
/// Zero-copy: the value a caller gets is a slice of the image already in memory, so listing a
/// store does not copy its values anywhere.
struct Records<'a> {
image: &'a [u8],
off: usize,
end: usize,
left: Option<u64>,
}
impl<'a> Records<'a> {
/// Begin at `off`, for a caller that has been told where a space starts.
///
/// The record count is dropped: it is a bound on the records ahead of the *first* record, and
/// starting anywhere else makes it a bound that no longer means anything. The image's extent
/// is the bound that still holds.
fn at(image: &'a [u8], header: &Header, off: usize) -> Self {
let mut records = Records::new(image, header);
if off > records.off && off <= records.end {
records.off = off;
records.left = None;
}
records
}
fn new(image: &'a [u8], header: &Header) -> Self {
let header_len = if header.version == VERSION_V1 {
HEADER_LEN_V1
} else {
HEADER_LEN_V2
};
let extent = match header.image_bytes {
Some(n) => core::cmp::min(n as usize, image.len()),
None => image.len(),
};
Records {
image,
off: core::cmp::min(header_len, extent),
end: extent,
left: header.record_count,
}
}
fn next(&mut self) -> Option<(&'a [u8; 32], &'a [u8; 24], &'a [u8])> {
if self.left == Some(0) || self.off + RECORD_FIXED > self.end {
return None;
}
let frame = &self.image[self.off..self.off + RECORD_FIXED];
// A v1 image states no count, so padding ends the walk — but only if every remaining byte
// is zero, because a record at the origin with an empty value is 64 zero bytes and is a
// record.
if self.left.is_none()
&& frame.iter().all(|b| *b == 0)
&& self.image[self.off..].iter().all(|b| *b == 0)
{
return None;
}
let mut word = [0u8; 8];
word.copy_from_slice(&frame[SPACE_ID_LEN + RAW_KEY_LEN..RECORD_FIXED]);
let value_len = u64::from_le_bytes(word) as usize;
let value_at = self.off + RECORD_FIXED;
if value_at + value_len > self.end {
return None;
}
let space: &[u8; SPACE_ID_LEN] = frame[..SPACE_ID_LEN].try_into().ok()?;
let key: &[u8; RAW_KEY_LEN] =
frame[SPACE_ID_LEN..SPACE_ID_LEN + RAW_KEY_LEN].try_into().ok()?;
let value = &self.image[value_at..value_at + value_len];
self.off = value_at + value_len;
if let Some(n) = self.left.as_mut() {
*n -= 1;
}
Some((space, key, value))
}
}
/// What the log says about one coordinate: written, or removed.
enum Effect<'a> {
/// The bytes it wrote, and the class mask it wrote them under.
Write(&'a [u8], u16),
Delete,
}
/// The log's newest word on `(space, key)`, if it says anything about it.
///
/// The log is the only thing that can override the image — it holds what has been written since
/// the last checkpoint — so this is what decides a read.
fn log_effect<'a>(log: &'a [u8], space: &[u8; SPACE_ID_LEN], key: &[u8; RAW_KEY_LEN]) -> Option<Effect<'a>> {
let mut entries = log_entries(log).ok()?;
let mut found = None;
while let Some((entry_space, entry_key, op, flags, value_at, value_len)) = entries.next() {
if entry_space == space && entry_key == key {
found = Some(if op == 2 {
Effect::Delete
} else {
Effect::Write(&log[value_at..value_at + value_len], flags)
});
}
}
found
}
/// One log entry, as the walk sees it: `(space, key, op, value)`.
struct LogEntries<'a> {
log: &'a [u8],
off: usize,
/// The entry stride for this log's version, and where its length sits.
fixed: usize,
len_at: usize,
/// Whether this log's entries carry a class mask (version 2). A version 1 entry has no mask
/// field, and reads as "no class" — the same rule the index uses for a v3 entry.
flagged: bool,
}
impl<'a> LogEntries<'a> {
/// The next entry: `(space, key, op, flags, where its value starts, its value's length)`.
///
/// The value's *position* rather than a slice, because a caller collecting edits keeps them in
/// a fixed-size vector and a borrow of the log would tie that vector's type to the log's life.
fn next(&mut self) -> Option<(&'a [u8], &'a [u8], u8, u16, usize, usize)> {
if self.off + self.fixed > self.log.len() {
return None;
}
let start = self.off;
let op = self.log[start];
if op != 1 && op != 2 {
return None;
}
let mut word = [0u8; 4];
word.copy_from_slice(&self.log[start + 1..start + 5]);
let crc = u32::from_le_bytes(word);
let space = &self.log[start + 5..start + 37];
let key = &self.log[start + 37..start + 61];
word.copy_from_slice(&self.log[start + self.len_at..start + self.len_at + 4]);
let len = u32::from_le_bytes(word) as usize;
let frame_end = start + self.fixed + len;
if frame_end > self.log.len() {
return None;
}
// The checksum covers space, key, the mask, length and value, so a torn tail is stopped
// at rather than applied — the same rule the fold uses.
if crc32(&self.log[start + 5..frame_end]) != crc {
return None;
}
let flags = if self.flagged {
cube_format::le_u16(self.log, start + 5 + SPACE_ID_LEN + RAW_KEY_LEN)
} else {
0
};
self.off = frame_end;
Some((space, key, op, flags, start + self.fixed, len))
}
}
/// The log's entries, when there is a valid log to read.
fn log_entries(log: &[u8]) -> Result<LogEntries<'_>, &'static str> {
if log.len() < WAL_HEADER_LEN || &log[0..4] != WAL_MAGIC {
return Err("no-log");
}
if log[4] != WAL_VERSION && log[4] != WAL_VERSION_V2 {
return Err("unsupported-log-version");
}
let (fixed, len_at) = wal_frame(log[4]);
Ok(LogEntries {
log,
off: WAL_HEADER_LEN,
fixed,
len_at,
flagged: log[4] == WAL_VERSION_V2,
})
}
/// One space's override of the image, on its way to being merged with it.
///
/// The key and the value borrow the log rather than copying it, so a walker can hand out slices of
/// bytes that are already in memory — a listing never copies a value.
#[derive(Clone, Copy)]
struct LogEdit<'a> {
key: &'a [u8; RAW_KEY_LEN],
/// Position in the log, so the later of two edits to one key is the one that counts.
seq: u32,
value: &'a [u8],
/// The class mask the writer stamped, carried through the merge so a flag scan sees the log's
/// word on a record and not only the image's.
flags: u16,
deleted: bool,
}
/// Every edit the log holds for one space, in key order, later edits last.
///
/// The image and the log are both in key order, so one merge pass over them is a listing — and a
/// space is what a walk names, so only its own edits are collected.
fn log_edits<'a>(
log: &'a [u8],
space: &[u8; SPACE_ID_LEN],
out: &mut KVVec<LogEdit<'a>>,
) -> Result<()> {
out.clear();
let mut entries = match log_entries(log) {
Ok(e) => e,
Err(_) => return Ok(()),
};
let mut seq: u32 = 0;
while let Some((edit_space, key, op, flags, value_at, value_len)) = entries.next() {
if edit_space == space {
let key: &[u8; RAW_KEY_LEN] = match key.try_into() {
Ok(k) => k,
Err(_) => break,
};
out.push(
LogEdit {
key,
seq,
value: &log[value_at..value_at + value_len],
flags,
deleted: op == 2,
},
GFP_KERNEL,
)?;
}
seq += 1;
}
sort_edits(out.as_mut_slice());
Ok(())
}
fn sort_edits(edits: &mut [LogEdit<'_>]) {
let len = edits.len();
if len < 2 {
return;
}
let mut start = len / 2;
while start > 0 {
start -= 1;
sift_down_edits(edits, start, len);
}
let mut end = len;
while end > 1 {
end -= 1;
edits.swap(0, end);
sift_down_edits(edits, 0, end);
}
}
/// Heapsort, for the same reason the entry sort is one: a constant stack frame, so the size of a
/// log cannot decide whether reading it is safe. `sort_unstable` is used nowhere in this module.
fn sift_down_edits(edits: &mut [LogEdit<'_>], mut root: usize, end: usize) {
loop {
let mut child = root * 2 + 1;
if child >= end {
return;
}
if child + 1 < end
&& (edits[child].key, edits[child].seq) < (edits[child + 1].key, edits[child + 1].seq)
{
child += 1;
}
if (edits[root].key, edits[root].seq) >= (edits[child].key, edits[child].seq) {
return;
}
edits.swap(root, child);
root = child;
}
}
/// The live records of one space: the image's records for it, with the log's edits applied, in key
/// order.
///
/// This is the whole of what a listing is — two sorted sequences and one pass — and it is why a
/// walk costs what it returns rather than what the store holds.
struct SpaceWalker<'a> {
records: Records<'a>,
edits: KVVec<LogEdit<'a>>,
edit_at: usize,
image_next: Option<(&'a [u8; RAW_KEY_LEN], &'a [u8])>,
wanted: [u8; SPACE_ID_LEN],
}
impl<'a> SpaceWalker<'a> {
fn new(view: &'a View, wanted: &[u8; SPACE_ID_LEN]) -> Result<Self> {
let mut edits = KVVec::<LogEdit<'a>>::new();
log_edits(view.log(), wanted, &mut edits)?;
// Start where the space starts: a listing of a small space in a large store must not read
// the records that are not in it.
let mut records = match view.space_start(wanted)? {
Some(at) => Records::at(view.image(), &view.header, at),
None => {
let mut empty = Records::new(view.image(), &view.header);
empty.off = empty.end; // nothing in the image: the log's edits are the space
empty
}
};
let image_next = next_in_space(&mut records, wanted);
Ok(SpaceWalker {
records,
edits,
edit_at: 0,
image_next,
wanted: *wanted,
})
}
/// The next live record, or nothing when the space is walked out.
/// The next live record: `(key, its class, its value)`.
///
/// The class comes back because a listing has to be able to say what a record *is*. A packed
/// image's records carry no mask field, so they read as "no class" — but a record the log
/// supplied carries the mask its writer stamped, and dropping it here is what left a
/// kernel-backed store unable to answer `entries_flagged`.
fn next(&mut self) -> Option<(&'a [u8; RAW_KEY_LEN], u16, &'a [u8])> {
loop {
// The entries of one key arrive in log order, so the last of a key's group is the
// newest word on it and the earlier ones are history.
while self.edit_at + 1 < self.edits.len()
&& self.edits[self.edit_at + 1].key == self.edits[self.edit_at].key
{
self.edit_at += 1;
}
let edit = self.edits.as_slice().get(self.edit_at).copied();
match (edit, self.image_next) {
// The log's word on a coordinate the image holds: it wins.
(Some(e), Some((key, _))) if e.key == key => {
self.edit_at += 1;
self.image_next = next_in_space(&mut self.records, &self.wanted);
if !e.deleted {
return Some((e.key, e.flags, e.value));
}
}
// A record only the log holds, in its key's place.
(Some(e), Some((key, value))) if e.key < key => {
self.edit_at += 1;
self.image_next = Some((key, value));
if !e.deleted {
return Some((e.key, e.flags, e.value));
}
}
// The image's own record, which the log says nothing about. A packed image has no
// mask field, so it answers none.
(_, Some((key, value))) => {
self.image_next = next_in_space(&mut self.records, &self.wanted);
return Some((key, 0, value));
}
// The image is walked out; the rest of the log still has records to report.
(Some(e), None) => {
self.edit_at += 1;
if !e.deleted {
return Some((e.key, e.flags, e.value));
}
}
(None, None) => return None,
}
}
}
}
/// Does this space hold at least one live record?
///
/// A space whose every record the log has removed is not a space to report, and a walk that stops
/// at the first record it finds is the cheapest way to know: it costs one record, not a listing.
fn space_has_records(view: &View, space: &[u8; SPACE_ID_LEN]) -> bool {
match SpaceWalker::new(view, space) {
Ok(mut walker) => walker.next().is_some(),
Err(_) => false,
}
}
/// The distinct spaces the log writes into, ascending. Small by construction: the log is what has
/// been written since the last checkpoint.
fn log_spaces(log: &[u8], out: &mut KVVec<[u8; SPACE_ID_LEN]>) -> Result<()> {
out.clear();
let mut entries = match log_entries(log) {
Ok(e) => e,
Err(_) => return Ok(()),
};
while let Some((space, _, _, _, _, _)) = entries.next() {
let mut s = [0u8; SPACE_ID_LEN];
s.copy_from_slice(&space[..SPACE_ID_LEN]);
if !out.as_slice().contains(&s) {
out.push(s, GFP_KERNEL)?;
}
}
sort_spaces(out.as_mut_slice());
Ok(())
}
/// Heapsort for a store's spaces. The log holds few, but `sort_unstable` is not used anywhere in
/// this module: it allocates about 3.5 KiB of kernel stack per level, which is how reading a
/// store once took the machine down.
fn sort_spaces(spaces: &mut [[u8; SPACE_ID_LEN]]) {
let len = spaces.len();
if len < 2 {
return;
}
let mut start = len / 2;
while start > 0 {
start -= 1;
sift_down_spaces(spaces, start, len);
}
let mut end = len;
while end > 1 {
end -= 1;
spaces.swap(0, end);
sift_down_spaces(spaces, 0, end);
}
}
fn sift_down_spaces(spaces: &mut [[u8; SPACE_ID_LEN]], mut root: usize, end: usize) {
loop {
let mut child = root * 2 + 1;
if child >= end {
return;
}
if child + 1 < end && spaces[child] < spaces[child + 1] {
child += 1;
}
if spaces[root] >= spaces[child] {
return;
}
spaces.swap(root, child);
root = child;
}
}
/// Hand the caller a value: its length, or -ENOENT when there is nothing there.
///
/// When the value does not fit in `len` the bytes are *not* copied and the length is returned
/// anyway, so the caller can size its buffer and ask again — a short read that silently truncated
/// would be worse than an error.
fn copy_out(value: &[u8], buf: *mut u8, len: usize) -> isize {
if value.len() > len {
return value.len() as isize;
}
if !value.is_empty() {
// SAFETY: the shim guarantees `buf` holds `len` bytes and `len >= value.len()`.
unsafe {
core::ptr::copy_nonoverlapping(value.as_ptr(), buf, value.len());
}
}
value.len() as isize
}
/// The next record of one space in an image walk, skipping the records of every other space.
fn next_in_space<'a>(
records: &mut Records<'a>,
wanted: &[u8; SPACE_ID_LEN],
) -> Option<(&'a [u8; RAW_KEY_LEN], &'a [u8])> {
while let Some((space, key, value)) = records.next() {
if space == wanted {
return Some((key, value));
}
}
None
}
/// Packs a walk's records into the caller's buffer, skipping what the cursor has covered.
struct Batch<'a> {
out: &'a mut [u8],
written: usize,
returned: u64,
seen: u64,
cursor: u64,
/// Set when a record did not fit: the size it needs, for the `-ERANGE` answer.
too_big: usize,
/// Whether each frame carries its own space.
///
/// A walk's frame leaves the space out because the caller named it — repeating 32 bytes per
/// record to say what the caller already said is a tax on every listing. A scan that names *no*
/// space cannot do that: a coordinate is meaningless without its space, so an every-space
/// answer has to say which space each record came from or the answer is unusable. Hence a frame
/// that depends on the scope, which is the scope's business and not the record's.
with_space: bool,
}
/// The fixed part of a walk's frame, before the value.
fn flagged_frame(with_space: bool) -> usize {
RAW_KEY_LEN + FLAGS_LEN + 4 + if with_space { SPACE_ID_LEN } else { 0 }
}
impl Batch<'_> {
/// Offer one record, class and all. `false` means nothing more will fit in this batch.
///
/// One frame for every walk. It used to be two — `key | len | value` for the plain walks and
/// `key | flags | len | value` for the scan — which meant a caller could not ask a listing what
/// class its records were, and a *store* could not answer `entries_flagged` from a kernel store
/// at all. A walk that cannot say what a record is forces a second reader of the format to find
/// out, and that is the thing this whole arrangement exists to avoid.
///
/// `space` is carried only when `with_space` is set — see [`Batch::with_space`].
fn offer(&mut self, space: &[u8; SPACE_ID_LEN], key: &[u8], flags: u16, value: &[u8]) -> bool {
self.seen += 1;
if self.seen <= self.cursor {
return true;
}
match pack_flagged(space, key, flags, value, self.out, self.written, self.with_space) {
Some(n) => {
self.written += n;
self.returned += 1;
true
}
None => {
self.too_big = flagged_frame(self.with_space) + value.len();
false
}
}
}
}
// ── The region walk: a box, its key span, and what decides membership ────────────────────
/// The inclusive box a region walk answers, and the key span that bounds it.
///
/// The span is the whole mechanism of `CUBE_OP_RANGE`: [`cube_format::key_span`] gives the two
/// corner keys, and **every** key in the box lies between them because the interleave is monotone on
/// each axis. So a sorted index can be binary-searched for the foot and read forward to the head.
///
/// It is a **bound and not the set**: a point *outside* the box can have a key inside the span (the
/// classic Z-order amplification), so [`holds`](Self::holds) is what decides and a caller must ask
/// it about every candidate. Confusing the bound for the set is the mistake that would make this
/// return records from outside the region the caller named — and there is no decomposition here to
/// blame it on, which is why the pair is checked in `cube-format`'s tests.
struct Region {
lo: [u64; 3],
hi: [u64; 3],
foot: [u8; RAW_KEY_LEN],
head: [u8; RAW_KEY_LEN],
}
impl Region {
/// `None` for an empty box — inverted on some axis. An empty box has no span worth building,
/// and building one would produce a range whose start is above its end.
fn new(lo: [u64; 3], hi: [u64; 3]) -> Option<Self> {
let mut i = 0;
while i < 3 {
if lo[i] > hi[i] {
return None;
}
i += 1;
}
let (foot, head) = cube_format::key_span(lo, hi);
Some(Region { lo, hi, foot, head })
}
/// Whether a key's point lies in the box — the test that keeps a bound from becoming a wrong
/// answer. This is why `morton_decode` exists.
fn holds(&self, key: &[u8; RAW_KEY_LEN]) -> bool {
let p = cube_format::morton_decode(key);
p[0] >= self.lo[0]
&& p[0] <= self.hi[0]
&& p[1] >= self.lo[1]
&& p[1] <= self.hi[1]
&& p[2] >= self.lo[2]
&& p[2] <= self.hi[2]
}
/// Whether a key sorts below the span, inside it, or above it.
///
/// A dedicated three-way rather than an `Ordering`, because `Ordering` has no room for it: a key
/// between the foot and the head is `Less` *than the head* and that is exactly the case a walk
/// must keep. Returning `key_cmp(key, head)` and matching on it conflates "inside the span" with
/// "below the foot", which silently skips every record in the region the caller asked for —
/// a mistake this code made for one edit, caught before it built.
fn place(&self, key: &[u8; RAW_KEY_LEN]) -> Span {
if cube_format::key_cmp(key, &self.foot) == core::cmp::Ordering::Less {
Span::Below
} else if cube_format::key_cmp(key, &self.head) == core::cmp::Ordering::Greater {
Span::Above
} else {
Span::Inside
}
}
}
/// Where a key sits relative to a region's key span. See [`Region::place`].
enum Span {
/// Sorts before the span's foot: not in the box, and skippable on a layout with no index.
Below,
/// Sorts within the span: a candidate, which still has to pass [`Region::holds`].
Inside,
/// Sorts past the span's head: the walk is finished, because the records are key-ordered.
Above,
}
// ── v3: the addressed reader ────────────────────────────────────────────────────────────
//
// A v3 image is read at the offsets the header describes, so an operation reads the bytes it names
// and nothing else. This is what the coordinate finally buys: a lookup is a binary search over a
// fixed stride, a listing starts where its space starts, and neither reads the store to find out.
/// A v3 image open for reading, with the few bytes an operation needs read on demand.
struct Addressed {
file: *mut bindings::file,
/// Where the image begins in the device.
image_off: u64,
header: HeaderV3,
/// The log window, which is what can override the image and is small.
log: KVVec<u8>,
/// Reused for every span read, so a lookup does not allocate per probe.
scratch: KVVec<u8>,
}
impl Addressed {
/// Open the store and read its control block, v3 header, and log window — not its records.
fn open() -> Result<Option<Self>> {
let file = store_file()?;
let mut opened = Addressed {
file,
image_off: 0,
header: HeaderV3 {
version: VERSION_V3,
image_bytes: 0,
record_count: 0,
space_count: 0,
index_off: 0,
values_off: 0,
},
log: KVVec::new(),
scratch: KVVec::new(),
};
opened.scratch = KVVec::<u8>::with_capacity(4096, GFP_KERNEL)?;
opened.scratch.resize(4096, 0, GFP_KERNEL)?;
let mut head = KVVec::<u8>::with_capacity(CTL_COPY_B + CTL_COPY_LEN, GFP_KERNEL)?;
read_exact_at(file, 0, CTL_COPY_B + CTL_COPY_LEN, &mut head, &mut opened.scratch)?;
if &head.as_slice()[0..4] != CTL_MAGIC.as_slice() {
return Ok(None); // a bare image, which the walking reader handles
}
let control = read_control(head.as_slice()).ok_or(EINVAL)?;
opened.image_off = control.image_off() as u64;
let mut raw = KVVec::<u8>::with_capacity(HEADER_LEN_V3, GFP_KERNEL)?;
read_exact_at(file, opened.image_off, HEADER_LEN_V3, &mut raw, &mut opened.scratch)?;
if raw.as_slice()[4] != VERSION_V3 && raw.as_slice()[4] != VERSION_V4 {
return Ok(None); // v1 or v2: the walking reader
}
opened.header = HeaderV3::decode(raw.as_slice()).map_err(|what| {
pr_err!("cubelinux: {}\n", what);
EINVAL
})?;
if control.log_used > 0 {
let len = core::cmp::min(WAL_HEADER_LEN + control.log_used as usize, MAX_BYTES);
read_exact_at(file, control.log_off() as u64, len, &mut opened.log, &mut opened.scratch)?;
}
Ok(Some(opened))
}
/// Read `len` bytes at an offset *within the image*.
fn read_image_at(&mut self, off: u64, len: usize, out: &mut KVVec<u8>) -> Result<()> {
if off + len as u64 > self.header.image_bytes {
return Err(EINVAL);
}
read_exact_at(self.file, self.image_off + off, len, out, &mut self.scratch)
}
/// The space table entry for a space: where its records start in the index, and how many.
///
/// The table is sorted by space and fixed-size, so this is a binary search over arithmetic
/// addresses — the same reason the index works.
fn space_entry(&mut self, space: &[u8; SPACE_ID_LEN]) -> Result<Option<(u64, u64)>> {
// One buffer for the whole search, not one per probe: a binary search over a thousand
// spaces allocates a thousand times otherwise, and an allocator stall is a latency spike
// with no bound on it — the opposite of what a control loop is allowed to see.
let mut buf = KVVec::<u8>::new();
let mut lo = 0u64;
let mut hi = self.header.space_count;
while lo < hi {
let mid = lo + (hi - lo) / 2;
let at = HEADER_LEN_V3 as u64 + mid * SPACE_ENTRY as u64;
self.read_image_at(at, SPACE_ENTRY, &mut buf)?;
let entry = buf.as_slice();
let found = &entry[..SPACE_ID_LEN];
if found < &space[..] {
lo = mid + 1;
} else if found > &space[..] {
hi = mid;
} else {
let mut w = [0u8; 8];
w.copy_from_slice(&entry[SPACE_ID_LEN..SPACE_ID_LEN + 8]);
let first = u64::from_le_bytes(w);
w.copy_from_slice(&entry[SPACE_ID_LEN + 8..SPACE_ENTRY]);
return Ok(Some((first, u64::from_le_bytes(w))));
}
}
Ok(None)
}
/// One space-table row by its position: the space, where its records start, and how many.
///
/// [`space_entry`](Self::space_entry) answers "where is *this* space" by binary search, which is
/// what a caller who named a space wants. A caller who named none has to visit every space in
/// the table's own order — which is sorted, so the rows come back in the order a checkpoint
/// wrote them and an every-space walk answers in (space, key) order for free.
fn space_row(&mut self, row: u64) -> Result<Option<([u8; SPACE_ID_LEN], u64, u64)>> {
if row >= self.header.space_count {
return Ok(None);
}
let mut buf = KVVec::<u8>::new();
self.read_image_at(HEADER_LEN_V3 as u64 + row * SPACE_ENTRY as u64, SPACE_ENTRY, &mut buf)?;
let entry = buf.as_slice();
let space: [u8; SPACE_ID_LEN] = match entry[..SPACE_ID_LEN].try_into() {
Ok(s) => s,
Err(_) => return Err(EINVAL),
};
let mut w = [0u8; 8];
w.copy_from_slice(&entry[SPACE_ID_LEN..SPACE_ID_LEN + 8]);
let first = u64::from_le_bytes(w);
w.copy_from_slice(&entry[SPACE_ID_LEN + 8..SPACE_ENTRY]);
Ok(Some((space, first, u64::from_le_bytes(w))))
}
/// One index entry, by its position in the index: `(value_off, value_len, flags)`.
///
/// The class mask comes back with the address because it is in the same stride and costs
/// nothing extra to read — and because a flag scan is exactly a walk that reads it and nothing
/// else. A v3 entry has no mask and answers zero, the same way its log entries do.
fn index_entry(&mut self, at: u64, buf: &mut KVVec<u8>) -> Result<(u64, u64, u16)> {
let stride = self.header.stride();
let off = self.header.index_off + at * stride as u64;
self.read_image_at(off, stride, buf)?;
let entry = buf.as_slice();
let flagged = self.header.version == VERSION_V4;
let flags = if flagged {
cube_format::le_u16(entry, RAW_KEY_LEN)
} else {
0
};
let vo = if flagged { RAW_KEY_LEN + FLAGS_LEN } else { RAW_KEY_LEN };
let mut w = [0u8; 8];
w.copy_from_slice(&entry[vo..vo + 8]);
let value_off = u64::from_le_bytes(w);
w.copy_from_slice(&entry[vo + 8..vo + 16]);
Ok((value_off, u64::from_le_bytes(w), flags))
}
/// The address of `key` inside a space's index range, or nothing if the space does not hold it:
/// `(value_off, value_len, flags)`.
///
/// The mask comes back with the address because it is in the same stride the search already
/// read, so a read that wants both does not search twice for one entry.
fn find(
&mut self,
key: &[u8; RAW_KEY_LEN],
first: u64,
records: u64,
) -> Result<Option<(u64, u64, u16)>> {
// One buffer for the whole search — see the note in `space_entry`.
let mut buf = KVVec::<u8>::new();
let mut lo = 0u64;
let mut hi = records;
while lo < hi {
let mid = lo + (hi - lo) / 2;
let stride = self.header.stride();
let off = self.header.index_off + (first + mid) * stride as u64;
self.read_image_at(off, stride, &mut buf)?;
let found = &buf.as_slice()[..RAW_KEY_LEN];
if found < &key[..] {
lo = mid + 1;
} else if found > &key[..] {
hi = mid;
} else {
return self.index_entry(first + mid, &mut buf).map(Some);
}
}
Ok(None)
}
/// The first index position at or **after** `key`, within a space's index range.
///
/// This is `find` answering the other question. `find` asks "where is this exact key" and is only
/// useful for a coordinate read; a region walk asks "where does this span begin", which is the
/// same binary search over the same arithmetic addresses with `>=` in place of `==`. The walk
/// then reads forward from here and stops when it passes the span's head, so the cost is the
/// span rather than the space — and it is the reason the span needed no decomposition.
fn lower_bound(
&mut self,
key: &[u8; RAW_KEY_LEN],
first: u64,
records: u64,
) -> Result<u64> {
// One buffer for the whole search — see the note in `space_entry`.
let mut buf = KVVec::<u8>::new();
let mut lo = 0u64;
let mut hi = records;
while lo < hi {
let mid = lo + (hi - lo) / 2;
let stride = self.header.stride();
let off = self.header.index_off + (first + mid) * stride as u64;
self.read_image_at(off, stride, &mut buf)?;
if &buf.as_slice()[..RAW_KEY_LEN] < &key[..] {
lo = mid + 1;
} else {
hi = mid;
}
}
Ok(first + lo)
}
/// The value of an index entry, read from where the index says it is.
fn value_at(&mut self, value_off: u64, value_len: u64) -> Result<KVVec<u8>> { let mut out = KVVec::<u8>::new();
self.read_image_at(value_off, value_len as usize, &mut out)?;
Ok(out)
}
/// Write the log's window back into `buf` — used to hand the log to the merge helpers, which
/// take a slice of a log that has already been read.
fn log(&self) -> &[u8] {
self.log.as_slice()
}
}
/// `CUBE_OP_GET` against a v3 image: the log's newest word, else a binary search.
///
/// # Safety
/// `buf` must hold `len` writable bytes.
unsafe fn v3_get(
mut image: Addressed,
sp: [u8; SPACE_ID_LEN],
key: [u8; RAW_KEY_LEN],
buf: *mut u8,
len: usize,
out_flags: *mut u16,
) -> isize {
// The log is owned for the duration: what it says borrows it, and reading the image needs
// `&mut image`. It is what a checkpoint has not folded, so it is small.
let mut log = KVVec::<u8>::new();
if log.extend_from_slice(image.log(), GFP_KERNEL).is_err() {
return -12; // -ENOMEM
}
match log_effect(log.as_slice(), &sp, &key) {
Some(Effect::Delete) => return -2, // -ENOENT
Some(Effect::Write(value, flags)) => {
// SAFETY: the caller guarantees a writable u16.
unsafe { *out_flags = flags };
return copy_out(value, buf, len);
}
None => {}
}
let (first, records) = match image.space_entry(&sp) {
Ok(Some(entry)) => entry,
Ok(None) => return -2,
Err(e) => return -(e.to_errno() as isize),
};
let (value_off, value_len, flags) = match image.find(&key, first, records) {
Ok(Some(entry)) => entry,
Ok(None) => return -2,
Err(e) => return -(e.to_errno() as isize),
};
let value = match image.value_at(value_off, value_len) {
Ok(v) => v,
Err(e) => return -(e.to_errno() as isize),
};
// The record's own class, from the index entry the address came from. A v3 entry has no mask
// and answers zero, which is what "no class" means everywhere here.
// SAFETY: the caller guarantees a writable u16.
unsafe { *out_flags = flags };
copy_out(value.as_slice(), buf, len)
}
/// `CUBE_OP_SPACES` against a v3 image: the space table, and whatever only the log writes.
///
/// # Safety
/// `space_out` must point to 32 writable bytes.
unsafe fn v3_spaces(mut image: Addressed, cursor: u64, space_out: *mut u8) -> i32 {
let mut log_only = KVVec::<[u8; SPACE_ID_LEN]>::new();
if log_spaces(image.log(), &mut log_only).is_err() {
return -12;
}
// The table is sorted and fixed-size: the cursor is an index into the *store's* spaces, which
// is the table's spaces merged with the log's. Reading the table a row at a time is a few dozen
// reads of 48 bytes, and it keeps the merge in one order.
let mut index: u64 = 0;
let mut log_at = 0usize;
let mut buf = KVVec::<u8>::new();
let mut row: u64 = 0;
while row < image.header.space_count {
if image.read_image_at(HEADER_LEN_V3 as u64 + row * SPACE_ENTRY as u64, SPACE_ENTRY, &mut buf).is_err() {
return -12;
}
let mut space = [0u8; SPACE_ID_LEN];
space.copy_from_slice(&buf.as_slice()[..SPACE_ID_LEN]);
while log_at < log_only.len() && log_only.as_slice()[log_at] < space {
let candidate = log_only.as_slice()[log_at];
log_at += 1;
if !space_has_records_v3(&mut image, &candidate) {
continue;
}
if index == cursor {
// SAFETY: the caller guarantees 32 writable bytes at `space_out`.
unsafe { core::ptr::copy_nonoverlapping(candidate.as_ptr(), space_out, SPACE_ID_LEN) };
return 0;
}
index += 1;
}
if log_at < log_only.len() && log_only.as_slice()[log_at] == space {
log_at += 1;
}
row += 1;
if !space_has_records_v3(&mut image, &space) {
continue;
}
if index == cursor {
// SAFETY: as above.
unsafe { core::ptr::copy_nonoverlapping(space.as_ptr(), space_out, SPACE_ID_LEN) };
return 0;
}
index += 1;
}
while log_at < log_only.len() {
let candidate = log_only.as_slice()[log_at];
log_at += 1;
if !space_has_records_v3(&mut image, &candidate) {
continue;
}
if index == cursor {
// SAFETY: as above.
unsafe { core::ptr::copy_nonoverlapping(candidate.as_ptr(), space_out, SPACE_ID_LEN) };
return 0;
}
index += 1;
}
-2 // -ENOENT: no such space; the walk is finished
}
/// Does a v3 image's space hold a live record? One index entry, not a listing.
fn space_has_records_v3(image: &mut Addressed, space: &[u8; SPACE_ID_LEN]) -> bool {
match image.space_entry(space) {
Ok(Some((_, records))) => records > 0,
_ => false,
}
}
/// `CUBE_OP_ENUM` against a v3 image.
///
/// With no log edits for the space — the ordinary case, and always the case just after a
/// checkpoint — the cursor is a count of records already returned and the index is fixed-size, so
/// the batch starts at `first + cursor`: arithmetic, and the cost is the records returned. When the
/// log does hold edits, they interleave with the image's records and the walk has to merge them, so
/// it streams the index from the space's start and counts.
///
/// # Safety
/// `buf` must hold `cap` writable bytes; `out_len` and `out_cursor` must point to writable `u64`s.
unsafe fn v3_enum(
mut image: Addressed,
wanted: [u8; SPACE_ID_LEN],
cursor: u64,
buf: *mut u8,
cap: usize,
out_len: *mut u64,
out_cursor: *mut u64,
) -> i32 {
let (first, records) = match image.space_entry(&wanted) {
Ok(Some(entry)) => entry,
Ok(None) => (0, 0),
Err(e) => return -(e.to_errno() as i32),
};
// Own the log: the edits borrow it, and walking the image needs `&mut image`.
let mut log = KVVec::<u8>::new();
if log.extend_from_slice(image.log(), GFP_KERNEL).is_err() {
return -12; // -ENOMEM
}
let mut edits = KVVec::<LogEdit<'_>>::new();
if log_edits(log.as_slice(), &wanted, &mut edits).is_err() {
return -12;
}
// SAFETY: the shim guarantees `cap` writable bytes at `buf`.
let out = unsafe { core::slice::from_raw_parts_mut(buf, cap) };
/*
* The two paths skip the already-returned records in two different ways, and each must do it
* exactly once.
*
* A batch's skip is `offer`'s `seen`, which counts the records it has been shown and starts at
* 0 — the cursor-th record is the first one offered that is returned. The merge path below has
* nothing but that count: it re-walks the space from its start every batch, because the log's
* edits interleave and the index position of the cursor-th match is not arithmetic.
*
* The no-edits path does not need it: it starts at the index entry the cursor names
* (`first + cursor`), which *is* the skip, exact and in one step. So it hands `offer` a cursor
* of 0 and lets the arithmetic do it — doing both would skip twice and lose records.
*
* What this used to do instead, on the merge path, was start `seen` *at* the cursor. `offer`
* compares after incrementing, so the first record it saw counted as `cursor + 1`, nothing was
* ever skipped, and every call re-served the space from its first record while `out_cursor`
* advanced by the records returned. A walk whose only end signal is a cursor that stops moving
* therefore never ended: a listing looped until the caller ran out of memory.
*/
let (written, returned, too_big) = if edits.is_empty() {
let mut batch = Batch {
out,
written: 0,
returned: 0,
seen: 0,
cursor: 0,
too_big: 0,
with_space: false,
};
// Straight through the index, from the record the cursor names, reading a page of index
// entries at a time and then the span of values they point at.
//
// The cursor is the caller's number, so it is added the saturating way: a wrapped sum
// would land back near the space's first record and re-serve the walk from there — the
// same endless listing this arithmetic exists to avoid.
let mut entry = KVVec::<u8>::new();
let mut at = first.saturating_add(cursor);
while at < first + records {
let (value_off, value_len, flags) = match image.index_entry(at, &mut entry) {
Ok(triple) => triple,
Err(e) => return -(e.to_errno() as i32),
};
let value = match image.value_at(value_off, value_len) {
Ok(v) => v,
Err(e) => return -(e.to_errno() as i32),
};
let key: &[u8; RAW_KEY_LEN] = match entry.as_slice()[..RAW_KEY_LEN].try_into() {
Ok(k) => k,
Err(_) => return -22,
};
if !batch.offer(&wanted, key, flags, value.as_slice()) {
break;
}
at += 1;
}
(batch.written, batch.returned, batch.too_big)
} else {
let mut batch = Batch {
out,
written: 0,
returned: 0,
seen: 0,
cursor,
too_big: 0,
with_space: false,
};
// The log's edits interleave, so the merged order is what the cursor counts. Stream the
// space's index entries rather than all of them: the space is what is being listed.
let mut entry = KVVec::<u8>::new();
let mut at = first;
let mut edit_at = 0usize;
loop {
while edit_at + 1 < edits.len() && edits[edit_at + 1].key == edits[edit_at].key {
edit_at += 1;
}
let edit = edits.as_slice().get(edit_at).copied();
let image_entry = if at < first + records {
match image.index_entry(at, &mut entry) {
Ok(pair) => Some(pair),
Err(e) => return -(e.to_errno() as i32),
}
} else {
None
};
let image_key: Option<&[u8; RAW_KEY_LEN]> = match image_entry {
Some(_) => entry.as_slice()[..RAW_KEY_LEN].try_into().ok(),
None => None,
};
match (edit, image_entry, image_key) {
(Some(e), Some((_, _, _)), Some(key)) if e.key == key => {
edit_at += 1;
at += 1;
if !e.deleted && !batch.offer(&wanted, e.key, e.flags, e.value) {
break;
}
}
(Some(e), Some((_, _, _)), Some(key)) if e.key < key => {
edit_at += 1;
if !e.deleted && !batch.offer(&wanted, e.key, e.flags, e.value) {
break;
}
}
(_, Some((value_off, value_len, flags)), Some(key)) => {
let value = match image.value_at(value_off, value_len) {
Ok(v) => v,
Err(e) => return -(e.to_errno() as i32),
};
at += 1;
if !batch.offer(&wanted, key, flags, value.as_slice()) {
break;
}
}
(Some(e), None, _) => {
edit_at += 1;
if !e.deleted && !batch.offer(&wanted, e.key, e.flags, e.value) {
break;
}
}
(None, None, _) => break,
_ => break,
}
}
(batch.written, batch.returned, batch.too_big)
};
// SAFETY: both out-pointers are writable under this function's contract.
unsafe {
if too_big > 0 && written == 0 {
*out_len = too_big as u64;
*out_cursor = cursor;
return -34; // -ERANGE
}
*out_len = written as u64;
*out_cursor = cursor + returned;
}
0
}
/// `CUBE_OP_RANGE` against a v3 image: **seek** the index to the span's foot, then read forward.
///
/// This is `v3_enum` with two differences, and both are the operation's whole point:
///
/// * it starts at the **span's foot** rather than at the space's first record, found by a binary
/// search over the fixed stride, so the cost is the span and not the space;
/// * it decides with [`Region::holds`] rather than returning everything, and stops the moment the
/// index walks past the span's head, because the index is sorted by key.
///
/// The log's edits interleave exactly as they do in a walk, so this is the same merge — with one
/// added step, which is the log's half of the seek: an edit sorting below the foot cannot be in the
/// box, because the span bounds every key in it.
///
/// Records **examined** can exceed records **returned**, and that excess is the honest cost of a
/// bound: the span contains keys of points outside the box, and `holds` rejects them.
///
/// # Safety
/// `buf` must hold `cap` writable bytes; `out_len` and `out_cursor` must point to writable `u64`s.
unsafe fn v3_range(
mut image: Addressed,
wanted: [u8; SPACE_ID_LEN],
region: Region,
cursor: u64,
buf: *mut u8,
cap: usize,
out_len: *mut u64,
out_cursor: *mut u64,
) -> i32 {
let (first, records) = match image.space_entry(&wanted) {
Ok(Some(entry)) => entry,
Ok(None) => (0, 0),
Err(e) => return -(e.to_errno() as i32),
};
// Own the log: the edits borrow it, and walking the image needs `&mut image`.
let mut log = KVVec::<u8>::new();
if log.extend_from_slice(image.log(), GFP_KERNEL).is_err() {
return -12; // -ENOMEM
}
let mut edits = KVVec::<LogEdit<'_>>::new();
if log_edits(log.as_slice(), &wanted, &mut edits).is_err() {
return -12;
}
// The log's half of the seek.
let mut edit_at = 0usize;
while edit_at < edits.len()
&& cube_format::key_cmp(edits.as_slice()[edit_at].key, &region.foot)
== core::cmp::Ordering::Less
{
edit_at += 1;
}
// The index's half.
let mut at = match image.lower_bound(&region.foot, first, records) {
Ok(p) => p,
Err(e) => return -(e.to_errno() as i32),
};
// SAFETY: the shim guarantees `cap` writable bytes at `buf`.
let out = unsafe { core::slice::from_raw_parts_mut(buf, cap) };
let mut batch = Batch {
out,
written: 0,
returned: 0,
// `seen` is what skips the `cursor` records already returned, and it must start at 0: this
// path re-seeks to the span's foot every batch (there is no `at = first + cursor` here, and
// there cannot be — a span holds records that are *not* returned, so the index position of
// the cursor-th match is not arithmetic), so `offer` is the only thing that counts. Starting
// it at `cursor` would re-serve the span from its foot forever.
seen: 0,
cursor,
too_big: 0,
with_space: false,
};
let mut entry = KVVec::<u8>::new();
loop {
// The entries of one key arrive in log order, so the last of a key's group is the newest
// word on it and the earlier ones are history.
while edit_at + 1 < edits.len()
&& edits.as_slice()[edit_at + 1].key == edits.as_slice()[edit_at].key
{
edit_at += 1;
}
let edit = edits.as_slice().get(edit_at).copied();
let image_entry = if at < first + records {
match image.index_entry(at, &mut entry) {
Ok(pair) => Some(pair),
Err(e) => return -(e.to_errno() as i32),
}
} else {
None
};
let image_key: Option<&[u8; RAW_KEY_LEN]> = match image_entry {
Some(_) => entry.as_slice()[..RAW_KEY_LEN].try_into().ok(),
None => None,
};
// The smallest key still ahead of either stream. If it is past the head, the walk is done:
// the index is sorted by key, so nothing further can be inside the span. This is what makes
// the seek stop at the span instead of reading to the end of the space.
let next_key = match (edit.map(|e| e.key), image_key) {
(Some(e), Some(k)) => {
if e < k {
e
} else {
k
}
}
(Some(e), None) => e,
(None, Some(k)) => k,
(None, None) => break,
};
if cube_format::key_cmp(next_key, &region.head) == core::cmp::Ordering::Greater {
break;
}
match (edit, image_entry, image_key) {
// The log's word on a coordinate the image holds: it wins.
(Some(e), Some(_), Some(key)) if e.key == key => {
edit_at += 1;
at += 1;
if !e.deleted
&& region.holds(e.key)
&& !batch.offer(&wanted, e.key, e.flags, e.value)
{
break;
}
}
// A record only the log holds, in its key's place.
(Some(e), Some(_), Some(key)) if e.key < key => {
edit_at += 1;
if !e.deleted
&& region.holds(e.key)
&& !batch.offer(&wanted, e.key, e.flags, e.value)
{
break;
}
}
// The image's own record, which the log says nothing about.
(_, Some((value_off, value_len, flags)), Some(key)) => {
let value = match image.value_at(value_off, value_len) {
Ok(v) => v,
Err(e) => return -(e.to_errno() as i32),
};
at += 1;
if region.holds(key) && !batch.offer(&wanted, key, flags, value.as_slice()) {
break;
}
}
(Some(e), None, _) => {
edit_at += 1;
if !e.deleted
&& region.holds(e.key)
&& !batch.offer(&wanted, e.key, e.flags, e.value)
{
break;
}
}
(None, None, _) => break,
_ => break,
}
}
// SAFETY: both out-pointers are writable under this function's contract.
unsafe {
if batch.too_big > 0 && batch.written == 0 {
*out_len = batch.too_big as u64;
*out_cursor = cursor;
return -34; // -ERANGE
}
*out_len = batch.written as u64;
*out_cursor = cursor + batch.returned;
}
0
}
// ── The flag scan: classify at write, retrieve by class ──────────────────────────────────
//
// The class mask lives in the index entry (v4) and in the log entry (WAL v2), so a scan by class is
// a walk that reads the mask and tests it. Nothing else changes: the same merge with the log, the
// same order. What it buys is that a record which does not match is never offered, so the caller
// pays for its own class rather than for the store.
/// A record matches when it shares *any* bit with the scan mask: "every error", "every Wi-Fi event".
const FLAG_MODE_ANY: u16 = 0;
/// A record matches when it carries *every* bit of the scan mask: "every Wi-Fi error".
const FLAG_MODE_ALL: u16 = 1;
/// The scan's scope: one named space, or every space. Two questions, not one widened.
const FLAG_SCOPE_ONE: u32 = 0;
const FLAG_SCOPE_EVERY: u32 = 1;
/// Whether a record's class mask answers a scan for `mask` under `mode`.
///
/// A mask of zero matches nothing. Asking "what is this" with no class named is asking no question,
/// and a scan that returned everything on an empty mask would be a walk wearing a scan's hat — so
/// this is a refusal rather than a wildcard. `cube-store`'s `scan_by_flag` states the same rule,
/// and the gate diffs the two.
fn flag_matches(flags: u16, mask: u16, mode: u16) -> bool {
if mask == 0 {
return false;
}
if mode == FLAG_MODE_ALL {
flags & mask == mask
} else {
flags & mask != 0
}
}
/// The answer a finished batch gives: how much was written and what cursor to pass next, or
/// `-ERANGE` with the size the record that did not fit would need.
///
/// # Safety
/// `out_len` and `out_cursor` must point to writable `u64`s.
unsafe fn finish_batch(
batch: &Batch<'_>,
cursor: u64,
out_len: *mut u64,
out_cursor: *mut u64,
) -> i32 {
// SAFETY: both out-pointers are writable under this function's contract.
unsafe {
if batch.too_big > 0 && batch.written == 0 {
*out_len = batch.too_big as u64;
*out_cursor = cursor;
return -34; // -ERANGE
}
*out_len = batch.written as u64;
*out_cursor = cursor + batch.returned;
}
0
}
/// Scan one space's live records into `batch`, in key order, offering only those whose mask matches.
///
/// This is [`v3_enum`]'s merge with one difference, and it is the operation's whole point: a record
/// is offered only when its class mask answers the scan, so the cursor counts **matches** rather
/// than records walked. That is why a scan re-walks from the space's first record on every batch (as
/// the region walk does) instead of starting at `first + cursor` the way the plain walk can: the
/// index position of the cursor-th match is not arithmetic, because the records before it are not
/// all matches.
///
/// Returns `false` when `batch` is full, so a caller walking several spaces can stop — which is what
/// the every-space scope needs: one batch, several spaces, and one cursor counting matches across
/// all of them.
fn flag_scan_space(
image: &mut Addressed,
log: &[u8],
wanted: &[u8; SPACE_ID_LEN],
first: u64,
records: u64,
mask: u16,
mode: u16,
batch: &mut Batch<'_>,
) -> Result<bool> {
let mut edits = KVVec::<LogEdit<'_>>::new();
log_edits(log, wanted, &mut edits)?;
let mut entry = KVVec::<u8>::new();
let mut at = first;
let mut edit_at = 0usize;
loop {
// The entries of one key arrive in log order, so the last of a key's group is the newest
// word on it — and the newest word carries the mask that counts.
while edit_at + 1 < edits.len() && edits[edit_at + 1].key == edits[edit_at].key {
edit_at += 1;
}
let edit = edits.as_slice().get(edit_at).copied();
let image_entry = if at < first + records {
Some(image.index_entry(at, &mut entry)?)
} else {
None
};
let image_key: Option<&[u8; RAW_KEY_LEN]> = match image_entry {
Some(_) => entry.as_slice()[..RAW_KEY_LEN].try_into().ok(),
None => None,
};
match (edit, image_entry, image_key) {
// The log's word on a coordinate the image holds: it wins, mask and all.
(Some(e), Some(_), Some(key)) if e.key == key => {
edit_at += 1;
at += 1;
if !e.deleted
&& flag_matches(e.flags, mask, mode)
&& !batch.offer(wanted, e.key, e.flags, e.value)
{
return Ok(false);
}
}
// A record only the log holds, in its key's place.
(Some(e), Some(_), Some(key)) if e.key < key => {
edit_at += 1;
if !e.deleted
&& flag_matches(e.flags, mask, mode)
&& !batch.offer(wanted, e.key, e.flags, e.value)
{
return Ok(false);
}
}
// The image's own record, which the log says nothing about.
(_, Some((value_off, value_len, flags)), Some(key)) => {
let value = image.value_at(value_off, value_len)?;
at += 1;
if flag_matches(flags, mask, mode)
&& !batch.offer(wanted, key, flags, value.as_slice())
{
return Ok(false);
}
}
(Some(e), None, _) => {
edit_at += 1;
if !e.deleted
&& flag_matches(e.flags, mask, mode)
&& !batch.offer(wanted, e.key, e.flags, e.value)
{
return Ok(false);
}
}
// One side ran out and the other has nothing left that could interleave: this space is
// done, not the batch.
_ => return Ok(true),
}
}
}
/// Offer the log's live word on each coordinate it holds for one space.
///
/// This is the whole answer for a store whose image has no index: a packed record carries no mask
/// field, so none of the image's records can answer a scan and only the log's can. That is not a
/// corner — it is the state of every store between the write that classified something and the fold
/// that moves the mask into the index.
///
/// Returns `false` when `batch` is full.
fn flag_scan_log_space(
log: &[u8],
space: &[u8; SPACE_ID_LEN],
mask: u16,
mode: u16,
batch: &mut Batch<'_>,
) -> Result<bool> {
let mut edits = KVVec::<LogEdit<'_>>::new();
log_edits(log, space, &mut edits)?;
let mut edit_at = 0usize;
while edit_at < edits.len() {
// The entries of one key arrive in log order, so the last of a key's group is the newest
// word on it — and only the newest can match, because it is what a read would return.
while edit_at + 1 < edits.len() && edits[edit_at + 1].key == edits[edit_at].key {
edit_at += 1;
}
let e = edits.as_slice()[edit_at];
edit_at += 1;
if !e.deleted
&& flag_matches(e.flags, mask, mode)
&& !batch.offer(space, e.key, e.flags, e.value)
{
return Ok(false);
}
}
Ok(true)
}
/// A batch ready to receive a flag scan's matches. `seen` starts at 0 for the reason the region
/// walk's does: this path re-walks from the beginning on every batch, so `offer_flagged` is the only
/// thing that counts. Starting it at `cursor` would skip the first `cursor` matches of a fresh walk
/// *and* still count from there, re-serving the first batch forever.
fn flag_batch<'a>(out: &'a mut [u8], cursor: u64, with_space: bool) -> Batch<'a> {
Batch {
out,
written: 0,
returned: 0,
seen: 0,
cursor,
too_big: 0,
with_space,
}
}
/// `CUBE_OP_FLAG_SCAN` against a v3/v4 image, scoped to one space.
///
/// A v3 image is a store whose records were written before the mask existed. Its entries read as
/// "no class", so a scan over it matches nothing and says so by returning an empty batch — which is
/// the honest answer, not an error.
///
/// # Safety
/// `buf` must hold `cap` writable bytes; `out_len` and `out_cursor` must point to writable `u64`s.
unsafe fn v3_flag_scan(
mut image: Addressed,
wanted: [u8; SPACE_ID_LEN],
mask: u16,
mode: u16,
cursor: u64,
buf: *mut u8,
cap: usize,
out_len: *mut u64,
out_cursor: *mut u64,
) -> i32 {
let (first, records) = match image.space_entry(&wanted) {
Ok(Some(entry)) => entry,
Ok(None) => (0, 0),
Err(e) => return -(e.to_errno() as i32),
};
// Own the log: the edits borrow it, and walking the image needs `&mut image`.
let mut log = KVVec::<u8>::new();
if log.extend_from_slice(image.log(), GFP_KERNEL).is_err() {
return -12; // -ENOMEM
}
// SAFETY: the shim guarantees `cap` writable bytes at `buf`.
let out = unsafe { core::slice::from_raw_parts_mut(buf, cap) };
// The caller named the space, so the frame need not repeat it per record.
let mut batch = flag_batch(out, cursor, false);
if let Err(e) = flag_scan_space(
&mut image,
log.as_slice(),
&wanted,
first,
records,
mask,
mode,
&mut batch,
) {
return -(e.to_errno() as i32);
}
// SAFETY: both out-pointers are writable under this function's contract.
unsafe { finish_batch(&batch, cursor, out_len, out_cursor) }
}
/// `CUBE_OP_FLAG_SCAN` with `CUBE_SPACE_EVERY`: the same question asked of every space.
///
/// One batch and one cursor across all of them, so a caller pages through the whole store's matches
/// as one list — which is what "every event anywhere" has to mean if it is to be usable in batches
/// at all. The spaces are visited in the space table's order, and a space the log writes into that
/// the image has never seen is visited in its place in that same order, so the answer comes back in
/// (space, key) order exactly as a checkpoint writes records.
///
/// # Safety
/// `buf` must hold `cap` writable bytes; `out_len` and `out_cursor` must point to writable `u64`s.
unsafe fn v3_flag_scan_every(
mut image: Addressed,
mask: u16,
mode: u16,
cursor: u64,
buf: *mut u8,
cap: usize,
out_len: *mut u64,
out_cursor: *mut u64,
) -> i32 {
let mut log = KVVec::<u8>::new();
if log.extend_from_slice(image.log(), GFP_KERNEL).is_err() {
return -12; // -ENOMEM
}
// The spaces the log writes into, sorted — including any the image has never seen, which are the
// ones a plain walk of the space table would miss entirely.
let mut from_log = KVVec::<[u8; SPACE_ID_LEN]>::new();
if log_spaces(log.as_slice(), &mut from_log).is_err() {
return -12;
}
// SAFETY: the shim guarantees `cap` writable bytes at `buf`.
let out = unsafe { core::slice::from_raw_parts_mut(buf, cap) };
// The caller named no space, so every frame has to say which one its record came from.
let mut batch = flag_batch(out, cursor, true);
// One closure would have to borrow `image` and `batch` mutably at once, so the walk steps by
// hand; `full` is the batch saying it has no room for another record.
let mut full = false;
let mut row = 0u64;
let mut log_at = 0usize;
while row < image.header.space_count && !full {
let (space, first, records) = match image.space_row(row) {
Ok(Some(triple)) => triple,
Ok(None) => break,
Err(e) => return -(e.to_errno() as i32),
};
// A space only the log writes sorts before this one: visit it here, in its place.
while log_at < from_log.len()
&& from_log.as_slice()[log_at] < space
&& !full
{
let candidate = from_log.as_slice()[log_at];
log_at += 1;
match flag_scan_space(&mut image, log.as_slice(), &candidate, 0, 0, mask, mode, &mut batch) {
Ok(keep_going) => full = !keep_going,
Err(e) => return -(e.to_errno() as i32),
}
}
// A space the log also writes into is the same space, not a second one.
if log_at < from_log.len() && from_log.as_slice()[log_at] == space {
log_at += 1;
}
if !full {
match flag_scan_space(&mut image, log.as_slice(), &space, first, records, mask, mode, &mut batch) {
Ok(keep_going) => full = !keep_going,
Err(e) => return -(e.to_errno() as i32),
}
}
row += 1;
}
// Whatever the log writes into that sorts after every space the image has.
while log_at < from_log.len() && !full {
let candidate = from_log.as_slice()[log_at];
log_at += 1;
match flag_scan_space(&mut image, log.as_slice(), &candidate, 0, 0, mask, mode, &mut batch) {
Ok(keep_going) => full = !keep_going,
Err(e) => return -(e.to_errno() as i32),
}
}
// SAFETY: both out-pointers are writable under this function's contract.
unsafe { finish_batch(&batch, cursor, out_len, out_cursor) }
}
///
/// Returns the record's length, or a negative errno. When the value does not fit in `len` the
/// bytes are *not* copied and the length is returned anyway, so the caller can size its buffer
/// and ask again — a short read that silently truncated would be worse than an error.
///
/// The log has the last word, because it holds what has been written since the last checkpoint;
/// only if it says nothing about the coordinate does the image answer. Nothing is copied out of
/// either: the value the caller gets is a slice of bytes already in memory.
///
/// # Safety
/// `space` must point to 32 readable bytes; `buf` to `len` writable bytes.
#[unsafe(no_mangle)]
pub unsafe extern "C" fn cubelinux_kernel_get(
space: *const u8,
x: u64,
y: u64,
z: u64,
buf: *mut u8,
len: usize,
out_flags: *mut u16,
) -> isize {
let (sp, key) = unsafe { coord_key(space, x, y, z) };
// A read that finds nothing must not leave a class behind for the caller to believe: the answer
// is "no record", and a stale mask beside it would be worse than none.
// SAFETY: the caller guarantees a writable u16.
unsafe { *out_flags = 0 };
// A v3 image is addressed, so a read is a binary search and a few bytes. v1 and v2 images are
// packed lists, which the walking reader below handles.
match Addressed::open() {
Ok(Some(image)) => return unsafe { v3_get(image, sp, key, buf, len, out_flags) },
Ok(None) => {}
Err(e) => return -(e.to_errno() as isize),
}
let view = match read_view() {
Ok(v) => v,
Err(e) => return -(e.to_errno() as isize),
};
match log_effect(view.log(), &sp, &key) {
Some(Effect::Delete) => return -2, // -ENOENT
Some(Effect::Write(value, flags)) => {
// SAFETY: the caller guarantees a writable u16.
unsafe { *out_flags = flags };
return copy_out(value, buf, len);
}
None => {}
}
// Walk the records where they lie, starting where this space's records start: the image is in
// `(space, key)` order, so the record for a coordinate is passed exactly once, and a coordinate
// in a late space does not pay for every record before it.
let at = match view.space_start(&sp) {
Ok(Some(at)) => at,
Ok(None) => return -2, // -ENOENT: the space is not in the image, and the log already spoke
Err(_) => return -12, // -ENOMEM
};
let mut records = Records::at(view.image(), &view.header, at);
while let Some((record_space, record_key, value)) = records.next() {
if record_space == &sp && record_key == &key {
return copy_out(value, buf, len);
}
}
-2 // -ENOENT
}
/// Pack one record the way the walk's uapi names it: `key(24) | value_len(u32, LE) | value`.
///
/// The space is not in the frame because the caller named it, and the order is the store's own
/// (space, then key), so a listing taken through the kernel and one taken in userspace are
/// byte-for-byte comparable — which is how this is tested.
fn pack_record(key: &[u8], value: &[u8], out: &mut [u8], at: usize) -> Option<usize> {
let need = RAW_KEY_LEN + 4 + value.len();
if at + need > out.len() {
return None;
}
out[at..at + RAW_KEY_LEN].copy_from_slice(&key[..RAW_KEY_LEN]);
out[at + RAW_KEY_LEN..at + RAW_KEY_LEN + 4]
.copy_from_slice(&(value.len() as u32).to_le_bytes());
out[at + RAW_KEY_LEN + 4..at + need].copy_from_slice(value);
Some(need)
}
/// Pack one record for a flag scan:
/// `[space(32) |] key(24) | flags(2, LE) | value_len(u32, LE) | value`.
///
/// The mask travels with the record, unlike the walk's frame, because a scan's answer has to say
/// *what class* each record answered with. A record can carry bits the scan did not ask for, and
/// no other op returns a mask — so dropping it here would make those bits unreachable, which is
/// the opposite of what a classification substrate is for. The field order echoes the v4 index
/// entry (`key | flags | …`), so the wire shape is the index's shape.
///
/// The space leads, and only when `with_space` says so, because it is the one field the frame
/// cannot imply when the caller named no space — and it leads rather than trails because the
/// (space, key) order it forms is the order records are stored in and returned in.
fn pack_flagged(
space: &[u8; SPACE_ID_LEN],
key: &[u8],
flags: u16,
value: &[u8],
out: &mut [u8],
at: usize,
with_space: bool,
) -> Option<usize> {
let need = flagged_frame(with_space) + value.len();
if at + need > out.len() {
return None;
}
let mut at = at;
if with_space {
out[at..at + SPACE_ID_LEN].copy_from_slice(&space[..SPACE_ID_LEN]);
at += SPACE_ID_LEN;
}
out[at..at + RAW_KEY_LEN].copy_from_slice(&key[..RAW_KEY_LEN]);
at += RAW_KEY_LEN;
out[at..at + FLAGS_LEN].copy_from_slice(&flags.to_le_bytes());
at += FLAGS_LEN;
out[at..at + 4].copy_from_slice(&(value.len() as u32).to_le_bytes());
at += 4;
out[at..at + value.len()].copy_from_slice(value);
Some(need)
}
/// `CUBE_OP_ENUM`: walk the records of one space into the caller's buffer, `cursor` records in.
///
/// The cursor is a COUNT OF RECORDS ALREADY RETURNED, not a position in the image. It is opaque to
/// the caller (pass back what you were given) and it survives an append, because the contract is
/// "the records I had not yet seen" rather than a snapshot. A record can be seen twice if the
/// image is rewritten underneath a walk; a caller that needs a snapshot takes one.
///
/// The walk walks. It used to read the whole device, parse every record of every space into a
/// merged pool and heapsort all of it, then discard the lot — so asking for one page of one space
/// cost the entire store, and every batch of a listing paid it again. Now the image is read once,
/// its records are walked where they lie in the order a checkpoint gave them, and the log is
/// consulted as an overlay. The cost is the records the walk touches.
///
/// `out_len` receives the bytes written, or with -ERANGE the size the first record that did not fit
/// would need — exactly as a read reports the size it wants.
///
/// # Safety
/// `space` must point to 32 readable bytes; `buf` must hold `cap` writable bytes; `out_len` and
/// `out_cursor` must each point to a writable `u64`.
#[unsafe(no_mangle)]
pub unsafe extern "C" fn cubelinux_kernel_enum(
space: *const u8,
cursor: u64,
buf: *mut u8,
cap: usize,
out_len: *mut u64,
out_cursor: *mut u64,
) -> i32 {
let mut wanted = [0u8; SPACE_ID_LEN];
// SAFETY: the caller guarantees 32 readable bytes at `space`.
unsafe { core::ptr::copy_nonoverlapping(space, wanted.as_mut_ptr(), SPACE_ID_LEN) };
match Addressed::open() {
Ok(Some(image)) => {
return unsafe { v3_enum(image, wanted, cursor, buf, cap, out_len, out_cursor) }
}
Ok(None) => {}
Err(e) => return -(e.to_errno() as i32),
}
let view = match read_view() {
Ok(v) => v,
Err(e) => return -(e.to_errno() as i32),
};
// The live records of this space: the image's, with the log's edits applied as an overlay.
// A space is what a walk names, so only its own edits are collected — and the log is what a
// checkpoint has not yet folded, so there are few.
let mut walker = match SpaceWalker::new(&view, &wanted) {
Ok(w) => w,
Err(_) => return -12, // -ENOMEM
};
// SAFETY: the shim guarantees `cap` writable bytes at `buf`.
let out = unsafe { core::slice::from_raw_parts_mut(buf, cap) };
let mut batch = Batch {
out,
written: 0,
returned: 0,
seen: 0,
cursor,
too_big: 0,
with_space: false,
};
while let Some((key, flags, value)) = walker.next() {
if !batch.offer(&wanted, key, flags, value) {
break;
}
}
// SAFETY: both out-pointers are writable under this function's contract.
unsafe {
if batch.too_big > 0 && batch.written == 0 {
// Not one whole record fits. Say how much it needs, the way a read does.
*out_len = batch.too_big as u64;
*out_cursor = cursor;
return -34; // -ERANGE
}
*out_len = batch.written as u64;
*out_cursor = cursor + batch.returned;
}
0
}
/// `CUBE_OP_RANGE`: the live records of one space whose point lies in the caller's box.
///
/// The cursor means what it means for `CUBE_OP_ENUM` — a **count of records already returned**, so it
/// is opaque, survives an append, and is the only end-of-walk signal — with one difference a caller
/// must know: it counts the records **in the box**, because those are the records this returns.
///
/// The box travels as its six numbers rather than as a struct, for the reason a coordinate does: the
/// key the box has to become is this side's business.
///
/// Two paths, and the difference between them is the honest difference between the layouts. A **v3**
/// image is addressed, so this **seeks**: the box's corner keys bound every key in it, the index is
/// sorted, and a binary search for the span's foot followed by a forward read to its head costs the
/// span rather than the space. A **v1/v2** image is a packed list with no index, so there is nothing
/// to search and the space is walked — the span still stops the walk early, but the cost is the walk.
/// Neither path tells the caller which it got, because the contract is the same either way.
///
/// # Safety
/// `space` must point to 32 readable bytes; `buf` must hold `cap` writable bytes; `out_len` and
/// `out_cursor` must each point to a writable `u64`.
#[unsafe(no_mangle)]
pub unsafe extern "C" fn cubelinux_kernel_range(
space: *const u8,
lo_x: u64,
lo_y: u64,
lo_z: u64,
hi_x: u64,
hi_y: u64,
hi_z: u64,
cursor: u64,
buf: *mut u8,
cap: usize,
out_len: *mut u64,
out_cursor: *mut u64,
) -> i32 {
let mut wanted = [0u8; SPACE_ID_LEN];
// SAFETY: the caller guarantees 32 readable bytes at `space`.
unsafe { core::ptr::copy_nonoverlapping(space, wanted.as_mut_ptr(), SPACE_ID_LEN) };
let region = match Region::new([lo_x, lo_y, lo_z], [hi_x, hi_y, hi_z]) {
Some(region) => region,
// An empty box. Nothing is in it, so the answer is nothing — and answering here keeps an
// inverted span away from the seek, where it would be a range that starts above its end.
// The shim answers this too; a second caller must not be able to reach the seek without it.
None => {
// SAFETY: both out-pointers are writable under this function's contract.
unsafe {
*out_len = 0;
*out_cursor = cursor;
}
return 0;
}
};
match Addressed::open() {
Ok(Some(image)) => {
return unsafe { v3_range(image, wanted, region, cursor, buf, cap, out_len, out_cursor) }
}
Ok(None) => {}
Err(e) => return -(e.to_errno() as i32),
}
let view = match read_view() {
Ok(v) => v,
Err(e) => return -(e.to_errno() as i32),
};
let mut walker = match SpaceWalker::new(&view, &wanted) {
Ok(w) => w,
Err(_) => return -12, // -ENOMEM
};
// SAFETY: the shim guarantees `cap` writable bytes at `buf`.
let out = unsafe { core::slice::from_raw_parts_mut(buf, cap) };
let mut batch = Batch {
out,
written: 0,
returned: 0,
// Same reason as the v3 path: this re-walks the space from its start every batch, and the
// cursor's skip is `offer`'s own `seen`, which counts the matches and starts at 0.
seen: 0,
cursor,
too_big: 0,
with_space: false,
};
while let Some((key, flags, value)) = walker.next() {
match region.place(key) {
// Below the span: not in the box, and skipping it is the seek's substitute on a layout
// that has no index to search.
Span::Below => continue,
// Above the span: the walk is sorted by key, so nothing further can be in the box.
Span::Above => break,
Span::Inside => {}
}
if region.holds(key) && !batch.offer(&wanted, key, flags, value) {
break;
}
}
// SAFETY: both out-pointers are writable under this function's contract.
unsafe {
if batch.too_big > 0 && batch.written == 0 {
// Not one whole record fits. Say how much it needs, the way a read does.
*out_len = batch.too_big as u64;
*out_cursor = cursor;
return -34; // -ERANGE
}
*out_len = batch.written as u64;
*out_cursor = cursor + batch.returned;
}
0
}
/// `CUBE_OP_FLAG_SCAN`: the live records whose class mask matches, in key order.
///
/// Its cursor counts **matches already returned**, exactly as the region walk's counts records in
/// the box — for the same reason, and with the same consequence: a batch holds as many whole
/// records as fit, most batches come back short, and reading a short batch as the end truncates the
/// answer. A finished scan answers with no records and the cursor unchanged.
///
/// The answer's frame is `key(24) | flags(2, LE) | value_len(u32, LE) | value` — the walk's frame
/// with the class mask in it, because a scan's answer has to say what class each record answered
/// with and no other op returns a mask.
///
/// `every_space` chooses between two questions rather than widening one: a space is a hard
/// partition, so "this class here" and "this class anywhere" are different answers, and the narrower
/// one is what a caller gets by leaving the field zero.
///
/// # Safety
/// `space` must point to 32 readable bytes; `buf` must hold `cap` writable bytes; `out_len` and
/// `out_cursor` must each point to a writable `u64`.
#[unsafe(no_mangle)]
pub unsafe extern "C" fn cubelinux_kernel_flag_scan(
space: *const u8,
mask: u16,
mode: u16,
every_space: u32,
cursor: u64,
buf: *mut u8,
cap: usize,
out_len: *mut u64,
out_cursor: *mut u64,
) -> i32 {
let mut wanted = [0u8; SPACE_ID_LEN];
// SAFETY: the caller guarantees 32 readable bytes at `space`.
unsafe { core::ptr::copy_nonoverlapping(space, wanted.as_mut_ptr(), SPACE_ID_LEN) };
// A mode this build does not know is refused rather than defaulted: a caller that asked for
// "all" and silently got "any" would get a superset, and a superset of a security question is
// the worst way to be wrong. The same for a scope it does not know — a caller that asked for
// "every space" and silently got one space would get a subset, which is the other way to be
// wrong and just as quiet.
if mode != FLAG_MODE_ANY && mode != FLAG_MODE_ALL {
return -22; // -EINVAL
}
if every_space != FLAG_SCOPE_ONE && every_space != FLAG_SCOPE_EVERY {
return -22;
}
let every = every_space == FLAG_SCOPE_EVERY;
match Addressed::open() {
Ok(Some(image)) => {
return unsafe {
if every {
v3_flag_scan_every(image, mask, mode, cursor, buf, cap, out_len, out_cursor)
} else {
v3_flag_scan(image, wanted, mask, mode, cursor, buf, cap, out_len, out_cursor)
}
}
}
Ok(None) => {}
Err(e) => return -(e.to_errno() as i32),
}
let view = match read_view() {
Ok(v) => v,
Err(e) => return -(e.to_errno() as i32),
};
// A v1/v2 image holds packed records, which have no mask field at all — they read as "no class"
// and so can never match a scan. What *can* match is the log: its entries are the flagged ones,
// and they carry the mask the writer stamped. So the answer here is the log's live word on each
// coordinate it holds, in key order — a walk's merge with the image's half of it known in
// advance to be empty of matches.
//
// This is not a corner: it is the state of every store between the write that classified
// something and the fold that moves the mask into the index — which is the ordinary state of a
// kernel that has just recorded an event. Answering it with "nothing" would make the substrate
// work only after a checkpoint, which is the opposite of "classify at write time".
let mut log = KVVec::<u8>::new();
if log.extend_from_slice(view.log(), GFP_KERNEL).is_err() {
return -12; // -ENOMEM
}
// SAFETY: the shim guarantees `cap` writable bytes at `buf`.
let out = unsafe { core::slice::from_raw_parts_mut(buf, cap) };
let mut batch = flag_batch(out, cursor, every);
if every {
// Every space the log writes into, sorted — which is every space that can hold a match,
// since none of the image's records can.
let mut from_log = KVVec::<[u8; SPACE_ID_LEN]>::new();
if log_spaces(log.as_slice(), &mut from_log).is_err() {
return -12;
}
for space in from_log.as_slice() {
match flag_scan_log_space(log.as_slice(), space, mask, mode, &mut batch) {
Ok(true) => {}
Ok(false) => break,
Err(e) => return -(e.to_errno() as i32),
}
}
} else if let Err(e) = flag_scan_log_space(log.as_slice(), &wanted, mask, mode, &mut batch) {
return -(e.to_errno() as i32);
}
// SAFETY: both out-pointers are writable under this function's contract.
unsafe { finish_batch(&batch, cursor, out_len, out_cursor) }
}
/// `CUBE_OP_SPACES`: the `cursor`-th distinct space that holds a record, or -ENOENT at the end.
///
/// An index rather than a count of records: a caller that wants a space's records walks it with
/// `CUBE_OP_ENUM` once it has learned the name. Entries are already in (space, key) order, so the
/// distinct spaces come out sorted.
///
/// # Safety
/// `space_out` must point to 32 writable bytes.
#[unsafe(no_mangle)]
pub unsafe extern "C" fn cubelinux_kernel_spaces(cursor: u64, space_out: *mut u8) -> i32 {
match Addressed::open() {
Ok(Some(image)) => return unsafe { v3_spaces(image, cursor, space_out) },
Ok(None) => {}
Err(e) => return -(e.to_errno() as i32),
}
let view = match read_view() {
Ok(v) => v,
Err(e) => return -(e.to_errno() as i32),
};
// Candidates, in ascending order: the spaces the image's records run through (a checkpoint
// writes them in key order, and a space is the first half of the key), and any space the log
// writes into that the image has never seen. One pass over the image collects the first kind;
// the second kind is whatever the log holds, which is small.
let mut log_only = KVVec::<[u8; SPACE_ID_LEN]>::new();
if log_spaces(view.log(), &mut log_only).is_err() {
return -12; // -ENOMEM
}
// A space is only worth reporting if it still holds a record: one whose records the log has all
// removed is not a space anybody can list.
let emit = |space: &[u8; SPACE_ID_LEN], index: &mut u64| -> bool {
if !space_has_records(&view, space) {
return false;
}
if *index == cursor {
// SAFETY: the caller guarantees 32 writable bytes at `space_out`.
unsafe {
core::ptr::copy_nonoverlapping(space.as_ptr(), space_out, SPACE_ID_LEN);
}
return true;
}
*index += 1;
false
};
// The image's own spaces are the table of where they start — no pass over the records, because
// that pass already happened when the table was built.
let starts = match space_starts(&view) {
Ok(s) => s,
Err(_) => return -12, // -ENOMEM
};
let mut index: u64 = 0;
let mut log_at = 0usize;
for (space, _) in starts.as_slice() {
let space: [u8; SPACE_ID_LEN] = *space;
while log_at < log_only.len() && log_only.as_slice()[log_at] < space {
let candidate = log_only.as_slice()[log_at];
log_at += 1;
if emit(&candidate, &mut index) {
return 0;
}
}
// A space the log also writes into is the same space, not a second one.
if log_at < log_only.len() && log_only.as_slice()[log_at] == space {
log_at += 1;
}
if emit(&space, &mut index) {
return 0;
}
}
while log_at < log_only.len() {
let candidate = log_only.as_slice()[log_at];
log_at += 1;
if emit(&candidate, &mut index) {
return 0;
}
}
-2 // -ENOENT: no such space; the walk is finished
}
/// `CUBE_OP_DEL`: remove the record at a coordinate.
///
/// A removal is a log entry, not an erasure: nothing in an append-only store is rewritten in
/// place, and the checkpoint is what finally drops it.
///
/// # Safety
/// `space` must point to 32 readable bytes.
#[unsafe(no_mangle)]
pub unsafe extern "C" fn cubelinux_kernel_del(space: *const u8, x: u64, y: u64, z: u64) -> i32 {
let _op = STORE_OP.lock();
ensure_boot_record();
let (sp, key) = unsafe { coord_key(space, x, y, z) };
let mutation = Mutation {
space: sp,
key,
flags: 0,
value: KVVec::new(),
};
// A delete is an entry with op=2; `append` writes op=1, so build it here from the same
// framing, and let the log's own reader be the judge of it. Like a put, it appends through the
// bounded read: a delete does not touch the records either.
match append_mutation(&mutation, 2) {
Ok(_) => 0,
Err(e) => -(e.to_errno() as i32),
}
}
/// `CUBE_OP_SYNC`: fold the log into the image.
#[unsafe(no_mangle)]
pub extern "C" fn cubelinux_kernel_sync() -> i32 {
let _op = STORE_OP.lock();
ensure_boot_record();
let (device, layout) = match device_and_layout() {
Ok(pair) => pair,
Err(e) => return -(e.to_errno() as i32),
};
match fold_now(&device, &layout) {
Ok(()) => 0,
Err(e) => -(e.to_errno() as i32),
}
}
/// The module's registration; holds the misc device for as long as the module lives.
#[pin_data]
struct CubeStoreModule {
#[pin]
_miscdev: MiscDeviceRegistration<CubeStore>,
}
impl kernel::InPlaceModule for CubeStoreModule {
fn init(_module: &'static ThisModule) -> impl PinInit<Self, Error> {
pr_info!("cubelinux: store reader registered at /dev/cubelinux\n");
// SAFETY: called exactly once, in the module initializer, before anything can take it.
unsafe { SPACE_STARTS.init() };
// STORE_FILE needs the same treatment, and it was the one that went without.
//
// The failure mode is why it hid for so long. A zeroed mutex satisfies the uncontended
// fast path — the count reads 0, which means "unlocked" — so one writer at a time works
// perfectly and nothing looks wrong. The first *contended* lock takes
// __mutex_lock_slowpath, which splices the task into the mutex's wait list; that list is
// uninitialised, so its head is NULL and the splice stores through it. A write to address
// 0 in kernel mode, after which the task returns with interrupts disabled and
// preemption held, which is a machine that cannot panic, log, or recover.
//
// Found by verify-file-store.sh MODE=race: twelve concurrent puts to a store that is a
// file on the root filesystem oopsed in cubelinux_store::store_file, while twelve
// sequential ones passed with exact log accounting. The same signature is on the box,
// where the store had several writers: a hard freeze with no panic despite panic=30, and
// a journal that simply stops.
// SAFETY: called exactly once, in the module initializer, before anything can take it.
unsafe { STORE_FILE.init() };
// SAFETY: called exactly once, in the module initializer, before anything can take it.
unsafe { STORE_OP.init() };
try_pin_init!(Self {
_miscdev <- MiscDeviceRegistration::register(MiscDeviceOptions {
name: c_str!("cubelinux"),
}),
})
}
}
#[pin_data]
struct CubeStore {
dev: ARef<Device>,
}
#[vtable]
impl MiscDevice for CubeStore {
type Ptr = Pin<KBox<Self>>;
fn open(_file: &File, misc: &MiscDeviceRegistration<Self>) -> Result<Pin<KBox<Self>>> {
KBox::try_pin_init(
try_pin_init! {
CubeStore { dev: ARef::from(misc.device()) }
},
GFP_KERNEL,
)
}
/// Write one operation, framed as the argument block the coordinate interface will pass.
///
/// ```text
/// put op=1 | space(32) | x(8) | y(8) | z(8) | len(4) | value[len]
/// sync op=3
/// ```
///
/// `put` is the byte plane — bytes at a coordinate, with no header written beside them —
/// which is what the store contract defines and what the differential gate compares
/// against userspace's `cell put`. The header tier sits above this.
///
/// `sync` folds the log into the image. It is a real operation rather than a test hook: a
/// caller that wants the log reclaimed — a shutdown, a snapshot, a handover — is entitled
/// to ask.
fn write_iter(kiocb: Kiocb<'_, Self::Ptr>, iov: &mut IovIterSource<'_>) -> Result<usize> {
let mut buf = KVVec::<u8>::new();
let len = iov.copy_from_iter_vec(&mut buf, GFP_KERNEL)?;
let bytes = buf.as_slice();
if bytes.is_empty() {
return Err(EINVAL);
}
// A write to the store device is the kernel taking the write path, so this is one of the
// places a boot gets recorded — the same hook the syscall's write operations call.
// The lock spans the op: everything below reads the layout and commits against it.
let _op = STORE_OP.lock();
ensure_boot_record();
let me = kiocb.file();
// A put needs the control block and the log's head; a fold needs everything. So the whole
// device is read inside the branch that folds, and not before the branch that does not.
match bytes[0] {
OP_PUT => {
if bytes.len() < 1 + 32 + 24 + 4 {
return Err(EINVAL);
}
let mut space = [0u8; 32];
space.copy_from_slice(&bytes[1..33]);
let axis = |at: usize| -> u64 {
let mut w = [0u8; 8];
w.copy_from_slice(&bytes[at..at + 8]);
u64::from_le_bytes(w)
};
let (x, y, z) = (axis(33), axis(41), axis(49));
let mut w = [0u8; 4];
w.copy_from_slice(&bytes[57..61]);
let value_len = u32::from_le_bytes(w) as usize;
if bytes.len() < 61 + value_len {
return Err(EINVAL);
}
let mut value = KVVec::<u8>::new();
value.extend_from_slice(&bytes[61..61 + value_len], GFP_KERNEL)?;
let mutation = Mutation {
space,
key: morton_encode(x, y, z),
flags: 0,
value,
};
match append_mutation(&mutation, 1) {
Ok(_) => {
dev_info!(
me.dev,
"cubelinux: appended {} bytes for a {} byte value\n",
ENTRY_FIXED + value_len,
value_len
);
Ok(len)
}
Err(e) => {
dev_err!(me.dev, "cubelinux: append failed: {:?}\n", e);
Err(e)
}
}
}
OP_SYNC => {
// The one write that reads the whole store: a fold rewrites the image, so it has to
// hold it. This is the tail of the latency budget, and it is why a fold belongs on a
// timer rather than on the path a caller waits behind.
let mut device = KVVec::<u8>::new();
read_image(&mut device)?;
let layout = match resolve_layout(&device) {
Ok(l) => l,
Err(what) => {
dev_err!(me.dev, "cubelinux: cannot resolve the store layout: {}\n", what);
return Err(EINVAL);
}
};
let ctl = match layout.control {
Some(c) => c,
None => {
// A bare image has one slot, so folding into it is not atomic and a
// crash in the middle would lose the store. Refusing is the honest
// answer; the fix is to format the device as a store.
dev_err!(
me.dev,
"cubelinux: cannot checkpoint a bare image (no control block)\n"
);
return Err(EINVAL);
}
};
// The store the image and its log describe, ready to be written out.
let live = &device[layout.image_off..];
let header = match parse_header(live) {
Ok(h) => h,
Err(what) => {
dev_err!(me.dev, "cubelinux: {}\n", what);
return Err(EINVAL);
}
};
let window = log_window(&device, &layout);
let mut merged = match build_merged(live, window, &header) {
Ok(m) => m,
Err(what) => {
dev_err!(me.dev, "cubelinux: cannot build the store: {}\n", what);
return Err(EINVAL);
}
};
match checkpoint(&ctl, &mut merged.0) {
Ok(bytes) => {
dev_info!(me.dev, "cubelinux: checkpointed {} bytes\n", bytes);
Ok(len)
}
Err(e) => {
dev_err!(me.dev, "cubelinux: checkpoint failed: {:?}\n", e);
Err(e)
}
}
}
other => {
dev_err!(me.dev, "cubelinux: unknown operation {}\n", other);
Err(EINVAL)
}
}
}
fn read_iter(mut kiocb: Kiocb<'_, Self::Ptr>, iov: &mut IovIterDest<'_>) -> Result<usize> {
// The work happens here, on demand, rather than at init: by the time anything can
// open this device, the block driver that provides the store is certainly up.
let mut image = KVVec::<u8>::new();
let line = match read_image(&mut image) {
Ok(()) => digest(&image),
Err(e) => {
let mut line = Line::new();
let _ = write!(line, "error={:?}", e);
line
}
};
pr_info!("cubelinux: {}\n", core::str::from_utf8(line.as_bytes()).unwrap_or("<bad utf8>"));
let me = kiocb.file();
dev_info!(me.dev, "cubelinux: store read complete\n");
iov.simple_read_from_buffer(kiocb.ki_pos_mut(), line.as_bytes())
}
}