cubesys: coalesce WAL fsync on commit; stop over-reporting durable seq

commit_txn issued an unconditional fsync per COMMIT, so N concurrent
writers serialized behind N disk syncs (p99 hit the 3s socket timeout
under 8 writers). Add Wal::sync_upto: committers queue on an fsync_gate,
the first one flushes the whole accumulated buffer, and waiters that find
committed_seq past their target return with zero I/O. N commits now cost
~1 fsync with the same durability guarantee.

Also fix a durability over-report: flush_pending stamped committed_seq
from the LIVE seq counter, so sequences taken by appenders that had not
yet buffered their bytes were reported durable. Track max_seq alongside
the pending buffer and advance committed_seq only to what was written.

Wire --wal-fsync-ms / --checkpoint-ms in cube-server (previously
hardcoded to defaults, so the documented knob did nothing).

Both regressions are mutation-verified: each test fails when its bug is
reintroduced.
This commit is contained in:
CUBELinux-2
2026-08-11 18:09:29 -04:00
parent c87fef0514
commit ec84fbc725
3 changed files with 221 additions and 24 deletions
+10
View File
@@ -110,6 +110,16 @@ impl TenantConfig {
durability: DurabilityConfig::default(),
}
}
/// Disk-backed tenant config at `store_dir` with explicit durability
/// tuning, so the daemon's `--wal-fsync-ms` / `--checkpoint-ms` flags
/// actually reach each per-tenant store's WAL.
pub fn disk_with(store_dir: impl Into<PathBuf>, durability: DurabilityConfig) -> Self {
TenantConfig::Disk {
store_dir: store_dir.into(),
durability,
}
}
}
/// A client's asserted identity, declared via `HELLO <tenant> <owner_local>`