fix(audit): eliminate O(n)/unbounded audit-path bottlenecks (100% ok under load)

Root cause of the ~4% error rate in audit-enabled runs (run-qc6newt3:
96.30% ok, op6 mean 367ms max 3003ms) was the audit path's three
compounding costs, isolated iteratively under the real 8-user x 150s
model-B-with-audit stress harness:

  1. append(): rewrote the whole log string on every op (O(n) read-modify-write
     under a per-store Mutex) -> op latency grew with log size.
  2. dump(): walked 1..=count re-reading every entry record (O(n)) -> became the
     new bottleneck once append was fixed (op6 still ~760-980ms).
  3. the  command returned the UNBOUNDED full log (~1MB at 12k entries)
     on every call -> ~1MB response serialized/sent/received = op6 ~978ms.

Fix (aligned with the PDF's 'access logs live in Null rows' time/stream-keyed
model):
  - AUDIT_HEAD stores only a decimal entry count (index); each entry is its own
    durable record at entry_coord(seq) -> append is O(1) (two put_record calls).
  - ConcurrentStore gains a per-store in-memory tail cache (audit_tail) shared by
    every Audit over that store; append extends it by one line, dump returns a
    clone -> dump is O(1) and never re-walks the store. Serialized under the
    cache guard so concurrent connections interleave correctly.
  - the interactive  command serves a bounded recent tail
    (Audit::AUDIT_TAIL_LIMIT = 200) instead of the full log; the full log stays
    available via Session::audit_dump()/Audit::dump() for export.

Verification (real, not assumed):
  - ./check gate GREEN (fmt + tests + clippy -D warnings), incl. R6 append/dump
    tests and pre-existing grant_and_revoke_emit_audit_entries.
  - hermes_verify_audit_o1: index=count (not log), distinct coords, ascending
    dump, concurrent interleave-correct. PASS.
  - model-B-with-audit re-run (8 users x 150s): 110,647 ops, 100.00% ok, 0
    failures; op6 mean 11.9ms (p99 46.9ms, max 124.9ms) vs 367ms pre-fix. Final
    run evidence: /root/cube-stress/run-kkaogy3b.
  - Auth confirmed a non-factor (zero rejections) across all runs.

docs/stress-comparison-20260811.md: corrected the bogus '~0.3ms audit op' claim
in S5 and replaced the placeholder S6 with the full root-cause/fix/verification
write-up including the iteration-to-100% table.
This commit is contained in:
CUBELinux-2
2026-08-11 21:31:42 -04:00
parent 15ce36d488
commit ad73f42e46
4 changed files with 284 additions and 50 deletions
+17
View File
@@ -421,6 +421,13 @@ pub struct ConcurrentStore {
stop: Arc<AtomicBool>,
flush_thread: Arc<Mutex<Option<JoinHandle<()>>>>,
cfg: DurabilityConfig,
/// In-memory mirror of the rendered audit log (one newline-joined string),
/// kept per-store so every `Audit` bound to this store shares it. `append`
/// extends it by one line and `dump` returns a clone — both O(1). The
/// durable per-entry records remain the source of truth; this is just a
/// fast read mirror, rebuilt from records on first `dump` if cold (e.g.
/// after a restart that loaded a durable store).
audit_tail: Mutex<String>,
}
impl ConcurrentStore {
@@ -435,9 +442,18 @@ impl ConcurrentStore {
stop: Arc::new(AtomicBool::new(true)),
flush_thread: Arc::new(Mutex::new(None)),
cfg: DurabilityConfig::default(),
audit_tail: Mutex::new(String::new()),
}
}
/// Shared, per-store in-memory mirror of the rendered audit log. All
/// `Audit` instances bound to this store use this one buffer so concurrent
/// connections interleave their entries exactly as the durable per-entry
/// records do. `audit.rs` holds the serialization contract.
pub(crate) fn audit_tail(&self) -> &Mutex<String> {
&self.audit_tail
}
/// Open (or create) a durable store at `db_path`, with the WAL at
/// `wal_path` and recovery events logged to `recovery_log`. Loads the last
/// checkpoint, replays any newer WAL entries, and starts the background
@@ -492,6 +508,7 @@ impl ConcurrentStore {
stop,
flush_thread: Arc::new(Mutex::new(None)),
cfg,
audit_tail: Mutex::new(String::new()),
};
// Background checkpoint thread.