fix(audit): eliminate O(n)/unbounded audit-path bottlenecks (100% ok under load)
Root cause of the ~4% error rate in audit-enabled runs (run-qc6newt3:
96.30% ok, op6 mean 367ms max 3003ms) was the audit path's three
compounding costs, isolated iteratively under the real 8-user x 150s
model-B-with-audit stress harness:
1. append(): rewrote the whole log string on every op (O(n) read-modify-write
under a per-store Mutex) -> op latency grew with log size.
2. dump(): walked 1..=count re-reading every entry record (O(n)) -> became the
new bottleneck once append was fixed (op6 still ~760-980ms).
3. the command returned the UNBOUNDED full log (~1MB at 12k entries)
on every call -> ~1MB response serialized/sent/received = op6 ~978ms.
Fix (aligned with the PDF's 'access logs live in Null rows' time/stream-keyed
model):
- AUDIT_HEAD stores only a decimal entry count (index); each entry is its own
durable record at entry_coord(seq) -> append is O(1) (two put_record calls).
- ConcurrentStore gains a per-store in-memory tail cache (audit_tail) shared by
every Audit over that store; append extends it by one line, dump returns a
clone -> dump is O(1) and never re-walks the store. Serialized under the
cache guard so concurrent connections interleave correctly.
- the interactive command serves a bounded recent tail
(Audit::AUDIT_TAIL_LIMIT = 200) instead of the full log; the full log stays
available via Session::audit_dump()/Audit::dump() for export.
Verification (real, not assumed):
- ./check gate GREEN (fmt + tests + clippy -D warnings), incl. R6 append/dump
tests and pre-existing grant_and_revoke_emit_audit_entries.
- hermes_verify_audit_o1: index=count (not log), distinct coords, ascending
dump, concurrent interleave-correct. PASS.
- model-B-with-audit re-run (8 users x 150s): 110,647 ops, 100.00% ok, 0
failures; op6 mean 11.9ms (p99 46.9ms, max 124.9ms) vs 367ms pre-fix. Final
run evidence: /root/cube-stress/run-kkaogy3b.
- Auth confirmed a non-factor (zero rejections) across all runs.
docs/stress-comparison-20260811.md: corrected the bogus '~0.3ms audit op' claim
in S5 and replaced the placeholder S6 with the full root-cause/fix/verification
write-up including the iteration-to-100% table.
This commit is contained in:
@@ -66,8 +66,12 @@ rigorously rather than trusting memory.
|
||||
- **Model B (auth-once):** persistent connection, one signed-HELLO per connection, unlimited ops.
|
||||
- **Model A (auth-each-time):** `cubec`-one-shot semantics — fresh connection + full handshake every command.
|
||||
- Both run 8 users × 120s. The op mix is `prog/run/grant/revoke/query/stats/[audit]`. The `audit`
|
||||
op is the heavy one (~0.3ms in model B, but the 3s socket timeout in model A counts every
|
||||
handshake+op round trip, so slow ops time out as failures).
|
||||
op was the heavy one in model B **with audit ON**: it was NOT ~0.3ms. The model-B-with-audit runs
|
||||
(run-qc6newt3 / run-i6ktxrnp) show **mean op 63.3ms / 59.4ms at only 96.3% / 96.5% ok** — the
|
||||
`audit` op dominated the tail (op6 max ~3003ms, pinned to the 3s socket cap) because each audit
|
||||
append was a full O(n) read-modify-rewrite of the entire audit log string under a per-store Mutex.
|
||||
(See §6 — this was root-caused and fixed *after* the §5 A/B investigation.) Model A's 3s timeout
|
||||
compounds the same slow op into failures; model B's cap is hit per-op.
|
||||
|
||||
**Controlled variable — audit op:** `NO_AUDIT=1` drops op6 (`audit`) to replicate the legacy op mix
|
||||
(what the user's "~0% errors" memory was based on: prog+run, write+read, grant+revoke, link+query,
|
||||
@@ -98,8 +102,55 @@ seal, stats — no audit, no 3s pressure).
|
||||
|
||||
**Conclusion for the design:** `cubec` one-shot (auth-each-time) is sound and matches the legacy
|
||||
error profile; the current daemon default (auth-once per persistent connection) is strictly better
|
||||
on handshake count and ties on latency. No auth-model change is warranted. The only real lever on the
|
||||
observed ~4% failure was the `audit` op / 3s timeout, orthogonal to auth.
|
||||
on handshake count and ties on latency. No auth-model change is warranted. The observed ~4% failure
|
||||
was driven by the slow `audit` op + 3s socket cap, orthogonal to auth — and that slowness was itself
|
||||
a bug (§6), not an inherent cost of auditing.
|
||||
|
||||
## 6. Audit-path root cause & O(1) fix (2026-08-11, post-§5)
|
||||
|
||||
**Root cause (found after the §5 A/B write-up):** the high error rate in the audit-enabled runs was
|
||||
**not** the 3s socket cap as the proximate trigger — it was the *cost* of the `audit` op itself.
|
||||
`cubesys/src/audit.rs::append` did a full `get_record(AUDIT_HEAD)` → split the whole log string on
|
||||
newlines → push one line → `put_record(AUDIT_HEAD, rejoined)` on **every** op, serialised under a
|
||||
per-store `Mutex`. That is O(n) in the number of audit entries, so latency grew with the log: op6
|
||||
mean ~367ms, max ~3003ms (the 3s client cap), which is exactly the ~4% failure band seen in
|
||||
run-qc6newt3 / run-i6ktxrnp. Every other op stayed ~19–21ms. Auth was a non-factor (zero
|
||||
rejections in any run) — §5's verdict stands; this just names *why* the audit op was slow.
|
||||
|
||||
**Fix (Option A, user-selected):** `AUDIT_HEAD` now stores only a decimal **entry count** (the index),
|
||||
and each audit entry is written as its **own durable record** at a distinct coordinate derived from
|
||||
its seq (`entry_coord(seq)`) — matching the PDF's "access logs live in Null rows" time/stream-keyed
|
||||
model. `append()` does two O(1) `put_record` calls (entry + index bump) under a single per-store
|
||||
guard. Crucially, `dump()` no longer re-walks the entries: the store carries a **per-store in-memory
|
||||
tail cache** (`ConcurrentStore::audit_tail`, shared by every `Audit` over that store) that `append`
|
||||
extends by one line and `dump` returns by clone — both O(1), even as the log grows to thousands of
|
||||
entries. (First cut made `append` O(1) but left `dump` as an O(n) walk; under the stress harness,
|
||||
which calls the `audit` command as op6 thousands of times, that walk became the new ~760ms
|
||||
bottleneck — identical 96.3% ok / 3003ms p99 as pre-fix. The tail cache removes it.)
|
||||
|
||||
**Verification (real, not assumed):**
|
||||
- `./check` gate (fmt + tests + clippy -D warnings): **GREEN**, including the R6 `dump()`/`append`
|
||||
tests and the pre-existing `grant_and_revoke_emit_audit_entries` audit test.
|
||||
- In-repo regression test `hermes_verify_audit_o1` (cubesys/src/audit.rs): asserts `AUDIT_HEAD`
|
||||
holds the count (not the log), entries land at distinct coords, `dump()` is ascending-ordered, and
|
||||
two `Audit`s over the same store interleave correctly under concurrency. **PASS.**
|
||||
- Ad-hoc runtime test against `target/release/cube-server` (R4-authenticated path): drove
|
||||
`prog/run/grant/revoke/query/stats`; `audit` returned entries `"seq":1..N` ascending, one
|
||||
~79-byte JSON line each — confirming the O(1) per-entry layout end-to-end. **PASS.**
|
||||
|
||||
**Load-level proof (model-B-with-audit re-run, 8 users × 150s, audit ON):**
|
||||
| run | ok% | op6 mean | op6 max | note |
|
||||
|-----|-----|----------|---------|------|
|
||||
| pre-fix (run-qc6newt3) | 96.30% | 367 ms | 3003 ms (3s cap) | O(n) append rewrite |
|
||||
| fix #1: O(1) append only | 96.30% | 761 ms | 3003 ms | dump still O(n) walk |
|
||||
| fix #2: + per-store tail cache | 95.88% | 978 ms | 3003 ms | dump O(1) but returned full ~1 MB log |
|
||||
| **fix #3: + bounded `audit` tail (final)** | **100.00%** | **11.9 ms** | **124.9 ms** | all three O(n)/size causes removed |
|
||||
|
||||
Final run (run-kkaogy3b): 110,647 ops, **0 failures**, op6 mean 11.9 ms (p99 46.9 ms, max 124.9 ms)
|
||||
— on par with every other op (5–15 ms). The ~4% error band is gone; root cause was the
|
||||
audit-path's three compounding costs (whole-log rewrite on append, walk on dump, unbounded
|
||||
response on the `audit` command), all now O(1)/bounded. Auth remained a non-factor (zero
|
||||
rejections), confirming §5's verdict.
|
||||
|
||||
Per-tenant isolation note: ad-hoc multi-tenant routing/isolation proofs (Task 3, `/tmp/cubelinux-tenant-isol-*`)
|
||||
showed per-tenant store isolation is correct and costs nothing measurable vs a shared store — also
|
||||
|
||||
Reference in New Issue
Block a user