ConcurrentStore.inner is now Arc<RwLock<CubeStore>>: all read paths take the read side, all mutations + checkpoint take the write side. Readers no longer exclude each other and overlap an active writer (verified by concurrent_reads_dont_block_on_writer + cube-bench Task 4 section). WAL, checkpoint, and coordinate encoding are untouched, so durability/replay is unchanged (.check green). Honest finding recorded in docs/task4-reader-writer-sharding.md: on this 8-core host std RwLock removes reader-vs-reader exclusion (correct) but shows no wall-clock speedup for short reads (cache-line bounce on one shared lock). Real read-throughput scaling would need sharded/lock-free storage, left as a follow-up decision rather than invented.
3.9 KiB
3.9 KiB
Task 4 — Reader/Writer Sharding: before / after
Plan: convert the single global Mutex<CubeStore> inside ConcurrentStore
into an RwLock<CubeStore> so many reader threads run in parallel while writes
stay serialized. Then demonstrate (via cube-bench, release build) that reads
stop blocking each other.
What changed (Task 4 commit)
cubesys/src/store.rs:ConcurrentStore.inner: Arc<Mutex<CubeStore>>->Arc<RwLock<CubeStore>>. All GET-side methods (get_raw,get_record,keys,scan_prefix,linked_to,query_doc_type,read_snapshot) take the read side; all mutation paths (put_raw,delete_raw,put_record,associate,with_mut,checkpoint, and the durability helpersincremental/fold_delta_into_base) take the write side.- The WAL, the checkpoint thread, and coordinate encoding are untouched — the
change is purely the lock granularity, so durability/replay behavior is
identical (proven by the existing
durable_checkpoint_and_replay,wal_recovery_after_crash,incremental_checkpoint_delta_modeltests). - Added
store::tests::concurrent_reads_dont_block_on_writerand acube-benchTask 4 section (behind abenchfeature flag oncubesysso the internals never ship in the daemon path).
Correctness result (verified)
- Reader-vs-reader exclusion is GONE: 8 concurrent reader threads + 1 writer thread all make progress; reads overlap an active writer instead of queuing behind it. No deadlock, no pathological stall.
- Durability tests still pass (
./checkgreen): the WAL boundary / delta model is unchanged.
Performance result (HONEST — see benchmarks below)
The RwLock delivers the correctness property the plan asked for (reads no
longer serialize behind one another, and a writer no longer stalls readers from
starting). It does NOT deliver a wall-clock speedup for short reads on this
8-core host, because std::sync::RwLock's per-lock atomic and the shared
CubeStore cache line bounce when 8 readers contend on the same lock.
cargo run -p cube-bench --release (this host, warm in-memory backend):
Task 4 concurrency (RwLock, 8 readers x 400000 reads)
reader solo 10.37 ms
8 readers (RwLock) 300.58 ms
8 readers (old Mutex) 82.95 ms (serialized bound = 8 x solo)
8 readers + writer 380.30 ms (writer active 150ms; readers overlapped)
aggregate read rate 10645.92 k reads/s
NOTE: for ~15ns reads, RwLock overhead > parallel gain (per-lock atomics);
the win is correctness (no reader-vs-reader exclusion) + writer overlap.
Task 4 concurrency, heavier reads (256B payload + ts, 30000 each)
reader solo 2.16 ms
8 readers (RwLock) 23.25 ms
8 readers (old Mutex) 17.31 ms (serialized bound)
RwLock vs Mutex ratio 0.74 x (1.0 = parity; <1 means RwLock slower here)
Conclusion / recommendation
- Task 4 as specified (RwLock sharding) is correct and merged. It removes reader-vs-reader mutual exclusion and lets reads overlap writes — a real concurrency improvement in the API/correctness sense.
- If the goal is actual read-throughput scaling (not just non-exclusion), the
single shared
CubeStoreunder oneRwLockis the bottleneck: all readers still serialize on the same lock word and contend on the same cache line. To scale reads you would need one of:- a sharded store (split the coordinate space across N
RwLock<CubeStore>shards, keyed byc/prefix) so concurrent reads land on different locks, or - a lock-free / MVCC snapshot map (readers clone an
Arc<Snapshot>and never touch a lock), or - the existing
read_snapshot()call pattern for bulk-read clients (VM, cubefs) so they take one consistent clone instead of per-call locks.
- a sharded store (split the coordinate space across N
- None of those were required by Task 4 (which asked for RwLock sharding), so they are left as a follow-up decision rather than silently invented.