|
ravel
Deterministic simulation testing for C++. Seed a bug, replay it exact.
|
Deterministic simulation testing (DST) for distributed C++ systems — seed a bug, replay it exact.
Status: pre-1.0 (0.4.0). The features below work and are tested on Linux, macOS and Windows, including under ASan, UBSan and TSan. It has not been used outside this repository yet, so the API, the meaning of a seed and the C ABI may still change between minor versions. Requires C++20 (coroutines).
ravel has not been used outside this repository yet, and I would like it to be. If you build distributed or stateful C++ (a queue, a replicated store, a consensus implementation, anything with retries, timeouts or crash recovery), I would be glad if you tried it on a small piece of your code and told me how it went. Short reports are fine, and so is "I gave up at step 3".
What helps most:
Open an issue with what you tried and what happened. A failing seed and the compiler and OS are plenty. Pull requests are welcome too: CONTRIBUTING.md says how to build, test and where things live, and docs/known-limitations.md lists what is known to be missing. For a security problem, see SECURITY.md instead of opening an issue.
FoundationDB/TigerBeetle-style deterministic simulation testing, for C++. Virtualize time, RNG, and network behind seams a Simulation owns; run under a controlled, seed-driven scheduler; on failure, the seed alone reproduces the exact same interleaving and faults.
Rust has this (turmoil, madsim, loom). C++ doesn't have a portable, permissively-licensed equivalent.
run_seeds runs many seeds in parallel; some interleavings expose the bug, and a failing seed replays identically. See examples/quickstart.cpp.
When a seed fails, shrink boils the run down. Every random decision (scheduling, message loss, delays, and anything your workload draws from sim.rng()) is recorded as a small integer where 0 is the simplest outcome. Shrinking edits that list, replays it, and keeps any change that still fails the same way with a shorter or smaller list. A 40-message run with a lost message shrinks to one that loses exactly one. Candidates are replayed on all hardware threads, and the result is the same whatever the thread count. Besides dropping and lowering single choices, it moves value between two draws that must add up to something (say, two delays) and lowers two draws together.
Messages travel over virtual channels whose faults come from the same seed:
Systems that never go quiet on their own (heartbeats, election timers) stop at SimulationOptions::time_limit, and their invariants are checked then.
Disks model what makes storage code hard: a write is visible at once but only durable after sync, and a crash loses, tears or keeps each unsynced write.
Tasks are C++20 coroutines. At every co_await scheduler.yield() or co_await scheduler.sleep(ticks) the seeded scheduler picks which runnable task goes next. Time is virtual: when every task sleeps, the clock jumps to the earliest wake-up.
examples/raft.hpp is a working Raft implementation (leader election and log replication, each node persisting its state to a virtual disk) that runs under ravel with nodes losing power at random, and a lossy, reordering network. The correct version survives every seed. Three deliberate bugs, of the kind real implementations have shipped with, are each found by ravel:
| Bug | What breaks | Seeds that fail (of 600) |
|---|---|---|
| state file renamed into place before its data is synced | a rebooted node forgets entries it acknowledged | ~200 |
rename never made durable (no sync_dir) | the same, but only if power fails right after a state change | ~1 |
| votes for candidates with stale logs | a new leader lacks committed entries | ~20 |
ravel_raft runs all four and shrinks the first bug from about 2000 random choices to a few hundred, printing what the checkers saw (a leader without an entry a node had already applied). The checkers are election safety, state machine safety and leader completeness.
| Seam | Today | Planned |
|---|---|---|
VirtualClock / VirtualRng | done, seed-deterministic | - |
Scheduler | seed-driven interleaving, virtual-time sleep | - |
Trace | event log with a replay digest; JSON Lines dump of failed runs (SimulationOptions::trace_dir) | - |
Channel / FaultSpec | one-way message channel with loss, latency, optional reordering | - |
Multi-seed runner (run_seeds) | parallel over seeds, thread-count-independent report | - |
Shrinking (shrink, replay) | minimizes a failing run; saves a replayable choices file | - |
Disk / DiskFaultSpec | virtual files and directories: write-back cache, sync, rename, remove, sync_dir, list; crash() loses, tears or keeps unsynced data and directory changes; ENOSPC, I/O errors, latency | - |
Do I have to change my code? Yes, a little, and that is the trade. Code under test takes its time, randomness, network and disk from ravel (sim.clock(), sim.rng(), Channel, Disk) instead of calling std::chrono, rand(), sockets or files directly, and runs its concurrency as ravel tasks. ravel does not intercept syscalls or processes the way rr, Antithesis or LD_PRELOAD tools do. What you give up is "works on a binary
you cannot modify". What you get is no ptrace, no root, no platform-specific magic, identical behavior on Linux, macOS and Windows, and a run you can step through in a normal debugger.
Is a seed stable across ravel versions? Not before 1.0. A change to how randomness is consumed changes what every seed does. That is why a shrunk failure is saved as a choices file (*.choices): a plain list of numbers that replays the same failure and can be checked in as a regression test. From 1.0, a seed will keep its meaning within a minor version line.
Can my code use threads? No. A simulation is single-threaded and cooperative: tasks are coroutines, and ravel decides the order they run in. Code that starts real threads, or reads a real clock or a real random source, makes runs unrepeatable, and shrink will refuse it (it checks that replaying a failure really reproduces it). run_seeds is parallel, but only across independent simulations that share nothing.
Is VirtualRng secure? No. It is a fast, portable, reproducible PRNG for simulation. Never use it for keys, nonces or tokens.
Which compilers? C++20 with coroutines. CI builds and runs the tests with GCC 10, 11, 13 (also as a 32-bit x86 build) and 14, Clang 14, 15, 18 and 19, MSVC (Visual Studio 2022) and Apple Clang. Older or other compilers may work but are not tested. GCC 10 needs -fcoroutines, which the CMake package adds for you. No big-endian target is tested. The digest and the RNG are defined on bytes and integers only, so results should not depend on byte order, but that is an argument, not a test.
Channel and Disk are virtual, and the only files ravel writes are the traces and choice lists you ask for.ravel_trace.py viewer and differ).run_sweep), pull-request and nightly workflows, and regression tests from saved reproducers.doxygen docs/Doxyfile, or the ravel_docs CMake target, output in build-docs/html).include/ravel/ravel.h exposes a minimal opaque-handle C surface for FFI from languages/runtimes that can't link C++ directly. Stability policy: docs/abi-policy.md.
MIT — see [LICENSE](LICENSE).
See SECURITY.md for scope and misuse boundaries.
See CONTRIBUTING.md. This project follows the Contributor Covenant.