|
ravel
Deterministic simulation testing for C++. Seed a bug, replay it exact.
|
ravel can only test code whose behavior depends on nothing but the seed. Real code reads the clock, waits in real time, starts threads and talks to real sockets and disks, and each of those makes a run unrepeatable. This guide shows how to get your code into a shape ravel can test, one step at a time, with a worked example.
Your code does two kinds of things: it decides (what to send, when to give up, which entry to keep) and it touches the world (reads the time, sends bytes, writes a file). ravel replaces the world with a virtual one that it controls, and leaves the decisions alone.
The virtual world has one source of randomness, and everything in it (task order, lost messages, delays, crashes) is drawn from that source. So a run is a pure function of its seed.
The one thing this asks of you: your logic must reach the world only through things ravel can replace. ravel does not intercept system calls (see the FAQ); you make the seam yourself, once, and it pays off in code that is easier to test in every other way too.
std::chrono, rand, std::random_device, sleep_for, std::thread, sockets and file streams.ravel::check_determinism, then run thousands of seeds.You do not have to port everything. Start with one component (a protocol, a cache, a queue) and simulate that.
Here is code that cannot be simulated. A monitor pings a peer every 100 ms and declares it dead if no answer has come for 300 ms:
Every line marked in the comments is a problem: it reads the real clock, waits in real time, runs on a real thread and talks to a real socket. Testing "what if three pings in a row are lost?" means unplugging cables and waiting seconds per try, and a failure could never be replayed.
The fix is to pull the logic out. Look at what the monitor actually decides: when to ping and whether the peer is alive. Both only need to know "what time is it?". So make time an argument:
This class has no clock, no thread, no socket. It is a plain state machine, and you can unit-test it by hand: call ponged(100), ask peer_alive(350). In production, a driver calls it with steady_clock time and real sockets. Under ravel, a driver calls it with virtual time and virtual channels:
The peer answers for 500 ticks and then goes silent, over a network that loses 20% of messages and delays the rest. Across 500 seeds the monitor always notices the dead peer by the end. Make peer_alive always return true and this run will tell you within a second.
The pattern to take away is **"sans-I/O"**: logic in the middle, no I/O in it, and thin drivers at the edges. It is the same shape as the Raft example (Node handlers change state and queue messages; the run loop does the I/O).
| Your code does | Under ravel |
|---|---|
std::chrono::steady_clock::now() | sim.clock().now() (virtual ticks), or take the time as an argument |
std::this_thread::sleep_for(d) | co_await sim.scheduler().sleep(d) |
| a timer or timeout | co_await channel.receive_within(d), or sleep |
rand(), std::mt19937, std::random_device | sim.rng().next_below(n), next_between(a, b), chance(p) |
| a socket, RPC or message queue | a Channel (send, co_await receive()), with a FaultSpec for loss, delay and reordering |
fstream, write, fsync, rename | a Disk: write, read, sync, rename, sync_dir; crash() for power loss |
std::thread / a thread pool | sim.scheduler().spawn("name", ...): cooperative tasks |
std::mutex / condition_variable | usually nothing: tasks only switch at co_await. (The interesting races are between those points, which is what ravel explores.) |
| "the process crashes and restarts" | disk.crash(), then a task that reboots from what is on the disk (see the Raft example) |
Two things worth knowing:
sim.rng(), including your own workload generators. Every draw is recorded and shrinkable, so a random workload shrinks too. Anything random from elsewhere breaks reproducibility.co_await between two steps cannot be interleaved between them. If you want ravel to try reordering operations, give it a co_await sim.scheduler().yield() between them.ravel's promise, that a seed always gives the same run, only holds if nothing else influences the code under test. The usual suspects:
| Leak | Why it breaks determinism |
|---|---|
std::chrono::...::now(), time(), clock() | different on every run |
rand(), std::random_device, address-derived seeds | different on every run |
std::thread, std::async, thread pools | the OS decides the interleaving |
| statics and globals that survive between runs | the second run of a seed starts in a different state |
std::unordered_map / unordered_set keyed by pointers | iteration order depends on addresses, which differ between runs |
| sorting or hashing by pointer value | same reason |
| uninitialized memory | whatever was there last |
| reading environment variables, files or the network | different on different machines |
p, this, addresses in logging that feed decisions | addresses differ |
You do not have to hunt for these by reading. Ask ravel:
It runs each seed twice (and once more replayed from the recorded choices) and compares the runs event by event. Here it is on code with a leak (a static that survives between runs):
The first message is the useful part: it names the first step where the runs part ways. Here, step 2 is when the worker wakes up; run 1 slept until t=10 and run 2 until t=5. Something that feeds the sleep time differs between the runs: look at where nap comes from and you find the static.
Make check_determinism part of your test suite, over a few dozen seeds. It is cheap, and it turns "my simulation is flaky" from a mystery into a line number. If you skip it, ravel::shrink still refuses to shrink a failure that does not replay, but that is a late and vague warning.
1. Do not capture locals of the setup function. The setup returns before the tasks run, so a captured local is already gone:
Capture references to things made with make_state, or to channels and disks from add_channel/add_disk (those live as long as the simulation). Capturing by value is always fine.
2. A task that simulates a process must stop when the power goes. After disk.crash(), an operation that completed a moment earlier still resumes its task. Read disk.crash_count() when the process starts and stop when it changes; the Raft example's dead() does this.
3. Never start real threads or block. A task that calls something that really waits (a real sleep, a blocking read) stalls the whole simulation, and one that starts a thread makes runs unrepeatable. ravel throws if a foreign thread touches a running simulation, but it cannot catch everything.
4. Do not keep state between runs. Statics, globals and singletons survive from one seed to the next. Keep per-run state in make_state. (check_determinism catches the ones you miss.)
5. Setup runs on several threads at once in run_seeds and shrink, so it must not touch shared mutable data. Read-only shared data is fine.
For a first port:
sim.rng() only.make_state, not in captures or statics.co_awaits (or yields) between the steps you want reordered.check_determinism passes on a few dozen seeds.run_seeds runs thousands of seeds, with shrink_first_failure and a trace_dir, so a failure leaves its reproducer behind.When something goes wrong along the way, debugging.md shows what each failure looks like and what to do about it.