|
ravel
Deterministic simulation testing for C++. Seed a bug, replay it exact.
|
A simulation test is an ordinary program that exits non-zero when a seed fails, so it fits any CI system. This page gives you the pieces: a test program with a ready-made command line, workflows for pull requests and nightly runs, a place for reproducers to live, and advice on how many seeds to run.
ravel::run_sweep_main gives your program a complete command line. This is the whole main:
(deposit_setup is the system from the tutorial, still with its bug, so we have something to find.) Run it with no arguments and you get a sweep of 1000 seeds. Here is the same program with 200 seeds:
The summary tells you what failed, gives you the reproducer and trace, and says exactly how to replay it. Exit status is 1. When everything passes it is 0, and 2 means a bad command line.
The options you will use in CI:
Two of them come from the environment too, which is convenient in CI: RAVEL_SEEDS and RAVEL_FIRST_SEED (a flag beats the variable).
On every pull request, run a modest number of seeds: enough to catch the common bugs, fast enough that nobody waits. With GitHub Actions:
What you get when a seed fails:
.choices reproducer are attached as a downloadable artifact.Use a Release build for the sweeps: they are compute-bound and run many times faster than Debug. Keep a separate, smaller job in Debug with sanitizers (ASan, UBSan) if you use them elsewhere, since a memory error inside a simulated task is a real bug.
New seeds are new chances to find something, and nobody is waiting overnight. Run a much bigger sweep on a schedule, and start at a different seed each night so you never repeat yourself:
github.run_number goes up by one each run, so run 41 covers seeds 41,000,000 to 41,499,999. Nothing overlaps, and any failure names its seed, so you can always come back to it.
When the nightly sweep finds something, you fix it and keep its reproducer, so the bug can never quietly return. Keep the shrunk .choices file in your repository (say under tests/regressions/) and replay it as a test. The --replay flag does exactly that, and exits non-zero if the run fails:
(This one fails, because the tutorial's server still has its bug. After the fix it passes, and stays a test.) With CMake, register every file in the folder as a test in one go:
Now ctest runs the sweep and every regression, locally and in CI. A new bug is a new file dropped into the folder. (Re-run CMake to pick it up.)
A reproducer replays the structure of the run, so it keeps working when ravel is upgraded, but it can stop failing if you change the code under test so much that the recorded choices no longer line up with what the code asks. That is not a regression to worry about: if the fix or a refactor makes a reproducer pass, it still guards the case it was written for as well as it can, and the sweep guards the rest.
run_sweep is the easy way in, but you can also call the pieces directly from gtest, Catch2, doctest or anything else. The two calls are ravel::run_seeds (a sweep) and ravel::replay (a saved reproducer):
The tutorial's regression example is the same idea in plain code. (The framework snippets above are not compiled by this repository's docs build, which has no dependency on gtest or Catch2; they use only calls that appear in the compiled examples.)
There is no magic number, only a trade of time against chances. Think in time budgets, not seed counts:
| Where | Budget | Why |
|---|---|---|
| Local, while developing | a few seconds | fast feedback: catches the common bugs |
| Every pull request | 1 to 5 minutes | most bugs need one unlucky moment and show up on a few percent of seeds |
| Nightly | 30 to 120 minutes | rare bugs: the rarest one in ravel's own Raft example fails about 1 seed in 600 |
To turn a budget into a number, measure once: run time ./my_sim_test --seeds 1000 and divide. Simple systems run many thousands of seeds a second; a much heavier one such as the Raft example (three nodes, crashes, disks) manages on the order of a hundred a second across all the cores of a laptop. Then pick the count that fills your budget. --threads defaults to every core, so a bigger CI machine helps almost linearly.
If a bug only shows up on rare seeds, more seeds is the answer. If it never shows up at all, ask whether your faults are reaching it: a system that never sees a lost message or a crash cannot fail because of one. Raise loss rates, add crashes, add latency variance. ravel's own examples use fairly harsh settings for that reason.
Two limits worth setting in the setup you give run_sweep_main (through SweepDefaults::simulation, or on the command line):
time_limit**, for systems that never go quiet (heartbeats, election timers): the run stops at that virtual time and the invariants are checked. --time-limit overrides it per invocation.max_steps** (a million by default) turns a livelock into a failure instead of a hung CI job. Raise it only if your system is legitimately that busy.ravel-traces artifact. It has the trace (*.replay.trace.jsonl, the minimal run) and the reproducer (*.choices)../my_sim_test --replay ravel-traces/ravel-seed-N.choices. Same choices, same run, every time..choices file into regressions/. It is now a test.If the failure does not reproduce locally, the code under test is not fully deterministic. Run ./my_sim_test --check-determinism: it names the first step where two runs of the same seed part ways. See the porting guide.