In about half an hour you will find a real bug, watch ravel shrink it to a handful of steps, read exactly what happened, fix it, and leave behind a test that guards against it forever.
The bug is a classic. A client asks a server to deposit money and retries if it hears nothing back. If the server's reply is lost, the retry makes the server deposit twice. It hides in production for months, because it needs a network fault at the wrong moment. ravel finds it in a fraction of a second.
Everything on this page is real: the code is compiled and the output is produced by the actual programs (see tools/check_docs.py). You can find them in [docs/snippets](snippets).
1. Install
ravel needs a C++20 compiler with coroutines (GCC 10+, Clang 14+, MSVC 2019+) and CMake 3.20+. The simplest way in is CMake's FetchContent:
cmake_minimum_required(VERSION 3.20)
project(my_tests LANGUAGES CXX)
include(FetchContent)
FetchContent_Declare(ravel
GIT_REPOSITORY https://github.com/FelixMiddelhoff/ravel.git
GIT_TAG v0.4.0)
FetchContent_MakeAvailable(ravel)
add_executable(my_test my_test.cpp)
target_link_libraries(my_test PRIVATE ravel::ravel)
target_compile_features(my_test PRIVATE cxx_std_20)
Prefer a package manager? The repository carries a vcpkg port (use it with --overlay-ports=packaging/vcpkg/ports) and a Conan recipe (conan create packaging/conan). After either, find_package(ravel CONFIG REQUIRED) gives you the same ravel::ravel target.
2. A first simulation
Start with the smallest thing that works, to see the moving parts:
#include <cstdio>
#include "ravel/runner.hpp"
int main() {
});
},
options);
std::printf(
"%llu seeds run, %zu failed\n",
static_cast<unsigned long long>(report.
seeds_run),
return report.
ok() ? 0 : 1;
}
Top-level harness: owns the seed, virtual clock, RNG, trace and scheduler for one deterministic run.
Definition simulation.hpp:70
void add_invariant(std::string name, InvariantFn invariant)
Invariants are checked once, after the scheduler has run to quiescence.
Scheduler & scheduler() noexcept
The scheduler: spawn tasks through it.
Definition simulation.hpp:84
The coroutine type for simulated tasks.
Definition task.hpp:19
How run_seeds runs.
Definition runner.hpp:14
std::uint64_t seed_count
How many seeds to run, starting at first_seed.
Definition runner.hpp:18
What run_seeds found.
Definition runner.hpp:36
std::vector< Result > failures
In ascending seed order.
Definition runner.hpp:41
std::uint64_t seeds_run
Seeds covered.
Definition runner.hpp:39
bool ok() const noexcept
True if no seed failed.
Definition runner.hpp:48
Build and run it:
What just happened:
- A simulation is one complete run of your system under one seed. It owns everything the code under test may use: a virtual clock, a random source, a scheduler, and (later) a network and disks.
- Tasks are C++20 coroutines.
co_await sim.scheduler().sleep(10) waits for ten virtual ticks. Nothing really waits: when every task is asleep, ravel jumps the clock forward. A simulated hour costs microseconds.
- An invariant is a property that must hold when the run ends.
run_seeds runs the same setup under 100 different seeds. Each seed makes different random choices (which task goes next, which message is lost), so each explores a different behavior of your system. Same seed, same run, every time, on every machine. That is what makes failures reproducible.
3. The system to test
Now the real system. It is short enough to read in one go:
#include <set>
#include <string>
#include "ravel/simulation.hpp"
struct Bank {
int balance = 0;
bool client_got_ack = false;
std::set<int> seen_requests;
};
inline ravel::SimulationSetup deposit_setup(bool dedupe) {
while (true) {
const ravel::Message request =
co_await to_server.
receive();
const int id = std::stoi(request.substr(8));
if (!dedupe || bank.seen_requests.insert(id).second) bank.balance += 10;
}
});
for (int attempt = 0; attempt < 5; ++attempt) {
to_server.
send(
"deposit 1");
bank.client_got_ack = true;
co_return;
}
}
});
sim.
add_invariant(
"deposit_applied_at_most_once", [&bank] {
return bank.balance <= 10; });
[&bank] { return !bank.client_got_ack || bank.balance == 10; });
};
}
A virtual, one-way, in-process channel between two named endpoints.
Definition network.hpp:43
TimedReceiveAwaiter receive_within(VirtualClock::Tick timeout) noexcept
Waits up to timeout ticks for a message: auto m = co_await channel.receive_within(50); The result is ...
Definition network.hpp:98
ReceiveAwaiter receive() noexcept
Waits for the next message: ravel::Message m = co_await channel.receive();
Definition network.hpp:78
void send(Message message)
Applies the fault spec: the message is dropped, or delivered after a random delay.
void spawn(std::string name, TaskFactory factory)
Registers a task.
T & make_state(Args &&... args)
Creates an object owned by the simulation and returns a reference to it.
Definition simulation.hpp:103
Channel & add_channel(std::string from, std::string to, FaultSpec fault)
Adds a one-way message channel from endpoint from to endpoint to.
Injectable fault behavior for a Channel.
Definition network.hpp:23
double loss_probability
In [0, 1]: chance that a message is dropped.
Definition network.hpp:24
The pieces worth noticing:
- **
SimulationSetup.** A function that builds one run from scratch: it adds the network, spawns the tasks and registers the invariants. ravel calls it once per seed, on a fresh Simulation.
- **
sim.make_state<Bank>().** State shared by tasks and invariants. Use it instead of locals: it is created fresh for each run and lives as long as the tasks do. (The setup function returns before the tasks run, so a local variable would already be gone. See the porting guide.)
- **
add_channel(...).** A one-way virtual network link. The FaultSpec makes it lose 30% of messages and delay the rest. Faults are random choices like any other, so they are reproducible.
- **
receive_within(50).** Wait for a message, but give up after 50 ticks. That is the client's retry timer.
- The server has no deduplication when
dedupe is false. That is the bug we are about to find. It is not obvious from reading the code, which is the point.
4. Hunt for the bug
Run the system under a thousand seeds. Ask ravel to shrink the first failure and save what happened:
int main(int argc, char** argv) {
if (report.
ok())
return 0;
if (print_saved_file(argc > 1 ? argv[1] : "", shrunk)) return 0;
std::printf(
"%llu seeds run, %zu failed\n",
static_cast<unsigned long long>(report.
seeds_run),
std::printf(
"first failure: seed %llu: %s\n",
static_cast<unsigned long long>(first.
seed),
std::printf("shrunk from %zu random choices (%llu steps) to %zu (%llu steps)\n",
std::printf("minimal choices:");
for (const auto choice : shrunk.choices) std::printf(" %llu", static_cast<unsigned long long>(choice));
std::printf(
"\nsaved: %s\n", forward_slashes(shrunk.
choices_path).c_str());
return 0;
}
How one run ended.
Definition simulation.hpp:56
std::string trace_path
The dumped trace; empty if none was written.
Definition simulation.hpp:64
std::string failure
What went wrong; empty when ok.
Definition simulation.hpp:61
std::uint64_t steps
Scheduler steps taken.
Definition simulation.hpp:62
std::uint64_t seed
The seed the run was made from.
Definition simulation.hpp:60
bool shrink_first_failure
Shrink the lowest failing seed once the sweep is done (see shrink()).
Definition runner.hpp:28
SimulationOptions simulation
Applied to every run.
Definition runner.hpp:32
std::optional< ShrinkResult > shrunk
The lowest failing seed, minimized.
Definition runner.hpp:45
What shrink() found.
Definition shrink.hpp:49
Result minimal
The smallest failing run found.
Definition shrink.hpp:52
Choices choices
Replay these to reproduce minimal.
Definition shrink.hpp:53
Choices original_choices
Its recorded choices.
Definition shrink.hpp:51
Result original
The seed's own run.
Definition shrink.hpp:50
std::string choices_path
Where choices was saved; empty if not.
Definition shrink.hpp:54
std::filesystem::path trace_dir
When set, a failed run writes its trace here as ravel-seed-<seed>.trace.jsonl.
Definition simulation.hpp:48
1000 seeds run, 281 failed
first failure: seed 6: invariant 'deposit_applied_at_most_once' failed
shrunk from 11 random choices (9 steps) to 4 (6 steps)
minimal choices: 0 0 0 1
saved: ravel-traces/ravel-seed-6.choices
trace: ravel-traces/ravel-seed-6.replay.trace.jsonl
(print_saved_file only does something when the program is run with --trace or --choices-file, as we will in a moment: it prints that saved file instead of the summary.)
Line by line:
- **
1000 seeds run, 281 failed.** More than a quarter of seeds expose the bug. That is typical: bugs that need a fault at the wrong moment show up on a fraction of seeds, and a thousand seeds run in well under a second.
- **
first failure: seed 6.** Seed 6 is the lowest failing seed. Run it again and you get exactly the same failure, today, tomorrow, on a colleague's laptop. The message names the invariant that broke.
- **
shrunk from 11 random choices (9 steps) to 4 (6 steps).** The failing run had 11 random decisions. ravel replayed edited versions of them, keeping any edit that still failed the same way, until it could not simplify further: 4 decisions, the smallest run that still has the bug.
- **
minimal choices: 0 0 0 1.** Those 4 decisions. Section 6 explains how to read them.
- **
saved: and trace:.** Two files, written because we set trace_dir. The first is a reproducer you can check in. The second is the story of the run.
5. Read what happened
The trace file has one line per event. Run the same program with --trace to print it:
{"format":"ravel-trace","trace_version":1,"ravel_version":"x.y.z","seed":6}
{"step":0,"time":0,"kind":"TaskSpawned","id":0,"name":"server"}
{"step":1,"time":0,"kind":"TaskSpawned","id":1,"name":"client"}
{"step":2,"time":0,"kind":"TaskResumed","id":0,"name":"server"}
{"step":3,"time":0,"kind":"TaskResumed","id":1,"name":"client"}
{"step":4,"time":0,"kind":"MessageSent","id":0,"name":"client->server"}
{"step":5,"time":1,"kind":"MessageDelivered","id":0,"name":"client->server"}
{"step":6,"time":1,"kind":"TaskResumed","id":0,"name":"server"}
{"step":7,"time":1,"kind":"MessageSent","id":1,"name":"server->client"}
{"step":8,"time":1,"kind":"MessageDropped","id":1,"name":"server->client"}
{"step":9,"time":50,"kind":"TaskResumed","id":1,"name":"client"}
{"step":10,"time":50,"kind":"MessageSent","id":0,"name":"client->server"}
{"step":11,"time":51,"kind":"MessageDelivered","id":0,"name":"client->server"}
{"step":12,"time":51,"kind":"TaskResumed","id":0,"name":"server"}
{"step":13,"time":51,"kind":"MessageSent","id":1,"name":"server->client"}
{"step":14,"time":52,"kind":"MessageDelivered","id":1,"name":"server->client"}
{"step":15,"time":52,"kind":"TaskResumed","id":1,"name":"client"}
{"step":16,"time":52,"kind":"TaskFinished","id":1,"name":"client"}
You can read this like a log. The important part:
step 4 t=0 MessageSent client->server the client asks for a deposit
step 5 t=1 MessageDelivered client->server the server gets it
step 6 t=1 TaskResumed server ...and deposits 10
step 7 t=1 MessageSent server->client the server replies "ok"
step 8 t=1 MessageDropped server->client the reply is lost
step 9 t=50 TaskResumed client the client waits 50 ticks, gets nothing
step 10 t=50 MessageSent client->server ...and asks again
step 11 t=51 MessageDelivered client->server the server gets the second request
step 12 t=51 TaskResumed server ...and deposits 10 again: balance 20
There it is. The reply was lost at step 8, so the client retried, and the server applied the same request twice. Every line of the trace is in virtual time (t= is ticks, never wall-clock), so it reads the same however fast your machine is.
Prefer a picture? tools/ravel_trace.py timeline lays the same trace out with one column per task and channel, like a sequence diagram, and jq slices it from the command line; see debugging.md for both.
6. What is a "choice"?
Everything random in a run (which task goes next, whether a message is lost, how long a delay is) is drawn from one source, and every draw is recorded as a small number where 0 means the simplest outcome. The list of those numbers describes the whole run. The saved file is exactly that list:
ravel-choices 1
4
0 0 0 1
Our four numbers, in the order the draws happened:
| Choice | Draw | 0 means | Value | So |
| 1st | which task runs first | the one waiting longest | 0 | the server |
| 2nd | is the request lost? | no | 0 | it is delivered |
| 3rd | how long is its delay? | the minimum (1 tick) | 0 | 1 tick |
| 4th | is the reply lost? | no | 1 | the reply is lost |
Everything after those four draws takes the default, 0. So this is the minimal recipe for the bug: deliver the request, lose the reply.
That is how shrinking works: it edits the list (drop some numbers, lower others), replays it, and keeps the edit if the bug is still there. A shorter list with smaller numbers is a simpler run. And because a run is its list, you can save it, check it in, and replay it any time, with no seed and no dependence on how ravel generates random numbers.
7. Fix it
The server must recognize a request it has already handled. That is one flag in our setup (dedupe), which makes the server remember request ids:
int main() {
std::printf(
"%llu seeds run, %zu failed\n",
static_cast<unsigned long long>(report.
seeds_run),
return report.
ok() ? 0 : 1;
}
A thousand seeds, none failing. (More seeds are cheap: change 1000 to 100000 and go for a coffee.)
8. Keep it fixed
Finding the bug once is not enough; you want a test that fails if it ever comes back. The reproducer from step 4 is perfect for that. Paste its numbers into a test and replay them against your code:
const ravel::Choices kDoubleDeposit = {0, 0, 0, 1};
int main() {
const ravel::Result fixed = ravel::replay(deposit_setup(
true), kDoubleDeposit);
const ravel::Result buggy = ravel::replay(deposit_setup(
false), kDoubleDeposit);
std::printf(
"fixed server: %s\n", fixed.
ok ?
"passes" :
"FAILS");
std::printf(
"buggy server: %s\n", buggy.
ok ?
"passes" : (
"fails: " + buggy.failure).c_str());
return fixed.
ok && !buggy.
ok ? 0 : 1;
}
bool ok
True if the run passed every check.
Definition simulation.hpp:58
fixed server: passes
buggy server: fails: invariant 'deposit_applied_at_most_once' failed
The fixed server passes; the buggy one fails on exactly this run, without needing to search for a lucky seed. In a real project this goes in your normal unit tests (gtest, Catch2, doctest, plain main: whatever you use). Keep the sweep over many seeds too: the sweep finds new bugs, the reproducer stops old ones from returning.
You can also keep the .choices file in your repository and load it with ravel::read_choices. The text format is documented in formats.md.
9. Run it in CI
A sweep is just a program that exits non-zero when a seed fails, so it drops into any CI system. On GitHub Actions:
- name: Simulation tests
run: |
cmake --build build --target my_test
./build/my_test
- name: Keep the evidence when it fails
if: failure()
uses: actions/upload-artifact@v4
with:
name: ravel-traces
path: ravel-traces/
Set options.simulation.trace_dir as in step 4 and a failing run leaves its trace and its reproducer in ravel-traces/; the artifact step keeps them so you can download them and read the story of the failure.
Some habits that pay off:
- Run many seeds in CI (thousands) and a few in local builds.
- Run a bigger sweep nightly with a different
first_seed each time. New seeds are new chances to find something.
- When a nightly run fails, commit the shrunk
.choices file with the fix.
Where next