Compare decoders on the same data
Sample once, replay matching and BP-OSD, then score their predictions against private answers.
Choose an input-compatible decoder
A decoder predicts observable flips from detector events and a detector error model. Minimum-weight perfect matching (MWPM, rmatching) uses graphlike error components; belief propagation with ordered-statistics decoding (BP-OSD, rbposd) works with parity-check models. Integer linear programming (ILP, rilpqec) is an optional solver path. None is universally best.
This experiment uses rsinter replay with its matching and BP-OSD runner features. The central installation section lists optional tools. Decoder choice and configuration are recorded in the replay statistics.
1. Freeze one model and one sample batch
Create circuit.stim in your working directory.
cat > circuit.stim <<'STIM'
R 0 1
X_ERROR(0.1) 0
X_ERROR(0.1) 1
CX 0 1
M 0 1
DETECTOR rec[-2]
DETECTOR rec[-1]
OBSERVABLE_INCLUDE(0) rec[-2] rec[-1]
STIM
rstim circuit dem --in circuit.stim --decompose-errors --out model.dem > /dev/null
rstim circuit detect --in circuit.stim --shots 128 --seed 7 \
--out detectors.b8 --out-format b8 \
--obs-out answers.b8 --obs-out-format b8 > /dev/null
The DEM has two detectors and one observable. b8 stores one byte per row for these widths; only the low two detector bits or low one observable bit are used. Keep answers.b8 private. Both decoders must read exactly the same detector file and DEM.
2. Replay both decoders
rsinter replay --dem model.dem --dets detectors.b8 --decoder rmatching \
--predictions-out matching.b8 --stats-out matching.json --shots 128
rsinter replay --dem model.dem --dets detectors.b8 --decoder rbposd \
--predictions-out bposd.b8 --stats-out bposd.json --shots 128
The commands read neither answers nor masks. Replay compiles once and streams batches; statistics include dimensions, normalized configuration, input/output hashes and timing. For a fair comparison, check those input hashes agree and separate compilation time from steady decoding time.
3. Score complete predictions
python3 - <<'PYTHON'
import json
from pathlib import Path
answers = Path('answers.b8').read_bytes()
assert len(answers) == 128 and all(b in (0, 1) for b in answers)
input_hashes = []
for name in ('matching', 'bposd'):
stats = json.loads(Path(name + '.json').read_text())
predictions = Path(name + '.b8').read_bytes()
assert stats['num_shots'] == 128
assert stats['num_detectors'] == 2 and stats['num_observables'] == 1
assert len(predictions) == len(answers)
assert all(b in (0, 1) for b in predictions)
input_hashes.append((stats['dem_sha256'], stats['detectors_sha256']))
failures = sum(p != a for p, a in zip(predictions, answers))
print(f'{name}: {failures} / {len(answers)} logical errors')
assert input_hashes[0] == input_hashes[1], 'Different decoder inputs'
PYTHON
matching: 0 / 128 logical errors
bposd: 0 / 128 logical errors
Each disagreement is one logical failure. This tiny graphlike case verifies wiring; it cannot establish a general ranking. An incomplete prediction file or nonzero replay exit is a failed experiment, not a decoder accuracy result.
Earlier Rust decoder example
The earlier two-row matching example remains available. For a complete experiment, use the Rust tutorial.
Scale the experiment
Use recorded decoder benchmarks for full campaigns, conditions and reproduction. For parameter sweeps, use the rsinter campaign reference; retain circuit, noise model, seeds, configuration, shot counts and stopping rules. For a Rust consumer, run the complete API experiment. The earlier two-row Rust decoder example remains downloadable.