Binary Bedrock
Recorded campaigns

Everything we have run. Everything we found.

This page is the raw record: every repository we pointed the tools at, every verdict, every limitation we hit and what we did about it. Every number comes from a recorded run with a receipt, and a machine holds us to it: CI recomputes every figure on this page from a ledger of the recorded runs and fails the build if a published number drifts. The flagship wins can be re-proven on your own machine. Where a run found nothing, or found our own bugs, that is on this page too. The discipline is the product, and the differentiators behind these rows are kept in a public register with re-verification triggers, so a claim can never quietly outlive its evidence.

The totals, computed by us before you compute them

Firewall (all C/C++ rows): 2,302 functions verified byte-identical after normalization; of the 126 non-identical comparable pairs, 21 received a decisive semantic verdict (18 PROVEN, 3 DIVERGES): a 16.7% decisive rate on changed code today, with 105 honest refusals reason-coded in the rows below. The identical class is itself a real answer (your commit touched 40 functions, 37 provably compile to the same behavior), but the decisive rate on the rest is the number we manage the product by, segmented by code class: integer-heavy kernels (CMSIS-NN reached 53% eligibility) decide far more often than general application C++, which is why the roadmap and the pilot wedge point at quantized inference, crypto, DSP and codec code first.

Audit (all rows with a full funnel): 1,735 functions submitted, 458 eligible (26.4%), 192 proven equivalent, 1 size win (the re-provable zlib receipt). Priced accordingly: the Audit is a provability-and-coverage report with receipts; optimization wins are upside, not the promise.

Firewall

Per-repo verification runs (C, C++, Rust)

One row per recorded run of verify-diff against a real commit. identical = same compiled IR; proven = solver-proved equivalent; diverges = proven behavior change with a counterexample; refusals = honest CANNOT-CERTIFY.

repo · commitidenticalprovendivergesrefusalserrorsnote
zstd cef5a561410100a live commit provably changes ZSTD_ldm_adjustParameters; counterexample in 49s
lodepng 22561886480000gates passed
leveldb 4a0c57263500140C++ at scale; demangled names; reproduced exactly across a 3-major compiler jump
tinyxml2 999a21f2960000GATE CLOSED (configuration): changed lines are #if-gated out; the gate refuses to vouch for code it never compiled
tinyxml2 8d8472f2740080destructor family refused honestly
libpng a22696be1240040refusals exactly on the touched allocation paths
cJSON b2890c88932260the null-fix proven to change behavior ONLY at null (with a domain proof)
FreeRTOS 5706e10cb460070tasks.c, the most deployed scheduler file in embedded
tinyexpr 9d5696c4160220
sds be182cb391010
linenoise e9fb8ed330050
miniz 384a12d21030the CVE-guard function PROVEN in 5.7s
zlib e3dc0a880030with project defines; varargs paths refused
brotli 0d1f62980030reproduced exactly on clang 22
tiny-AES-c 6be2e1133020AES CBC/CTR loop optimization proven (conditional on the unroll bound, and it says so)
tiny-AES-c 2ca3e8150030int-to-size_t signature change; later proven via the declared-conversion adapter
inih c75edb812030
log.c f9ea34970010
CMSIS-DSP 918014fconfig-dependentone commit, two configurations, two OPPOSITE verdicts, both correct (DIVERGES under LOOPUNROLL, identical elsewhere)
libsodium e6324db702000the maintainer's own optimization judgment on production crypto, formally confirmed (980ms)
libsodium d4c60aee20000follow-up commit; both touched functions byte-identical
lwip 3d896ba0config-requiredevery unit needs the app-provided lwipopts.h; refused honestly rather than compiled wrong
BLAKE3 f3149ecno C/C++ changefreshest commit touches no C sources; recorded as-is
rust crate · commitidenticalprovenrefusalserrorsnote
hashbrown 8d9d6c5351050the std HashMap internals, monomorphized via the crate's own tests
aho-corasick 6c0abf55880140
semver 7625c7a13180first field solver-proofs on monomorphized Rust; 56 added/removed reported honestly
byteorder (tests target)842068201,524 monomorphized functions, zero frontend errors (was 1,525 errors before the LLVM 22 toolchain)
libm 1c64b167891210the fmod rename PROVEN across a module move in 19ms (rename-aware pairing)
memchr bd6068c509020
anyhow bf3ed91242000gate passed
itoa af773857000gate passed
ryu f0b52bb0030function surface changed; gate failed honestly
bitflags 0fc3762174000gate passed
rust-base64 7cffce630501590
heapless 4d515c9not comparableno_std crates need target/test setup the harness does not provide yet; named boundary

Rust support is experimental and labeled as such: these rows are early evidence, not a track record.

Audit

Whole-repo optimization audits

The audit compiles every function, searches with three independent engines, and keeps only what the prover certifies. Wins ship as re-provable receipts; everything that drops out is named.

repo · targetsubmittedeligibleprovenwinsnote
zlib · host11151241inflateValidate 11.2% smaller (89 → 79 bytes), receipt re-proves with one command
littlefs · host22286430eligibility doubled by closure recovery
littlefs · cortex-m419859310real MCU filesystem, sized in the target's own bytes
xxHash · host762131310xxHash is already tight; 4 non-terminating references named as oracle-timeouts, not hung
xxHash · cortex-m4762cross-compiled3 oracle-timeouts named
printf (embedded) · cortex-m411330embedded printf audited on the target class it is written for
BLAKE3 c/ · host, full search40930hand-tuned by world experts; the zero is the tool telling the truth. First recorded full-depth cost: $21.25, solver median 90s
FreeRTOS-Kernel · cortex-m412721100the most deployed scheduler in embedded, audited on its real target with a 23-line config stub via --cflags (was 0 eligible before config provisioning shipped)
CMSIS-NN · cortex-m414979340ARM's production quantized neural-network kernels on their real target; 53% eligibility, the campaign's highest: quantized ML is integer math, and integer math is home turf
lwip · cortex-m411519130the TCP/IP stack on its real target, unlocked with three stub headers via --cflags
Adversarial record

The gate under attack

A second battery attacks everything AROUND the prover

The trust boundary is the whole chain, not just the solver: a pipeline adversarial battery replays every recorded incident class end-to-end (preprocessor-excluded changes, renames, dropped target flags, non-terminating references, missing configs) and asserts the system-level verdict is never a false approval. Latest run 2026-08-24: 7 attacks, 0 false approvals (2 clean runs). The defect-discovery curve is published; it flattens only if campaigns stop finding new classes.

0 unsound accepts across 15 consecutive battery runs (as of 2026-08-24)

Each battery plants thousands of deliberately broken rewrites (an independent concrete oracle confirms they change behavior) and checks the gate never blesses one. 17,554 recorded adversarial verdicts and counting. The battery re-runs after every change that touches a verdict path; nothing ships on a red battery.

The bounded-loop condition, found by attacking ourselves

Loop proofs cover executions up to an unroll bound; a divergence only beyond that bound is invisible to the prover. We demonstrated this against our own tool, then made every loop-bearing PROVEN carry the condition explicitly, on every surface. Population check: 80 of 80 above-bound mutants received the conditional label, zero slipped through unlabeled.

Frontiers, measured (experimental)

Floating point: 4 rewrite identities PROVEN bit-exact (x*2 to x+x, x/2 to x*0.5, commutativity both ways) and 6 subtle traps REFUTED with counterexamples (signaling-NaN quieting, signed zero, NaN propagation). GPU: per-thread device math now PROVEN on BOTH major targets (NVPTX and AMDGPU, integer and float, after SROA normalization); barriers and shared memory stay out of scope by model, stated plainly.

Live production checks

Against the hosted endpoint on the current stack: PROVEN in 128ms on a real rewrite, DIVERGES with a concrete counterexample in 170ms on a sabotaged constant.

Findings and fixes

What the campaigns broke, and what that bought

Every new repo class stresses a different part of the machine. When something breaks, the failure is diagnosed, fixed, requalified against the battery, and recorded. A selection:

Full methodology, raw transcripts and win bundles are in the repository's campaign book and artifacts directory, available to beta partners. If a claim here has no receipt, tell us: that is a bug in the page.