Three Failed Harnesses, Then a Provable Consensus Panic
Static analysis found 321 panic sites in a BFT consensus engine. Most are unreachable noise. This is the process that separated one provable, network-triggered crash from the rest, including the two attempts that honestly failed.
The haystack
A BFT consensus engine written in Rust is full of panic! calls. That is by design: the engine asserts invariants that should be impossible to violate, and crashing loudly is considered safer than continuing with corrupted state. On the engine I was auditing, a static pass counted 321 panic sites across the core crates.
Most of those are defensive programming. The valuable question for a security report is narrower: can a peer on the network drive an honest validator into one of these panics with ordinary consensus messages? If yes, that is a remote liveness kill against a validator that did nothing wrong.
Three candidates stood out. Two were in the vote bookkeeping layer, guarding the invariant that precommit certificates are only stored when the votes behind them still exist. One was in the decision handler: it builds a commit certificate from locally stored precommits, verifies it through the host engine, and panics if verification fails.
Attempt one and two: the honest failures
The first harness reproduced the obvious race: record a precommit threshold, prune the votes, then deliver a delayed output that reads them. In a synchronous unit test of the driver, the panic never fires. The state reads that check the threshold and the state reads that use the votes are the same map, so pruning removes both conditions at once. There is no window.
The second harness drove the full consensus path, real inputs, real timeouts. Same result: once two-thirds of validators precommit, the state machine moves straight to a decision, and the delayed-output interleaving I was trying to force simply does not occur in a synchronous run.
Both results went into the notes as not reachable in a synchronous driver. That is an unsatisfying answer and an honest one. The panic sites are real, the guard is really missing, but no interleaving I could construct in-process reached it.
Attempt three: read how the host actually calls in
The consensus engine is not a plain library. It runs as a generator-style state machine that yields effects to its host, and the host resolves them and resumes execution with an answer. Verification is one such effect. The host can answer a verification effect with a failure, and the decision handler treats a failed verification of its own locally-built certificate as an unrecoverable invariant break.
The test suite for this engine has a pattern for exactly that: drive the state machine with a handler that answers every effect, and control what the verification effect returns. The third harness used it.
Start a height with three validators. Submit a valid proposal, then the two-thirds prevote, then the two-thirds precommit. The state machine reaches its decision and builds the commit certificate from the precommits it accumulated. When the verification effect fires, answer it with a failure: a signer whose address is not in the validator set.
The result:
thread 'zz_fuzz_decide_certificate_invalid_panics' panicked at
core-consensus/src/handle/decide.rs:56:13:
Decide: Commit certificate is not valid: UnknownValidator(...)
test result: ok. 1 passed; 0 failed
The test passes because it expects the panic. The panic is reached from ordinary peer input: three valid precommits bring an honest validator to a decision, and a single failed verification of its own certificate crashes the whole node.
What the failure answers mean in production
When can verification of a locally-built certificate actually fail? The signatures were already checked when the votes arrived. The realistic windows are validator set changes between heights, an unbonding race that removes a signer mid-flight, or a corrupted certificate in state. All of those are supposed to be impossible. The engine treats them as impossible in a way that turns them into a remote liveness kill: an honest validator crashes at the worst possible moment, during finalization.
The earlier failures had also narrowed the search space. In a synchronous driver there is no window between the threshold check and the vote read. That is precisely why the reachable path goes through the effect boundary, where the host, not the driver, controls the timing of the answer.
The general method
- Static pass to enumerate panic sites, then rank by how close they sit to network input.
- Reproduce the suspected race directly against the component. Expect failure and record it honestly; a failed reproduction is real information about where the boundary is.
- Read how the host drives the component. Generators, effect handlers, and trait-object callbacks each have a place where an attacker-controlled or environment-controlled answer enters synchronously.
- Write the harness in the project's own test idiom. A passing
#[should_panic]test in the upstream crate, at the upstream commit, is a receipt that is hard to argue with.
The same audit later produced a small bounded model checker for this class of state machine, which found the exact interleaving the first two harnesses missed: a pending-output queue sitting between threshold detection and delivery. The harness proved one path; the model checker explains the family. Both matter when you are asking a team to fix a panic that their tests say is unreachable.