Logprob Verification with Majority Consensus
Verifiers independently check output-token log probabilities in parallel, without regenerating the answer, and reach majority consensus. In Qwen2.5-7B tests, three replays together used 1.9%–14.8% of generation GPU time. Similar efficiency gains are expected for larger dense models under comparable conditions, especially with long outputs; exact ratios vary.

