2026-08-29 · rev v1

Pauper Consensus: Two Pre-Registered LLM Studies

The pre-registered capability gate passes at +0.20289 nats over a frozen model-free bar — a comparison against a specified model-free feature bar, not evidence of ensemble superiority.

preprintconfidence: likelycode

At a glance

Design
Two pre-registered studies built one on top of the other
Study 1
An identical frozen proposition-aggregation protocol on two independently selected three-model panels returned opposite registered verdicts (−0.159 vs +0.122 nats, both 95% CIs excluding zero)
Diagnosis
A calibration map with no intercept measured against a baseline that had one, and a corpus that could barely contain falsehoods (0 of 607 negative-polarity propositions ever scored by the alignment layer)
Repair
A second registration, git-tagged before any new generation, confirmed the central repair on two panel–corpus configurations sharing no model family (+0.220 and +0.272 nats)
Protocol
Pauper Consensus — a jury of twelve 3–4B local language-model configurations (four families × three arms) voting PASS / FAIL / NOT_STATED, aggregated by a Dawid–Skene model with per-arm calibration maps applied frozen
Corpus
8,000 labeled claims constructed from 200 real news articles — 96,000 votes in 13.2 h on one Mac Studio
Capability gate
+0.20289 nats of log-loss over a frozen model-free bar at the registered 25% FAIL-mass design point (95% article-block bootstrap CI [0.18737, 0.21740]), all four test cells green
System gate
Passes against a single near-frontier 27B model via its cheaper-and-indistinguishable branch; the jury route is 0.781× length-adjusted cost

Abstract

Agreement across large language models is a measurement instrument, and this paper treats the whole instrument as the object under test, under pre-registration, in two studies built one on top of the other. In Study 1, an identical frozen proposition-aggregation protocol run on two independently selected three-model panels returned opposite registered verdicts on synthetic reasoning traces (−0.159 vs +0.122 nats, both 95% CIs excluding zero); a post-hoc diagnosis located the disagreement in the instrument twice over — a calibration map with no intercept measured against a baseline that had one, and a corpus that could barely contain falsehoods (0 of 607 negative-polarity propositions ever scored by the alignment layer) — and a second registration, git-tagged before any new generation, confirmed the central repair on two panel–corpus configurations sharing no model family (+0.220 and +0.272 nats) while falsifying two companion predictions. A 2026-08-29 correction then retracted a further claim of the draft (the "one vote per source" invariant, whose frozen arm was actually capped and unsigned, so it measured polarity, not capping: no unsigned arm exceeds AUROC 0.6063 anywhere, every signed arm is at least 0.9001) and first reported, post-hoc, the contrast that separates pooling from choosing well: the panel beats the calibration-selected best single source by +0.0448 to +0.0887 nats on 3 of 4 panel-cycles, inconclusively on the fourth. A third cycle (prereg-v3) is registered to test that margin as its primary. Study 2 stands on the repaired instrument: a jury of twelve 3–4B local language-model configurations (four families × three arms) votes PASS / FAIL / NOT_STATED on 8,000 labeled claims constructed from 200 real news articles (96,000 votes in 13.2 h on one Mac Studio), and the votes are aggregated by a Dawid–Skene model with per-arm calibration maps fitted once on a smaller prior corpus and applied frozen (the EM weights themselves are refit unsupervised on the new votes) — the protocol we name Pauper Consensus. Every design rule is inherited from Study 1's corrections: tag before inference (R1), intercept-bearing calibration maps (R2), a corpus with a registered 25% FAIL mass (R3), one signed observation per source (R4). The pre-registered capability gate passes: +0.20289 nats of log-loss over a frozen model-free bar at the registered 25% FAIL-mass design point (95% article-block bootstrap CI [0.18737, 0.21740]; refitting the bar on v2 alone costs the headline −0.020 nats, §13.3), all four test cells green — while the companion registered prediction that the model-free bar's own value would transfer fails. The headline is a comparison against a specified model-free feature bar, not evidence of ensemble superiority: the panel-versus-best-single-source margin is small and unregistered (+0.0448 to +0.0887 nats on 3 of 4 Study 1 cycles, §8.3), and the registered third cycle takes it as its primary. The system gate against a single near-frontier 27B model passes via its cheaper-and-indistinguishable branch: both routes co-fail the only two unsupported claims in the defendant stream (12–0 jury agreement across all twelve configurations, PASS from the 27B on both); with two false claims in the stream the false-claim rates are statistically indistinguishable (95% CI [0.0, 5e-05]) — a cost result, not a quality-equivalence claim — and the jury route is 0.781× length-adjusted cost; the 27B is nonetheless the more accurate single model on the full pool (95.887% vs 93.912% three-state), and the binding hardware barrier is 48 GB GDDR7 memory, not arithmetic. Across the two studies, every "failure" — the cycle-1 flip, the two failed cycle-2 predictions, P9, the mis-specified null, the 0/2 catch — was an instrument property or a registered bound, surfaced under frozen registration, not a failed hypothesis. The honest unit of account for LLM-based verification is the instrument: protocol, corpus, calibration, hardware, cost.

Method

Treat agreement across large language models as a measurement instrument, and treat the whole instrument as the object under test, under pre-registration, in two studies built one on top of the other.

Study 1 runs an identical frozen proposition-aggregation protocol on two independently selected three-model panels over synthetic reasoning traces; a second registration, git-tagged before any new generation, then confirms the central repair.

Study 2 stands on the repaired instrument: a jury of twelve 3–4B local language-model configurations (four families × three arms) votes PASS / FAIL / NOT_STATED on 8,000 labeled claims constructed from 200 real news articles (96,000 votes in 13.2 h on one Mac Studio), and the votes are aggregated by a Dawid–Skene model with per-arm calibration maps fitted once on a smaller prior corpus and applied frozen, while the EM weights themselves are refit unsupervised on the new votes — the protocol named Pauper Consensus.

Every design rule is inherited from Study 1's corrections: tag before inference (R1), intercept-bearing calibration maps (R2), a corpus with a registered 25% FAIL mass (R3), and one signed observation per source (R4).

What we found

In Study 1 the identical protocol returned opposite registered verdicts on the two panels (−0.159 vs +0.122 nats, both 95% CIs excluding zero). A post-hoc diagnosis located the disagreement in the instrument twice over — a calibration map with no intercept measured against a baseline that had one, and a corpus that could barely contain falsehoods (0 of 607 negative-polarity propositions ever scored by the alignment layer).

The second registration confirmed the central repair on two panel–corpus configurations sharing no model family (+0.220 and +0.272 nats), while falsifying two companion predictions.

In Study 2 the pre-registered capability gate passes: +0.20289 nats of log-loss over a frozen model-free bar at the registered 25% FAIL-mass design point (95% article-block bootstrap CI [0.18737, 0.21740]), with all four test cells green.

The system gate against a single near-frontier 27B model passes via its cheaper-and-indistinguishable branch: both routes co-fail the only two unsupported claims in the defendant stream (12–0 jury agreement across all twelve configurations, PASS from the 27B on both), the false-claim rates are statistically indistinguishable (95% CI [0.0, 5e-05]), and the jury route is 0.781× length-adjusted cost.

Across the two studies, every "failure" — the cycle-1 flip, the two failed cycle-2 predictions, P9, the mis-specified null, the 0/2 catch — was an instrument property or a registered bound, surfaced under frozen registration, not a failed hypothesis. The honest unit of account for LLM-based verification is the instrument: protocol, corpus, calibration, hardware, cost.

What's uncertain

The headline is a comparison against a specified model-free feature bar, not evidence of ensemble superiority. The panel-versus-best-single-source margin is small and unregistered (+0.0448 to +0.0887 nats on 3 of 4 Study 1 cycles), and a registered third cycle (prereg-v3) takes it as its primary.

A 2026-08-29 correction retracted a further claim of the draft — the "one vote per source" invariant, whose frozen arm was actually capped and unsigned, so it measured polarity, not capping (no unsigned arm exceeds AUROC 0.6063 anywhere, every signed arm is at least 0.9001).

The companion registered prediction that the model-free bar's own value would transfer fails, and refitting the bar on v2 alone costs the headline −0.020 nats.

The indistinguishable false-claim rates are a cost result, not a quality-equivalence claim: the 27B is nonetheless the more accurate single model on the full pool (95.887% vs 93.912% three-state), and the binding hardware barrier is 48 GB GDDR7 memory, not arithmetic.

Authors

Jeremiah Mannings

Founder · applied LLM research

Andryo Marzuki

Founder · applied LLM research