The Flip Was in the Instrument: Two Pre-Registered Cycles of Cross-Model Proposition Aggregation
The same frozen protocol returned opposite verdicts on two panels; a second pre-registered cycle located the flip in the instrument and confirmed the central prediction on both.
At a glance
- Design
- Two pre-registered cycles; the second registration frozen and git-tagged before any of its generation ran
- First cycle
- Opposite verdicts from an identical frozen protocol: −0.159 nats [−0.249, −0.070], a decisive stop, on one panel; +0.122 [+0.037, +0.197], a decisive go, on the other
- Diagnosis
- The instrument, twice over: a calibration map with no intercept, and test splits that held 14–16 negatives
- Second cycle
- Central prediction confirmed on two panel–corpus configurations sharing no model family: +0.220 [+0.160, +0.280] and +0.272 [+0.180, +0.353], both go, at 81 and 65 test negatives
- Registered failures
- Consensus quality does not degrade monotonically with proof depth; a deny-vote filter that helped one panel hurt the other
- Invariant
- Counting distinct agents rather than claim instances is the difference between AUROC ≈ 0.5–0.6 and AUROC 0.89–1.00
Abstract
Methods that pool several language models treat agreement as evidence, almost always at the level of the final answer. We tested the proposition-level version of that premise under pre-registration, twice. In the first cycle, an identical frozen protocol run on two independently selected three-model panels returned opposite verdicts on the registered primary, held-out Δ log-loss against a covariate baseline: −0.159 nats [−0.249, −0.070], a decisive stop, on one panel and +0.122 [+0.037, +0.197], a decisive go, on the other, with both intervals excluding zero. A paper reporting either panel alone would have been confident and wrong about the method. Post-hoc diagnosis located the disagreement not in the panels but in the instrument, twice over. First, the registered calibration map, a single fitted temperature, has no intercept, while the baseline it was measured against is a logistic regression fitted with one; giving both sides the same two parameters moved the failing panel from stop to go and turned all twelve panel × stratum × arm combinations positive, while every rank-based measurement was invariant by construction. Second, the corpus could barely contain a falsehood: 48% of propositions came from theories in which nothing can be false, only positive-polarity falsehoods ever survived alignment (0 of 607 negative-polarity propositions were scored), and the test splits held 14–16 negatives. Rather than reporting a post-hoc rescue, we froze a second registration (negation-family corpus with depth-5 enrichment, an intercept-bearing calibration map, three falsifiable predictions) and git-tagged it before any of its generation ran. The central prediction was confirmed on two panel–corpus configurations sharing no model family with each other, one panel entirely new and one continuing from cycle 1, both on the enriched corpus: +0.220 [+0.160, +0.280] and +0.272 [+0.180, +0.353], both go, at 81 and 65 test negatives. The other two registered predictions failed: consensus quality does not degrade monotonically with proof depth, and a deny-vote filter that helped one panel hurt the other. One finding is invariant across both cycles, both panels, and every stratum: counting distinct agents rather than claim instances is the difference between AUROC ≈ 0.5–0.6 and AUROC 0.89–1.00. Agreement measures who asserts a proposition, not how much text asserts it.
Method
Run an identical frozen protocol on two independently selected three-model panels and test, under pre-registration, whether proposition-level agreement is evidence: the registered primary is held-out Δ log-loss against a covariate baseline.
Rather than report a post-hoc rescue, the second registration froze a negation-family corpus with depth-5 enrichment, an intercept-bearing calibration map and three falsifiable predictions, git-tagged before any of its generation ran.
What we found
In the first cycle, the registered primary returned opposite verdicts: −0.159 nats [−0.249, −0.070], a decisive stop, on one panel and +0.122 [+0.037, +0.197], a decisive go, on the other, with both intervals excluding zero. A paper reporting either panel alone would have been confident and wrong about the method.
Post-hoc diagnosis located the disagreement in the instrument, twice over: the registered calibration map, a single fitted temperature, has no intercept while the baseline is a logistic regression fitted with one, and the corpus could barely contain a falsehood — 0 of 607 negative-polarity propositions were ever scored, and the test splits held 14–16 negatives.
The central prediction of the second registration was confirmed on two panel–corpus configurations sharing no model family with each other, one panel entirely new and one continuing from cycle 1, both on the enriched corpus: +0.220 [+0.160, +0.280] and +0.272 [+0.180, +0.353], both go, at 81 and 65 test negatives.
Counting distinct agents rather than claim instances is the difference between AUROC ≈ 0.5–0.6 and AUROC 0.89–1.00, across both cycles, both panels, and every stratum. Agreement measures who asserts a proposition, not how much text asserts it.
What's uncertain
The other two registered predictions failed: consensus quality does not degrade monotonically with proof depth, and a deny-vote filter that helped one panel hurt the other.