INDEPENDENT LLM RESEARCH · MELBOURNE

Maximum energy in one direction. Deliberate nulls everywhere else.

Main Lobe Labs studies how large language models actually behave under pressure: pruning, memorisation, and the gap between benchmark scores and mechanism. We publish everything, with code, controls, and stated confidence. Small lab, narrow beam.

ULA · d = λ/2 · θ₀ = 90°
Steering angle 90 degrees

RESEARCH INDEX

2026-08-29
rev v1

Pauper Consensus: Two Pre-Registered LLM Studies

Agreement across large language models is a measurement instrument, and this paper treats the whole instrument as the object under test, under pre-registration, in two studies built one on top of the other. In Study 1, an identical frozen proposition-aggregation protocol run on two independently selected three-model panels returned opposite registered verdicts on synthetic reasoning traces (−0.159 vs +0.122 nats, both 95% CIs excluding zero); a post-hoc diagnosis located the disagreement in the instrument twice over — a calibration map with no intercept measured against a baseline that had one, and a corpus that could barely contain falsehoods (0 of 607 negative-polarity propositions ever scored by the alignment layer) — and a second registration, git-tagged before any new generation, confirmed the central repair on two panel–corpus configurations sharing no model family (+0.220 and +0.272 nats) while falsifying two companion predictions. A 2026-08-29 correction then retracted a further claim of the draft (the "one vote per source" invariant, whose frozen arm was actually capped and unsigned, so it measured polarity, not capping: no unsigned arm exceeds AUROC 0.6063 anywhere, every signed arm is at least 0.9001) and first reported, post-hoc, the contrast that separates pooling from choosing well: the panel beats the calibration-selected best single source by +0.0448 to +0.0887 nats on 3 of 4 panel-cycles, inconclusively on the fourth. A third cycle (prereg-v3) is registered to test that margin as its primary. Study 2 stands on the repaired instrument: a jury of twelve 3–4B local language-model configurations (four families × three arms) votes PASS / FAIL / NOT_STATED on 8,000 labeled claims constructed from 200 real news articles (96,000 votes in 13.2 h on one Mac Studio), and the votes are aggregated by a Dawid–Skene model with per-arm calibration maps fitted once on a smaller prior corpus and applied frozen (the EM weights themselves are refit unsupervised on the new votes) — the protocol we name Pauper Consensus. Every design rule is inherited from Study 1's corrections: tag before inference (R1), intercept-bearing calibration maps (R2), a corpus with a registered 25% FAIL mass (R3), one signed observation per source (R4). The pre-registered capability gate passes: +0.20289 nats of log-loss over a frozen model-free bar at the registered 25% FAIL-mass design point (95% article-block bootstrap CI [0.18737, 0.21740]; refitting the bar on v2 alone costs the headline −0.020 nats, §13.3), all four test cells green — while the companion registered prediction that the model-free bar's own value would transfer fails. The headline is a comparison against a specified model-free feature bar, not evidence of ensemble superiority: the panel-versus-best-single-source margin is small and unregistered (+0.0448 to +0.0887 nats on 3 of 4 Study 1 cycles, §8.3), and the registered third cycle takes it as its primary. The system gate against a single near-frontier 27B model passes via its cheaper-and-indistinguishable branch: both routes co-fail the only two unsupported claims in the defendant stream (12–0 jury agreement across all twelve configurations, PASS from the 27B on both); with two false claims in the stream the false-claim rates are statistically indistinguishable (95% CI [0.0, 5e-05]) — a cost result, not a quality-equivalence claim — and the jury route is 0.781× length-adjusted cost; the 27B is nonetheless the more accurate single model on the full pool (95.887% vs 93.912% three-state), and the binding hardware barrier is 48 GB GDDR7 memory, not arithmetic. Across the two studies, every "failure" — the cycle-1 flip, the two failed cycle-2 predictions, P9, the mis-specified null, the 0/2 catch — was an instrument property or a registered bound, surfaced under frozen registration, not a failed hypothesis. The honest unit of account for LLM-based verification is the instrument: protocol, corpus, calibration, hardware, cost.

preprintconfidence: likelycode
2026-08-19
rev v1

Do Instructional Fingerprints Produce Stable Expert-Routing Signatures in a Mixture-of-Experts Model?

Mixture-of-Experts (MoE) language models route each token to a small set of experts. This raises a practical question for model fingerprinting: if an instructional trigger-response fingerprint is implanted in an MoE model, does the fingerprint localise to a stable set of experts? If it does, targeted expert removal could be a direct defence. If it does not, pruning-based defences need a different criterion. We test this question on Qwen1.5-MoE-A2.7B. We implant three trigger-response fingerprints using router-trainable full fine-tuning, so that both expert MLPs and routing weights are allowed to change. We then use a routing-only router-activation localiser and compare trigger overlap against same-syntax null prompts. The implant succeeds: all triggers fire, and routing changes measurably. However, we do not detect a stable trigger-specific expert-routing signature above the matched null distribution. In the matched across-model null reruns for seeds 42 and 123, the P95 three-trigger overlap contains 6 experts in both seeds, below the matched null CI-high value of 11. In the full 27,000-triple matched null, only 0.61% and 0.27% of triples fall below that overlap, and 7.57% and 4.32% are at or below it, for seeds 42 and 123 respectively. A weaker-implant sweep does not reach a clean partial-firing regime; the firing confirmatory weaker config remains FFR-saturated, and the negative persists under it. A graded injected router-boost sensitivity control shows that the localiser recovers a known three-expert boost at routing weight 0.10 or higher, while also showing that routing-loss implants can contaminate null prompts. Our contribution is a controlled negative and a methodological warning. Under the tested implant and routing-only localiser, we did not detect a stable trigger-specific expert-routing signature above the matched across-model null distribution in Qwen1.5-MoE-A2.7B. We argue that MoE fingerprint-localisation claims should include matched null controls, specificity-validated localisers, stated localiser channels, sensitivity estimates, and capability-preserving implant regimes before they can support targeted expert-removal defences.

preprintconfidence: likelycode
2026-08-17
rev v3

The Flip Was in the Instrument: Two Pre-Registered Cycles of Cross-Model Proposition Aggregation

Methods that pool several language models treat agreement as evidence, almost always at the level of the final answer. We tested the proposition-level version of that premise under pre-registration, twice. In the first cycle, an identical frozen protocol run on two independently selected three-model panels returned opposite verdicts on the registered primary, held-out Δ log-loss against a covariate baseline: −0.159 nats [−0.249, −0.070], a decisive stop, on one panel and +0.122 [+0.037, +0.197], a decisive go, on the other, with both intervals excluding zero. A paper reporting either panel alone would have been confident and wrong about the method. Post-hoc diagnosis located the disagreement not in the panels but in the instrument, twice over. First, the registered calibration map, a single fitted temperature, has no intercept, while the baseline it was measured against is a logistic regression fitted with one; giving both sides the same two parameters moved the failing panel from stop to go and turned all twelve panel × stratum × arm combinations positive, while every rank-based measurement was invariant by construction. Second, the corpus could barely contain a falsehood: 48% of propositions came from theories in which nothing can be false, only positive-polarity falsehoods ever survived alignment (0 of 607 negative-polarity propositions were scored), and the test splits held 14–16 negatives. Rather than reporting a post-hoc rescue, we froze a second registration (negation-family corpus with depth-5 enrichment, an intercept-bearing calibration map, three falsifiable predictions) and git-tagged it before any of its generation ran. The central prediction was confirmed on two panel–corpus configurations sharing no model family with each other, one panel entirely new and one continuing from cycle 1, both on the enriched corpus: +0.220 [+0.160, +0.280] and +0.272 [+0.180, +0.353], both go, at 81 and 65 test negatives. The other two registered predictions failed: consensus quality does not degrade monotonically with proof depth, and a deny-vote filter that helped one panel hurt the other. One finding is invariant across both cycles, both panels, and every stratum: counting distinct agents rather than claim instances is the difference between AUROC ≈ 0.5–0.6 and AUROC 0.89–1.00. Agreement measures who asserts a proposition, not how much text asserts it.

preprintconfidence: likely
2026-08-04
rev v4

Measured Pruning Damage Depends on the Evaluation Corpus: A Renaming Control for Mixture-of-Experts Expert Pruning

Pruning 32 experts per layer from a 256-expert MoE costs 44% of base NLL on CPython stdlib and 16% on the same files with identifiers renamed: the reported damage figure moves 1.6× in absolute nats under a rewrite that changes no program structure. Flip rates at confidently predicted positions halve. An extractability probe confirms the corpora differ in memorisation as intended, but per-file extractability does not predict per-file damage, so memorisation is reported as the hypothesised mechanism rather than a demonstrated one.

preprintconfidence: likelycodebibtex

AGENDA

THE MAIN LOBE

Mechanistic questions about model compression and memorisation: what pruning, quantisation, and distillation actually remove, and how memorised data distorts every measurement we make of them.

Everything ships with code, controls, and a stated confidence level. Negative results are results.

THE NULLS

No frontier-scale pretraining. No product roadmap. No benchmark chasing, thought leadership, or papers whose headline number lacks a control.

An antenna gains directivity by choosing where not to radiate. So does a lab.

ABOUT

Main Lobe Labs is an independent research lab in Melbourne, Australia, run by practitioners who build and govern production AI systems by day and take models apart by night.

We think the interesting problems in applied LLM research are measurement problems, and that a small lab with a narrow beam can move faster on them than a large one with a wide one.