2026-08-19 · rev v1

Do Instructional Fingerprints Produce Stable Expert-Routing Signatures in a Mixture-of-Experts Model?

The implant succeeds and routing changes measurably, but no stable trigger-specific expert-routing signature is detected above the matched null distribution.

preprintconfidence: likelycode

At a glance

Model
Qwen1.5-MoE-A2.7B
Intervention
Three trigger-response fingerprints implanted with router-trainable full fine-tuning
Localiser
A routing-only router-activation localiser; trigger overlap compared against same-syntax null prompts
Headline
No stable trigger-specific expert-routing signature detected above the matched null distribution
Key numbers
P95 three-trigger overlap of 6 experts in both seeds, below the matched null CI-high value of 11
Contribution
A controlled negative and a methodological warning

Abstract

Mixture-of-Experts (MoE) language models route each token to a small set of experts. This raises a practical question for model fingerprinting: if an instructional trigger-response fingerprint is implanted in an MoE model, does the fingerprint localise to a stable set of experts? If it does, targeted expert removal could be a direct defence. If it does not, pruning-based defences need a different criterion. We test this question on Qwen1.5-MoE-A2.7B. We implant three trigger-response fingerprints using router-trainable full fine-tuning, so that both expert MLPs and routing weights are allowed to change. We then use a routing-only router-activation localiser and compare trigger overlap against same-syntax null prompts. The implant succeeds: all triggers fire, and routing changes measurably. However, we do not detect a stable trigger-specific expert-routing signature above the matched null distribution. In the matched across-model null reruns for seeds 42 and 123, the P95 three-trigger overlap contains 6 experts in both seeds, below the matched null CI-high value of 11. In the full 27,000-triple matched null, only 0.61% and 0.27% of triples fall below that overlap, and 7.57% and 4.32% are at or below it, for seeds 42 and 123 respectively. A weaker-implant sweep does not reach a clean partial-firing regime; the firing confirmatory weaker config remains FFR-saturated, and the negative persists under it. A graded injected router-boost sensitivity control shows that the localiser recovers a known three-expert boost at routing weight 0.10 or higher, while also showing that routing-loss implants can contaminate null prompts. Our contribution is a controlled negative and a methodological warning. Under the tested implant and routing-only localiser, we did not detect a stable trigger-specific expert-routing signature above the matched across-model null distribution in Qwen1.5-MoE-A2.7B. We argue that MoE fingerprint-localisation claims should include matched null controls, specificity-validated localisers, stated localiser channels, sensitivity estimates, and capability-preserving implant regimes before they can support targeted expert-removal defences.

Method

Test whether an instructional trigger-response fingerprint implanted in an MoE model localises to a stable set of experts: if it does, targeted expert removal could be a direct defence; if it does not, pruning-based defences need a different criterion.

Implant three trigger-response fingerprints in Qwen1.5-MoE-A2.7B using router-trainable full fine-tuning, so that both expert MLPs and routing weights are allowed to change. Then use a routing-only router-activation localiser and compare trigger overlap against same-syntax null prompts, with matched across-model null reruns for seeds 42 and 123 and a graded injected router-boost sensitivity control.

What we found

The implant succeeds: all triggers fire, and routing changes measurably. However, no stable trigger-specific expert-routing signature is detected above the matched null distribution.

In the matched across-model null reruns for seeds 42 and 123, the P95 three-trigger overlap contains 6 experts in both seeds, below the matched null CI-high value of 11. In the full 27,000-triple matched null, only 0.61% and 0.27% of triples fall below that overlap, and 7.57% and 4.32% are at or below it.

A graded injected router-boost sensitivity control shows that the localiser recovers a known three-expert boost at routing weight 0.10 or higher.

The contribution is a controlled negative and a methodological warning: MoE fingerprint-localisation claims should include matched null controls, specificity-validated localisers, stated localiser channels, sensitivity estimates, and capability-preserving implant regimes before they can support targeted expert-removal defences.

What's uncertain

A weaker-implant sweep does not reach a clean partial-firing regime; the firing confirmatory weaker config remains FFR-saturated, and the negative persists under it.

Routing-loss implants can contaminate null prompts.

Authors

Andryo Marzuki

Founder · applied LLM research

Jeremiah Mannings

Founder · applied LLM research