Do Instructional Fingerprints Produce Stable Expert-Routing Signatures in a Mixture-of-Experts Model?
The implant succeeds and routing changes measurably, but no stable trigger-specific expert-routing signature is detected above the matched null distribution.
At a glance
- Model
- Qwen1.5-MoE-A2.7B
- Intervention
- Three trigger-response fingerprints implanted with router-trainable full fine-tuning
- Localiser
- A routing-only router-activation localiser; trigger overlap compared against same-syntax null prompts
- Headline
- No stable trigger-specific expert-routing signature detected above the matched null distribution
- Key numbers
- P95 three-trigger overlap of 6 experts in both seeds, below the matched null CI-high value of 11
- Contribution
- A controlled negative and a methodological warning
Abstract
Mixture-of-Experts (MoE) language models route each token to a small set of experts. This raises a practical question for model fingerprinting: if an instructional trigger-response fingerprint is implanted in an MoE model, does the fingerprint localise to a stable set of experts? If it does, targeted expert removal could be a direct defence. If it does not, pruning-based defences need a different criterion. We test this question on Qwen1.5-MoE-A2.7B. We implant three trigger-response fingerprints using router-trainable full fine-tuning, so that both expert MLPs and routing weights are allowed to change. We then use a routing-only router-activation localiser and compare trigger overlap against same-syntax null prompts. The implant succeeds: all triggers fire, and routing changes measurably. However, we do not detect a stable trigger-specific expert-routing signature above the matched null distribution. In the matched across-model null reruns for seeds 42 and 123, the P95 three-trigger overlap contains 6 experts in both seeds, below the matched null CI-high value of 11. In the full 27,000-triple matched null, only 0.61% and 0.27% of triples fall below that overlap, and 7.57% and 4.32% are at or below it, for seeds 42 and 123 respectively. A weaker-implant sweep does not reach a clean partial-firing regime; the firing confirmatory weaker config remains FFR-saturated, and the negative persists under it. A graded injected router-boost sensitivity control shows that the localiser recovers a known three-expert boost at routing weight 0.10 or higher, while also showing that routing-loss implants can contaminate null prompts. Our contribution is a controlled negative and a methodological warning. Under the tested implant and routing-only localiser, we did not detect a stable trigger-specific expert-routing signature above the matched across-model null distribution in Qwen1.5-MoE-A2.7B. We argue that MoE fingerprint-localisation claims should include matched null controls, specificity-validated localisers, stated localiser channels, sensitivity estimates, and capability-preserving implant regimes before they can support targeted expert-removal defences.
Method
Test whether an instructional trigger-response fingerprint implanted in an MoE model localises to a stable set of experts: if it does, targeted expert removal could be a direct defence; if it does not, pruning-based defences need a different criterion.
Implant three trigger-response fingerprints in Qwen1.5-MoE-A2.7B using router-trainable full fine-tuning, so that both expert MLPs and routing weights are allowed to change. Then use a routing-only router-activation localiser and compare trigger overlap against same-syntax null prompts, with matched across-model null reruns for seeds 42 and 123 and a graded injected router-boost sensitivity control.
What we found
The implant succeeds: all triggers fire, and routing changes measurably. However, no stable trigger-specific expert-routing signature is detected above the matched null distribution.
In the matched across-model null reruns for seeds 42 and 123, the P95 three-trigger overlap contains 6 experts in both seeds, below the matched null CI-high value of 11. In the full 27,000-triple matched null, only 0.61% and 0.27% of triples fall below that overlap, and 7.57% and 4.32% are at or below it.
A graded injected router-boost sensitivity control shows that the localiser recovers a known three-expert boost at routing weight 0.10 or higher.
The contribution is a controlled negative and a methodological warning: MoE fingerprint-localisation claims should include matched null controls, specificity-validated localisers, stated localiser channels, sensitivity estimates, and capability-preserving implant regimes before they can support targeted expert-removal defences.
What's uncertain
A weaker-implant sweep does not reach a clean partial-firing regime; the firing confirmatory weaker config remains FFR-saturated, and the negative persists under it.
Routing-loss implants can contaminate null prompts.