# Preprint identifies a more effective way to weaken MoE model refusals

_Published Monday, October 5, 2026 at 7:58 AM EDT · Science, AI · Latest · Tier 2 — Notable_

Researchers report that selecting and suppressing model components by their influence on loss weakens safety refusals more effectively than selecting them by activation frequency. The preprint tests five mixture-of-experts language model architectures, which route inputs through selected components called experts, without retraining them.

The researchers ranked experts using 500 benign and 500 malicious prompts, then tested refusals on 100 held-out malicious prompts. Their gradient-based method reduced refusals more than frequency-based selection in 24 of 25 conditions under each of two suppression budgets. In OLMoE, refusals fell from 34 to 9 of 100 prompts without degraded outputs. Results support the method under the tested budgets.

## Sources

- [cs.LG updates on arXiv.org](https://arxiv.org/abs/2610.02910)

---
Canonical: https://techandbusiness.org/newswire/ZkIyw-lE-7AOUPlvkdiPwD
Published: 2026-10-05T11:58:25.442Z
Story chronology: 2026-10-05T04:00:00.000Z
Retrieved: 2026-10-05T13:55:18.205Z
Publisher: Tech & Business (techandbusiness.org)
