{"schema":"hottub-news/story/v1","story":"st_smudjyxalcju3i4pkv6a","title":"RAZOR: Pruning Replaceable Experts in LLMs","title_en":null,"outlets":1,"brief":null,"page":"https://hottub.news/stories/razor-pruning-replaceable-experts-in-llms-smudjyxalc","count":1,"items":[{"id":"n_smudjyxalcju3i4pkv6a","seq":341235,"kind":"article","title":"RAZOR: Pruning Replaceable Experts in LLMs","url":"https://arxiv.org/abs/2609.30465","summary":"arXiv:2609.30465v1 Announce Type: cross Abstract: Mixture-of-experts (MoE) models activate few experts per token but store the full expert pool. Expert pruning reduces this storage burden; at a fixed pruning budget, the goal is to preserve the original model's output distribution as closely as possible. Yet an expert's usage or contribution magnitude does not by itself determine the damage caused by its removal. What matters is whether the surviving computation can replace…","published":"2026-09-28T04:00:00Z","seen":"2026-09-28T04:38:25Z","lang":"en","source":{"id":"kite-export-arxiv-org-4e3a51","name":"export.arxiv.org","domain":"export.arxiv.org"},"via":"kite-export-arxiv-org-4e3a51","topics":["ai","cs.lg","cs.cl"],"authors":["Mingyang Song, Mao Zheng"],"entities":[{"id":"Q13422881","name":"razor","type":"other"}],"mentions":2,"story":"st_smudjyxalcju3i4pkv6a","category":"science-technology","category_p":0.7099999785423279,"sentiment":"neutral","tone":-0.05000000074505806,"political":0.09000000357627869}]}