Skip to main navigation Skip to search Skip to main content

ResMoE: Space-efficient Compression of Mixture of Experts LLMs via Residual Restoration

  • Mengting Ai (Co-first author)
  • , Tianxin Wei (Co-first author)
  • , Yifan Chen* (Co-first author)
  • , Zhichen Zeng
  • , Ritchie Zhao
  • , Girish Varatkar
  • , Bita Darvish Rouhani
  • , Xianfeng Tang
  • , Hanghang Tong
  • , Jingrui He*
  • *Corresponding author for this work

Research output: Chapter in book/report/conference proceedingConference proceedingpeer-review

10 Citations (Scopus)

Abstract

Mixture-of-Experts (MoE) Transformer, the backbone architecture of multiple phenomenal language models, leverages sparsity by activating only a fraction of model parameters for each input token. The sparse structure, while allowing constant time costs, results in space inefficiency: we still need to load all the model parameters during inference. We introduce ResMoE, an innovative MoE approximation framework that utilizes Wasserstein barycenter to extract a common expert (barycenter expert) and approximate the residuals between this barycenter expert and the original ones. ResMoE enhances the space efficiency for inference of large-scale MoE Transformers in a one-shot and data-agnostic manner without retraining while maintaining minimal accuracy loss, thereby paving the way for broader accessibility to large language models. We demonstrate the effectiveness of ResMoE through extensive experiments on Switch Transformer, Mixtral, and DeepSeekMoE models. The results show that ResMoE can reduce the number of parameters in an expert by up to 75% while maintaining comparable performance. The code is available at https://github.com/iDEA-iSAIL-Lab-UIUC/ResMoE, and the supplementary appendix is available at https://famous-blue-raincoat.github.io/mengtingai/files/ResMoE_Appendix.pdf.
Original languageEnglish
Title of host publicationKDD 2025 - Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining
Place of PublicationNew York
PublisherAssociation for Computing Machinery (ACM)
Pages1–12
Number of pages12
Volume1
ISBN (Electronic)9798400712456
DOIs
Publication statusPublished - 20 Jul 2025
Event31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining, KDD 2025 - Toronto Convention Centre, Toronto, Canada
Duration: 3 Aug 20257 Aug 2025
https://dl.acm.org/doi/proceedings/10.1145/3690624 (Conference proceeding)
https://kdd2025.kdd.org/ (Conference website)
https://kdd2025.kdd.org/schedule-at-a-glance/ (Conference schedule)

Publication series

NameProceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining
ISSN (Print)2154-817X

Conference

Conference31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining, KDD 2025
Abbreviated titleKDD 2025
Country/TerritoryCanada
CityToronto
Period3/08/257/08/25
Internet address

User-Defined Keywords

  • compression
  • mixture-of-experts
  • optimal transport
  • wasserstein barycenter

Fingerprint

Dive into the research topics of 'ResMoE: Space-efficient Compression of Mixture of Experts LLMs via Residual Restoration'. Together they form a unique fingerprint.

Cite this