{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,3]],"date-time":"2026-06-03T07:54:19Z","timestamp":1780473259508,"version":"3.54.1"},"publisher-location":"New York, NY, USA","reference-count":34,"publisher":"ACM","license":[{"start":{"date-parts":[[2024,11,20]],"date-time":"2024-11-20T00:00:00Z","timestamp":1732060800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by-nc-sa\/4.0\/"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2024,11,20]]},"DOI":"10.1145\/3698038.3698521","type":"proceedings-article","created":{"date-parts":[[2024,11,14]],"date-time":"2024-11-14T06:32:43Z","timestamp":1731565963000},"page":"434-442","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["MoEsaic: Shared Mixture of Experts"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-9506-0003","authenticated-orcid":false,"given":"Umesh","family":"Deshpande","sequence":"first","affiliation":[{"name":"IBM Research, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-3033-3964","authenticated-orcid":false,"given":"Travis","family":"Janssen","sequence":"additional","affiliation":[{"name":"IBM Research, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5874-3750","authenticated-orcid":false,"given":"Mudhakar","family":"Srivatsa","sequence":"additional","affiliation":[{"name":"IBM Research, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4468-3061","authenticated-orcid":false,"given":"Swaminathan","family":"Sundararaman","sequence":"additional","affiliation":[{"name":"IBM Research, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2024,11,20]]},"reference":[{"key":"e_1_3_2_1_1_1","unstructured":"2024. Create Mixtures of Experts with MergeKit. https:\/\/huggingface.co\/blog\/mlabonne\/frankenmoe"},{"key":"e_1_3_2_1_2_1","unstructured":"2024. Measuring the GPU Occupancy of Multi-stream Workloads. https:\/\/developer.nvidia.com\/blog\/measuring-the-gpu-occupancy-of-multi-stream-workloads"},{"key":"e_1_3_2_1_3_1","unstructured":"2024. mergekit\/docs\/moe.md at main \u00b7 arcee-ai\/mergekit. https:\/\/github.com\/arcee-ai\/mergekit\/blob\/main\/docs\/moe.md"},{"key":"e_1_3_2_1_4_1","unstructured":"2024. Open Release of Grok-1. https:\/\/x.ai\/blog\/grok-os"},{"key":"e_1_3_2_1_5_1","unstructured":"2024. vllm-project\/vllm. https:\/\/github.com\/vllm-project\/vllmoriginal-date: 2023-02-09T11:23:20Z."},{"key":"e_1_3_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCBB.2022.3175456"},{"key":"e_1_3_2_1_7_1","doi-asserted-by":"crossref","unstructured":"Mikel Artetxe Shruti Bhosale Naman Goyal Todor Mihaylov Myle Ott Sam Shleifer Xi Victoria Lin Jingfei Du Srinivasan Iyer Ramakanth Pasunuru Giri Anantharaman Xian Li Shuohui Chen Halil Akin Mandeep Baines Louis Martin Xing Zhou Punit Singh Koura Brian O'Horo Jeff Wang Luke Zettlemoyer Mona Diab Zornitsa Kozareva and Ves Stoyanov. 2022. Efficient Large Scale Language Modeling with Mixtures of Experts. arXiv:2112.10684 [cs.CL] https:\/\/arxiv.org\/abs\/2112.10684","DOI":"10.18653\/v1\/2022.emnlp-main.804"},{"key":"e_1_3_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.02153"},{"key":"e_1_3_2_1_9_1","volume-title":"Punica: Multi-Tenant LoRA Serving. arXiv:2310.18547 [cs.DC] https:\/\/arxiv.org\/abs\/2310.18547","author":"Chen Lequn","year":"2023","unstructured":"Lequn Chen, Zihao Ye, Yongji Wu, Danyang Zhuo, Luis Ceze, and Arvind Krishnamurthy. 2023. Punica: Multi-Tenant LoRA Serving. arXiv:2310.18547 [cs.DC] https:\/\/arxiv.org\/abs\/2310.18547"},{"key":"e_1_3_2_1_10_1","unstructured":"Tianyu Chen Shaohan Huang Yuan Xie Binxing Jiao Daxin Jiang Haoyi Zhou Jianxin Li and Furu Wei. 2022. Task-Specific Expert Pruning for Sparse Mixture-of-Experts. arXiv:2206.00277 [cs.LG] https:\/\/arxiv.org\/abs\/2206.00277"},{"key":"e_1_3_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.01138"},{"key":"e_1_3_2_1_12_1","doi-asserted-by":"publisher","unstructured":"Damai Dai Chengqi Deng Chenggang Zhao R. X. Xu Huazuo Gao Deli Chen Jiashi Li Wangding Zeng Xingkai Yu Y. Wu Zhenda Xie Y. K. Li Panpan Huang Fuli Luo Chong Ruan Zhifang Sui and Wenfeng Liang. 2024. DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models. https:\/\/doi.org\/10.48550\/arXiv.2401.06066 arXiv:2401.06066 [cs].","DOI":"10.48550\/arXiv.2401.06066"},{"key":"e_1_3_2_1_13_1","volume-title":"Chi","author":"Hazimeh Hussein","year":"2021","unstructured":"Hussein Hazimeh, Zhe Zhao, Aakanksha Chowdhery, Maheswaran Sathiamoorthy, Yihua Chen, Rahul Mazumder, Lichan Hong, and Ed Chi. 2021. DSelect-k: Differentiable Selection in the Mixture of Experts with Applications to Multi-Task Learning. In Advances in Neural Information Processing Systems, Vol. 34. Curran Associates, Inc., 29335--29347. https:\/\/papers.nips.cc\/paper\/2021\/hash\/f5ac21cd0ef1b88e9848571aeb53551a-Abstract.html"},{"key":"e_1_3_2_1_14_1","unstructured":"Shwai He Daize Dong Liang Ding and Ang Li. 2024. Demystifying the Compression of Mixture-of-Experts Through a Unified Framework. http:\/\/arxiv.org\/abs\/2406.02500 arXiv:2406.02500 [cs]."},{"key":"e_1_3_2_1_15_1","unstructured":"Albert Q. Jiang Alexandre Sablayrolles Antoine Roux Arthur Mensch Blanche Savary Chris Bamford Devendra Singh Chaplot Diego de las Casas Emma Bou Hanna Florian Bressand Gianna Lengyel Guillaume Bour Guillaume Lample L\u00e9lio Renard Lavaud Lucile Saulnier MarieAnne Lachaux Pierre Stock Sandeep Subramanian Sophia Yang Szymon Antoniak Teven Le Scao Th\u00e9ophile Gervet Thibaut Lavril Thomas Wang Timoth\u00e9e Lacroix and William El Sayed. 2024. Mixtral of Experts. arXiv:2401.04088 [cs.LG] https:\/\/arxiv.org\/abs\/2401.04088"},{"key":"e_1_3_2_1_16_1","unstructured":"Jared Kaplan Sam McCandlish Tom Henighan Tom B. Brown Benjamin Chess Rewon Child Scott Gray Alec Radford Jeffrey Wu and Dario Amodei. 2020. Scaling Laws for Neural Language Models. http:\/\/arxiv.org\/abs\/2001.08361 arXiv:2001.08361 [cs stat]."},{"key":"e_1_3_2_1_17_1","unstructured":"Young Jin Kim Raffy Fahim and Hany Hassan Awadalla. 2023. Mixture of Quantized Experts (MoQE): Complementary Effect of Low-bit Quantization and Robustness. arXiv:2310.02410 [cs.LG] https:\/\/arxiv.org\/abs\/2310.02410"},{"key":"e_1_3_2_1_18_1","unstructured":"Jakub Krajewski Jan Ludziejewski Kamil Adamczewski Maciej Pi\u00f3ro Micha\u0142 Krutul Szymon Antoniak Kamil Ciebiera Krystian Kr\u00f3l Tomasz Odrzyg\u00f3\u017ad\u017a Piotr Sankowski Marek Cygan and Sebastian Jaszczur. 2024. Scaling Laws for Fine-Grained Mixture of Experts. arXiv:2402.07871 [cs.LG] https:\/\/arxiv.org\/abs\/2402.07871"},{"key":"e_1_3_2_1_19_1","unstructured":"Leeroo-AI. 2024. mergoo: A library for easily merging multiple LLM experts and efficiently train the merged LLM. https:\/\/github.com\/Leeroo-AI\/mergoo."},{"key":"e_1_3_2_1_20_1","unstructured":"Pingzhi Li Xiaolong Jin Yu Cheng and Tianlong Chen. 2024. Examining Post-Training Quantization for Mixture-of-Experts: A Benchmark. arXiv:2406.08155 [cs.LG] https:\/\/arxiv.org\/abs\/2406.08155"},{"key":"e_1_3_2_1_21_1","volume-title":"The Twelfth International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=eFWG9Cy3WK","author":"Li Pingzhi","year":"2024","unstructured":"Pingzhi Li, Zhenyu Zhang, Prateek Yadav, Yi-Lin Sung, Yu Cheng, Mohit Bansal, and Tianlong Chen. 2024. Merge, Then Compress: Demystify Efficient SMoE with Hints from Its Routing Policy. In The Twelfth International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=eFWG9Cy3WK"},{"key":"e_1_3_2_1_22_1","unstructured":"Suyi Li Hanfeng Lu Tianyuan Wu Minchen Yu Qizhen Weng Xusheng Chen Yizhou Shan Binhang Yuan and Wei Wang. 2024. CaraServe: CPU-Assisted and Rank-Aware LoRA Serving for Generative LLM Inference. https:\/\/api.semanticscholar.org\/CorpusID:267068417"},{"key":"e_1_3_2_1_23_1","unstructured":"Yunxin Li Shenyuan Jiang Baotian Hu Longyue Wang Wanqi Zhong Wenhan Luo Lin Ma and Min Zhang. [n. d.]. Uni-MoE: Scaling Unified Multimodal LLMs with Mixture of Experts. ([n.d.])."},{"key":"e_1_3_2_1_24_1","volume-title":"M3ViT: Mixture-of-Experts Vision Transformer for Efficient Multi-task Learning with Model-Accelerator Co-design. Advances in Neural Information Processing Systems 35 (Dec","author":"Liang Hanxue","year":"2022","unstructured":"Hanxue Liang, Zhiwen Fan, Rishov Sarkar, Ziyu Jiang, Tianlong Chen, Kai Zou, Yu Cheng, Cong Hao, and Zhangyang Wang. 2022. M3ViT: Mixture-of-Experts Vision Transformer for Efficient Multi-task Learning with Model-Accelerator Co-design. Advances in Neural Information Processing Systems 35 (Dec. 2022), 28441--28457. https:\/\/papers.nips.cc\/paper_files\/paper\/2022\/hash\/b653f34d576d1790481e3797cb740214-Abstract-Conference.html"},{"key":"e_1_3_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/3219819.3220007"},{"key":"e_1_3_2_1_26_1","unstructured":"Mohammed Muqeeth Haokun Liu and Colin Raffel. 2024. Soft Merging of Experts with Adaptive Routing. arXiv:2306.03745 [cs.LG] https:\/\/arxiv.org\/abs\/2306.03745"},{"key":"e_1_3_2_1_27_1","volume-title":"Oh (Eds.)","volume":"35","author":"Mustafa Basil","year":"2022","unstructured":"Basil Mustafa, Carlos Riquelme, Joan Puigcerver, Rodolphe Jenatton, and Neil Houlsby. 2022. Multimodal Contrastive Learning with LIMoE: the Language-Image Mixture of Experts. In Advances in Neural Information Processing Systems, S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh (Eds.), Vol. 35. Curran Associates, Inc., 9564--9576. https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2022\/file\/3e67e84abf900bb2c7cbd5759bfce62d-Paper-Conference.pdf"},{"key":"e_1_3_2_1_28_1","unstructured":"Noam Shazeer Azalia Mirhoseini Krzysztof Maziarz Andy Davis Quoc Le Geoffrey Hinton and Jeff Dean. 2017. Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer. (2017)."},{"key":"e_1_3_2_1_29_1","doi-asserted-by":"publisher","unstructured":"Ying Sheng Shiyi Cao Dacheng Li Coleman Hooper Nicholas Lee Shuo Yang Christopher Chou Banghua Zhu Lianmin Zheng Kurt Keutzer Joseph E. Gonzalez and Ion Stoica. 2024. S-LoRA: Serving Thousands of Concurrent LoRA Adapters. https:\/\/doi.org\/10.48550\/arXiv.2311.03285 arXiv:2311.03285 [cs].","DOI":"10.48550\/arXiv.2311.03285"},{"key":"e_1_3_2_1_30_1","volume-title":"Baptiste Rozi\u00e8re, Jacob Kahn, Daniel Li, Wen-tau Yih, Jason Weston, and Xian Li.","author":"Sukhbaatar Sainbayar","year":"2024","unstructured":"Sainbayar Sukhbaatar, Olga Golovneva, Vasu Sharma, Hu Xu, Xi Victoria Lin, Baptiste Rozi\u00e8re, Jacob Kahn, Daniel Li, Wen-tau Yih, Jason Weston, and Xian Li. 2024. Branch-Train-MiX: Mixing Expert LLMs into a Mixture-of-Experts LLM. arXiv preprint arXiv:2403.07816 (2024)."},{"key":"e_1_3_2_1_31_1","doi-asserted-by":"publisher","unstructured":"Gemini Team. 2024. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context. https:\/\/doi.org\/10.48550\/arXiv.2403.05530 arXiv:2403.05530 [cs] version: 1.","DOI":"10.48550\/arXiv.2403.05530"},{"key":"e_1_3_2_1_32_1","volume-title":"His Last: Expert-Specialized Fine-Tuning for Sparse Architectural Large Language Models.","author":"Wang Zihan","year":"2024","unstructured":"Zihan Wang, Deli Chen, Damai Dai, Runxin Xu, Zhuoshu Li, and Y. Wu. 2024. Let the Expert Stick to His Last: Expert-Specialized Fine-Tuning for Sparse Architectural Large Language Models. (2024). arXiv:2407.01906 [cs.CL] https:\/\/arxiv.org\/abs\/2407.01906"},{"key":"e_1_3_2_1_33_1","volume-title":"Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush.","author":"Wolf Thomas","year":"2020","unstructured":"Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Perric Cistac, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush. 2020. Transformers: State-of-the-Art Natural Language Processing. Association for Computational Linguistics, 38--45. https:\/\/www.aclweb.org\/anthology\/2020.emnlp-demos.6"},{"key":"e_1_3_2_1_34_1","doi-asserted-by":"crossref","unstructured":"Tong Zhu Xiaoye Qu Daize Dong Jiacheng Ruan Jingqi Tong Conghui He and Yu Cheng. 2024. LLaMA-MoE: Building Mixture-of-Experts from LLaMA with Continual Pre-training. arXiv:2406.16554 [cs.CL] https:\/\/arxiv.org\/abs\/2406.16554","DOI":"10.18653\/v1\/2024.emnlp-main.890"}],"event":{"name":"SoCC '24: ACM Symposium on Cloud Computing","location":"Redmond WA USA","acronym":"SoCC '24","sponsor":["SIGMOD ACM Special Interest Group on Management of Data","SIGOPS ACM Special Interest Group on Operating Systems"]},"container-title":["Proceedings of the ACM Symposium on Cloud Computing"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3698038.3698521","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3698038.3698521","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,8,22]],"date-time":"2025-08-22T18:59:25Z","timestamp":1755889165000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3698038.3698521"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,11,20]]},"references-count":34,"alternative-id":["10.1145\/3698038.3698521","10.1145\/3698038"],"URL":"https:\/\/doi.org\/10.1145\/3698038.3698521","relation":{},"subject":[],"published":{"date-parts":[[2024,11,20]]},"assertion":[{"value":"2024-11-20","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}