{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,11,17]],"date-time":"2025-11-17T12:09:09Z","timestamp":1763381349499,"version":"3.45.0"},"publisher-location":"New York, NY, USA","reference-count":20,"publisher":"ACM","funder":[{"DOI":"10.13039\/100000001","name":"NSF (National Science Foundation)","doi-asserted-by":"publisher","award":["CCF-2326606"],"award-info":[{"award-number":["CCF-2326606"]}],"id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2025,11,17]]},"DOI":"10.1145\/3772356.3772379","type":"proceedings-article","created":{"date-parts":[[2025,11,17]],"date-time":"2025-11-17T12:02:48Z","timestamp":1763380968000},"page":"131-138","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Lost in Translation: The Search for Meaning in Network-Attached AI Accelerator Disaggregation"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0009-0002-7769-9836","authenticated-orcid":false,"given":"Jaewan","family":"Hong","sequence":"first","affiliation":[{"name":"UC Berkeley, Berkeley, California, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0003-3651-6973","authenticated-orcid":false,"given":"Yifan","family":"Qiao","sequence":"additional","affiliation":[{"name":"UC Berkeley, Berkeley, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-1449-1447","authenticated-orcid":false,"given":"Soujanya","family":"Ponnapalli","sequence":"additional","affiliation":[{"name":"UC Berkeley, Berkeley, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0009-0825-6178","authenticated-orcid":false,"given":"Shu","family":"Liu","sequence":"additional","affiliation":[{"name":"UC Berkeley, Berkeley, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3489-2468","authenticated-orcid":false,"given":"Marcos K.","family":"Aguilera","sequence":"additional","affiliation":[{"name":"NVIDIA, Santa Clara, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7683-208X","authenticated-orcid":false,"given":"Vincent","family":"Liu","sequence":"additional","affiliation":[{"name":"University of Pennsylvania, Philadelphia, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0329-3657","authenticated-orcid":false,"given":"Christopher J.","family":"Rossbach","sequence":"additional","affiliation":[{"name":"UT Austin and Microsoft, Austin, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5373-0088","authenticated-orcid":false,"given":"Ion","family":"Stoica","sequence":"additional","affiliation":[{"name":"UC Berkeley, Berkeley, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2025,11,17]]},"reference":[{"key":"e_1_3_2_1_1_1","first-page":"150","volume-title":"Proceedings of the 50th Annual IEEE\/ACM International Symposium on Microarchitecture","author":"Ausavarungnirun Rachata","year":"2017","unstructured":"Rachata Ausavarungnirun, Joshua Landgraf, Vance Miller, Saugata Ghose, Jayneel Gandhi, Christopher J Rossbach, and Onur Mutlu. Mosaic: a gpu memory manager with application-transparent support for multiple page sizes. In Proceedings of the 50th Annual IEEE\/ACM International Symposium on Microarchitecture, pages 136\u2013150, 2017."},{"key":"e_1_3_2_1_2_1","volume-title":"January","year":"2025","unstructured":"Broadcom. End of availability (eoa) for vsphere bitfusion, January 2025. Accessed: 2025-01-15."},{"key":"e_1_3_2_1_3_1","first-page":"219","article-title":"Pytorch rpc: Distributed deep learning built on tensor-optimized remote procedure calls","volume":"5","author":"Damania Pritam","year":"2023","unstructured":"Pritam Damania, Shen Li, Alban Desmaison, Alisson Azzolini, Brian Vaughan, Edward Yang, Gregory Chanan, Guoqiang Jerry Chen, Hongyi Jia, Howard Huang, et al. Pytorch rpc: Distributed deep learning built on tensor-optimized remote procedure calls. Proceedings of Machine Learning and Systems, 5:219\u2013231, 2023.","journal-title":"Proceedings of Machine Learning and Systems"},{"key":"e_1_3_2_1_4_1","first-page":"231","volume-title":"2010 International Conference on High Performance Computing & Simulation","author":"Duato Jos\u00e9","unstructured":"Jos\u00e9 Duato, Antonio J Pena, Federico Silla, Rafael Mayo, and Enrique S Quintana-Ort\u00ed. rcuda: Reducing the number of gpu-based accelerators in high performance clusters. In 2010 International Conference on High Performance Computing & Simulation, pages 224\u2013231. IEEE, 2010."},{"key":"e_1_3_2_1_5_1","first-page":"750","volume-title":"2022 IEEE International Parallel and Distributed Processing Symposium (IPDPS)","author":"Fingler Henrique","unstructured":"Henrique Fingler, Zhiting Zhu, Esther Yoon, Zhipeng Jia, Emmett Witchel, and Christopher J Rossbach. Dgsf: Disaggregated gpus for serverless functions. In 2022 IEEE International Parallel and Distributed Processing Symposium (IPDPS), pages 739\u2013750. IEEE, 2022."},{"key":"e_1_3_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/3617995"},{"key":"e_1_3_2_1_7_1","volume-title":"The dawn of disaggregation and the coherence conundrum: A call for federated coherence","author":"Hong Jaewan","year":"2025","unstructured":"Jaewan Hong, Marcos K. Aguilera, Emmanuel Amaro, Vincent Liu, Aurojit Panda, and Ion Stoica. The dawn of disaggregation and the coherence conundrum: A call for federated coherence, 2025."},{"key":"e_1_3_2_1_8_1","first-page":"12","volume-title":"IEEE International Symposium on High-Performance Comp Architecture","author":"Lim Kevin","unstructured":"Kevin Lim, Yoshio Turner, Jose Renato Santos, Alvin AuYoung, Jichuan Chang, Parthasarathy Ranganathan, and Thomas F Wenisch. System-level implications of disaggregated memory. In IEEE International Symposium on High-Performance Comp Architecture, pages 1\u201312. IEEE, 2012."},{"key":"e_1_3_2_1_9_1","first-page":"122","volume-title":"Hwan Doh, and Arvind Krishnamurthy. Gimbal: enabling multi-tenant storage disaggregation on smartnic jbofs. In Proceedings of the 2021 ACM SIGCOMM 2021 Conference","author":"Min Jaehong","year":"2021","unstructured":"Jaehong Min, Ming Liu, Tapan Chugh, Chenxingyu Zhao, Andrew Wei, In Hwan Doh, and Arvind Krishnamurthy. Gimbal: enabling multi-tenant storage disaggregation on smartnic jbofs. In Proceedings of the 2021 ACM SIGCOMM 2021 Conference, pages 106\u2013122, 2021."},{"key":"e_1_3_2_1_10_1","first-page":"577","volume-title":"Proceedings of the 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI)","author":"Moritz Philipp","unstructured":"Philipp Moritz, Robert Nishihara, Stephanie Wang, Alexey Tumanov, Richard Liaw, Eric Liang, Melih Elibol, Zongheng Yang, William Paul, Michael I. Jordan, and Ion Stoica. Ray: A distributed framework for emerging ai applications. In Proceedings of the 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI), pages 561\u2013577. USENIX Association, 2018."},{"key":"e_1_3_2_1_11_1","volume-title":"Splitwise: Efficient generative llm inference using phase splitting. arXiv preprint arXiv:2311.18677","author":"Patel Pratyush","year":"2023","unstructured":"Pratyush Patel, Esha Choukse, Chaojie Zhang, Aashaka Shah, \u00cd\u00f1igo Goiri, Saeed Maleki, and Ricardo Bianchini. Splitwise: Efficient generative llm inference using phase splitting. arXiv preprint arXiv:2311.18677, 2023."},{"key":"e_1_3_2_1_12_1","volume-title":"September","author":"Markets Research","year":"2024","unstructured":"Research and Markets. Data center accelerators strategic business research report 2024: Global market to reach $395 billion by 2030 - ai-powered accelerators mark promising upheaval with energy-efficient datacenters, September 2024. Accessed: 2025-01-15."},{"key":"e_1_3_2_1_13_1","volume-title":"January","author":"Rogers Owen","year":"2024","unstructured":"Owen Rogers. Tens of thousands of gpus go under-utilized in the cloud, January 2024. Accessed: 2025-01-15."},{"key":"e_1_3_2_1_14_1","first-page":"87","volume-title":"13th USENIX Symposium on Operating Systems Design and Implementation (OSDI 18)","author":"Shan Yizhou","year":"2018","unstructured":"Yizhou Shan, Yutong Huang, Yilun Chen, and Yiying Zhang. {LegoOS}: A disseminated, distributed {OS} for hardware resource disaggregation. In 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI 18), pages 69\u201387, 2018."},{"key":"e_1_3_2_1_15_1","first-page":"11","volume-title":"2009 IEEE International Symposium on Parallel & Distributed Processing","author":"Shi Lin","year":"2009","unstructured":"Lin Shi, Hao Chen, Jianhua Sun, and Kenli Li. vcuda: Gpu accelerated high performance computing in virtual machines. 2009 IEEE International Symposium on Parallel & Distributed Processing, pages 1\u201311, 2009."},{"key":"e_1_3_2_1_16_1","volume-title":"Tensorpipe: A tensor-aware point-to-point communication library. https:\/\/github.com\/pytorch\/tensorpipe","author":"Team PyTorch","year":"2020","unstructured":"PyTorch Team. Tensorpipe: A tensor-aware point-to-point communication library. https:\/\/github.com\/pytorch\/tensorpipe, 2020. Commit reference: main branch, accessed October 2025."},{"key":"e_1_3_2_1_17_1","first-page":"1010","volume-title":"Proceedings of the 18th USENIX Symposium on Networked Systems Design and Implementation (NSDI)","author":"Wang Stephanie","unstructured":"Stephanie Wang, William Paul, Eric Liang, Robert Nishihara, Philipp Moritz, Ion Stoica, and Alexey Tumanov. Ownership: A distributed futures system for fine-grained tasks. In Proceedings of the 18th USENIX Symposium on Networked Systems Design and Implementation (NSDI), pages 995\u20131010. USENIX Association, 2021."},{"key":"e_1_3_2_1_18_1","first-page":"863","volume-title":"22nd USENIX Symposium on Networked Systems Design and Implementation (NSDI 25)","author":"Yang Lingyun","year":"2025","unstructured":"Lingyun Yang, Yongchen Wang, Yinghao Yu, Qizhen Weng, Jianbo Dong, Kan Liu, Chi Zhang, Yanyi Zi, Hao Li, Zechao Zhang, et al. {GPU-Disaggregated} serving for deep learning recommendation models at scale. In 22nd USENIX Symposium on Networked Systems Design and Implementation (NSDI 25), pages 847\u2013863, 2025."},{"key":"e_1_3_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/2934664"},{"key":"e_1_3_2_1_20_1","first-page":"348","volume-title":"Proceedings of the 18th USENIX Symposium on Operating Systems Design and Implementation, OSDI '24","author":"Zhong Yinmin","unstructured":"Yinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu, Yibo Zhu, Xuanzhe Liu, Xin Jin, and Hao Zhang. Distserve: Disaggregating prefill and decoding for goodput-optimized large language model serving. In Proceedings of the 18th USENIX Symposium on Operating Systems Design and Implementation, OSDI '24, pages 331\u2013348. USENIX Association, 2024."}],"event":{"name":"HotNets '25: 24th ACM Workshop on Hot Topics in Networks","location":"UMD Campus College Park MD USA","acronym":"HotNets '25","sponsor":["SIGCOMM ACM Special Interest Group on Data Communication"]},"container-title":["Proceedings of the 24th ACM Workshop on Hot Topics in Networks"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3772356.3772379","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,11,17]],"date-time":"2025-11-17T12:07:27Z","timestamp":1763381247000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3772356.3772379"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,11,17]]},"references-count":20,"alternative-id":["10.1145\/3772356.3772379","10.1145\/3772356"],"URL":"https:\/\/doi.org\/10.1145\/3772356.3772379","relation":{},"subject":[],"published":{"date-parts":[[2025,11,17]]},"assertion":[{"value":"2025-11-17","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}