{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,11,7]],"date-time":"2025-11-07T14:05:53Z","timestamp":1762524353332,"version":"build-2065373602"},"reference-count":74,"publisher":"Association for Computing Machinery (ACM)","issue":"1","funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62272291 and 62132014"],"award-info":[{"award-number":["62272291 and 62132014"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Comput. Syst."],"published-print":{"date-parts":[[2026,2,28]]},"abstract":"<jats:p>\n                    This article presents\n                    <jats:sc>UGache<\/jats:sc>\n                    , a unified multi-GPU cache system designed for embedding-based deep learning (EmbDL).\n                    <jats:sc>UGache<\/jats:sc>\n                    is primarily motivated by the unique characteristics of EmbDL applications, namely read-only and skewed embedding accesses with affinity and predictability.\n                    <jats:sc>UGache<\/jats:sc>\n                    introduces a novel factored extraction mechanism that avoids bandwidth congestion to fully exploit high-speed cross-GPU interconnects (e.g., NVLink and NVSwitch). Based on a\n                    <jats:italic toggle=\"yes\">hotness<\/jats:italic>\n                    metric,\n                    <jats:sc>UGache<\/jats:sc>\n                    also provides a near-optimal cache policy that balances local and remote access to minimize the extraction time for diverse GPU interconnect topologies. We have implemented\n                    <jats:sc>UGache<\/jats:sc>\n                    and integrated it into two representative frameworks, TensorFlow and PyTorch. Evaluation using two typical types of EmbDL applications, namely graph neural network (GNN) training and deep learning recommendation (DLR) inference, shows that\n                    <jats:sc>UGache<\/jats:sc>\n                    outperforms state-of-the-art replication and partition designs by an average of 1.93\u00d7 and 1.63\u00d7 (up to 5.25\u00d7 and 3.45\u00d7), respectively. Furthermore, we demonstrate the applicability of\n                    <jats:sc>UGache<\/jats:sc>\n                    \u2019s principle beyond embedding-based deep learning, with an example of text-to-image generation on an inference cluster.\n                  <\/jats:p>","DOI":"10.1145\/3767725","type":"journal-article","created":{"date-parts":[[2025,9,13]],"date-time":"2025-09-13T07:28:41Z","timestamp":1757748521000},"page":"1-32","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Unified and Near-optimal Multi-GPU Cache for Embedding-based Deep Learning"],"prefix":"10.1145","volume":"44","author":[{"ORCID":"https:\/\/orcid.org\/0009-0009-1887-2636","authenticated-orcid":false,"given":"Xiaoniu","family":"Song","sequence":"first","affiliation":[{"name":"Institute of Parallel and Distributed Systems, Shanghai Jiao Tong University","place":["Shanghai, China"]},{"name":"Engineering Research Center for Domain-specific Operating Systems","place":["Shanghai, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6115-8130","authenticated-orcid":false,"given":"Rong","family":"Chen","sequence":"additional","affiliation":[{"name":"Institute of Parallel and Distributed Systems, Shanghai Jiao Tong University","place":["Shanghai, China"]},{"name":"Engineering Research Center for Domain-specific Operating Systems","place":["Shanghai, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2113-2224","authenticated-orcid":false,"given":"Haitao","family":"Song","sequence":"additional","affiliation":[{"name":"Engineering Research Center for Domain-specific Operating Systems","place":["Shanghai, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-0462-1288","authenticated-orcid":false,"given":"Yiwen","family":"Zhang","sequence":"additional","affiliation":[{"name":"Institute of Parallel and Distributed Systems, Shanghai Jiao Tong University","place":["Shanghai, China"]},{"name":"Engineering Research Center for Domain-specific Operating Systems","place":["Shanghai, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9720-0361","authenticated-orcid":false,"given":"Haibo","family":"Chen","sequence":"additional","affiliation":[{"name":"Institute of Parallel and Distributed Systems, Shanghai Jiao Tong University","place":["Shanghai, China"]},{"name":"Engineering Research Center for Domain-specific Operating Systems","place":["Shanghai, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2025,11,7]]},"reference":[{"key":"e_1_3_2_2_2","unstructured":"2011. Peer-to-Peer and Unified Virtual Addressing. Retrieved from https:\/\/developer.download.nvidia.com\/CUDA\/training\/cuda_webinars_GPUDirect_uva.pdf. (2011)."},{"key":"e_1_3_2_3_2","unstructured":"2021. Open Graph Benchmark: Benchmark datasets data loaders and evaluators for graph machine learning. Retrieved from https:\/\/ogb.stanford.edu\/. (2021)."},{"key":"e_1_3_2_4_2","unstructured":"2021. Open Graph Benchmark: The MAG240M dataset. Retrieved from https:\/\/ogb.stanford.edu\/docs\/lsc\/mag240m\/. (2021)."},{"key":"e_1_3_2_5_2","unstructured":"2021. Open Graph Benchmark: The ogbn-papers100M dataset. Retrieved from https:\/\/ogb.stanford.edu\/docs\/nodeprop\/#ogbn-papers100M. (2021)."},{"key":"e_1_3_2_6_2","unstructured":"2022. Distributed Embeddings \\(\\cdot\\) NVIDIA-Merlin. (Nov.2022). Retrieved from https:\/\/github.com\/NVIDIA-Merlin\/distributed-embeddings"},{"key":"e_1_3_2_7_2","unstructured":"2022. Download Criteo 1TB Click Logs dataset - Criteo AI Lab. Retrieved from https:\/\/ailab.criteo.com\/download-criteo-1tb-click-logs-dataset\/. (2022)."},{"key":"e_1_3_2_8_2","unstructured":"2022. neuralworm\/stable-diffusion-discord-prompts \\(\\cdot\\) Datasets at Hugging Face. (2022). Retrieved from https:\/\/huggingface.co\/datasets\/neuralworm\/stable-diffusion-discord-prompts"},{"key":"e_1_3_2_9_2","unstructured":"2022. quiver-team\/torch-quiver. (Nov.2022). Retrieved from https:\/\/github.com\/quiver-team\/torch-quiver"},{"key":"e_1_3_2_10_2","unstructured":"2022. Sparse Operation Kit NVIDIA-Merlin. (2022). Retrieved from https:\/\/github.com\/NVIDIA-Merlin\/HugeCTR"},{"key":"e_1_3_2_11_2","unstructured":"2022. zilliztech\/GPTCache. (2022).Retrieved from https:\/\/github.com\/zilliztech\/GPTCache"},{"key":"e_1_3_2_12_2","unstructured":"2023. Gurobi Optimizer. Retrieved from https:\/\/www.gurobi.com\/. (2023)."},{"key":"e_1_3_2_13_2","unstructured":"2023. Multi-Process Service. (2023). Retrieved from http:\/\/docs.nvidia.com\/deploy\/mps\/index.html"},{"key":"e_1_3_2_14_2","unstructured":"2023. Nvidia Collective Communication Library (NCCL). Retrieved from https:\/\/developer.nvidia.com\/nccl. (2023)."},{"key":"e_1_3_2_15_2","unstructured":"2023. NVIDIA Nsight Systems. Retrieved from https:\/\/developer.nvidia.com\/nsight-systems. (2023)."},{"key":"e_1_3_2_16_2","unstructured":"2023. Open Graph Benchmark Leaderboards. Retrieved from https:\/\/ogb.stanford.edu\/docs\/leader_overview\/. (2023)."},{"key":"e_1_3_2_17_2","unstructured":"2023. stable-diffusion-v1-5\/stable-diffusion-v1-5 \\(\\cdot\\) Hugging Face. (2023). Retrieved from https:\/\/huggingface.co\/stable-diffusion-v1-5\/stable-diffusion-v1-5"},{"key":"e_1_3_2_18_2","unstructured":"Martin Abadi Paul Barham Jianmin Chen Zhifeng Chen Andy Davis Jeffrey Dean Matthieu Devin Sanjay Ghemawat Geoffrey Irving Michael Isard et\u00a0al. 2016. TensorFlow: A System for Large-Scale Machine Learning. 265\u2013283. Retrieved from https:\/\/www.usenix.org\/conference\/osdi16\/technical-sessions\/presentation\/abadi"},{"issue":"5461","key":"e_1_3_2_19_2","doi-asserted-by":"crossref","first-page":"2115","DOI":"10.1126\/science.287.5461.2115a","article-title":"Power-law distribution of the world wide web","volume":"287","author":"Adamic Lada A.","year":"2000","unstructured":"Lada A. Adamic and Bernardo A. Huberman. 2000. Power-law distribution of the world wide web. Science 287, 5461 (2000), 2115\u20132115.","journal-title":"Science"},{"issue":"1","key":"e_1_3_2_20_2","doi-asserted-by":"crossref","first-page":"127","DOI":"10.14778\/3485450.3485462","article-title":"Accelerating recommendation system training by leveraging popular choices","volume":"15","author":"Adnan Muhammad","year":"2021","unstructured":"Muhammad Adnan, Yassaman Ebrahimzadeh Maboud, Divya Mahajan, and Prashant J. Nair. 2021. Accelerating recommendation system training by leveraging popular choices. Proceedings of the VLDB Endowment 15, 1 (Sept.2021), 127\u2013140.","journal-title":"Proceedings of the VLDB Endowment"},{"key":"e_1_3_2_21_2","doi-asserted-by":"crossref","first-page":"348","DOI":"10.1145\/3600006.3613142","volume-title":"Proceedings of the 29th Symposium on Operating Systems Principles (SOSP\u201923)","author":"Agarwal Saurabh","year":"2023","unstructured":"Saurabh Agarwal, Chengpo Yan, Ziyi Zhang, and Shivaram Venkataraman. 2023. Bagpipe: Accelerating deep recommendation model training. In Proceedings of the 29th Symposium on Operating Systems Principles (SOSP\u201923). Association for Computing Machinery, New York, NY, USA, 348\u2013363. DOI:10.1145\/3600006.3613142"},{"key":"e_1_3_2_22_2","first-page":"129","volume-title":"Proceedings of the 2nd Conference on Symposium on Networked Systems Design & Implementation-Volume 2","author":"Annapureddy Siddhartha","year":"2005","unstructured":"Siddhartha Annapureddy, Michael J. Freedman, and David Mazieres. 2005. Shark: Scaling file servers via cooperative caching. In Proceedings of the 2nd Conference on Symposium on Networked Systems Design & Implementation-Volume 2. 129\u2013142."},{"key":"e_1_3_2_23_2","first-page":"263","volume-title":"Proceedings of the 15th ACM Conference on Recommender Systems (RecSys\u201921)","author":"Balasubramanian Keshav","year":"2021","unstructured":"Keshav Balasubramanian, Abdulla Alshabanah, Joshua D Choe, and Murali Annavaram. 2021. cDLRM: Look ahead caching for scalable training of recommendation models. In Proceedings of the 15th ACM Conference on Recommender Systems (RecSys\u201921). Association for Computing Machinery, New York, NY, USA, 263\u2013272."},{"key":"e_1_3_2_24_2","unstructured":"James Briggs and Nima Boscarino. 2023. Making Stable Diffusion Faster with Intelligent Caching | Pinecone. (2023). Retrieved from https:\/\/www.pinecone.io\/learn\/faster-stable-diffusion\/"},{"issue":"2","key":"e_1_3_2_25_2","doi-asserted-by":"crossref","first-page":"264","DOI":"10.1145\/1150019.1136509","article-title":"Cooperative caching for chip multiprocessors","volume":"34","author":"Chang Jichuan","year":"2006","unstructured":"Jichuan Chang and Gurindar S. Sohi. 2006. Cooperative caching for chip multiprocessors. ACM SIGARCH Computer Architecture News 34, 2 (2006), 264\u2013276.","journal-title":"ACM SIGARCH Computer Architecture News"},{"key":"e_1_3_2_26_2","volume-title":"35th Conference on Neural Information Processing Systems (NeurIPS 2021)","author":"Chen Qi","year":"2021","unstructured":"Qi Chen, Bing Zhao, Haidong Wang, Mingqin Li, Chuanjie Liu, Zengzhong Li, Mao Yang, and Jingdong Wang. 2021. SPANN: Highly-efficient Billion-scale Approximate Nearest Neighbor Search. In 35th Conference on Neural Information Processing Systems (NeurIPS 2021)."},{"key":"e_1_3_2_27_2","first-page":"1","volume-title":"Proceedings of the Tenth European Conference on Computer Systems","author":"Chen Rong","year":"2015","unstructured":"Rong Chen, Jiaxin Shi, Yanzhe Chen, and Haibo Chen. 2015. PowerLyra: Differentiated graph computation and partitioning on skewed graphs. In Proceedings of the Tenth European Conference on Computer Systems. 1\u201315."},{"key":"e_1_3_2_28_2","volume-title":"First Symposium on Operating Systems Design and Implementation (OSDI 94)","author":"Dahlin Michael D.","year":"1994","unstructured":"Michael D. Dahlin, Randolph Y. Wang, Thomas E. Anderson, and David A. Patterson. 1994. Cooperative Caching: Using Remote Client Memory to Improve File System Performance. In First Symposium on Operating Systems Design and Implementation (OSDI 94). USENIX Association, Monterey, CA. Retrieved from https:\/\/www.usenix.org\/conference\/osdi-94\/cooperative-caching-using-remote-client-memory-improve-file-system-performance"},{"key":"e_1_3_2_29_2","doi-asserted-by":"crossref","first-page":"1754","DOI":"10.1109\/ICDE53745.2022.00177","volume-title":"2022 IEEE 38th International Conference on Data Engineering (ICDE)","author":"Dong Sicong","year":"2022","unstructured":"Sicong Dong, Xupeng Miao, Pengkai Liu, Xin Wang, Bin Cui, and Jianxin Li. 2022. HET-KG: Communication-efficient knowledge graph embedding training via hotness-aware cache. In 2022 IEEE 38th International Conference on Data Engineering (ICDE). 1754\u20131766. DOI:10.1109\/ICDE53745.2022.00177"},{"key":"e_1_3_2_30_2","doi-asserted-by":"crossref","first-page":"2","DOI":"10.1109\/HPCA.2007.346180","volume-title":"2007 IEEE 13th International Symposium on High Performance Computer Architecture","author":"Dybdahl Haakon","year":"2007","unstructured":"Haakon Dybdahl and Per Stenstrom. 2007. An adaptive shared\/private nuca cache partitioning scheme for chip multiprocessors. In 2007 IEEE 13th International Symposium on High Performance Computer Architecture. IEEE, 2\u201312."},{"issue":"4","key":"e_1_3_2_31_2","doi-asserted-by":"crossref","first-page":"254","DOI":"10.1145\/285243.285287","article-title":"Summary cache: A scalable wide-area web cache sharing protocol","volume":"28","author":"Fan Li","year":"1998","unstructured":"Li Fan, Pei Cao, Jussara Almeida, and Andrei Z Broder. 1998. Summary cache: A scalable wide-area web cache sharing protocol. ACM SIGCOMM Computer Communication Review 28, 4 (1998), 254\u2013265.","journal-title":"ACM SIGCOMM Computer Communication Review"},{"key":"e_1_3_2_32_2","doi-asserted-by":"crossref","first-page":"201","DOI":"10.1145\/224056.224072","volume-title":"Proceedings of the Fifteenth ACM Symposium on Operating Systems Principles (SOSP\u201995)","author":"Feeley M. J.","year":"1995","unstructured":"M. J. Feeley, W. E. Morgan, E. P. Pighin, A. R. Karlin, H. M. Levy, and C. A. Thekkath. 1995. Implementing global memory management in a workstation cluster. In Proceedings of the Fifteenth ACM Symposium on Operating Systems Principles (SOSP\u201995). Association for Computing Machinery, New York, NY, USA, 201\u2013212. DOI:10.1145\/224056.224072"},{"key":"e_1_3_2_33_2","volume-title":"MPI: A Message-Passing Interface Standard","author":"Forum Message P.","year":"1994","unstructured":"Message P. Forum. 1994. MPI: A Message-Passing Interface Standard. Technical Report. University of Tennessee, USA."},{"key":"e_1_3_2_34_2","volume-title":"Proceedings of the 15th USENIX Conference on Operating Systems Design and Implementation (OSDI\u201921)","author":"Gandhi Swapnil","year":"2021","unstructured":"Swapnil Gandhi and Anand Padmanabha Iyer. 2021. P3: Distributed deep graph learning at scale. In Proceedings of the 15th USENIX Conference on Operating Systems Design and Implementation (OSDI\u201921)."},{"key":"e_1_3_2_35_2","first-page":"17","volume-title":"10th USENIX Symposium on Operating Systems Design and Implementation (OSDI\u201912)","author":"Gonzalez Joseph E.","year":"2012","unstructured":"Joseph E. Gonzalez, Yucheng Low, Haijie Gu, Danny Bickson, and Carlos Guestrin. 2012. Powergraph: Distributed graph-parallel computation on natural graphs. In 10th USENIX Symposium on Operating Systems Design and Implementation (OSDI\u201912). 17\u201330."},{"key":"e_1_3_2_36_2","first-page":"1269","volume-title":"Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR\u201921)","author":"Guo Huifeng","year":"2021","unstructured":"Huifeng Guo, Wei Guo, Yong Gao, Ruiming Tang, Xiuqiang He, and Wenzhi Liu. 2021. ScaleFreeCTR: MixCache-based distributed training system for CTR models with huge embedding table. In Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR\u201921). Association for Computing Machinery, New York, NY, USA, 1269\u20131278."},{"key":"e_1_3_2_37_2","first-page":"1025","volume-title":"Proceedings of the 31st International Conference on Neural Information Processing Systems (NeurIPS\u201917)","author":"Hamilton William L.","year":"2017","unstructured":"William L. Hamilton, Rex Ying, and Jure Leskovec. 2017. Inductive representation learning on large graphs. In Proceedings of the 31st International Conference on Neural Information Processing Systems (NeurIPS\u201917). 1025\u20131035."},{"key":"e_1_3_2_38_2","first-page":"1","volume-title":"The International Symposium on Memory Systems (MEMSYS 2021)","author":"Ibrahim Mohamed Assem","year":"2022","unstructured":"Mohamed Assem Ibrahim, Onur Kayiran, and Shaizeen Aga. 2022. Efficient cache utilization via model-aware data placement for recommendation models. In The International Symposium on Memory Systems (MEMSYS 2021). Association for Computing Machinery, New York, NY, USA, 1\u201311."},{"key":"e_1_3_2_39_2","doi-asserted-by":"crossref","first-page":"213","DOI":"10.1145\/571825.571861","volume-title":"Proceedings of the Twenty-First Annual Symposium on Principles of Distributed Computing","author":"Iyer Sitaram","year":"2002","unstructured":"Sitaram Iyer, Antony Rowstron, and Peter Druschel. 2002. Squirrel: A decentralized peer-to-peer web cache. In Proceedings of the Twenty-First Annual Symposium on Principles of Distributed Computing. 213\u2013222."},{"key":"e_1_3_2_40_2","doi-asserted-by":"crossref","first-page":"327","DOI":"10.1109\/BigData47090.2019.9006396","volume-title":"2019 IEEE International Conference on Big Data (Big Data)","author":"Kaynar Emine Ugur","year":"2019","unstructured":"Emine Ugur Kaynar, Mania Abdi, Mohammad Hossein Hajkazemi, Ata Turk, Raja R. Sambasivan, David Cohen, Larry Rudolph, Peter Desnoyers, and Orran Krieger. 2019. D3N: A multi-layer cache for the rest of us. In 2019 IEEE International Conference on Big Data (Big Data). IEEE, 327\u2013338."},{"key":"e_1_3_2_41_2","unstructured":"Daya Khudia Jianyu Huang Protonu Basu Summer Deng Haixin Liu Jongsoo Park and Mikhail Smelyanskiy. 2021. FBGEMM: Enabling high-performance low-precision deep learning inference. arXiv:2101.05615. Retrieved from https:\/\/arxiv.org\/abs\/\/2101.05615"},{"key":"e_1_3_2_42_2","volume-title":"Proceedings of the 5th International Conference on Learning Representations (ICLR\u201917)","author":"Kipf Thomas N.","year":"2017","unstructured":"Thomas N. Kipf and Max Welling. 2017. Semi-supervised classification with graph convolutional networks. In Proceedings of the 5th International Conference on Learning Representations (ICLR\u201917)."},{"key":"e_1_3_2_43_2","doi-asserted-by":"crossref","first-page":"1343","DOI":"10.1145\/2487788.2488173","volume-title":"Proceedings of the 22nd International Conference on World Wide Web (WWW\u201913 Companion)","author":"Kunegis J\u00e9r\u00f4me","year":"2013","unstructured":"J\u00e9r\u00f4me Kunegis. 2013. KONECT: The Koblenz network collection. In Proceedings of the 22nd International Conference on World Wide Web (WWW\u201913 Companion). Association for Computing Machinery, New York, NY, USA, 1343\u20131350."},{"key":"e_1_3_2_44_2","first-page":"281","volume-title":"Proceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2 (ASPLOS 2023)","author":"Kurniawan Daniar H.","year":"2023","unstructured":"Daniar H. Kurniawan, Ruipu Wang, Kahfi S. Zulkifli, Fandi A. Wiranata, John Bent, Ymir Vigfusson, and Haryadi S. Gunawi. 2023. EVStore: Storage and caching capabilities for scaling embedding tables in deep recommendation systems. In Proceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2 (ASPLOS 2023). Association for Computing Machinery, New York, NY, USA, 281\u2013294."},{"key":"e_1_3_2_45_2","first-page":"817","volume-title":"17th USENIX Symposium on Operating Systems Design and Implementation (OSDI 23)","author":"Lai Fan","year":"2023","unstructured":"Fan Lai, Wei Zhang, Rui Liu, William Tsai, Xiaohan Wei, Yuxi Hu, Sabin Devkota, Jianyu Huang, Jongsoo Park, Xing Liu, et\u00a0al. 2023. AdaEmbed: Adaptive embedding for large-scale recommendation models. In 17th USENIX Symposium on Operating Systems Design and Implementation (OSDI 23). 817\u2013831."},{"key":"e_1_3_2_46_2","doi-asserted-by":"publisher","DOI":"10.5334\/tismir.39"},{"key":"e_1_3_2_47_2","first-page":"401","volume-title":"Proceedings of the 11th ACM Symposium on Cloud Computing","author":"Lin Zhiqi","year":"2020","unstructured":"Zhiqi Lin, Cheng Li, Youshan Miao, Yunxin Liu, and Yinlong Xu. 2020. PaGraph: Scaling GNN training on large graphs via computation-aware caching. In Proceedings of the 11th ACM Symposium on Cloud Computing. ACM, Virtual Event USA, 401\u2013415."},{"key":"e_1_3_2_48_2","first-page":"401","volume-title":"Proceedings of the 11th ACM Symposium on Cloud Computing (SoCC\u201920)","author":"Lin Zhiqi","year":"2020","unstructured":"Zhiqi Lin, Cheng Li, Youshan Miao, Yunxin Liu, and Yinlong Xu. 2020. Pagraph: Scaling GNN training on large graphs via computation-aware caching. In Proceedings of the 11th ACM Symposium on Cloud Computing (SoCC\u201920). 401\u2013415."},{"key":"e_1_3_2_49_2","doi-asserted-by":"publisher","DOI":"10.1016\/S0166-5316(00)00032-8"},{"key":"e_1_3_2_50_2","doi-asserted-by":"crossref","first-page":"470","DOI":"10.1145\/3514221.3517902","volume-title":"Proceedings of the 2022 International Conference on Management of Data (SIGMOD\u201922)","author":"Miao Xupeng","year":"2022","unstructured":"Xupeng Miao, Yining Shi, Hailin Zhang, Xin Zhang, Xiaonan Nie, Zhi Yang, and Bin Cui. 2022. HET-GMP: A graph-based system approach to scaling large embedding model training. In Proceedings of the 2022 International Conference on Management of Data (SIGMOD\u201922). Association for Computing Machinery, New York, NY, USA, 470\u2013480."},{"issue":"2","key":"e_1_3_2_51_2","doi-asserted-by":"crossref","first-page":"312","DOI":"10.14778\/3489496.3489511","article-title":"HET: Scaling out huge embedding model training via cache-enabled distributed framework","volume":"15","author":"Miao Xupeng","year":"2021","unstructured":"Xupeng Miao, Hailin Zhang, Yining Shi, Xiaonan Nie, Zhi Yang, Yangyu Tao, and Bin Cui. 2021. HET: Scaling out huge embedding model training via cache-enabled distributed framework. Proceedings of the VLDB Endowment 15, 2 (Oct.2021), 312\u2013320.","journal-title":"Proceedings of the VLDB Endowment"},{"issue":"11","key":"e_1_3_2_52_2","doi-asserted-by":"crossref","first-page":"2087","DOI":"10.14778\/3476249.3476264","article-title":"Large graph convolutional network training with GPU-oriented data communication architecture","volume":"14","author":"Min Seung Won","year":"2021","unstructured":"Seung Won Min, Kun Wu, Sitao Huang, Mert Hidayeto\u011flu, Jinjun Xiong, Eiman Ebrahimi, Deming Chen, and Wen-mei Hwu. 2021. Large graph convolutional network training with GPU-oriented data communication architecture. Proceedings of the VLDB Endowment 14, 11 (Oct.2021), 2087\u20132100.","journal-title":"Proceedings of the VLDB Endowment"},{"key":"e_1_3_2_53_2","doi-asserted-by":"crossref","first-page":"993","DOI":"10.1145\/3470496.3533727","volume-title":"Proceedings of the 49th Annual International Symposium on Computer Architecture (ISCA\u201922)","author":"Mudigere Dheevatsa","year":"2022","unstructured":"Dheevatsa Mudigere, Yuchen Hao, Jianyu Huang, Zhihao Jia, Andrew Tulloch, Srinivas Sridharan, Xing Liu, Mustafa Ozdal, Jade Nie, Jongsoo Park, et\u00a0al. , 2022. Software-hardware co-design for fast and scalable training of deep learning recommendation models. In Proceedings of the 49th Annual International Symposium on Computer Architecture (ISCA\u201922). Association for Computing Machinery, New York, NY, USA, 993\u20131011."},{"key":"e_1_3_2_54_2","unstructured":"Maxim Naumov Dheevatsa Mudigere Hao-Jun Michael Shi Jianyu Huang Narayanan Sundaraman Jongsoo Park Xiaodong Wang Udit Gupta Carole-Jean Wu Alisson G. Azzolini et\u00a0al. 2019. Deep Learning Recommendation Model for Personalization and Recommendation Systems. (2019)."},{"key":"e_1_3_2_55_2","unstructured":"Maxim Naumov Dheevatsa Mudigere Hao-Jun Michael Shi Jianyu Huang Narayanan Sundaraman Jongsoo Park Xiaodong Wang Udit Gupta Carole-Jean Wu Alisson G. Azzolini et\u00a0al. 2019. Deep Learning Recommendation Model for Personalization and Recommendation Systems. (May2019).arXiv:1906.00091. Retrieved from https:\/\/arxiv.org\/abs\/\/1906.00091"},{"key":"e_1_3_2_56_2","unstructured":"Adam Paszke Sam Gross Francisco Massa Adam Lerer James Bradbury Gregory Chanan Trevor Killeen Zeming Lin Natalia Gimelshein Luca Antiga et\u00a0al. 2019. Pytorch: An imperative style high-performance deep learning library. In Proceedings of the 33rd International Conference on Neural Information Processing Systems. Curran Associates Inc. Red Hook NY USA Article 721 8026\u20138037."},{"key":"e_1_3_2_57_2","first-page":"10684","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Rombach Robin","year":"2022","unstructured":"Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj\u00f6rn Ommer. 2022. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 10684\u201310695."},{"issue":"4","key":"e_1_3_2_58_2","doi-asserted-by":"crossref","first-page":"387","DOI":"10.1145\/362670.362675","article-title":"Hint-based cooperative caching","volume":"18","author":"Sarkar Prasenjit","year":"2000","unstructured":"Prasenjit Sarkar and John H. Hartman. 2000. Hint-based cooperative caching. ACM Transactions on Computer Systems (TOCS) 18, 4 (2000), 387\u2013419.","journal-title":"ACM Transactions on Computer Systems (TOCS)"},{"key":"e_1_3_2_59_2","first-page":"344","volume-title":"Proceedings of the 27th ACM International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS\u201922)","author":"Sethi Geet","year":"2022","unstructured":"Geet Sethi, Bilge Acun, Niket Agarwal, Christos Kozyrakis, Caroline Trippel, and Carole-Jean Wu. 2022. RecShard: Statistical feature-based memory optimization for industry-scale neural recommendation. In Proceedings of the 27th ACM International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS\u201922). Association for Computing Machinery, New York, NY, USA, 344\u2013358."},{"key":"e_1_3_2_60_2","first-page":"1","volume-title":"Proceedings of the 36th ACM International Conference on Supercomputing (ICS\u201922)","author":"Song Shihui","year":"2022","unstructured":"Shihui Song and Peng Jiang. 2022. Rethinking graph data placement for graph neural network training on multiple GPUs. In Proceedings of the 36th ACM International Conference on Supercomputing (ICS\u201922). Association for Computing Machinery, New York, NY, USA, 1\u201310."},{"key":"e_1_3_2_61_2","doi-asserted-by":"crossref","first-page":"627","DOI":"10.1145\/3600006.3613169","volume-title":"Proceedings of the 29th Symposium on Operating Systems Principles (SOSP\u201923)","author":"Song Xiaoniu","year":"2023","unstructured":"Xiaoniu Song, Yiwen Zhang, Rong Chen, and Haibo Chen. 2023. UGACHE: A Unified GPU cache for embedding-based deep learning. In Proceedings of the 29th Symposium on Operating Systems Principles (SOSP\u201923). Association for Computing Machinery, New York, NY, USA, 627\u2013641. DOI:10.1145\/3600006.3613169"},{"key":"e_1_3_2_62_2","first-page":"165","volume-title":"2023 USENIX Annual Technical Conference (USENIX ATC 23)","author":"Sun Jie","year":"2023","unstructured":"Jie Sun, Li Su, Zuocheng Shi, Wenting Shen, Zeke Wang, Lei Wang, Jie Zhang, Yong Li, Wenyuan Yu, Jingren Zhou, et\u00a0al. 2023. Legion: Automatically pushing the envelope of Multi-GPU system for billion-scale GNN training. In 2023 USENIX Annual Technical Conference (USENIX ATC 23). USENIX Association, Boston, MA, 165\u2013179. Retrieved from https:\/\/www.usenix.org\/conference\/atc23\/presentation\/sun"},{"issue":"5","key":"e_1_3_2_63_2","doi-asserted-by":"crossref","first-page":"1105","DOI":"10.14778\/3641204.3641219","article-title":"XGNN: Boosting Multi-GPU GNN training via global GNN memory store","volume":"17","author":"Tang Dahai","year":"2024","unstructured":"Dahai Tang, Jiali Wang, Rong Chen, Lei Wang, Wenyuan Yu, Jingren Zhou, and Kenli Li. 2024. XGNN: Boosting Multi-GPU GNN training via global GNN memory store. Proc. VLDB Endow. 17, 5 (2024), 1105\u20131118. Retrieved from https:\/\/www.vldb.org\/pvldb\/vol17\/p1105-chen.pdf","journal-title":"Proc. VLDB Endow."},{"key":"e_1_3_2_64_2","unstructured":"Minjie Wang Da Zheng Zihao Ye Quan Gan Mufei Li Xiang Song Jinjing Zhou Chao Ma Lingfan Yu Yu Gai et\u00a0al. 2019. Deep graph library: A graph-centric highly-performant package for graph neural networks. arXiv:1909.01315. Retrieved from https:\/\/arxiv.org\/abs\/\/1909.01315"},{"key":"e_1_3_2_65_2","first-page":"1","volume-title":"Proceedings of the ADKDD\u201917 (ADKDD\u201917)","author":"Wang Ruoxi","year":"2017","unstructured":"Ruoxi Wang, Bin Fu, Gang Fu, and Mingliang Wang. 2017. Deep & cross network for Ad click predictions. In Proceedings of the ADKDD\u201917 (ADKDD\u201917). Association for Computing Machinery, New York, NY, USA, 1\u20137."},{"key":"e_1_3_2_66_2","doi-asserted-by":"crossref","first-page":"534","DOI":"10.1145\/3523227.3547405","volume-title":"Proceedings of the 16th ACM Conference on Recommender Systems (RecSys\u201922)","author":"Wang Zehuan","year":"2022","unstructured":"Zehuan Wang, Yingcan Wei, Minseok Lee, Matthias Langer, Fan Yu, Jie Liu, Shijie Liu, Daniel G. Abel, Xu Guo, Jianbing Dong, et\u00a0al. 2022. Merlin HugeCTR: GPU-accelerated recommender system training and inference. In Proceedings of the 16th ACM Conference on Recommender Systems (RecSys\u201922). Association for Computing Machinery, New York, NY, USA, 534\u2013537."},{"key":"e_1_3_2_67_2","first-page":"408","volume-title":"Proceedings of the 16th ACM Conference on Recommender Systems (RecSys\u201922)","author":"Wei Yingcan","year":"2022","unstructured":"Yingcan Wei, Matthias Langer, Fan Yu, Minseok Lee, Jie Liu, Ji Shi, and Zehuan Wang. 2022. A GPU-specialized inference parameter server for large-scale deep recommendation models. In Proceedings of the 16th ACM Conference on Recommender Systems (RecSys\u201922). Association for Computing Machinery, New York, NY, USA, 408\u2013419."},{"key":"e_1_3_2_68_2","first-page":"402","volume-title":"Proceedings of the Seventeenth European Conference on Computer Systems","author":"Xie Minhui","year":"2022","unstructured":"Minhui Xie, Youyou Lu, Jiazhen Lin, Qing Wang, Jian Gao, Kai Ren, and Jiwu Shu. 2022. Fleche: An efficient GPU embedding cache for personalized recommendations. In Proceedings of the Seventeenth European Conference on Computer Systems. ACM, Rennes France, 402\u2013416."},{"key":"e_1_3_2_69_2","doi-asserted-by":"publisher","DOI":"10.1109\/MSST.2013.6558446"},{"key":"e_1_3_2_70_2","doi-asserted-by":"crossref","unstructured":"Dongxu Yang Junhong Liu Jiaxing Qi and Junjie Lai. 2022. WholeGraph: A fast graph neural network training framework with Multi-GPU distributed shared memory architecture. In Proceedings of the International Conference on High Performance Computing Networking Storage and Analysis (SC\u201922). IEEE Press Article 54 1\u201314.","DOI":"10.1109\/SC41404.2022.00059"},{"key":"e_1_3_2_71_2","doi-asserted-by":"crossref","first-page":"417","DOI":"10.1145\/3492321.3519557","volume-title":"Proceedings of the Seventeenth European Conference on Computer Systems (EuroSys\u201922)","author":"Yang Jianbang","year":"2022","unstructured":"Jianbang Yang, Dahai Tang, Xiaoniu Song, Lei Wang, Qiang Yin, Rong Chen, Wenyuan Yu, and Jingren Zhou. 2022. GNNLab: A factored system for sample-based GNN training over GPUs. In Proceedings of the Seventeenth European Conference on Computer Systems (EuroSys\u201922). Association for Computing Machinery, New York, NY, USA, 417\u2013434."},{"key":"e_1_3_2_72_2","first-page":"4461","volume-title":"Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD\u201922)","author":"Zha Daochen","year":"2022","unstructured":"Daochen Zha, Louis Feng, Bhargav Bhushanam, Dhruv Choudhary, Jade Nie, Yuandong Tian, Jay Chae, Yinbin Ma, Arun Kejariwal, and Xia Hu. 2022. AutoShard: Automated embedding table sharding for recommender systems. In Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD\u201922). Association for Computing Machinery, New York, NY, USA, 4461\u20134471."},{"key":"e_1_3_2_73_2","doi-asserted-by":"crossref","first-page":"3453","DOI":"10.1109\/ICDE53745.2022.00324","volume-title":"2022 IEEE 38th International Conference on Data Engineering (ICDE)","author":"Zhang Yuanxing","year":"2022","unstructured":"Yuanxing Zhang, Langshi Chen, Siran Yang, Man Yuan, Huimin Yi, Jie Zhang, Jiamang Wang, Jianbo Dong, Yunlong Xu, Yue Song, et\u00a0al. 2022. PICASSO: Unleashing the potential of GPU-centric training for wide-and-deep recommender systems. In 2022 IEEE 38th International Conference on Data Engineering (ICDE). 3453\u20133466."},{"key":"e_1_3_2_74_2","doi-asserted-by":"crossref","unstructured":"Carolina Zheng Minhui Huang Dmitrii Pedchenko Kaushik Rangadurai Siyu Wang Gaby Nahum Jie Lei Yang Yang Tao Liu Zutian Luo et\u00a0al. 2025. Enhancing Embedding Representation Stability in Recommendation Systems with Semantic ID. (2025). arXiv:2504.02137. Retrieved from https:\/\/arxiv.org\/abs\/\/2504.02137","DOI":"10.1145\/3705328.3748123"},{"key":"e_1_3_2_75_2","first-page":"36","volume-title":"Proceedings of the 10th IEEE\/ACM Workshop on Irregular Applications: Architectures and Algorithms (IA3\u201920)","author":"Zheng Da","year":"2020","unstructured":"Da Zheng, Chao Ma, Minjie Wang, Jinjing Zhou, Qidong Su, Xiang Song, Quan Gan, Zheng Zhang, and George Karypis. 2020. DistDGL: Distributed graph neural network training for billion-scale graphs. In Proceedings of the 10th IEEE\/ACM Workshop on Irregular Applications: Architectures and Algorithms (IA3\u201920). 36\u201344."}],"container-title":["ACM Transactions on Computer Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3767725","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,11,7]],"date-time":"2025-11-07T14:00:57Z","timestamp":1762524057000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3767725"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,11,7]]},"references-count":74,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2026,2,28]]}},"alternative-id":["10.1145\/3767725"],"URL":"https:\/\/doi.org\/10.1145\/3767725","relation":{},"ISSN":["0734-2071","1557-7333"],"issn-type":[{"type":"print","value":"0734-2071"},{"type":"electronic","value":"1557-7333"}],"subject":[],"published":{"date-parts":[[2025,11,7]]},"assertion":[{"value":"2024-04-21","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-09-05","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-11-07","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}