{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,11]],"date-time":"2026-07-11T15:47:18Z","timestamp":1783784838248,"version":"3.55.0"},"reference-count":90,"publisher":"Association for Computing Machinery (ACM)","issue":"6","content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["Proc. VLDB Endow."],"published-print":{"date-parts":[[2025,2]]},"abstract":"<jats:p>Deep learning (DL) system research is often impeded by the limited availability and expensive costs of GPUs. In this paper, we introduce GPEmu, a GPU emulator for faster and cheaper prototyping and evaluation of deep learning system research without using real GPUs. GPEmu comes with four novel features: time emulation, memory emulation, distributed system support, and sharing support. We support over 30 DL models and 6 GPU models, the largest scale to date. We demonstrate the power of GPEmu by successfully reproducing the main results of nine recent publications and easily prototyping three new micro-optimizations.<\/jats:p>","DOI":"10.14778\/3725688.3725716","type":"journal-article","created":{"date-parts":[[2025,8,29]],"date-time":"2025-08-29T14:19:21Z","timestamp":1756477161000},"page":"1919-1932","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["GPEmu: A GPU Emulator for Faster and Cheaper Prototyping and Evaluation of Deep Learning System Research"],"prefix":"10.14778","volume":"18","author":[{"given":"Meng","family":"Wang","sequence":"first","affiliation":[{"name":"University of Chicago"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Gus","family":"Waldspurger","sequence":"additional","affiliation":[{"name":"University of Chicago"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Naufal","family":"Ananda","sequence":"additional","affiliation":[{"name":"Telkom University"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yuyang","family":"Huang","sequence":"additional","affiliation":[{"name":"University of Chicago"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Kemas","family":"Wiharja","sequence":"additional","affiliation":[{"name":"Telkom University"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"John","family":"Bent","sequence":"additional","affiliation":[{"name":"LANL"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Swaminathan","family":"Sundararaman","sequence":"additional","affiliation":[{"name":"IBM Research"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Vijay","family":"Chidambaram","sequence":"additional","affiliation":[{"name":"UT Austin"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Haryadi S.","family":"Gunawi","sequence":"additional","affiliation":[{"name":"University of Chicago"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,8,29]]},"reference":[{"key":"e_1_2_1_1_1","unstructured":"Accessed in November 2024. 8.4. Configuration Tools Red Hat Enterprise Linux 7 | Red Hat Customer Portal. https:\/\/access.redhat.com\/documentation\/en-us\/red_hat_enterprise_linux\/7\/html\/performance_tuning_guide."},{"key":"e_1_2_1_2_1","unstructured":"Accessed in November 2024. Allox GitHub. https:\/\/github.com\/lenhattan86\/allox."},{"key":"e_1_2_1_3_1","unstructured":"Accessed in November 2024. Amazon EC2 P3 Instances. https:\/\/aws.amazon.com\/ec2\/instance-types\/p3."},{"key":"e_1_2_1_4_1","unstructured":"Accessed in November 2024. AWS service quotas. https:\/\/docs.aws.amazon.com\/general\/latest\/gr\/aws_service_limits.html."},{"key":"e_1_2_1_5_1","unstructured":"Accessed in November 2024. Azure and AWS's 'GPU general availability' lies. https:\/\/www.fast.ai\/posts\/2016-12-19-gpu-lies.html."},{"key":"e_1_2_1_6_1","unstructured":"Accessed in November 2024. Azure Machine Learning pricing. https:\/\/azure.microsoft.com\/en-us\/pricing\/details\/machine-learning\/."},{"key":"e_1_2_1_7_1","unstructured":"Accessed in November 2024. Cannot Extend GPU Quota on Google Cloud. https:\/\/stackoverflow.com\/questions\/48362544\/cannot-extend-gpu-quota-on-google-cloud."},{"key":"e_1_2_1_8_1","unstructured":"Accessed in November 2024. Chameleon - A configurable experimental environment for large-scale cloud research. https:\/\/www.chameleoncloud.org."},{"key":"e_1_2_1_9_1","unstructured":"Accessed in November 2024. Cloudbank Website. https:\/\/www.cloudbank.org."},{"key":"e_1_2_1_10_1","unstructured":"Accessed in November 2024. DALI. https:\/\/developer.nvidia.com\/dali."},{"key":"e_1_2_1_11_1","unstructured":"Accessed in November 2024. Device Plugins | Kubernetes. https:\/\/kubernetes.io\/docs\/concepts\/extend-kubernetes\/compute-storage-net\/device-plugins\/."},{"key":"e_1_2_1_12_1","unstructured":"Accessed in November 2024. Emulating multipath wireless links on CloudLab and FABRIC. https:\/\/witestlab.poly.edu\/blog\/emulating-multipath-wireless\/."},{"key":"e_1_2_1_13_1","unstructured":"Accessed in November 2024. FastFlow GitHub. https:\/\/github.com\/SamsungLabs\/FastFlow."},{"key":"e_1_2_1_14_1","unstructured":"Accessed in November 2024. GCE Discussion: GPU Quota. https:\/\/groups.google.com\/g\/gce-discussion\/c\/mtHV1NlKKBo\/m\/i9uyk-PeAgAJ."},{"key":"e_1_2_1_15_1","unstructured":"Accessed in November 2024. GCE Discussion: No P100 GPUs in any us zone. https:\/\/groups.google.com\/g\/gce-discussion\/c\/34zBBmTV8Tg."},{"key":"e_1_2_1_16_1","unstructured":"Accessed in November 2024. GCE Discussion: Not enough resources to fulfill for the past 14 hours. https:\/\/groups.google.com\/g\/gce-discussion\/c\/8vCwUKaGs2o."},{"key":"e_1_2_1_17_1","unstructured":"Accessed in November 2024. Getting Started with Distributed Data Parallel. https:\/\/pytorch.org\/tutorials\/intermediate\/ddp_tutorial.html."},{"key":"e_1_2_1_18_1","unstructured":"Accessed in November 2024. Google Cloud GPU Pricing. https:\/\/cloud.google.com\/compute\/gpus-pricing."},{"key":"e_1_2_1_19_1","unstructured":"Accessed in November 2024. How to Optimize Data Transfers in CUDA C\/C++. https:\/\/developer.nvidia.com\/blog\/how-optimize-data-transfers-cuda-cc\/."},{"key":"e_1_2_1_20_1","unstructured":"Accessed in November 2024. ImageNet training in PyTorch. https:\/\/github.com\/pytorch\/examples\/tree\/main\/imagenet."},{"key":"e_1_2_1_21_1","unstructured":"Accessed in November 2024. Linux I\/O schedulers. https:\/\/wiki.ubuntu.com\/Kernel\/Reference\/IOSchedulers."},{"key":"e_1_2_1_22_1","unstructured":"Accessed in November 2024. MinIO GitHub. https:\/\/github.com\/msr-fiddle\/CoorDL."},{"key":"e_1_2_1_23_1","unstructured":"Accessed in November 2024. MLPerf Storage Benchmark Suite GitHub. https:\/\/github.com\/mlcommons\/storage."},{"key":"e_1_2_1_24_1","unstructured":"Accessed in November 2024. Muri GitHub. https:\/\/github.com\/pkusys\/Muri."},{"key":"e_1_2_1_25_1","unstructured":"Accessed in November 2024. Object detection reference training scripts. https:\/\/github.com\/pytorch\/vision\/tree\/main\/references\/detection."},{"key":"e_1_2_1_26_1","unstructured":"Accessed in November 2024. PyTorch CUDA Asynchronous Execution. https:\/\/pytorch.org\/docs\/master\/notes\/cuda.html#asynchronous-execution."},{"key":"e_1_2_1_27_1","unstructured":"Accessed in November 2024. RabbitMQ: easy to use flexible messaging and streaming \u2014 RabbitMQ. https:\/\/www.rabbitmq.com."},{"key":"e_1_2_1_28_1","unstructured":"Accessed in November 2024. RAMSSD GitHub. https:\/\/github.com\/thustorage\/ramssd."},{"key":"e_1_2_1_29_1","unstructured":"Accessed in November 2024. Synergy GitHub. https:\/\/github.com\/msr-fiddle\/synergy."},{"key":"e_1_2_1_30_1","unstructured":"Accessed in November 2024. Time-sharing GPUs on GKE | Google Kubernetes Engine (GKE) | Google Cloud. https:\/\/cloud.google.com\/kubernetes-engine\/docs\/concepts\/timesharing-gpus."},{"key":"e_1_2_1_31_1","unstructured":"Accessed in November 2024. Time-Slicing GPUs in Kubernetes \u2014 NVIDIA GPU Operator. https:\/\/docs.nvidia.com\/datacenter\/cloud-native\/gpu-operator\/latest\/gpu-sharing.html."},{"key":"e_1_2_1_32_1","unstructured":"Accessed in November 2024. What is io-uring? https:\/\/unixism.net\/loti\/what_is_io_uring.html."},{"key":"e_1_2_1_33_1","doi-asserted-by":"crossref","first-page":"127","DOI":"10.14778\/3485450.3485462","article-title":"Accelerating Recommendation System Training by Leveraging Popular Choices","volume":"15","author":"Adnan Muhammad","year":"2021","unstructured":"Muhammad Adnan, Yassaman Ebrahimzadeh Maboud, Divya Mahajan, and Prashant J Nair. 2021. Accelerating Recommendation System Training by Leveraging Popular Choices. Proceedings of the VLDB Endowment (PVLDB) 15, 1 (2021), 127\u2013140.","journal-title":"Proceedings of the VLDB Endowment (PVLDB)"},{"key":"e_1_2_1_34_1","volume-title":"Proceedings of the 2023 ACM Symposium on Cloud Computing (SoCC). 358\u2013375","author":"Audibert Andrew","year":"2023","unstructured":"Andrew Audibert, Yang Chen, Dan Graur, Ana Klimovic, Ji\u0159\u00ed \u0160im\u0161a, and Chandramohan A Thekkath. 2023. tf. data service: A case for disaggregating ML input data processing. In Proceedings of the 2023 ACM Symposium on Cloud Computing (SoCC). 358\u2013375."},{"key":"e_1_2_1_35_1","volume-title":"2009 IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS). IEEE, 163\u2013174","author":"Bakhoda Ali","year":"2009","unstructured":"Ali Bakhoda, George L Yuan, Wilson WL Fung, Henry Wong, and Tor M Aamodt. 2009. Analyzing CUDA workloads using a detailed GPU simulator. In 2009 IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS). IEEE, 163\u2013174."},{"key":"e_1_2_1_36_1","volume-title":"Once-for-all: Train one network and specialize it for efficient deployment. arXiv preprint arXiv:1908.09791","author":"Cai Han","year":"2019","unstructured":"Han Cai, Chuang Gan, Tianzhe Wang, Zhekai Zhang, and Song Han. 2019. Once-for-all: Train one network and specialize it for efficient deployment. arXiv preprint arXiv:1908.09791 (2019)."},{"key":"e_1_2_1_37_1","volume-title":"Proxylessnas: Direct neural architecture search on target task and hardware. arXiv preprint arXiv:1812.00332","author":"Cai Han","year":"2018","unstructured":"Han Cai, Ligeng Zhu, and Song Han. 2018. Proxylessnas: Direct neural architecture search on target task and hardware. arXiv preprint arXiv:1812.00332 (2018)."},{"key":"e_1_2_1_38_1","volume-title":"2023 IEEE International Symposium on High-Performance Computer Architecture (HPCA). IEEE, 220\u2013232","author":"Chen Weijian","year":"2023","unstructured":"Weijian Chen, Shuibing He, Yaowen Xu, Xuechen Zhang, Siling Yang, Shuang Hu, Xian-He Sun, and Gang Chen. 2023. icache: An importance-sampling-informed cache for accelerating i\/o-bound dnn model training. In 2023 IEEE International Symposium on High-Performance Computer Architecture (HPCA). IEEE, 220\u2013232."},{"key":"e_1_2_1_39_1","volume-title":"Automation & Test in Europe Conference & Exhibition (DATE). IEEE, 1375\u20131380","author":"Duan Zhuohui","year":"2018","unstructured":"Zhuohui Duan, Haikun Liu, Xiaofei Liao, and Hai Jin. 2018. HME: A lightweight emulator for hybrid memory. In 2018 Design, Automation & Test in Europe Conference & Exhibition (DATE). IEEE, 1375\u20131380."},{"key":"e_1_2_1_40_1","doi-asserted-by":"crossref","unstructured":"Dmitry Duplyakin Robert Ricci Aleksander Maricq Gary Wong Jonathon Duerig Eric Eide Leigh Stoller Mike Hibler David Johnson Kirk Webb et al. 2019. The Design and Operation of CloudLab. In 2019 USENIX annual technical conference (USENIX ATC). 1\u201314.","DOI":"10.1109\/ICNP.2019.8888128"},{"key":"e_1_2_1_41_1","volume-title":"Catastrophic forgetting in connectionist networks. Trends in cognitive sciences 3, 4","author":"French Robert M","year":"1999","unstructured":"Robert M French. 1999. Catastrophic forgetting in connectionist networks. Trends in cognitive sciences 3, 4 (1999), 128\u2013135."},{"key":"e_1_2_1_42_1","volume-title":"Deep learning","author":"Goodfellow Ian","unstructured":"Ian Goodfellow, Yoshua Bengio, and Aaron Courville. 2016. Deep learning. MIT press."},{"key":"e_1_2_1_43_1","volume-title":"2022 USENIX Annual Technical Conference (USENIX ATC). 689\u2013706","author":"Graur Dan","year":"2022","unstructured":"Dan Graur, Damien Aymon, Dan Kluser, Tanguy Albrici, Chandramohan A Thekkath, and Ana Klimovic. 2022. Cachew: Machine learning input data processing as a service. In 2022 USENIX Annual Technical Conference (USENIX ATC). 689\u2013706."},{"key":"e_1_2_1_44_1","volume-title":"16th USENIX Symposium on Networked Systems Design and Implementation (NSDI). 485\u2013500","author":"Gu Juncheng","year":"2019","unstructured":"Juncheng Gu, Mosharaf Chowdhury, Kang G Shin, Yibo Zhu, Myeongjae Jeon, Junjie Qian, Hongqiang Liu, and Chuanxiong Guo. 2019. Tiresias: A GPU cluster manager for distributed deep learning. In 16th USENIX Symposium on Networked Systems Design and Implementation (NSDI). 485\u2013500."},{"key":"e_1_2_1_45_1","volume-title":"14th USENIX Symposium on Operating Systems Design and Implementation (OSDI). 443\u2013462","author":"Gujarati Arpan","year":"2020","unstructured":"Arpan Gujarati, Reza Karimi, Safya Alzayat, Wei Hao, Antoine Kaufmann, Ymir Vigfusson, and Jonathan Mace. 2020. Serving DNNs like Clockwork: Performance Predictability from the Bottom Up. In 14th USENIX Symposium on Operating Systems Design and Implementation (OSDI). 443\u2013462."},{"key":"e_1_2_1_46_1","volume-title":"2008 USENIX Annual Technical Conference (USENIX ATC).","author":"Hibler Mike","year":"2008","unstructured":"Mike Hibler, Robert Ricci, Leigh Stoller, Jonathon Duerig, Shashi Guruprasad, Tim Stack, Kirk Webb, and Jay Lepreau. 2008. Large-scale virtualization in the emulab network testbed. In 2008 USENIX Annual Technical Conference (USENIX ATC)."},{"key":"e_1_2_1_47_1","volume-title":"Proceedings of the 29th Symposium on Operating Systems Principles (SOSP). 642\u2013657","author":"Subramanya Suhas Jayaram","year":"2023","unstructured":"Suhas Jayaram Subramanya, Daiyaan Arfeen, Shouxu Lin, Aurick Qiao, Zhihao Jia, and Gregory R Ganger. 2023. Sia: Heterogeneity-aware, goodput-optimized ML-cluster scheduling. In Proceedings of the 29th Symposium on Operating Systems Principles (SOSP). 642\u2013657."},{"key":"e_1_2_1_48_1","volume-title":"11th USENIX Workshop on Hot Topics in Cloud Computing (HotCloud 19)","author":"Kakaraparthy Aarati","year":"2019","unstructured":"Aarati Kakaraparthy, Abhay Venkatesh, Amar Phanishayee, and Shivaram Venkataraman. 2019. The case for unifying data loading in machine learning clusters. In 11th USENIX Workshop on Hot Topics in Cloud Computing (HotCloud 19)."},{"key":"e_1_2_1_49_1","volume-title":"Dataset Placement and Data Loading Optimizations for Cloud-Native Deep Learning Workloads. In 2023 IEEE 26th International Symposium on Real-Time Distributed Computing (ISORC). IEEE, 107\u2013116","author":"Kang Zhuangwei","year":"2023","unstructured":"Zhuangwei Kang, Ziran Min, Shuang Zhou, Yogesh D Barve, and Aniruddha Gokhale. 2023. Dataset Placement and Data Loading Optimizations for Cloud-Native Deep Learning Workloads. In 2023 IEEE 26th International Symposium on Real-Time Distributed Computing (ISORC). IEEE, 107\u2013116."},{"key":"e_1_2_1_50_1","unstructured":"Kate Keahey Jason Anderson Zhuo Zhen Pierre Riteau Paul Ruth Dan Stanzione Mert Cevik Jacob Colleran Haryadi S Gunawi Cody Hammock et al. 2020. Lessons learned from the chameleon testbed. In 2020 USENIX annual technical conference (USENIX ATC). 219\u2013233."},{"key":"e_1_2_1_51_1","volume-title":"2020 ACM\/IEEE 47th Annual International Symposium on Computer Architecture (ISCA). IEEE, 473\u2013486","author":"Khairy Mahmoud","year":"2020","unstructured":"Mahmoud Khairy, Zhesheng Shen, Tor M Aamodt, and Timothy G Rogers. 2020. Accel-Sim: An extensible simulation framework for validated GPU modeling. In 2020 ACM\/IEEE 47th Annual International Symposium on Computer Architecture (ISCA). IEEE, 473\u2013486."},{"key":"e_1_2_1_52_1","volume-title":"SHADE: Enable Fundamental Cacheability for Distributed Deep Learning Training. In 21st USENIX Conference on File and Storage Technologies (FAST). 135\u2013152","author":"Seraj Khan Redwan Ibne","year":"2023","unstructured":"Redwan Ibne Seraj Khan, Ahmad Hossein Yazdani, Yuqi Fu, Arnab K Paul, Bo Ji, Xun Jian, Yue Cheng, and Ali R Butt. 2023. SHADE: Enable Fundamental Cacheability for Distributed Deep Learning Training. In 21st USENIX Conference on File and Storage Technologies (FAST). 135\u2013152."},{"key":"e_1_2_1_53_1","volume-title":"Macsim: A cpu-gpu heterogeneous simulation framework user guide","author":"Kim Hyesoon","year":"2012","unstructured":"Hyesoon Kim, Jaekyu Lee, Nagesh B Lakshminarayana, Jaewoong Sim, Jieun Lim, and Tri Pho. 2012. Macsim: A cpu-gpu heterogeneous simulation framework user guide. Georgia Institute of Technology (2012), 1\u201357."},{"key":"e_1_2_1_54_1","volume-title":"Proceedings of the Fifteenth European Conference on Computer Systems (EuroSys). 1\u201316","author":"Le Tan N","year":"2020","unstructured":"Tan N Le, Xiao Sun, Mosharaf Chowdhury, and Zhenhua Liu. 2020. Allox: compute allocation in hybrid clusters. In Proceedings of the Fifteenth European Conference on Computer Systems (EuroSys). 1\u201316."},{"key":"e_1_2_1_55_1","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 12011\u201312020","author":"Leclerc Guillaume","year":"2023","unstructured":"Guillaume Leclerc, Andrew Ilyas, Logan Engstrom, Sung Min Park, Hadi Salman, and Aleksander M\u0105dry. 2023. FFCV: Accelerating training by removing data bottlenecks. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 12011\u201312020."},{"key":"e_1_2_1_56_1","volume-title":"16th USENIX Conference on File and Storage Technologies (FAST). 83\u201390","author":"Li Huaicheng","year":"2018","unstructured":"Huaicheng Li, Mingzhe Hao, Michael Hao Tong, Swaminathan Sundararaman, Matias Bj\u00f8rling, and Haryadi S Gunawi. 2018. The CASE of FEMU: Cheap, accurate, scalable and extensible flash emulator. In 16th USENIX Conference on File and Storage Technologies (FAST). 83\u201390."},{"key":"e_1_2_1_57_1","doi-asserted-by":"crossref","first-page":"3005","DOI":"10.14778\/3415478.3415530","article-title":"PyTorch distributed: experiences on accelerating data parallel training","volume":"13","author":"Li Shen","year":"2020","unstructured":"Shen Li, Yanli Zhao, Rohan Varma, Omkar Salpekar, Pieter Noordhuis, Teng Li, Adam Paszke, Jeff Smith, Brian Vaughan, Pritam Damania, et al. 2020. PyTorch distributed: experiences on accelerating data parallel training. Proceedings of the VLDB Endowment (PVLDB) 13, 12 (2020), 3005\u20133018.","journal-title":"Proceedings of the VLDB Endowment (PVLDB)"},{"key":"e_1_2_1_58_1","volume-title":"Proceedings of the 51st International Conference on Parallel Processing (ICPP). 1\u201311","author":"Liu Jie","year":"2022","unstructured":"Jie Liu, Bogdan Nicolae, and Dong Li. 2022. Lobster: Load balance-aware I\/O for distributed DNN training. In Proceedings of the 51st International Conference on Parallel Processing (ICPP). 1\u201311."},{"key":"e_1_2_1_59_1","volume-title":"2015 IEEE 23rd International Symposium on Modeling, Analysis, and Simulation of Computer and Telecommunication Systems (MASCOTS). IEEE, 43\u201346","author":"Malladi Krishna T","year":"2015","unstructured":"Krishna T Malladi, Mu-Tien Chang, John Ping, and Hongzhong Zheng. 2015. FAME: A fast and accurate memory emulator for new memory system architecture exploration. In 2015 IEEE 23rd International Symposium on Modeling, Analysis, and Simulation of Computer and Telecommunication Systems (MASCOTS). IEEE, 43\u201346."},{"key":"e_1_2_1_60_1","volume-title":"Looking Beyond GPUs for DNN Scheduling on Multi-Tenant Clusters. In 16th USENIX Symposium on Operating Systems Design and Implementation (OSDI). 579\u2013596","author":"Mohan Jayashree","year":"2022","unstructured":"Jayashree Mohan, Amar Phanishayee, Janardhan Kulkarni, and Vijay Chidambaram. 2022. Looking Beyond GPUs for DNN Scheduling on Multi-Tenant Clusters. In 16th USENIX Symposium on Operating Systems Design and Implementation (OSDI). 579\u2013596."},{"key":"e_1_2_1_61_1","doi-asserted-by":"crossref","first-page":"771","DOI":"10.14778\/3446095.3446100","article-title":"Analyzing and mitigating data stalls in DNN training","volume":"14","author":"Mohan Jayashree","year":"2021","unstructured":"Jayashree Mohan, Amar Phanishayee, Ashish Raniwala, and Vijay Chidambaram. 2021. Analyzing and mitigating data stalls in DNN training. Proceedings of the VLDB Endowment (PVLDB) 14, 5 (2021), 771\u2013784.","journal-title":"Proceedings of the VLDB Endowment (PVLDB)"},{"key":"e_1_2_1_62_1","doi-asserted-by":"crossref","first-page":"2945","DOI":"10.14778\/3476311.3476374","article-title":"tf. data: a machine learning data processing framework","volume":"14","author":"Murray Derek G","year":"2021","unstructured":"Derek G Murray, Ji\u0159\u00ed \u0160im\u0161a, Ana Klimovic, and Ihor Indyk. 2021. tf. data: a machine learning data processing framework. Proceedings of the VLDB Endowment (PVLDB) 14, 12 (2021), 2945\u20132958.","journal-title":"Proceedings of the VLDB Endowment (PVLDB)"},{"key":"e_1_2_1_63_1","volume-title":"Heterogeneity-Aware Cluster Scheduling Policies for Deep Learning Workloads. In 14th USENIX Symposium on Operating Systems Design and Implementation (OSDI). 481\u2013498","author":"Narayanan Deepak","year":"2020","unstructured":"Deepak Narayanan, Keshav Santhanam, Fiodar Kazhamiaka, Amar Phanishayee, and Matei Zaharia. 2020. Heterogeneity-Aware Cluster Scheduling Policies for Deep Learning Workloads. In 14th USENIX Symposium on Operating Systems Design and Implementation (OSDI). 481\u2013498."},{"key":"e_1_2_1_64_1","volume-title":"Proceedings of the 29th Symposium on Operating Systems Principles (SOSP). 595\u2013610","author":"Ng Kelvin KW","year":"2023","unstructured":"Kelvin KW Ng, Henri Maxime Demoulin, and Vincent Liu. 2023. Paella: Low-latency model serving with software-defined gpu scheduling. In Proceedings of the 29th Symposium on Operating Systems Principles (SOSP). 595\u2013610."},{"key":"e_1_2_1_65_1","volume-title":"Naomi Alterman, Rob Fatland, Sarah Stone, Amanda Tan, et al.","author":"Norman Michael","year":"2021","unstructured":"Michael Norman, Vince Kellen, Shava Smallen, Brian DeMeulle, Shawn Strande, Ed Lazowska, Naomi Alterman, Rob Fatland, Sarah Stone, Amanda Tan, et al. 2021. Cloudbank: Managed services to simplify cloud access for computer science research and education. In Practice and Experience in Advanced Research Computing (PEARC). 1\u20134."},{"key":"e_1_2_1_66_1","volume-title":"Pytorch: An imperative style, high-performance deep learning library. Advances in Neural Information Processing Systems (NIPS) 32","author":"Paszke Adam","year":"2019","unstructured":"Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. 2019. Pytorch: An imperative style, high-performance deep learning library. Advances in Neural Information Processing Systems (NIPS) 32 (2019)."},{"key":"e_1_2_1_67_1","volume-title":"15th USENIX Symposium on Operating Systems Design and Implementation (OSDI).","author":"Qiao Aurick","year":"2021","unstructured":"Aurick Qiao, Sang Keun Choe, Suhas Jayaram Subramanya, Willie Neiswanger, Qirong Ho, Hao Zhang, Gregory R Ganger, and Eric P Xing. 2021. Pollux: Co-adaptive cluster scheduling for goodput-optimized deep learning. In 15th USENIX Symposium on Operating Systems Design and Implementation (OSDI)."},{"key":"e_1_2_1_68_1","doi-asserted-by":"crossref","unstructured":"Olga Russakovsky Jia Deng Hao Su Jonathan Krause Sanjeev Satheesh Sean Ma Zhiheng Huang Andrej Karpathy Aditya Khosla Michael Bernstein et al. 2015. Imagenet large scale visual recognition challenge. International journal of computer vision (IJCV) 115 (2015) 211\u2013252.","DOI":"10.1007\/s11263-015-0816-y"},{"key":"e_1_2_1_69_1","volume-title":"Megatron-lm: Training multi-billion parameter language models using model parallelism. arXiv preprint arXiv:1909.08053","author":"Shoeybi Mohammad","year":"2019","unstructured":"Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. 2019. Megatron-lm: Training multi-billion parameter language models using model parallelism. arXiv preprint arXiv:1909.08053 (2019)."},{"key":"e_1_2_1_70_1","volume-title":"2018 IEEE International Conference on Cluster Computing (CLUSTER). IEEE, 392\u2013401","author":"Sreedhar Dheeraj","year":"2018","unstructured":"Dheeraj Sreedhar, Vaibhav Saxena, Yogish Sabharwal, Ashish Verma, and Sameer Kumar. 2018. Efficient training of convolutional neural nets on large distributed systems. In 2018 IEEE International Conference on Cluster Computing (CLUSTER). IEEE, 392\u2013401."},{"key":"e_1_2_1_71_1","volume-title":"Proceedings of the Nineteenth European Conference on Computer Systems (EuroSys). 1075\u20131092","author":"Strati Foteini","year":"2024","unstructured":"Foteini Strati, Xianzhe Ma, and Ana Klimovic. 2024. Orion: Interference-aware, Fine-grained GPU Sharing for ML Applications. In Proceedings of the Nineteenth European Conference on Computer Systems (EuroSys). 1075\u20131092."},{"key":"e_1_2_1_72_1","volume-title":"Proceedings of the 46th International Symposium on Computer Architecture (ISCA). 197\u2013209","author":"Sun Yifan","year":"2019","unstructured":"Yifan Sun, Trinayan Baruah, Saiful A Mojumder, Shi Dong, Xiang Gong, Shane Treadway, Yuhui Bao, Spencer Hance, Carter McCardwell, Vincent Zhao, et al. 2019. MGPUSim: Enabling multi-GPU performance modeling and optimization. In Proceedings of the 46th International Symposium on Computer Architecture (ISCA). 197\u2013209."},{"key":"e_1_2_1_73_1","doi-asserted-by":"crossref","first-page":"1086","DOI":"10.14778\/3579075.3579083","article-title":"Fastflow: Accelerating deep learning model training with smart offloading of input data pipeline","volume":"16","author":"Um Taegeon","year":"2023","unstructured":"Taegeon Um, Byungsoo Oh, Byeongchan Seo, Minhyeok Kweun, Goeun Kim, and Woo-Yeon Lee. 2023. Fastflow: Accelerating deep learning model training with smart offloading of input data pipeline. Proceedings of the VLDB Endowment (PVLDB) 16, 5 (2023), 1086\u20131099.","journal-title":"Proceedings of the VLDB Endowment (PVLDB)"},{"key":"e_1_2_1_74_1","first-page":"696","article-title":"Wavelet: Efficient DNN training with tick-tock scheduling","volume":"3","author":"Wang Guanhua","year":"2021","unstructured":"Guanhua Wang, Kehan Wang, Kenan Jiang, Xiangjun Li, and Ion Stoica. 2021. Wavelet: Efficient DNN training with tick-tock scheduling. Proceedings of Machine Learning and Systems (MLSys) 3 (2021), 696\u2013710.","journal-title":"Proceedings of Machine Learning and Systems (MLSys)"},{"key":"e_1_2_1_75_1","volume-title":"Proceedings of the 16th ACM Workshop on Hot Topics in Storage and File Systems (HotStorage). 63\u201370","author":"Wang Meng","year":"2024","unstructured":"Meng Wang, Gus Waldspurger, and Swaminathan Sundararaman. 2024. A Selective Preprocessing Offloading Framework for Reducing Data Traffic in DL Training. In Proceedings of the 16th ACM Workshop on Hot Topics in Storage and File Systems (HotStorage). 63\u201370."},{"key":"e_1_2_1_76_1","volume-title":"Transparent GPU Sharing in Container Clouds for Deep Learning Workloads. In 20th USENIX Symposium on Networked Systems Design and Implementation (NSDI). 69\u201385","author":"Wu Bingyang","year":"2023","unstructured":"Bingyang Wu, Zili Zhang, Zhihao Bai, Xuanzhe Liu, and Xin Jin. 2023. Transparent GPU Sharing in Container Clouds for Deep Learning Workloads. In 20th USENIX Symposium on Networked Systems Design and Implementation (NSDI). 69\u201385."},{"key":"e_1_2_1_77_1","volume-title":"13th USENIX Symposium on Operating Systems Design and Implementation (OSDI). 595\u2013610","author":"Xiao Wencong","year":"2018","unstructured":"Wencong Xiao, Romil Bhardwaj, Ramachandran Ramjee, Muthian Sivathanu, Nipun Kwatra, Zhenhua Han, Pratyush Patel, Xuan Peng, Hanyu Zhao, Quanlu Zhang, et al. 2018. Gandiva: Introspective cluster scheduling for deep learning. In 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI). 595\u2013610."},{"key":"e_1_2_1_78_1","volume-title":"AntMan: Dynamic Scaling on GPU Clusters for Deep Learning. In 14th USENIX Symposium on Operating Systems Design and Implementation (OSDI). 533\u2013548","author":"Xiao Wencong","year":"2020","unstructured":"Wencong Xiao, Shiru Ren, Yong Li, Yang Zhang, Pengyang Hou, Zhi Li, Yihui Feng, Wei Lin, and Yangqing Jia. 2020. AntMan: Dynamic Scaling on GPU Clusters for Deep Learning. In 14th USENIX Symposium on Operating Systems Design and Implementation (OSDI). 533\u2013548."},{"key":"e_1_2_1_79_1","volume-title":"Proceedings of the 2022 International Conference on Management of Data (SIGMOD). 1286\u20131300","author":"Xu Lijie","year":"2022","unstructured":"Lijie Xu, Shuang Qiu, Binhang Yuan, Jiawei Jiang, Cedric Renggli, Shaoduo Gan, Kaan Kara, Guoliang Li, Ji Liu, Wentao Wu, et al. 2022. In-database machine learning with corgipile: Stochastic gradient descent without full data shuffle. In Proceedings of the 2022 International Conference on Management of Data (SIGMOD). 1286\u20131300."},{"key":"e_1_2_1_80_1","volume-title":"2019 IEEE 26th International Conference on High Performance Computing, Data, and Analytics (HiPC). IEEE, 235\u2013245","author":"Yang Chih-Chieh","year":"2019","unstructured":"Chih-Chieh Yang and Guojing Cong. 2019. Accelerating data loading in deep neural network training. In 2019 IEEE 26th International Conference on High Performance Computing, Data, and Analytics (HiPC). IEEE, 235\u2013245."},{"key":"e_1_2_1_81_1","volume-title":"Proceedings of Machine Learning and Systems (MLSys), I. Dhillon, D. Papailiopoulos, and V. Sze (Eds.)","volume":"2","author":"Yu Peifeng","year":"2020","unstructured":"Peifeng Yu and Mosharaf Chowdhury. 2020. Fine-Grained GPU Sharing Primitives for Deep Learning Applications. In Proceedings of Machine Learning and Systems (MLSys), I. Dhillon, D. Papailiopoulos, and V. Sze (Eds.), Vol. 2. 98\u2013111."},{"key":"e_1_2_1_82_1","volume-title":"20th USENIX Symposium on Networked Systems Design and Implementation (NSDI). 787\u2013808","author":"Zhang Hong","year":"2023","unstructured":"Hong Zhang, Yupeng Tang, Anurag Khandelwal, and Ion Stoica. 2023. SHEPHERD: Serving DNNs in the wild. In 20th USENIX Symposium on Networked Systems Design and Implementation (NSDI). 787\u2013808."},{"key":"e_1_2_1_83_1","volume-title":"Proceedings of the Eighteenth European Conference on Computer Systems (EuroSys). 883\u2013898","author":"Zhao Hanyu","year":"2023","unstructured":"Hanyu Zhao, Zhenhua Han, Zhi Yang, Quanlu Zhang, Mingxia Li, Fan Yang, Qianxi Zhang, Binyang Li, Yuqing Yang, Lili Qiu, et al. 2023. Silod: A co-design of caching and scheduling for deep learning clusters. In Proceedings of the Eighteenth European Conference on Computer Systems (EuroSys). 883\u2013898."},{"key":"e_1_2_1_84_1","volume-title":"14th USENIX Symposium on Operating Systems Design and Implementation (OSDI). 515\u2013532","author":"Zhao Hanyu","year":"2020","unstructured":"Hanyu Zhao, Zhenhua Han, Zhi Yang, Quanlu Zhang, Fan Yang, Lidong Zhou, Mao Yang, Francis CM Lau, Yuqi Wang, Yifan Xiong, et al. 2020. HiveD: Sharing a GPU Cluster for Deep Learning with Guarantees. In 14th USENIX Symposium on Operating Systems Design and Implementation (OSDI). 515\u2013532."},{"key":"e_1_2_1_85_1","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3589773","article-title":"Goldminer: Elastic scaling of training data pre-processing pipelines for deep learning","volume":"1","author":"Zhao Hanyu","year":"2023","unstructured":"Hanyu Zhao, Zhi Yang, Yu Cheng, Chao Tian, Shiru Ren, Wencong Xiao, Man Yuan, Langshi Chen, Kaibo Liu, Yang Zhang, et al. 2023. Goldminer: Elastic scaling of training data pre-processing pipelines for deep learning. Proceedings of the ACM on Management of Data 1, 2 (2023), 1\u201325.","journal-title":"Proceedings of the ACM on Management of Data"},{"key":"e_1_2_1_86_1","volume-title":"cedar: Composable and Optimized Machine Learning Input Data Pipelines. arXiv preprint arXiv:2401.08895","author":"Zhao Mark","year":"2024","unstructured":"Mark Zhao, Emanuel Adamiak, and Christos Kozyrakis. 2024. cedar: Composable and Optimized Machine Learning Input Data Pipelines. arXiv preprint arXiv:2401.08895 (2024)."},{"key":"e_1_2_1_87_1","volume-title":"Proceedings of the 49th Annual International Symposium on Computer Architecture (ISCA). 1042\u20131057","author":"Zhao Mark","year":"2022","unstructured":"Mark Zhao, Niket Agarwal, Aarti Basant, Bu\u011fra Gedik, Satadru Pan, Mustafa Ozdal, Rakesh Komuravelli, Jerry Pan, Tianshu Bao, Haowei Lu, et al. 2022. Understanding data storage and ingestion for large-scale deep recommendation model training: Industrial product. In Proceedings of the 49th Annual International Symposium on Computer Architecture (ISCA). 1042\u20131057."},{"key":"e_1_2_1_88_1","doi-asserted-by":"crossref","unstructured":"Yanli Zhao Andrew Gu Rohan Varma Liang Luo Chien-Chin Huang Min Xu Less Wright Hamid Shojanazeri Myle Ott Sam Shleifer et al. 2023. Pytorch fsdp: experiences on scaling fully sharded data parallel. arXiv preprint arXiv:2304.11277 (2023).","DOI":"10.14778\/3611540.3611569"},{"key":"e_1_2_1_89_1","volume-title":"Proceedings of the ACM SIGCOMM 2022 Conference (SIGCOMM). 428\u2013440","author":"Zhao Yihao","year":"2022","unstructured":"Yihao Zhao, Yuanqiang Liu, Yanghua Peng, Yibo Zhu, Xuanzhe Liu, and Xin Jin. 2022. Multi-resource interleaving for deep learning training. In Proceedings of the ACM SIGCOMM 2022 Conference (SIGCOMM). 428\u2013440."},{"key":"e_1_2_1_90_1","volume-title":"2018 IEEE 26th International Symposium on Modeling, Analysis, and Simulation of Computer and Telecommunication Systems (MASCOTS). IEEE, 145\u2013156","author":"Zhu Yue","year":"2018","unstructured":"Yue Zhu, Fahim Chowdhury, Huansong Fu, Adam Moody, Kathryn Mohror, Kento Sato, and Weikuan Yu. 2018. Entropy-aware I\/O pipelining for large-scale deep learning on HPC systems. In 2018 IEEE 26th International Symposium on Modeling, Analysis, and Simulation of Computer and Telecommunication Systems (MASCOTS). IEEE, 145\u2013156."}],"container-title":["Proceedings of the VLDB Endowment"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.14778\/3725688.3725716","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,8,29]],"date-time":"2025-08-29T14:23:52Z","timestamp":1756477432000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.14778\/3725688.3725716"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,2]]},"references-count":90,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2025,2]]}},"alternative-id":["10.14778\/3725688.3725716"],"URL":"https:\/\/doi.org\/10.14778\/3725688.3725716","relation":{},"ISSN":["2150-8097"],"issn-type":[{"value":"2150-8097","type":"print"}],"subject":[],"published":{"date-parts":[[2025,2]]},"assertion":[{"value":"2025-08-29","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}