{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,11]],"date-time":"2026-07-11T15:38:47Z","timestamp":1783784327772,"version":"3.55.0"},"reference-count":173,"publisher":"Association for Computing Machinery (ACM)","issue":"10","license":[{"start":{"date-parts":[[2025,5,7]],"date-time":"2025-05-07T00:00:00Z","timestamp":1746576000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/100000015","name":"U.S. Department of Energy","doi-asserted-by":"crossref","id":[{"id":"10.13039\/100000015","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Office of Science, Office of Advanced Scientific Computing Research","award":["DE-AC02-05CH11231"],"award-info":[{"award-number":["DE-AC02-05CH11231"]}]},{"DOI":"10.13039\/100017223","name":"National Energy Research Scientific Computing Center","doi-asserted-by":"crossref","id":[{"id":"10.13039\/100017223","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Office of Science of the U.S. Department of Energy","award":["DE-AC02-05CH11231"],"award-info":[{"award-number":["DE-AC02-05CH11231"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Comput. Surv."],"published-print":{"date-parts":[[2025,10,31]]},"abstract":"<jats:p>Growing interest in Artificial Intelligence (AI) has resulted in a surge in demand for faster methods of Machine Learning (ML) model training and inference. This demand for speed has prompted the use of high performance computing (HPC) systems that excel in managing distributed workloads. Because data is the main fuel for AI applications, the performance of the storage and I\/O subsystem of HPC systems is critical. In the past, HPC applications accessed large portions of data written by simulations or experiments or ingested data for visualizations or analysis tasks. ML workloads perform small reads spread across a large number of random files. This shift of I\/O access patterns poses several challenges to modern parallel storage systems. In this article, we survey I\/O in ML applications on HPC systems, and target literature within a 6-year time window from 2019 to 2024. We define the scope of the survey, provide an overview of the common phases of ML, review available profilers and benchmarks, examine the I\/O patterns encountered during offline data preparation, training, and inference, and explore I\/O optimizations utilized in modern ML frameworks and proposed in recent literature. Lastly, we seek to expose research gaps that could spawn further R&amp;D.<\/jats:p>","DOI":"10.1145\/3722215","type":"journal-article","created":{"date-parts":[[2025,3,7]],"date-time":"2025-03-07T11:24:51Z","timestamp":1741346691000},"page":"1-41","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":12,"title":["I\/O in Machine Learning Applications on HPC Systems: A 360-degree Survey"],"prefix":"10.1145","volume":"57","author":[{"ORCID":"https:\/\/orcid.org\/0009-0005-1732-7776","authenticated-orcid":false,"given":"Noah","family":"Lewis","sequence":"first","affiliation":[{"name":"Louisiana State University and A&amp;M College, Baton Rouge, United States and The Ohio State University, Columbus, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3915-1135","authenticated-orcid":false,"given":"Jean Luca","family":"Bez","sequence":"additional","affiliation":[{"name":"Lawrence Berkeley National Laboratory, Berkeley, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3048-3448","authenticated-orcid":false,"given":"Suren","family":"Byna","sequence":"additional","affiliation":[{"name":"The Ohio State University, Columbus, United States and Lawrence Berkeley National Laboratory, Berkeley, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,5,7]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.aej.2022.05.052"},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1142\/9789813235533_0011"},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.1109\/BigData47090.2019.9005703"},{"key":"e_1_3_1_5_2","unstructured":"Takuya Akiba Shuji Suzuki and Keisuke Fukuda. 2017. Extremely large minibatch SGD: Training ResNet-50 on imageNet in 15 Minutes. Retrieved from https:\/\/arxiv.org\/abs\/1711.04325"},{"key":"e_1_3_1_6_2","unstructured":"Amazon Web Services. 2024. 1000 Genomes Project. Retrieved September 12 2024 from https:\/\/registry.opendata.aws\/1000-genomes\/"},{"key":"e_1_3_1_7_2","unstructured":"Amazon Web Services. 2024. AWS Public Datasets. Retrieved September 12 2024 from https:\/\/registry.opendata.aws\/"},{"key":"e_1_3_1_8_2","volume-title":"Amazon Simple Storage Service (S3)","author":"Services Amazon Web","unstructured":"Amazon Web Services. n.d.. Amazon Simple Storage Service (S3). Retrieved March 20, 2025 from https:\/\/aws.amazon.com\/s3\/"},{"key":"e_1_3_1_9_2","first-page":"441","volume-title":"Proceedings of the 20th European Symposium on Artificial Neural Networks, Computational Intelligence and Machine Learning (ESANN)","author":"Anguita Davide","year":"2012","unstructured":"Davide Anguita, Luca Ghelardoni, Alessandro Ghio, Luca Oneto, Sandro Ridella. 2012. The\u2019K\u2019in K-fold cross validation. In Proceedings of the 20th European Symposium on Artificial Neural Networks, Computational Intelligence and Machine Learning (ESANN). University of Genova. 441\u2013446."},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/SCWS55283.2021.00018"},{"key":"e_1_3_1_11_2","volume-title":"RecordIO: Apache Mesos Documentation","author":"Mesos Apache","year":"2024","unstructured":"Apache Mesos. 2024. RecordIO: Apache Mesos Documentation. Retrieved February 19, 2024 from https:\/\/mesos.apache.org\/documentation\/latest\/recordio\/"},{"key":"e_1_3_1_12_2","unstructured":"Apache Software Foundation. 2024. Apache Spark. Retrieved March 20 2025 from https:\/\/spark.apache.org\/"},{"key":"e_1_3_1_13_2","volume-title":"Hadoop","author":"Foundation Apache Software","year":"2024","unstructured":"Apache Software Foundation. 2024. Hadoop. Retrieved March 20, 2025 from https:\/\/hadoop.apache.org"},{"key":"e_1_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2018.2798607"},{"key":"e_1_3_1_15_2","unstructured":"Olivier Beaumont Lionel Eyraud-Dubois Alena Shilova and Xunyi Zhao. 2022. Weight Offloading Strategies for Training Large DNN Models. (Feb.2022). Retrieved March 20 2025 from https:\/\/inria.hal.science\/hal-03580767. working paper or preprint."},{"key":"e_1_3_1_16_2","doi-asserted-by":"crossref","unstructured":"L\u00e9on Bottou. 2012. Stochastic gradient tricks. In Neural Networks Tricks of the Trade Reloaded Gr\u00e9goire Montavon Genevieve B. Orr and Klaus-Robert M\u00fcller (Eds.). Springer 430\u2013445. Retrieved from http:\/\/leon.bottou.org\/papers\/bottou-tricks-2012","DOI":"10.1007\/978-3-642-35289-8_25"},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","unstructured":"L\u00e9on Bottou Frank E. Curtis and Jorge Nocedal. 2018. Optimization methods for large-scale machine learning. SIAM Review 60 2 (2018) 223-311. DOI:10.1137\/16M1080173","DOI":"10.1137\/16M1080173"},{"key":"e_1_3_1_18_2","volume-title":"Simple Object Access Protocol (SOAP) 1.1","author":"Box Don","year":"2000","unstructured":"Don Box, David Ehnebuske, Gopal Kakivaya, Andrew Layman, Noah Mendelsohn, Henrik Frystyk Nielsen, Satish Thatte, and Dave Winer. 2000. Simple Object Access Protocol (SOAP) 1.1. W3C Note. World Wide Web Consortium. Retrieved from http:\/\/www.w3.org\/TR\/SOAP\/"},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPEC43674.2020.9286169"},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPEC43674.2020.9286138"},{"key":"e_1_3_1_21_2","unstructured":"Jason Brownlee. 2018. What is the difference between a batch and an epoch in a neural network. Machine Learning Mastery 20 1 (2018) 1\u201315. Retrieved from https:\/\/lipingyang.org"},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","unstructured":"Suren Byna M. Scot Breitenfeld Bin Dong Quincey Koziol Elena Pourmal Dana Robinson Jerome Soumagne Houjun Tang Venkatram Vishwanath and Richard Warren. 2020. ExaHDF5: Delivering efficient parallel I\/O on exascale computing systems. J. Comput. Sci. Technol. 35 1 (January 2020) 145\u2013160. DOI:10.1007\/s11390-020-9822-9","DOI":"10.1007\/s11390-020-9822-9"},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.1109\/CloudCom.2010.115"},{"key":"e_1_3_1_24_2","unstructured":"Tianqi Chen Mu Li Yutian Li Min Lin Naiyan Wang Minjie Wang Tianjun Xiao Bing Xu Chiyuan Zhang and Zheng Zhang. 2015. MXNet: A flexible and efficient machine learning library for heterogeneous distributed systems. Retrieved from https:\/\/arxiv.org\/abs\/1512.01274"},{"key":"e_1_3_1_25_2","unstructured":"Yu Chen Zhenming Liu Bin Ren and Xin Jin. 2020. On efficient constructions of checkpoints. In Proceedings of the 37th International Conference on Machine Learning (ICML\u201920). JMLR.org."},{"key":"e_1_3_1_26_2","unstructured":"Peng Cheng and Haryadi S. Gunawi. 2021. Storage benchmarking with deep learning workloads. Retrieved March 20 2025 from https:\/\/api.semanticscholar.org\/CorpusID:231845927"},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1371\/journal.pone.0293338"},{"key":"e_1_3_1_28_2","volume-title":"Search Results for Distributed-Training: Search Query for Distributed-training up to 2015: Distributed Training Created:..2015 and up to 2024: Distributed Training Created:..2024","author":"Community Stack Overflow","year":"2024","unstructured":"Stack Overflow Community. 2024. Search Results for Distributed-Training: Search Query for Distributed-training up to 2015: Distributed Training Created:..2015 and up to 2024: Distributed Training Created:..2024."},{"key":"e_1_3_1_29_2","unstructured":"Donglai Dai. 2022. SCR-Exa: Enhanced scalable checkpoint restart (SCR) library for next generation exascale computing. (22022). Retrieved from https:\/\/www.osti.gov\/biblio\/1847927"},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.1109\/CCGrid54584.2022.00011"},{"key":"e_1_3_1_31_2","unstructured":"Tri Dao Daniel Y. Fu Stefano Ermon Atri Rudra and Christopher R\u00e9. 2022. FlashAttention: Fast and memory-efficient exact attention with IO-awareness. Retrieved from https:\/\/arxiv.org\/abs\/2205.14135"},{"key":"e_1_3_1_32_2","volume-title":"Dask: Scalable Analytics in Python","year":"2024","unstructured":"Dask. 2024. Dask: Scalable Analytics in Python. https:\/\/www.dask.org\/. Accessed on February 5, 2024."},{"key":"e_1_3_1_33_2","unstructured":"Hoang Anh Dau Anthony Bagnall Kaveh Kamgar Chin-Chia Michael Yeh Yan Zhu Shaghayegh Gharghabi Chotirat Ann Ratanamahatana and Eamonn Keogh. 2019. The UCR Time Series Archive. Retrieved from https:\/\/arxiv.org\/abs\/1810.07758"},{"key":"e_1_3_1_34_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"e_1_3_1_35_2","doi-asserted-by":"publisher","DOI":"10.1109\/MSP.2012.2211477"},{"key":"e_1_3_1_36_2","unstructured":"Jacob Devlin Ming-Wei Chang Kenton Lee and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. Retrieved from https:\/\/arxiv.org\/abs\/1810.04805"},{"key":"e_1_3_1_37_2","unstructured":"Mohammad Shoeybi Mostofa Patwary Raul Puri Patrick LeGresley Jared Casper and Bryan Catanzaro. 2020. Megatron-LM: Training multi-billion parameter language models using model parallelism. Retrieved from https:\/\/arxiv.org\/abs\/1909.08053"},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","unstructured":"Markus Dreseler Jan Kossmann Martin Boissier Stefan Klauck Matthias Uflacker and Hasso Plattner. 2019. Hyrise Re-engineered: An extensible database system for research in relational In-memory data management. DOI:10.5441\/002\/edbt.2019.28","DOI":"10.5441\/002\/edbt.2019.28"},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","DOI":"10.1145\/3620678.3624666"},{"key":"e_1_3_1_40_2","doi-asserted-by":"publisher","DOI":"10.1145\/1376616.1376712"},{"key":"e_1_3_1_41_2","first-page":"929","volume-title":"Proceedings of the 19th USENIX Symposium on Networked Systems Design and Implementation (NSDI\u201922)","author":"Annavaram Assaf Eisenman, Kiran Kumar Matam, Steven Ingram, Dheevatsa Mudigere, Raghuraman Krishnamoorthi, Krishnakumar Nair, Misha Smelyanskiy, and Murali","year":"2022","unstructured":"Assaf Eisenman, Kiran Kumar Matam, Steven Ingram, Dheevatsa Mudigere, Raghuraman Krishnamoorthi, Krishnakumar Nair, Misha Smelyanskiy, and Murali Annavaram. 2022. Check-N-Run: A checkpointing system for training deep learning recommendation models. In Proceedings of the 19th USENIX Symposium on Networked Systems Design and Implementation (NSDI\u201922). USENIX Association, 929\u2013943. Retrieved from https:\/\/www.usenix.org\/conference\/nsdi22\/presentation\/eisenman"},{"key":"e_1_3_1_42_2","unstructured":"Anton Lozhkov Raymond Li Loubna Ben Allal Federico Cassano Joel Lamy-Poirier Nouamane Tazi Ao Tang Dmytro Pykhtar Jiawei Liu Yuxiang Wei Tianyang Liu Max Tian Denis Kocetkov Arthur Zucker Younes Belkada Zijian Wang Qian Liu Dmitry Abulkhanov Indraneil Paul Zhuang Li Wen-Ding Li Megan Risdal Jia Li Jian Zhu Terry Yue Zhuo Evgenii Zheltonozhskii Nii Osae Osae Dade Wenhao Yu Lucas Krau\u00df Naman Jain Yixuan Su Xuanli He Manan Dey Edoardo Abati Yekun Chai Niklas Muennighoff Xiangru Tang Muhtasham Oblokulov Christopher Akiki Marc Marone Chenghao Mou Mayank Mishra Alex Gu Binyuan Hui Tri Dao Armel Zebaze Olivier Dehaene Nicolas Patry Canwen Xu Julian McAuley Han Hu Torsten Scholak Sebastien Paquet Jennifer Robinson Carolyn Jane Anderson Nicolas Chapados Mostofa Patwary Nima Tajbakhsh Yacine Jernite Carlos Mu\u00f1oz Ferrandis Lingming Zhang Sean Hughes Thomas Wolf Arjun Guha Leandro von Werra and Harm de Vries. 2024. StarCoder 2 and The stack v2: The next generation. Retrieved from https:\/\/arxiv.org\/abs\/2402.19173"},{"key":"e_1_3_1_43_2","doi-asserted-by":"publisher","DOI":"10.1109\/PDP52278.2021.00026"},{"issue":"1","key":"e_1_3_1_44_2","first-page":"1","article-title":"Data normalization and standardization: A technical report","volume":"1","author":"Faraj Ali, Peshawa Jamal Muhammad, Rezhna Hassan Faraj, Erbil Koya, Peshawa J. Muhammad Ali, and Rezhna H.","year":"2014","unstructured":"Ali, Peshawa Jamal Muhammad, Rezhna Hassan Faraj, Erbil Koya, Peshawa J. Muhammad Ali, and Rezhna H. Faraj. 2014. Data normalization and standardization: A technical report. Machine Learning Technical Reports 1, 1 (2014), 1\u20136.","journal-title":"Machine Learning Technical Reports"},{"key":"e_1_3_1_45_2","doi-asserted-by":"publisher","DOI":"10.1109\/PDSW54622.2021.00008"},{"key":"e_1_3_1_46_2","doi-asserted-by":"crossref","unstructured":"Chia-Yu Chen Jungwook Choi Daniel Brand Ankur Agrawal Wei Zhang and Kailash Gopalakrishnan. 2017. AdaComp: Adaptive residual gradient compression for data-parallel distributed training. Retrieved from https:\/\/arxiv.org\/abs\/1712.02679","DOI":"10.1609\/aaai.v32i1.11728"},{"key":"e_1_3_1_47_2","doi-asserted-by":"publisher","DOI":"10.1145\/3337821.3337902"},{"key":"e_1_3_1_48_2","doi-asserted-by":"publisher","DOI":"10.1109\/CLUSTER49012.2020.00046"},{"key":"e_1_3_1_49_2","doi-asserted-by":"publisher","DOI":"10.1109\/IISWC.2012.6402898"},{"key":"e_1_3_1_50_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA56546.2023.10070964"},{"key":"e_1_3_1_51_2","doi-asserted-by":"publisher","DOI":"10.1109\/CCGrid51090.2021.00018"},{"key":"e_1_3_1_52_2","doi-asserted-by":"publisher","DOI":"10.1145\/3458817.3476181"},{"key":"e_1_3_1_53_2","unstructured":"Daniel Smilkov Nikhil Thorat Yannick Assogba Charles Nicholson Nick Kreeger Ping Yu Shanqing Cai Eric Nielsen David Soegel Stan Bileschi Michael Terry Ann Yuan Kangyi Zhang Sandeep Gupta Sarah Sirajuddin D. Sculley Rajat Monga Greg Corrado Fernanda Viegas and Martin M. Wattenberg. 2019. TensorFlow.js: Machine learning for the web and beyond. In Proceedings of Machine Learning and Systems. 309\u2013321. Retrieved from https:\/\/proceedings.mlsys.org\/paper_files\/paper\/2019\/file\/acd593d2db87a799a8d3da5a860c028e-Paper.pdf"},{"key":"e_1_3_1_54_2","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPSW50202.2020.00174"},{"key":"e_1_3_1_55_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-84825-5_1"},{"key":"e_1_3_1_56_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.01156"},{"key":"e_1_3_1_57_2","unstructured":"Guilherme Penedo Quentin Malartic Daniel Hesslow Ruxandra Cojocaru Alessandro Cappelli Hamza Alobeidli Baptiste Pannier Ebtesam Almazrouei and Julien Launay. 2023. The RefinedWeb dataset for Falcon LLM: Outperforming curated corpora with web data and web data only. Retrieved from https:\/\/arxiv.org\/abs\/2306.01116"},{"key":"e_1_3_1_58_2","doi-asserted-by":"publisher","DOI":"10.1109\/DSD57027.2022.00048"},{"key":"e_1_3_1_59_2","doi-asserted-by":"publisher","DOI":"10.1145\/3394486.3406477"},{"key":"e_1_3_1_60_2","doi-asserted-by":"publisher","DOI":"10.1145\/3624062.3624171"},{"key":"e_1_3_1_61_2","first-page":"2320","volume-title":"Proceedings of the 33rd International Conference on Machine Learning","author":"Shmakov Liberty, Edo, Kevin Lang, Konstantin","year":"2016","unstructured":"Liberty, Edo, Kevin Lang, Konstantin Shmakov. 2016. Stratified sampling meets machine learning. In Proceedings of the 33rd International Conference on Machine Learning. Proceedings of Machine Learning Research, Vol. 48, PMLR, New York, New York, USA, 2320\u20132329. Retrieved from https:\/\/proceedings.mlr.press\/v48\/liberty16.html"},{"key":"e_1_3_1_62_2","doi-asserted-by":"publisher","DOI":"10.1109\/HiPC53243.2021.00046"},{"key":"e_1_3_1_63_2","doi-asserted-by":"publisher","DOI":"10.1145\/2493123.2462919"},{"key":"e_1_3_1_64_2","doi-asserted-by":"publisher","DOI":"10.1109\/SC.2018.00068"},{"key":"e_1_3_1_65_2","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2020.2974843"},{"key":"e_1_3_1_66_2","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2019.00099"},{"key":"e_1_3_1_67_2","doi-asserted-by":"publisher","DOI":"10.1109\/CCGrid49817.2020.00-76"},{"key":"e_1_3_1_68_2","doi-asserted-by":"publisher","unstructured":"Norman P. Jouppi Cliff Young Nishant Patil David Patterson Gaurav Agrawal Raminder Bajwa Sarah Bates Suresh Bhatia Nan Boden Al Borchers Rick Boyle Pierre-luc Cantin Clifford Chao Chris Clark Jeremy Coriell Mike Daley Matt Dau Jeffrey Dean Ben Gelb Tara Vazir Ghaemmaghami Rajendra Gottipati William Gulland Robert Hagmann C. Richard Ho Doug Hogberg John Hu Robert Hundt Dan Hurt Julian Ibarz Aaron Jaffey Alek Jaworski Alexander Kaplan Harshit Khaitan Daniel Killebrew Andy Koch Naveen Kumar Steve Lacy James Laudon James Law Diemthu Le Chris Leary Zhuyuan Liu Kyle Lucke Alan Lundin Gordon MacKean Adriana Maggiore Maire Mahony Kieran Miller Rahul Nagarajan Ravi Narayanaswami Ray Ni Kathy Nix Thomas Norrie Mark Omernick Narayana Penukonda Andy Phelps Jonathan Ross Matt Ross Amir Salek Emad Samadiani Chris Severn Gregory Sizikov Matthew Snelham Jed Souter Dan Steinberg Andy Swing Mercedes Tan Gregory Thorson Bo Tian Horia Toma Erick Tuttle Vijay Vasudevan Richard Walter Walter Wang Eric Wilcox and Doe Hyun Yoon. 2017. In-Datacenter performance analysis of a tensor processing unit. In Proceedings of the 44th Annual International Symposium on Computer Architecture (ISCA\u201917). Association for Computing Machinery Toronto ON Canada 1\u201312. DOI:10.1145\/3079856.3080246","DOI":"10.1145\/3079856.3080246"},{"key":"e_1_3_1_69_2","doi-asserted-by":"publisher","unstructured":"Truong Thao Nguyen Fran\u00e7ois Trahay Jens Domke Aleksandr Drozd Emil Vatai Jianwei Liao Mohamed Wahib and Balazs Gerofi. 2022. Why Globally Re-shuffle? Revisiting data shuffling in large scale deep learning. In 2022 IEEE International Parallel and Distributed Processing Symposium (IPDPS). 1085\u20131096. DOI:10.1109\/IPDPS53621.2022.00109","DOI":"10.1109\/IPDPS53621.2022.00109"},{"key":"e_1_3_1_70_2","doi-asserted-by":"publisher","DOI":"10.1109\/HiPC50609.2020.00034"},{"key":"e_1_3_1_71_2","doi-asserted-by":"publisher","DOI":"10.1109\/MASCOTS53633.2021.9614303"},{"key":"e_1_3_1_72_2","unstructured":"Pablo Villalobos Jaime Sevilla Tamay Besiroglu Lennart Heim Anson Ho and Marius Hobbhahn. 2022. Machine learning model sizes and the parameter gap. Retrieved from https:\/\/arxiv.org\/abs\/2207.02852"},{"key":"e_1_3_1_73_2","first-page":"1","article-title":"Fathom: Reference workloads for modern deep learning methods","author":"Brooks Robert Adolf, Saketh Rama, Brandon Reagen, Gu-Yeon Wei, and David","year":"2016","unstructured":"Robert Adolf, Saketh Rama, Brandon Reagen, Gu-Yeon Wei, and David Brooks. 2016. Fathom: Reference workloads for modern deep learning methods. In Proceedings of the 2016 IEEE International Symposium on Workload Characterization (IISWC\u201916). 1\u201310. Retrieved from https:\/\/api.semanticscholar.org\/CorpusID:1809844","journal-title":"Proceedings of the 2016 IEEE International Symposium on Workload Characterization (IISWC\u201916)"},{"key":"e_1_3_1_74_2","doi-asserted-by":"publisher","DOI":"10.1145\/3394486.3406703"},{"key":"e_1_3_1_75_2","doi-asserted-by":"crossref","unstructured":"Steven Farrell Murali Emani Jacob Balma Lukas Drescher Aleksandr Drozd Andreas Fink Geoffrey Fox David Kanter Thorsten Kurth Peter Mattson Dawei Mu Amit Ruhela Kento Sato Koichi Shirahata Tsuguchika Tabaru Aristeidis Tsaris Jan Balewski Ben Cumming Takumi Danjo Jens Domke Takaaki Fukai Naoto Fukumoto Tatsuya Fukushi Balazs Gerofi Takumi Honda Toshiyuki Imamura Akihiko Kasagi Kentaro Kawakami Shuhei Kudo Akiyoshi Kuroda Maxime Martinasso Satoshi Matsuoka Henrique Mendon\u00e7a Kazuki Minami Prabhat Ram Takashi Sawada Mallikarjun Shankar Tom St. John Akihiro Tabuchi Venkatram Vishwanath Mohamed Wahib Masafumi Yamazaki and Junqi Yin. 2021. MLPerf HPC: A holistic benchmark suite for scientific machine learning on HPC systems. Retrieved from https:\/\/arxiv.org\/abs\/2110.11466","DOI":"10.1109\/MLHPC54614.2021.00009"},{"key":"e_1_3_1_76_2","doi-asserted-by":"publisher","DOI":"10.1109\/ESPT.2016.006"},{"key":"e_1_3_1_77_2","volume-title":"Proceedings of the 37th International Conference on Machine Learning","volume":"119","author":"De Smith, Samuel, Erich Elsen, and Soham","year":"2020","unstructured":"Smith, Samuel, Erich Elsen, and Soham De. 2020. On the generalization benefit of noise in stochastic gradient descent. In Proceedings of the 37th International Conference on Machine Learning. Hal Daum\u00e9 III and Aarti Singh (Eds.), Vol. 119, PMLR, Cornell University. Retrieved from https:\/\/proceedings.mlr.press\/v119\/smith20a.html"},{"key":"e_1_3_1_78_2","doi-asserted-by":"publisher","DOI":"10.1109\/SC41404.2022.00076"},{"key":"e_1_3_1_79_2","unstructured":"Tom B. Brown Benjamin Mann Nick Ryder Melanie Subbiah Jared Kaplan Prafulla Dhariwal Arvind Neelakantan Pranav Shyam Girish Sastry Amanda Askell Sandhini Agarwal Ariel Herbert-Voss Gretchen Krueger Tom Henighan Rewon Child Aditya Ramesh Daniel M. Ziegler Jeffrey Wu Clemens Winter Christopher Hesse Mark Chen Eric Sigler Mateusz Litwin Scott Gray Benjamin Chess Jack Clark Christopher Berner Sam McCandlish Alec Radford Ilya Sutskever and Dario Amodei. 2020. Language models are few-shot learners. In Proceedings of the 34th International Conference on Neural Information Processing Systems (NIPS\u201920) Curran Associates Inc. Vancouver BC Canada."},{"key":"e_1_3_1_80_2","unstructured":"Vijay Janapa Reddi Christine Cheng David Kanter Peter Mattson Guenther Schmuelling Carole-Jean Wu Brian Anderson Maximilien Breughe Mark Charlebois William Chou Ramesh Chukka Cody Coleman Sam Davis Pan Deng Greg Diamos Jared Duke Dave Fick J. Scott Gardner Itay Hubara Sachin Idgunji Thomas B. Jablin Jeff Jiao Tom St. John Pankaj Kanwar David Lee Jeffery Liao Anton Lokhmotov Francisco Massa Peng Meng Paulius Micikevicius Colin Osborne Gennady Pekhimenko Arun Tejusve Raghunath Rajan Dilip Sequeira Ashish Sirasao Fei Sun Hanlin Tang Michael Thomson Frank Wei Ephrem Wu Lingjie Xu Koichi Yamada Bing Yu George Yuan Aaron Zhong Peizhao Zhang and Yuchen Zhou. 2020. MLPerf inference benchmark. Retrieved from https:\/\/arxiv.org\/abs\/1911.02549"},{"key":"e_1_3_1_81_2","unstructured":"Weihua Hu Matthias Fey Marinka Zitnik Yuxiao Dong Hongyu Ren Bowen Liu Michele Catasta and Jure Leskovec. 2020. Open graph benchmark: Datasets for machine learning on graphs. In Proceedings of the 34th International Conference on Neural Information Processing Systems (NIPS\u201920). Curran Associates Inc. Vancouver BC Canada."},{"key":"e_1_3_1_82_2","first-page":"10","volume-title":"CUG2020","author":"al Wan, Lipeng, Matthew Wolf, Feiyi Wang, Jong Youl Choi, George Ostrouchov, Jieyang Chen, Norbert Podhorszki, Jeremy Logan, Kshitij Mehta, Scott Klasky, et","year":"2020","unstructured":"Wan, Lipeng, Matthew Wolf, Feiyi Wang, Jong Youl Choi, George Ostrouchov, Jieyang Chen, Norbert Podhorszki, Jeremy Logan, Kshitij Mehta, Scott Klasky, et al. 2020. I\/O performance characterization and prediction through machine learning on HPC systems. In CUG2020. CUG, 10."},{"key":"e_1_3_1_83_2","first-page":"17146","volume-title":"Advances in Neural Information Processing Systems","author":"Tong Xie, Jian, Jingwei Xu, Guochang Wang, Yuan Yao, Zenan Li, Chun Cao, Hanghang","year":"2022","unstructured":"Xie, Jian, Jingwei Xu, Guochang Wang, Yuan Yao, Zenan Li, Chun Cao, Hanghang Tong. 2022. A deep learning dataloader with shared data preparation. In Advances in Neural Information Processing Systems. S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh (Eds.), Vol. 35, Curran Associates, Inc., Virtual, 17146\u201317156."},{"key":"e_1_3_1_84_2","doi-asserted-by":"publisher","DOI":"10.1145\/3384419.3430898"},{"key":"e_1_3_1_85_2","doi-asserted-by":"publisher","DOI":"10.1109\/TVLSI.2018.2791442"},{"key":"e_1_3_1_86_2","doi-asserted-by":"publisher","DOI":"10.1109\/MASCOTS.2018.00023"},{"key":"e_1_3_1_87_2","unstructured":"Facebook Incubator. 2024. Gloo. Retrieved March 20 2025 from https:\/\/github.com\/facebookincubator\/gloo"},{"key":"e_1_3_1_88_2","unstructured":"Roy Fielding. 2000. Architectural styles and the design of network-based software architectures. phdthesis."},{"key":"e_1_3_1_89_2","volume-title":"Apache Parquet","author":"Foundation Apache Software","year":"2024","unstructured":"Apache Software Foundation. 2024. Apache Parquet. Apache Software Foundation. Retrieved from https:\/\/parquet.apache.org\/. Version 2.10.0."},{"key":"e_1_3_1_90_2","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2015.116"},{"key":"e_1_3_1_91_2","doi-asserted-by":"crossref","unstructured":"Ana Gainaru Dmitry Ganyushin Bing Xie Tahsin Kurc Joel Saltz Sarp Oral Norbert Podhorszki Franz Poeschel Axel Huebl and Scott Klasky. 2022. Understanding and leveraging the I\/O patterns of emerging machine learning analytics. In Driving Scientific and Engineering Discoveries Through the Integration of Experiment Big Data and Modeling and Simulation Springer International Publishing Cham 119\u2013138.","DOI":"10.1007\/978-3-030-96498-6_7"},{"key":"e_1_3_1_92_2","unstructured":"Leo Gao Stella Biderman Sid Black Laurence Golding Travis Hoppe Charles Foster Jason Phang Horace He Anish Thite Noa Nabeshima Shawn Presser and Connor Leahy. 2020. The Pile: An 800GB dataset of diverse text for language modeling. Retrieved from https:\/\/arxiv.org\/abs\/2101.00027"},{"key":"e_1_3_1_93_2","doi-asserted-by":"publisher","DOI":"10.5555\/1753228.1753234"},{"key":"e_1_3_1_94_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2017.7952261"},{"key":"e_1_3_1_95_2","doi-asserted-by":"publisher","unstructured":"Anjus George Rick Mohr James Simmons and Sarp Oral. 2021. Understanding Lustre Internals. Second Edition. Oak Ridge National Laboratory (ORNL) Oak Ridge TN (United States). DOI:10.2172\/1824954","DOI":"10.2172\/1824954"},{"key":"e_1_3_1_96_2","unstructured":"Google Cloud. 2024. Google Analytics Data on BigQuery. Retrieved September 12 2024 from https:\/\/cloud.google.com\/bigquery\/public-data\/google-analytics"},{"key":"e_1_3_1_97_2","unstructured":"Google Cloud. 2024. Google BigQuery Public Datasets. Retrieved September 12 2024 from https:\/\/cloud.google.com\/bigquery\/public-data"},{"key":"e_1_3_1_98_2","volume-title":"Google Cloud Storage","author":"Cloud Google","unstructured":"Google Cloud. n.d.. Google Cloud Storage. Retrieved from https:\/\/cloud.google.com\/storage"},{"key":"e_1_3_1_99_2","volume-title":"Enhancing Data Movement and Access for GPUs","author":"GPUDirect NVIDIA","year":"2024","unstructured":"NVIDIA GPUDirect. 2024. Enhancing Data Movement and Access for GPUs. Retrieved February 26, 2024 from https:\/\/developer.nvidia.com\/gpudirect"},{"key":"e_1_3_1_100_2","unstructured":"Sasun Hambardzumyan Abhinav Tuli Levon Ghukasyan Fariz Rahman Hrant Topchyan David Isayan Mark McQuade Mikayel Harutyunyan Tatevik Hakobyan Ivo Stranic and Davit Buniatyan. 2022. Deep lake: A lakehouse for deep learning. Retrieved from https:\/\/arxiv.org\/abs\/2209.10785"},{"key":"e_1_3_1_101_2","doi-asserted-by":"publisher","unstructured":"Kaveen Hiniduma Suren Byna and Jean Luca Bez. 2025. Data readiness for AI: A 360-degree survey. ACM Comput. Surv. (March 2025). DOI:10.1145\/3722214","DOI":"10.1145\/3722214"},{"key":"e_1_3_1_102_2","doi-asserted-by":"crossref","unstructured":"Hongsun Jang Jaeyong Song Jaewon Jung Jaeyoung Park Youngsok Kim and Jinho Lee. 2024. Smart-infinity: fast large language model training using near-storage processing on a real system. Retrieved from https:\/\/arxiv.org\/abs\/2403.06664","DOI":"10.1109\/HPCA57654.2024.00034"},{"key":"e_1_3_1_103_2","doi-asserted-by":"publisher","DOI":"10.1109\/COMST.2014.2354398"},{"key":"e_1_3_1_104_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICKG59574.2023.00014"},{"key":"e_1_3_1_105_2","doi-asserted-by":"publisher","DOI":"10.1109\/TC.2023.3336161"},{"key":"e_1_3_1_106_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10115-008-0142-6"},{"key":"e_1_3_1_107_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2021.3110745"},{"key":"e_1_3_1_108_2","doi-asserted-by":"publisher","DOI":"10.1109\/eScience.2016.7870922"},{"key":"e_1_3_1_109_2","unstructured":"Alex Krizhevsky Geoffrey Hinton and others. 2009. Learning multiple layers of features from tiny images. (2009)."},{"key":"e_1_3_1_110_2","doi-asserted-by":"publisher","DOI":"10.14778\/3476249.3476308"},{"key":"e_1_3_1_111_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2011.6126543"},{"key":"e_1_3_1_112_2","doi-asserted-by":"publisher","unstructured":"Alexander Lavin and Subutai Ahmad. 2015. Evaluating real-time anomaly detection algorithms \u2013 the numenta anomaly benchmark. In 2015 IEEE 14th International Conference on Machine Learning and Applications (ICMLA) IEEE. DOI:10.1109\/icmla.2015.141","DOI":"10.1109\/icmla.2015.141"},{"key":"e_1_3_1_113_2","unstructured":"Choonghwan Lee Muqun Yang and Ruth A. Aydt. 2008. NetCDF-4 performance report. Retrieved from https:\/\/api.semanticscholar.org\/CorpusID:15588867"},{"key":"e_1_3_1_114_2","doi-asserted-by":"publisher","DOI":"10.1109\/CCGrid57682.2023.00017"},{"key":"e_1_3_1_115_2","doi-asserted-by":"publisher","DOI":"10.3390\/app132212102"},{"key":"e_1_3_1_116_2","doi-asserted-by":"publisher","DOI":"10.1109\/TWC.2019.2946140"},{"key":"e_1_3_1_117_2","doi-asserted-by":"publisher","DOI":"10.1109\/BIGCOM61073.2023.00036"},{"key":"e_1_3_1_118_2","doi-asserted-by":"publisher","DOI":"10.1109\/JIOT.2020.2970110"},{"key":"e_1_3_1_119_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.ijforecast.2019.04.014"},{"key":"e_1_3_1_120_2","unstructured":"Priyanka Mary Mammen. 2021. Federated Learning: Opportunities and challenges. Retrieved from https:\/\/arxiv.org\/abs\/2101.05428"},{"key":"e_1_3_1_121_2","doi-asserted-by":"publisher","DOI":"10.1145\/3625549.3658685"},{"key":"e_1_3_1_122_2","doi-asserted-by":"publisher","unstructured":"Andrew McCallum. 2017. Cora Dataset. Texas Data Repository. DOI:10.18738\/T8\/HUIG48","DOI":"10.18738\/T8\/HUIG48"},{"key":"e_1_3_1_123_2","unstructured":"H. Brendan McMahan Eider Moore Daniel Ramage Seth Hampson and Blaise Ag\u00fcera y Arcas. 2023. Communication-Efficient Learning of Deep Networks from Decentralized Data. Retrieved from https:\/\/arxiv.org\/abs\/1602.05629"},{"key":"e_1_3_1_124_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2019.01.037"},{"key":"e_1_3_1_125_2","volume-title":"Proceedings of the 30th International Conference on Machine Learning","volume":"28","author":"Meng Xiangrui","year":"2013","unstructured":"Xiangrui Meng. 2013. Scalable simple random sampling and stratified sampling. In Proceedings of the 30th International Conference on Machine Learning. Sanjoy Dasgupta and David McAllester (Eds.), Vol. 28, PMLR. Retrieved from https:\/\/proceedings.mlr.press\/v28\/meng13a.html"},{"key":"e_1_3_1_126_2","volume-title":"Open Database Connectivity (ODBC)","year":"2024","unstructured":"Microsoft 2024. Open Database Connectivity (ODBC). Microsoft. Retrieved March 20, 2025 from https:\/\/docs.microsoft.com\/en-us\/sql\/odbc\/reference\/odbc"},{"key":"e_1_3_1_127_2","unstructured":"Microsoft Azure. 2024. Azure Open Datasets. Retrieved September 12 2024 from https:\/\/azure.microsoft.com\/en-us\/services\/open-datasets\/"},{"key":"e_1_3_1_128_2","unstructured":"Microsoft Azure. 2024. NYC Taxi & Limousine Commission Data. Retrieved September 12 2024 from https:\/\/azure.microsoft.com\/en-us\/services\/open-datasets\/nyc-taxi-limousine\/"},{"key":"e_1_3_1_129_2","unstructured":"Derek G. Murray Jiri Simsa Ana Klimovic and Ihor Indyk. 2021. tf.data: A machine learning data processing framework. Retrieved from https:\/\/arxiv.org\/abs\/2101.12127"},{"key":"e_1_3_1_130_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2009.05.017"},{"key":"e_1_3_1_131_2","doi-asserted-by":"publisher","DOI":"10.1109\/WSC.2018.8632351"},{"key":"e_1_3_1_132_2","volume-title":"GPUDirect RDMA: Direct Communcation between NVIDIA GPUs","year":"2024","unstructured":"NVIDIA. 2024. GPUDirect RDMA: Direct Communcation between NVIDIA GPUs. Retrieved February 26, 2024 from https:\/\/developer.nvidia.com\/blog\/gpudirect-storage"},{"key":"e_1_3_1_133_2","volume-title":"GPUDirect Storage: A Direct Path Between Storage and GPU Memory","year":"2024","unstructured":"NVIDIA. 2024. GPUDirect Storage: A Direct Path Between Storage and GPU Memory. Retrieved February 26, 2024 from https:\/\/docs.nvidia.com\/cuda\/gpudirect-rdma\/index.html"},{"key":"e_1_3_1_134_2","volume-title":"NVIDIA DALI","author":"Corporation NVIDIA","year":"2024","unstructured":"NVIDIA Corporation. 2024. NVIDIA DALI. Retrieved March 20, 2025 from https:\/\/docs.nvidia.com\/deeplearning\/dali\/user-guide\/docs\/index.html"},{"key":"e_1_3_1_135_2","doi-asserted-by":"publisher","DOI":"10.1109\/IEMECONX.2019.8877011"},{"key":"e_1_3_1_136_2","article-title":"NumPy: A guide to NumPy","author":"Oliphant Travis E.","year":"2006","unstructured":"Travis E. Oliphant. 2006. NumPy: A guide to NumPy. Trelgol Publishing. Retrieved from https:\/\/numpy.org\/","journal-title":"Trelgol Publishing"},{"key":"e_1_3_1_137_2","doi-asserted-by":"crossref","unstructured":"Vassil Panayotov Guoguo Chen Daniel Povey and Sanjeev Khudanpur. 2015. LibriSpeech: An ASR Corpus based on Public Domain Audio Books. Retrieved February 26 2024 from http:\/\/www.openslr.org\/12","DOI":"10.1109\/ICASSP.2015.7178964"},{"key":"e_1_3_1_138_2","volume-title":"Inside the Lustre File System","author":"Petersen Torben Kling","year":"2015","unstructured":"Torben Kling Petersen. 2015. Inside the Lustre File System. Technical Paper. SEAGATE Technology."},{"key":"e_1_3_1_139_2","volume-title":"PyTorch 2.2.0 Documentation","year":"2024","unstructured":"Pytorch. 2024. PyTorch 2.2.0 Documentation. PyTorch. https:\/\/pytorch.org\/docs\/stable\/index.html. Accessed on February 1, 2024."},{"key":"e_1_3_1_140_2","doi-asserted-by":"crossref","unstructured":"Samyam Rajbhandari Olatunji Ruwase Jeff Rasley Shaden Smith and Yuxiong He. 2021. ZeRO-Infinity: Breaking the GPU memory wall for extreme scale deep learning. Retrieved from https:\/\/arxiv.org\/abs\/2104.07857","DOI":"10.1145\/3458817.3476205"},{"key":"e_1_3_1_141_2","doi-asserted-by":"publisher","DOI":"10.1017\/CBO9780511812651"},{"key":"e_1_3_1_142_2","unstructured":"Elvis Rojas Albert Njoroge Kahira Esteban Meneses Leonardo Bautista Gomez and Rosa M. Badia. 2021. A study of checkpointing in large scale training of deep neural networks. Retrieved from https:\/\/arxiv.org\/abs\/2012.00825"},{"key":"e_1_3_1_143_2","unstructured":"Sebastian Ruder. 2017. An Overview of Gradient Descent Optimization Algorithms. Retrieved from https:\/\/arxiv.org\/abs\/1609.04747"},{"key":"e_1_3_1_144_2","doi-asserted-by":"publisher","DOI":"10.1145\/2647868.2655045"},{"key":"e_1_3_1_145_2","doi-asserted-by":"publisher","DOI":"10.1109\/Cluster48925.2021.00062"},{"key":"e_1_3_1_146_2","volume-title":"scikit-learn: Machine Learning in Python","year":"2024","unstructured":"Scikit-learn. 2024. scikit-learn: Machine Learning in Python. https:\/\/scikit-learn.org\/stable\/. Accessed on February 5, 2024."},{"key":"e_1_3_1_147_2","unstructured":"Alexander Sergeev and Mike Del Balso. 2018. Horovod: fast and easy distributed deep learning in TensorFlow. Retrieved from https:\/\/arxiv.org\/abs\/1802.05799"},{"key":"e_1_3_1_148_2","doi-asserted-by":"publisher","DOI":"10.1145\/3365109.3368768"},{"key":"e_1_3_1_149_2","unstructured":"Mohammad Shoeybi Mostofa Patwary Raul Puri Patrick LeGresley Jared Casper and Bryan Catanzaro. 2020. Megatron-LM: Training multi-billion parameter language models using model parallelism. Retrieved from https:\/\/arxiv.org\/abs\/1909.08053"},{"key":"e_1_3_1_150_2","doi-asserted-by":"publisher","DOI":"10.1186\/s40537-019-0197-0"},{"key":"e_1_3_1_151_2","doi-asserted-by":"publisher","DOI":"10.1186\/s40537-021-00492-0"},{"key":"e_1_3_1_152_2","doi-asserted-by":"crossref","unstructured":"Gunnar A. Sigurdsson G\u00fcl Varol Xiaolong Wang Ali Farhadi Ivan Laptev and Abhinav Gupta. 2016. Hollywood in homes: Crowdsourcing data collection for activity understanding. Retrieved from https:\/\/arxiv.org\/abs\/1604.01753","DOI":"10.1007\/978-3-319-46448-0_31"},{"key":"e_1_3_1_153_2","doi-asserted-by":"publisher","DOI":"10.5555\/3018823.3018825"},{"key":"e_1_3_1_154_2","doi-asserted-by":"crossref","unstructured":"Xinying Song Alex Salcianu Yang Song Dave Dopson and Denny Zhou. 2021. Fast WordPiece tokenization. Retrieved from https:\/\/arxiv.org\/abs\/2012.15524","DOI":"10.18653\/v1\/2021.emnlp-main.160"},{"key":"e_1_3_1_155_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICDAR.2005.251"},{"key":"e_1_3_1_156_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.compchemeng.2009.10.007"},{"key":"e_1_3_1_157_2","volume-title":"High-Performance Storage Architecture and Scalable Cluster File System","year":"2007","unstructured":"SUN. 2007. High-Performance Storage Architecture and Scalable Cluster File System. Technical Report. Sun Microsystems, Inc."},{"key":"e_1_3_1_158_2","unstructured":"Ivan Svogor Christian Eichenberger Markus Spanring Moritz Neun and Michael Kopp. 2022. Profiling and Improving the PyTorch Dataloader for high-latency Storage: A Technical Report. Retrieved from https:\/\/arxiv.org\/abs\/2211.04908"},{"key":"e_1_3_1_159_2","unstructured":"Vivienne Sze Yu-Hsin Chen Tien-Ju Yang and Joel Emer. 2017. Efficient processing of deep neural networks: A tutorial and survey. Retrieved from https:\/\/arxiv.org\/abs\/1703.09039"},{"key":"e_1_3_1_160_2","volume-title":"TensorFlow 2.15.0 Documentation","year":"2024","unstructured":"TensorFlow. 2024. TensorFlow 2.15.0 Documentation. TensorFlow. Retrieved February 5, 2024 from https:\/\/www.tensorflow.org\/versions\/r2.15\/api_docs"},{"key":"e_1_3_1_161_2","unstructured":"TensorFlow. 2024. TensorFlow Hub. Retrieved March 20 2025 from https:\/\/www.tensorflow.org\/hub"},{"key":"e_1_3_1_162_2","unstructured":"TensorFlow. 2024. TensorFlow Lite. Retrieved March 20 2025 from https:\/\/www.tensorflow.org\/lite"},{"key":"e_1_3_1_163_2","volume-title":"TFRecord and tf.train.Example","year":"2024","unstructured":"TensorFlow. 2024. TFRecord and tf.train.Example. TensorFlow. Retrieved February 19, 2024 from https:\/\/www.tensorflow.org\/tutorials\/load_data\/tfrecord"},{"key":"e_1_3_1_164_2","doi-asserted-by":"publisher","unstructured":"Rajeev Thakur. 2024. AuroraGPT: A large-scale foundation model for advancing science. Zenodo. DOI:10.5281\/zenodo.13345059","DOI":"10.5281\/zenodo.13345059"},{"key":"e_1_3_1_165_2","volume-title":"Hierarchical Data Format, version 5","author":"Group The HDF","unstructured":"The HDF Group. 1997-2022. Hierarchical Data Format, version 5. The HDF Group. Retrieved from https:\/\/www.hdfgroup.org\/HDF5\/"},{"key":"e_1_3_1_166_2","unstructured":"THUMOS Challenge Workshop. 2012. UFC-101 Dataset. Retrieved February 26 2024 from http:\/\/www.thumos.info\/download.html"},{"key":"e_1_3_1_167_2","doi-asserted-by":"publisher","DOI":"10.1016\/S0164-1212(99)00062-X"},{"key":"e_1_3_1_168_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2021.3104252"},{"key":"e_1_3_1_169_2","doi-asserted-by":"publisher","DOI":"10.1109\/PDSW54622.2021.00009"},{"key":"e_1_3_1_170_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.cosrev.2023.100595"},{"key":"e_1_3_1_171_2","doi-asserted-by":"publisher","unstructured":"Jaewon Yang and Jure Leskovec. 2012. Defining and evaluating network communities based on ground-truth. In Proceedings of the ACM SIGKDD Workshop on Mining Data Semantics (MDS\u201912) Association for Computing Machinery Beijing China. DOI:10.1145\/2350190.2350193","DOI":"10.1145\/2350190.2350193"},{"key":"e_1_3_1_172_2","doi-asserted-by":"publisher","unstructured":"Benbo Zha and Hong Shen. 2022. Adaptively periodic I\/O scheduling for Concurrent HPC Applications. Electronics 11 9 (2022). DOI:10.3390\/electronics11091318","DOI":"10.3390\/electronics11091318"},{"key":"e_1_3_1_173_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.11"},{"key":"e_1_3_1_174_2","doi-asserted-by":"crossref","unstructured":"\u00d6zg\u00fcn \u00c7i\u00e7ek Ahmed Abdulkadir Soeren S. Lienkamp Thomas Brox and Olaf Ronneberger. 2016. 3D U-Net: Learning dense volumetric segmentation from sparse annotation. Retrieved from https:\/\/arxiv.org\/abs\/1606.06650","DOI":"10.1007\/978-3-319-46723-8_49"}],"container-title":["ACM Computing Surveys"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3722215","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3722215","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T18:43:51Z","timestamp":1750272231000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3722215"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,5,7]]},"references-count":173,"journal-issue":{"issue":"10","published-print":{"date-parts":[[2025,10,31]]}},"alternative-id":["10.1145\/3722215"],"URL":"https:\/\/doi.org\/10.1145\/3722215","relation":{},"ISSN":["0360-0300","1557-7341"],"issn-type":[{"value":"0360-0300","type":"print"},{"value":"1557-7341","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,5,7]]},"assertion":[{"value":"2024-04-16","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-02-19","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-05-07","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}