{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,28]],"date-time":"2026-07-28T03:07:02Z","timestamp":1785208022617,"version":"3.55.0"},"reference-count":61,"publisher":"Association for Computing Machinery (ACM)","issue":"1","funder":[{"name":"Iran National Science Foundation","award":["4025937"],"award-info":[{"award-number":["4025937"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["Proc. ACM Meas. Anal. Comput. Syst."],"published-print":{"date-parts":[[2026,3,26]]},"abstract":"<jats:p>\n                    In modern multi-accelerator nodes, GPU throughput is increasingly constrained by storage and I\/O bottlenecks, leaving accelerators idle as data transfer is restricted by software. In this study, we focus on single-node, multi-GPU systems with small-to-medium-scale models and present a systematic, phase-aware evaluation of datapath performance for\n                    <jats:italic toggle=\"yes\">Large Language Models<\/jats:italic>\n                    (LLMs), covering technologies such as in-kernel\n                    <jats:italic toggle=\"yes\">libaio,<\/jats:italic>\n                    hybrid user-kernel\n                    <jats:italic toggle=\"yes\">io_uring,<\/jats:italic>\n                    user-space NVMe via the\n                    <jats:italic toggle=\"yes\">Storage Performance Development Kit<\/jats:italic>\n                    (SPDK), and\n                    <jats:italic toggle=\"yes\">GPUDirect Storage<\/jats:italic>\n                    (GDS). These approaches are evaluated across various storage media including SATA Solid State Drives (SSDs), NVMe SSDs, Optane NVMe, and\n                    <jats:italic toggle=\"yes\">Optane Persistent Memory<\/jats:italic>\n                    (PMem). Leveraging an automated evaluation framework, we explore over 25,000 configurations, measuring throughput, latency,\n                    <jats:italic toggle=\"yes\">I\/O per second<\/jats:italic>\n                    (IOPS), and CPU cost.\n                  <\/jats:p>\n                  <jats:p>\n                    Our study offers LLM storage scenarios in both standardized benchmarks and real-world production traces, ensuring that our workload models accurately reflect the I\/O demands across pre-training, fine-tuning, and inference. We find that for inference,\n                    <jats:italic toggle=\"yes\">io_uring<\/jats:italic>\n                    achieves the lowest latency and competitive IOPS for small random I\/O on NVMe. In contrast, SPDK is limited to raw block-device evaluation due to its lack of POSIX file-system support. For pre-training and fine-tuning, workloads are dominated by coarse-grained sequential reads and writes, where GDS excels in reducing load times and host CPU usage. Among CPU-mediated datapaths, CPU efficiency\u2014measured as GB\/s per core\u2014emerges as the key differentiator.\n                  <\/jats:p>\n                  <jats:p>\n                    Taken together, these results yield actionable design guidelines: align the choice of datapath with the LLM pipeline phase. Use\n                    <jats:italic toggle=\"yes\">io_uring<\/jats:italic>\n                    for inference to optimize data transfer efficiency and minimize latency, and leverage GDS for pre-training and fine-tuning to improve throughput per core, thereby narrowing the storage-to-compute gap in GPU LLM clusters.\n                  <\/jats:p>","DOI":"10.1145\/3788106","type":"journal-article","created":{"date-parts":[[2026,3,26]],"date-time":"2026-03-26T18:49:47Z","timestamp":1774550987000},"page":"1-26","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":2,"title":["Towards Scalable Storage Architectures for GPU Clusters Running Large Language Models"],"prefix":"10.1145","volume":"10","author":[{"ORCID":"https:\/\/orcid.org\/0009-0007-4698-3486","authenticated-orcid":false,"given":"Ali","family":"Sedaghatgoo","sequence":"first","affiliation":[{"name":"Sharif University of Technology, Tehran, Tehran, Iran"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3786-7102","authenticated-orcid":false,"given":"Reza","family":"Salkhordeh","sequence":"additional","affiliation":[{"name":"Johannes Gutenberg-Universit\u00e4t Mainz, Mainz, Rhineland Palatinate, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3083-2775","authenticated-orcid":false,"given":"Andr\u00e9","family":"Brinkmann","sequence":"additional","affiliation":[{"name":"Saarland University, Saarbr\u00fccken, Saarland, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0264-3865","authenticated-orcid":false,"given":"Hossein","family":"Asadi","sequence":"additional","affiliation":[{"name":"Sharif University of Technology, Tehran, Tehran, Iran"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,3,26]]},"reference":[{"key":"e_1_2_1_1_1","volume-title":"TensorFlow: A System for Large-Scale Machine Learning. In 12th USENIX Symposium on Operating Systems Design and Implementation, OSDI 2016","author":"Abadi Mart\u00edn","year":"2016","unstructured":"Mart\u00edn Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, Manjunath Kudlur, Josh Levenberg, Rajat Monga, Sherry Moore, Derek Gordon Murray, Benoit Steiner, Paul A. Tucker, Vijay Vasudevan, Pete Warden, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng. 2016. TensorFlow: A System for Large-Scale Machine Learning. In 12th USENIX Symposium on Operating Systems Design and Implementation, OSDI 2016, Savannah, GA, USA, November 2-4, 2016, Kimberly Keeton and Timothy Roscoe (Eds.). USENIX Association, 265-283. https:\/\/www.usenix.org\/conference\/osdi16\/technical-sessions\/presentation\/abadi."},{"key":"e_1_2_1_2_1","volume-title":"Unsloth: Fast and Memory-Efficient Fine-Tuning of Large Language Models. https:\/\/github.com\/unslothai\/unsloth.","author":"Unsloth","year":"2024","unstructured":"Unsloth AI. 2024. Unsloth: Fast and Memory-Efficient Fine-Tuning of Large Language Models. https:\/\/github.com\/unslothai\/unsloth."},{"key":"e_1_2_1_3_1","first-page":"67","volume-title":"ACM SIGMETRICS\/IFIP PERFORMANCE Joint International Conference on Measurement and Modeling of Computer Systems","author":"Akbarzadeh Negar","year":"2024","unstructured":"Negar Akbarzadeh, Sina Darabi, Atiyeh Gheibi-Fetrat, Amir Mirzaei, Mohammad Sadrosadati, and Hamid Sarbazi-Azad. June 2024. A High-bandwidth High-capacity Hybrid 3D Memory for GPUs. In ACM SIGMETRICS\/IFIP PERFORMANCE Joint International Conference on Measurement and Modeling of Computer Systems, Venice, Italy, Michele Garetto, Andrea Marin, Florin Ciucu, Giulia Fanti, and Rhonda Righter (Eds.). ACM, 67-68. https:\/\/doi.org\/10.1145\/3652963.3655057."},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2501.00203"},{"key":"e_1_2_1_5_1","volume-title":"Proceedings of the Linux Storage and Filesystems Conference (Vault '19)","author":"Axboe Jens","year":"2019","unstructured":"Jens Axboe. 2019. Efficient I\/O with io_uring. In Proceedings of the Linux Storage and Filesystems Conference (Vault '19). https:\/\/kernel.dk\/io_uring.pdf."},{"key":"e_1_2_1_6_1","unstructured":"Jens Axboe et al. 2020. Commit 01ee1941aba1202f0926d5047a2a4cf84d0e45d: io_uring enhancements (v5.7-rc7). https:\/\/git.kernel.org\/pub\/scm\/linux\/kernel\/git\/stable\/linux-stable.git\/commit\/?id=01ee1941aba1. Introduced parts of hybrid polling and interrupt handling in io_uring."},{"key":"e_1_2_1_7_1","unstructured":"Jens Axboe and Vincent Fu. 2024. FIO Repository. https:\/\/github.com\/axboe\/fio."},{"key":"e_1_2_1_8_1","first-page":"387","volume-title":"FlashNeuron: SSD-Enabled Large-Batch Training of Very Deep Neural Networks. In 19th USENIX Conference on File and Storage Technologies, FAST 2021","author":"Bae Jonghyun","year":"2021","unstructured":"Jonghyun Bae, Jongsung Lee, Yunho Jin, Sam Son, Shine Kim, Hakbeom Jang, Tae Jun Ham, and Jae W. Lee. 2021. FlashNeuron: SSD-Enabled Large-Batch Training of Very Deep Neural Networks. In 19th USENIX Conference on File and Storage Technologies, FAST 2021, February 23-25, 2021, Marcos K. Aguilera and Gala Yadgar (Eds.). USENIX Association, 387-401. https:\/\/www.usenix.org\/conference\/fast21\/presentation\/bae."},{"key":"e_1_2_1_9_1","unstructured":"Compute Express Link Consortium. 2024. Compute Express Link (CXL) Specification 3.1. https:\/\/computeexpresslink.org\/."},{"key":"e_1_2_1_10_1","volume-title":"Proceedings of the 6th International Workshop on Runtime and Operating Systems for Supercomputers","author":"Daoud Feras","year":"2016","unstructured":"Feras Daoud, Amir Wated, and Mark Silberstein. 2016. GPUrdma: GPU-side library for high performance networking from GPU kernels. In Proceedings of the 6th International Workshop on Runtime and Operating Systems for Supercomputers, Kyoto, Japan, June 1, 2016, Kamil Iskra and Torsten Hoefler (Eds.). ACM, 6:1-6:8. https:\/\/doi.org\/10.1145\/2931088.2931091."},{"key":"e_1_2_1_11_1","unstructured":"Google DeepMind. 2024. Gemma 3: Open Models Based on Gemini Research. https:\/\/ai.google.dev\/gemma."},{"key":"e_1_2_1_12_1","volume-title":"DeepSeek-Coder: When Code Meets Efficiency. arXiv preprint arXiv:2501.04648","author":"AI.","year":"2025","unstructured":"DeepSeek-AI. 2025. DeepSeek-Coder: When Code Meets Efficiency. arXiv preprint arXiv:2501.04648 (2025). https:\/\/arxiv.org\/abs\/2501.04648."},{"key":"e_1_2_1_13_1","doi-asserted-by":"crossref","first-page":"118","DOI":"10.1109\/CLUSTER52292.2023.00018","volume-title":"IEEE International Conference on Cluster Computing, CLUSTER 2023","author":"Doekemeijer Krijn","year":"2023","unstructured":"Krijn Doekemeijer, Nick Tehrany, Balakrishnan Chandrasekaran, Matias Bj\u00f8rling, and Animesh Trivedi. 2023. Performance Characterization of NVMe Flash Devices with Zoned Namespaces (ZNS). In IEEE International Conference on Cluster Computing, CLUSTER 2023, Santa Fe, NM, USA, October 31 - Nov. 3, 2023. IEEE, 118-131. https:\/\/doi.org\/10.1109\/CLUSTER52292.2023.00018."},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2401.14351"},{"key":"e_1_2_1_15_1","unstructured":"Sebastien Godard. 2024. sysstat: Performance Monitoring Tools for Linux. https:\/\/github.com\/sysstat\/sysstat."},{"key":"e_1_2_1_16_1","volume-title":"LoRA: Low-Rank Adaptation of Large Language Models. CoRR abs\/2106.09685","author":"Hu Edward J.","year":"2021","unstructured":"Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, and Weizhu Chen. 2021. LoRA: Low-Rank Adaptation of Large Language Models. CoRR abs\/2106.09685 (2021). https:\/\/arxiv.org\/abs\/2106.09685."},{"key":"e_1_2_1_17_1","volume-title":"Characterization of Large Language Model Development in the Datacenter. In 21st USENIX Symposium on Networked Systems Design and Implementation, NSDI 2024","author":"Hu Qinghao","year":"2024","unstructured":"Qinghao Hu, Zhisheng Ye, Zerui Wang, Guoteng Wang, Meng Zhang, Qiaoling Chen, Peng Sun, Dahua Lin, Xiaolin Wang, Yingwei Luo, Yonggang Wen, and Tianwei Zhang. 2024. Characterization of Large Language Model Development in the Datacenter. In 21st USENIX Symposium on Networked Systems Design and Implementation, NSDI 2024, Santa Clara, CA, April 15-17, 2024, Laurent Vanbever and Irene Zhang (Eds.). USENIX Association, 709-729. https:\/\/www.usenix.org\/conference\/nsdi24\/presentation\/hu."},{"key":"e_1_2_1_18_1","unstructured":"HuggingFace. 2024. Multilingual reasoning and chain-of-thought dataset. https:\/\/huggingface.co\/datasets\/HuggingFaceH4\/Multilingual-Thinking."},{"key":"e_1_2_1_19_1","unstructured":"HuggingFaceH4. 2023. Large-scale multi-turn dialogue dataset. https:\/\/huggingface.co\/datasets\/HuggingFaceH4\/ultrachat_200k"},{"key":"e_1_2_1_20_1","first-page":"1","volume-title":"Quantifying Performance Gains of GPUDirect Storage. In IEEE International Conference on Networking, Architecture and Storage, NAS 2022","author":"Inupakutika Devasena","year":"2022","unstructured":"Devasena Inupakutika, Bridget Davis, Qirui Yang, Daniel Kim, and David Akopian. 2022. Quantifying Performance Gains of GPUDirect Storage. In IEEE International Conference on Networking, Architecture and Storage, NAS 2022, Philadelphia, PA, USA, October 3-4, 2022. IEEE, 1-9. https:\/\/doi.org\/10.1109\/NAS55553.2022.9925516."},{"key":"e_1_2_1_21_1","volume-title":"ACL 2025","author":"Jiang Chaoyi","year":"2025","unstructured":"Chaoyi Jiang, Lei Gao, Hossein Entezari Zarch, and Murali Annavaram. 2025. KVPR: Efficient LLM Inference with I\/O-Aware KV Cache Partial Recomputation. In Findings of the Association for Computational Linguistics, ACL 2025, Vienna, Austria, July 27 - August 1, 2025, Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mohammad Taher Pilehvar (Eds.). Association for Computational Linguistics, 19474-19488. https:\/\/aclanthology.org\/2025.findings-acl.997.pdf."},{"key":"e_1_2_1_22_1","volume-title":"22nd USENIX Conference on File and Storage Technologies, FAST 2024","author":"Joshi Kanchan","year":"2024","unstructured":"Kanchan Joshi, Anuj Gupta, Javier Gonz\u00e1lez, Ankit Kumar, Krishna Kanth Reddy, Arun George, Simon Andreas Frimann Lund, and Jens Axboe. 2024. I\/O Passthru: Upstreaming a flexible and efficient I\/O Path in Linux. In 22nd USENIX Conference on File and Storage Technologies, FAST 2024, Santa Clara, CA, USA, February 27-29, 2024, Xiaosong Ma and Youjip Won (Eds.). USENIX Association, 107-121. https:\/\/www.usenix.org\/conference\/fast24\/presentation\/joshi."},{"key":"e_1_2_1_23_1","doi-asserted-by":"crossref","first-page":"324","DOI":"10.1109\/CLUSTER51413.2022.00044","volume-title":"IEEE International Conference on Cluster Computing, CLUSTER 2022","author":"Khan Awais","year":"2022","unstructured":"Awais Khan, Arnab K. Paul, Christopher Zimmer, Sarp Oral, Sajal Dash, Scott Atchley, and Feiyi Wang. 2022. Hvac: Removing I\/O Bottleneck for Large-Scale Deep Learning Applications. In IEEE International Conference on Cluster Computing, CLUSTER 2022, Heidelberg, Germany, September 5-8, 2022. IEEE, 324-335. https:\/\/doi.org\/10.1109\/CLUSTER51413.2022.00044."},{"key":"e_1_2_1_24_1","first-page":"48","volume-title":"Proceedings of the 30th ACM International Conference on Architectural Support for Programming Languages and Operating Systems","volume":"2","author":"Kumar Abhishek Vijaya","year":"2025","unstructured":"Abhishek Vijaya Kumar, Gianni Antichi, and Rachee Singh. 2025. Aqua: Network-Accelerated Memory Offloading for LLMs in Scale-Up GPU Domains. In Proceedings of the 30th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2, ASPLOS 2025, Rotterdam, Netherlands, 30 March 2025 - 3 April 2025, Lieven Eeckhout, Georgios Smaragdakis, Katai Liang, Adrian Sampson, Martha A. Kim, and Christopher J. Rossbach (Eds.). ACM, 48-62. https:\/\/doi.org\/10.1145\/3676641.3715983."},{"key":"e_1_2_1_25_1","doi-asserted-by":"crossref","first-page":"611","DOI":"10.1145\/3600006.3613165","volume-title":"Proceedings of the 29th Symposium on Operating Systems Principles, SOSP 2023","author":"Kwon Woosuk","year":"2023","unstructured":"Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Stoica. 2023. Efficient Memory Management for Large Language Model Serving with PagedAttention. In Proceedings of the 29th Symposium on Operating Systems Principles, SOSP 2023, Koblenz, Germany, October 23-26, 2023, Jason Flinn, Margo I. Seltzer, Peter Druschel, Antoine Kaufmann, and Jonathan Mace (Eds.). ACM, 611-626. https:\/\/doi.org\/10.1145\/3600006.3613165."},{"key":"e_1_2_1_26_1","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3722215","article-title":"I\/O in Machine Learning Applications on HPC Systems: A 360-Degree","volume":"57","author":"Lewis Noah","year":"2025","unstructured":"Noah Lewis, Jean Luca Bez, and Suren Byna. 2025. I\/O in Machine Learning Applications on HPC Systems: A 360-Degree Survey. Comput. Surveys 57, 10 (2025), 1-41. https:\/\/dl.acm.org\/doi\/full\/10.1145\/3722215.","journal-title":"Survey. Comput. Surveys"},{"key":"e_1_2_1_27_1","unstructured":"Linux man-pages project. 2024. Linux Native Asynchronous I\/O (libaio) Manual Pages. https:\/\/man7.org\/linux\/man-pages\/man7\/aio.7.html."},{"key":"e_1_2_1_28_1","volume-title":"Proceedings of the International Conference for High Performance Computing, Networking, Storage, and Analysis, SC 2018","author":"Markthub Pak","year":"2018","unstructured":"Pak Markthub, Mehmet E. Belviranli, Seyong Lee, Jeffrey S. Vetter, and Satoshi Matsuoka. 2018. DRAGON: breaking GPU memory capacity limits with direct NVM access. In Proceedings of the International Conference for High Performance Computing, Networking, Storage, and Analysis, SC 2018, Dallas, TX, USA, November 11-16, 2018. IEEE \/ ACM, 32:1-32:13. http:\/\/dl.acm.org\/citation.cfm?id=3291699."},{"key":"e_1_2_1_29_1","unstructured":"MLCommons. 2024. MLPerf Storage Benchmark v2.0. https:\/\/github.com\/mlcommons\/storage."},{"key":"e_1_2_1_30_1","doi-asserted-by":"crossref","first-page":"771","DOI":"10.14778\/3446095.3446100","article-title":"Analyzing and Mitigating Data Stalls in DNN Training","volume":"14","author":"Mohan Jayashree","year":"2021","unstructured":"Jayashree Mohan, Amar Phanishayee, Ashish Raniwala, and Vijay Chidambaram. 2021. Analyzing and Mitigating Data Stalls in DNN Training. Proc. VLDB Endow. 14, 5 (2021), 771-784. http:\/\/www.vldb.org\/pvldb\/vol14\/p771-mohan.pdf.","journal-title":"Proc. VLDB Endow."},{"key":"e_1_2_1_31_1","unstructured":"NVIDIA. 2022. NVIDIA Nsight Compute. https:\/\/developer.nvidia.com\/nsight-compute"},{"key":"e_1_2_1_32_1","unstructured":"NVIDIA Corporation. 2023. NVIDIA GPUDirect Storage: Architecture and Performance. Technical Whitepaper. https:\/\/developer.nvidia.com\/gpudirect-storage."},{"key":"e_1_2_1_33_1","unstructured":"NVIDIA Corporation. 2024. NVIDIA GPUDirect RDMA Overview Guide. https:\/\/docs.nvidia.com\/cuda\/gpudirect-rdma\/"},{"key":"e_1_2_1_34_1","unstructured":"NVIDIA Corporation. 2025. NVIDIA GPUDirect Storage Design Guide. https:\/\/docs.nvidia.com\/gpudirect-storage\/pdf\/design-guide.pdf\/ Release r1.12."},{"key":"e_1_2_1_35_1","unstructured":"NVIDIA Corporation. 2025. NVIDIA Magnum IO GPUDirect Storage Overview Guide. https:\/\/docs.nvidia.com\/gpudirect-storage\/overview-guide\/index.html. Release r1.12."},{"key":"e_1_2_1_36_1","volume-title":"High-Performance Deep Learning Library. In Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019","author":"Paszke Adam","year":"2019","unstructured":"Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas K\u00f6pf, Edward Z. Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. 2019. PyTorch: An Imperative Style, High-Performance Deep Learning Library. In Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada, Hanna M. Wallach, Hugo Larochelle, Alina Beygelzimer, Florence d'Alch\u00e9-Buc, Emily B. Fox, and Roman Garnett (Eds.). 8024-8035. https:\/\/proceedings.neurips.cc\/paper\/2019\/hash\/bdbca288fee7f92f2bfa9f7012727740-Abstract.html."},{"key":"e_1_2_1_37_1","first-page":"1133","volume-title":"Proceedings of the 30th ACM International Conference on Architectural Support for Programming Languages and Operating Systems","volume":"1","author":"Prabhu Ramya","year":"2025","unstructured":"Ramya Prabhu, Ajay Nayak, Jayashree Mohan, Ramachandran Ramjee, and Ashish Panwar. 2025. vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention. In Proceedings of the 30th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 1, ASPLOS 2025, Rotterdam, The Netherlands, 30 March 2025 - 3 April 2025, Lieven Eeckhout, Georgios Smaragdakis, Kaitai Liang, Adrian Sampson, Martha A. Kim, and Christopher J. Rossbach (Eds.). ACM, 1133-1150. https:\/\/doi.org\/10.1145\/3669940.3707256."},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1145\/3575693.3575748"},{"key":"e_1_2_1_39_1","article-title":"Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer","volume":"21","author":"Raffel Colin","year":"2020","unstructured":"Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer. J. Mach. Learn. Res. 21 (2020), 140:1-140:67. https:\/\/jmlr.org\/papers\/v21\/20-074.html.","journal-title":"J. Mach. Learn. Res."},{"key":"e_1_2_1_40_1","volume-title":"International Conference for High Performance Computing, Networking, Storage and Analysis, SC 2021","author":"Rajbhandari Samyam","year":"2021","unstructured":"Samyam Rajbhandari, Olatunji Ruwase, Jeff Rasley, Shaden Smith, and Yuxiong He. 2021. ZeRO-infinity: breaking the GPU memory wall for extreme scale deep learning. In International Conference for High Performance Computing, Networking, Storage and Analysis, SC 2021, St. Louis, Missouri, USA, November 14-19, 2021, Bronis R. de Supinski, Mary W. Hall, and Todd Gamblin (Eds.). ACM, 59. https:\/\/doi.org\/10.1145\/3458817.3476205."},{"key":"e_1_2_1_41_1","first-page":"28","volume-title":"Fifth IEEE\/ACM International Parallel Data Systems Workshop, PDSW@SC 2020","author":"Ravi John","year":"2020","unstructured":"John Ravi, Suren Byna, and Quincey Koziol. 2020. GPU Direct I\/O with HDF5. In Fifth IEEE\/ACM International Parallel Data Systems Workshop, PDSW@SC 2020, Atlanta, GA, USA, November 12, 2020. IEEE, 28-33. https:\/\/doi.org\/10.1109\/PDSW51947.2020.00010."},{"key":"e_1_2_1_42_1","volume-title":"Proceedings of the 5th Workshop on Challenges and Opportunities of Efficient and Performant Storage Systems. ACM. https:\/\/dl.acm.org\/doi\/pdf\/10","author":"Ren Zheng","year":"2025","unstructured":"Zheng Ren, Koen Doekemeijer, Tommaso De Matteis, Yousra Zaidan, and Alexandru Iosup. 2025. An I\/O Characterizing Study of Offloading LLM Models and KV Caches to NVMe SSD. In Proceedings of the 5th Workshop on Challenges and Opportunities of Efficient and Performant Storage Systems. ACM. https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3719330.3721230."},{"key":"e_1_2_1_43_1","volume-title":"Proceedings of the 3rd Workshop on Challenges and Opportunities of Efficient and Performant Storage Systems, CHEOPS 2023","author":"Ren Zebin","year":"2023","unstructured":"Zebin Ren and Animesh Trivedi. 2023. Performance Characterization of Modern Storage Stacks: POSIX I\/O, libaio, SPDK, and io_uring. In Proceedings of the 3rd Workshop on Challenges and Opportunities of Efficient and Performant Storage Systems, CHEOPS 2023, Rome, Italy, 8 May 2023, Jean-Thomas Acquaviva, Shadi Ibrahim, and Suren Byna (Eds.). ACM, 35-45. https:\/\/doi.org\/10.1145\/3578353.3589545."},{"key":"e_1_2_1_44_1","unstructured":"Ali Sedaghatgoo Reza Salkhordeh Andr\u00e9 Brinkmann and Hossein Asadi. 2026. Towards Scalable Storage Architectures for GPU Clusters Running Large Language Models. http:\/\/dsn.ce.sharif.edu\/traces\/codes\/scalable-storage-llm.zip. Open-source release of benchmarking framework configuration scripts and analysis tools."},{"key":"e_1_2_1_45_1","volume-title":"International Conference on Machine Learning, ICML 2023","volume":"31116","author":"Sheng Ying","year":"2023","unstructured":"Ying Sheng, Lianmin Zheng, Binhang Yuan, Zhuohan Li, Max Ryabinin, Beidi Chen, Percy Liang, Christopher R\u00e9, Ion Stoica, and Ce Zhang. 2023. FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU. In International Conference on Machine Learning, ICML 2023, 23-29 July 2023, Honolulu, Hawaii, USA (Proceedings of Machine Learning Research, Vol. 202), Andreas Krause, Emma Brunskill, Kyunghyun Cho, Barbara Engelhardt, Sivan Sabato, and Jonathan Scarlett (Eds.). PMLR, 31094-31116. https:\/\/proceedings.mlr.press\/v202\/sheng23a.html."},{"key":"e_1_2_1_46_1","volume-title":"Stanford Alpaca: An Instruction-Following LLaMA Model. GitHub Repository. https:\/\/github.com\/tatsu-lab\/stanford_alpaca.","author":"Taori Rohan","year":"2023","unstructured":"Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Percy Liang, et al. 2023. Stanford Alpaca: An Instruction-Following LLaMA Model. GitHub Repository. https:\/\/github.com\/tatsu-lab\/stanford_alpaca."},{"key":"e_1_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2403.08295"},{"key":"e_1_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2407.21783"},{"key":"e_1_2_1_49_1","unstructured":"Hugo Touvron Louis Martin Thibaut Lavril and the Meta AI Team. 2024. LLaMA 3 Model Card. Meta AI Research Release. https:\/\/ai.meta.com\/research\/publications\/llama3\/."},{"key":"e_1_2_1_50_1","volume-title":"HuggingFace's Transformers: State-of-the-art Natural Language Processing. CoRR abs\/1910.03771","author":"Wolf Thomas","year":"2019","unstructured":"Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, R\u00e9mi Louf, Morgan Funtowicz, and Jamie Brew. 2019. HuggingFace's Transformers: State-of-the-art Natural Language Processing. CoRR abs\/1910.03771 (2019). http:\/\/arxiv.org\/abs\/1910.03771."},{"key":"e_1_2_1_51_1","first-page":"196","volume-title":"IEEE International Conference on Cluster Computing, CLUSTER - Workshops 2024","author":"Wu Du","year":"2024","unstructured":"Du Wu, Peng Chen, Yiyu Tan, Yusuke Tanimura, Toshio Endo, Satoshi Matsuoka, and Mohamed Wahib. 2024. Asynchronous I\/O Optimization for X-Ray Imaging via GPUDirect Storage. In IEEE International Conference on Cluster Computing, CLUSTER - Workshops 2024, Kobe, Japan, September 24-27, 2024. IEEE, 196-197. https:\/\/doi.org\/10.1109\/CLUSTERWorkshops61563.2024.00056."},{"key":"e_1_2_1_52_1","first-page":"1","volume-title":"62nd ACM\/IEEE Design Automation Conference, DAC 2025","author":"Wu Kun","year":"2025","unstructured":"Kun Wu, Jeongmin Brian Park, Xiaofan Zhang, Mert Hidayetoglu, Vikram Sharma Mailthody, Sitao Huang, Steven S. Lumetta, and Wen-Mei Hwu. 2025. SSDTrain: An Activation Offloading Framework to SSDs for Faster Large Language Model Training. In 62nd ACM\/IEEE Design Automation Conference, DAC 2025, San Francisco, CA, USA, June 22-25, 2025. IEEE, 1-7. https:\/\/doi.org\/10.1109\/DAC63849.2025.11132754."},{"key":"e_1_2_1_53_1","first-page":"473","volume-title":"Proceedings of the 2015 ACM SIGMETRICS International Conference on Measurement and Modeling of Computer Systems","author":"Xu Qiumin","year":"2015","unstructured":"Qiumin Xu, Huzefa Siyamwala, Mrinmoy Ghosh, Manu Awasthi, Tameesh Suri, Zvika Guz, Anahita Shayesteh, and Vijay Balakrishnan. 2015. Performance Characterization of Hyperscale Applications on on NVMe SSDs. In Proceedings of the 2015 ACM SIGMETRICS International Conference on Measurement and Modeling of Computer Systems, Portland, OR, USA, June 15-19, 2015, Bill Lin, Jun (Jim) Xu, Sudipta Sengupta, and Devavrat Shah (Eds.). ACM, 473-474. https:\/\/doi.org\/10.1145\/2745844.2745901."},{"key":"e_1_2_1_54_1","unstructured":"An Yang Anfeng Li Baosong Yang Beichen Zhang Binyuan Hui Bo Zheng Bowen Yu Chang Gao Chengen Huang Chenxu Lv Chujie Zheng Dayiheng Liu Fan Zhou Fei Huang Feng Hu Hao Ge Haoran Wei Huan Lin Jialong Tang Jian Yang Jianhong Tu Jianwei Zhang Jian Yang Jiaxi Yang Jingren Zhou Junyang Lin Kai Dang Keqin Bao Kexin Yang Le Yu Lianghao Deng Mei Li Mingfeng Xue Mingze Li Pei Zhang Peng Wang Qin Zhu Rui Men Ruize Gao Shixuan Liu Shuang Luo Tianhao Li Tianyi Tang Wenbiao Yin Xingzhang Ren Xinyu Wang Xinyu Zhang Xuancheng Ren Yang Fan Yang Su Yichang Zhang Yinger Zhang Yu Wan Yuqiong Liu Zekun Wang Zeyu Cui Zhenru Zhang Zhipeng Zhou and Zihan Qiu. 2025. Qwen3 Technical Report. CoRR abs\/2505.09388 (2025). https:\/\/doi.org\/10.48550\/arXiv.2505.09388."},{"key":"e_1_2_1_55_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2412.15115"},{"key":"e_1_2_1_56_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2503.17864"},{"key":"e_1_2_1_57_1","first-page":"154","volume-title":"IEEE International Conference on Cloud Computing Technology and Science, CloudCom 2017","author":"Yang Ziye","year":"2017","unstructured":"Ziye Yang, James R. Harris, Benjamin Walker, Daniel Verkamp, Changpeng Liu, Cunyin Chang, Gang Cao, Jonathan Stern, Vishal Verma, and Luse E. Paul. 2017. SPDK: A Development Kit to Build High Performance Storage Applications. In IEEE International Conference on Cloud Computing Technology and Science, CloudCom 2017, Hong Kong, December 11-14, 2017. IEEE Computer Society, 154-161. https:\/\/doi.org\/10.1109\/CloudCom.2017.14."},{"key":"e_1_2_1_58_1","doi-asserted-by":"crossref","first-page":"1028","DOI":"10.1145\/3712285.3759778","volume-title":"Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis, SC 2025","author":"Yang Zhuoping","year":"2025","unstructured":"Zhuoping Yang, Jinming Zhuang, Xingzhen Chen, Alex K. Jones, and Peipei Zhou. 2025. AGILE: Lightweight and Efficient Asynchronous GPU-SSD Integration. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis, SC 2025, St. Louis, MO, USA, November 16-21, 2025. ACM, 1028-1042. https:\/\/doi.org\/10.1145\/3712285.3759778."},{"key":"e_1_2_1_59_1","doi-asserted-by":"publisher","DOI":"10.1145\/605521.605524"},{"key":"e_1_2_1_60_1","first-page":"395","volume-title":"Proceedings of the 56th Annual IEEE\/ACM International Symposium on Microarchitecture, MICRO 2023","author":"Zhang Haoyang","year":"2023","unstructured":"Haoyang Zhang, Yirui Eric Zhou, Yuqi Xue, Yiqi Liu, and Jian Huang. 2023. G10: Enabling An Efficient Unified GPU Memory and Storage Architecture with Smart Tensor Migrations. In Proceedings of the 56th Annual IEEE\/ACM International Symposium on Microarchitecture, MICRO 2023, Toronto, ON, Canada, 28 October 2023 - 1 November 2023. ACM, 395-410. https:\/\/doi.org\/10.1145\/3613424.3614309."},{"key":"e_1_2_1_61_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2205.01068"}],"container-title":["Proceedings of the ACM on Measurement and Analysis of Computing Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3788106","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,3,26]],"date-time":"2026-03-26T18:49:58Z","timestamp":1774550998000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3788106"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,3,26]]},"references-count":61,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2026,3,26]]}},"alternative-id":["10.1145\/3788106"],"URL":"https:\/\/doi.org\/10.1145\/3788106","relation":{},"ISSN":["2476-1249"],"issn-type":[{"value":"2476-1249","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,3,26]]},"assertion":[{"value":"2026-03-26","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}