{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,11]],"date-time":"2026-07-11T15:42:01Z","timestamp":1783784521596,"version":"3.55.0"},"reference-count":110,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2024,4,4]],"date-time":"2024-04-04T00:00:00Z","timestamp":1712188800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"National Science Foundation","award":["CCF-1919113, CNS-1405697, CNS-1615411, CNS-1565314\/1838271 OAC-1835890, CSR-2312785, CSR-2106634\/2312785, and CCF-1919113\/1919075"],"award-info":[{"award-number":["CCF-1919113, CNS-1405697, CNS-1615411, CNS-1565314\/1838271 OAC-1835890, CSR-2312785, CSR-2106634\/2312785, and CCF-1919113\/1919075"]}]},{"name":"Oak Ridge Leadership Computing Facility"},{"name":"National Center for Computational Sciences"},{"name":"Office of Science of the DOE","award":["DE-AC05-00OR22725"],"award-info":[{"award-number":["DE-AC05-00OR22725"]}]},{"name":"European High-Performance Computing Joint Undertaking","award":["101033975"],"award-info":[{"award-number":["101033975"]}]},{"name":"European Union\u2019s Horizon 2020"},{"DOI":"10.13039\/501100006464","name":"BITS Pilani","doi-asserted-by":"crossref","award":["BBF\/BITS(G)\/FY2022-23\/BCPS-123, GOA\/ACG\/2022-2023\/Oct\/11, and BPGC\/RIG\/2021-22\/06-2022\/02"],"award-info":[{"award-number":["BBF\/BITS(G)\/FY2022-23\/BCPS-123, GOA\/ACG\/2022-2023\/Oct\/11, and BPGC\/RIG\/2021-22\/06-2022\/02"]}],"id":[{"id":"10.13039\/501100006464","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Storage"],"published-print":{"date-parts":[[2024,5,31]]},"abstract":"<jats:p>\n            The imbalanced I\/O load on large parallel file systems affects the parallel I\/O performance of high-performance computing (HPC) applications. One of the main reasons for I\/O imbalances is the lack of a global view of system-wide resource consumption. While approaches to address the problem already exist, the diversity of HPC workloads combined with different file striping patterns prevents widespread adoption of these approaches. In addition, load-balancing techniques should be transparent to client applications. To address these issues, we propose\n            <jats:monospace>Tarazu<\/jats:monospace>\n            , an end-to-end control plane where clients transparently and adaptively write to a set of selected I\/O servers to achieve balanced data placement. Our control plane leverages real-time load statistics for global data placement on distributed storage servers, while our design model employs trace-based optimization techniques to minimize latency for I\/O load requests between clients and servers and to handle multiple striping patterns in files. We evaluate our proposed system on an experimental cluster for two common use cases: the synthetic I\/O benchmark IOR and the scientific application I\/O kernel HACC-I\/O. We also use a discrete-time simulator with real HPC application traces from emerging workloads running on the Summit supercomputer to validate the effectiveness and scalability of\n            <jats:monospace>Tarazu<\/jats:monospace>\n            in large-scale storage environments. The results show improvements in load balancing and read performance of up to 33% and 43%, respectively, compared to the state-of-the-art.\n          <\/jats:p>","DOI":"10.1145\/3641885","type":"journal-article","created":{"date-parts":[[2024,2,1]],"date-time":"2024-02-01T11:56:43Z","timestamp":1706788603000},"page":"1-42","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":5,"title":["Tarazu: An Adaptive End-to-end I\/O Load-balancing Framework for Large-scale Parallel File Systems"],"prefix":"10.1145","volume":"20","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-3694-5511","authenticated-orcid":false,"given":"Arnab K.","family":"Paul","sequence":"first","affiliation":[{"name":"BITS Pilani, KK Birla Goa Campus, Zuarinagar, India"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7409-153X","authenticated-orcid":false,"given":"Sarah","family":"Neuwirth","sequence":"additional","affiliation":[{"name":"Johannes Gutenberg University Mainz, Mainz, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9178-2831","authenticated-orcid":false,"given":"Bharti","family":"Wadhwa","sequence":"additional","affiliation":[{"name":"IBM Research, Yorktown Heights, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0099-1559","authenticated-orcid":false,"given":"Feiyi","family":"Wang","sequence":"additional","affiliation":[{"name":"Oak Ridge National Laboratory, Oak Ridge, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8745-7078","authenticated-orcid":false,"given":"Sarp","family":"Oral","sequence":"additional","affiliation":[{"name":"Oak Ridge National Laboratory, Oak Ridge, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0871-7263","authenticated-orcid":false,"given":"Ali R.","family":"Butt","sequence":"additional","affiliation":[{"name":"Virginia Tech, Blacksburg, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2024,4,4]]},"reference":[{"key":"e_1_3_1_2_2","first-page":"265","volume-title":"12th USENIX Symposium on Operating Systems Design and Implementation (OSDI\u201916)","author":"Abadi Mart\u00edn","year":"2016","unstructured":"Mart\u00edn Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, Manjunath Kudlur, Josh Levenberg, Rajat Monga, Sherry Moore, Derek G. Murray, Benoit Steiner, Paul Tucker, Vijay Vasudevan, Pete Warden, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng. 2016. TensorFlow: A system for Large-Scale machine learning. In 12th USENIX Symposium on Operating Systems Design and Implementation (OSDI\u201916). USENIX Association, 265\u2013283. Retrieved from https:\/\/www.usenix.org\/conference\/osdi16\/technical-sessions\/presentation\/abadi"},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1109\/PDSW49588.2019.00007"},{"key":"e_1_3_1_4_2","volume-title":"Network Flows: Theory, Algorithms, and Applications","author":"Ahuja Ravindra K.","year":"2017","unstructured":"Ravindra K. Ahuja. 2017. Network Flows: Theory, Algorithms, and Applications. Pearson Education, Chennai, India."},{"key":"e_1_3_1_5_2","volume-title":"Towards Efficient and Flexible Object Storage Using Resource and Functional Partitioning","author":"Anwar Ali","year":"2018","unstructured":"Ali Anwar. 2018. Towards Efficient and Flexible Object Storage Using Resource and Functional Partitioning. Ph. D. Dissertation. Virginia Tech."},{"key":"e_1_3_1_6_2","doi-asserted-by":"publisher","DOI":"10.1145\/2907294.2907304"},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPSW50202.2020.00138"},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPSW52791.2021.00118"},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1145\/3309205"},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.1145\/3611007"},{"key":"e_1_3_1_11_2","volume-title":"The Lustre Storage Architecture (Tech. Rep.)","author":"Braam P. J.","year":"2004","unstructured":"P. J. Braam. 2004. The Lustre Storage Architecture (Tech. Rep.). Technical Report. Retrieved from http:\/\/wiki.lustre.org\/"},{"key":"e_1_3_1_12_2","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-319-29854-2","volume-title":"Introduction to Time Series and Forecasting (3rd ed.)","author":"Brockwell Peter J.","year":"2016","unstructured":"Peter J. Brockwell and Richard A. Davis. 2016. Introduction to Time Series and Forecasting (3rd ed.). Springer International Publishing, Cham, Switzerland."},{"key":"e_1_3_1_13_2","first-page":"12","volume-title":"IEEE International Conference on Cluster Computing and Workshops","author":"Carns Philip","year":"2009","unstructured":"Philip Carns, Robert Latham, Robert Ross, Kamil Iskra, Samuel Lang, and Katherine Riley. 2009. 24\/7 characterization of petascale I\/O workloads. In IEEE International Conference on Cluster Computing and Workshops. IEEE, 12 pages. DOI:10.1109\/CLUSTR.2009.5289150"},{"key":"e_1_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.1145\/347823.347828"},{"key":"e_1_3_1_15_2","volume-title":"48th International Conference on Parallel Processing (ICPP\u201919)","author":"Chowdhury Fahim","year":"2019","unstructured":"Fahim Chowdhury, Yue Zhu, Todd Heer, Saul Paredes, Adam Moody, Robin Goldstone, Kathryn Mohror, and Weikuan Yu. 2019. I\/O characterization and performance evaluation of BeeGFS for deep learning. In 48th International Conference on Parallel Processing (ICPP\u201919). ACM, New York, NY. DOI:10.1145\/3337821.3337902"},{"key":"e_1_3_1_16_2","article-title":"CODES: Enabling Co-design of multi-layer exascale storage architectures","author":"Cope J.","year":"2011","unstructured":"J. Cope, N. Liu, S. Lang, P. Carns, C. Carothers, and R. Ross. 2011. CODES: Enabling Co-design of multi-layer exascale storage architectures. In Workshop on Emerging Supercomputing Technologies.","journal-title":"Workshop on Emerging Supercomputing Technologies"},{"key":"e_1_3_1_17_2","article-title":"Workflows community summit: Advancing the state-of-the-art of scientific workflows management systems research and development","volume":"2106","author":"Silva Rafael Ferreira da","year":"2021","unstructured":"Rafael Ferreira da Silva, Henri Casanova, Kyle Chard, Tain\u00e3 Coleman, Dan Laney, Dong Ahn, Shantenu Jha, Dorran Howell, Stian Soiland-Reyes, Ilkay Altintas, Douglas Thain, Rosa Filgueira, Yadu N. Babuji, Rosa M. Badia, Bartosz Balis, Silvina Ca\u00edno-Lores, Scott Callaghan, Frederik Coppens, Michael R. Crusoe, Kaushik De, Frank Di Natale, Tu Mai Anh Do, Bjoern Enders, Thomas Fahringer, Anne Fouilloux, Grigori Fursin, Alban Gaignard, Alex Ganose, Daniel Garijo, Sandra Gesing, Carole A. Goble, Adil Hasan, Sebastiaan Huber, Daniel S. Katz, Ulf Leser, Douglas Lowe, Bertram Lud\u00e4scher, Ketan Maheshwari, Maciej Malawski, Rajiv Mayani, Kshitij Mehta, Andr\u00e9 Merzky, Todd S. Munson, Jonathan Ozik, Lo\u00efc Pottier, Sashko Ristov, Mehdi Roozmeh, Renan Souza, Fr\u00e9d\u00e9ric Suter, Benjam\u00edn Tovar, Matteo Turilli, Karan Vahi, Alvaro Vidal-Torreira, Wendy R. Whitcup, Michael Wilde, Alan Williams, Matthew Wolf, and Justin M. Wozniak. 2021. Workflows community summit: Advancing the state-of-the-art of scientific workflows management systems research and development. CoRR abs\/2106.05177 (2021)","journal-title":"CoRR"},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.future.2017.12.022"},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.jpdc.2012.05.006"},{"key":"e_1_3_1_20_2","first-page":"155","volume-title":"IEEE International Conference on Cluster Computing","author":"Dorier Matthieu","year":"2012","unstructured":"Matthieu Dorier, Gabriel Antoniu, Franck Cappello, Marc Snir, and Leigh Orf. 2012. Damaris: How to efficiently leverage multicore parallelism to achieve scalable, jitter-free I\/O. In IEEE International Conference on Cluster Computing. IEEE, 155\u2013163. DOI:10.1109\/CLUSTER.2012.26"},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2015.2485980"},{"key":"e_1_3_1_22_2","first-page":"1","volume-title":"IEEE\/IFIP International Conference on Dependable Systems and Networks (DSN\u201912)","author":"Erazo Miguel A.","year":"2012","unstructured":"Miguel A. Erazo, Ting Li, Jason Liu, and Stephan Eidenbenz. 2012. Toward comprehensive and accurate simulation performance prediction of parallel file systems. In IEEE\/IFIP International Conference on Dependable Systems and Networks (DSN\u201912). IEEE, 1\u201312. DOI:10.1109\/DSN.2012.6263930"},{"key":"e_1_3_1_23_2","volume-title":"Titan Supercomputer","author":"Facility Oak Ridge Leadership","year":"2022","unstructured":"Oak Ridge Leadership Facility. 2022. Titan Supercomputer. Oak Ridge National Laboratory. Retrieved from https:\/\/www.olcf.ornl.gov\/olcf-resources\/compute-systems\/titan\/"},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.1109\/PDSW.2014.12"},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.future.2017.02.026"},{"key":"e_1_3_1_26_2","volume-title":"EDBT\/ICDT\u201911 Workshop on Array Databases","author":"Folk Mike","year":"2011","unstructured":"Mike Folk, Gerd Heber, Quincey Koziol, Elena Pourmal, and Dana Robinson. 2011. An overview of the HDF5 technology suite and its applications. In EDBT\/ICDT\u201911 Workshop on Array Databases. ACM, New York, NY."},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1002\/cpe.5649"},{"key":"e_1_3_1_28_2","unstructured":"Jim Garlick. 2010. Lustre Monitoring Tool (LMT). https:\/\/github.com\/LLNL\/lmt"},{"key":"e_1_3_1_29_2","volume-title":"International Conference on High Performance Computing, Networking, Storage and Analysis (SC\u201913)","author":"Habib S.","year":"2013","unstructured":"S. Habib, V. Morozov, N. Frontiere, H. Finkel, A. Pope, and K. Heitmann. 2013. HACC: Extreme scaling and performance across diverse architectures. In International Conference on High Performance Computing, Networking, Storage and Analysis (SC\u201913). ACM, New York, NY. DOI:10.1145\/2503210.2504566"},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2015.23"},{"key":"e_1_3_1_31_2","unstructured":"Jan Heichler. 2014. An Introduction to BeeGFS v1.1. https:\/\/www.beegfs.de\/docs\/whitepapers\/Introduction_to_BeeGFS_by_ThinkParQ.pdf. Accessed February 16 2024."},{"key":"e_1_3_1_32_2","volume-title":"ZeroMQ: Messaging for Many Applications","author":"Hintjens Pieter","year":"2013","unstructured":"Pieter Hintjens. 2013. ZeroMQ: Messaging for Many Applications. O\u2019Reilly Media, Inc."},{"key":"e_1_3_1_33_2","volume-title":"IEEE 26th International Conference on Parallel and Distributed Systems (ICPADS\u201920)","author":"Huang L.","year":"2020","unstructured":"L. Huang and S. Liu. 2020. OOOPS: An innovative tool for IO workload management on supercomputers. In IEEE 26th International Conference on Parallel and Distributed Systems (ICPADS\u201920). IEEE. DOI:10.1109\/ICPADS51040.2020.00069"},{"key":"e_1_3_1_34_2","first-page":"265","volume-title":"17th USENIX Conference on File and Storage Technologies (FAST\u201919)","author":"Ji Xu","year":"2019","unstructured":"Xu Ji, Bin Yang, Tianyu Zhang, Xiaosong Ma, Xiupeng Zhu, Xiyang Wang, Nosayba El-Sayed, Jidong Zhai, Weiguo Liu, and Wei Xue. 2019. Automatic, application-aware I\/O forwarding resource allocation. In 17th USENIX Conference on File and Storage Technologies (FAST\u201919). USENIX Association, 265\u2013279. Retrieved from https:\/\/www.usenix.org\/conference\/fast19\/presentation\/ji"},{"key":"e_1_3_1_35_2","doi-asserted-by":"publisher","DOI":"10.1145\/3369583.3392678"},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","DOI":"10.1109\/SC.Companion.2012.14"},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.1109\/FAS-W.2016.30"},{"key":"e_1_3_1_38_2","first-page":"294","volume-title":"International Conference on Information Technology (ICIT\u201916)","author":"Kumar Anoop S.","year":"2016","unstructured":"Anoop S. Kumar and Somnath Mazumdar. 2016. Forecasting HPC workload using ARMA models and SSA. In International Conference on Information Technology (ICIT\u201916). IEEE, 294\u2013297. DOI:10.1109\/ICIT.2016.065"},{"key":"e_1_3_1_39_2","volume-title":"The DiskSim Simulation Environment (V4.0)","author":"Lab Parallel Data","year":"2024","unstructured":"Parallel Data Lab. 2024. The DiskSim Simulation Environment (V4.0). Carnegie Mellon University. Retrieved from: https:\/\/www.pdl.cmu.edu\/DiskSim\/"},{"key":"e_1_3_1_40_2","article-title":"TOKIO: Total Knowledge of I\/O","author":"Laboratory Lawrence Berkeley National","year":"2022","unstructured":"Lawrence Berkeley National Laboratory and Argonne National Laboratory. 2022. TOKIO: Total Knowledge of I\/O. Retrieved from https:\/\/www.nersc.gov\/research-and-development\/storage-and-i-o-technologies\/tokio\/","journal-title":"R"},{"key":"e_1_3_1_41_2","unstructured":"Lawrence Livermore National Laboratory. 2021. IOR Benchmark Summary. Retrieved from https:\/\/asc.llnl.gov\/sequoia\/benchmarks\/IORsummaryv1.0.pdf"},{"key":"e_1_3_1_42_2","article-title":"Summit Supercomputer","author":"Laboratory Oak Ridge National","year":"2021","unstructured":"Oak Ridge National Laboratory. 2021. Summit Supercomputer. Retrieved from https:\/\/www.olcf.ornl.gov\/summit\/","journal-title":"R"},{"key":"e_1_3_1_43_2","unstructured":"Oak Ridge National Laboratory. 2022. Frontier Supercomputer. Retrieved from https:\/\/www.olcf.ornl.gov\/frontier\/"},{"key":"e_1_3_1_44_2","article-title":"Oak Ridge National Laboratory Storage Ecosystem","author":"Leverman Dustin","year":"2022","unstructured":"Dustin Leverman. 2022. Oak Ridge National Laboratory Storage Ecosystem. In Platform for Advanced Scientific Computing Conference (PASC\u201922). Retrieved from https:\/\/linklings.s3.amazonaws.com\/organizations\/pasc\/pasc22\/submissions\/stype117\/PYFgV-msa274s1.pdf","journal-title":"Platform for Advanced Scientific Computing Conference (PASC\u201922)."},{"key":"e_1_3_1_45_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11227-013-1077-6"},{"key":"e_1_3_1_46_2","first-page":"1","volume-title":"International Conference for High Performance Computing, Networking, Storage and Analysis (SC\u201917)","author":"Lim Seung-Hwan","year":"2017","unstructured":"Seung-Hwan Lim, Hyogi Sim, Raghul Gunasekaran, and Sudharshan S. Vazhkudai. 2017. Scientific user behavior and data-sharing trends in a petascale file system. In International Conference for High Performance Computing, Networking, Storage and Analysis (SC\u201917). ACM, New York, NY, 1\u201312. DOI:10.1145\/3126908.3126924"},{"key":"e_1_3_1_47_2","first-page":"12","volume-title":"7th IEEE International Workshop on Storage Network Architecture and Parallel I\/O (SNAPI\u201911)","author":"Liu Yonggang","year":"2011","unstructured":"Yonggang Liu, Renato Figueiredo, Dulcardo Clavijo, Yiqi Xu, and Ming Zhao. 2011. Towards simulation of parallel file system scheduling algorithms with PFSsim. In 7th IEEE International Workshop on Storage Network Architecture and Parallel I\/O (SNAPI\u201911). IEEE, 12 pages."},{"key":"e_1_3_1_48_2","doi-asserted-by":"publisher","DOI":"10.1109\/MSST.2013.6558438"},{"key":"e_1_3_1_49_2","doi-asserted-by":"publisher","DOI":"10.1145\/3149393.3149395"},{"key":"e_1_3_1_50_2","first-page":"5","volume-title":"IEEE International Conference on Cluster Computing (CLUSTER\u201913)","author":"Luu Huong","year":"2013","unstructured":"Huong Luu, Babak Behzad, Ruth Aydt, and Marianne Winslett. 2013. A multi-level approach for understanding I\/O activity in HPC applications. In IEEE International Conference on Cluster Computing (CLUSTER\u201913). IEEE, 5 pages. DOI:10.1109\/CLUSTER.2013.6702690"},{"key":"e_1_3_1_51_2","first-page":"11","volume-title":"International Conference for High Performance Computing, Networking, Storage and Analysis (SC\u201918)","author":"Mathuriya A.","year":"2018","unstructured":"A. Mathuriya, D. Bard, P. Mendygral, L. Meadows, J. Arnemann, L. Shao, S. He, T. K\u00e4rn\u00e4, D. Moise, S. J. Pennycook, K. Maschhoff, J. Sewall, N. Kumar, S. Ho, M. F. Ringenburg, P. Prabhat, and V. Lee. 2018. CosmoFlow: Using deep learning to learn the universe at scale. In International Conference for High Performance Computing, Networking, Storage and Analysis (SC\u201918). IEEE, 11 pages. DOI:10.1109\/SC.2018.00068"},{"key":"e_1_3_1_52_2","volume-title":"Cray User Group Conference (CUG\u201916)","author":"Mohr Rick","year":"2016","unstructured":"Rick Mohr, Michael Brim, Sarp Oral, and Andreas Dilger. 2016. Evaluating Progressive File Layouts for Lustre. In Cray User Group Conference (CUG\u201916)."},{"key":"e_1_3_1_53_2","doi-asserted-by":"publisher","DOI":"10.1088\/1742-6596\/180\/1\/012050"},{"key":"e_1_3_1_54_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-28145-7_40"},{"key":"e_1_3_1_55_2","unstructured":"Lustre Networking. 2008. High-Performance Features and Flexible Support for a Wide Array of Networks."},{"key":"e_1_3_1_56_2","volume-title":"Accelerating Network Communication and I\/O in Scientific High Performance Computing Environments","author":"Neuwirth Sarah","year":"2018","unstructured":"Sarah Neuwirth. 2018. Accelerating Network Communication and I\/O in Scientific High Performance Computing Environments. Ph. D. Dissertation. Heidelberg University, Germany."},{"key":"e_1_3_1_57_2","article-title":"An I\/O load balancing framework for large-scale applications (BPIO 2.0)","author":"Neuwirth Sarah","year":"2016","unstructured":"Sarah Neuwirth, S. Oral, F. Wang, and Ulrich Bruening. 2016. An I\/O load balancing framework for large-scale applications (BPIO 2.0). In Poster at SC\u201916.","journal-title":"Poster at SC\u201916"},{"key":"e_1_3_1_58_2","first-page":"671","volume-title":"IEEE International Conference on Cluster Computing (CLUSTER\u201921)","author":"Neuwirth Sarah","year":"2021","unstructured":"Sarah Neuwirth and Arnab K. Paul. 2021. Parallel I\/O evaluation techniques and emerging HPC workloads: A perspective. In IEEE International Conference on Cluster Computing (CLUSTER\u201921). IEEE, 671\u2013679. DOI:10.1109\/Cluster48925.2021.00100"},{"key":"e_1_3_1_59_2","first-page":"604","volume-title":"IEEE 23rd International Conference on Parallel and Distributed Systems (ICPADS\u201917)","author":"Neuwirth Sarah","year":"2017","unstructured":"Sarah Neuwirth, Feiyi Wang, Sarp Oral, and Ulrich Bruening. 2017. Automatic and transparent resource contention mitigation for improving large-scale parallel file system performance. In IEEE 23rd International Conference on Parallel and Distributed Systems (ICPADS\u201917). IEEE, 604\u2013613. DOI:10.1109\/ICPADS.2017.00084"},{"key":"e_1_3_1_60_2","volume-title":"NS-3 Network Simulator","year":"2023","unstructured":"NSNAM. 2023. NS-3 Network Simulator. Retrieved from https:\/\/www.nsnam.org\/"},{"key":"e_1_3_1_61_2","first-page":"147","volume-title":"16th International Conference on Supercomputing (ICS\u201902)","author":"Oly James","year":"2002","unstructured":"James Oly and Daniel A. Reed. 2002. Markov model prediction of I\/O requests for scientific applications. In 16th International Conference on Supercomputing (ICS\u201902). ACM, New York, NY, 147\u2013155. DOI:10.1145\/514191.514214"},{"key":"e_1_3_1_62_2","unstructured":"OpenSFS and EOFS. 2021. Lustre Operations Manual 2.x. Retrieved from https:\/\/www.lustre.org\/documentation\/"},{"key":"e_1_3_1_63_2","doi-asserted-by":"publisher","unstructured":"Sarp Oral James Simmons Jason Hill Dustin Leverman Feiyi Wang Matt Ezell Ross Miller Douglas Fuller Raghul Gunasekaran Youngjae Kim Saurabh Gupta Devesh Tiwari Sudharshan S. Vazhkudai James H. Rogers David Dillow Galen M. Shipman and Arthur S. Bland. 2014. Best practices and lessons learned from deploying and operating large-scale data-centric parallel file systems. SC\u201914: Proceedings of the International Conference for High Performance Computing Networking Storage and Analysis IEEE 217\u2013228. DOI:10.1109\/SC.2014.23","DOI":"10.1109\/SC.2014.23"},{"key":"e_1_3_1_64_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-981-13-9008-1_18"},{"key":"e_1_3_1_65_2","first-page":"91","volume-title":"18th USENIX Conference on File and Storage Technologies (FAST\u201920)","author":"Patel Tirthak","year":"2020","unstructured":"Tirthak Patel and Suren Byna. 2020. Uncovering access, reuse, and sharing characteristics of I\/O-intensive files on large-scale production HPC systems. In 18th USENIX Conference on File and Storage Technologies (FAST\u201920). USENIX Association, 91\u2013101. Retrieved from https:\/\/www.usenix.org\/conference\/fast20\/presentation\/patel-hpc-systems"},{"key":"e_1_3_1_66_2","volume-title":"International Conference for High Performance Computing, Networking, Storage and Analysis (SC\u201919)","author":"Patel Tirthak","year":"2019","unstructured":"Tirthak Patel, Suren Byna, Glenn K. Lockwood, and Devesh Tiwari. 2019. Revisiting I\/O behavior in large-scale storage systems: The expected and the unexpected. In International Conference for High Performance Computing, Networking, Storage and Analysis (SC\u201919). ACM, New York, NY, Article 65, 13 pages. DOI:10.1145\/3295500.3356183"},{"key":"e_1_3_1_67_2","first-page":"202","volume-title":"IEEE 27th International Conference on High Performance Computing, Data, and Analytics (HiPC\u201920)","author":"Paul Arnab K.","year":"2020","unstructured":"Arnab K. Paul, Olaf Faaland, Adam Moody, Elsa Gonsiorowski, Kathryn Mohror, and Ali R. Butt. 2020. Understanding HPC application I\/O behavior using system level statistics. In IEEE 27th International Conference on High Performance Computing, Data, and Analytics (HiPC\u201920). IEEE, 202\u2013211. DOI:10.1109\/HiPC50609.2020.00034"},{"key":"e_1_3_1_68_2","first-page":"233","volume-title":"IEEE International Conference on Big Data (Big Data\u201917)","author":"Paul Arnab K.","year":"2017","unstructured":"Arnab K. Paul, Arpit Goyal, Feiyi Wang, Sarp Oral, Ali R. Butt, Michael J. Brim, and Sangeetha B. Srinivasa. 2017. I\/O load balancing for big data HPC applications. In IEEE International Conference on Big Data (Big Data\u201917). IEEE, 233\u2013242. DOI:10.1109\/BigData.2017.8257931"},{"key":"e_1_3_1_69_2","doi-asserted-by":"publisher","DOI":"10.1109\/MASCOTS53633.2021.9614303"},{"key":"e_1_3_1_70_2","doi-asserted-by":"publisher","DOI":"10.4018\/978-1-5225-1721-4.ch006"},{"key":"e_1_3_1_71_2","doi-asserted-by":"publisher","DOI":"10.1145\/3149393.3149402"},{"key":"e_1_3_1_72_2","first-page":"110","volume-title":"IEEE International Conference on Cluster Computing (CLUSTER\u201916)","author":"Paul Arnab Kumar","year":"2016","unstructured":"Arnab Kumar Paul, Wenjie Zhuang, Luna Xu, Min Li, M. Mustafa Rafique, and Ali R. Butt. 2016. CHOPPER: Optimizing data partitioning for in-memory data analytics frameworks. In IEEE International Conference on Cluster Computing (CLUSTER\u201916). IEEE, 110\u2013119. DOI:10.1109\/CLUSTER.2016.41"},{"key":"e_1_3_1_73_2","first-page":"617","article-title":"An introduction to the InfiniBand\u2122 architecture","volume":"42","author":"Pfister Gregory F.","year":"2001","unstructured":"Gregory F. Pfister. 2001. An introduction to the InfiniBand\u2122 architecture. High Perform. Mass Stor. Parallel I\/O 42 (2001), 617\u2013632.","journal-title":"High Perform. Mass Stor. Parallel I\/O"},{"key":"e_1_3_1_74_2","volume-title":"High Performance Parallel I\/O","year":"2014","unstructured":"Prabhat and Quincey Koziol. 2014. High Performance Parallel I\/O. CRC Press."},{"key":"e_1_3_1_75_2","doi-asserted-by":"publisher","DOI":"10.1007\/s00450-009-0073-9"},{"key":"e_1_3_1_76_2","doi-asserted-by":"publisher","DOI":"10.1145\/3431379.3460640"},{"key":"e_1_3_1_77_2","doi-asserted-by":"publisher","DOI":"10.1016\/S1389-1286(00)00044-X"},{"key":"e_1_3_1_78_2","first-page":"231","volume-title":"1st USENIX Conference on File and Storage Technologies (FAST\u201902)","author":"Schmuck Frank B.","year":"2002","unstructured":"Frank B. Schmuck and Roger L. Haskin. 2002. GPFS: A shared-disk file system for large computing clusters. In 1st USENIX Conference on File and Storage Technologies (FAST\u201902). USENIX Association, 231\u2013244."},{"key":"e_1_3_1_79_2","doi-asserted-by":"publisher","DOI":"10.1109\/ESPT.2016.006"},{"key":"e_1_3_1_80_2","doi-asserted-by":"publisher","DOI":"10.1145\/2832087.2832091"},{"key":"e_1_3_1_81_2","doi-asserted-by":"publisher","DOI":"10.1109\/CCGrid.2011.26"},{"key":"e_1_3_1_82_2","doi-asserted-by":"publisher","DOI":"10.1145\/301816.301826"},{"key":"e_1_3_1_83_2","unstructured":"TOP500.org. 2022. TOP500 List. Retrieved from https:\/\/www.top500.org\/lists\/top500\/2022\/06\/"},{"key":"e_1_3_1_84_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-58667-0_17"},{"key":"e_1_3_1_85_2","doi-asserted-by":"publisher","DOI":"10.1287\/moor.2.2.143"},{"key":"e_1_3_1_86_2","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2010.5470424"},{"key":"e_1_3_1_87_2","first-page":"1","volume-title":"International Conference for High Performance Computing, Networking, Storage and Analysis (SC\u201917)","author":"Vazhkudai Sudharshan S.","year":"2017","unstructured":"Sudharshan S. Vazhkudai, Ross Miller, Devesh Tiwari, Christopher Zimmer, Feiyi Wang, Sarp Oral, Raghul Gunasekaran, and Deryl Steinert. 2017. GUIDE: A scalable information directory service to collect, federate, and analyze logs for operational insights into a leadership HPC facility. In International Conference for High Performance Computing, Networking, Storage and Analysis (SC\u201917). ACM, New York, NY, 1\u201312. DOI:10.1145\/3126908.3126946"},{"key":"e_1_3_1_88_2","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2019.00070"},{"issue":"1","key":"e_1_3_1_89_2","doi-asserted-by":"crossref","first-page":"11","DOI":"10.1145\/210308.210315","article-title":"The POSIX family of standards","volume":"3","author":"Walli Stephen R.","year":"1995","unstructured":"Stephen R. Walli. 1995. The POSIX family of standards. StandardView 3, 1 (Mar.1995), 11\u201317.","journal-title":"StandardView"},{"key":"e_1_3_1_90_2","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPSW50202.2020.00176"},{"key":"e_1_3_1_91_2","doi-asserted-by":"publisher","DOI":"10.1145\/2538542.2538562"},{"key":"e_1_3_1_92_2","first-page":"656","volume-title":"20th IEEE International Conference on Parallel and Distributed Systems (ICPADS\u201914)","author":"Wang Feiyi","year":"2014","unstructured":"Feiyi Wang, Sarp Oral, Saurabh Gupta, Devesh Tiwari, and Sudharshan S. Vazhkudai. 2014. Improving large-scale storage system performance via topology-aware and balanced data placement. In 20th IEEE International Conference on Parallel and Distributed Systems (ICPADS\u201914). IEEE, 656\u2013663. DOI:10.1109\/PADSW.2014.7097866"},{"key":"e_1_3_1_93_2","first-page":"307","volume-title":"7th Conference on Operating Systems Design and Implementation (OSDI\u201906)","author":"Weil Sage A.","year":"2006","unstructured":"Sage A. Weil, Scott A. Brandt, Ethan L. Miller, Darrell D. E. Long, and Carlos Maltzahn. 2006. Ceph: A scalable, high-performance distributed file system. In 7th Conference on Operating Systems Design and Implementation (OSDI\u201906). USENIX Association, 307\u2013320."},{"key":"e_1_3_1_94_2","doi-asserted-by":"publisher","unstructured":"Marc C. Wiedemann Julian M. Kunkel Michaela Zimmer Thomas Ludwig Michael Resch Thomas B\u00f6nisch Xuan Wang Andriy Chut Alvaro Aguilera Wolfgang E. Nagel Michael Kluge and Holger Mickler. 2013. Towards I\/O analysis of HPC systems and a generic architecture to collect access patterns. Computer Science-Research and Development 28 (2013) 241\u2013251. DOI:10.1007\/s00450-012-0221-5","DOI":"10.1007\/s00450-012-0221-5"},{"key":"e_1_3_1_95_2","doi-asserted-by":"publisher","DOI":"10.1093\/comjnl\/bxs044"},{"key":"e_1_3_1_96_2","first-page":"59","volume-title":"27th International ACM Conference on International Conference on Supercomputing (ICS\u201913)","author":"Wu Xing","year":"2013","unstructured":"Xing Wu and Frank Mueller. 2013. Elastic and scalable tracing and accurate replay of non-deterministic events. In 27th International ACM Conference on International Conference on Supercomputing (ICS\u201913). ACM, New York, NY, 59\u201368. DOI:10.1145\/2464996.2465001"},{"key":"e_1_3_1_97_2","first-page":"196","volume-title":"International Conference on Parallel Processing","author":"Wu Xing","year":"2011","unstructured":"Xing Wu, Karthik Vijayakumar, Frank Mueller, Xiaosong Ma, and Philip C. Roth. 2011. Probabilistic communication and I\/O tracing with deterministic replay at scale. In International Conference on Parallel Processing. IEEE, 196\u2013205. DOI:10.1109\/ICPP.2011.50"},{"key":"e_1_3_1_98_2","first-page":"2286","volume-title":"IEEE International Conference on Big Data (Big Data\u201916)","author":"Xenopoulos Peter","year":"2016","unstructured":"Peter Xenopoulos, Jamison Daniel, Michael Matheson, and Sreenivas Sukumar. 2016. Big data analytics on HPC architectures: Performance and cost. In IEEE International Conference on Big Data (Big Data\u201916). IEEE, 2286\u20132295. DOI:10.1109\/BigData.2016.7840861"},{"key":"e_1_3_1_99_2","doi-asserted-by":"publisher","unstructured":"Shaocheng Xie Wuyin Lin Philip J. Rasch Po-Lun Ma Richard Neale Vincent E. Larson Yun Qian Peter A. Bogenschutz Peter Caldwell Philip Cameron-Smith Jean-Christophe Golaz Salil Mahajan Balwinder Singh Qi Tang Hailong Wang Jin-Ho Yoon Kai Zhang and Yuying Zhang. 2018. Understanding cloud and convective characteristics in version 1 of the E3SM atmosphere model. Journal of Advances in Modeling Earth Systems 10 10 (2018) 2618\u20132644. DOI:10.1029\/2018MS001350","DOI":"10.1029\/2018MS001350"},{"key":"e_1_3_1_100_2","article-title":"LIOProf: Exposing Lustre File System Behavior for I\/O Middleware","author":"Xu Cong","year":"2016","unstructured":"Cong Xu, Suren Byna, Vishwanath Venkatesan, Robert Sisneros, Omkar Kulkarni, Mohamad Chaarawi, and Kalyana Chadalavada. 2016. LIOProf: Exposing Lustre File System Behavior for I\/O Middleware. InCray User Group Meeting (CUG\u201916).","journal-title":"In"},{"key":"e_1_3_1_101_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.parco.2016.08.001"},{"key":"e_1_3_1_102_2","first-page":"379","volume-title":"16th USENIX Symposium on Networked Systems Design and Implementation (NSDI\u201919)","author":"Yang Bin","year":"2019","unstructured":"Bin Yang, Xu Ji, Xiaosong Ma, Xiyang Wang, Tianyu Zhang, Xiupeng Zhu, Nosayba El-Sayed, Haidong Lan, Yibo Yang, Jidong Zhai, Weiguo Liu, and Wei Xue. 2019. End-to-end I\/O monitoring on a leading supercomputer. In 16th USENIX Symposium on Networked Systems Design and Implementation (NSDI\u201919). USENIX Association379\u2013394. Retrieved from https:\/\/www.usenix.org\/conference\/nsdi19\/presentation\/yang"},{"key":"e_1_3_1_103_2","doi-asserted-by":"publisher","DOI":"10.1145\/3568425"},{"key":"e_1_3_1_104_2","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS53621.2022.00128"},{"key":"e_1_3_1_105_2","unstructured":"Qian Yingjin Wang Di and Nirant Puntambekar. 2009. Lustre Simulator. Retrieved from https:\/\/github.com\/yingjinqian\/Lustre-Simulator"},{"key":"e_1_3_1_106_2","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2007.190694"},{"key":"e_1_3_1_107_2","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2004.1303013"},{"key":"e_1_3_1_108_2","first-page":"511","volume-title":"IEEE 10th International Conference on High Performance Computing and Communications and IEEE International Conference on Embedded and Ubiquitous Computing (HPCC & EUC\u201913)","author":"Zhu Mingfa","year":"2013","unstructured":"Mingfa Zhu, Guoying Li, Li Ruan, Ke Xie, and Limin Xiao. 2013. HySF: A striped file assignment strategy for parallel file system with hybrid storage. In IEEE 10th International Conference on High Performance Computing and Communications and IEEE International Conference on Embedded and Ubiquitous Computing (HPCC & EUC\u201913). IEEE, 511\u2013517. DOI:10.1109\/HPCC.and.EUC.2013.79"},{"key":"e_1_3_1_109_2","doi-asserted-by":"publisher","DOI":"10.13140\/RG.2.2.10671.92325"},{"key":"e_1_3_1_110_2","first-page":"581","volume-title":"IEEE International Conference on Cluster Computing (CLUSTER\u201922)","author":"Zhu Zhaobin","year":"2022","unstructured":"Zhaobin Zhu, Sarah Neuwirth, and Thomas Lippert. 2022. A comprehensive I\/O knowledge cycle for modular and automated HPC workload analysis. In IEEE International Conference on Cluster Computing (CLUSTER\u201922). IEEE, 581\u2013588. DOI:10.1109\/CLUSTER51413.2022.00076"},{"key":"e_1_3_1_111_2","unstructured":"Laura Zingaretti and Miguel P\u00e9rez-Enciso. 2022. Deep Learning for Genomic Prediction (DeepGP). Retrieved from https:\/\/github.com\/lauzingaretti\/DeepGP"}],"container-title":["ACM Transactions on Storage"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3641885","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3641885","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T00:04:03Z","timestamp":1750291443000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3641885"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,4,4]]},"references-count":110,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2024,5,31]]}},"alternative-id":["10.1145\/3641885"],"URL":"https:\/\/doi.org\/10.1145\/3641885","relation":{},"ISSN":["1553-3077","1553-3093"],"issn-type":[{"value":"1553-3077","type":"print"},{"value":"1553-3093","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,4,4]]},"assertion":[{"value":"2023-01-18","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-12-26","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-04-04","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}