{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,8,24]],"date-time":"2025-08-24T00:02:39Z","timestamp":1755993759177,"version":"3.44.0"},"publisher-location":"New York, NY, USA","reference-count":44,"publisher":"ACM","license":[{"start":{"date-parts":[[2024,7,8]],"date-time":"2024-07-08T00:00:00Z","timestamp":1720396800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100006374","name":"Alfred P. Sloan Foundation","doi-asserted-by":"publisher","id":[{"id":"10.13039\/501100006374","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100006374","name":"VMware","doi-asserted-by":"publisher","id":[{"id":"10.13039\/501100006374","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2024,7,8]]},"DOI":"10.1145\/3655038.3665943","type":"proceedings-article","created":{"date-parts":[[2024,6,27]],"date-time":"2024-06-27T00:19:48Z","timestamp":1719447588000},"page":"23-30","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Rethinking Erasure-Coding Libraries in the Age of Optimized Machine Learning"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-1899-9062","authenticated-orcid":false,"given":"Jiyu","family":"Hu","sequence":"first","affiliation":[{"name":"Carnegie Mellon University, Pittsburgh, Pennsylvania, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8812-7847","authenticated-orcid":false,"given":"Jack","family":"Kosaian","sequence":"additional","affiliation":[{"name":"Carnegie Mellon University, Pittsburgh, Pennsylvania, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2227-7460","authenticated-orcid":false,"given":"K. V.","family":"Rashmi","sequence":"additional","affiliation":[{"name":"Carnegie Mellon University, Pittsburgh, Pennsylvania, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2024,7,8]]},"reference":[{"volume-title":"Intel Galois Field New Instructions (GFNI) Technology Guide. https:\/\/tinyurl.com\/35w8pebx. Last accessed","year":"2022","key":"e_1_3_2_1_1_1","unstructured":"2021. Intel Galois Field New Instructions (GFNI) Technology Guide. https:\/\/tinyurl.com\/35w8pebx. Last accessed 12 December 2022."},{"volume-title":"Intel Intelligent Storage Acceleration Library. https:\/\/www.intel.com\/content\/www\/us\/en\/developer\/tools\/isa-l\/overview.html. Last accessed","year":"2022","key":"e_1_3_2_1_2_1","unstructured":"2021. Intel Intelligent Storage Acceleration Library. https:\/\/www.intel.com\/content\/www\/us\/en\/developer\/tools\/isa-l\/overview.html. Last accessed 17 September 2022."},{"key":"e_1_3_2_1_3_1","volume-title":"https:\/\/developer.nvidia.com\/gpudirect. Last accessed","author":"Direct NVIDIA","year":"2024","unstructured":"2024?. NVIDIA GPUDirect. https:\/\/developer.nvidia.com\/gpudirect. Last accessed 28 May 2024."},{"key":"e_1_3_2_1_4_1","volume-title":"Evaluating Multi-Level Checkpointing for Distributed Deep Neural Network Training. In 2021 SC Workshops Supplementary Proceedings (SCWS 21)","author":"Anthony Quentin","year":"2021","unstructured":"Quentin Anthony and Donglai Dai. 2021. Evaluating Multi-Level Checkpointing for Distributed Deep Neural Network Training. In 2021 SC Workshops Supplementary Proceedings (SCWS 21)."},{"key":"e_1_3_2_1_5_1","volume-title":"RAJA: Portable Performance for Large-Scale Scientific Applications. In 2019 IEEE\/ACM International Workshop on Performance, Portability and Productivity in HPC (P3HPC).","author":"Beckingsale David A","year":"2019","unstructured":"David A Beckingsale, Jason Burmark, Rich Hornung, Holger Jones, William Killian, Adam J Kunen, Olga Pearce, Peter Robinson, Brian S Ryujin, and Thomas RW Scogland. 2019. RAJA: Portable Performance for Large-Scale Scientific Applications. In 2019 IEEE\/ACM International Workshop on Performance, Portability and Productivity in HPC (P3HPC)."},{"key":"e_1_3_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/18.746771"},{"key":"e_1_3_2_1_7_1","unstructured":"Johannes Bloemer Malik Kalfane Richard Karp Marek Karpinski Michael Luby and David Zuckerman. 1995. An XOR-Based Erasure-Resilient Coding Scheme. Technical Report TR-95-048. University of California Berkeley."},{"key":"e_1_3_2_1_8_1","volume-title":"OpenCL-Based Erasure Coding on Heterogeneous Architectures. In 2016 IEEE 27th International Conference on Application-Specific Systems, Architectures and Processors (ASAP 16)","author":"Chen Guoyang","year":"2016","unstructured":"Guoyang Chen, Huiyang Zhou, Xipeng Shen, Josh Gahm, Narayan Venkat, Skip Booth, and John Marshall. 2016. OpenCL-Based Erasure Coding on Heterogeneous Architectures. In 2016 IEEE 27th International Conference on Application-Specific Systems, Architectures and Processors (ASAP 16)."},{"key":"e_1_3_2_1_9_1","volume-title":"TVM: An Automated End-to-End Optimizing Compiler for Deep Learning. In 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI 18)","author":"Chen Tianqi","year":"2018","unstructured":"Tianqi Chen, Thierry Moreau, Ziheng Jiang, Lianmin Zheng, Eddie Yan, Haichen Shen, Meghan Cowan, Leyuan Wang, Yuwei Hu, Luis Ceze, Carlos Guestrin, and Arvind Krishnamurthy. 2018. TVM: An Automated End-to-End Optimizing Compiler for Deep Learning. In 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI 18)."},{"key":"e_1_3_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1002\/cpe.1810"},{"key":"e_1_3_2_1_11_1","volume-title":"Proceedings of the USENIX Annual Technical Conference (ATC). 1--14","author":"Duplyakin Dmitry","year":"2019","unstructured":"Dmitry Duplyakin, Robert Ricci, Aleksander Maricq, Gary Wong, Jonathon Duerig, Eric Eide, Leigh Stoller, Mike Hibler, David Johnson, Kirk Webb, Aditya Akella, Kuangching Wang, Glenn Ricart, Larry Landweber, Chip Elliott, Michael Zink, Emmanuel Cecchet, Snigdhaswin Kar, and Prabodh Mishra. 2019. The Design and Operation of CloudLab. In Proceedings of the USENIX Annual Technical Conference (ATC). 1--14. https:\/\/www.flux.utah.edu\/paper\/duplyakin-atc19"},{"key":"e_1_3_2_1_12_1","volume-title":"Check-N-Run: A Checkpointing System for Training Deep Learning Recommendation Models. In 19th USENIX Symposium on Networked Systems Design and Implementation (NSDI 22)","author":"Eisenman Assaf","year":"2022","unstructured":"Assaf Eisenman, Kiran Kumar Matam, Steven Ingram, Dheevatsa Mudigere, Raghuraman Krishnamoorthi, Krishnakumar Nair, Misha Smelyanskiy, and Murali Annavaram. 2022. Check-N-Run: A Checkpointing System for Training Deep Learning Recommendation Models. In 19th USENIX Symposium on Networked Systems Design and Implementation (NSDI 22)."},{"key":"e_1_3_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1145\/945445.945450"},{"key":"e_1_3_2_1_14_1","volume-title":"Erasure Coding in Windows Azure Storage. In 2012 USENIX Annual Technical Conference (USENIX ATC 12)","author":"Huang Cheng","year":"2012","unstructured":"Cheng Huang, Huseyin Simitci, Yikang Xu, Aaron Ogus, Brad Calder, Parikshit Gopalan, Jin Li, and Sergey Yekhanin. 2012. Erasure Coding in Windows Azure Storage. In 2012 USENIX Annual Technical Conference (USENIX ATC 12)."},{"key":"e_1_3_2_1_15_1","volume-title":"Dissecting the NVIDIA Turing T4 GPU via Microbenchmarking. arXiv preprint arXiv:1903.07486","author":"Jia Zhe","year":"2019","unstructured":"Zhe Jia, Marco Maggioni, Jeffrey Smith, and Daniele Paolo Scarpazza. 2019. Dissecting the NVIDIA Turing T4 GPU via Microbenchmarking. arXiv preprint arXiv:1903.07486 (2019)."},{"key":"e_1_3_2_1_16_1","volume-title":"Tiger: Disk-Adaptive Redundancy Without Placement Restrictions. In 16th USENIX Symposium on Operating Systems Design and Implementation (OSDI 22)","author":"Kadekodi Saurabh","year":"2022","unstructured":"Saurabh Kadekodi, Francisco Maturana, Sanjith Athlur, Arif Merchant, KV Rashmi, and Gregory R Ganger. 2022. Tiger: Disk-Adaptive Redundancy Without Placement Restrictions. In 16th USENIX Symposium on Operating Systems Design and Implementation (OSDI 22)."},{"key":"e_1_3_2_1_17_1","volume-title":"PACEMAKER: Avoiding HeART Attacks in Storage Clusters with Disk-Adaptive Redundancy. In 14th USENIX Symposium on Operating Systems Design and Implementation (OSDI 20)","author":"Kadekodi Saurabh","year":"2020","unstructured":"Saurabh Kadekodi, Francisco Maturana, Suhas Jayaram Subramanya, Juncheng Yang, KV Rashmi, and Gregory R Ganger. 2020. PACEMAKER: Avoiding HeART Attacks in Storage Clusters with Disk-Adaptive Redundancy. In 14th USENIX Symposium on Operating Systems Design and Implementation (OSDI 20)."},{"key":"e_1_3_2_1_18_1","volume-title":"Boosting Full-Node Repair in Erasure-Coded Storage. In 2021 USENIX Annual Technical Conference (USENIX ATC 21)","author":"Lin Shiyao","year":"2021","unstructured":"Shiyao Lin, Guowen Gong, Zhirong Shen, Patrick PC Lee, and Jiwu Shu. 2021. Boosting Full-Node Repair in Erasure-Coded Storage. In 2021 USENIX Annual Technical Conference (USENIX ATC 21)."},{"key":"e_1_3_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2018.2791438"},{"key":"e_1_3_2_1_20_1","volume-title":"ECRM: Efficient Fault Tolerance for Recommendation Model Training via Erasure Coding. arXiv preprint arXiv:2104.01981","author":"Liu Kaige","year":"2021","unstructured":"Kaige Liu, Jack Kosaian, and KV Rashmi. 2021. ECRM: Efficient Fault Tolerance for Recommendation Model Training via Erasure Coding. arXiv preprint arXiv:2104.01981 (2021)."},{"key":"e_1_3_2_1_21_1","article-title":"Efficient Encoding Schedules for XOR-based Erasure Codes","author":"Luo J.","year":"2013","unstructured":"J. Luo, M. Shrestha, L. Xu, and J. S. Plank. 2013. Efficient Encoding Schedules for XOR-based Erasure Codes. IEEE Transactions on Computing (May 2013).","journal-title":"IEEE Transactions on Computing"},{"volume-title":"The Theory of Error Correcting Codes","author":"MacWilliams Florence Jessie","key":"e_1_3_2_1_22_1","unstructured":"Florence Jessie MacWilliams and Neil James Alexander Sloane. 1977. The Theory of Error Correcting Codes. Vol. 16. Elsevier."},{"key":"e_1_3_2_1_23_1","volume-title":"CPR: Understanding and Improving Failure Tolerant Training for Deep Learning Recommendation with Partial Recovery. In The Fourth Conference on Systems and Machine Learning (MLSys 21)","author":"Maeng Kiwan","year":"2021","unstructured":"Kiwan Maeng, Shivam Bharuka, Isabel Gao, Mark C Jeffrey, Vikram Saraph, Bor-Yiing Su, Caroline Trippel, Jiyan Yang, Mike Rabbat, Brandon Lucia, et al. 2021. CPR: Understanding and Improving Failure Tolerant Training for Deep Learning Recommendation with Partial Recovery. In The Fourth Conference on Systems and Machine Learning (MLSys 21)."},{"key":"e_1_3_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1109\/SC.2010.18"},{"key":"e_1_3_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/3458817.3476209"},{"key":"e_1_3_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2019.00099"},{"volume-title":"Proceedings of the ACM SIGMOD International Conference on Management of Data (SIGMOD 88)","author":"Patterson David A.","key":"e_1_3_2_1_27_1","unstructured":"David A. Patterson, Garth Gibson, and Randy H. Katz. 1988. A Case for Redundant Arrays of Inexpensive Disks (RAID). In Proceedings of the ACM SIGMOD International Conference on Management of Data (SIGMOD 88)."},{"key":"e_1_3_2_1_28_1","volume-title":"Jerasure: A Library in C Facilitating Erasure Coding for Storage Applications. Univ","author":"Plank J","year":"2014","unstructured":"J Plank and K Greenan. 2014. Jerasure: A Library in C Facilitating Erasure Coding for Storage Applications. Univ. Tennessee, Knoxville, TN, USA, Tech. Rep. CS-07-603 (2014)."},{"key":"e_1_3_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.5555\/1364813.1364820"},{"key":"e_1_3_2_1_30_1","volume-title":"Technical Report UT-CS-13-717. University of Tennessee.","author":"Plank J. S.","year":"2013","unstructured":"J. S. Plank, K. M. Greenan, and E. L. Miller. 2013. A Complete Treatment of Software Implementations of Finite Field Arithmetic for Erasure Coding Applications. Technical Report UT-CS-13-717. University of Tennessee."},{"key":"e_1_3_2_1_31_1","volume-title":"Screaming Fast Galois Field Arithmetic Using Intel SIMD Instructions. In 11th USENIX Conference on File and Storage Technologies (FAST 13)","author":"Plank James S","year":"2013","unstructured":"James S Plank, Kevin M Greenan, and Ethan L Miller. 2013. Screaming Fast Galois Field Arithmetic Using Intel SIMD Instructions. In 11th USENIX Conference on File and Storage Technologies (FAST 13)."},{"key":"e_1_3_2_1_32_1","first-page":"2571","article-title":"NetEC: Accelerating Erasure Coding Reconstruction With In-Network Aggregation","volume":"33","author":"Qiao Yi","year":"2022","unstructured":"Yi Qiao, Menghao Zhang, Yu Zhou, Xiao Kong, Han Zhang, Mingwei Xu, Jun Bi, and Jilong Wang. 2022. NetEC: Accelerating Erasure Coding Reconstruction With In-Network Aggregation. IEEE Transactions on Parallel and Distributed Systems 33, 10 (2022), 2571--2583.","journal-title":"IEEE Transactions on Parallel and Distributed Systems"},{"key":"e_1_3_2_1_33_1","volume-title":"Isaac Gelado, Seung Won Min, Amna Masood, Jeongmin Park, Jinjun Xiong, CJ Newburn, Dmitri Vainbrand, I Chung, et al.","author":"Qureshi Zaid","year":"2022","unstructured":"Zaid Qureshi, Vikram Sharma Mailthody, Isaac Gelado, Seung Won Min, Amna Masood, Jeongmin Park, Jinjun Xiong, CJ Newburn, Dmitri Vainbrand, I Chung, et al. 2022. BaM: A Case for Enabling Fine-grain High Throughput GPU-Orchestrated Access to Storage. arXiv preprint arXiv:2203.04910 (2022)."},{"key":"e_1_3_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1145\/2619239.2626325"},{"key":"e_1_3_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.14778\/2535573.2488339"},{"key":"e_1_3_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1145\/3295500.3356178"},{"key":"e_1_3_2_1_37_1","volume-title":"INEC: Fast and Coherent In-Network Erasure Coding. In International Conference for High Performance Computing, Networking, Storage and Analysis (SC 20)","author":"Shi Haiyang","year":"2020","unstructured":"Haiyang Shi and Xiaoyi Lu. 2020. INEC: Fast and Coherent In-Network Erasure Coding. In International Conference for High Performance Computing, Networking, Storage and Analysis (SC 20)."},{"key":"e_1_3_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2021.3097283"},{"key":"e_1_3_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1145\/3458817.3476204"},{"key":"e_1_3_2_1_40_1","volume-title":"Replication: A Quantitative Comparison. In International Workshop on Peer-to-Peer Systems (IPTPS","author":"Weatherspoon Hakim","year":"2002","unstructured":"Hakim Weatherspoon and John D Kubiatowicz. 2002. Erasure Coding vs. Replication: A Quantitative Comparison. In International Workshop on Peer-to-Peer Systems (IPTPS 2002)."},{"key":"e_1_3_2_1_41_1","volume-title":"Hubertus JJ Van Dam, and Chao Yang.","author":"Williams-Young David B","year":"2020","unstructured":"David B Williams-Young, Wibe A De Jong, Hubertus JJ Van Dam, and Chao Yang. 2020. On the Efficient Evaluation of the Exchange Correlation Potential on Graphics Processing Unit Clusters. Frontiers in chemistry 8 (2020), 581058."},{"key":"e_1_3_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1109\/SRDS51746.2020.00032"},{"key":"e_1_3_2_1_43_1","volume-title":"Ansor: Generating High-Performance Tensor Programs for Deep Learning. In 14th USENIX Symposium on Operating Systems Design and Implementation (OSDI 20)","author":"Zheng Lianmin","year":"2020","unstructured":"Lianmin Zheng, Chengfan Jia, Minmin Sun, Zhao Wu, Cody Hao Yu, Ameer Haj-Ali, Yida Wang, Jun Yang, Danyang Zhuo, Koushik Sen, et al. 2020. Ansor: Generating High-Performance Tensor Programs for Deep Learning. In 14th USENIX Symposium on Operating Systems Design and Implementation (OSDI 20)."},{"key":"e_1_3_2_1_44_1","volume-title":"Fast Erasure Coding for Data Storage: A Comprehensive Study of the Acceleration Techniques. In 17th USENIX Conference on File and Storage Technologies (FAST 19)","author":"Zhou Tianli","year":"2019","unstructured":"Tianli Zhou and Chao Tian. 2019. Fast Erasure Coding for Data Storage: A Comprehensive Study of the Acceleration Techniques. In 17th USENIX Conference on File and Storage Technologies (FAST 19)."}],"event":{"name":"HOTSTORAGE '24: 16th ACM Workshop on Hot Topics in Storage and File Systems","sponsor":["SIGOPS ACM Special Interest Group on Operating Systems"],"location":"Santa Clara CA USA","acronym":"HOTSTORAGE '24"},"container-title":["Proceedings of the 16th ACM Workshop on Hot Topics in Storage and File Systems"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3655038.3665943","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3655038.3665943","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,8,23]],"date-time":"2025-08-23T02:10:14Z","timestamp":1755915014000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3655038.3665943"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,7,8]]},"references-count":44,"alternative-id":["10.1145\/3655038.3665943","10.1145\/3655038"],"URL":"https:\/\/doi.org\/10.1145\/3655038.3665943","relation":{},"subject":[],"published":{"date-parts":[[2024,7,8]]},"assertion":[{"value":"2024-07-08","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}