{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,11]],"date-time":"2026-03-11T01:35:01Z","timestamp":1773192901961,"version":"3.50.1"},"reference-count":52,"publisher":"Association for Computing Machinery (ACM)","issue":"11","content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["Proc. VLDB Endow."],"published-print":{"date-parts":[[2023,7]]},"abstract":"<jats:p>Deep-learning-based recommendation models (DLRMs) are widely deployed to serve personalized content. In addition to using neural networks, DLRMs have large, sparsely-accessed embedding tables, which map categorical features to a learned dense representation. Due to the large sizes of embedding tables, DLRM training is typically distributed across the memory of tens or hundreds of nodes. Node failures are common in such large systems and must be mitigated to enable training to complete within production deadlines. Checkpointing is the primary approach used for fault tolerance in these systems, but incurs significant time overhead both during normal operation and when recovering from failures. As these overheads increase with DLRM size, checkpointing is slated to become an even larger overhead for future DLRMs, which are expected to grow. This calls for rethinking fault tolerance in DLRM training.<\/jats:p>\n          <jats:p>We present ECRec, a DLRM training system that achieves efficient fault tolerance by coupling erasure coding with the unique characteristics of DLRM training. ECRec takes a hybrid approach between erasure coding and replicating different DLRM parameters, correctly and efficiently updates redundant parameters, and enables training to proceed without pauses, while maintaining the consistency of the recovered parameters. We implement ECRec atop XDL, an open-source, industrial-scale DLRM training system. Compared to checkpointing, ECRec reduces training-time overhead on large DLRMs by up to 66%, recovers from failure up to 9.8\u00d7 faster, and continues training during recovery with only a 7--13% drop in throughput (whereas checkpointing must pause).<\/jats:p>","DOI":"10.14778\/3611479.3611514","type":"journal-article","created":{"date-parts":[[2023,8,25]],"date-time":"2023-08-25T02:08:08Z","timestamp":1692929288000},"page":"3137-3150","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":9,"title":["Efficient Fault Tolerance for Recommendation Model Training via Erasure Coding"],"prefix":"10.14778","volume":"16","author":[{"given":"Tianyu","family":"Zhang","sequence":"first","affiliation":[{"name":"Carnegie Mellon University, Pittsburgh, PA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Kaige","family":"Liu","sequence":"additional","affiliation":[{"name":"Carnegie Mellon University, Pittsburgh, PA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jack","family":"Kosaian","sequence":"additional","affiliation":[{"name":"Carnegie Mellon University, Pittsburgh, PA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Juncheng","family":"Yang","sequence":"additional","affiliation":[{"name":"Carnegie Mellon University, Pittsburgh, PA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Rashmi","family":"Vinayak","sequence":"additional","affiliation":[{"name":"Carnegie Mellon University, Pittsburgh, PA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2023,8,24]]},"reference":[{"key":"e_1_2_1_1_1","volume-title":"Download Terabyte Click Logs https:\/\/labs.criteo.com\/2013\/12\/download-terabyte-click-logs\/. Last accessed","author":"Labs Criteo","year":"2023","unstructured":"Criteo Labs : Download Terabyte Click Logs https:\/\/labs.criteo.com\/2013\/12\/download-terabyte-click-logs\/. Last accessed 10 July 2023 . Criteo Labs: Download Terabyte Click Logs https:\/\/labs.criteo.com\/2013\/12\/download-terabyte-click-logs\/. Last accessed 10 July 2023."},{"key":"e_1_2_1_2_1","volume-title":"A Training Framework Dedicated to Recommender Systems. https:\/\/tinyurl.com\/yy82pd2l. Last accessed","author":"Merlin Introducing NVIDIA","year":"2023","unstructured":"Introducing NVIDIA Merlin HugeCTR : A Training Framework Dedicated to Recommender Systems. https:\/\/tinyurl.com\/yy82pd2l. Last accessed 10 July 2023 . Introducing NVIDIA Merlin HugeCTR: A Training Framework Dedicated to Recommender Systems. https:\/\/tinyurl.com\/yy82pd2l. Last accessed 10 July 2023."},{"key":"e_1_2_1_3_1","volume-title":"https:\/\/www.kaggle.com\/c\/avazu-ctr-prediction. Last accessed","author":"Prediction Contest Kaggle Avazu CTR","year":"2023","unstructured":"Kaggle Avazu CTR Prediction Contest . https:\/\/www.kaggle.com\/c\/avazu-ctr-prediction. Last accessed 10 July 2023 . Kaggle Avazu CTR Prediction Contest. https:\/\/www.kaggle.com\/c\/avazu-ctr-prediction. Last accessed 10 July 2023."},{"key":"e_1_2_1_4_1","volume-title":"https:\/\/github.com\/mlperf\/inference. Last accessed","author":"Github Repository Perf Inference","year":"2023","unstructured":"ML Perf Inference Github Repository . https:\/\/github.com\/mlperf\/inference. Last accessed 10 July 2023 . MLPerf Inference Github Repository. https:\/\/github.com\/mlperf\/inference. Last accessed 10 July 2023."},{"key":"e_1_2_1_5_1","volume-title":"Xiaoqiang Zheng. TensorFlow: A System for Large-Scale Machine Learning. In 12th USENIX Symposium on Operating Systems Design and Implementation (OSDI 16)","author":"Abadi Mart\u00edn","year":"2016","unstructured":"Mart\u00edn Abadi , Paul Barham , Jianmin Chen , Zhifeng Chen , Andy Davis , Jeffrey Dean , Matthieu Devin , Sanjay Ghemawat , Geoffrey Irving , Michael Isard , Manjunath Kudlur , Josh Levenberg , Rajat Monga , Sherry Moore , Derek G. Murray , Benoit Steiner , Paul Tucker , Vijay Vasudevan , Pete Warden , Martin Wicke , Yuan Yu , and Xiaoqiang Zheng. TensorFlow: A System for Large-Scale Machine Learning. In 12th USENIX Symposium on Operating Systems Design and Implementation (OSDI 16) , 2016 . Mart\u00edn Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, Manjunath Kudlur, Josh Levenberg, Rajat Monga, Sherry Moore, Derek G. Murray, Benoit Steiner, Paul Tucker, Vijay Vasudevan, Pete Warden, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng. TensorFlow: A System for Large-Scale Machine Learning. In 12th USENIX Symposium on Operating Systems Design and Implementation (OSDI 16), 2016."},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA51647.2021.00072"},{"key":"e_1_2_1_7_1","volume-title":"Divya Mahajan, and Prashant J. Nair. Accelerating Recommendation System Training by Leveraging Popular Choices","author":"Adnan Muhammad","year":"2022","unstructured":"Muhammad Adnan , Yassaman Ebrahimzadeh Maboud , Divya Mahajan, and Prashant J. Nair. Accelerating Recommendation System Training by Leveraging Popular Choices , 2022 . Muhammad Adnan, Yassaman Ebrahimzadeh Maboud, Divya Mahajan, and Prashant J. Nair. Accelerating Recommendation System Training by Leveraging Popular Choices, 2022."},{"key":"e_1_2_1_8_1","volume-title":"BagPipe: Accelerating Deep Recommendation Model Training. arXiv preprint arXiv:2202.12429","author":"Agarwal Saurabh","year":"2022","unstructured":"Saurabh Agarwal , Ziyi Zhang , and Shivaram Venkataraman . BagPipe: Accelerating Deep Recommendation Model Training. arXiv preprint arXiv:2202.12429 , 2022 . Saurabh Agarwal, Ziyi Zhang, and Shivaram Venkataraman. BagPipe: Accelerating Deep Recommendation Model Training. arXiv preprint arXiv:2202.12429, 2022."},{"key":"e_1_2_1_9_1","volume-title":"Xin Jin. On Efficient Constructions of Checkpoints. In Proceedings of the International Conference on Machine Learning (ICML 20)","author":"Chen Yu","year":"2020","unstructured":"Yu Chen , Zhenming Liu , Bin Ren , and Xin Jin. On Efficient Constructions of Checkpoints. In Proceedings of the International Conference on Machine Learning (ICML 20) , 2020 . Yu Chen, Zhenming Liu, Bin Ren, and Xin Jin. On Efficient Constructions of Checkpoints. In Proceedings of the International Conference on Machine Learning (ICML 20), 2020."},{"key":"e_1_2_1_10_1","volume-title":"Emre Sargin. Deep Neural Networks for YouTube Recommendations. In Proceedings of the 10th ACM Conference on Recommender Systems","author":"Covington Paul","year":"2016","unstructured":"Paul Covington , Jay Adams , and Emre Sargin. Deep Neural Networks for YouTube Recommendations. In Proceedings of the 10th ACM Conference on Recommender Systems , 2016 . Paul Covington, Jay Adams, and Emre Sargin. Deep Neural Networks for YouTube Recommendations. In Proceedings of the 10th ACM Conference on Recommender Systems, 2016."},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.5555\/1134241.1708449"},{"key":"e_1_2_1_12_1","volume-title":"Kingma and Jimmy Ba. Adam: A Method for Stochastic Optimization. In International Conference on Learning Representations (ICLR 15)","author":"Diederik","year":"2015","unstructured":"Diederik P. Kingma and Jimmy Ba. Adam: A Method for Stochastic Optimization. In International Conference on Learning Representations (ICLR 15) , 2015 . Diederik P. Kingma and Jimmy Ba. Adam: A Method for Stochastic Optimization. In International Conference on Learning Representations (ICLR 15), 2015."},{"issue":"7","key":"e_1_2_1_13_1","article-title":"Adaptive Subgradient Methods for Online Learning and Stochastic Optimization","volume":"12","author":"Duchi John","year":"2011","unstructured":"John Duchi , Elad Hazan , and Yoram Singer . Adaptive Subgradient Methods for Online Learning and Stochastic Optimization . Journal of Machine Learning Research , 12 ( 7 ), 2011 . John Duchi, Elad Hazan, and Yoram Singer. Adaptive Subgradient Methods for Online Learning and Stochastic Optimization. Journal of Machine Learning Research, 12(7), 2011.","journal-title":"Journal of Machine Learning Research"},{"key":"e_1_2_1_14_1","volume-title":"Prabodh Mishra. The Design and Operation of CloudLab. In 2019 USENIX Annual Technical Conference (USENIX ATC 19)","author":"Duplyakin Dmitry","year":"2019","unstructured":"Dmitry Duplyakin , Robert Ricci , Aleksander Maricq , Gary Wong , Jonathon Duerig , Eric Eide , Leigh Stoller , Mike Hibler , David Johnson , Kirk Webb , Aditya Akella , Kuangching Wang , Glenn Ricart , Larry Landweber , Chip Elliott , Michael Zink , Emmanuel Cecchet , Snigdhaswin Kar , and Prabodh Mishra. The Design and Operation of CloudLab. In 2019 USENIX Annual Technical Conference (USENIX ATC 19) , 2019 . Dmitry Duplyakin, Robert Ricci, Aleksander Maricq, Gary Wong, Jonathon Duerig, Eric Eide, Leigh Stoller, Mike Hibler, David Johnson, Kirk Webb, Aditya Akella, Kuangching Wang, Glenn Ricart, Larry Landweber, Chip Elliott, Michael Zink, Emmanuel Cecchet, Snigdhaswin Kar, and Prabodh Mishra. The Design and Operation of CloudLab. In 2019 USENIX Annual Technical Conference (USENIX ATC 19), 2019."},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISIT.2018.8437852"},{"key":"e_1_2_1_16_1","volume-title":"Mu-rali Annavaram. Check-N-Run: A Checkpointing System for Training Deep Learning Recommendation Models. In 19th USENIX Symposium on Networked Systems Design and Implementation (NSDI 22)","author":"Eisenman Assaf","year":"2022","unstructured":"Assaf Eisenman , Kiran Kumar Matam , Steven Ingram , Dheevatsa Mudigere , Raghuraman Krishnamoorthi , Krishnakumar Nair , Misha Smelyanskiy , and Mu-rali Annavaram. Check-N-Run: A Checkpointing System for Training Deep Learning Recommendation Models. In 19th USENIX Symposium on Networked Systems Design and Implementation (NSDI 22) , 2022 . Assaf Eisenman, Kiran Kumar Matam, Steven Ingram, Dheevatsa Mudigere, Raghuraman Krishnamoorthi, Krishnakumar Nair, Misha Smelyanskiy, and Mu-rali Annavaram. Check-N-Run: A Checkpointing System for Training Deep Learning Recommendation Models. In 19th USENIX Symposium on Networked Systems Design and Implementation (NSDI 22), 2022."},{"key":"e_1_2_1_17_1","volume-title":"Sachin Katti. Bandana: Using NonVolatile Memory for Storing Deep Learning Models. In The Second Conference on Systems and Machine Learning (SysML 19)","author":"Eisenman Assaf","year":"2019","unstructured":"Assaf Eisenman , Maxim Naumov , Darryl Gardner , Misha Smelyanskiy , Sergey Pupyrev , Kim Hazelwood , Asaf Cidon , and Sachin Katti. Bandana: Using NonVolatile Memory for Storing Deep Learning Models. In The Second Conference on Systems and Machine Learning (SysML 19) , 2019 . Assaf Eisenman, Maxim Naumov, Darryl Gardner, Misha Smelyanskiy, Sergey Pupyrev, Kim Hazelwood, Asaf Cidon, and Sachin Katti. Bandana: Using NonVolatile Memory for Storing Deep Learning Models. In The Second Conference on Systems and Machine Learning (SysML 19), 2019."},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISIT45174.2021.9517710"},{"key":"e_1_2_1_19_1","volume-title":"Carole-Jean Wu. DeepRecSys: A System for Optimizing End-to-End At-Scale Neural Recommendation Inference. In 2020 ACM\/IEEE 47th Annual International Symposium on Computer Architecture (ISCA 20)","author":"Gupta Udit","year":"2020","unstructured":"Udit Gupta , Samuel Hsia , Vikram Saraph , Xiaodong Wang , Brandon Reagen , Gu-Yeon Wei , Hsien-Hsin S Lee , David Brooks , and Carole-Jean Wu. DeepRecSys: A System for Optimizing End-to-End At-Scale Neural Recommendation Inference. In 2020 ACM\/IEEE 47th Annual International Symposium on Computer Architecture (ISCA 20) , 2020 . Udit Gupta, Samuel Hsia, Vikram Saraph, Xiaodong Wang, Brandon Reagen, Gu-Yeon Wei, Hsien-Hsin S Lee, David Brooks, and Carole-Jean Wu. DeepRecSys: A System for Optimizing End-to-End At-Scale Neural Recommendation Inference. In 2020 ACM\/IEEE 47th Annual International Symposium on Computer Architecture (ISCA 20), 2020."},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA47549.2020.00047"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/3326937.3341255"},{"key":"e_1_2_1_22_1","volume-title":"The Fourth Conference on Systems and Machine Learning (MLSys 21)","author":"Jiang Wenqi","year":"2021","unstructured":"Wenqi Jiang , Zhenhao He , Shuai Zhang , Thomas B Preu\u00dfer , Kai Zeng , Liang Feng , Jiansong Zhang , Tongxuan Liu , Yong Li , Jingren Zhou , : Accelerating Deep Recommendation Systems to Microseconds by Hardware and Data Structure Solutions . In The Fourth Conference on Systems and Machine Learning (MLSys 21) , 2021 . Wenqi Jiang, Zhenhao He, Shuai Zhang, Thomas B Preu\u00dfer, Kai Zeng, Liang Feng, Jiansong Zhang, Tongxuan Liu, Yong Li, Jingren Zhou, et al. MicroRec: Accelerating Deep Recommendation Systems to Microseconds by Hardware and Data Structure Solutions. In The Fourth Conference on Systems and Machine Learning (MLSys 21), 2021."},{"key":"e_1_2_1_23_1","volume-title":"Proceedings of the 14th USENIX Conference on Operating Systems Design and Implementation, OSDI'20, USA","author":"Kadekodi Saurabh","year":"2020","unstructured":"Saurabh Kadekodi , Francisco Maturana , Suhas Jayaram Subramanya , Juncheng Yang , K. V. Rashmi , and Gregory R. Ganger . Pacemaker: Avoiding heart attacks in storage clusters with disk-adaptive redundancy . In Proceedings of the 14th USENIX Conference on Operating Systems Design and Implementation, OSDI'20, USA , 2020 . USENIX Association. Saurabh Kadekodi, Francisco Maturana, Suhas Jayaram Subramanya, Juncheng Yang, K. V. Rashmi, and Gregory R. Ganger. Pacemaker: Avoiding heart attacks in storage clusters with disk-adaptive redundancy. In Proceedings of the 14th USENIX Conference on Operating Systems Design and Implementation, OSDI'20, USA, 2020. USENIX Association."},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.5555\/3433701.3433758"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1109\/TSE.1987.232562"},{"key":"e_1_2_1_26_1","volume-title":"Shivaram Venkataraman. Parity Models: Erasure-Coded Resilience for Prediction Serving Systems. In Proceedings of the 27th ACM Symposium on Operating Systems Principles (SOSP 19)","author":"Kosaian Jack","year":"2019","unstructured":"Jack Kosaian , K. V. Rashmi , and Shivaram Venkataraman. Parity Models: Erasure-Coded Resilience for Prediction Serving Systems. In Proceedings of the 27th ACM Symposium on Operating Systems Principles (SOSP 19) , 2019 . Jack Kosaian, K. V. Rashmi, and Shivaram Venkataraman. Parity Models: Erasure-Coded Resilience for Prediction Serving Systems. In Proceedings of the 27th ACM Symposium on Operating Systems Principles (SOSP 19), 2019."},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIT.2017.2736066"},{"key":"e_1_2_1_28_1","volume-title":"Mark Hempstead. Understanding Capacity-Driven Scale-Out Neural Recommendation Inference. In 2021 IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS 21)","author":"Lui Michael","year":"2021","unstructured":"Michael Lui , Yavuz Yetim , \u00d6zg\u00fcr \u00d6zkan , Zhuoran Zhao , Shin-Yeh Tsai , Carole-Jean Wu , and Mark Hempstead. Understanding Capacity-Driven Scale-Out Neural Recommendation Inference. In 2021 IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS 21) , 2021 . Michael Lui, Yavuz Yetim, \u00d6zg\u00fcr \u00d6zkan, Zhuoran Zhao, Shin-Yeh Tsai, Carole-Jean Wu, and Mark Hempstead. Understanding Capacity-Driven Scale-Out Neural Recommendation Inference. In 2021 IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS 21), 2021."},{"key":"e_1_2_1_29_1","volume-title":"CPR: Understanding and Improving Failure Tolerant Training for Deep Learning Recommendation with Partial Recovery. In The Fourth Conference on Systems and Machine Learning (MLSys 21)","author":"Maeng Kiwan","year":"2021","unstructured":"Kiwan Maeng , Shivam Bharuka , Isabel Gao , Mark C Jeffrey , Vikram Saraph , Bor-Yiing Su , Caroline Trippel , Jiyan Yang , Mike Rabbat , Brandon Lucia , CPR: Understanding and Improving Failure Tolerant Training for Deep Learning Recommendation with Partial Recovery. In The Fourth Conference on Systems and Machine Learning (MLSys 21) , 2021 . Kiwan Maeng, Shivam Bharuka, Isabel Gao, Mark C Jeffrey, Vikram Saraph, Bor-Yiing Su, Caroline Trippel, Jiyan Yang, Mike Rabbat, Brandon Lucia, et al. CPR: Understanding and Improving Failure Tolerant Training for Deep Learning Recommendation with Partial Recovery. In The Fourth Conference on Systems and Machine Learning (MLSys 21), 2021."},{"key":"e_1_2_1_30_1","volume-title":"MLPerf Training Benchmark. In The Third Conference on Systems and Machine Learning (MLSys 20)","author":"Mattson Peter","year":"2020","unstructured":"Peter Mattson , Christine Cheng , Cody Coleman , Greg Diamos , Paulius Micikevicius , David Patterson , Hanlin Tang , Gu-Yeon Wei , Peter Bailis , Victor Bittorf , MLPerf Training Benchmark. In The Third Conference on Systems and Machine Learning (MLSys 20) , 2020 . Peter Mattson, Christine Cheng, Cody Coleman, Greg Diamos, Paulius Micikevicius, David Patterson, Hanlin Tang, Gu-Yeon Wei, Peter Bailis, Victor Bittorf, et al. MLPerf Training Benchmark. In The Third Conference on Systems and Machine Learning (MLSys 20), 2020."},{"key":"e_1_2_1_31_1","volume-title":"Fine-Grained DNN Checkpointing. In 19th USENIX Conference on File and Storage Technologies (FAST 21)","author":"Mohan Jayashree","year":"2021","unstructured":"Jayashree Mohan , Amar Phanishayee , and Vijay Chidambaram . CheckFreq : Frequent , Fine-Grained DNN Checkpointing. In 19th USENIX Conference on File and Storage Technologies (FAST 21) , 2021 . Jayashree Mohan, Amar Phanishayee, and Vijay Chidambaram. CheckFreq: Frequent, Fine-Grained DNN Checkpointing. In 19th USENIX Conference on File and Storage Technologies (FAST 21), 2021."},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1109\/SC.2010.18"},{"key":"e_1_2_1_33_1","volume-title":"Software-Hardware Co-design for Fast and Scalable Training of Deep Learning Recommendation Models. arXiv preprint arXiv:2104.05158","author":"Mudigere Dheevatsa","year":"2021","unstructured":"Dheevatsa Mudigere , Yuchen Hao , Jianyu Huang , Zhihao Jia , Andrew Tulloch , Srinivas Sridharan , Xing Liu , Mustafa Ozdal , Jade Nie , Jongsoo Park , Software-Hardware Co-design for Fast and Scalable Training of Deep Learning Recommendation Models. arXiv preprint arXiv:2104.05158 , 2021 . Dheevatsa Mudigere, Yuchen Hao, Jianyu Huang, Zhihao Jia, Andrew Tulloch, Srinivas Sridharan, Xing Liu, Mustafa Ozdal, Jade Nie, Jongsoo Park, et al. Software-Hardware Co-design for Fast and Scalable Training of Deep Learning Recommendation Models. arXiv preprint arXiv:2104.05158, 2021."},{"key":"e_1_2_1_34_1","volume-title":"Deep Learning Training in Facebook Data Centers: Design of Scale-up and Scale-out Systems. arXiv preprint arXiv:2003.09518","author":"Naumov Maxim","year":"2020","unstructured":"Maxim Naumov , John Kim , Dheevatsa Mudigere , Srinivas Sridharan , Xiaodong Wang , Whitney Zhao , Serhat Yilmaz , Changkyu Kim , Hector Yuen , Mustafa Ozdal , Deep Learning Training in Facebook Data Centers: Design of Scale-up and Scale-out Systems. arXiv preprint arXiv:2003.09518 , 2020 . Maxim Naumov, John Kim, Dheevatsa Mudigere, Srinivas Sridharan, Xiaodong Wang, Whitney Zhao, Serhat Yilmaz, Changkyu Kim, Hector Yuen, Mustafa Ozdal, et al. Deep Learning Training in Facebook Data Centers: Design of Scale-up and Scale-out Systems. arXiv preprint arXiv:2003.09518, 2020."},{"key":"e_1_2_1_35_1","unstructured":"Maxim Naumov Dheevatsa Mudigere Hao-Jun Michael Shi Jianyu Huang Narayanan Sundaraman Jongsoo Park Xiaodong Wang Udit Gupta Carole-Jean Wu Alisson G Azzolini et al. Deep Learning Recommendation Model for Personalization and Recommendation Systems. arXiv preprint arXiv:1906.00091 2019.  Maxim Naumov Dheevatsa Mudigere Hao-Jun Michael Shi Jianyu Huang Narayanan Sundaraman Jongsoo Park Xiaodong Wang Udit Gupta Carole-Jean Wu Alisson G Azzolini et al. Deep Learning Recommendation Model for Personalization and Recommendation Systems. arXiv preprint arXiv:1906.00091 2019."},{"key":"e_1_2_1_36_1","volume-title":"Franck Cappello. DeepFreeze: Towards Scalable Asynchronous Checkpointing of Deep Learning Models. In 2020 20th IEEE\/ACM International Symposium on Cluster, Cloud and Internet Computing (CCGrid 20)","author":"Nicolae Bogdan","year":"2020","unstructured":"Bogdan Nicolae , Jiali Li , Justin M Wozniak , George Bosilca , Matthieu Dorier , and Franck Cappello. DeepFreeze: Towards Scalable Asynchronous Checkpointing of Deep Learning Models. In 2020 20th IEEE\/ACM International Symposium on Cluster, Cloud and Internet Computing (CCGrid 20) , 2020 . Bogdan Nicolae, Jiali Li, Justin M Wozniak, George Bosilca, Matthieu Dorier, and Franck Cappello. DeepFreeze: Towards Scalable Asynchronous Checkpointing of Deep Learning Models. In 2020 20th IEEE\/ACM International Symposium on Cluster, Cloud and Internet Computing (CCGrid 20), 2020."},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1145\/50202.50214"},{"key":"e_1_2_1_38_1","first-page":"5220","volume-title":"Eric Xing. Fault Tolerance in Iterative-Convergent Machine Learning. In International Conference on Machine Learning","author":"Qiao Aurick","year":"2019","unstructured":"Aurick Qiao , Bryon Aragam , Bingjing Zhang , and Eric Xing. Fault Tolerance in Iterative-Convergent Machine Learning. In International Conference on Machine Learning , pages 5220 -- 5230 , 2019 . Aurick Qiao, Bryon Aragam, Bingjing Zhang, and Eric Xing. Fault Tolerance in Iterative-Convergent Machine Learning. In International Conference on Machine Learning, pages 5220--5230, 2019."},{"key":"e_1_2_1_39_1","volume-title":"Low-Latency Cluster Caching with Online Erasure Coding. In 12th USENIX Symposium on Operating Systems Design and Implementation (OSDI 16)","author":"Rashmi K. V.","year":"2016","unstructured":"K. V. Rashmi , Mosharaf Chowdhury , Jack Kosaian , Ion Stoica , and Kannan Ramchandran . EC-Cache : Load-Balanced , Low-Latency Cluster Caching with Online Erasure Coding. In 12th USENIX Symposium on Operating Systems Design and Implementation (OSDI 16) , 2016 . K. V. Rashmi, Mosharaf Chowdhury, Jack Kosaian, Ion Stoica, and Kannan Ramchandran. EC-Cache: Load-Balanced, Low-Latency Cluster Caching with Online Erasure Coding. In 12th USENIX Symposium on Operating Systems Design and Implementation (OSDI 16), 2016."},{"key":"e_1_2_1_40_1","volume-title":"Fast and Efficient Data Reconstruction in Erasure-Coded Data Centers. In Proceedings of the 2014 ACM SIGCOMM Conference (SIGCOMM 14)","author":"Rashmi K. V.","year":"2014","unstructured":"K. V. Rashmi , Nihar B Shah , Dikang Gu , Hairong Kuang , Dhruba Borthakur , and Kannan Ramchandran . A Hitchhiker's Guide to Fast and Efficient Data Reconstruction in Erasure-Coded Data Centers. In Proceedings of the 2014 ACM SIGCOMM Conference (SIGCOMM 14) , 2014 . K. V. Rashmi, Nihar B Shah, Dikang Gu, Hairong Kuang, Dhruba Borthakur, and Kannan Ramchandran. A Hitchhiker's Guide to Fast and Efficient Data Reconstruction in Erasure-Coded Data Centers. In Proceedings of the 2014 ACM SIGCOMM Conference (SIGCOMM 14), 2014."},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1145\/263876.263881"},{"key":"e_1_2_1_42_1","volume-title":"Proceedings of the VLDB Endowment, 6(5)","author":"Sathiamoorthy Mahesh","year":"2013","unstructured":"Mahesh Sathiamoorthy , Megasthenis Asteris , Dimitris Papailiopoulos , Alexan-dros G Dimakis , Ramkumar Vadali , Scott Chen , and Dhruba Borthakur. XORing Elephants: Novel Erasure Codes for Big Data . Proceedings of the VLDB Endowment, 6(5) , 2013 . Mahesh Sathiamoorthy, Megasthenis Asteris, Dimitris Papailiopoulos, Alexan-dros G Dimakis, Ramkumar Vadali, Scott Chen, and Dhruba Borthakur. XORing Elephants: Novel Erasure Codes for Big Data. Proceedings of the VLDB Endowment, 6(5), 2013."},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1145\/3503222.3507777"},{"key":"e_1_2_1_44_1","volume-title":"Nikos Karampatziakis. Gradient Coding: Avoiding Stragglers in Distributed Learning. In International Conference on Machine Learning (ICML 17)","author":"Tandon Rashish","year":"2017","unstructured":"Rashish Tandon , Qi Lei , Alexandros G Dimakis , and Nikos Karampatziakis. Gradient Coding: Avoiding Stragglers in Distributed Learning. In International Conference on Machine Learning (ICML 17) , 2017 . Rashish Tandon, Qi Lei, Alexandros G Dimakis, and Nikos Karampatziakis. Gradient Coding: Avoiding Stragglers in Distributed Learning. In International Conference on Machine Learning (ICML 17), 2017."},{"key":"e_1_2_1_45_1","volume-title":"Replication: A Quantitative Comparison. In International Workshop on Peer-to-Peer Systems (IPTPS 2002)","author":"Weatherspoon Hakim","year":"2002","unstructured":"Hakim Weatherspoon and John D Kubiatowicz . Erasure Coding vs . Replication: A Quantitative Comparison. In International Workshop on Peer-to-Peer Systems (IPTPS 2002) , 2002 . Hakim Weatherspoon and John D Kubiatowicz. Erasure Coding vs. Replication: A Quantitative Comparison. In International Workshop on Peer-to-Peer Systems (IPTPS 2002), 2002."},{"key":"e_1_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.1145\/3445814.3446763"},{"key":"e_1_2_1_47_1","volume-title":"Haryadi S Gunawi. Tiny-Tail Flash: Near-Perfect Elimination of Garbage Collection Tail Latencies in NAND SSDs. In 15th USENIX Conference on File and Storage Technologies (FAST 17)","author":"Yan Shiqin","year":"2017","unstructured":"Shiqin Yan , Huaicheng Li , Mingzhe Hao , Michael Hao Tong , Swaminathan Sundararaman , Andrew A Chien , and Haryadi S Gunawi. Tiny-Tail Flash: Near-Perfect Elimination of Garbage Collection Tail Latencies in NAND SSDs. In 15th USENIX Conference on File and Storage Technologies (FAST 17) , 2017 . Shiqin Yan, Huaicheng Li, Mingzhe Hao, Michael Hao Tong, Swaminathan Sundararaman, Andrew A Chien, and Haryadi S Gunawi. Tiny-Tail Flash: Near-Perfect Elimination of Garbage Collection Tail Latencies in NAND SSDs. In 15th USENIX Conference on File and Storage Technologies (FAST 17), 2017."},{"key":"e_1_2_1_48_1","volume-title":"Ping Tak Peter Tang, and Andrew Tulloch. Mixed-Precision Embedding Using a Cache. arXiv preprint arXiv:2010.11305","author":"Yang Jie Amy","year":"2020","unstructured":"Jie Amy Yang , Jianyu Huang , Jongsoo Park , Ping Tak Peter Tang, and Andrew Tulloch. Mixed-Precision Embedding Using a Cache. arXiv preprint arXiv:2010.11305 , 2020 . Jie Amy Yang, Jianyu Huang, Jongsoo Park, Ping Tak Peter Tang, and Andrew Tulloch. Mixed-Precision Embedding Using a Cache. arXiv preprint arXiv:2010.11305, 2020."},{"key":"e_1_2_1_49_1","first-page":"1159","volume-title":"19th USENIX Symposium on Networked Systems Design and Implementation (NSDI 22)","author":"Yang Juncheng","year":"2022","unstructured":"Juncheng Yang , Anirudh Sabnis , Daniel S. Berger , K. V. Rashmi , and Ramesh K. Sitaraman . C2DN: How to harness erasure codes at the edge for efficient content delivery . In 19th USENIX Symposium on Networked Systems Design and Implementation (NSDI 22) , pages 1159 -- 1177 , Renton, WA , April 2022 . USENIX Association. Juncheng Yang, Anirudh Sabnis, Daniel S. Berger, K. V. Rashmi, and Ramesh K. Sitaraman. C2DN: How to harness erasure codes at the edge for efficient content delivery. In 19th USENIX Symposium on Networked Systems Design and Implementation (NSDI 22), pages 1159--1177, Renton, WA, April 2022. USENIX Association."},{"key":"e_1_2_1_50_1","volume-title":"Carole-Jean Wu. TT-Rec: Tensor Train Compression for Deep Learning Recommendation Models. In The Fourth Conference on Systems and Machine Learning (MLSys 21)","author":"Yin Chunxing","year":"2021","unstructured":"Chunxing Yin , Bilge Acun , Xing Liu , and Carole-Jean Wu. TT-Rec: Tensor Train Compression for Deep Learning Recommendation Models. In The Fourth Conference on Systems and Machine Learning (MLSys 21) , 2021 . Chunxing Yin, Bilge Acun, Xing Liu, and Carole-Jean Wu. TT-Rec: Tensor Train Compression for Deep Learning Recommendation Models. In The Fourth Conference on Systems and Machine Learning (MLSys 21), 2021."},{"key":"e_1_2_1_51_1","volume-title":"Proceedings of the 22nd International Conference on Artificial Intelligence and Statistics (AISTATS 19)","author":"Yu Qian","year":"2019","unstructured":"Qian Yu , Netanel Raviv , Jinhyun So , and A Salman Avestimehr . Lagrange Coded Computing: Optimal Design for Resiliency, Security and Privacy . In Proceedings of the 22nd International Conference on Artificial Intelligence and Statistics (AISTATS 19) , 2019 . Qian Yu, Netanel Raviv, Jinhyun So, and A Salman Avestimehr. Lagrange Coded Computing: Optimal Design for Resiliency, Security and Privacy. In Proceedings of the 22nd International Conference on Artificial Intelligence and Statistics (AISTATS 19), 2019."},{"key":"e_1_2_1_52_1","doi-asserted-by":"publisher","DOI":"10.1145\/3357384.3358045"}],"container-title":["Proceedings of the VLDB Endowment"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.14778\/3611479.3611514","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,9,23]],"date-time":"2023-09-23T22:20:53Z","timestamp":1695507653000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.14778\/3611479.3611514"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,7]]},"references-count":52,"journal-issue":{"issue":"11","published-print":{"date-parts":[[2023,7]]}},"alternative-id":["10.14778\/3611479.3611514"],"URL":"https:\/\/doi.org\/10.14778\/3611479.3611514","relation":{},"ISSN":["2150-8097"],"issn-type":[{"value":"2150-8097","type":"print"}],"subject":[],"published":{"date-parts":[[2023,7]]},"assertion":[{"value":"2023-08-24","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}