{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T04:20:56Z","timestamp":1750220456550,"version":"3.41.0"},"publisher-location":"New York, NY, USA","reference-count":70,"publisher":"ACM","license":[{"start":{"date-parts":[[2021,6,3]],"date-time":"2021-06-03T00:00:00Z","timestamp":1622678400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"University of California, Riverside Academic Senate Committee on Research (CoR)"},{"name":"National Science Foundation (NSF)","award":["1305624"],"award-info":[{"award-number":["1305624"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2021,6,3]]},"DOI":"10.1145\/3447818.3460364","type":"proceedings-article","created":{"date-parts":[[2021,6,4]],"date-time":"2021-06-04T15:09:36Z","timestamp":1622819376000},"page":"127-138","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":7,"title":["FT-BLAS"],"prefix":"10.1145","author":[{"given":"Yujia","family":"Zhai","sequence":"first","affiliation":[{"name":"University of California, Riverside"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Elisabeth","family":"Giem","sequence":"additional","affiliation":[{"name":"University of California, Riverside"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Quan","family":"Fan","sequence":"additional","affiliation":[{"name":"University of California, Riverside"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Kai","family":"Zhao","sequence":"additional","affiliation":[{"name":"University of California, Riverside"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jinyang","family":"Liu","sequence":"additional","affiliation":[{"name":"University of California, Riverside"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zizhong","family":"Chen","sequence":"additional","affiliation":[{"name":"University of California, Riverside"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2021,6,4]]},"reference":[{"volume-title":"Intel Math Kernel Library. Reference Manual","key":"e_1_3_2_1_1_1","unstructured":"2009. Intel Math Kernel Library. Reference Manual . Intel Corporation , Santa Clara, USA. ISBN 630813-054US. 2009. Intel Math Kernel Library. Reference Manual. Intel Corporation, Santa Clara, USA. ISBN 630813-054US."},{"key":"e_1_3_2_1_2_1","unstructured":"Mart\u00edn Abadi Ashish Agarwal Paul Barham Eugene Brevdo Zhifeng Chen Craig Citro Greg S. Corrado Andy Davis Jeffrey Dean Matthieu Devin Sanjay Ghemawat Ian Goodfellow Andrew Harp Geoffrey Irving Michael Isard Yangqing Jia Rafal Jozefowicz Lukasz Kaiser Manjunath Kudlur Josh Levenberg Dandelion Man\u00e9 Rajat Monga Sherry Moore Derek Murray Chris Olah Mike Schuster Jonathon Shlens Benoit Steiner Ilya Sutskever Kunal Talwar Paul Tucker Vincent Vanhoucke Vijay Vasudevan Fernanda Vi\u00e9gas Oriol Vinyals Pete Warden Martin Wattenberg Martin Wicke Yuan Yu and Xiaoqiang Zheng. 2015. TensorFlow: Large-Scale Machine Learning on Heterogeneous Systems. https:\/\/www.tensorflow.org\/ Software available from tensorflow.org.  Mart\u00edn Abadi Ashish Agarwal Paul Barham Eugene Brevdo Zhifeng Chen Craig Citro Greg S. Corrado Andy Davis Jeffrey Dean Matthieu Devin Sanjay Ghemawat Ian Goodfellow Andrew Harp Geoffrey Irving Michael Isard Yangqing Jia Rafal Jozefowicz Lukasz Kaiser Manjunath Kudlur Josh Levenberg Dandelion Man\u00e9 Rajat Monga Sherry Moore Derek Murray Chris Olah Mike Schuster Jonathon Shlens Benoit Steiner Ilya Sutskever Kunal Talwar Paul Tucker Vincent Vanhoucke Vijay Vasudevan Fernanda Vi\u00e9gas Oriol Vinyals Pete Warden Martin Wattenberg Martin Wicke Yuan Yu and Xiaoqiang Zheng. 2015. TensorFlow: Large-Scale Machine Learning on Heterogeneous Systems. https:\/\/www.tensorflow.org\/ Software available from tensorflow.org."},{"key":"e_1_3_2_1_3_1","unstructured":"AGNER. 2019. https:\/\/www.agner.org\/optimize\/instruction_tables.pdf. Online.  AGNER. 2019. https:\/\/www.agner.org\/optimize\/instruction_tables.pdf. Online."},{"key":"e_1_3_2_1_4_1","doi-asserted-by":"crossref","unstructured":"E. Anderson Z. Bai C. Bischof S. Blackford J. Demmel J. Dongarra J. Du Croz A. Greenbaum S. Hammarling A. McKenney and D. Sorensen. 1999. LAPACK Users' Guide (third ed.). Society for Industrial and Applied Mathematics Philadelphia PA.  E. Anderson Z. Bai C. Bischof S. Blackford J. Demmel J. Dongarra J. Du Croz A. Greenbaum S. Hammarling A. McKenney and D. Sorensen. 1999. LAPACK Users' Guide (third ed.). Society for Industrial and Applied Mathematics Philadelphia PA.","DOI":"10.1137\/1.9780898719604"},{"key":"e_1_3_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.5555\/156137.2813183"},{"key":"e_1_3_2_1_6_1","volume-title":"Soft errors in commercial semiconductor technology: Overview and scaling trends","author":"Baumann Robert","year":"2002","unstructured":"Robert Baumann . 2002. Soft errors in commercial semiconductor technology: Overview and scaling trends . IEEE 2002 Reliability Physics Tutorial Notes, Reliability Fundamentals 7 (2002). Robert Baumann. 2002. Soft errors in commercial semiconductor technology: Overview and scaling trends. IEEE 2002 Reliability Physics Tutorial Notes, Reliability Fundamentals 7 (2002)."},{"key":"e_1_3_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/3078597.3078617"},{"key":"e_1_3_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2016.81"},{"key":"e_1_3_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/PADSW.2014.7097827"},{"key":"e_1_3_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2008.4536158"},{"key":"e_1_3_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/2442516.2442533"},{"key":"e_1_3_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2008.58"},{"key":"e_1_3_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/HASE.2008.13"},{"key":"e_1_3_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/2903150.2903170"},{"key":"e_1_3_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/SC.2014.53"},{"key":"e_1_3_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.procs.2015.05.187"},{"key":"e_1_3_2_1_17_1","volume-title":"Intel 64 and IA-32 Architectures Optimization Reference Manual","author":"Intel Corporation","year":"2019","unstructured":"Intel Corporation . 2019. Intel 64 and IA-32 Architectures Optimization Reference Manual . Intel Corporation , Sept ( 2019 ). Intel Corporation. 2019. Intel 64 and IA-32 Architectures Optimization Reference Manual. Intel Corporation, Sept (2019)."},{"key":"e_1_3_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2016.2517639"},{"key":"e_1_3_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1177\/1094342010391989"},{"key":"e_1_3_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/MSPEC.2016.7420396"},{"key":"e_1_3_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1109\/CLUSTER.2015.108"},{"key":"e_1_3_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/1377603.1377607"},{"key":"e_1_3_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2014.128"},{"key":"e_1_3_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1109\/DSN.2001.941390"},{"key":"e_1_3_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2014.2320502"},{"key":"e_1_3_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1145\/17356.17401"},{"key":"e_1_3_2_1_27_1","volume-title":"Algorithm-based fault tolerance for matrix operations","author":"Huang Kuang-Hua","year":"1984","unstructured":"Kuang-Hua Huang and Jacob A Abraham . 1984. Algorithm-based fault tolerance for matrix operations . IEEE transactions on computers 100, 6 ( 1984 ), 518--528. Kuang-Hua Huang and Jacob A Abraham. 1984. Algorithm-based fault tolerance for matrix operations. IEEE transactions on computers 100, 6 (1984), 518--528."},{"key":"e_1_3_2_1_28_1","unstructured":"Intel. 2014. https:\/\/software.intel.com\/content\/www\/us\/en\/develop\/articles\/disclosure-of-hw-prefetcher-control-on-some-intel-processors.html. Online.  Intel. 2014. https:\/\/software.intel.com\/content\/www\/us\/en\/develop\/articles\/disclosure-of-hw-prefetcher-control-on-some-intel-processors.html. Online."},{"key":"e_1_3_2_1_29_1","volume-title":"Dependable computing and fault-tolerance. Digest of Papers FTCS-15","author":"Laprie Jean-Claude","year":"1985","unstructured":"Jean-Claude Laprie . 1985. Dependable computing and fault-tolerance. Digest of Papers FTCS-15 ( 1985 ), 2--11. Jean-Claude Laprie. 1985. Dependable computing and fault-tolerance. Digest of Papers FTCS-15 (1985), 2--11."},{"key":"e_1_3_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.5555\/2388996.2389074"},{"key":"e_1_3_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1145\/3126908.3126964"},{"key":"e_1_3_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1145\/3295500.3356195"},{"key":"e_1_3_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1145\/3126908.3126915"},{"key":"e_1_3_2_1_34_1","doi-asserted-by":"crossref","unstructured":"Robert Lucas James Ang Keren Bergman Shekhar Borkar William Carlson Laura Carrington George Chiu Robert Colwell William Dally Jack Dongarra etal 2014. DOE advanced scientific computing advisory subcommittee (ASCAC) report: top ten exascale research challenges. Technical Report. USDOE Office of Science (SC)(United States).  Robert Lucas James Ang Keren Bergman Shekhar Borkar William Carlson Laura Carrington George Chiu Robert Colwell William Dally Jack Dongarra et al. 2014. DOE advanced scientific computing advisory subcommittee (ASCAC) report: top ten exascale research challenges . Technical Report. USDOE Office of Science (SC)(United States).","DOI":"10.2172\/1222713"},{"key":"e_1_3_2_1_35_1","volume-title":"Vijay Janapa Reddi, and Kim Hazelwood","author":"Luk Chi-Keung","year":"2005","unstructured":"Chi-Keung Luk , Robert Cohn , Robert Muth , Harish Patil , Artur Klauser , Geoff Lowney , Steven Wallace , Vijay Janapa Reddi, and Kim Hazelwood . 2005 . Pin: building customized program analysis tools with dynamic instrumentation. Acm sigplan notices 40, 6 (2005), 190--200. Chi-Keung Luk, Robert Cohn, Robert Muth, Harish Patil, Artur Klauser, Geoff Lowney, Steven Wallace, Vijay Janapa Reddi, and Kim Hazelwood. 2005. Pin: building customized program analysis tools with dynamic instrumentation. Acm sigplan notices 40, 6 (2005), 190--200."},{"volume-title":"Analyzing software requirements errors in safety-critical, embedded systems. In [1993] Proceedings of the IEEE International Symposium on Requirements Engineering","author":"Lutz Robyn R","key":"e_1_3_2_1_36_1","unstructured":"Robyn R Lutz . 1993. Analyzing software requirements errors in safety-critical, embedded systems. In [1993] Proceedings of the IEEE International Symposium on Requirements Engineering . IEEE , 126--133. Robyn R Lutz. 1993. Analyzing software requirements errors in safety-critical, embedded systems. In [1993] Proceedings of the IEEE International Symposium on Requirements Engineering. IEEE, 126--133."},{"key":"e_1_3_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1109\/T-ED.1979.19370"},{"key":"e_1_3_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1109\/VLSI-TSA.2014.6839639"},{"key":"e_1_3_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1109\/VTEST.1999.766651"},{"key":"e_1_3_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1109\/24.994926"},{"key":"e_1_3_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1109\/24.994913"},{"key":"e_1_3_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1145\/3075564.3075598"},{"key":"e_1_3_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1145\/3126908.3126960"},{"key":"e_1_3_2_1_44_1","volume-title":"Retrieved","author":"BLAS.","year":"2021","unstructured":"Open BLAS. Retrieved in 2021 . https:\/\/github.com\/xianyi\/OpenBLAS\/blob\/develop\/common.h\\#L530. Online . OpenBLAS. Retrieved in 2021. https:\/\/github.com\/xianyi\/OpenBLAS\/blob\/develop\/common.h\\#L530. Online."},{"volume-title":"PyTorch: An Imperative Style","author":"Paszke Adam","key":"e_1_3_2_1_45_1","unstructured":"Adam Paszke , Sam Gross , Francisco Massa , Adam Lerer , James Bradbury , Gregory Chanan , Trevor Killeen , Zeming Lin , Natalia Gimelshein , Luca Antiga , Alban Desmaison , Andreas Kopf , Edward Yang , Zachary DeVito , Martin Raison , Alykhan Tejani , Sasank Chilamkurthy , Benoit Steiner , Lu Fang , Junjie Bai , and Soumith Chintala . 2019. PyTorch: An Imperative Style , High-Performance Deep Learning Library . In Advances in Neural Information Processing Systems 32, H. Wallach, H. Larochelle, A. Beygelzimer, F. dAlch\u00e9-Buc, E. Fox, and R. Garnett (Eds.). Curran Associates, Inc., 8024--8035. http:\/\/papers.neurips.cc\/paper\/9015-pytorch-an-imperative-style-high-performance-deep-learning-library.pdf Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. 2019. PyTorch: An Imperative Style, High-Performance Deep Learning Library. In Advances in Neural Information Processing Systems 32, H. Wallach, H. Larochelle, A. Beygelzimer, F. dAlch\u00e9-Buc, E. Fox, and R. Garnett (Eds.). Curran Associates, Inc., 8024--8035. http:\/\/papers.neurips.cc\/paper\/9015-pytorch-an-imperative-style-high-performance-deep-learning-library.pdf"},{"key":"e_1_3_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.1002\/jcc.20289"},{"key":"e_1_3_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.1145\/339647.339652"},{"key":"e_1_3_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1109\/CGO.2005.34"},{"key":"e_1_3_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.1145\/2530268.2530272"},{"key":"e_1_3_2_1_50_1","unstructured":"William C Skamarock Joseph B Klemp Jimy Dudhia David O Gill Dale M Barker Wei Wang and Jordan G Powers. 2008. A description of the Advanced Research WRF version 3. NCAR Technical note-475+ STR. (2008).  William C Skamarock Joseph B Klemp Jimy Dudhia David O Gill Dale M Barker Wei Wang and Jordan G Powers. 2008. A description of the Advanced Research WRF version 3. NCAR Technical note-475+ STR. (2008)."},{"key":"e_1_3_2_1_51_1","doi-asserted-by":"publisher","DOI":"10.1109\/DSN.2012.6263938"},{"key":"e_1_3_2_1_52_1","unstructured":"Tyler M Smith Robert A van de Geijn Mikhail Smelyanskiy and Enrique S Quintana-Ort\u0131. [n.d.]. Toward ABFT for BLIS GEMM.  Tyler M Smith Robert A van de Geijn Mikhail Smelyanskiy and Enrique S Quintana-Ort\u0131. [n.d.]. Toward ABFT for BLIS GEMM."},{"key":"e_1_3_2_1_53_1","doi-asserted-by":"publisher","DOI":"10.1177\/1094342014522573"},{"key":"e_1_3_2_1_54_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.parco.2014.09.001"},{"key":"e_1_3_2_1_55_1","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2015.108"},{"key":"e_1_3_2_1_56_1","volume-title":"CutQC: Using Small Quantum Computers for Large Quantum Circuit Evaluations. arXiv preprint arXiv:2012.02333","author":"Tang Wei","year":"2020","unstructured":"Wei Tang , Teague Tomesh , Jeffrey Larson , Martin Suchara , and Margaret Martonosi . 2020. CutQC: Using Small Quantum Computers for Large Quantum Circuit Evaluations. arXiv preprint arXiv:2012.02333 ( 2020 ). Wei Tang, Teague Tomesh, Jeffrey Larson, Martin Suchara, and Margaret Martonosi. 2020. CutQC: Using Small Quantum Computers for Large Quantum Circuit Evaluations. arXiv preprint arXiv:2012.02333 (2020)."},{"key":"e_1_3_2_1_57_1","doi-asserted-by":"publisher","DOI":"10.1145\/3208040.3208050"},{"key":"e_1_3_2_1_58_1","doi-asserted-by":"publisher","DOI":"10.1109\/71.207595"},{"key":"e_1_3_2_1_59_1","doi-asserted-by":"publisher","DOI":"10.1145\/2907294.2907306"},{"key":"e_1_3_2_1_60_1","doi-asserted-by":"publisher","DOI":"10.1145\/2764454"},{"key":"e_1_3_2_1_61_1","doi-asserted-by":"publisher","DOI":"10.1145\/2503210.2503219"},{"key":"e_1_3_2_1_62_1","volume-title":"SC'98: Proceedings of the 1998 ACM\/IEEE conference on Supercomputing. IEEE, 38--38","author":"Clinton Whaley R","year":"1998","unstructured":"R Clinton Whaley and Jack J Dongarra . 1998 . Automatically tuned linear algebra software . In SC'98: Proceedings of the 1998 ACM\/IEEE conference on Supercomputing. IEEE, 38--38 . R Clinton Whaley and Jack J Dongarra. 1998. Automatically tuned linear algebra software. In SC'98: Proceedings of the 1998 ACM\/IEEE conference on Supercomputing. IEEE, 38--38."},{"key":"e_1_3_2_1_63_1","volume-title":"Automated empirical optimizations of software and the ATLAS project. Parallel computing 27, 1-2","author":"Whaley R Clint","year":"2001","unstructured":"R Clint Whaley , Antoine Petitet , and Jack J Dongarra . 2001. Automated empirical optimizations of software and the ATLAS project. Parallel computing 27, 1-2 ( 2001 ), 3--35. R Clint Whaley, Antoine Petitet, and Jack J Dongarra. 2001. Automated empirical optimizations of software and the ATLAS project. Parallel computing 27, 1-2 (2001), 3--35."},{"key":"e_1_3_2_1_64_1","doi-asserted-by":"publisher","DOI":"10.1145\/2600212.2600232"},{"key":"e_1_3_2_1_65_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.jocs.2013.05.002"},{"key":"e_1_3_2_1_66_1","doi-asserted-by":"publisher","DOI":"10.1145\/2907294.2907315"},{"key":"e_1_3_2_1_67_1","doi-asserted-by":"publisher","DOI":"10.1145\/2907294.2907321"},{"key":"e_1_3_2_1_68_1","doi-asserted-by":"publisher","DOI":"10.1109\/CGO.2009.14"},{"key":"e_1_3_2_1_69_1","doi-asserted-by":"publisher","DOI":"10.1145\/3330345.3330373"},{"key":"e_1_3_2_1_70_1","volume-title":"Algorithm-based fault tolerance for convolutional neural networks","author":"Zhao Kai","year":"2020","unstructured":"Kai Zhao , Sheng Di , Sihuan Li , Xin Liang , Yujia Zhai , Jieyang Chen , Kaiming Ouyang , Franck Cappello , and Zizhong Chen . 2020. Algorithm-based fault tolerance for convolutional neural networks . IEEE Transactions on Parallel and Distributed Systems ( 2020 ). Kai Zhao, Sheng Di, Sihuan Li, Xin Liang, Yujia Zhai, Jieyang Chen, Kaiming Ouyang, Franck Cappello, and Zizhong Chen. 2020. Algorithm-based fault tolerance for convolutional neural networks. IEEE Transactions on Parallel and Distributed Systems (2020)."}],"event":{"name":"ICS '21: 2021 International Conference on Supercomputing","sponsor":["SIGARCH ACM Special Interest Group on Computer Architecture"],"location":"Virtual Event USA","acronym":"ICS '21"},"container-title":["Proceedings of the ACM International Conference on Supercomputing"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3447818.3460364","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3447818.3460364","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T20:48:06Z","timestamp":1750193286000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3447818.3460364"}},"subtitle":["a high performance BLAS implementation with online fault tolerance"],"short-title":[],"issued":{"date-parts":[[2021,6,3]]},"references-count":70,"alternative-id":["10.1145\/3447818.3460364","10.1145\/3447818"],"URL":"https:\/\/doi.org\/10.1145\/3447818.3460364","relation":{},"subject":[],"published":{"date-parts":[[2021,6,3]]},"assertion":[{"value":"2021-06-04","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}