{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,23]],"date-time":"2026-04-23T05:56:15Z","timestamp":1776923775261,"version":"3.51.2"},"reference-count":37,"publisher":"Institute of Electronics, Information and Communications Engineers (IEICE)","issue":"12","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["IEICE Trans. Inf. &amp; Syst."],"published-print":{"date-parts":[[2017]]},"DOI":"10.1587\/transinf.2017pap0002","type":"journal-article","created":{"date-parts":[[2017,11,30]],"date-time":"2017-11-30T22:26:35Z","timestamp":1512080795000},"page":"2749-2760","source":"Crossref","is-referenced-by-count":4,"title":["Energy-Performance Modeling of Speculative Checkpointing for Exascale Systems"],"prefix":"10.1587","volume":"E100.D","author":[{"given":"Muhammad","family":"ALFIAN AMRIZAL","sequence":"first","affiliation":[{"name":"Cyberscience Center, Tohoku University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Atsuya","family":"UNO","sequence":"additional","affiliation":[{"name":"Next-Generation Supercomputer R&D Center, RIKEN"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yukinori","family":"SATO","sequence":"additional","affiliation":[{"name":"Global Science and Information Center, Tokyo Institute of Technology"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hiroyuki","family":"TAKIZAWA","sequence":"additional","affiliation":[{"name":"Cyberscience Center, Tohoku University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hiroaki","family":"KOBAYASHI","sequence":"additional","affiliation":[{"name":"Graduate School of Information Sciences, Tohoku University"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"532","reference":[{"key":"1","unstructured":"[1] S. Amarasinghe, D. Campbell, W. Carlson, A. Chien, W. Dally, E. Elnohazy, M. Hall, R. Harrison, W. Harrod, and K. Hill, \u201cExaScale Software Study: Software Challenges in Extreme Scale Systems,\u201d DARPA IPTO, Air Force Research Labs, Tech. Rep, pp.1-153, 2009."},{"key":"2","doi-asserted-by":"publisher","unstructured":"[2] B. Schroeder and G.A. Gibson, \u201cUnderstanding Failures in Petascale Computers,\u201d Journal of Physics: Conference Series, vol.78, no.1, pp.12-22, 2007. 10.1088\/1742-6596\/78\/1\/012022","DOI":"10.1088\/1742-6596\/78\/1\/012022"},{"key":"3","doi-asserted-by":"publisher","unstructured":"[3] E.N.M. Elnozahy, L. Alvisi, Y.M. Wang, and D.B. Johnson, \u201cA Survey of Rollbackrecovery Protocols in Message-passing Systems,\u201d ACM Comput. Surv., vol.34, no.3, pp.375-408, 2002. 10.1145\/568522.568525","DOI":"10.1145\/568522.568525"},{"key":"4","unstructured":"[4] J. Hursey, \u201cCoordinated Checkpoint\/Restart Process Fault Tolerance for MPI Applications on HPC Systems,\u201d PhD thesis, Indiana University, 2010."},{"key":"5","doi-asserted-by":"crossref","unstructured":"[5] J.C. Sancho, F. Pertini, G. Johnson, J. Fernandez, and E. Frachtenberg, \u201cOn the Feasibility of Incremental Checkpointing for Scientific Computing,\u201d Proceedings of IPDPS 2014, pp.58-67, 2004. 10.1109\/ipdps.2004.1302982","DOI":"10.1109\/IPDPS.2004.1302982"},{"key":"6","unstructured":"[6] I. Yamagata, S. Matsuoka, H. Jitsumoto, and H. Nakada, \u201cSpeculative Checkpointing: Exploiting Temporal Affinity of Memory Operation,\u201d HPC Asia 2009, pp.390-396, 2009."},{"key":"7","doi-asserted-by":"crossref","unstructured":"[7] A. Moody, G. Bronevetsky, K. Mohror, and B.R. de Supinski, \u201cDesign, Modeling, and Evaluation of a Scalable Multi-level Checkpointing System,\u201d Proceedings of SC10, pp.1-11, 2010.","DOI":"10.2172\/984082"},{"key":"8","unstructured":"[8] TOP500.org, \u201cTop 500 Supercomputer Sites,\u201d available on the WWW, 2016. http:\/\/www.top500.org\/."},{"key":"9","doi-asserted-by":"crossref","unstructured":"[9] R. Hedges, B. Loewe, T. McLarty, and C. Morrone, \u201cParallel File System Testing for the Lunatic Fringe: The Care and Feeding of Restless I\/O Power Users,\u201d Proceedings of the 22nd IEEE\/13th NASA Goddard Conference on Mass Storage Systems and Technologies, pp.3-17, 2005. 10.1109\/msst.2005.22","DOI":"10.1109\/MSST.2005.22"},{"key":"10","unstructured":"[10] T.H. Cormen and D. Kotz, \u201cIntegrating Theory and Practice in Parallel File Systems,\u201d Proceedings of the 1993 DAGS\/PC Symposium, vol.7, pp.64-74, 1993."},{"key":"11","unstructured":"[11] H. Shan and J. Shalf, \u201cUsing IOR to Analyze the I\/O Performance for HPC Platforms,\u201d Proceedings of the 49th Cray User Group (CUG) Conference 2007, pp.1-15, 2007."},{"key":"12","doi-asserted-by":"crossref","unstructured":"[12] J. Shalf, S. Dosanjh, and J. Morrison, \u201cExascale Computing Technology Challenges,\u201d VECPAR 2010, LNCS, vol.6449, pp.1-25, 2011. 10.1007\/978-3-642-19328-6_1","DOI":"10.1007\/978-3-642-19328-6_1"},{"key":"13","doi-asserted-by":"crossref","unstructured":"[13] S. Agarwal, R. Garg, M.S. Gupta, and J.E. Moreira, \u201cAdaptive Incremental Checkpointing for Massively Parallel Systems,\u201d Proceedings of the 18th Annual International Conference on Supercomputing (ICS), pp.277-286, 2004. 10.1145\/1006209.1006248","DOI":"10.1145\/1006209.1006248"},{"key":"14","doi-asserted-by":"crossref","unstructured":"[14] B. Nicolae and F. Capello, \u201cAI-Ckpt: Leveraging Memory Access Patterns for Adaptive Asynchronous Incremental Checkpointing,\u201d Proceedings of the 22nd International Symposium onHigh-performance Parallel and Distributed Computing, pp.155-166, 2013. 10.1145\/2493123.2462918","DOI":"10.1145\/2493123.2462918"},{"key":"15","doi-asserted-by":"crossref","unstructured":"[15] Y. Matsubara and Y. Sato, \u201cOnline Memory Access Pattern Analysis on an Application Profiling Tool,\u201d Proceedings of the 2nd International Symposium on Computing and Networking, pp.602-604, 2014. 10.1109\/candar.2014.86","DOI":"10.1109\/CANDAR.2014.86"},{"key":"16","doi-asserted-by":"crossref","unstructured":"[16] Y. Oyama, S. Ishiguro, J. Murakami, S. Sasaki, R. Matsumiya, and O. Tatebe, \u201cReduction of Operating System Jitter Caused by Page Reclaim,\u201d Proceedings of the 4th International Workshop on Runtime and Operating Systems for Supercomputers, Article no.9, 2014.","DOI":"10.1145\/2612262.2612270"},{"key":"17","doi-asserted-by":"publisher","unstructured":"[17] P. Beckman, K. Iskra, K. Yoshii, and S. Coghlan, \u201cOperating System Issues for Petascale Systems,\u201d ACM SIGOPS Operating Systems Review, vol.40, no.2, p.29, 2006. 10.1145\/1131322.1131332","DOI":"10.1145\/1131322.1131332"},{"key":"18","doi-asserted-by":"crossref","unstructured":"[18] F. Petrini, D.J. Kerbyson, and S. Pakin, \u201cThe Case of the Missing Supercomputer Performance: Achieving Optimal Performance on the 8192 Processors of ASCI Q,\u201d ACM Supercomputing, pp.1-17, 2003.","DOI":"10.1145\/1048935.1050204"},{"key":"19","doi-asserted-by":"crossref","unstructured":"[19] S. Di, M.S. Bouguerra, L. Bautista-Gomez, and F. Cappello, \u201cOptimization of multi-level checkpoint model for large scale HPC applications,\u201d Proceedings of IPDPS 2014, pp.1181-1190, 2014. 10.1109\/ipdps.2014.122","DOI":"10.1109\/IPDPS.2014.122"},{"key":"20","doi-asserted-by":"publisher","unstructured":"[20] J.W. Young, \u201cA First Order Approximation to The Optimum Checkpoint Interval,\u201d Communications of the ACM, vol.17, no.9, pp.530-531, 1974. 10.1145\/361147.361115","DOI":"10.1145\/361147.361115"},{"key":"21","doi-asserted-by":"publisher","unstructured":"[21] J.T. Daly, \u201cA Higher Order Estimate of The Optimum Checkpoint Interval for Restart Dumps,\u201d Future Generation Computer Systems, vol.22, no.3, pp.303-312, 2006. 10.1016\/j.future.2004.11.016","DOI":"10.1016\/j.future.2004.11.016"},{"key":"22","unstructured":"[22] P. Luszczek, J. Dongarra, D. Koester, B. Rabenseifner, J. Lucas, J. Kepner, J. McCalpin, D. Bailey, and D. Takahashi, \u201cIntroduction to the HPC Challenge Benchmark Suite,\u201d available on the WWW, 2015. http:\/\/icl.cs.utk.edu\/hpcc\/pubs"},{"key":"23","doi-asserted-by":"publisher","unstructured":"[23] Z. Zheng, L. Yu, and Z. Lan, \u201cReliability-Aware Speedup Models for Parallel Applications with Coordinated Checkpointing\/Restart,\u201d IEEE Trans. Comput., vol.64, no.5, pp.1402-1415, 2015. 10.1109\/tc.2014.2317182","DOI":"10.1109\/TC.2014.2317182"},{"key":"24","doi-asserted-by":"publisher","unstructured":"[24] J. Dongarra, P. Beckman, P. Aerts, F. Cappello, T. Lippert, S. Matsuoka, P. Messina, T. Moore, R. Stevens, A. Trefethen, and M. Valero, \u201cThe International Exascale Software Project: A Call to Cooperative Action by The Global High-performance Community,\u201d International Journal of High Performance Computing Applications, vol.23, no.4, pp.309-322, 2009. 10.1177\/1094342009347714","DOI":"10.1177\/1094342009347714"},{"key":"25","doi-asserted-by":"crossref","unstructured":"[25] V. Sarkar, et al., \u201cExascale Software Study: Software Challenges in Extreme Scale Systems,\u201d 2009, White paper available at; ExascaleComputingStudyReports\/ECSS%20report%20101909.pdf","DOI":"10.1088\/1742-6596\/180\/1\/012045"},{"key":"26","unstructured":"[26] US Department of Energy, \u201cExascale Computing Initiative Update 2012,\u201d available at; http:\/\/science.energy.gov\/~\/media\/ascr\/ascac\/pdf\/meetings\/aug12\/2012-ECI-ASCAC-v4.pdf."},{"key":"27","doi-asserted-by":"crossref","unstructured":"[27] M. el Mehdi Diouri, O. Gluck, L. Lefevre, and F. Cappello, \u201cECOFIT: A framework to estimate energy consumption of fault tolerance protocols for HPC applications,\u201d 13th IEEE\/ACM International Symposium on Cluster, Cloud and Grid Computing(CCGrid13), pp.522-529, 2013. 10.1109\/ccgrid.2013.80","DOI":"10.1109\/CCGrid.2013.80"},{"key":"28","doi-asserted-by":"crossref","unstructured":"[28] E. Meneses, O. Sarood, and L.V. Kal&apos;e, \u201cEnergy profile of rollback-recovery strategies in high performance computing,\u201d Parallel Computing, pp.536-547, 2014.","DOI":"10.1016\/j.parco.2014.03.005"},{"key":"29","doi-asserted-by":"crossref","unstructured":"[29] X. Fan, W.D. Weber, and L.A. Barroso, \u201cPower Provisioning for a Warehouse-sized Computer,\u201d Proceedings of the 34th Annual International Symposium on Computer Architecture, pp.13-23, 2007.","DOI":"10.1145\/1250662.1250665"},{"key":"30","unstructured":"[30] F. Shoji, S. Matsui, M. Okamoto, F. Sueyasu, T. Tsukamoto, A. Uno, and K. Yamamoto, \u201cLong Term Failure Analysis of 10 Petascale Supercomputer,\u201d HPC in Asia Poster, ISC, 2015."},{"key":"31","unstructured":"[31] D.H. Bailey, J.T. Barton, T.A. Lasinski, and H.D. Simon, \u201cThe NAS Parallel Benchmarks,\u201d Technical Report RNR-91-002 Revision 2, NASA Ames Research Laboratory, pp.1-71, 1991."},{"key":"32","unstructured":"[32] G.H. Bryan and R. Rotunno, \u201cThe Maximum Intensity of Tropical Cyclones in Axisymmetric Numerical Model Simulations,\u201d Journal of the American Meteorological Society, vol.137, pp.1770-1789, 2009."},{"key":"33","unstructured":"[33] Nersc 6 Procurement Benchmark, available on the WWW. http:\/\/www.nersc.gov"},{"key":"34","doi-asserted-by":"publisher","unstructured":"[34] A. Duda, \u201cThe Effects of Checkpointing on Program Execution Time,\u201d Information Processing Letters, vol.16, no.5, pp.221-229, 1983. 10.1016\/0020-0190(83)90093-5","DOI":"10.1016\/0020-0190(83)90093-5"},{"key":"35","doi-asserted-by":"publisher","unstructured":"[35] J.S. Plank and M.G. Thomason, \u201cProcessor Allocation and Checkpoint Interval Selection in Cluster Computing Systems,\u201d Journal of Parallel Distributed Computing, vol.61, no.11, pp.1570-1590, 2001. 10.1006\/jpdc.2001.1757","DOI":"10.1006\/jpdc.2001.1757"},{"key":"36","doi-asserted-by":"crossref","unstructured":"[36] P. Balaprakash, L.A.B. Gomez, M.-S. Bouguerra, S.M. Wild, F. Capello, and P.D. Hovland, \u201cAnalysis of the Tradeoffs Between Energy and Run Time for Multilevel Checkpointing,\u201d Performance Modeling, Benchmarking, and Simulation of High Performance Computer Systems, pp.249-263, 2014. 10.1007\/978-3-319-17248-4_13","DOI":"10.1007\/978-3-319-17248-4_13"},{"key":"37","doi-asserted-by":"crossref","unstructured":"[37] G. Aupy, A. Benoit, T. Herault, Y. Robert, and J. Dongarra, \u201cOptimal Checkpointing Period: Time vs. Energy,\u201d CoRR, abs\/310.8456, 2013.","DOI":"10.1007\/978-3-319-10214-6_10"}],"container-title":["IEICE Transactions on Information and Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.jstage.jst.go.jp\/article\/transinf\/E100.D\/12\/E100.D_2017PAP0002\/_pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,27]],"date-time":"2025-06-27T22:17:53Z","timestamp":1751062673000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.jstage.jst.go.jp\/article\/transinf\/E100.D\/12\/E100.D_2017PAP0002\/_article"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2017]]},"references-count":37,"journal-issue":{"issue":"12","published-print":{"date-parts":[[2017]]}},"URL":"https:\/\/doi.org\/10.1587\/transinf.2017pap0002","relation":{},"ISSN":["0916-8532","1745-1361"],"issn-type":[{"value":"0916-8532","type":"print"},{"value":"1745-1361","type":"electronic"}],"subject":[],"published":{"date-parts":[[2017]]}}}