{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,6]],"date-time":"2026-06-06T07:04:42Z","timestamp":1780729482689,"version":"3.54.1"},"reference-count":38,"publisher":"SAGE Publications","issue":"5","license":[{"start":{"date-parts":[[2016,9,26]],"date-time":"2016-09-26T00:00:00Z","timestamp":1474848000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/journals.sagepub.com\/page\/policies\/text-and-data-mining-license"}],"content-domain":{"domain":["journals.sagepub.com"],"crossmark-restriction":true},"short-container-title":["The International Journal of High Performance Computing Applications"],"published-print":{"date-parts":[[2018,9]]},"abstract":"<jats:p>We present a unified fault-tolerance framework for task-parallel message-passing applications to mitigate transient errors. First, we propose a fault-tolerant message-logging protocol that only requires the restart of the task that experienced the error and transparently handles any message passing interface calls inside the task. In our experiments we demonstrate that our fault-tolerant solution has a reasonable overhead, with a maximum observed overhead of 4.5%. We also show that fine-grained parallelization is important for hiding the overheads related to the protocol as well as the recovery of tasks. Secondly, we develop a mathematical model to unify task-level checkpointing and our protocol with system-wide checkpointing in order to provide complete failure coverage. We provide closed formulas for the optimal checkpointing interval and the performance score of the unified scheme. Experimental results show that the performance improvement can be as high as 98% with the unified scheme.<\/jats:p>","DOI":"10.1177\/1094342016669416","type":"journal-article","created":{"date-parts":[[2016,9,27]],"date-time":"2016-09-27T20:11:29Z","timestamp":1475007089000},"page":"641-657","update-policy":"https:\/\/doi.org\/10.1177\/sage-journals-update-policy","source":"Crossref","is-referenced-by-count":10,"title":["Unified fault-tolerance framework for hybrid task-parallel message-passing applications"],"prefix":"10.1177","volume":"32","author":[{"given":"Omer","family":"Subasi","sequence":"first","affiliation":[{"name":"Barcelona Supercomputing Center, Spain"},{"name":"Universitat Politecnica de Catalunya, Spain"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Tatiana","family":"Martsinkevich","sequence":"additional","affiliation":[{"name":"INRIA, University of Paris Sud, France"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ferad","family":"Zyulkyarov","sequence":"additional","affiliation":[{"name":"Barcelona Supercomputing Center, Spain"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Osman","family":"Unsal","sequence":"additional","affiliation":[{"name":"Barcelona Supercomputing Center, Spain"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jesus","family":"Labarta","sequence":"additional","affiliation":[{"name":"Barcelona Supercomputing Center, Spain"},{"name":"Universitat Politecnica de Catalunya, Spain"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Franck","family":"Cappello","sequence":"additional","affiliation":[{"name":"Argonne National Laboratory, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"179","published-online":{"date-parts":[[2016,9,26]]},"reference":[{"key":"bibr1-1094342016669416","unstructured":"Amarasinghe S, Campbell D, Carlson W. (2009) Exa-Scale software study: Software challenges in extreme scale systems. Available at: http:\/\/users.ece.gatech.edu\/mrichard\/ExascaleComputingStudyReports\/ECSS%20report%20101909.pdf (accessed 1 July 2015)."},{"key":"bibr2-1094342016669416","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-38750-0_19"},{"key":"bibr3-1094342016669416","first-page":"278","author":"Bautista\u2013Gomez L","year":"2014","journal-title":"CLUSTER"},{"key":"bibr4-1094342016669416","doi-asserted-by":"publisher","DOI":"10.1145\/2063384.2063427"},{"key":"bibr5-1094342016669416","doi-asserted-by":"publisher","DOI":"10.1002\/cpe.3173"},{"key":"bibr6-1094342016669416","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-23397-5_6"},{"key":"bibr7-1094342016669416","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2009.5160999"},{"key":"bibr8-1094342016669416","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2015.81"},{"issue":"1","key":"bibr9-1094342016669416","first-page":"5","volume":"1","author":"Cappello F","year":"2014","journal-title":"Supercomputing Frontiers and Innovations"},{"key":"bibr10-1094342016669416","first-page":"1","volume-title":"Proceedings of the international conference on high performance computing, networking, storage and analysis","author":"Chung J"},{"key":"bibr11-1094342016669416","doi-asserted-by":"publisher","DOI":"10.1016\/j.future.2004.11.016"},{"key":"bibr12-1094342016669416","doi-asserted-by":"publisher","DOI":"10.1109\/SC.2014.79"},{"key":"bibr13-1094342016669416","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2014.122"},{"key":"bibr14-1094342016669416","doi-asserted-by":"publisher","DOI":"10.1177\/1094342010391989"},{"key":"bibr15-1094342016669416","doi-asserted-by":"publisher","DOI":"10.1142\/S0129626411000151"},{"key":"bibr16-1094342016669416","doi-asserted-by":"publisher","DOI":"10.1109\/FTCS.1994.315630"},{"key":"bibr17-1094342016669416","first-page":"1","author":"Fu H","year":"2010","journal-title":"2010 International Conference on Information Science and Applications ICISA"},{"key":"bibr18-1094342016669416","doi-asserted-by":"publisher","DOI":"10.1145\/2597917.2597942"},{"key":"bibr19-1094342016669416","doi-asserted-by":"publisher","DOI":"10.1109\/ICPP.2010.80"},{"key":"bibr20-1094342016669416","doi-asserted-by":"publisher","DOI":"10.1145\/359545.359563"},{"key":"bibr21-1094342016669416","doi-asserted-by":"publisher","DOI":"10.1109\/DSNW.2012.6264673"},{"key":"bibr22-1094342016669416","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2014.2342228"},{"key":"bibr23-1094342016669416","doi-asserted-by":"publisher","DOI":"10.1109\/SC.2010.18"},{"key":"bibr24-1094342016669416","doi-asserted-by":"crossref","first-page":"111","DOI":"10.1109\/ISCA.2002.1003567","author":"Prvulovic M","year":"2002","journal-title":"Proceedings of the 29th Annual International Symposium on Computer Architecture ISCA"},{"key":"bibr25-1094342016669416","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2008.4658655"},{"key":"bibr26-1094342016669416","doi-asserted-by":"publisher","DOI":"10.1109\/SC.2012.18"},{"key":"bibr27-1094342016669416","doi-asserted-by":"publisher","DOI":"10.1145\/2503210.2503271"},{"key":"bibr28-1094342016669416","doi-asserted-by":"publisher","DOI":"10.1177\/1094342014522573"},{"key":"bibr29-1094342016669416","doi-asserted-by":"crossref","first-page":"123","DOI":"10.1109\/ISCA.2002.1003568","author":"Sorin DJ","year":"2002","journal-title":"Proceedings of the 29th Annual International Symposium on Computer Architecture, ISCA"},{"key":"bibr30-1094342016669416","doi-asserted-by":"publisher","DOI":"10.1504\/IJHPSA.2007.013289"},{"key":"bibr31-1094342016669416","doi-asserted-by":"publisher","DOI":"10.1109\/PDP.2015.17"},{"key":"bibr32-1094342016669416","first-page":"470","author":"Subasi O","year":"2015","journal-title":"High Performance Computing and Communications"},{"key":"bibr33-1094342016669416","volume":"7179","author":"Tahan O","year":"2012","journal-title":"Architecture of Computing Systems"},{"key":"bibr34-1094342016669416","first-page":"256","volume-title":"Proceedings of the 2007 Conference of the Center for Advanced Studies on Collaborative Research","author":"Teruel X","year":"2007"},{"key":"bibr35-1094342016669416","first-page":"1","volume":"501002","author":"Treaster M","year":"2005","journal-title":"ACM Computing Research Repository"},{"key":"bibr36-1094342016669416","doi-asserted-by":"publisher","DOI":"10.1109\/DSN.2005.67"},{"key":"bibr37-1094342016669416","doi-asserted-by":"publisher","DOI":"10.1145\/361147.361115"},{"key":"bibr38-1094342016669416","doi-asserted-by":"publisher","DOI":"10.1109\/CLUSTR.2009.5289177"}],"container-title":["The International Journal of High Performance Computing Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/1094342016669416","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/full-xml\/10.1177\/1094342016669416","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/1094342016669416","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,29]],"date-time":"2026-04-29T08:15:35Z","timestamp":1777450535000},"score":1,"resource":{"primary":{"URL":"https:\/\/journals.sagepub.com\/doi\/10.1177\/1094342016669416"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2016,9,26]]},"references-count":38,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2018,9]]}},"alternative-id":["10.1177\/1094342016669416"],"URL":"https:\/\/doi.org\/10.1177\/1094342016669416","relation":{},"ISSN":["1094-3420","1741-2846"],"issn-type":[{"value":"1094-3420","type":"print"},{"value":"1741-2846","type":"electronic"}],"subject":[],"published":{"date-parts":[[2016,9,26]]}}}