{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,25]],"date-time":"2026-04-25T14:48:17Z","timestamp":1777128497297,"version":"3.51.4"},"publisher-location":"New York, NY, USA","reference-count":43,"publisher":"ACM","content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2007,6,17]]},"DOI":"10.1145\/1274971.1274978","type":"proceedings-article","created":{"date-parts":[[2010,4,7]],"date-time":"2010-04-07T02:57:04Z","timestamp":1270609024000},"page":"23-32","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":239,"title":["Proactive fault tolerance for HPC with Xen virtualization"],"prefix":"10.1145","author":[{"given":"Arun Babu","family":"Nagarajan","sequence":"first","affiliation":[{"name":"North Carolina State University, Raleigh, NC"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Frank","family":"Mueller","sequence":"additional","affiliation":[{"name":"North Carolina State University, Raleigh, NC"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Christian","family":"Engelmann","sequence":"additional","affiliation":[{"name":"Oak Ridge National Laboratory, Oak Ridge, TN"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Stephen L.","family":"Scott","sequence":"additional","affiliation":[{"name":"Oak Ridge National Laboratory, Oak Ridge, TN"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2007,6,17]]},"reference":[{"key":"e_1_3_2_1_1_1","unstructured":"Ganglia. http:\/\/ganglia.sourceforge.net\/.  Ganglia. http:\/\/ganglia.sourceforge.net\/."},{"key":"e_1_3_2_1_2_1","unstructured":"OpenIPMI. http:\/\/openipmi.sourceforge.net\/.  OpenIPMI. http:\/\/openipmi.sourceforge.net\/."},{"key":"e_1_3_2_1_3_1","volume-title":"http:\/\/www.acpi.info\/","author":"Advanced","year":"2004","unstructured":"Advanced configuration & power interface. http:\/\/www.acpi.info\/ , 2004 . Advanced configuration & power interface. http:\/\/www.acpi.info\/, 2004."},{"key":"e_1_3_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2004.1302920"},{"key":"e_1_3_2_1_5_1","first-page":"101","volume-title":"Proceedings of the","author":"Barak A.","year":"1989","unstructured":"A. Barak and R. Wheeler . MOSIX: An integrated multiprocessor UNIX. In USENIX Association, editor , Proceedings of the Winter 1989 USENIX Conference : January 30--February 3, 1989, San Diego, California, USA, pages 101 -- 112 , Berkeley, CA, USA, Winter 1989. USENIX. A. Barak and R. Wheeler. MOSIX: An integrated multiprocessor UNIX. In USENIX Association, editor, Proceedings of the Winter 1989 USENIX Conference: January 30--February 3, 1989, San Diego, California, USA, pages 101--112, Berkeley, CA, USA, Winter 1989. USENIX."},{"key":"e_1_3_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/945445.945462"},{"key":"e_1_3_2_1_7_1","volume-title":"Supercomputing","author":"Bosilca G.","year":"2002","unstructured":"G. Bosilca , A. Boutellier , and F. Cappello . MPICH-V: Toward a scalable fault tolerant MPI for volatile nodes . In Supercomputing , Nov. 2002 . G. Bosilca, A. Boutellier, and F. Cappello. MPICH-V: Toward a scalable fault tolerant MPI for volatile nodes. In Supercomputing, Nov. 2002."},{"key":"e_1_3_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.5555\/648137.746492"},{"key":"e_1_3_2_1_9_1","volume-title":"HPCRI: 1st Workshop on High Performance Computing Reliability Issues, in Proceedings of the 11th International Symposium on High Performance Computer Architecture (HPCA-11)","author":"Chakravorty S.","year":"2005","unstructured":"S. Chakravorty , C. Mendes , and L. Kale . Proactive fault tolerance in large systems . In HPCRI: 1st Workshop on High Performance Computing Reliability Issues, in Proceedings of the 11th International Symposium on High Performance Computer Architecture (HPCA-11) . IEEE Computer Society , 2005 . S. Chakravorty, C. Mendes, and L. Kale. Proactive fault tolerance in large systems. In HPCRI: 1st Workshop on High Performance Computing Reliability Issues, in Proceedings of the 11th International Symposium on High Performance Computer Architecture (HPCA-11). IEEE Computer Society, 2005."},{"key":"e_1_3_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1007\/11945918_47"},{"key":"e_1_3_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2007.370310"},{"key":"e_1_3_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.5555\/1251203.1251223"},{"key":"e_1_3_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1002\/spe.4380210802"},{"key":"e_1_3_2_1_14_1","volume-title":"Lawrence Berkeley National Laboratory","author":"Duell J.","year":"2000","unstructured":"J. Duell . The design and implementation of berkeley lab's linux checkpoint\/restart. Tr , Lawrence Berkeley National Laboratory , 2000 . J. Duell. The design and implementation of berkeley lab's linux checkpoint\/restart. Tr, Lawrence Berkeley National Laboratory, 2000."},{"key":"e_1_3_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/12.142678"},{"key":"e_1_3_2_1_16_1","first-page":"346","volume-title":"FT-MPI: Fault Tolerant MPI, supporting dynamic applications in a dynamic world","author":"Fagg G. E.","year":"2000","unstructured":"G. E. Fagg and J. J. Dongarra . FT-MPI: Fault Tolerant MPI, supporting dynamic applications in a dynamic world . In Euro PVM\/MPI User's Group Meeting , Lecture Notes in Computer Science, volume 1908 , pages 346 -- 353 , 2000 . G. E. Fagg and J. J. Dongarra. FT-MPI: Fault Tolerant MPI, supporting dynamic applications in a dynamic world. In Euro PVM\/MPI User's Group Meeting, Lecture Notes in Computer Science, volume 1908, pages 346--353, 2000."},{"key":"e_1_3_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/945445.945450"},{"key":"e_1_3_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/1133572.1133616"},{"key":"e_1_3_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/268998.266660"},{"key":"e_1_3_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/SC.2005.3"},{"key":"e_1_3_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/1183401.1183421"},{"key":"e_1_3_2_1_22_1","volume-title":"Ruud Haring","author":"Watson IBM T.J.","year":"2005","unstructured":"IBM T.J. Watson . Personal communications . Ruud Haring , July 2005 . IBM T.J. Watson. Personal communications. Ruud Haring, July 2005."},{"key":"e_1_3_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/35037.42182"},{"key":"e_1_3_2_1_24_1","first-page":"40","volume-title":"IEEE Workshop on Mobile Computing Systems and Applications","author":"Kozuch M.","unstructured":"M. Kozuch and M. Satyanarayanan . Internet suspend\/resume . In IEEE Workshop on Mobile Computing Systems and Applications , pages 40 -, 2002. M. Kozuch and M. Satyanarayanan. Internet suspend\/resume. In IEEE Workshop on Mobile Computing Systems and Applications, pages 40-, 2002."},{"key":"e_1_3_2_1_25_1","volume-title":"USENIX Conference","author":"Liu J.","year":"2006","unstructured":"J. Liu , W. Huang , B. Abali , and D. Panda . High performance vmm-bypass i\/o in virtual machines . In USENIX Conference , June 2006 . J. Liu, W. Huang, B. Abali, and D. Panda. High performance vmm-bypass i\/o in virtual machines. In USENIX Conference, June 2006."},{"key":"e_1_3_2_1_26_1","volume-title":"USENIX Conference","author":"Menon A.","year":"2006","unstructured":"A. Menon , A. Cox , and W. Zwaenepoel . Optimizing network virtualization in xen . In USENIX Conference , June 2006 . A. Menon, A. Cox, and W. Zwaenepoel. Optimizing network virtualization in xen. In USENIX Conference, June 2006."},{"key":"e_1_3_2_1_27_1","volume-title":"International Parallel and Distributed Processing Symposium","author":"Oliner A.","year":"2004","unstructured":"A. Oliner , R. Sahoo , J. Moreira , M. Gupta , and A. Sivasubramaniam . Fault-aware job scheduling for bluegene\/l systems . In International Parallel and Distributed Processing Symposium , 2004 . A. Oliner, R. Sahoo, J. Moreira, M. Gupta, and A. Sivasubramaniam. Fault-aware job scheduling for bluegene\/l systems. In International Parallel and Distributed Processing Symposium, 2004."},{"key":"e_1_3_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1145\/1183401.1183406"},{"key":"e_1_3_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.5555\/1060289.1060323"},{"key":"e_1_3_2_1_30_1","volume-title":"HPCRI: 1st Workshop on High Performance Computing Reliability Issues, in Proceedings of the 11th International Symposium on High Performance Computer Architecture (HPCA-11)","author":"Philp I.","year":"2005","unstructured":"I. Philp . Software failures and the road to a petaflop machine . In HPCRI: 1st Workshop on High Performance Computing Reliability Issues, in Proceedings of the 11th International Symposium on High Performance Computer Architecture (HPCA-11) . IEEE Computer Society , 2005 . I. Philp. Software failures and the road to a petaflop machine. In HPCRI: 1st Workshop on High Performance Computing Reliability Issues, in Proceedings of the 11th International Symposium on High Performance Computer Architecture (HPCA-11). IEEE Computer Society, 2005."},{"key":"e_1_3_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1145\/800217.806619"},{"key":"e_1_3_2_1_32_1","first-page":"2006","volume-title":"High Availability and Performance Computing Workshop","author":"Rani S.","unstructured":"S. Rani , C. Leangsuksun , A. Tikotekar , V. Rampure , and S. Scott . Toward efficient failre detection and recovery in hpc . In High Availability and Performance Computing Workshop , page (accepted), 2006 . S. Rani, C. Leangsuksun, A. Tikotekar, V. Rampure, and S. Scott. Toward efficient failre detection and recovery in hpc. In High Availability and Performance Computing Workshop, page (accepted), 2006."},{"key":"e_1_3_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1145\/956750.956799"},{"key":"e_1_3_2_1_34_1","volume-title":"Proceedings, LACSI Symposium","author":"Sankaran S.","year":"2003","unstructured":"S. Sankaran , J. M. Squyres , B. Barrett , A. Lumsdaine , J. Duell , P. Hargrove , and E. Roman . The LAM\/MPI checkpoint\/restart framework: System-initiated checkpointing . In Proceedings, LACSI Symposium , Oct. 2003 . S. Sankaran, J. M. Squyres, B. Barrett, A. Lumsdaine, J. Duell, P. Hargrove, and E. Roman. The LAM\/MPI checkpoint\/restart framework: System-initiated checkpointing. In Proceedings, LACSI Symposium, Oct. 2003."},{"key":"e_1_3_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.5555\/1060289.1060324"},{"key":"e_1_3_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1109\/DSN.2006.5"},{"key":"e_1_3_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1109\/ARES.2006.37"},{"key":"e_1_3_2_1_38_1","volume-title":"Proceedings of IPPS '96. The 10th International Parallel Processing Symposium: Honolulu, HI, USA, 15--19 April 1996, pages 526--531, 1109 Spring Street, Suite 300, Silver Spring, MD 20910, USA","author":"Stellner G.","year":"1996","unstructured":"G. Stellner . CoCheck : checkpointing and process migration for MPI. In IEEE, editor , Proceedings of IPPS '96. The 10th International Parallel Processing Symposium: Honolulu, HI, USA, 15--19 April 1996, pages 526--531, 1109 Spring Street, Suite 300, Silver Spring, MD 20910, USA , 1996 . IEEE Computer Society Press. G. Stellner. CoCheck: checkpointing and process migration for MPI. In IEEE, editor, Proceedings of IPPS '96. The 10th International Parallel Processing Symposium: Honolulu, HI, USA, 15--19 April 1996, pages 526--531, 1109 Spring Street, Suite 300, Silver Spring, MD 20910, USA, 1996. IEEE Computer Society Press."},{"key":"e_1_3_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1145\/323647.323629"},{"key":"e_1_3_2_1_40_1","first-page":"2007","volume-title":"International Parallel and Distributed Processing Symposium","author":"Wang C.","unstructured":"C. Wang , F. Mueller , C. Engelmann , and S. Scott . A job pause service under lam\/mpi+blcr for transparent fault tolerance . In International Parallel and Distributed Processing Symposium , page (accepted), Apr. 2007 . C. Wang, F. Mueller, C. Engelmann, and S. Scott. A job pause service under lam\/mpi+blcr for transparent fault tolerance. In International Parallel and Distributed Processing Symposium, page (accepted), Apr. 2007."},{"key":"e_1_3_2_1_41_1","first-page":"169","volume-title":"Symposium on Networked Systems Design and Implementation","author":"Whitaker A.","year":"2004","unstructured":"A. Whitaker , R. S. Cox , M. Shaw , and S. D. Gribble . Constructing services with interposable virtual hardware . In Symposium on Networked Systems Design and Implementation , pages 169 -- 182 , 2004 . A. Whitaker, R. S. Cox, M. Shaw, and S. D. Gribble. Constructing services with interposable virtual hardware. In Symposium on Networked Systems Design and Implementation, pages 169--182, 2004."},{"key":"e_1_3_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1145\/331532.331573"},{"key":"e_1_3_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1145\/41457.37503"}],"event":{"name":"ICS07: International Conference on Supercomputing","location":"Seattle Washington","acronym":"ICS07","sponsor":["SIGARCH ACM Special Interest Group on Computer Architecture"]},"container-title":["Proceedings of the 21st annual international conference on Supercomputing"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/1274971.1274978","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,1,11]],"date-time":"2023-01-11T20:08:03Z","timestamp":1673467683000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/1274971.1274978"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2007,6,17]]},"references-count":43,"alternative-id":["10.1145\/1274971.1274978","10.1145\/1274971"],"URL":"https:\/\/doi.org\/10.1145\/1274971.1274978","relation":{},"subject":[],"published":{"date-parts":[[2007,6,17]]},"assertion":[{"value":"2007-06-17","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}