{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T04:39:51Z","timestamp":1750307991454,"version":"3.41.0"},"reference-count":45,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2006,4,1]],"date-time":"2006-04-01T00:00:00Z","timestamp":1143849600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["SIGOPS Oper. Syst. Rev."],"published-print":{"date-parts":[[2006,4]]},"abstract":"<jats:p>MOLAR is a multi-institutional research effort that concentrates on adaptive, reliable, and efficient operating and runtime system (OS\/R) solutions for ultra-scale high-end scientific computing on the next generation of supercomputers. This research addresses the challenges outlined in FAST-OS (forum to address scalable technology for runtime and operating systems) and HECRTF (high-end computing revitalization task force) activities by exploring the use of advanced monitoring and adaptation to improve application performance and predictability of system interruptions, and by advancing computer reliability, availability and serviceability (RAS) management systems to work cooperatively with the OS\/R to identify and preemptively resolve system issues. This paper describes recent research of the MOLAR team in advancing RAS for high-end computing OS\/Rs.<\/jats:p>","DOI":"10.1145\/1131322.1131337","type":"journal-article","created":{"date-parts":[[2006,7,24]],"date-time":"2006-07-24T17:00:26Z","timestamp":1153760426000},"page":"63-72","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":5,"title":["MOLAR"],"prefix":"10.1145","volume":"40","author":[{"given":"Christian","family":"Engelmann","sequence":"first","affiliation":[{"name":"Oak Ridge National Laboratory, Oak Ridge, TN"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Stephen L.","family":"Scott","sequence":"additional","affiliation":[{"name":"Oak Ridge National Laboratory, Oak Ridge, TN"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"David E.","family":"Bernholdt","sequence":"additional","affiliation":[{"name":"Oak Ridge National Laboratory, Oak Ridge, TN"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Narasimha R.","family":"Gottumukkala","sequence":"additional","affiliation":[{"name":"Louisiana Tech University, Ruston, LA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Chokchai","family":"Leangsuksun","sequence":"additional","affiliation":[{"name":"Louisiana Tech University, Ruston, LA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jyothish","family":"Varma","sequence":"additional","affiliation":[{"name":"North Carolina State University, Raleigh, NC"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Chao","family":"Wang","sequence":"additional","affiliation":[{"name":"North Carolina State University, Raleigh, NC"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Frank","family":"Mueller","sequence":"additional","affiliation":[{"name":"North Carolina State University, Raleigh, NC"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Aniruddha G.","family":"Shet","sequence":"additional","affiliation":[{"name":"The Ohio State University, Columbus, OH"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"P.","family":"Sadayappan","sequence":"additional","affiliation":[{"name":"The Ohio State University, Columbus, OH"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2006,4]]},"reference":[{"key":"e_1_2_1_1_1","volume-title":"Recovery patterns for iterative methods in a parallel unstable environment. Submitted to SIAM Journal on Scientific Computing","author":"Bosilca G.","year":"2005","unstructured":"G. Bosilca , Z. Chen , J. J. Dongarra , and J. Langou . Recovery patterns for iterative methods in a parallel unstable environment. Submitted to SIAM Journal on Scientific Computing , 2005 .]] G. Bosilca, Z. Chen, J. J. Dongarra, and J. Langou. Recovery patterns for iterative methods in a parallel unstable environment. Submitted to SIAM Journal on Scientific Computing, 2005.]]"},{"key":"e_1_2_1_2_1","volume-title":"Proceedings of the Symposium on Principles and Practice of Parallel Programming (PPoPP)","author":"Chen Z.","year":"2005","unstructured":"Z. Chen , G. E. Fagg , E. Gabriel , J. Langou , T. Angskun , G. Bosilca , and J. J. Dongarra . Building fault survivable MPI programs with FT-MPI using diskless checkpointing . Proceedings of the Symposium on Principles and Practice of Parallel Programming (PPoPP) , 2005 .]] Z. Chen, G. E. Fagg, E. Gabriel, J. Langou, T. Angskun, G. Bosilca, and J. J. Dongarra. Building fault survivable MPI programs with FT-MPI using diskless checkpointing. Proceedings of the Symposium on Principles and Practice of Parallel Programming (PPoPP), 2005.]]"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/155332.155333"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1145\/1041680.1041682"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.5555\/792760.793177"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2005.34"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1007\/11428831_39"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1007\/11758525_77"},{"key":"e_1_2_1_9_1","volume-title":"Proceedings of High Availability and Performance Computing Workshop (HAPCW)","author":"Engelmann C.","year":"2005","unstructured":"C. Engelmann and S. L. Scott . Concepts for high availability in scientific high-end computing . Proceedings of High Availability and Performance Computing Workshop (HAPCW) , 2005 .]] C. Engelmann and S. L. Scott. Concepts for high availability in scientific high-end computing. Proceedings of High Availability and Performance Computing Workshop (HAPCW), 2005.]]"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.5555\/645458.655464"},{"key":"e_1_2_1_11_1","volume-title":"Proceedings of High Availability and Performance Computing Workshop (HAPCW)","author":"Engelmann C.","year":"2004","unstructured":"C. Engelmann , S. L. Scott , and G. A. Geist . High availability through distributed control . Proceedings of High Availability and Performance Computing Workshop (HAPCW) , 2004 .]] C. Engelmann, S. L. Scott, and G. A. Geist. High availability through distributed control. Proceedings of High Availability and Performance Computing Workshop (HAPCW), 2004.]]"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/ARES.2006.23"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1016\/S0167-8191(01)00100-4"},{"key":"e_1_2_1_14_1","unstructured":"Forum to Address Scalable Technology for Runtime and Operating Systems. FAST-OS at http:\/\/www.fastos.org.]]  Forum to Address Scalable Technology for Runtime and Operating Systems. FAST-OS at http:\/\/www.fastos.org.]]"},{"key":"e_1_2_1_15_1","unstructured":"Forum to Address Scalable Technology for Runtime and Operating Systems (FAST-OS). MOLAR project at http:\/\/www.fastos.org\/molar.]]  Forum to Address Scalable Technology for Runtime and Operating Systems (FAST-OS). MOLAR project at http:\/\/www.fastos.org\/molar.]]"},{"key":"e_1_2_1_16_1","unstructured":"Fault Tolerant MPI (FT-MPI) Project at University of Tennessee Knoxville TN USA. At http:\/\/icl.cs.utk.edu\/ftmpi.]]  Fault Tolerant MPI (FT-MPI) Project at University of Tennessee Knoxville TN USA. At http:\/\/icl.cs.utk.edu\/ftmpi.]]"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-540-30218-6_19"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.5555\/207505"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1142\/S0129626499000244"},{"key":"e_1_2_1_20_1","volume-title":"Proceedings of 2nd Workshop on High Performance Computing Reliability Issues (HPCRI)","author":"Gottumukkala N. R.","year":"2006","unstructured":"N. R. Gottumukkala , C. Leangsuksun , and S. L. Scott . Reliability-aware approach to improve job completion time for large-scale parallel applications . Proceedings of 2nd Workshop on High Performance Computing Reliability Issues (HPCRI) , 2006 .]] N. R. Gottumukkala, C. Leangsuksun, and S. L. Scott. Reliability-aware approach to improve job completion time for large-scale parallel applications. Proceedings of 2nd Workshop on High Performance Computing Reliability Issues (HPCRI), 2006.]]"},{"key":"e_1_2_1_21_1","unstructured":"HA-OSCAR at Louisiana Tech University Ruston LA USA. http:\/\/xcr.cenit.latech.edu\/ha-oscar.]]  HA-OSCAR at Louisiana Tech University Ruston LA USA. http:\/\/xcr.cenit.latech.edu\/ha-oscar.]]"},{"key":"e_1_2_1_22_1","article-title":"HA-OSCAR: Towards highly available linux clusters","author":"Haddad I.","year":"2004","unstructured":"I. Haddad , C. Leangsuksun , and S. L. Scott . HA-OSCAR: Towards highly available linux clusters . Linux World Magazine , March 2004 .]] I. Haddad, C. Leangsuksun, and S. L. Scott. HA-OSCAR: Towards highly available linux clusters. Linux World Magazine, March 2004.]]","journal-title":"Linux World Magazine"},{"key":"e_1_2_1_23_1","volume-title":"Proceedings of High Availability and Performance Computing Workshop (HAPCW)","author":"He X.","year":"2004","unstructured":"X. He , L. Ou , S. L. Scott , and C. Engelmann . A highly available cluster storage system using scavenging . Proceedings of High Availability and Performance Computing Workshop (HAPCW) , 2004 .]] X. He, L. Ou, S. L. Scott, and C. Engelmann. A highly available cluster storage system using scavenging. Proceedings of High Availability and Performance Computing Workshop (HAPCW), 2004.]]"},{"key":"e_1_2_1_24_1","unstructured":"High-End Computing Revitalization Task Force. HECRTF at http:\/\/www.nitrd.gov\/subcommittee\/hec\/hecrtf-outreach.]]  High-End Computing Revitalization Task Force. HECRTF at http:\/\/www.nitrd.gov\/subcommittee\/hec\/hecrtf-outreach.]]"},{"key":"e_1_2_1_25_1","unstructured":"InfiniBand. http:\/\/www.infinibandta.org\/home.]]  InfiniBand. http:\/\/www.infinibandta.org\/home.]]"},{"key":"e_1_2_1_26_1","unstructured":"Lawrence Berkeley National Laboratory Berkeley CA USA. Berkeley Lab Checkpoint Restart (BLCR) Project at http:\/\/ftg.lbl.gov\/checkpoint.]]  Lawrence Berkeley National Laboratory Berkeley CA USA. Berkeley Lab Checkpoint Restart (BLCR) Project at http:\/\/ftg.lbl.gov\/checkpoint.]]"},{"key":"e_1_2_1_27_1","unstructured":"Lawrence Livermore National Laboratory Livermore CA USA. Trace logs at http:\/\/www.llnl.gov\/asci\/platforms\/white.]]  Lawrence Livermore National Laboratory Livermore CA USA. Trace logs at http:\/\/www.llnl.gov\/asci\/platforms\/white.]]"},{"key":"e_1_2_1_28_1","volume-title":"Proceedings of 2nd International Workshop on Operating Systems, Programming Environments and Management Tools for High-Performance Computing on Clusters (COSET-2)","author":"Leangsuksun C.","year":"2005","unstructured":"C. Leangsuksun , V. K. Munganuru , T. Liu , S. L. Scott , and C. Engelmann . Asymmetric active-active high availability for high-end computing . Proceedings of 2nd International Workshop on Operating Systems, Programming Environments and Management Tools for High-Performance Computing on Clusters (COSET-2) , 2005 .]] C. Leangsuksun, V. K. Munganuru, T. Liu, S. L. Scott, and C. Engelmann. Asymmetric active-active high availability for high-end computing. Proceedings of 2nd International Workshop on Operating Systems, Programming Environments and Management Tools for High-Performance Computing on Clusters (COSET-2), 2005.]]"},{"key":"e_1_2_1_29_1","volume-title":"September","author":"Moore S.","year":"2001","unstructured":"S. Moore , D. Cronk , K. London , and J. Dongarra . Review of Performance Analysis Tools for MPI Parallel Programs. Lecture Notes in Computer Science: 8th European PVM\/MPI Users, 2331:241--248 , September 2001 .]] S. Moore, D. Cronk, K. London, and J. Dongarra. Review of Performance Analysis Tools for MPI Parallel Programs. Lecture Notes in Computer Science: 8th European PVM\/MPI Users, 2331:241--248, September 2001.]]"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDCS.1994.302392"},{"key":"e_1_2_1_31_1","unstructured":"MPICH-V Project at University of Paris - South France. http:\/\/www.Iri.fr\/~gk\/mpich-v.]]  MPICH-V Project at University of Paris - South France. http:\/\/www.Iri.fr\/~gk\/mpich-v.]]"},{"key":"e_1_2_1_32_1","unstructured":"MPICH2. http:\/\/www-unix.mcs.anl.gov\/mpi\/mpich2.]]  MPICH2. http:\/\/www-unix.mcs.anl.gov\/mpi\/mpich2.]]"},{"key":"e_1_2_1_33_1","unstructured":"MVAPICH2 MPI over InfiniBand Project. http:\/\/nowlab.cse.ohio-state.edu\/projects\/mpi-iba.]]  MVAPICH2 MPI over InfiniBand Project. http:\/\/nowlab.cse.ohio-state.edu\/projects\/mpi-iba.]]"},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1109\/CLUSTR.2003.1253309"},{"key":"e_1_2_1_35_1","unstructured":"Oak Ridge National Laboratory TN USA. Harness project at http:\/\/www.csm.ornl.gov\/harness.]]  Oak Ridge National Laboratory TN USA. Harness project at http:\/\/www.csm.ornl.gov\/harness.]]"},{"key":"e_1_2_1_36_1","unstructured":"Open MPI Project. http:\/\/www.open-mpi.org.]]  Open MPI Project. http:\/\/www.open-mpi.org.]]"},{"key":"e_1_2_1_37_1","unstructured":"OpenPBS resource manager at Altair Engineering Troy MI USA. http:\/\/www.openpbs.org.]]  OpenPBS resource manager at Altair Engineering Troy MI USA. http:\/\/www.openpbs.org.]]"},{"key":"e_1_2_1_38_1","unstructured":"PVFS at Clemson University Clemson SC USA. http:\/\/www.parl.clemson.edu\/pvfs.]]  PVFS at Clemson University Clemson SC USA. http:\/\/www.parl.clemson.edu\/pvfs.]]"},{"key":"e_1_2_1_39_1","unstructured":"PVM Project at Oak Ridge National Laboratory. Oak Ridge TN USA. http:\/\/www.csm.ornl.gov\/pvm.]]  PVM Project at Oak Ridge National Laboratory. Oak Ridge TN USA. http:\/\/www.csm.ornl.gov\/pvm.]]"},{"key":"e_1_2_1_40_1","volume-title":"A modern taxonomy of high availability","author":"Resnick R. I.","year":"1996","unstructured":"R. I. Resnick . A modern taxonomy of high availability , 1996 . http:\/\/www.generalconcepts.com\/resources\/reliability\/resnick\/HA.htm.]] R. I. Resnick. A modern taxonomy of high availability, 1996. http:\/\/www.generalconcepts.com\/resources\/reliability\/resnick\/HA.htm.]]"},{"key":"e_1_2_1_41_1","unstructured":"Science Case for Large-scale Simulation. SCaLeS at http:\/\/www.pnl.gov\/scales.]]  Science Case for Large-scale Simulation. SCaLeS at http:\/\/www.pnl.gov\/scales.]]"},{"key":"e_1_2_1_43_1","unstructured":"SLURM resource manager at Lawrence Livermore National Laboratory Livermore CA USA. http:\/\/www.llnl.gov\/linux\/slurm.]]  SLURM resource manager at Lawrence Livermore National Laboratory Livermore CA USA. http:\/\/www.llnl.gov\/linux\/slurm.]]"},{"key":"e_1_2_1_44_1","unstructured":"TORQUE resource manager at Cluster Resources Inc. Spanish Fork UT USA. http:\/\/www.clusterresources.com.]]  TORQUE resource manager at Cluster Resources Inc. Spanish Fork UT USA. http:\/\/www.clusterresources.com.]]"},{"key":"e_1_2_1_45_1","volume-title":"UK","author":"Uhlemann K.","year":"2006","unstructured":"K. Uhlemann . High availability for ultra-scale high-end scientific computing. Master Thesis at the Department of Computer Science of the University of Reading , UK , March 2006 .]] K. Uhlemann. High availability for ultra-scale high-end scientific computing. Master Thesis at the Department of Computer Science of the University of Reading, UK, March 2006.]]"},{"key":"e_1_2_1_46_1","volume-title":"An Analysis of Popular MPI Implementations. Third MPI Developers' and Users' Conference","author":"White J. B.","year":"1999","unstructured":"J. B. White and S. W. Bova . Where's the Overlap ? An Analysis of Popular MPI Implementations. Third MPI Developers' and Users' Conference , March 1999 .]] J. B. White and S. W. Bova. Where's the Overlap? An Analysis of Popular MPI Implementations. Third MPI Developers' and Users' Conference, March 1999.]]"}],"container-title":["ACM SIGOPS Operating Systems Review"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/1131322.1131337","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/1131322.1131337","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T15:06:16Z","timestamp":1750259176000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/1131322.1131337"}},"subtitle":["adaptive runtime support for high-end computing operating and runtime systems"],"short-title":[],"issued":{"date-parts":[[2006,4]]},"references-count":45,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2006,4]]}},"alternative-id":["10.1145\/1131322.1131337"],"URL":"https:\/\/doi.org\/10.1145\/1131322.1131337","relation":{},"ISSN":["0163-5980"],"issn-type":[{"type":"print","value":"0163-5980"}],"subject":[],"published":{"date-parts":[[2006,4]]},"assertion":[{"value":"2006-04-01","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}