{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,1]],"date-time":"2026-05-01T23:00:59Z","timestamp":1777676459985,"version":"3.51.4"},"reference-count":22,"publisher":"SAGE Publications","issue":"4","license":[{"start":{"date-parts":[[2010,12,29]],"date-time":"2010-12-29T00:00:00Z","timestamp":1293580800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/journals.sagepub.com\/page\/policies\/text-and-data-mining-license"}],"content-domain":{"domain":["journals.sagepub.com"],"crossmark-restriction":true},"short-container-title":["The International Journal of High Performance Computing Applications"],"published-print":{"date-parts":[[2011,11]]},"abstract":"<jats:p>Performance analysis of applications on modern high-end petascale systems is increasingly challenging due to the rising complexity and quantity of the computing units. This paper presents a performance-analysis study using the Vampir performance-analysis tool suite, which examines application behavior as well as the fundamental system properties. This study was carried out on the Jaguar system at Oak Ridge National Laboratory, the fastest computer on the November 2009 Top500 list. We analyzed the FLASH simulation code that is designed to be scaled with tens of thousands of CPU cores, which means that using existing performance-analysis tools is very complex. The study reveals two classes of performance problems that are relevant for very high CPU counts: MPI communication and scalable I\/O. For both, solutions are presented and verified. Finally, the paper proposes improvements and extensions for event tracing tools in order to allow scalability of the tools towards higher degrees of parallelism.<\/jats:p>","DOI":"10.1177\/1094342010387806","type":"journal-article","created":{"date-parts":[[2010,12,29]],"date-time":"2010-12-29T20:25:33Z","timestamp":1293654333000},"page":"428-439","update-policy":"https:\/\/doi.org\/10.1177\/sage-journals-update-policy","source":"Crossref","is-referenced-by-count":2,"title":["Trace-based performance analysis for the petascale simulation code FLASH"],"prefix":"10.1177","volume":"25","author":[{"given":"Heike","family":"Jagode","sequence":"first","affiliation":[{"name":"The University of Tennessee, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Andreas","family":"Kn\u00fcpfer","sequence":"additional","affiliation":[{"name":"Technische Universit\u00e4t Dresden, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jack","family":"Dongarra","sequence":"additional","affiliation":[{"name":"The University of Tennessee, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Matthias","family":"Jurenz","sequence":"additional","affiliation":[{"name":"Technische Universit\u00e4t Dresden, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Matthias S","family":"M\u00fcller","sequence":"additional","affiliation":[{"name":"Technische Universit\u00e4t Dresden, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Wolfgang E","family":"Nagel","sequence":"additional","affiliation":[{"name":"Technische Universit\u00e4t Dresden, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"179","published-online":{"date-parts":[[2010,12,29]]},"reference":[{"key":"bibr1-1094342010387806","unstructured":"Brunst H (2008) Integrative concepts for scalable distributed performance analysis and visualization of parallel programs, PhD thesis, Shaker Verlag."},{"key":"bibr2-1094342010387806","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-540-74466-5_2"},{"key":"bibr3-1094342010387806","volume-title":"Proceedings of Workshop on Productivity and Performance (PROPER 2009) in conjunction with Euro-Par 2009","author":"F\u00fcrlinger K","year":"2009"},{"key":"bibr4-1094342010387806","doi-asserted-by":"publisher","DOI":"10.1145\/1378533.1378554"},{"key":"bibr5-1094342010387806","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-540-75416-9_22"},{"key":"bibr6-1094342010387806","doi-asserted-by":"crossref","unstructured":"Jagode H, Dongarra J, Alam S, Vetter J, Spear W, Malony A (2009) A Holistic Approach for Performance Measurement and Analysis for Petascale Applications, ICCS 2009, Part II, LNCS 5545. Berlin, Heidelberg: Springer-Verlag, 686\u201395.","DOI":"10.1007\/978-3-642-01973-9_77"},{"key":"bibr7-1094342010387806","doi-asserted-by":"publisher","DOI":"10.1086\/588269"},{"key":"bibr8-1094342010387806","unstructured":"Kn\u00fcpfer A (2008) Advanced memory data structures for scalable event trace analysis, PhD thesis, Technische Universit\u00e4t Dresden."},{"key":"bibr9-1094342010387806","doi-asserted-by":"publisher","DOI":"10.1016\/j.future.2004.11.021"},{"key":"bibr10-1094342010387806","doi-asserted-by":"publisher","DOI":"10.1007\/11758525_71"},{"key":"bibr11-1094342010387806","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-540-68564-7_9"},{"key":"bibr12-1094342010387806","unstructured":"Larkin J, Fahey M (2007) Guidelines for efficient parallel I\/O on the Cray XT3\/XT4. In: Proceedings of Cray User Group."},{"key":"bibr13-1094342010387806","unstructured":"MacNeice P, Olson KM, Mobarry C, deFainchtein R, Packer C (1999) PARAMESH: A parallel adaptive mesh refinement community toolkit, NASA\/CR-1999-209483."},{"key":"bibr14-1094342010387806","volume-title":"Proceedings of ParCo 2005","author":"Malony AD","year":"2006"},{"key":"bibr15-1094342010387806","first-page":"287","volume-title":"Euro-Par 2008 Workshops \u2013 Parallel Processing","author":"Mickler H","year":"2008"},{"key":"bibr16-1094342010387806","doi-asserted-by":"publisher","DOI":"10.1145\/1654059.1654115"},{"key":"bibr17-1094342010387806","doi-asserted-by":"publisher","DOI":"10.1016\/j.jpdc.2008.09.001"},{"key":"bibr18-1094342010387806","unstructured":"Roth PC, Elford C, Fin B, Huber J, Madhyastha T, Schwartz B, Shields K (1996) Etrusca: Event Trace Reduction Using Statistical Data Clustering Analysis, PhD thesis."},{"key":"bibr19-1094342010387806","doi-asserted-by":"publisher","DOI":"10.1145\/75108.75382"},{"key":"bibr20-1094342010387806","doi-asserted-by":"publisher","DOI":"10.1145\/1654059.1654097"},{"key":"bibr21-1094342010387806","unstructured":"Yang M, Koziol Q. Using collective IO inside a high performance IO software package \u2013 HDF5, www.hdfgroup.uiuc.edu\/papers\/papers\/ParallelIO\/HDF5-CollectiveChunkIO.pdf."},{"key":"bibr22-1094342010387806","doi-asserted-by":"publisher","DOI":"10.1109\/CCGRID.2007.51"}],"container-title":["The International Journal of High Performance Computing Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/1094342010387806","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/full-xml\/10.1177\/1094342010387806","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/1094342010387806","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,29]],"date-time":"2026-04-29T08:19:00Z","timestamp":1777450740000},"score":1,"resource":{"primary":{"URL":"https:\/\/journals.sagepub.com\/doi\/10.1177\/1094342010387806"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2010,12,29]]},"references-count":22,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2011,11]]}},"alternative-id":["10.1177\/1094342010387806"],"URL":"https:\/\/doi.org\/10.1177\/1094342010387806","relation":{},"ISSN":["1094-3420","1741-2846"],"issn-type":[{"value":"1094-3420","type":"print"},{"value":"1741-2846","type":"electronic"}],"subject":[],"published":{"date-parts":[[2010,12,29]]}}}