{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,3]],"date-time":"2026-04-03T15:04:39Z","timestamp":1775228679065,"version":"3.50.1"},"reference-count":47,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2019,4,23]],"date-time":"2019-04-23T00:00:00Z","timestamp":1555977600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"National Technology and Engineering Solutions of Sandia LLC, a wholly owned subsidiary of Honeywell International Inc."},{"DOI":"10.13039\/100006234","name":"Sandia National Laboratories","doi-asserted-by":"crossref","id":[{"id":"10.13039\/100006234","id-type":"DOI","asserted-by":"crossref"}]},{"name":"U.S. Department of Energy's National Nuclear Security Administration","award":["DE-NA0003525"],"award-info":[{"award-number":["DE-NA0003525"]}]},{"name":"U.S. Department of Energy or the United States Government"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Model. Perform. Eval. Comput. Syst."],"published-print":{"date-parts":[[2019,6,30]]},"abstract":"<jats:p>In this article, we present an approach to streaming collection of application performance data. Practical application performance tuning and troubleshooting in production high-performance computing (HPC) environments requires an understanding of how applications interact with the platform, including (but not limited to) parallel programming libraries such as Message Passing Interface (MPI). Several profiling and tracing tools exist that collect heavy runtime data traces either in memory (released only at application exit) or on a file system (imposing an I\/O load that may interfere with the performance being measured). Although these approaches are beneficial in development stages and post-run analysis, a systemwide and low-overhead method is required to monitor deployed applications continuously. This method must be able to collect information at both the application and system levels to yield a complete performance picture.<\/jats:p><jats:p>In our approach, an application profiler collects application event counters. A sampler uses an efficient inter-process communication method to periodically extract the application counters and stream them into an infrastructure for performance data collection. We implement a tool-set based on our approach and integrate it with the Lightweight Distributed Metric Service (LDMS) system, a monitoring system used on large-scale computational platforms. LDMS provides the infrastructure to create and gather streams of performance data in a low overhead manner. We demonstrate our approach using applications implemented with MPI, as it is one of the most common standards for the development of large-scale scientific applications.<\/jats:p><jats:p>We utilize our tool-set to study the impact of our approach on an open source HPC application, Nalu. Our tool-set enables us to efficiently identify patterns in the behavior of the application without source-level knowledge. We leverage LDMS to collect system-level performance data and explore the correlation between the system and application events. Also, we demonstrate how our tool-set can help detect anomalies with a low latency. We run tests on two different architectures: a system enabled with Intel Xeon Phi and another system equipped with Intel Xeon processor. Our overhead study shows our method imposes at most 0.5% CPU usage overhead on the application in realistic deployment scenarios.<\/jats:p>","DOI":"10.1145\/3319498","type":"journal-article","created":{"date-parts":[[2019,4,23]],"date-time":"2019-04-23T15:24:23Z","timestamp":1556033063000},"page":"1-25","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":3,"title":["Production Application Performance Data Streaming for System Monitoring"],"prefix":"10.1145","volume":"4","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-7253-084X","authenticated-orcid":false,"given":"Ramin","family":"Izadpanah","sequence":"first","affiliation":[{"name":"University of Central Florida, Orlando, FL"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Benjamin A.","family":"Allan","sequence":"additional","affiliation":[{"name":"Sandia National Laboratories, Albuquerque, NM"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Damian","family":"Dechev","sequence":"additional","affiliation":[{"name":"University of Central Florida, Orlando, FL"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jim","family":"Brandt","sequence":"additional","affiliation":[{"name":"Sandia National Laboratories, Albuquerque, NM"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2019,4,23]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.5555\/1753228.1753233"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1109\/SC.2014.18"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.parco.2016.05.009"},{"key":"e_1_2_1_4_1","volume-title":"Cray User Group CUG","author":"Agelastos A. M.","year":"2017"},{"key":"e_1_2_1_5_1","volume-title":"Lin","author":"Agelastos Anthony Michael","year":"2013"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10586-007-0047-2"},{"key":"e_1_2_1_7_1","volume-title":"Periscope: An online-based distributed performance analysis tool. In Tools for High Performance Computing","author":"Benedict Shajulin","year":"2010"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/2503210.2503247"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPSW.2016.188"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1177\/109434200001400303"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/3126908.3126926"},{"key":"e_1_2_1_12_1","volume-title":"Anthony Michael Agelastos, and others. Design and Implementation of a Scalable Monitoring System for Trinity. Sandia National Lab.(SNL-NM)","author":"DeConinck Adam","year":"2016"},{"key":"e_1_2_1_13_1","volume-title":"David Morton, Amanda Bonnie, Cory Lueninghoener, James M. Brandt, Ann C. Gentile, Kevin Pedretti, Anthony Michael Agelastos, Courtenay T. Vaughan and others.","author":"DeConinck Adam","year":"2017"},{"key":"e_1_2_1_14_1","volume-title":"Monitoring with Graphite: Tracking Dynamic Host and Application Metrics at Scale. O\u2019Reilly Media","author":"Dixon Jason"},{"key":"e_1_2_1_15_1","unstructured":"S. Domino. 2018. Nalu\u2019s milestoneRun test at master. Retrieved from https:\/\/github.com\/NaluCFD\/Nalu\/tree\/master\/reg_tests\/test_files\/milestoneRun. S. Domino. 2018. Nalu\u2019s milestoneRun test at master. Retrieved from https:\/\/github.com\/NaluCFD\/Nalu\/tree\/master\/reg_tests\/test_files\/milestoneRun."},{"key":"e_1_2_1_16_1","unstructured":"S. Domino. 2018. Sierra Low Mach Module: Nalu Theory Manual 1.0. Retrieved from https:\/\/github.com\/NaluCFD\/NaluDoc. S. Domino. 2018. Sierra Low Mach Module: Nalu Theory Manual 1.0. Retrieved from https:\/\/github.com\/NaluCFD\/NaluDoc."},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1109\/CLUSTER.2015.125"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2011.67"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.5555\/1753228.1753234"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/872726.806987"},{"key":"e_1_2_1_22_1","doi-asserted-by":"crossref","first-page":"181","DOI":"10.1080\/00031305.1998.10480559","article-title":"Violin plots: A box plot-density trace synergism","volume":"52","author":"Hintze Jerry L.","year":"1998","journal-title":"Am. Stat."},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/2807591.2807644"},{"key":"e_1_2_1_24_1","volume-title":"Tools for High Performance Computing","author":"Ilsche Thomas","year":"2014"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISORC.2016.16"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1145\/3225058.3225086"},{"key":"e_1_2_1_27_1","unstructured":"Jim Jeffers and James Reinders. 2015. High Performance Parallelism Pearls Volume Two: Multicore and Many-core Programming Approaches. Morgan Kaufmann. Jim Jeffers and James Reinders. 2015. High Performance Parallelism Pearls Volume Two: Multicore and Many-core Programming Approaches. Morgan Kaufmann."},{"key":"e_1_2_1_28_1","volume-title":"Proceedings of the CEUR Workshop","volume":"1787","author":"Kadochnikov I. S."},{"key":"e_1_2_1_29_1","volume-title":"Proceedings of the Workshop on Environments and Tools for Parallel Scientific Computing. 195--200","author":"Karrels Edward","year":"1994"},{"key":"e_1_2_1_30_1","unstructured":"Michael Kerrisk. 2010. The Linux Programming Interface. No Starch Press. Michael Kerrisk. 2010. The Linux Programming Interface. No Starch Press."},{"key":"e_1_2_1_31_1","volume-title":"Nagel","author":"Kn\u00fcpfer Andreas","year":"2008"},{"key":"e_1_2_1_32_1","volume-title":"Scott Biersdorff, Kai Diethelm, Dominic Eschweiler, Markus Geimer, Michael Gerndt, Daniel Lorenz, Allen Malony, et al.","author":"Kn\u00fcpfer Andreas","year":"2012"},{"key":"e_1_2_1_33_1","unstructured":"Los Alamos National Laboratory. 2018. Trinity. Retrieved from https:\/\/www.lanl.gov\/projects\/trinity. Los Alamos National Laboratory. 2018. Trinity. Retrieved from https:\/\/www.lanl.gov\/projects\/trinity."},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPSW.2016.199"},{"key":"e_1_2_1_35_1","unstructured":"S. S. Antman J. E. Marsden L. Sirovich S. Wiggins L. Glass R. V. Kohn and S. S. Sastry. 1993. Interdisciplinary Applied Mathematics. Vol. 3. Springer Berlin. S. S. Antman J. E. Marsden L. Sirovich S. Wiggins L. Glass R. V. Kohn and S. S. Sastry. 1993. Interdisciplinary Applied Mathematics. Vol. 3. Springer Berlin."},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.parco.2004.04.001"},{"key":"e_1_2_1_37_1","volume-title":"TAU: A portable parallel program analysis environment for C++. In Proceedings of the Parallel Processing: CONPAR 94\u2014VAPP","author":"Mohr Bernd","year":"1994"},{"key":"e_1_2_1_38_1","unstructured":"National Center for Supercomputing Applications. 2018. Blue Waters. Retrieved from https:\/\/bluewaters.ncsa.illinois.edu. National Center for Supercomputing Applications. 2018. Blue Waters. Retrieved from https:\/\/bluewaters.ncsa.illinois.edu."},{"key":"e_1_2_1_39_1","unstructured":"Prometheus authors. 2018. Prometheus\u2014Monitoring System and Time Series Database. Retrieved from https:\/\/prometheus.io. Prometheus authors. 2018. Prometheus\u2014Monitoring System and Time Series Database. Retrieved from https:\/\/prometheus.io."},{"key":"e_1_2_1_40_1","volume-title":"Strategies, Tools, and Techniques: Proceedings of the 26th Large Installation System Administration Conference (LISA\u201912). 153--162.","author":"Reams J."},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1109\/CLUSTER.2017.115"},{"key":"e_1_2_1_42_1","volume-title":"Koretsky","author":"Sarwar Syed Mansoor","year":"2016"},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1155\/2008\/713705"},{"key":"e_1_2_1_44_1","volume-title":"Proceedings of the European Conference on Parallel Processing. Springer, 185--198","author":"Servat Harald","year":"2009"},{"key":"e_1_2_1_45_1","volume-title":"Performance monitoring of parallel scientific applications. Ernest Orlando Lawrence Berkeley National Laboratory","author":"Skinner David"},{"key":"e_1_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.1145\/3126908.3126946"},{"key":"e_1_2_1_47_1","unstructured":"Jeffrey Vetter and Chris Chambreau. 2005. mpiP: Lightweight Scalable MPIProfiling. http:\/\/mpip.sourceforge.net\/ Jeffrey Vetter and Chris Chambreau. 2005. mpiP: Lightweight Scalable MPIProfiling. http:\/\/mpip.sourceforge.net\/"},{"key":"e_1_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1186\/s13677-014-0024-2"}],"container-title":["ACM Transactions on Modeling and Performance Evaluation of Computing Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3319498","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3319498","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3319498","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T22:38:21Z","timestamp":1750199901000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3319498"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019,4,23]]},"references-count":47,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2019,6,30]]}},"alternative-id":["10.1145\/3319498"],"URL":"https:\/\/doi.org\/10.1145\/3319498","relation":{},"ISSN":["2376-3639","2376-3647"],"issn-type":[{"value":"2376-3639","type":"print"},{"value":"2376-3647","type":"electronic"}],"subject":[],"published":{"date-parts":[[2019,4,23]]},"assertion":[{"value":"2018-03-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2019-02-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2019-04-23","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}