{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,25]],"date-time":"2026-02-25T13:39:54Z","timestamp":1772026794572,"version":"3.50.1"},"reference-count":37,"publisher":"Springer Science and Business Media LLC","issue":"5","license":[{"start":{"date-parts":[[2018,4,25]],"date-time":"2018-04-25T00:00:00Z","timestamp":1524614400000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2018,4,25]],"date-time":"2018-04-25T00:00:00Z","timestamp":1524614400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Neural Comput &amp; Applic"],"published-print":{"date-parts":[[2019,5,3]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Performance diagnosing for HPC applications can be extremely difficult due to their complicated performance behaviors. One hand, developers used to identify the potential performance bottlenecks by conducting detailed instrumentation, which may introduce significant performance overheads or even performance deviations. On the other hand, developers can only conduct small numbers of application runs for profiling the performance with the limitations on both computing resources and time duration. Meanwhile, the performance bottlenecks of HPC applications may vary with the degree of parallelism. To address these challenges, our paper proposes a systematic performance diagnosing method focusing on building an accurate and interpretable performance model with performance counters. Our method is able to diagnose the HPC application scaling issues by predicting its runtime and performance behaviors in different functions. After applying this modeling method on three real-world HPC applications, HOMME, CICE and OpenFoam, our evaluations show that our diagnosing method based on the performance model has the ability to diagnose the potential scaling issues, which is typically missed by the traditional performance diagnosing method and achieves about 10% prediction errors in a scale of 4096 MPI ranks on two problem sizes.<\/jats:p>","DOI":"10.1007\/s00521-018-3496-z","type":"journal-article","created":{"date-parts":[[2018,4,25]],"date-time":"2018-04-25T07:46:04Z","timestamp":1524642364000},"page":"1563-1575","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":10,"title":["Using hardware counter-based performance model to diagnose scaling issues of HPC applications"],"prefix":"10.1007","volume":"31","author":[{"given":"Nan","family":"Ding","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Shiming","family":"Xu","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zhenya","family":"Song","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Baoquan","family":"Zhang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jingmei","family":"Li","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zhigao","family":"Zheng","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2018,4,25]]},"reference":[{"key":"3496_CR1","doi-asserted-by":"crossref","unstructured":"Calotoiu A, Hoefler T, Poke M, Wolf F (2013) Using automated performance modeling to find scalability bugs in complex codes. In: Proceedings of the international conference on high performance computing, networking, storage and analysis. ACM, p 45","DOI":"10.1145\/2503210.2503277"},{"key":"3496_CR2","doi-asserted-by":"crossref","unstructured":"Bhattacharyya A, Kwasniewski G, Hoefler T (2015) Using compiler techniques to improve automatic performance modeling. In: Proceedings of the 24th international conference on parallel architectures and compilation. ACM","DOI":"10.1109\/PACT.2015.39"},{"issue":"5","key":"3496_CR3","first-page":"1740008","volume":"25","author":"H Wang","year":"2017","unstructured":"Wang H, Jingchao LI, Guo L, Dou Z, Lin Y, Zhou R (2017) Fractal complexity-based feature extraction algorithm of communication signals. Fract Complex Geom Patterns Scaling Nat Soc 25(5):1740008","journal-title":"Fract Complex Geom Patterns Scaling Nat Soc"},{"issue":"10","key":"3496_CR4","doi-asserted-by":"publisher","first-page":"1675","DOI":"10.3390\/s16101675","volume":"16","author":"Y Lin","year":"2016","unstructured":"Lin Y, Wang C, Wang J, Dou Z (2016) A novel dynamic spectrum access framework based on reinforcement learning for cognitive radio sensor networks. Sensors 16(10):1675","journal-title":"Sensors"},{"key":"3496_CR5","doi-asserted-by":"publisher","first-page":"2874","DOI":"10.1007\/s11227-016-1681-3","volume":"72","author":"Y Lin","year":"2016","unstructured":"Lin Y, Wang C, Ma C, Dou Z, Ma X (2016) A new combination method for multisensor conflict information. J Supercomput 72:2874\u20132890","journal-title":"J Supercomput"},{"key":"3496_CR6","doi-asserted-by":"crossref","unstructured":"Kn\u00fcpfer A, R\u00f6ssel C, an\u00a0Mey D, Biersdorff S, Diethelm K, Eschweiler D, Geimer M, Gerndt M, Lorenz D, Malony A et\u00a0al. (2012) Score-P: a joint performance measurement run-time infrastructure for periscope, Scalasca, TAU, and Vampir. In: Tools for high performance computing 2011. Springer, pp 79\u201391","DOI":"10.1007\/978-3-642-31476-6_7"},{"issue":"6","key":"3496_CR7","doi-asserted-by":"publisher","first-page":"702","DOI":"10.1002\/cpe.1556","volume":"22","author":"M Geimer","year":"2010","unstructured":"Geimer M, Wolf F, Wylie BJ, \u00c1brah\u00e1m E, Becker D, Mohr B (2010) The Scalasca performance toolset architecture. Concurr Comput Pract Exp 22(6):702","journal-title":"Concurr Comput Pract Exp"},{"issue":"4","key":"3496_CR8","doi-asserted-by":"publisher","first-page":"65","DOI":"10.1145\/1498765.1498785","volume":"52","author":"S Williams","year":"2009","unstructured":"Williams S, Waterman A, Patterson D (2009) Roofline: an insightful visual performance model for multicore architectures. Commun ACM 52(4):65","journal-title":"Commun ACM"},{"key":"3496_CR9","doi-asserted-by":"crossref","unstructured":"Stengel H, Treibig J, Hager G, Wellein G (2015) Quantifying performance bottlenecks of stencil computations using the execution-cache-memory model. In: Proceedings of the 29th ACM on international conference on supercomputing. ACM, pp 207\u2013216","DOI":"10.1145\/2751205.2751240"},{"issue":"3","key":"3496_CR10","doi-asserted-by":"publisher","first-page":"261","DOI":"10.1177\/1094342005056114","volume":"19","author":"DJ Kerbyson","year":"2005","unstructured":"Kerbyson DJ, Jones PW (2005) A performance model of the parallel ocean program. Int J High Perform Comput Appl 19(3):261","journal-title":"Int J High Perform Comput Appl"},{"key":"3496_CR11","doi-asserted-by":"crossref","unstructured":"Bauer G, Gottlieb S, Hoefler T (2012) Performance modeling and comparative analysis of the MILC lattice QCD application su3_rmd. In: 12th IEEE\/ACM international symposium on cluster, cloud and grid computing (CCGrid), 2012. IEEE, pp 652\u2013659","DOI":"10.1109\/CCGrid.2012.123"},{"key":"3496_CR12","doi-asserted-by":"crossref","unstructured":"Mondragon OH, Bridges PG, Levy S, Ferreira KB, Widener P (2016) Understanding performance interference in next-generation HPC systems. In: High performance computing, networking, storage and analysis, SC16: international conference for IEEE, pp 384\u2013395","DOI":"10.1109\/SC.2016.32"},{"key":"3496_CR13","doi-asserted-by":"crossref","unstructured":"Jayakumar A, Murali P, Vadhiyar S (2015) Matching application signatures for performance predictions using a single execution. In: Parallel and distributed processing symposium (IPDPS), 2015 IEEE International. IEEE, pp 1161\u20131170","DOI":"10.1109\/IPDPS.2015.20"},{"key":"3496_CR14","doi-asserted-by":"crossref","unstructured":"Alexandrov A, Ionescu MF, Schauser KE, Scheiman C (1995) LogGP: incorporating long messages into the LogP model-one step closer towards a realistic model for parallel computation. In: Proceedings of the seventh annual ACM symposium on parallel algorithms and architectures. ACM, pp 95\u2013105","DOI":"10.1145\/215399.215427"},{"issue":"1","key":"3496_CR15","doi-asserted-by":"publisher","first-page":"74","DOI":"10.1177\/1094342011428142","volume":"26","author":"JM Dennis","year":"2012","unstructured":"Dennis JM, Edwards J, Evans KJ, Guba O, Lauritzen PH, Mirin AA, St-Cyr A, Taylor MA, Worley PH (2012) CAM-SE: a scalable spectral element dynamical core for the Community Atmosphere Model. Int J High Perform Comput Appl 26(1):74","journal-title":"Int J High Perform Comput Appl"},{"key":"3496_CR16","unstructured":"Hunke EC, Lipscomb WH, Turner AK et\u00a0al. (2010) CICE: the Los Alamos sea ice model documentation and software users manual version 4.1 LA-CC-06-012. T-3 Fluid Dynamics Group, Los Alamos National Laboratory 675"},{"key":"3496_CR17","unstructured":"Jasak H, Jemcov A, Tukovic Z, et\u00a0al. (2007) OpenFOAM: A C++ library for complex physics simulations. In: International workshop on coupled methods in numerical dynamics, IUC Dubrovnik, Croatia, vol 1000, pp 1\u201320"},{"key":"3496_CR18","unstructured":"Uh GR, Cohn R, Yadavalli B, Peri R, Ayyagari R (2006) Analyzing dynamic binary instrumentation overhead. In: WBIA workshop at ASPLOS"},{"key":"3496_CR19","unstructured":"Weaver VM (2016) Advanced hardware profiling and sampling (PEBS, IBS, etc.): creating a new PAPI sampling interface"},{"issue":"4","key":"3496_CR20","doi-asserted-by":"publisher","first-page":"421","DOI":"10.1007\/s00165-010-0163-2","volume":"23","author":"A Stewart","year":"2011","unstructured":"Stewart A (2011) A programming model for BSP with partitioned synchronisation. Form Asp Comput 23(4):421","journal-title":"Form Asp Comput"},{"key":"3496_CR21","doi-asserted-by":"crossref","unstructured":"Clapp R, Dimitrov M, Kumar K, Viswanathan V, Willhalm T (2015) Quantifying the performance impact of memory latency and bandwidth for big data workloads. In: IEEE international symposium on workload characterization (IISWC), 2015. IEEE, pp 213\u2013224","DOI":"10.1109\/IISWC.2015.32"},{"key":"3496_CR22","doi-asserted-by":"crossref","unstructured":"Chatzopoulos G, Dragojevi\u0107 A, Guerraoui R (2016) Estima: extrapolating scalability of in-memory applications. In: Proceedings of the 21st ACM SIGPLAN symposium on principles and practice of parallel programming. ACM, p 27","DOI":"10.1145\/2851141.2851159"},{"issue":"2","key":"3496_CR23","doi-asserted-by":"publisher","first-page":"154","DOI":"10.1177\/1094342014548771","volume":"29","author":"AP Craig","year":"2015","unstructured":"Craig AP, Mickelson SA, Hunke EC et al (2015) Improved parallel performance of the CICE model in CESM1. J High Perform Comput Appl 29(2):154\u2013165","journal-title":"J High Perform Comput Appl"},{"key":"3496_CR24","doi-asserted-by":"crossref","unstructured":"Prakash S, Bagrodia RL (1998) MPI-SIM: using parallel simulation to evaluate MPI programs. In: Proceedings of the 30th conference on Winter simulation. IEEE Computer Society Press, pp 467\u2013474","DOI":"10.1109\/WSC.1998.745023"},{"key":"3496_CR25","unstructured":"Zheng G, Kakulapati G, V. Kal\u00e9 L (2004) Bigsim: a parallel simulator for performance prediction of extremely large parallel machines. In: Parallel and distributed processing symposium, 2004. Proceedings. 18th International. IEEE, p 78"},{"key":"3496_CR26","doi-asserted-by":"crossref","unstructured":"Zhai J, Chen W, Zheng W (2010) Phantom: predicting performance of parallel applications on large-scale parallel machines using a single node. In: ACM sigplan notices, vol 45. ACM, pp 305\u2013314","DOI":"10.1145\/1837853.1693493"},{"key":"3496_CR27","doi-asserted-by":"crossref","unstructured":"Wu X, Mueller F (2011) Scalaextrap: trace-based communication extrapolation for spmd programs. In: ACM SIGPLAN notices, vol 46. ACM, pp 113\u2013122","DOI":"10.1145\/2038037.1941569"},{"key":"3496_CR28","doi-asserted-by":"publisher","first-page":"59","DOI":"10.1016\/j.future.2013.04.014","volume":"30","author":"C Engelmann","year":"2014","unstructured":"Engelmann C (2014) Scaling to a million cores and beyond: using light-weight simulation to understand the challenges ahead on the road to exascale. Future Gener Comput Syst 30:59","journal-title":"Future Gener Comput Syst"},{"key":"3496_CR29","doi-asserted-by":"crossref","unstructured":"Yin H, Hu Z, Zhou X, Wang H, Zheng K, Nguyen QVH, Sadiq S (2016) Discovering interpretable geo-social communities for user behavior prediction. In: IEEE international conference on data engineering, pp 942\u2013953","DOI":"10.1109\/ICDE.2016.7498303"},{"issue":"2","key":"3496_CR30","first-page":"11","volume":"35","author":"H Yin","year":"2016","unstructured":"Yin H, Cui B, Zhou X, Wang W, Huang Z, Sadiq S (2016) Joint modeling of user check-in behaviors for real-time point-of-interest recommendation. ACM Trans Inf Syst (TOIS) 35(2):11","journal-title":"ACM Trans Inf Syst (TOIS)"},{"key":"3496_CR31","doi-asserted-by":"crossref","unstructured":"Hoefler T, Gropp W, Kramer W, Snir M (2011) Performance modeling for systematic performance tuning. In: State of the practice reports, ACM, p 6","DOI":"10.1145\/2063348.2063356"},{"key":"3496_CR32","doi-asserted-by":"crossref","unstructured":"Bhattacharyya A, Hoefler T (2014) Pemogen: Automatic adaptive performance modeling during program runtime. In: Proceedings of the 23rd international conference on parallel architectures and compilation. ACM, pp 393\u2013404","DOI":"10.1145\/2628071.2628100"},{"key":"3496_CR33","unstructured":"Wang S, Wang S, Wang S, Wang S, Wang S, Wang S (2016) Learning graph-based POI embedding for location-based recommendation. In: ACM international on conference on information and knowledge management, pp 15\u201324"},{"issue":"99","key":"3496_CR34","first-page":"1","volume":"PP","author":"H Yin","year":"2017","unstructured":"Yin H, Wang W, Wang H, Chen L, Zhou X (2017) Spatial-aware hierarchical collaborative deep learning for POI recommendation. IEEE Trans Knowl Data Eng PP(99):1","journal-title":"IEEE Trans Knowl Data Eng"},{"issue":"3","key":"3496_CR35","doi-asserted-by":"publisher","first-page":"11","DOI":"10.1145\/2629461","volume":"32","author":"H Yin","year":"2014","unstructured":"Yin H, Cui B, Sun Y, Hu Z, Chen L (2014) LCARS:aA spatial item recommender system. Acm Trans Inf Syst 32(3):11","journal-title":"Acm Trans Inf Syst"},{"key":"3496_CR36","doi-asserted-by":"crossref","unstructured":"Martinasso M, Kwasniewski G, Alam SR, Schulthess TC, Hoefler T (2016) A PCIe congestion-aware performance model for densely populated accelerator servers. In: Proceedings of the international conference for high performance computing, networking, storage and analysis. IEEE Press, p 63","DOI":"10.1109\/SC.2016.62"},{"key":"3496_CR37","doi-asserted-by":"crossref","unstructured":"Yang LT, Ma X, Mueller F (2005) Cross-platform performance prediction of parallel applications using partial execution. In: Supercomputing, 2005. Proceedings of the ACM\/IEEE SC 2005 conference. IEEE, pp 40\u201340","DOI":"10.1109\/SC.2005.20"}],"updated-by":[{"DOI":"10.1007\/s00521-024-10074-9","type":"retraction","label":"Retraction","source":"publisher","updated":{"date-parts":[[2024,7,11]],"date-time":"2024-07-11T00:00:00Z","timestamp":1720656000000}}],"container-title":["Neural Computing and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s00521-018-3496-z.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s00521-018-3496-z\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s00521-018-3496-z.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,7,11]],"date-time":"2024-07-11T09:03:42Z","timestamp":1720688622000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s00521-018-3496-z"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2018,4,25]]},"references-count":37,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2019,5,3]]}},"alternative-id":["3496"],"URL":"https:\/\/doi.org\/10.1007\/s00521-018-3496-z","relation":{},"ISSN":["0941-0643","1433-3058"],"issn-type":[{"value":"0941-0643","type":"print"},{"value":"1433-3058","type":"electronic"}],"subject":[],"published":{"date-parts":[[2018,4,25]]},"assertion":[{"value":"5 December 2017","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"18 April 2018","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"25 April 2018","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"11 July 2024","order":4,"name":"change_date","label":"Change Date","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"Correction","order":5,"name":"change_type","label":"Change Type","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"This article has been retracted. Please see the Retraction Notice for more detail:","order":6,"name":"change_details","label":"Change Details","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"https:\/\/doi.org\/10.1007\/s00521-024-10074-9","URL":"https:\/\/doi.org\/10.1007\/s00521-024-10074-9","order":7,"name":"change_details","label":"Change Details","group":{"name":"ArticleHistory","label":"Article History"}}]}}