{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,25]],"date-time":"2026-06-25T02:58:33Z","timestamp":1782356313301,"version":"3.54.5"},"reference-count":17,"publisher":"MDPI AG","issue":"10","license":[{"start":{"date-parts":[[2023,10,10]],"date-time":"2023-10-10T00:00:00Z","timestamp":1696896000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Information"],"abstract":"<jats:p>Top-Down models are defined by hardware architects to provide information on the utilization of different hardware components. The target is to isolate the users from the complexity of the hardware architecture while giving them insight into how efficiently the code uses the resources. In this paper, we explore the applicability of four Top-Down models defined for different hardware architectures powering state-of-the-art HPC clusters (Intel Skylake, Fujitsu A64FX, IBM Power9, and Huawei Kunpeng 920) and propose a model for AMD Zen 2. We study a parallel CFD code used for scientific production to compare these five Top-Down models. We evaluate the level of insight achieved, the clarity of the information, the ease of use, and the conclusions each allows us to reach. Our study indicates that the Top-Down model makes it very difficult for a performance analyst to spot inefficiencies in complex scientific codes without delving deep into micro-architecture details.<\/jats:p>","DOI":"10.3390\/info14100554","type":"journal-article","created":{"date-parts":[[2023,10,10]],"date-time":"2023-10-10T10:23:42Z","timestamp":1696933422000},"page":"554","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":2,"title":["Top-Down Models across CPU Architectures: Applicability and Comparison in a High-Performance Computing Environment"],"prefix":"10.3390","volume":"14","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-9809-0857","authenticated-orcid":false,"given":"Fabio","family":"Banchelli","sequence":"first","affiliation":[{"name":"Barcelona Supercomputing Center, Pla\u00e7a Eusebi G\u00fcell, 1-3, 08034 Barcelona, Spain"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3682-9905","authenticated-orcid":false,"given":"Marta","family":"Garcia-Gasulla","sequence":"additional","affiliation":[{"name":"Barcelona Supercomputing Center, Pla\u00e7a Eusebi G\u00fcell, 1-3, 08034 Barcelona, Spain"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3559-4825","authenticated-orcid":false,"given":"Filippo","family":"Mantovani","sequence":"additional","affiliation":[{"name":"Barcelona Supercomputing Center, Pla\u00e7a Eusebi G\u00fcell, 1-3, 08034 Barcelona, Spain"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2023,10,10]]},"reference":[{"key":"ref_1","unstructured":"(2022, November 01). Top500 List. Available online: https:\/\/www.top500.org\/lists\/top500\/2022\/11\/."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"65","DOI":"10.1145\/1498765.1498785","article-title":"Roofline: An Insightful Visual Performance Model for Multicore Architectures","volume":"52","author":"Williams","year":"2009","journal-title":"Commun. ACM"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Ofenbeck, G., Steinmann, R., Caparros, V., Spampinato, D.G., and P\u00fcschel, M. (2014, January 23\u201325). Applying the roofline model. Proceedings of the 2014 IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS), Monterey, CA, USA.","DOI":"10.1109\/ISPASS.2014.6844463"},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"21","DOI":"10.1109\/L-CA.2013.6","article-title":"Cache-aware Roofline model: Upgrading the loft","volume":"13","author":"Ilic","year":"2014","journal-title":"IEEE Comput. Archit. Lett."},{"key":"ref_5","unstructured":"Banchelli, F., Garcia-Gasulla, M., Houzeaux, G., and Mantovani, F. (July, January 29). Benchmarking of State-of-the-Art HPC Clusters with a Production CFD Code. Proceedings of the Platform for Advanced Scientific Computing Conference, Geneva, Switzerland."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Yasin, A. (2014, January 23\u201325). A Top-Down method for performance analysis and counters architecture. Proceedings of the 2014 IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS), Monterey, CA, USA.","DOI":"10.1109\/ISPASS.2014.6844459"},{"key":"ref_7","unstructured":"Intel (2022). Top-Down Microarchitecture Analysis Method, Intel."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"146","DOI":"10.1016\/j.simpat.2016.08.006","article-title":"Top-Down Characterization Approximation based on performance counters architecture for AMD processors","volume":"68","author":"Jarus","year":"2016","journal-title":"Simul. Model. Pract. Theory"},{"key":"ref_9","unstructured":"Banchelli, F., Oyarzun, G., Garcia-Gasulla, M., Mantovani, F., Both, A., Houzeaux, G., and Mira, D. (2022). A portable coding strategy to exploit vectorization on combustion simulations. arXiv."},{"key":"ref_10","unstructured":"Fog, A. (2022). The Microarchitecture of Intel, AMD, and VIA CPUs\u2014An Optimization Guide for Assembly Programmers and Compiler Makers, Copenhagen University College of Engineering."},{"key":"ref_11","unstructured":"(2023, August 07). A64FX Microarchitecture Manual. Available online: https:\/\/raw.githubusercontent.com\/fujitsu\/A64FX\/master\/doc\/A64FX_Microarchitecture_Manual_en_1.6.pdf."},{"key":"ref_12","unstructured":"(2023, August 07). POWER9 Performance Monitor Unit User\u2019s Guide. Available online: https:\/\/wiki.raptorcs.com\/w\/images\/6\/6b\/POWER9_PMU_UG_v12_28NOV2018_pub.pdf."},{"key":"ref_13","unstructured":"(2023, August 07). Unified European Applications Benchmark Suite. Available online: https:\/\/prace-ri.eu\/training-support\/technical-documentation\/benchmark-suites\/."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"M\u00fcller, M.S., Resch, M.M., Schulz, A., and Nagel, W.E. (2010). Tools for High Performance Computing 2009: Proceedings of the 3rd International Workshop on Parallel Tools for High Performance Computing, September 2009, ZIH, Dresden, Springer.","DOI":"10.1007\/978-3-642-11261-4"},{"key":"ref_15","unstructured":"(2023, August 07). Extrae. Available online: https:\/\/tools.bsc.es\/extrae."},{"key":"ref_16","unstructured":"Pillet, V., Pillet, V., Labarta, J., Cortes, T., Cortes, T., Girona, S., Girona, S., and Computadors, D.D.D. (1995). Proceedings of WoTUG-18: Transputer and Occam Developments, IOS Press. Technical Report."},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Gonzalez, J., Gimenez, J., and Labarta, J. (2009, January 10\u201312). Automatic detection of parallel applications computation phases. Proceedings of the 2009 IEEE International Symposium on Parallel & Distributed Processing, Chengdu, China.","DOI":"10.1109\/IPDPS.2009.5161027"}],"container-title":["Information"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2078-2489\/14\/10\/554\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T21:04:30Z","timestamp":1760130270000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2078-2489\/14\/10\/554"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,10,10]]},"references-count":17,"journal-issue":{"issue":"10","published-online":{"date-parts":[[2023,10]]}},"alternative-id":["info14100554"],"URL":"https:\/\/doi.org\/10.3390\/info14100554","relation":{},"ISSN":["2078-2489"],"issn-type":[{"value":"2078-2489","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,10,10]]}}}