{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,11]],"date-time":"2026-06-11T16:20:28Z","timestamp":1781194828366,"version":"3.54.1"},"publisher-location":"Cham","reference-count":25,"publisher":"Springer International Publishing","isbn-type":[{"value":"9783030507428","type":"print"},{"value":"9783030507435","type":"electronic"}],"license":[{"start":{"date-parts":[[2020,1,1]],"date-time":"2020-01-01T00:00:00Z","timestamp":1577836800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2020,6,15]],"date-time":"2020-06-15T00:00:00Z","timestamp":1592179200000},"content-version":"vor","delay-in-days":166,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2020]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Hardware platforms in high performance computing are constantly getting more complex to handle even when considering multicore CPUs alone. Numerous features and configuration options in the hardware and the software environment that are relevant for performance are not even known to most application users or developers. Microbenchmarks, i.e., simple codes that fathom a particular aspect of the hardware, can help to shed light on such issues, but only if they are well understood and if the results can be reconciled with known facts or performance models. The insight gained from microbenchmarks may then be applied to real applications for performance analysis or optimization. In this paper we investigate two modern Intel x86 server CPU architectures in depth: Broadwell EP and Cascade Lake SP. We highlight relevant hardware configuration settings that can have a decisive impact on code performance and show how to properly measure on-chip and off-chip data transfer bandwidths. The new victim L3 cache of Cascade Lake and its advanced replacement policy receive due attention. Finally we use DGEMM, sparse matrix-vector multiplication, and the HPCG benchmark to make a connection to relevant application scenarios.<\/jats:p>","DOI":"10.1007\/978-3-030-50743-5_21","type":"book-chapter","created":{"date-parts":[[2020,6,15]],"date-time":"2020-06-15T19:03:45Z","timestamp":1592247825000},"page":"412-433","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":14,"title":["Understanding HPC Benchmark Performance on Intel Broadwell and\u00a0Cascade Lake Processors"],"prefix":"10.1007","author":[{"given":"Christie L.","family":"Alappat","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Johannes","family":"Hofmann","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Georg","family":"Hager","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Holger","family":"Fehske","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Alan R.","family":"Bishop","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Gerhard","family":"Wellein","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2020,6,15]]},"reference":[{"key":"21_CR1","unstructured":"Intel 64 and IA-32 Architectures Optimization Reference Manual. Intel Press, 2016 June 2016. http:\/\/www.intel.com\/content\/dam\/www\/public\/us\/en\/documents\/manuals\/64-ia-32-architectures-optimization-manual.pdf"},{"key":"21_CR2","doi-asserted-by":"crossref","unstructured":"Afzal, A., Hager, G., Wellein, G.: Desynchronization and wave pattern formation in MPI-parallel and hybrid memory-bound programs (2020). https:\/\/arxiv.org\/abs\/2002.02989. Accepted for ISC High Performance 2020","DOI":"10.1007\/978-3-030-50743-5_20"},{"key":"21_CR3","doi-asserted-by":"publisher","unstructured":"Alappat, C.L., et al.: A recursive algebraic coloring technique for hardware-efficient symmetric sparse matrix-vector multiplication (2020). Accepted for publication in ACM Transactions on Parallel Computing.https:\/\/doi.org\/10.1145\/3399732","DOI":"10.1145\/3399732"},{"key":"21_CR4","unstructured":"ARM: ARM Cortex-A75 Core Technical Reference Manual - Write streaming mode. http:\/\/infocenter.arm.com\/help\/index.jsp?topic=\/com.arm.doc.100403_0200_00_en\/lto1473834732563.html. Accessed 26 Mar 2020"},{"issue":"1","key":"21_CR5","doi-asserted-by":"publisher","first-page":"1:1","DOI":"10.1145\/2049662.2049663","volume":"38","author":"TA Davis","year":"2011","unstructured":"Davis, T.A., Hu, Y.: The University of Florida sparse matrix collection. ACM Trans. Math. Softw. 38(1), 1:1\u20131:25 (2011). http:\/\/doi.acm.org\/10.1145\/2049662.2049663","journal-title":"ACM Trans. Math. Softw."},{"key":"21_CR6","doi-asserted-by":"crossref","unstructured":"Hammond, S., et al.: Evaluating the Marvell ThunderX2 server processor for HPC workloads. In: The 6th Special Session on High-Performance Computing Benchmarking and Optimization (HPBench 2019) (2019)","DOI":"10.1109\/HPCS48598.2019.9188171"},{"key":"21_CR7","doi-asserted-by":"publisher","unstructured":"Hammond, S., Vaughan, C., Hughes, C.: Evaluating the Intel Skylake Xeon processor for HPC workloads. In: 2018 International Conference on High Performance Computing Simulation (HPCS), pp. 342\u2013349, July 2018. https:\/\/doi.org\/10.1109\/HPCS.2018.00064","DOI":"10.1109\/HPCS.2018.00064"},{"key":"21_CR8","unstructured":"Wong, H.: Intel Ivy Bridge Cache replacement policy. http:\/\/blog.stuffedcow.net\/2013\/01\/ivb-cache-replacement\/"},{"key":"21_CR9","series-title":"Lecture Notes in Computer Science","doi-asserted-by":"publisher","first-page":"210","DOI":"10.1007\/978-3-319-30695-7_16","volume-title":"Architecture of Computing Systems \u2013 ARCS 2016","author":"J Hofmann","year":"2016","unstructured":"Hofmann, J., Fey, D., Eitzinger, J., Hager, G., Wellein, G.: Analysis of Intel\u2019s haswell microarchitecture using the ECM model and microbenchmarks. In: Hannig, F., Cardoso, J.M.P., Pionteck, T., Fey, D., Schr\u00f6der-Preikschat, W., Teich, J. (eds.) ARCS 2016. LNCS, vol. 9637, pp. 210\u2013222. Springer, Cham (2016). https:\/\/doi.org\/10.1007\/978-3-319-30695-7_16"},{"key":"21_CR10","series-title":"Lecture Notes in Computer Science","doi-asserted-by":"publisher","first-page":"22","DOI":"10.1007\/978-3-319-92040-5_2","volume-title":"High Performance Computing","author":"J Hofmann","year":"2018","unstructured":"Hofmann, J., Hager, G., Fey, D.: On the accuracy and usefulness of analytic energy models for contemporary multicore processors. In: Yokota, R., Weiland, M., Keyes, D., Trinitis, C. (eds.) ISC High Performance 2018. LNCS, vol. 10876, pp. 22\u201343. Springer, Cham (2018). https:\/\/doi.org\/10.1007\/978-3-319-92040-5_2"},{"key":"21_CR11","series-title":"Lecture Notes in Computer Science","doi-asserted-by":"publisher","first-page":"294","DOI":"10.1007\/978-3-319-58667-0_16","volume-title":"High Performance Computing","author":"J Hofmann","year":"2017","unstructured":"Hofmann, J., Hager, G., Wellein, G., Fey, D.: An analysis of core- and chip-level architectural features in four generations of intel server processors. In: Kunkel, J.M., Yokota, R., Balaji, P., Keyes, D. (eds.) ISC 2017. LNCS, vol. 10266, pp. 294\u2013314. Springer, Cham (2017). https:\/\/doi.org\/10.1007\/978-3-319-58667-0_16"},{"issue":"3","key":"21_CR12","doi-asserted-by":"publisher","first-page":"12:1","DOI":"10.1145\/3155290","volume":"4","author":"TM Malas","year":"2017","unstructured":"Malas, T.M., Hager, G., Ltaief, H., Keyes, D.E.: Multidimensional intratile parallelization for memory-starved stencil computations. ACM Trans. Parallel Comput. 4(3), 12:1\u201312:32 (2017). http:\/\/doi.acm.org\/10.1145\/3155290","journal-title":"ACM Trans. Parallel Comput."},{"key":"21_CR13","series-title":"Lecture Notes in Computer Science","doi-asserted-by":"publisher","first-page":"172","DOI":"10.1007\/978-3-319-17248-4_9","volume-title":"High Performance Computing Systems. Performance Modeling, Benchmarking, and Simulation","author":"V Marjanovi\u0107","year":"2015","unstructured":"Marjanovi\u0107, V., Gracia, J., Glass, C.W.: Performance modeling of the HPCG benchmark. In: Jarvis, S.A., Wright, S.A., Hammond, S.D. (eds.) PMBS 2014. LNCS, vol. 8966, pp. 172\u2013192. Springer, Cham (2015). https:\/\/doi.org\/10.1007\/978-3-319-17248-4_9"},{"key":"21_CR14","first-page":"19","volume":"2","author":"JD McCalpin","year":"1995","unstructured":"McCalpin, J.D.: Memory bandwidth and machine balance in current high performance computers. IEEE Comput. Soc. Tech. Comm. Comput. Archit. (TCCA) Newsl. 2, 19\u201325 (1995)","journal-title":"IEEE Comput. Soc. Tech. Comm. Comput. Archit. (TCCA) Newsl."},{"issue":"16","key":"21_CR15","doi-asserted-by":"publisher","first-page":"e5110","DOI":"10.1002\/cpe.5110","volume":"31","author":"S McIntosh-Smith","year":"2019","unstructured":"McIntosh-Smith, S., Price, J., Deakin, T., Poenaru, A.: A performance analysis of the first generation of HPC-optimized arm processors. Concurr. Comput.: Pract. Exp. 31(16), e5110 (2019). https:\/\/onlinelibrary.wiley.com\/doi\/abs\/10.1002\/cpe.5110. e5110 cpe.5110","journal-title":"Concurr. Comput.: Pract. Exp."},{"key":"21_CR16","unstructured":"McVoy, L., Staelin, C.: Lmbench: portable tools for performance analysis. In: Proceedings of the 1996 Annual Conference on USENIX Annual Technical Conference ATEC 1996, pp. 23\u201323. USENIX Association, Berkeley (1996). http:\/\/dl.acm.org\/citation.cfm?id=1268299.1268322"},{"key":"21_CR17","doi-asserted-by":"crossref","unstructured":"Molka, D., Hackenberg, D., Sch\u00f6ne, R.: Main memory and cache performance of Intel Sandy Bridge and AMD Bulldozer. In: Proceedings of the Workshop on Memory Systems Performance and Correctness MSPC 2014, pp. 4:1\u20134:10. ACM, New York (2014). http:\/\/doi.acm.org\/10.1145\/2618128.2618129","DOI":"10.1145\/2618128.2618129"},{"key":"21_CR18","doi-asserted-by":"publisher","first-page":"226","DOI":"10.1016\/j.jcp.2016.08.027","volume":"325","author":"A Pieper","year":"2016","unstructured":"Pieper, A., et al.: High-performance implementation of Chebyshev filter diagonalization for interior eigenvalue computations. J. Comput. Phys. 325, 226\u2013243 (2016). http:\/\/www.sciencedirect.com\/science\/article\/pii\/S0021999116303837","journal-title":"J. Comput. Phys."},{"key":"21_CR19","doi-asserted-by":"crossref","unstructured":"Qureshi, M.K., Jaleel, A., Patt, Y.N., Steely, S.C., Emer, J.: Adaptive insertion policies for high performance caching. In: Proceedings of the 34th Annual International Symposium on Computer Architecture ISCA 2007, pp. 381\u2013391. ACM, New York (2007). http:\/\/doi.acm.org\/10.1145\/1250662.1250709","DOI":"10.1145\/1250662.1250709"},{"key":"21_CR20","doi-asserted-by":"publisher","unstructured":"Saini, S., Hood, R.: Performance evaluation of Intel Broadwell nodes based supercomputer using computational fluid dynamics and climate applications. In: 2017 IEEE 19th International Conference on High Performance Computing and Communications Workshops (HPCCWS), pp. 58\u201365, December 2017. https:\/\/doi.org\/10.1109\/HPCCWS.2017.00015","DOI":"10.1109\/HPCCWS.2017.00015"},{"key":"21_CR21","doi-asserted-by":"publisher","unstructured":"Saini, S., Hood, R., Chang, J., Baron, J.: Performance evaluation of an Intel Haswell- and Ivy Bridge-based supercomputer using scientific and engineering applications. In: 2016 IEEE 18th International Conference on High Performance Computing and Communications; IEEE 14th International Conference on Smart City; IEEE 2nd International Conference on Data Science and Systems (HPCC\/SmartCity\/DSS), pp. 1196\u20131203, December 2016. https:\/\/doi.org\/10.1109\/HPCC-SmartCity-DSS.2016.0167","DOI":"10.1109\/HPCC-SmartCity-DSS.2016.0167"},{"key":"21_CR22","doi-asserted-by":"publisher","unstructured":"Staar, P.W.J., et al.: Stochastic matrix-function estimators: scalable big-data kernels with high performance. In: 2016 IEEE International Parallel and Distributed Processing Symposium (IPDPS), pp. 812\u2013821, May 2016. https:\/\/doi.org\/10.1109\/IPDPS.2016.34","DOI":"10.1109\/IPDPS.2016.34"},{"issue":"3","key":"21_CR23","doi-asserted-by":"publisher","first-page":"684","DOI":"10.1007\/s10915-013-9813-x","volume":"60","author":"AY Suhov","year":"2014","unstructured":"Suhov, A.Y.: An accurate polynomial approximation of exponential integrators. J. Sci. Comput. 60(3), 684\u2013698 (2014). https:\/\/doi.org\/10.1007\/s10915-013-9813-x","journal-title":"J. Sci. Comput."},{"key":"21_CR24","doi-asserted-by":"publisher","first-page":"27","DOI":"10.1007\/978-3-642-31476-6_3","volume-title":"Parallel Tools Workshop","author":"J Treibig","year":"2011","unstructured":"Treibig, J., Hager, G., Wellein, G.: likwid-bench: an extensible microbenchmarking platform for x86 multicore compute nodes. In: Brunst, H., M\u00fcller, M., Nagel, W., Resch, M. (eds.) Parallel Tools Workshop, pp. 27\u201336. Springer, Heidelberg (2011). https:\/\/doi.org\/10.1007\/978-3-642-31476-6_3"},{"key":"21_CR25","doi-asserted-by":"publisher","unstructured":"Wellein, G., Hager, G., Zeiser, T., Wittmann, M., Fehske, H.: Efficient temporal blocking for stencil computations by multicore-aware wavefront parallelization. In: 2009 33rd Annual IEEE International Computer Software and Applications Conference, vol. 1, pp. 579\u2013586, July 2009. https:\/\/doi.org\/10.1109\/COMPSAC.2009.82","DOI":"10.1109\/COMPSAC.2009.82"}],"container-title":["Lecture Notes in Computer Science","High Performance Computing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/978-3-030-50743-5_21","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,12,18]],"date-time":"2023-12-18T20:05:39Z","timestamp":1702929939000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/978-3-030-50743-5_21"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020]]},"ISBN":["9783030507428","9783030507435"],"references-count":25,"URL":"https:\/\/doi.org\/10.1007\/978-3-030-50743-5_21","relation":{},"ISSN":["0302-9743","1611-3349"],"issn-type":[{"value":"0302-9743","type":"print"},{"value":"1611-3349","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020]]},"assertion":[{"value":"15 June 2020","order":1,"name":"first_online","label":"First Online","group":{"name":"ChapterHistory","label":"Chapter History"}},{"value":"ISC High Performance","order":1,"name":"conference_acronym","label":"Conference Acronym","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"International Conference on High Performance Computing","order":2,"name":"conference_name","label":"Conference Name","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"Frankfurt am Main","order":3,"name":"conference_city","label":"Conference City","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"Germany","order":4,"name":"conference_country","label":"Conference Country","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"2020","order":5,"name":"conference_year","label":"Conference Year","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"22 June 2020","order":7,"name":"conference_start_date","label":"Conference Start Date","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"25 June 2020","order":8,"name":"conference_end_date","label":"Conference End Date","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"35","order":9,"name":"conference_number","label":"Conference Number","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"supercomputing2020","order":10,"name":"conference_id","label":"Conference ID","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"https:\/\/www.isc-hpc.com\/","order":11,"name":"conference_url","label":"Conference URL","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"Double-blind","order":1,"name":"type","label":"Type","group":{"name":"ConfEventPeerReviewInformation","label":"Peer Review Information (provided by the conference organizers)"}},{"value":"Linklings","order":2,"name":"conference_management_system","label":"Conference Management System","group":{"name":"ConfEventPeerReviewInformation","label":"Peer Review Information (provided by the conference organizers)"}},{"value":"87","order":3,"name":"number_of_submissions_sent_for_review","label":"Number of Submissions Sent for Review","group":{"name":"ConfEventPeerReviewInformation","label":"Peer Review Information (provided by the conference organizers)"}},{"value":"27","order":4,"name":"number_of_full_papers_accepted","label":"Number of Full Papers Accepted","group":{"name":"ConfEventPeerReviewInformation","label":"Peer Review Information (provided by the conference organizers)"}},{"value":"0","order":5,"name":"number_of_short_papers_accepted","label":"Number of Short Papers Accepted","group":{"name":"ConfEventPeerReviewInformation","label":"Peer Review Information (provided by the conference organizers)"}},{"value":"31% - The value is computed by the equation \"Number of Full Papers Accepted \/ Number of Submissions Sent for Review * 100\" and then rounded to a whole number.","order":6,"name":"acceptance_rate_of_full_papers","label":"Acceptance Rate of Full Papers","group":{"name":"ConfEventPeerReviewInformation","label":"Peer Review Information (provided by the conference organizers)"}},{"value":"3.73","order":7,"name":"average_number_of_reviews_per_paper","label":"Average Number of Reviews per Paper","group":{"name":"ConfEventPeerReviewInformation","label":"Peer Review Information (provided by the conference organizers)"}},{"value":"4.33","order":8,"name":"average_number_of_papers_per_reviewer","label":"Average Number of Papers per Reviewer","group":{"name":"ConfEventPeerReviewInformation","label":"Peer Review Information (provided by the conference organizers)"}},{"value":"No","order":9,"name":"external_reviewers_involved","label":"External Reviewers Involved","group":{"name":"ConfEventPeerReviewInformation","label":"Peer Review Information (provided by the conference organizers)"}},{"value":"The conference was held virtually due to the COVID-19 pandemic.","order":10,"name":"additional_info_on_review_process","label":"Additional Info on Review Process","group":{"name":"ConfEventPeerReviewInformation","label":"Peer Review Information (provided by the conference organizers)"}}]}}