{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,13]],"date-time":"2026-05-13T19:10:54Z","timestamp":1778699454865,"version":"3.51.4"},"reference-count":59,"publisher":"Association for Computing Machinery (ACM)","issue":"1","license":[{"start":{"date-parts":[[2020,12,8]],"date-time":"2020-12-08T00:00:00Z","timestamp":1607385600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by-sa\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100002347","name":"Bundesministerium f\u00fcr Bildung und Forschung","doi-asserted-by":"crossref","award":["01IH16003C"],"award-info":[{"award-number":["01IH16003C"]}],"id":[{"id":"10.13039\/501100002347","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100001659","name":"Deutsche Forschungsgemeinschaft","doi-asserted-by":"publisher","award":["EXC 2181\/1 - 390900948"],"award-info":[{"award-number":["EXC 2181\/1 - 390900948"]}],"id":[{"id":"10.13039\/501100001659","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Math. Softw."],"published-print":{"date-parts":[[2021,3,31]]},"abstract":"<jats:p>SIMD vectorization has lately become a key challenge in high-performance computing. However, hand-written explicitly vectorized code often poses a threat to the software\u2019s sustainability. In this publication, we solve this sustainability and performance portability issue by enriching the simulation framework dune-pdelab with a code generation approach. The approach is based on the well-known domain-specific language UFL but combines it with loopy, a more powerful intermediate representation for the computational kernel. Given this flexible tool, we present and implement a new class of vectorization strategies for the assembly of Discontinuous Galerkin methods on hexahedral meshes exploiting the finite element\u2019s tensor product structure. The performance-optimal variant from this class is chosen by the code generator through an auto-tuning approach. The implementation is done within the open source PDE software framework Dune and the discretization module dune-pdelab. The strength of the proposed approach is illustrated with performance measurements for DG schemes for a scalar diffusion reaction equation and the Stokes equation. In our measurements, we utilize both the AVX2 and the AVX512 instruction set, achieving 30% to 40% of the machine\u2019s theoretical peak performance for one matrix-free application of the operator.<\/jats:p>","DOI":"10.1145\/3424144","type":"journal-article","created":{"date-parts":[[2020,12,9]],"date-time":"2020-12-09T23:11:26Z","timestamp":1607555486000},"page":"1-31","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":17,"title":["Automatic Code Generation for High-performance Discontinuous Galerkin Methods on Modern Architectures"],"prefix":"10.1145","volume":"47","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-6140-2332","authenticated-orcid":false,"given":"Dominic","family":"Kempf","sequence":"first","affiliation":[{"name":"Heidelberg University, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ren\u00e9","family":"He\u00df","sequence":"additional","affiliation":[{"name":"Heidelberg University, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Steffen","family":"M\u00fcthing","sequence":"additional","affiliation":[{"name":"Heidelberg University, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Peter","family":"Bastian","sequence":"additional","affiliation":[{"name":"Heidelberg University, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2020,12,8]]},"reference":[{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-41321-1_2"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1137\/11082539X"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1145\/2566630"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/2566630"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/1268776.1268779"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10596-014-9426-y"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1007\/s007910050003"},{"key":"e_1_2_1_9_1","first-page":"2","article-title":"A generic grid interface for parallel and adaptive scientific computing. part II: Implementation and tests in DUNE","volume":"82","author":"Bastian Peter","year":"2008","unstructured":"Peter Bastian , Markus Blatt , Andreas Dedner , Christian Engwer , Robert Kl\u00f6fkorn , Ralf Kornhuber , Mario Ohlberger , and Oliver Sander . 2008 . A generic grid interface for parallel and adaptive scientific computing. part II: Implementation and tests in DUNE . Computing 82 , 2 \u2013 3 (2008), 121--138. Peter Bastian, Markus Blatt, Andreas Dedner, Christian Engwer, Robert Kl\u00f6fkorn, Ralf Kornhuber, Mario Ohlberger, and Oliver Sander. 2008. A generic grid interface for parallel and adaptive scientific computing. part II: Implementation and tests in DUNE. Computing 82, 2\u20133 (2008), 121--138.","journal-title":"Computing"},{"key":"e_1_2_1_10_1","first-page":"2","article-title":"A generic grid interface for parallel and adaptive scientific computing. part I: Abstract framework","volume":"82","author":"Bastian Peter","year":"2008","unstructured":"Peter Bastian , Markus Blatt , Andreas Dedner , Christian Engwer , Robert Kl\u00f6fkorn , Mario Ohlberger , and Oliver Sander . 2008 . A generic grid interface for parallel and adaptive scientific computing. part I: Abstract framework . Computing 82 , 2 \u2013 3 (2008), 103--119. Peter Bastian, Markus Blatt, Andreas Dedner, Christian Engwer, Robert Kl\u00f6fkorn, Mario Ohlberger, and Oliver Sander. 2008. A generic grid interface for parallel and adaptive scientific computing. part I: Abstract framework. Computing 82, 2\u20133 (2008), 103--119.","journal-title":"Computing"},{"key":"e_1_2_1_11_1","first-page":"294","article-title":"Generic implementation of finite element methods in the distributed and unified numerics environment (DUNE)","volume":"46","author":"Bastian Peter","year":"2010","unstructured":"Peter Bastian , Felix Heimann , and Sven Marnach . 2010 . Generic implementation of finite element methods in the distributed and unified numerics environment (DUNE) . Kybernetika 46 , 2 (2010), 294 -- 315 . Peter Bastian, Felix Heimann, and Sven Marnach. 2010. Generic implementation of finite element methods in the distributed and unified numerics environment (DUNE). Kybernetika 46, 2 (2010), 294--315.","journal-title":"Kybernetika"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.jcp.2019.06.001"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1145\/225545.225548"},{"key":"e_1_2_1_14_1","volume-title":"Computation and Applications. Lecture Notes in Computational Science and Engineering","volume":"11","author":"Cockburn B.","year":"2000","unstructured":"B. Cockburn , S. Y. Lin , and C.-W. Shu ( Eds .). 2000 . Discontinuous Galerkin Methods. Theory , Computation and Applications. Lecture Notes in Computational Science and Engineering , Vol. 11 . Springer-Verlag. B. Cockburn, S. Y. Lin, and C.-W. Shu (Eds.). 2000. Discontinuous Galerkin Methods. Theory, Computation and Applications. Lecture Notes in Computational Science and Engineering, Vol. 11. Springer-Verlag."},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.5555\/1413370.1413375"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1109\/JSSC.1974.1050511"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11227-017-2177-5"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-58312-4_11"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1093\/imanum\/drm050"},{"key":"e_1_2_1_20_1","volume-title":"et\u00a0al","author":"Fischer Paul","year":"2020","unstructured":"Paul Fischer , Misun Min , Thilina Rathnayake , Som Dutta , Tzanio Kolev , Veselin Dobrev , Jean-Sylvain Camier , Martin Kronbichler , Tim Warburton , Kasia Swirydowicz , et\u00a0al . 2020 . Scalability of high-performance PDE solvers. Arxiv Preprint Arxiv :2004.06722 (2020). Paul Fischer, Misun Min, Thilina Rathnayake, Som Dutta, Tzanio Kolev, Veselin Dobrev, Jean-Sylvain Camier, Martin Kronbichler, Tim Warburton, Kasia Swirydowicz, et\u00a0al. 2020. Scalability of high-performance PDE solvers. Arxiv Preprint Arxiv:2004.06722 (2020)."},{"key":"e_1_2_1_21_1","volume-title":"Proceedings of the 22nd AIAA Computational Fluid Dynamics Conference","author":"Fischer P. F.","unstructured":"P. F. Fischer , K. Heisey , and M. Min . 2015. Scaling limits for PDE-based simulation . In Proceedings of the 22nd AIAA Computational Fluid Dynamics Conference . Dallas, TX. P. F. Fischer, K. Heisey, and M. Min. 2015. Scaling limits for PDE-based simulation. In Proceedings of the 22nd AIAA Computational Fluid Dynamics Conference. Dallas, TX."},{"key":"e_1_2_1_22_1","unstructured":"Agner Fog. [n.d.]. VCL C++ vector class library v 1.30. Retrieved from http:\/\/www.agner.org\/optimize\/vectorclass.pdf.  Agner Fog. [n.d.]. VCL C++ vector class library v 1.30. Retrieved from http:\/\/www.agner.org\/optimize\/vectorclass.pdf."},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1109\/JPROC.2004.840491"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1090\/S0025-5718-04-01652-7"},{"key":"e_1_2_1_25_1","unstructured":"Google. 2020. Benchmark. Retrieved from https:\/\/github.com\/google\/benchmark.  Google. 2020. Benchmark. Retrieved from https:\/\/github.com\/google\/benchmark."},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1109\/SC.2016.83"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1137\/17M1130642"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.cma.2013.04.010"},{"key":"e_1_2_1_29_1","volume-title":"Sherwin","author":"Karniadakis George Em","year":"2005","unstructured":"George Em Karniadakis and Spencer J . Sherwin . 2005 . Spectral\/hp Element Methods for CFD. Oxford University Press . George Em Karniadakis and Spencer J. Sherwin. 2005. Spectral\/hp Element Methods for CFD. Oxford University Press."},{"key":"e_1_2_1_30_1","unstructured":"Dominic Kempf and Ren\u00e9 He\u00df. 2020. Automatic Code Generation for High-Performance Discontinuous Galerkin Methods on Modern Architectures\u2014Software Stack: Retrieved from https:\/\/doi.org\/10.5281\/zenodo.377926. DOI:https:\/\/doi.org\/10.5281\/zenodo.3779266  Dominic Kempf and Ren\u00e9 He\u00df. 2020. Automatic Code Generation for High-Performance Discontinuous Galerkin Methods on Modern Architectures\u2014Software Stack: Retrieved from https:\/\/doi.org\/10.5281\/zenodo.377926. DOI:https:\/\/doi.org\/10.5281\/zenodo.3779266"},{"key":"e_1_2_1_31_1","first-page":"151","article-title":"System testing in scientific numerical software frameworks using the example of DUNE","volume":"5","author":"Kempf Dominic","year":"2017","unstructured":"Dominic Kempf and Timo Koch . 2017 . System testing in scientific numerical software frameworks using the example of DUNE . Arch. Numer. Softw. 5 , 1 (2017), 151 -- 168 . Dominic Kempf and Timo Koch. 2017. System testing in scientific numerical software frameworks using the example of DUNE. Arch. Numer. Softw. 5, 1 (2017), 151--168.","journal-title":"Arch. Numer. Softw."},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1145\/3126908.3126941"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1145\/1163641.1163644"},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1007\/s00211-011-0431-y"},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1145\/2627373.2627387"},{"key":"e_1_2_1_36_1","volume-title":"Proceedings of the 3rd ACM SIGPLAN International Workshop on Libraries, Languages, and Compilers for Array Programming (ARRAY\u201916)","author":"Kl\u00f6ckner Andreas","unstructured":"Andreas Kl\u00f6ckner , Lucas C. Wilcox , and T. Warburton . 2016. Array program transformation with Loo.Py by example: High-order finite elements . In Proceedings of the 3rd ACM SIGPLAN International Workshop on Libraries, Languages, and Compilers for Array Programming (ARRAY\u201916) . ACM, New York, NY, 9--16. DOI:https:\/\/doi.org\/10.1145\/2935323.2935325 Andreas Kl\u00f6ckner, Lucas C. Wilcox, and T. Warburton. 2016. Array program transformation with Loo.Py by example: High-order finite elements. In Proceedings of the 3rd ACM SIGPLAN International Workshop on Libraries, Languages, and Compilers for Array Programming (ARRAY\u201916). ACM, New York, NY, 9--16. DOI:https:\/\/doi.org\/10.1145\/2935323.2935325"},{"key":"e_1_2_1_37_1","unstructured":"Tzanio Kolev et\u00a0al. [n.d.]. MFEM: Modular finite element methods. Retrieved from http:\/\/mfem.org.  Tzanio Kolev et\u00a0al. [n.d.]. MFEM: Modular finite element methods. Retrieved from http:\/\/mfem.org."},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.jcp.2017.07.039"},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1002\/spe.1149"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1137\/130930352"},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.compfluid.2012.04.012"},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1145\/3325864"},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1137\/16M110455X"},{"key":"e_1_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.cma.2014.10.048"},{"key":"e_1_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.1016\/S0377-0427(00)00393-9"},{"key":"e_1_2_1_46_1","volume-title":"et\u00a0al","author":"Logg Anders","year":"2012","unstructured":"Anders Logg , Kent-Andre Mardal , Garth N. Wells , et\u00a0al . 2012 . Automated Solution of Differential Equations by the Finite Element Method. Springer . DOI:https:\/\/doi.org\/10.1007\/978-3-642-23099-8 Anders Logg, Kent-Andre Mardal, Garth N. Wells, et\u00a0al. 2012. Automated Solution of Differential Equations by the Finite Element Method. Springer. DOI:https:\/\/doi.org\/10.1007\/978-3-642-23099-8"},{"key":"e_1_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.1137\/15M1021167"},{"key":"e_1_2_1_48_1","unstructured":"Steffen M\u00fcthing Marian Piatkowski and Peter Bastian. 2017. High-performance implementation of matrix-free high-order discontinuous Galerkin methods. Retrieved from https:\/\/Arxiv:1711.10885.  Steffen M\u00fcthing Marian Piatkowski and Peter Bastian. 2017. High-performance implementation of matrix-free high-order discontinuous Galerkin methods. Retrieved from https:\/\/Arxiv:1711.10885."},{"key":"e_1_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.1016\/0021-9991(80)90005-4"},{"key":"e_1_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.jcp.2017.10.030"},{"key":"e_1_2_1_51_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.jcp.2017.11.035"},{"key":"e_1_2_1_52_1","volume-title":"Kelly","author":"Rathgeber Florian","year":"2015","unstructured":"Florian Rathgeber , David A. Ham , Lawrence Mitchell , Michael Lange , Fabio Luporini , Andrew T. T. McRae , Gheorghe-Teodor Bercea , Graham R. Markall , and Paul H. J . Kelly . 2015 . Firedrake : Automating the finite element method by composing abstractions. Retrieved from http:\/\/arxiv.org\/abs\/1501.01809. Florian Rathgeber, David A. Ham, Lawrence Mitchell, Michael Lange, Fabio Luporini, Andrew T. T. McRae, Gheorghe-Teodor Bercea, Graham R. Markall, and Paul H. J. Kelly. 2015. Firedrake: Automating the finite element method by composing abstractions. Retrieved from http:\/\/arxiv.org\/abs\/1501.01809."},{"key":"e_1_2_1_53_1","unstructured":"J. Sch\u00f6berl A. Arnold J. Erb J. M. Melenk and T. P. Wihler. 2017. C++11 Implementation of Finite Elements in NGSolve. Technical Report.  J. Sch\u00f6berl A. Arnold J. Erb J. M. Melenk and T. P. Wihler. 2017. C++11 Implementation of Finite Elements in NGSolve. Technical Report."},{"key":"e_1_2_1_54_1","doi-asserted-by":"publisher","DOI":"10.1177\/1094342004041295"},{"key":"e_1_2_1_55_1","volume-title":"Kelly","author":"Sun Tianjiao","year":"2019","unstructured":"Tianjiao Sun , Lawrence Mitchell , Kaushik Kulkarni , Andreas Kl\u00f6ckner , David A. Ham , and Paul H. J . Kelly . 2019 . A study of vectorization for matrix-free finite element methods. Retrieved from https:\/\/Arxiv:1903.08243. Tianjiao Sun, Lawrence Mitchell, Kaushik Kulkarni, Andreas Kl\u00f6ckner, David A. Ham, and Paul H. J. Kelly. 2019. A study of vectorization for matrix-free finite element methods. Retrieved from https:\/\/Arxiv:1903.08243."},{"key":"e_1_2_1_56_1","volume-title":"The free lunch is over. Dr. Dobb\u2019s J. 30, 3","author":"Sutter Herb","year":"2005","unstructured":"Herb Sutter . 2005. The free lunch is over. Dr. Dobb\u2019s J. 30, 3 ( 2005 ). Herb Sutter. 2005. The free lunch is over. Dr. Dobb\u2019s J. 30, 3 (2005)."},{"key":"e_1_2_1_57_1","doi-asserted-by":"publisher","DOI":"10.1177\/1094342018816368"},{"key":"e_1_2_1_58_1","unstructured":"Ulrich Trottenberg Cornelius W. Oosterlee and Anton Schuller. 2000. Multigrid. Elsevier.  Ulrich Trottenberg Cornelius W. Oosterlee and Anton Schuller. 2000. Multigrid. Elsevier."},{"key":"e_1_2_1_59_1","doi-asserted-by":"publisher","DOI":"10.1145\/373574.373576"},{"key":"e_1_2_1_60_1","doi-asserted-by":"publisher","DOI":"10.1145\/1498765.1498785"}],"container-title":["ACM Transactions on Mathematical Software"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3424144","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3424144","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T22:01:51Z","timestamp":1750197711000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3424144"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,12,8]]},"references-count":59,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2021,3,31]]}},"alternative-id":["10.1145\/3424144"],"URL":"https:\/\/doi.org\/10.1145\/3424144","relation":{},"ISSN":["0098-3500","1557-7295"],"issn-type":[{"value":"0098-3500","type":"print"},{"value":"1557-7295","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,12,8]]},"assertion":[{"value":"2018-12-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2020-09-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2020-12-08","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}