{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,27]],"date-time":"2026-06-27T07:11:38Z","timestamp":1782544298245,"version":"3.54.5"},"reference-count":54,"publisher":"Springer Science and Business Media LLC","issue":"6","license":[{"start":{"date-parts":[[2025,4,23]],"date-time":"2025-04-23T00:00:00Z","timestamp":1745366400000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2025,4,23]],"date-time":"2025-04-23T00:00:00Z","timestamp":1745366400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100001659","name":"Deutsche Forschungsgemeinschaft","doi-asserted-by":"publisher","award":["AI 117\/7-1"],"award-info":[{"award-number":["AI 117\/7-1"]}],"id":[{"id":"10.13039\/501100001659","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001659","name":"Deutsche Forschungsgemeinschaft","doi-asserted-by":"publisher","award":["KE 2844\/1-1"],"award-info":[{"award-number":["KE 2844\/1-1"]}],"id":[{"id":"10.13039\/501100001659","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001659","name":"Deutsche Forschungsgemeinschaft","doi-asserted-by":"publisher","award":["KE 2844\/1-1"],"award-info":[{"award-number":["KE 2844\/1-1"]}],"id":[{"id":"10.13039\/501100001659","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001659","name":"Deutsche Forschungsgemeinschaft","doi-asserted-by":"publisher","award":["AI 117\/7-1"],"award-info":[{"award-number":["AI 117\/7-1"]}],"id":[{"id":"10.13039\/501100001659","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100020618","name":"Universit\u00e4t Bayreuth","doi-asserted-by":"crossref","id":[{"id":"10.13039\/100020618","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J Supercomput"],"abstract":"<jats:title>Abstract<\/jats:title>\n          <jats:p>SYCL is an open standard for targeting heterogeneous hardware from C++. In this work, we evaluate a SYCL implementation for a discontinuous Galerkin discretization of the 2D shallow water equations targeting CPUs, GPUs, and also FPGAs. The discretization uses polynomial orders zero to two on unstructured triangular meshes. Separating memory accesses from the numerical code allow us to optimize data accesses for the target architecture. A performance analysis shows good portability across x86 and ARM CPUs, GPUs from different vendors, and even two variants of Intel Stratix 10 FPGAs. Measuring the energy to solution shows that GPUs yield an up to 10x higher energy efficiency in terms of degrees of freedom per joule compared to CPUs. With custom designed caches, FPGAs offer a meaningful complement to the other architectures with particularly good computational performance on smaller meshes. FPGAs with High Bandwidth Memory are less affected by bandwidth issues and have similar energy efficiency as latest generation CPUs.<\/jats:p>","DOI":"10.1007\/s11227-025-07063-7","type":"journal-article","created":{"date-parts":[[2025,4,23]],"date-time":"2025-04-23T12:24:35Z","timestamp":1745411075000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":3,"title":["Analyzing performance portability for a SYCL implementation of the 2D shallow water equations"],"prefix":"10.1007","volume":"81","author":[{"given":"Markus","family":"B\u00fcttner","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Christoph","family":"Alt","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Tobias","family":"Kenter","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Harald","family":"K\u00f6stler","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Christian","family":"Plessl","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Vadym","family":"Aizinger","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2025,4,23]]},"reference":[{"issue":"5","key":"7063_CR1","doi-asserted-by":"publisher","first-page":"28","DOI":"10.1109\/MCSE.2021.3097276","volume":"23","author":"SJ Pennycook","year":"2021","unstructured":"Pennycook SJ, Sewall JD, Jacobsen DW, Deakin T, McIntosh-Smith S (2021) Navigating performance, portability, and productivity. Computi Sci Eng 23(5):28\u201338. https:\/\/doi.org\/10.1109\/MCSE.2021.3097276","journal-title":"Computi Sci Eng"},{"key":"7063_CR2","doi-asserted-by":"publisher","unstructured":"Pennycook SJ, Sewall JD (2021) Revisiting a metric for performance portability. In: 2021 International Workshop on Performance, Portability and Productivity in HPC (P3HPC), pp. 1\u20139 . https:\/\/doi.org\/10.1109\/P3HPC54578.2021.00004","DOI":"10.1109\/P3HPC54578.2021.00004"},{"issue":"1","key":"7063_CR3","doi-asserted-by":"publisher","first-page":"154","DOI":"10.1002\/spe.3002","volume":"52","author":"A Marowka","year":"2022","unstructured":"Marowka A (2022) Reformulation of the performance portability metric. Softw Pract Exp 52(1):154\u2013171. https:\/\/doi.org\/10.1002\/spe.3002","journal-title":"Softw Pract Exp"},{"issue":"25","key":"7063_CR4","doi-asserted-by":"publisher","first-page":"7868","DOI":"10.1002\/cpe.7868","volume":"35","author":"A Marowka","year":"2023","unstructured":"Marowka A (2023) A comparison of two performance portability metrics. Concurr Comput Pract Exp 35(25):7868. https:\/\/doi.org\/10.1002\/cpe.7868","journal-title":"Concurr Comput Pract Exp"},{"issue":"4","key":"7063_CR5","doi-asserted-by":"publisher","first-page":"805","DOI":"10.1109\/TPDS.2021.3097283","volume":"33","author":"CR Trott","year":"2022","unstructured":"Trott CR, Lebrun-Grandi\u00e9 D, Arndt D, Ciesko J, Dang V, Ellingwood N, Gayatri R, Harvey E, Hollman DS, Ibanez D, Liber N, Madsen J, Miles J, Poliakoff D, Powell A, Rajamanickam S, Simberg M, Sunderland D, Turcksin B, Wilke J (2022) Kokkos 3: Programming model extensions for the exascale era. IEEE Trans Parallel Distrib Syst (TPDS) 33(4):805\u2013817. https:\/\/doi.org\/10.1109\/TPDS.2021.3097283","journal-title":"IEEE Trans Parallel Distrib Syst (TPDS)"},{"key":"7063_CR6","doi-asserted-by":"publisher","unstructured":"Deakin T, McIntosh-Smith S (2020) Evaluating the performance of HPC-style SYCL applications. In: Proceedings of the International Workshop on OpenCL. IWOCL \u201920, pp. 1\u201311. Association for Computing Machinery, New York, NY, USA . https:\/\/doi.org\/10.1145\/3388333.3388643","DOI":"10.1145\/3388333.3388643"},{"key":"7063_CR7","doi-asserted-by":"publisher","unstructured":"Reguly IZ (2023) Evaluating the performance portability of SYCL across CPUs and GPUs on bandwidth-bound applications. In: Proceedings of the SC \u201923 Workshops of The International Conference on High Performance Computing, Network, Storage, and Analysis. SC-W \u201923, pp. 1038\u20131047. Association for Computing Machinery, New York, NY, USA. https:\/\/doi.org\/10.1145\/3624062.3624180","DOI":"10.1145\/3624062.3624180"},{"key":"7063_CR8","doi-asserted-by":"publisher","unstructured":"Apanasevich L, Kale Y, Sharma H, Sokovic AM (2024) A Comparison of the performance of the molecular dynamics simulation package gromacs implemented in the sycl and cuda programming models https:\/\/doi.org\/10.48550\/ARXIV.2406.10362 . Publisher: arXiv, Version Number: 1. Accessed 2025-01-09","DOI":"10.48550\/ARXIV.2406.10362"},{"key":"7063_CR9","doi-asserted-by":"publisher","unstructured":"Rangel EM, Pennycook SJ, Pope A, Frontiere N, Ma Z, Madananth V (2023) A Performance-Portable SYCL Implementation of CRK-HACC for Exascale. In: Proceedings of the SC \u201923 Workshops of The International Conference on High Performance Computing, Network, Storage, and Analysis, pp. 1114\u20131125. ACM, Denver CO USA. https:\/\/doi.org\/10.1145\/3624062.3624187","DOI":"10.1145\/3624062.3624187"},{"key":"7063_CR10","doi-asserted-by":"publisher","unstructured":"Nichols NS, Childers JT, Burch TJ, Field L (2024) Improving Performance Portability of the Procedurally Generated High Energy Physics Event Generator MadGraph Using SYCL. In: Proceedings of the 12th International Workshop on OpenCL And SYCL. IWOCL \u201924, pp. 1\u20138. Association for Computing Machinery, New York, NY, USA. https:\/\/doi.org\/10.1145\/3648115.3648116","DOI":"10.1145\/3648115.3648116"},{"key":"7063_CR11","doi-asserted-by":"publisher","unstructured":"Alpay A, Heuveline V (2023) One pass to bind them: The first single-pass sycl compiler with unified code representation across backends. In: Proc Int Workshop on OpenCL (IWOCL). IWOCL \u201923. Association for Computing Machinery, New York, NY, USA. https:\/\/doi.org\/10.1145\/3585341.3585351","DOI":"10.1145\/3585341.3585351"},{"key":"7063_CR12","doi-asserted-by":"publisher","unstructured":"Meyer J, Alpay A, Hack S, Fr\u00f6ning H, Heuveline V (2023) Implementation techniques for SPMD kernels on CPUs. In: Proc Int Workshop on OpenCL (IWOCL). IWOCL \u201923. Association for Computing Machinery, New York, NY, USA. https:\/\/doi.org\/10.1145\/3585341.3585342","DOI":"10.1145\/3585341.3585342"},{"key":"7063_CR13","doi-asserted-by":"publisher","unstructured":"B\u00fcttner M, Alt C, Kenter T, K\u00f6stler H, Plessl C, Aizinger V (2024) Enabling Performance Portability for Shallow Water Equations on CPUs, GPUs, and FPGAs with SYCL. In: Proceedings of the platform for advanced scientific computing conference. ACM, Zurich Switzerland, pp 1\u201312. https:\/\/doi.org\/10.1145\/3659914.3659925","DOI":"10.1145\/3659914.3659925"},{"issue":"134110","key":"7063_CR14","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1063\/5.0018516","volume":"153","author":"S P\u00e1ll","year":"2020","unstructured":"P\u00e1ll S, Zhmurov A, Bauer P, Abraham M, Lundborg M, Gray A, Hess B, Lindahl E (2020) Heterogeneous parallelization and acceleration of molecular dynamics simulations in GROMACS. J Chem Phys 153(134110):1\u201315. https:\/\/doi.org\/10.1063\/5.0018516","journal-title":"J Chem Phys"},{"key":"7063_CR15","doi-asserted-by":"publisher","unstructured":"Dufek AS, Gayatri R, Mehta N, Doerfler D, Cook B, Ghadar Y, DeTar C (2021) Case study of using kokkos and sycl as performance-portable frameworks for milc-dslash benchmark on nvidia, amd and Intel gpus. In: 2021 International workshop on performance, portability and productivity in HPC (P3HPC), pp 57\u201367. https:\/\/doi.org\/10.1109\/P3HPC54578.2021.00009","DOI":"10.1109\/P3HPC54578.2021.00009"},{"key":"7063_CR16","doi-asserted-by":"publisher","unstructured":"Alpay A, Heuveline V (2020) SYCL beyond opencl: The architecture, current state and future direction of hipSYCL. In: Proc Int Workshop on OpenCL and SYCL (IWOCL). IWOCL \u201920. Association for Computing Machinery, New York, NY, USA. https:\/\/doi.org\/10.1145\/3388333.3388658","DOI":"10.1145\/3388333.3388658"},{"key":"7063_CR17","unstructured":"Pennycook SJ, Sewall JD, Lee VW (2016) A metric for performance portability. https:\/\/arxiv.org\/abs\/1611.07409"},{"key":"7063_CR18","doi-asserted-by":"publisher","first-page":"947","DOI":"10.1016\/j.future.2017.08.007","volume":"92","author":"SJ Pennycook","year":"2019","unstructured":"Pennycook SJ, Sewall JD, Lee VW (2019) Implications of a metric for performance portability. Future Gener Comput Syst 92:947\u2013958. https:\/\/doi.org\/10.1016\/j.future.2017.08.007","journal-title":"Future Gener Comput Syst"},{"key":"7063_CR19","doi-asserted-by":"crossref","unstructured":"Marowka A (2024) Portability efficiency approach for calculating performance portability . https:\/\/arxiv.org\/abs\/2407.00232","DOI":"10.1016\/j.future.2025.107826"},{"key":"7063_CR20","doi-asserted-by":"publisher","unstructured":"Weckert C, Solis-Vasquez L, Oppermann J, Koch A, Sinnen O (2023) Altis-SYCL: migrating altis benchmarking suite from CUDA to SYCL for GPUs and FPGAs. In: Proc workshop on heterogeneous high-performance reconfigurable computing (H2RC), Held in Conjuction with Int Conf on High Performance Computing, Networking, Storage and Analysis (SC). SC-W \u201923. Association for Computing Machinery, New York, NY, USA, pp 547\u2013555. https:\/\/doi.org\/10.1145\/3624062.3624542","DOI":"10.1145\/3624062.3624542"},{"key":"7063_CR21","doi-asserted-by":"publisher","unstructured":"Hu B, Rossbach CJ (2020) Altis: Modernizing GPGPU benchmarks. In: Proc IEEE Int Symp on Performance Analysis of Systems and Software (ISPASS), pp 1\u201311. https:\/\/doi.org\/10.1109\/ISPASS48437.2020.00011","DOI":"10.1109\/ISPASS48437.2020.00011"},{"key":"7063_CR22","doi-asserted-by":"publisher","unstructured":"Wang Y, Zhou Y, Wang QS, Wang Y, Xu Q, Wang C, Peng B, Zhu Z, Takuya K, Wang D (2021) Developing medical ultrasound beamforming application on GPU and FPGA using oneAPI. In: Proc Int Symp on Parallel and Distributed Processing Workshops (IPDPSW), pp. 360\u2013370 . https:\/\/doi.org\/10.1109\/IPDPSW52791.2021.00064","DOI":"10.1109\/IPDPSW52791.2021.00064"},{"key":"7063_CR23","doi-asserted-by":"publisher","unstructured":"Kenter T, Shambhu A, Faghih-Naini S, Aizinger V (2021) Algorithm-hardware co-design of a discontinuous Galerkin shallow-water model for a dataflow architecture on FPGA. In: Proceedings of the Platform for Advanced Scientific Computing Conference, pp. 1\u201311. ACM, Geneva Switzerland . https:\/\/doi.org\/10.1145\/3468267.3470617","DOI":"10.1145\/3468267.3470617"},{"key":"7063_CR24","doi-asserted-by":"publisher","unstructured":"Wu X, Kenter T, Schade R, K\u00fchne TD, Plessl C (2023) Computing and compressing electron repulsion integrals on FPGAs. In: Proc IEEE Symp on Field-Programmable Custom Computing Machines (FCCM), pp 162\u2013173. https:\/\/doi.org\/10.1109\/FCCM57271.2023.00026","DOI":"10.1109\/FCCM57271.2023.00026"},{"key":"7063_CR25","doi-asserted-by":"publisher","unstructured":"Opdenh\u00f6vel J-O, Plessl C, Kenter T (2023) Mutation tree reconstruction of tumor cells on FPGAs using a bit-level matrix representation. In: Proc Int Symp Highly-Efficient Accelerators and Reconfigurable Technologies (HEART). HEART \u201923, pp. 27\u201334. Association for Computing Machinery, New York, NY, USA. https:\/\/doi.org\/10.1145\/3597031.3597050","DOI":"10.1145\/3597031.3597050"},{"key":"7063_CR26","doi-asserted-by":"publisher","unstructured":"Olgu K, Kenter T, Nunez-Yanez J, Mcintosh-Smith S (2024) Optimisation and evaluation of breadth first search with oneAPI\/SYCL on Intel FPGAs: from describing algorithms to describing architectures. In: Proc Int Workshop on OpenCL and SYCL (IWOCL). IWOCL \u201924. Association for Computing Machinery, New York, NY, USA. https:\/\/doi.org\/10.1145\/3648115.3648134","DOI":"10.1145\/3648115.3648134"},{"key":"7063_CR27","doi-asserted-by":"publisher","unstructured":"Opdenh\u00f6vel J-O, Alt C, Plessl C, Kenter T (2024) Stencilstream: A SYCL-based stencil simulation framework targeting FPGAs. In: Proc Int Conf on Field Programmable Logic and Applications (FPL), pp. 100\u2013108. https:\/\/doi.org\/10.1109\/FPL64840.2024.00023","DOI":"10.1109\/FPL64840.2024.00023"},{"issue":"1","key":"7063_CR28","doi-asserted-by":"publisher","first-page":"37","DOI":"10.4208\/cicp.070114.271114a","volume":"18","author":"R Gandham","year":"2015","unstructured":"Gandham R, Medina D, Warburton T (2015) GPU accelerated discontinuous Galerkin methods for shallow water equations. Commun Comput Phys 18(1):37\u201364. https:\/\/doi.org\/10.4208\/cicp.070114.271114a","journal-title":"Commun Comput Phys"},{"key":"7063_CR29","doi-asserted-by":"publisher","unstructured":"Caviedes-Voulli\u00e8me D, Morales-Hern\u00e1ndez M, Norman MR, \u00d6zgen-Xian I (2023) SERGHEI (SERGHEI-SWE) v1.0: a performance-portable high-performance parallel-computing shallow-water solver for hydrology and environmental hydraulics. Geosci Model Dev 16(3):977\u20131008. https:\/\/doi.org\/10.5194\/gmd-16-977-2023","DOI":"10.5194\/gmd-16-977-2023"},{"issue":"3","key":"7063_CR30","doi-asserted-by":"publisher","first-page":"637","DOI":"10.3390\/w12030637","volume":"12","author":"F Aureli","year":"2020","unstructured":"Aureli F, Prost F, Vacondio R, Dazzi S, Ferrari A (2020) A GPU-accelerated shallow-water scheme for surface runoff simulations. Water 12(3):637. https:\/\/doi.org\/10.3390\/w12030637","journal-title":"Water"},{"key":"7063_CR31","doi-asserted-by":"publisher","DOI":"10.1016\/j.envsoft.2021.105034","volume":"141","author":"M Morales-Hern\u00e1ndez","year":"2021","unstructured":"Morales-Hern\u00e1ndez M, Sharif MB, Kalyanapu A, Ghafoor SK, Dullo TT, Gangrade S, Kao S-C, Norman MR, Evans KJ (2021) TRITON: a multi-GPU open source 2D hydrodynamic flood model. Environ Modell Softw 141:105034. https:\/\/doi.org\/10.1016\/j.envsoft.2021.105034","journal-title":"Environ Modell Softw"},{"key":"7063_CR32","doi-asserted-by":"publisher","DOI":"10.1016\/j.cpc.2020.107251","volume":"254","author":"A Reinarz","year":"2020","unstructured":"Reinarz A, Charrier DE, Bader M, Bovard L, Dumbser M, Duru K, Fambri F, Gabriel A-A, Gallard J-M, K\u00f6ppel S, Krenz L, Rannabauer L, Rezzolla L, Samfass P, Tavelli M, Weinzierl T (2020) ExaHyPE: an engine for parallel dynamically adaptive simulations of wave problems. Comput Phys Commun 254:107251. https:\/\/doi.org\/10.1016\/j.cpc.2020.107251","journal-title":"Comput Phys Commun"},{"key":"7063_CR33","doi-asserted-by":"publisher","unstructured":"Loi CM, Bockhorst H, Weinzierl T (2024) SYCL compute kernels for ExaHyPE, pp. 90\u2013103 . https:\/\/doi.org\/10.1137\/1.9781611977967.8","DOI":"10.1137\/1.9781611977967.8"},{"key":"7063_CR34","doi-asserted-by":"publisher","unstructured":"H\u00e9l\u00e8ne C, Minh-Hoang L, S\u00e9bastien L (2013) Parallelization of shallow-water equations with the algorithmic skeleton library SkelGIS. Procedia Computer Science 18:591\u2013600 https:\/\/doi.org\/10.1016\/j.procs.2013.05.223 . 2013 International Conference on Computational Science","DOI":"10.1016\/j.procs.2013.05.223"},{"key":"7063_CR35","doi-asserted-by":"publisher","unstructured":"Faghih-Naini S, Kuckuk S, Zint D, Kemmler S, K\u00f6stler H, Aizinger V (2023) Discontinuous Galerkin method for the shallow water equations on complex domains using masked block-structured grids. Advances in Water Resources 182, 104584 https:\/\/doi.org\/10.1016\/j.advwatres.2023.104584","DOI":"10.1016\/j.advwatres.2023.104584"},{"key":"7063_CR36","doi-asserted-by":"publisher","unstructured":"Alt C, Kenter T, Faghih-Naini S, Faj J, Opdenh\u00f6vel J-O, Plessl C, Aizinger V, H\u00f6nig J, K\u00f6stler H (2023) Shallow water DG simulations on FPGAs: Design and comparison of a novel code generation pipeline. In: Proc Int Conf on High Performance Computing (ISC High Performance). Lecture Notes in Computer Science (LNCS), pp. 86\u2013105. Springer, Cham . https:\/\/doi.org\/10.1007\/978-3-031-32041-5_5","DOI":"10.1007\/978-3-031-32041-5_5"},{"key":"7063_CR37","doi-asserted-by":"publisher","unstructured":"Faj J, Plessl C, Kenter T, Faghih-Naini S, Aizinger V (2023) Scalable Multi-FPGA Design of a Discontinuous Galerkin Shallow-Water Model on Unstructured Meshes. In: Proc Platform for Advanced Scientific Computing Conf (PASC), pp. 1\u201312 .https:\/\/doi.org\/10.1145\/3592979.3593407","DOI":"10.1145\/3592979.3593407"},{"key":"7063_CR38","doi-asserted-by":"publisher","unstructured":"Kenter T, Mahale G, Alhaddad S, Grynko Y, Schmitt C, Afzal A, Hannig F, F\u00f6rstner J, Plessl C (2018) OpenCL-based FPGA design to accelerate the nodal discontinuous Galerkin method for unstructured meshes. In: Proc IEEE Symp on Field-Programmable Custom Computing Machines (FCCM), pp. 189\u2013196. IEEE, New York, NY, USA . https:\/\/doi.org\/10.1109\/FCCM.2018.00037","DOI":"10.1109\/FCCM.2018.00037"},{"key":"7063_CR39","doi-asserted-by":"publisher","unstructured":"Gourounas D, Hanindhito B, Fathi A, Trenev D, John LK, Gerstlauer A (2023) FAWS: Fpga acceleration of large-scale wave simulations. In: Proc IEEE Int Conf on Application-Specific Systems, Architectures, and Processors (ASAP), pp. 76\u201384 . https:\/\/doi.org\/10.1109\/ASAP57973.2023.00025","DOI":"10.1109\/ASAP57973.2023.00025"},{"issue":"6","key":"7063_CR40","doi-asserted-by":"publisher","first-page":"2396","DOI":"10.1016\/j.jcp.2011.11.018","volume":"231","author":"PD D\u00fcben","year":"2012","unstructured":"D\u00fcben PD, Korn P, Aizinger V (2012) A discontinuous\/continuous low order finite element shallow water model on the sphere. J Comput Phys 231(6):2396\u20132413. https:\/\/doi.org\/10.1016\/j.jcp.2011.11.018","journal-title":"J Comput Phys"},{"issue":"1","key":"7063_CR41","doi-asserted-by":"publisher","first-page":"67","DOI":"10.1016\/S0309-1708(01)00019-7","volume":"25","author":"V Aizinger","year":"2002","unstructured":"Aizinger V, Dawson C (2002) A discontinuous Galerkin method for two-dimensional flow and transport in shallow water. Adv Water Resour 25(1):67\u201384. https:\/\/doi.org\/10.1016\/S0309-1708(01)00019-7","journal-title":"Adv Water Resour"},{"key":"7063_CR42","doi-asserted-by":"publisher","first-page":"185","DOI":"10.1016\/j.envsoft.2018.01.003","volume":"102","author":"H Hajduk","year":"2018","unstructured":"Hajduk H, Hodges BR, Aizinger V, Reuter B (2018) Locally filtered transport for computational efficiency in multi-component advection-reaction models. Environ Model Softw 102:185\u2013198. https:\/\/doi.org\/10.1016\/j.envsoft.2018.01.003","journal-title":"Environ Model Softw"},{"key":"7063_CR43","doi-asserted-by":"publisher","first-page":"411","DOI":"10.1090\/S0025-5718-1989-0983311-4","volume":"52","author":"B Cockburn","year":"1989","unstructured":"Cockburn B, Shu C-W (1989) TVB Runge-Kutta local projection discontinuous Galerkin finite element method for conservation laws. II. General framework Math Comp 52:411\u2013435. https:\/\/doi.org\/10.1090\/S0025-5718-1989-0983311-4","journal-title":"General framework Math Comp"},{"issue":"1","key":"7063_CR44","doi-asserted-by":"publisher","first-page":"105","DOI":"10.1016\/S0045-7825(97)00108-4","volume":"151","author":"S Chippada","year":"1998","unstructured":"Chippada S, Dawson CN, Martinez ML, Wheeler MF (1998) A Godunov-type finite volume method for the system of shallow water equations. Comput Methods Appl Mech Eng 151(1):105\u2013129. https:\/\/doi.org\/10.1016\/S0045-7825(97)00108-4","journal-title":"Comput Methods Appl Mech Eng"},{"key":"7063_CR45","doi-asserted-by":"publisher","DOI":"10.1016\/j.advwatres.2020.103552","volume":"138","author":"S Faghih-Naini","year":"2020","unstructured":"Faghih-Naini S, Kuckuk S, Aizinger V, Zint D, Grosso R, K\u00f6stler H (2020) Quadrature-free discontinuous Galerkin method with code generation features for shallow water equations on automatically generated block-structured meshes. Adv Water Resour 138:103552","journal-title":"Adv Water Resour"},{"key":"7063_CR46","doi-asserted-by":"crossref","unstructured":"Westerink JJ, Stolzenbach KD, Connor JJ (1989) General spectral computations of the nonlinear shallow water tidal interactions within the bight of Abaco. J Phys Oceanogr 19(9):1348\u20131371","DOI":"10.1175\/1520-0485(1989)019<1348:GSCOTN>2.0.CO;2"},{"key":"7063_CR47","doi-asserted-by":"publisher","unstructured":"Sakiotis I, Arumugam K, Paterno M, Ranjan D, Terzi\u0107 B, Zubair M (2023) Porting Numerical Integration Codes from CUDA to oneAPI: a case study. In: Bhatele, A., Hammond, J., Baboulin, M., Kruse, C. (eds.) High Performance Computing, pp. 339\u2013358. Springer, Cham . https:\/\/doi.org\/10.1007\/978-3-031-32041-5_18","DOI":"10.1007\/978-3-031-32041-5_18"},{"key":"7063_CR48","doi-asserted-by":"publisher","DOI":"10.17815\/jlsrf-8-187","author":"C Bauer","year":"2024","unstructured":"Bauer C, Kenter T, Lass M, Mazur L, Meyer M, Nitsche H, Riebler H, Schade R, Schwarz M, Winnwa N, Wiens A, Wu X, Plessl C, Simon J (2024) Noctua 2 supercomputer. J Large-scale Res Facil (JLSRF). https:\/\/doi.org\/10.17815\/jlsrf-8-187","journal-title":"J Large-scale Res Facil (JLSRF)"},{"key":"7063_CR49","doi-asserted-by":"publisher","first-page":"79","DOI":"10.1016\/j.jpdc.2021.10.007","volume":"160","author":"M Meyer","year":"2022","unstructured":"Meyer M, Kenter T, Plessl C (2022) In-depth FPGA accelerator performance evaluation with single node benchmarks from the HPC challenge benchmark suite for Intel and Xilinx FPGAs using OpenCL. J Parallel Distrib Comput 160:79\u201389. https:\/\/doi.org\/10.1016\/j.jpdc.2021.10.007","journal-title":"J Parallel Distrib Comput"},{"key":"7063_CR50","doi-asserted-by":"publisher","unstructured":"Gruber, T., Eitzinger, J., Hager, G., Wellein, G.: LIKWID. https:\/\/doi.org\/10.5281\/zenodo.10105559","DOI":"10.5281\/zenodo.10105559"},{"key":"7063_CR51","doi-asserted-by":"publisher","unstructured":"Sch\u00f6ne R, Ilsche T, Bielert M, Velten M, Schmidl M, Hackenberg D (2021) Energy efficiency aspects of the AMD Zen 2 architecture. In: 2021 IEEE International Conference on Cluster Computing (CLUSTER), pp. 562\u2013571. IEEE, Portland, OR, USA . https:\/\/doi.org\/10.1109\/Cluster48925.2021.00087","DOI":"10.1109\/Cluster48925.2021.00087"},{"key":"7063_CR52","unstructured":"NVIDIA: NVIDIA Grace Performance Tuning Guide. (2024). https:\/\/docs.nvidia.com\/grace-performance-tuning-guide.pdf Accessed 2024-06-19"},{"key":"7063_CR53","doi-asserted-by":"publisher","unstructured":"Corda, S., Veenboer, B., Tolley, E.: PMT: power measurement toolkit. In: 2022 IEEE\/ACM International Workshop on HPC User Support Tools (HUST), pp. 44\u201347 (2022). https:\/\/doi.org\/10.1109\/HUST56722.2022.00011","DOI":"10.1109\/HUST56722.2022.00011"},{"key":"7063_CR54","doi-asserted-by":"publisher","unstructured":"Alhaddad, S., F\u00f6rstner, J., Groth, S., Gr\u00fcnewald, D., Grynko, Y., Hannig, F., Kenter, T., Pfreundt, F.-J., Plessl, C., Schotte, M., Steinke, T., Teich, J., Weiser, M., Wende, F.: The HighPerMeshes framework for numerical algorithms on unstructured grids. Concurrency and Computation: Practice and Experience 34, 1\u201315 (2021) https:\/\/doi.org\/10.1002\/cpe.6616","DOI":"10.1002\/cpe.6616"}],"container-title":["The Journal of Supercomputing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11227-025-07063-7.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11227-025-07063-7\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11227-025-07063-7.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,4,23]],"date-time":"2025-04-23T12:24:38Z","timestamp":1745411078000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11227-025-07063-7"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,4,23]]},"references-count":54,"journal-issue":{"issue":"6","published-online":{"date-parts":[[2025,4]]}},"alternative-id":["7063"],"URL":"https:\/\/doi.org\/10.1007\/s11227-025-07063-7","relation":{},"ISSN":["1573-0484"],"issn-type":[{"value":"1573-0484","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,4,23]]},"assertion":[{"value":"12 February 2025","order":1,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"23 April 2025","order":2,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare no competing interests.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"772"}}