{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,18]],"date-time":"2026-06-18T11:35:55Z","timestamp":1781782555882,"version":"3.54.5"},"reference-count":47,"publisher":"SAGE Publications","issue":"5","license":[{"start":{"date-parts":[[2019,2,4]],"date-time":"2019-02-04T00:00:00Z","timestamp":1549238400000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/journals.sagepub.com\/page\/policies\/text-and-data-mining-license"}],"funder":[{"DOI":"10.13039\/100008393","name":"Teknologi og Produktion, Det Frie Forskningsr\u00e5d","doi-asserted-by":"publisher","award":["09-070032"],"award-info":[{"award-number":["09-070032"]}],"id":[{"id":"10.13039\/100008393","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["journals.sagepub.com"],"crossmark-restriction":true},"short-container-title":["The International Journal of High Performance Computing Applications"],"published-print":{"date-parts":[[2019,9]]},"abstract":"<jats:p>The focus of this article is on the parallel scalability of a distributed multigrid framework, known as the DTU Compute GPUlab Library, for execution on graphics processing unit (GPU)-accelerated supercomputers. We demonstrate near-ideal weak scalability for a high-order fully nonlinear potential flow (FNPF) time domain model on the Oak Ridge Titan supercomputer, which is equipped with a large number of many-core CPU-GPU nodes. The high-order finite difference scheme for the solver is implemented to expose data locality and scalability, and the linear Laplace solver is based on an iterative multilevel preconditioned defect correction method designed for high-throughput processing and massive parallelism. In this work, the FNPF discretization is based on a multi-block discretization that allows for large-scale simulations. In this setup, each grid block is based on a logically structured mesh with support for curvilinear representation of horizontal block boundaries to allow for an accurate representation of geometric features such as surface-piercing bottom-mounted structures\u2014for example, mono-pile foundations as demonstrated. Unprecedented performance and scalability results are presented for a system of equations that is historically known as being too expensive to solve in practical applications. A novel feature of the potential flow model is demonstrated, being that a modest number of multigrid restrictions is sufficient for fast convergence, improving overall parallel scalability as the coarse grid problem diminishes. In the numerical benchmarks presented, we demonstrate using 8192 modern Nvidia GPUs enabling large-scale and high-resolution nonlinear marine hydrodynamics applications.<\/jats:p>","DOI":"10.1177\/1094342019826662","type":"journal-article","created":{"date-parts":[[2019,2,4]],"date-time":"2019-02-04T22:51:03Z","timestamp":1549320663000},"page":"855-868","update-policy":"https:\/\/doi.org\/10.1177\/sage-journals-update-policy","source":"Crossref","is-referenced-by-count":7,"title":["A massively scalable distributed multigrid framework for nonlinear marine hydrodynamics"],"prefix":"10.1177","volume":"33","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-4893-2328","authenticated-orcid":false,"given":"Stefan Lemvig","family":"Glimberg","sequence":"first","affiliation":[{"name":"Department of Applied Mathematics and Computer Science, Technical University of Denmark, Kongens Lyngby, Denmark"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Allan Peter","family":"Engsig-Karup","sequence":"additional","affiliation":[{"name":"Department of Applied Mathematics and Computer Science, Technical University of Denmark, Kongens Lyngby, Denmark"},{"name":"Center for Energy Resources Engineering, Technical University of Denmark, Kongens Lyngby, Denmark"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Luke N","family":"Olson","sequence":"additional","affiliation":[{"name":"Department of Computer Science, University of Illinois at Urbana\u2013Champaign, Urbana, IL, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"179","published-online":{"date-parts":[[2019,2,4]]},"reference":[{"key":"bibr1-1094342019826662","unstructured":"Acklam E, Langtangen HP (1998) Parallelization of explicit finite difference schemes via domain decomposition. Report for Oslo Scientific Computing Archive. Report no. 1998-2, February."},{"key":"bibr2-1094342019826662","unstructured":"Asanovic K, Bodik R, Catanzaro BC, et al. (2006) The landscape of parallel computing research: a view from Berkeley. Report, University of California, US, Dec, UCB\/EECS-2006-183."},{"key":"bibr3-1094342019826662","first-page":"359","volume-title":"GPU Computing Gems Jade Edition","author":"Bell N","year":"2011"},{"key":"bibr4-1094342019826662","doi-asserted-by":"publisher","DOI":"10.1137\/110838844"},{"key":"bibr5-1094342019826662","doi-asserted-by":"publisher","DOI":"10.1007\/s10665-016-9848-8"},{"key":"bibr6-1094342019826662","doi-asserted-by":"publisher","DOI":"10.1007\/s10665-006-9108-4"},{"key":"bibr7-1094342019826662","doi-asserted-by":"publisher","DOI":"10.1016\/j.advwatres.2004.11.004"},{"key":"bibr8-1094342019826662","doi-asserted-by":"publisher","DOI":"10.1016\/j.fluiddyn.2005.08.007"},{"key":"bibr9-1094342019826662","doi-asserted-by":"publisher","DOI":"10.1016\/S1001-6058(09)60198-0"},{"key":"bibr10-1094342019826662","doi-asserted-by":"crossref","unstructured":"Engsig-Karup AP (2006) Unstructured nodal DG-FEM solution of high-order boussinesq-type equations. PhD Thesis, Technical University of Denmark, DK.","DOI":"10.1007\/s10665-006-9064-z"},{"key":"bibr11-1094342019826662","doi-asserted-by":"publisher","DOI":"10.1002\/fld.3873"},{"key":"bibr12-1094342019826662","doi-asserted-by":"publisher","DOI":"10.1016\/j.jcp.2008.11.028"},{"key":"bibr13-1094342019826662","first-page":"251","volume-title":"Designing Scientific Applications on GPUs","author":"Engsig-Karup AP","year":"2013"},{"key":"bibr14-1094342019826662","doi-asserted-by":"publisher","DOI":"10.1002\/fld.2675"},{"key":"bibr15-1094342019826662","doi-asserted-by":"publisher","DOI":"10.3997\/2214-4609.20141771"},{"key":"bibr16-1094342019826662","first-page":"47","volume-title":"Proceedings of the 2004 ACM\/IEEE conference on Supercomputing","author":"Fan Z","year":"2004"},{"key":"bibr17-1094342019826662","doi-asserted-by":"publisher","DOI":"10.1016\/S1001-6058(11)60239-4"},{"key":"bibr18-1094342019826662","unstructured":"Glimberg SL (2013) Designing Scientific Software for Heterogenous Computing - With application to large-scale water wave simulations. PhD Thesis, Technical University of Denmark, DK."},{"key":"bibr19-1094342019826662","first-page":"645","volume-title":"9th European Conference on Numerical Mathematics and Advanced Applications, Numerical Mathematics and Advanced Applications","author":"Glimberg SL","year":"2011"},{"key":"bibr20-1094342019826662","first-page":"73","volume-title":"Designing Scientific Applications on GPUs","author":"Glimberg SL","year":"2013"},{"key":"bibr21-1094342019826662","unstructured":"Goddeke D (2011) Fast and accurate finite-element multigrid solvers for PDE simulations on GPU clusters. PhD Thesis, Technischen Universit\u00e3t Dortmund, DE."},{"key":"bibr22-1094342019826662","doi-asserted-by":"publisher","DOI":"10.1016\/j.parco.2007.09.002"},{"key":"bibr23-1094342019826662","doi-asserted-by":"publisher","DOI":"10.1016\/S0955-7997(01)00113-8"},{"key":"bibr24-1094342019826662","volume-title":"Using MPI: Portable Parallel Programming with the Message Passing Interface","author":"Gropp W","year":"1999","edition":"2"},{"key":"bibr25-1094342019826662","first-page":"522","volume-title":"Proceedings of the 27th International Towing Tank Conference (ITTC)","author":"Hino T","year":"2014"},{"key":"bibr26-1094342019826662","doi-asserted-by":"publisher","DOI":"10.1145\/2304576.2304619"},{"key":"bibr27-1094342019826662","unstructured":"Kazolea M, Ricchiuto M (2016) Wave breaking for Boussinesq-type models using a turbulence kinetic energy model. Report RR-8781, Inria Bordeaux Sud-Ouest, FR."},{"key":"bibr28-1094342019826662","doi-asserted-by":"publisher","DOI":"10.1016\/0029-8018(90)90029-6"},{"key":"bibr29-1094342019826662","doi-asserted-by":"publisher","DOI":"10.1016\/0029-8018(92)90048-9"},{"key":"bibr30-1094342019826662","doi-asserted-by":"publisher","DOI":"10.1016\/S0378-3839(96)00046-4"},{"key":"bibr31-1094342019826662","doi-asserted-by":"publisher","DOI":"10.1061\/(ASCE)0733-950X(2001)127:1(16)"},{"key":"bibr32-1094342019826662","doi-asserted-by":"publisher","DOI":"10.1061\/(ASCE)0733-950X(2001)127:3(152)"},{"key":"bibr33-1094342019826662","doi-asserted-by":"crossref","DOI":"10.1201\/9781482265910","volume-title":"Numerical Modeling of Water Waves","author":"Lin P","year":"2008"},{"key":"bibr34-1094342019826662","first-page":"125","volume-title":"Proceedings of The 28th International Workshop on Water Waves and Floating Bodies","author":"Lindberg O","year":"2013"},{"key":"bibr35-1094342019826662","unstructured":"MacCamy RC, Fuchs RA (1954) Wave forces on piles: A Diffraction Theory. Report, Technical memorandum no. 69, Beach Erosion Board, December."},{"key":"bibr36-1094342019826662","doi-asserted-by":"publisher","DOI":"10.1145\/1513895.1513905"},{"key":"bibr37-1094342019826662","doi-asserted-by":"publisher","DOI":"10.1137\/140980260"},{"key":"bibr38-1094342019826662","doi-asserted-by":"publisher","DOI":"10.1115\/1.4007597"},{"key":"bibr39-1094342019826662","doi-asserted-by":"publisher","DOI":"10.1002\/fld.1650210803"},{"key":"bibr40-1094342019826662","doi-asserted-by":"publisher","DOI":"10.1016\/S0378-3839(00)00067-3"},{"key":"bibr41-1094342019826662","volume-title":"Domain decomposition: Parallel multilevel methods for elliptic partial differential equations","author":"Smith BF","year":"1996"},{"key":"bibr42-1094342019826662","doi-asserted-by":"publisher","DOI":"10.1016\/j.proeng.2013.08.031"},{"key":"bibr43-1094342019826662","volume-title":"Hydrodynamics of coastal regions. Vol. 3","author":"Svendsen IA","year":"1976"},{"key":"bibr44-1094342019826662","volume-title":"Multigrid","author":"Trottenberg U","year":"2000","edition":"1"},{"key":"bibr45-1094342019826662","doi-asserted-by":"publisher","DOI":"10.1201\/9781315229256-9"},{"key":"bibr46-1094342019826662","doi-asserted-by":"publisher","DOI":"10.1146\/annurev.fl.14.010182.002143"},{"key":"bibr47-1094342019826662","doi-asserted-by":"publisher","DOI":"10.1016\/j.coastaleng.2005.02.004"}],"container-title":["The International Journal of High Performance Computing Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/1094342019826662","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/full-xml\/10.1177\/1094342019826662","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/1094342019826662","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,29]],"date-time":"2026-04-29T08:15:46Z","timestamp":1777450546000},"score":1,"resource":{"primary":{"URL":"https:\/\/journals.sagepub.com\/doi\/10.1177\/1094342019826662"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019,2,4]]},"references-count":47,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2019,9]]}},"alternative-id":["10.1177\/1094342019826662"],"URL":"https:\/\/doi.org\/10.1177\/1094342019826662","relation":{},"ISSN":["1094-3420","1741-2846"],"issn-type":[{"value":"1094-3420","type":"print"},{"value":"1741-2846","type":"electronic"}],"subject":[],"published":{"date-parts":[[2019,2,4]]}}}