{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,16]],"date-time":"2026-04-16T07:16:21Z","timestamp":1776323781571,"version":"3.50.1"},"reference-count":39,"publisher":"Springer Science and Business Media LLC","issue":"4","license":[{"start":{"date-parts":[[2024,6,21]],"date-time":"2024-06-21T00:00:00Z","timestamp":1718928000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2024,6,21]],"date-time":"2024-06-21T00:00:00Z","timestamp":1718928000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100004869","name":"Universit\u00e4t M\u00fcnster","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100004869","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Int J Parallel Prog"],"published-print":{"date-parts":[[2024,8]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>Complex algorithms and enormous data sets require parallel execution of programs to attain results in a reasonable amount of time. Both aspects are combined in the domain of three-dimensional stencil operations, for example, computational fluid dynamics. This work contributes to the research on high-level parallel programming by discussing the generalizable implementation of a three-dimensional stencil skeleton that works in heterogeneous computing environments. Two exemplary programs, a gas simulation with the Lattice Boltzmann method, and a mean blur, are executed in a multi-node multi-graphics processing units environment, proving the runtime improvements in heterogeneous computing environments compared to a sequential program.<\/jats:p>","DOI":"10.1007\/s10766-024-00769-w","type":"journal-article","created":{"date-parts":[[2024,6,21]],"date-time":"2024-06-21T17:01:12Z","timestamp":1718989272000},"page":"274-297","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":1,"title":["Optimizing Three-Dimensional Stencil-Operations on Heterogeneous Computing Environments"],"prefix":"10.1007","volume":"52","author":[{"given":"Nina","family":"Herrmann","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Justus","family":"Dieckmann","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Herbert","family":"Kuchen","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2024,6,21]]},"reference":[{"key":"769_CR1","unstructured":"MPI Standard: https:\/\/www.mpi-forum.org\/docs\/. Accessed 24 Feb 2023"},{"key":"769_CR2","unstructured":"The OpenMP API specification for parallel programming. https:\/\/www.openmp.org\/. Accessed 24 Feb 2023"},{"key":"769_CR3","unstructured":"NVIDIA: CUDA: https:\/\/developer.nvidia.com\/cuda-zone. Accessed 24 Feb 2023"},{"key":"769_CR4","unstructured":"Cole, M.I.: Algorithmic skeletons: structured management of parallel computation. Computer science thesis. Pitman, London (1989)"},{"issue":"2","key":"769_CR5","doi-asserted-by":"publisher","first-page":"283","DOI":"10.1007\/s10766-016-0416-7","volume":"45","author":"S Ernsting","year":"2017","unstructured":"Ernsting, S., Kuchen, H.: Data parallel algorithmic skeletons with accelerator support. Int. J. Parallel Prog. 45(2), 283\u2013299 (2017)","journal-title":"Int. J. Parallel Prog."},{"key":"769_CR6","doi-asserted-by":"publisher","first-page":"761","DOI":"10.1007\/11549468_83","volume-title":"Euro-Par 2005 Parallel Processing","author":"A Benoit","year":"2005","unstructured":"Benoit, A., Cole, M., Gilmore, S., Hillston, J.: Flexible skeletal programming with eskel. In: Cunha, J.C., Medeiros, P.D. (eds.) Euro-Par 2005 Parallel Processing, pp. 761\u2013770. Springer, Berlin, Heidelberg (2005)"},{"issue":"6","key":"769_CR7","doi-asserted-by":"publisher","first-page":"846","DOI":"10.1007\/s10766-021-00704-3","volume":"49","author":"A Ernstsson","year":"2021","unstructured":"Ernstsson, A., Ahlqvist, J., Zouzoula, S., Kessler, C.: Skepu 3: portable high-level programming of heterogeneous systems and hpc clusters. Int. J. Parallel Prog. 49(6), 846\u2013866 (2021)","journal-title":"Int. J. Parallel Prog."},{"key":"769_CR8","doi-asserted-by":"publisher","first-page":"261","DOI":"10.1002\/9781119332015.ch13","volume-title":"Programming multi-core and many-core computing systems, parallel and distributed computing","author":"M Aldinucci","year":"2017","unstructured":"Aldinucci, M., Danelutto, M., Kilpatrick, P., Torquati, M.: Fastflow: high-level and efficient streaming on multi-core. In: Pllana, S., Xhafa, F. (eds.) Programming multi-core and many-core computing systems, parallel and distributed computing, pp. 261\u2013280. Wiley, London (2017)"},{"issue":"7","key":"769_CR9","doi-asserted-by":"publisher","first-page":"5098","DOI":"10.1007\/s11227-019-02825-6","volume":"76","author":"F Wrede","year":"2020","unstructured":"Wrede, F., Rieger, C., Kuchen, H.: Generation of high-performance code based on a domain-specific language for algorithmic skeletons. J. Supercomput. 76(7), 5098\u20135116 (2020)","journal-title":"J. Supercomput."},{"issue":"3\u20134","key":"769_CR10","doi-asserted-by":"publisher","first-page":"341","DOI":"10.1007\/s10766-022-00731-8","volume":"50","author":"P Thoman","year":"2022","unstructured":"Thoman, P., Tischler, F., Salzmann, P., Fahringer, T.: The celerity high-level api: C++ 20 for accelerator clusters. Int. J. Parallel Prog. 50(3\u20134), 341\u2013359 (2022)","journal-title":"Int. J. Parallel Prog."},{"key":"769_CR11","doi-asserted-by":"crossref","unstructured":"Goli, M., Gonz\u00e1lez-V\u00e9lez, H.: Heterogeneous algorithmic skeletons for fast flow with seamless coordination over hybrid architectures. In: 2013 21st euromicro international conference on parallel, distributed, and network-based processing, pp. 148\u2013156 (2013)","DOI":"10.1109\/PDP.2013.29"},{"key":"769_CR12","doi-asserted-by":"crossref","unstructured":"Hagedorn, B., Stoltzfus, L., Steuwer, M., Gorlatch, S., Dubach, C.: High performance stencil code generation with lift. In: Proceedings of the 2018 international symposium on code generation and optimization. CGO 2018, pp. 100\u2013112. Association for Computing Machinery, New York, NY (2018)","DOI":"10.1145\/3179541.3168824"},{"issue":"1","key":"769_CR13","doi-asserted-by":"publisher","first-page":"25","DOI":"10.1007\/s11227-014-1213-y","volume":"69","author":"M Steuwer","year":"2014","unstructured":"Steuwer, M., Gorlatch, S.: Skelcl: a high-level extension of opencl for multi-gpu systems. J. Supercomput. 69(1), 25\u201333 (2014)","journal-title":"J. Supercomput."},{"key":"769_CR14","doi-asserted-by":"publisher","first-page":"478","DOI":"10.1016\/j.camwa.2020.01.007","volume":"81","author":"M Bauer","year":"2021","unstructured":"Bauer, M., Eibl, S., Godenschwager, C., Kohl, N., Kuron, M., Rettinger, C., Schornbaum, F., Schwarzmeier, C., Th\u00f6nnes, D., K\u00f6stler, H., R\u00fcde, U.: walberla: A block-structured high-performance framework for multiphysics simulations. Comput. Math. Appl. 81, 478\u2013501 (2021)","journal-title":"Comput. Math. Appl."},{"issue":"9","key":"769_CR15","doi-asserted-by":"publisher","first-page":"9409","DOI":"10.1007\/s11227-022-05040-y","volume":"79","author":"M Castro","year":"2023","unstructured":"Castro, M., Santamaria-Valenzuela, I., Torres, Y., Gonzalez-Escribano, A., Llanos, D.R.: Epsilod: efficient parallel skeleton for generic iterative stencil computations in distributed gpus. J. Supercomput. 79(9), 9409\u20139442 (2023)","journal-title":"J. Supercomput."},{"key":"769_CR16","doi-asserted-by":"crossref","unstructured":"Kuckuk, S., K\u00f6stler, H.: Whole program generation of massively parallel shallow water equation solvers. In: 2018 IEEE international conference on cluster computing (CLUSTER), pp. 78\u201387 (2018)","DOI":"10.1109\/CLUSTER.2018.00020"},{"key":"769_CR17","doi-asserted-by":"publisher","first-page":"75","DOI":"10.1016\/j.camwa.2020.06.007","volume":"81","author":"P Bastian","year":"2021","unstructured":"Bastian, P., Blatt, M., Dedner, A., Dreier, N.-A., Engwer, C., Fritze, R., Gr\u00e4ser, C., Gr\u00fcninger, C., Kempf, D., Kl\u00f6fkorn, R., Ohlberger, M., Sander, O.: The dune framework: basic concepts and recent developments. Comput. Math. Appl. 81, 75\u2013112 (2021)","journal-title":"Comput. Math. Appl."},{"issue":"4","key":"769_CR18","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/2400682.2400718","volume":"9","author":"T Lutz","year":"2013","unstructured":"Lutz, T., Fensch, C., Cole, M.: PARTANS: an autotuning framework for stencil computation on multi-GPU systems. ACM Trans. Archit. Code Optim. 9(4), 1\u201324 (2013)","journal-title":"ACM Trans. Archit. Code Optim."},{"issue":"17","key":"769_CR19","doi-asserted-by":"publisher","first-page":"4938","DOI":"10.1002\/cpe.3479","volume":"27","author":"AD Pereira","year":"2015","unstructured":"Pereira, A.D., Ramos, L., G\u00f3es, L.F.W.: Pskel: a stencil programming framework for cpu-gpu systems. Concurr. Comput. Pract. Exp. 27(17), 4938\u20134953 (2015)","journal-title":"Concurr. Comput. Pract. Exp."},{"key":"769_CR20","doi-asserted-by":"crossref","unstructured":"Pereira, A.D., Castro, M., Dantas, M.A., Rocha, R.C., G\u00f3es, L.F.: Extending openacc for efficient stencil code generation and execution by skeleton frameworks. In: 2017 international conference on high performance computing & simulation (HPCS), pp. 719\u2013726. IEEE (2017)","DOI":"10.1109\/HPCS.2017.110"},{"key":"769_CR21","doi-asserted-by":"publisher","first-page":"334","DOI":"10.1016\/j.camwa.2020.03.022","volume":"81","author":"J Latt","year":"2021","unstructured":"Latt, J., Malaspinas, O., Kontaxakis, D., Parmigiani, A., Lagrava, D., Brogi, F., Belgacem, M.B., Thorimbert, Y., Leclaire, S., Li, S., Marson, F., Lemus, J., Kotsalos, C., Conradin, R., Coreixas, C., Petkantchin, R., Raynaud, F., Beny, J., Chopard, B.: Palabos: parallel lattice Boltzmann solver. Comput. Math. Appl. 81, 334\u2013350 (2021)","journal-title":"Comput. Math. Appl."},{"key":"769_CR22","doi-asserted-by":"publisher","DOI":"10.1016\/j.parco.2021.102757","volume":"103","author":"R Gonzales","year":"2021","unstructured":"Gonzales, R., Gryazin, Y., Lee, Y.T.: Parallel fft algorithms for high-order approximations on three-dimensional compact stencils. Parallel Comput. 103, 102757 (2021)","journal-title":"Parallel Comput."},{"key":"769_CR23","unstructured":"Skepu: https:\/\/github.com\/skepu\/skepu\/. Accessed 13 Feb 2024"},{"key":"769_CR24","unstructured":"Celerity: https:\/\/github.com\/celerity\/celerity-runtime. Accessed 13 Feb 2024"},{"key":"769_CR25","unstructured":"Lift: https:\/\/github.com\/lift-project\/lift\/tree\/master. Accessed 13 Feb 2024"},{"key":"769_CR26","unstructured":"SkelCL: https:\/\/github.com\/skelcl\/skelcl. Accessed 13 Feb 2024"},{"key":"769_CR27","doi-asserted-by":"publisher","DOI":"10.1016\/j.jocs.2020.101269","volume":"49","author":"M Bauer","year":"2021","unstructured":"Bauer, M., K\u00f6stler, H., R\u00fcde, U.: lbmpy: automatic code generation for efficient parallel lattice boltzmann methods. J. Comput. Sci. 49, 101269 (2021)","journal-title":"J. Comput. Sci."},{"key":"769_CR28","unstructured":"WaLBerla: https:\/\/i10git.cs.fau.de\/walberla\/walberla. Accessed 13 Feb 2024"},{"key":"769_CR29","unstructured":"Lbmpy: https:\/\/i10git.cs.fau.de\/pycodegen\/lbmpy. Accessed 13 Feb 2024"},{"key":"769_CR30","unstructured":"EPSILOD: https:\/\/gitlab.com\/trasgo-group-valladolid\/controllers\/-\/tree\/epsilod_JoS22. Accessed 13 Feb 2024"},{"key":"769_CR31","unstructured":"ExaStencil: https:\/\/github.com\/lssfau\/ExaStencils. Accessed 13 Feb 2024"},{"key":"769_CR32","unstructured":"DUNE: https:\/\/gitlab.dune-project.org\/core. Accessed 13 Feb 2024"},{"key":"769_CR33","unstructured":"PSkel: https:\/\/github.com\/pskel\/pskel. Accessed 13 Feb 2024"},{"key":"769_CR34","unstructured":"Palabos: https:\/\/gitlab.com\/unigespc\/palabos. Accessed 13 Feb 2024"},{"key":"769_CR35","doi-asserted-by":"crossref","unstructured":"Kotsalos, C., Latt, J., Chopard, B.: Palabos-npfem: software for the simulation of cellular blood flow (digital blood). arXiv preprint arXiv:2011.04332 (2020)","DOI":"10.5334\/jors.343"},{"key":"769_CR36","doi-asserted-by":"publisher","first-page":"874","DOI":"10.1007\/978-3-642-40047-6_86","volume-title":"Euro-Par 2013 parallel processing","author":"R Marques","year":"2013","unstructured":"Marques, R., Paulino, H., Alexandre, F., Medeiros, P.D.: Algorithmic skeleton framework for the orchestration of gpu computations. In: Wolf, F., Mohr, B., Mey, D. (eds.) Euro-Par 2013 parallel processing, pp. 874\u2013885. Springer, Berlin, Heidelberg (2013)"},{"issue":"1","key":"769_CR37","doi-asserted-by":"publisher","first-page":"74","DOI":"10.1504\/IJICA.2007.013403","volume":"1","author":"E Alba","year":"2007","unstructured":"Alba, E., Luque, G., Garcia-Nieto, J., Ordonez, G., Leguizamon, G.: Mallba: a software library to design efficient optimisation algorithms. Int. J. Innovative Comput. Appl. 1(1), 74\u201385 (2007)","journal-title":"Int. J. Innovative Comput. Appl."},{"key":"769_CR38","doi-asserted-by":"crossref","unstructured":"Karasawa, Y., Iwasaki, H.: A parallel skeleton library for multi-core clusters. In: 2009 international conference on parallel processing, pp. 84\u201391 (2009)","DOI":"10.1109\/ICPP.2009.18"},{"key":"769_CR39","volume-title":"The lattice Boltzmann method: Principles and Practice. Graduate texts in physics","author":"T Kr\u00fcger","year":"2016","unstructured":"Kr\u00fcger, T., Kusumaatmaja, H., Kuzmin, A., Shardt, O., Silva, G., Viggen, E.M.: The lattice Boltzmann method: Principles and Practice. Graduate texts in physics. Springer, Cham (2016)"}],"container-title":["International Journal of Parallel Programming"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10766-024-00769-w.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10766-024-00769-w\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10766-024-00769-w.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,7,20]],"date-time":"2024-07-20T10:04:42Z","timestamp":1721469882000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10766-024-00769-w"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,6,21]]},"references-count":39,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2024,8]]}},"alternative-id":["769"],"URL":"https:\/\/doi.org\/10.1007\/s10766-024-00769-w","relation":{"has-preprint":[{"id-type":"doi","id":"10.21203\/rs.3.rs-3345798\/v1","asserted-by":"object"}]},"ISSN":["0885-7458","1573-7640"],"issn-type":[{"value":"0885-7458","type":"print"},{"value":"1573-7640","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,6,21]]},"assertion":[{"value":"11 September 2023","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"1 May 2024","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"21 June 2024","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare that they have no Conflict of interest.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}}]}}