{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,16]],"date-time":"2026-04-16T07:16:30Z","timestamp":1776323790109,"version":"3.50.1"},"reference-count":52,"publisher":"Springer Science and Business Media LLC","issue":"9","license":[{"start":{"date-parts":[[2023,1,14]],"date-time":"2023-01-14T00:00:00Z","timestamp":1673654400000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2023,1,14]],"date-time":"2023-01-14T00:00:00Z","timestamp":1673654400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"name":"Ministerio de Econom\u00eda, Industria y Competitividad of Spain, European Regional Development Fund (ERDF) program","award":["TIN2017-88614-R"],"award-info":[{"award-number":["TIN2017-88614-R"]}]},{"name":"Conserjer\u00eda de Educaci\u00f3n, Junta de Castilla y Le\u00f3n, Spain","award":["VA226P20"],"award-info":[{"award-number":["VA226P20"]}]},{"name":"Red Espa\u00f1ola de Supercomputaci\u00f3n, Spain","award":["RES-IM-2022-1-0014"],"award-info":[{"award-number":["RES-IM-2022-1-0014"]}]},{"DOI":"10.13039\/501100007515","name":"Universidad de Valladolid","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100007515","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J Supercomput"],"published-print":{"date-parts":[[2023,6]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Iterative stencil computations are widely used in numerical simulations. They present a high degree of parallelism, high locality and mostly-coalesced memory access patterns. Therefore, GPUs are good candidates to speed up their computation. However, the development of stencil programs that can work with huge grids in distributed systems with multiple GPUs is not straightforward, since it requires solving problems related to the partition of the grid across nodes and devices, and the synchronization and data movement across remote GPUs. In this work, we present EPSILOD, a high-productivity parallel programming skeleton for iterative stencil computations on distributed multi-GPUs, of the same or different vendors that supports any type of n-dimensional geometric stencils of any order. It uses an abstract specification of the stencil pattern (neighbors and weights) to internally derive the data partition, synchronizations and communications. Computation is split to better overlap with communications. This paper describes the underlying architecture of EPSILOD, its main components, and presents an experimental evaluation to show the benefits of our approach, including a comparison with another state-of-the-art solution. The experimental results show that EPSILOD is faster and shows good strong and weak scalability for platforms with both homogeneous and heterogeneous types of GPU.<\/jats:p>","DOI":"10.1007\/s11227-022-05040-y","type":"journal-article","created":{"date-parts":[[2023,1,14]],"date-time":"2023-01-14T13:04:34Z","timestamp":1673701474000},"page":"9409-9442","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":5,"title":["EPSILOD: efficient parallel skeleton for generic iterative stencil computations in distributed GPUs"],"prefix":"10.1007","volume":"79","author":[{"given":"Manuel","family":"de Castro","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Inmaculada","family":"Santamaria-Valenzuela","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yuri","family":"Torres","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Arturo","family":"Gonzalez-Escribano","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Diego R.","family":"Llanos","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2023,1,14]]},"reference":[{"key":"5040_CR1","doi-asserted-by":"publisher","unstructured":"Ao Y, Yang C, Wang X, Xue W, Fu H, Liu F, Gan L, Xu P, Ma W (2017) 26 pflops stencil computations for atmospheric modeling on sunway taihulight. In: 2017 IEEE International parallel and Distributed Processing symposium (IPDPS), pp 535\u2013544. https:\/\/doi.org\/10.1109\/IPDPS.2017.9","DOI":"10.1109\/IPDPS.2017.9"},{"key":"5040_CR2","doi-asserted-by":"publisher","unstructured":"Rossinelli D, Hejazialhosseini B, Hadjidoukas P, Bekas C, Curioni A, Bertsch A, Futral S, Schmidt SJ, Adams NA, Koumoutsakos P (2013) 11 pflop\/s simulations of cloud cavitation collapse. In: Proceedings of the International Conference on High Performance Computing, Networking, Storage and Analysis. SC \u201913. Association for Computing Machinery, New York, NY, USA. https:\/\/doi.org\/10.1145\/2503210.2504565","DOI":"10.1145\/2503210.2504565"},{"key":"5040_CR3","doi-asserted-by":"publisher","unstructured":"Shimokawabe T, Aoki T, Muroi C, Ishida J, Kawano K, Endo T, Nukada A, Maruyama N, Matsuoka S (2010) An 80-fold speedup, 15.0 TFlops full GPU acceleration of non-hydrostatic weather model ASUCA production code\u2019. In: SC \u201910: Proceedings of the 2010 ACM\/IEEE International Conference for High Performance Computing, Networking, Storage and Analysis, pp 1\u201311. https:\/\/doi.org\/10.1109\/SC.2010.9","DOI":"10.1109\/SC.2010.9"},{"key":"5040_CR4","doi-asserted-by":"publisher","unstructured":"Shimokawabe T, Aoki T, Takaki T, Endo T, Yamanaka A, Maruyama N, Nukada A, Matsuoka S (2011) Peta-scale phase-field simulation for dendritic solidification on the tsubame 2.0 supercomputer. In: Proceedings of 2011 International Conference for High Performance Computing, Networking, Storage and Analysis. SC \u201911. Association for Computing Machinery, New York, NY, USA. https:\/\/doi.org\/10.1145\/2063384.2063388","DOI":"10.1145\/2063384.2063388"},{"key":"5040_CR5","unstructured":"TOP500.org (2022) TOP 500 Main Page. https:\/\/www.top500.org\/lists\/top500\/"},{"key":"5040_CR6","unstructured":"NVIDIA (2022) CUDA Toolkit Documentation v11.7.0. http:\/\/docs.nvidia.com\/cuda\/, Last visit: May, 2022"},{"key":"5040_CR7","unstructured":"Khronos (2022) Open Computing Language (OpenCL). http:\/\/www.khronos.org\/opencl\/, Last visit: May, 2022"},{"key":"5040_CR8","unstructured":"Forum M (2022) Message Passing Interface (MPI). https:\/\/www.mpi-forum.org\/, Last visit: May, 2022"},{"key":"5040_CR9","unstructured":"de Castro M, Santamaria-Valenzuela I, Miguel-Lopez S, Torres Y, Gonzalez-Escribano A (2021) Towards an efficient parallel skeleton for generic iterative stencil computations in distributed gpus. In: SC21\u2014ACM\/IEEE Conference on High Performance Networking and Computing. https:\/\/sc21.supercomputing.org\/proceedings\/tech_poster\/tech_poster_pages\/rpost167.html"},{"issue":"6","key":"5040_CR10","doi-asserted-by":"publisher","first-page":"838","DOI":"10.1177\/1094342017702962","volume":"32","author":"A Moreton-Fernandez","year":"2018","unstructured":"Moreton-Fernandez A, Ortega-Arranz H, Gonzalez-Escribano A (2018) Controllers: an abstraction to ease the use of hardware accelerators. Int J High Perform Comput Appl (IJHPCA) 32(6):838\u2013853. https:\/\/doi.org\/10.1177\/1094342017702962","journal-title":"Int J High Perform Comput Appl (IJHPCA)"},{"issue":"5","key":"5040_CR11","doi-asserted-by":"publisher","first-page":"1145","DOI":"10.1109\/TPDS.2013.83","volume":"25","author":"A Gonzalez-Escribano","year":"2014","unstructured":"Gonzalez-Escribano A, Torres Y, Fresno J, Llanos DR (2014) An extensible system for multilevel automatic data partition and mapping. IEEE Trans Parallel Distrib Syst 25(5):1145\u20131154. https:\/\/doi.org\/10.1109\/TPDS.2013.83","journal-title":"IEEE Trans Parallel Distrib Syst"},{"key":"5040_CR12","doi-asserted-by":"publisher","first-page":"291","DOI":"10.1007\/978-3-030-29400-7_21","volume-title":"Euro-Par 2019: parallel processing","author":"P Thoman","year":"2019","unstructured":"Thoman P, Salzmann P, Cosenza B, Fahringer T (2019) Celerity: high-level C++ for accelerator clusters. In: Yahyapour R (ed) Euro-Par 2019: parallel processing. Springer, Cham, pp 291\u2013303. https:\/\/doi.org\/10.1007\/978-3-030-29400-7_21"},{"key":"5040_CR13","doi-asserted-by":"publisher","unstructured":"Sourouri M, Langguth J, Spiga F, Baden SB, Cai X (2015) CPU+ GPU programming of stencil computations for resource-efficient use of GPU clusters. In: 2015 IEEE 18th International Conference on Computational Science and Engineering, pp 17\u201326. https:\/\/doi.org\/10.1109\/CSE.2015.33","DOI":"10.1109\/CSE.2015.33"},{"issue":"9","key":"5040_CR14","doi-asserted-by":"publisher","first-page":"536","DOI":"10.1016\/j.parco.2011.03.005","volume":"37","author":"C Feichtinger","year":"2011","unstructured":"Feichtinger C, Habich J, K\u00f6Stler H, Hager G, R\u00fcDe U, Wellein G (2011) A flexible patch-based lattice Boltzmann parallelization approach for heterogeneous GPU\u2013CPU clusters. Parallel Comput 37(9):536\u2013549. https:\/\/doi.org\/10.1016\/j.parco.2011.03.005","journal-title":"Parallel Comput"},{"key":"5040_CR15","doi-asserted-by":"publisher","unstructured":"Shimokawabe T, Aoki T, Ishida J, Kawano K, Muroi C (2011) 145 TFlops performance on 3990 GPUs of TSUBAME 2.0 supercomputer for an operational weather prediction. In: Proceedings of the International Conference on Computational Science, ICCS 2011, Nanyang Technological University, Singapore, 1-3 June, 2011, pp 1535\u20131544. https:\/\/doi.org\/10.1016\/j.procs.2011.04.166","DOI":"10.1016\/j.procs.2011.04.166"},{"key":"5040_CR16","doi-asserted-by":"publisher","unstructured":"Shimokawabe T, Aoki T, Muroi C, Ishida J, Kawano K, Endo T, Nukada A, Maruyama N, Matsuoka S (2010) An 80-fold speedup, 15.0 TFlops full GPU acceleration of non-hydrostatic weather model ASUCA production code. In: Proceedings of the 2010 ACM\/IEEE International Conference for High Performance Computing, Networking, Storage and Analysis. SC \u201910, pp 1\u201311. IEEE Computer Society, Washington, DC, USA. https:\/\/doi.org\/10.1109\/SC.2010.9","DOI":"10.1109\/SC.2010.9"},{"key":"5040_CR17","doi-asserted-by":"publisher","unstructured":"Shimokawabe T, Aoki T, Takaki T, Endo T, Yamanaka A, Maruyama N, Nukada A, Matsuoka S (2011) Peta-scale phase-field simulation for dendritic solidification on the TSUBAME 2.0 supercomputer. In: Proceedings of 2011 International Conference for High Performance Computing, Networking, Storage and Analysis. SC \u201911, pp 3\u20131311. ACM, New York, NY, USA. https:\/\/doi.org\/10.1145\/2063384.2063388","DOI":"10.1145\/2063384.2063388"},{"key":"5040_CR18","doi-asserted-by":"publisher","first-page":"285","DOI":"10.1007\/978-3-540-87475-1_39","volume-title":"Recent Advances in Parallel Virtual Machine and Message Passing Interface","author":"A Sch\u00e4fer","year":"2008","unstructured":"Sch\u00e4fer A, Fey D (2008) libgeodecomp: a grid-enabled library for geometric decomposition codes. In: Lastovetsky A, Kechadi T, Dongarra J (eds) Recent Advances in Parallel Virtual Machine and Message Passing Interface. Springer, Berlin, pp 285\u2013294. https:\/\/doi.org\/10.1007\/978-3-540-87475-1_39"},{"key":"5040_CR19","doi-asserted-by":"publisher","unstructured":"Stark DT, Barrett RF, Grant RE, Olivier SL, Pedretti KT, Vaughan CT (2014) Early experiences co-scheduling work and communication tasks for hybrid MPI+ X applications. In: 2014 Workshop on Exascale MPI at Supercomputing Conference, pp 9\u201319. https:\/\/doi.org\/10.1109\/ExaMPI.2014.6","DOI":"10.1109\/ExaMPI.2014.6"},{"key":"5040_CR20","unstructured":"Chakroun I, Vander Aa T, De Fraine B, Haber T, Wuyts R, Demeuter W (2015) Exashark: A scalable hybrid array kit for exascale simulation. In: Proceedings of the Symposium on High Performance Computing. HPC \u201915, pp 41\u201348. Society for Computer Simulation International, San Diego, CA, USA. http:\/\/dl.acm.org\/citation.cfm?id=2872599.2872605"},{"key":"5040_CR21","doi-asserted-by":"publisher","unstructured":"Baskaran M, Pradelle B, Meister B, Konstantinidis A, Lethin R (2016) Automatic code generation and data management for an asynchronous task-based runtime. In: 2016 5th Workshop on Extreme-Scale Programming Tools (ESPT), pp 34\u201341. https:\/\/doi.org\/10.1109\/ESPT.2016.009","DOI":"10.1109\/ESPT.2016.009"},{"key":"5040_CR22","doi-asserted-by":"publisher","unstructured":"Bachan J, Bonachea D, Hargrove PH, Hofmeyr S, Jacquelin M, Kamil A, van Straalen B, Baden SB (2017) The UPC++ PGAS library for exascale computing. In: Proceedings of the Second Annual PGAS Applications Workshop. PAW17, pp 7\u2013174. ACM, New York, NY, USA. https:\/\/doi.org\/10.1145\/3144779.3169108","DOI":"10.1145\/3144779.3169108"},{"key":"5040_CR23","doi-asserted-by":"publisher","unstructured":"Tanaka H, Ishihara Y, Sakamoto R, Nakamura T, Kimura Y, Nitadori K, Tsubouchi M, Makino J (2018) Automatic generation of high-order finite-difference code with temporal blocking for extreme-scale many-core systems. In: 2018 IEEE\/ACM 4th International Workshop on Extreme Scale Programming Models and Middleware (ESPM2), pp. 29\u201336. https:\/\/doi.org\/10.1109\/ESPM2.2018.00008","DOI":"10.1109\/ESPM2.2018.00008"},{"issue":"4","key":"5040_CR24","doi-asserted-by":"publisher","first-page":"40","DOI":"10.1145\/3274653","volume":"15","author":"S Kronawitter","year":"2018","unstructured":"Kronawitter S, Lengauer C (2018) Polyhedral search space exploration in the exastencils code generator. ACM Trans Archit Code Optim 15(4):40\u201314025. https:\/\/doi.org\/10.1145\/3274653","journal-title":"ACM Trans Archit Code Optim"},{"key":"5040_CR25","doi-asserted-by":"publisher","DOI":"10.1145\/3374916","author":"F Luporini","year":"2020","unstructured":"Luporini F, Louboutin M, Lange M, Kukreja N, Witte P, H\u00fcckelheim J, Yount C, Kelly PHJ, Herrmann FJ, Gorman GJ (2020) Architecture and performance of devito, a system for automated stencil computation. ACM Trans Math Softw. https:\/\/doi.org\/10.1145\/3374916","journal-title":"ACM Trans Math Softw"},{"key":"5040_CR26","doi-asserted-by":"publisher","unstructured":"Hagedorn B, Stoltzfus L, Steuwer M, Gorlatch S, Dubach C (2018) High performance stencil code generation with lift. In: Proceedings of the 2018 International Symposium on Code Generation and Optimization. CGO 2018, pp 100\u2013112. ACM, New York, NY, USA. https:\/\/doi.org\/10.1145\/3168824","DOI":"10.1145\/3168824"},{"key":"5040_CR27","doi-asserted-by":"publisher","unstructured":"Pereira AD, Castro M, Dantas MAR, Rocha RCO, G\u00f3es LFW (2017) Extending OpenACC for efficient stencil code generation and execution by skeleton frameworks. In: 2017 International Conference on High Performance Computing Simulation (HPCS), pp 719\u2013726. https:\/\/doi.org\/10.1109\/HPCS.2017.110","DOI":"10.1109\/HPCS.2017.110"},{"key":"5040_CR28","doi-asserted-by":"publisher","first-page":"2027","DOI":"10.1016\/j.procs.2011.04.221","volume":"4","author":"A Sch\u00e4fer","year":"2011","unstructured":"Sch\u00e4fer A, Fey D (2011) High performance stencil code algorithms for GPGPUs. Procedia Comput Sci 4:2027\u20132036. https:\/\/doi.org\/10.1016\/j.procs.2011.04.221. (Proceedings of the International Conference on Computational Science, ICCS 2011)","journal-title":"Procedia Comput Sci"},{"key":"5040_CR29","doi-asserted-by":"publisher","unstructured":"Anjum O, Simon GdG, Hidayetoglu M, Hwu W-M (2019) An efficient GPU implementation technique for higher-order 3d stencils. In: 2019 IEEE 21st International Conference on High Performance Computing and Communications; IEEE 17th International Conference on Smart City; IEEE 5th International Conference on Data Science and Systems (HPCC\/SmartCity\/DSS), pp 552\u2013561. https:\/\/doi.org\/10.1109\/HPCC\/SmartCity\/DSS.2019.00086","DOI":"10.1109\/HPCC\/SmartCity\/DSS.2019.00086"},{"key":"5040_CR30","doi-asserted-by":"publisher","unstructured":"Matsumura K, Zohouri HR, Wahib M, Endo T, Matsuoka S (2020) AN5D: automated stencil framework for high-degree temporal blocking on GPUs. In: Proceedings of the 18th International Symposium on Code Generation and Optimization, pp 199\u2013211. Association for Computing Machinery, New York, NY, USA. https:\/\/doi.org\/10.1145\/3368826.3377904","DOI":"10.1145\/3368826.3377904"},{"key":"5040_CR31","doi-asserted-by":"publisher","unstructured":"Rawat PS, Vaidya M, Sukumaran-Rajam A, Rountev A, Pouchet L-N, Sadayappan P (2019) On optimizing complex stencils on GPUs. In: 2019 IEEE International Parallel and Distributed Processing Symposium (IPDPS), pp 641\u2013652. https:\/\/doi.org\/10.1109\/IPDPS.2019.00073","DOI":"10.1109\/IPDPS.2019.00073"},{"key":"5040_CR32","doi-asserted-by":"publisher","unstructured":"Oh C, Zheng Z, Shen X, Zhai J, Yi Y (2020) Gopipe: A granularity-oblivious programming framework for pipelined stencil executions on GPU. In: Proceedings of the ACM International Conference on Parallel Architectures and Compilation Techniques. PACT \u201920, pp 43\u201354. Association for Computing Machinery, New York, NY, USA. https:\/\/doi.org\/10.1145\/3410463.3414656","DOI":"10.1145\/3410463.3414656"},{"issue":"17","key":"5040_CR33","doi-asserted-by":"publisher","first-page":"4938","DOI":"10.1002\/cpe.3479","volume":"27","author":"AD Pereira","year":"2015","unstructured":"Pereira AD, Ramos L, G\u00f3es LFW (2015) Pskel: a stencil programming framework for CPU\u2013GPU systems. Concurr Comput Pract Exper 27(17):4938\u20134953. https:\/\/doi.org\/10.1002\/cpe.3479","journal-title":"Concurr Comput Pract Exper"},{"issue":"12","key":"5040_CR34","doi-asserted-by":"publisher","first-page":"4152","DOI":"10.1002\/cpe.4152","volume":"29","author":"M Vi\u00f1as","year":"2017","unstructured":"Vi\u00f1as M, Fraguela BB, Andrade D, Doallo R (2017) Facilitating the development of stencil applications using the heterogeneous programming library. Concurr Comput Pract Exp 29(12):4152. https:\/\/doi.org\/10.1002\/cpe.4152","journal-title":"Concurr Comput Pract Exp"},{"issue":"03","key":"5040_CR35","doi-asserted-by":"publisher","first-page":"1441005","DOI":"10.1142\/S0129626414410059","volume":"24","author":"M Steuwer","year":"2014","unstructured":"Steuwer M, Haidl M, Breuer S, Gorlatch S (2014) High-level programming of stencil computations on multi-GPU systems using the SkelCL library. Parallel Process Lett 24(03):1441005. https:\/\/doi.org\/10.1142\/S0129626414410059","journal-title":"Parallel Process Lett"},{"key":"5040_CR36","doi-asserted-by":"publisher","unstructured":"Maruyama N, Sato K, Nomura T, Matsuoka S (2011) Physis: an implicitly parallel programming model for stencil computations on large-scale GPU-accelerated supercomputers. In: SC \u201911: Proceedings of 2011 International Conference for High Performance Computing, Networking, Storage and Analysis, pp 1\u201312. https:\/\/doi.org\/10.1145\/2063384.2063398","DOI":"10.1145\/2063384.2063398"},{"issue":"4","key":"5040_CR37","doi-asserted-by":"publisher","first-page":"59","DOI":"10.1145\/2400682.2400718","volume":"9","author":"T Lutz","year":"2013","unstructured":"Lutz T, Fensch C, Cole M (2013) Partans: an autotuning framework for stencil computation on multi-GPU systems. ACM Trans Archit Code Optim 9(4):59\u201315924. https:\/\/doi.org\/10.1145\/2400682.2400718","journal-title":"ACM Trans Archit Code Optim"},{"key":"5040_CR38","unstructured":"Shimokawabe T, Aoki T, Onodera N (2014) A high-productivity framework for multi-gpu computation of mesh-based applications. In: Gr\u00f6sslinger A, K\u00f6stler H (eds), Proceedings of the 1st International Workshop on High-Performance Stencil Computations, Vienna, Austria, pp 23\u201330. https:\/\/hgpu.org\/?p=11286"},{"key":"5040_CR39","unstructured":"Breuer S, Steuwer M, Gorlatch S (2014) Extending the SkelCL skeleton library for stencil computations on multi-GPU systems. In: HiStencils 2014, First International Workshop on High-Performance Stencil Computations, pp 1\u201313. https:\/\/hgpu.org\/?p=11368"},{"issue":"11","key":"5040_CR40","doi-asserted-by":"publisher","first-page":"5690","DOI":"10.1007\/s11227-016-1871-z","volume":"74","author":"M Aldinucci","year":"2018","unstructured":"Aldinucci M, Danelutto M, Drocco M, Kilpatrick P, Misale C, Peretti Pezzi G, Torquati M (2018) A parallel pattern for iterative stencil + reduce. J Supercomput 74(11):5690\u20135705. https:\/\/doi.org\/10.1007\/s11227-016-1871-z","journal-title":"J Supercomput"},{"issue":"3","key":"5040_CR41","doi-asserted-by":"publisher","first-page":"32","DOI":"10.1145\/3232521","volume":"15","author":"H Kim","year":"2018","unstructured":"Kim H, Hadidi R, Nai L, Kim H, Jayasena N, Eckert Y, Kayiran O, Loh G (2018) Coda: enabling co-location of computation and data for multiple GPU systems. ACM Trans Archit Code Optim 15(3):32\u201313223. https:\/\/doi.org\/10.1145\/3232521","journal-title":"ACM Trans Archit Code Optim"},{"issue":"5","key":"5040_CR42","doi-asserted-by":"publisher","first-page":"433","DOI":"10.1007\/s10766-022-00735-4","volume":"50","author":"N Herrmann","year":"2022","unstructured":"Herrmann N, de Melo Menezes BA, Kuchen H (2022) Stencil calculations with algorithmic skeletons for heterogeneous computing environments. Int J Parallel Program 50(5):433\u2013453. https:\/\/doi.org\/10.1007\/s10766-022-00735-4","journal-title":"Int J Parallel Program"},{"key":"5040_CR43","volume-title":"Parallel Programming in OpenMP","author":"R Chandra","year":"2001","unstructured":"Chandra R, Dagum L, Kohr D, Maydan D, McDonald J, Menon R (2001) Parallel Programming in OpenMP. Morgan Kaufmann Publishers Inc., San Francisco"},{"key":"5040_CR44","unstructured":"Tian S, Doerfert J, Chapman B (2020) Extending the SkelCL skeleton library for stencil computations on multi-GPU systems. In: Fourth LLVM Performance Workshop at CGO. https:\/\/llvm.org\/devmtg\/2020-02-23\/"},{"key":"5040_CR45","doi-asserted-by":"publisher","unstructured":"Beckingsale DA, Burmark J, Hornung R, Jones H, Killian W, Kunen AJ, Pearce O, Robinson P, Ryujin BS, Scogland TR (2019) Raja: Portable performance for large-scale scientific applications. In: 2019 IEEE\/ACM International Workshop on Performance, Portability and Productivity in HPC (P3HPC), pp 71\u201381. https:\/\/doi.org\/10.1109\/P3HPC49587.2019.00012","DOI":"10.1109\/P3HPC49587.2019.00012"},{"key":"5040_CR46","doi-asserted-by":"publisher","unstructured":"Beckingsale DA, Burmark J, Hornung R, Jones H, Killian W, Kunen AJ, Pearce O, Robinson P, Ryujin BS, Scogland TR (2019) Raja: Portable performance for large-scale scientific applications. In: 2019 IEEE\/ACM International Workshop on Performance, Portability and Productivity in HPC (P3HPC). IEEE, New York, NY, USA. https:\/\/doi.org\/10.1109\/P3HPC49587.2019. IEEE\/ACM","DOI":"10.1109\/P3HPC49587.2019"},{"issue":"12","key":"5040_CR47","doi-asserted-by":"publisher","first-page":"3202","DOI":"10.1016\/j.jpdc.2014.07.003","volume":"74","author":"HC Edwards","year":"2014","unstructured":"Edwards HC, Trott CR, Sunderland D (2014) Kokkos: enabling manycore performance portability through polymorphic memory access patterns. J Parallel Distrib Comput 74(12):3202\u20133216. https:\/\/doi.org\/10.1016\/j.jpdc.2014.07.003. (Domain-Specific Languages and High-Level Frameworks for High-Performance Computing)","journal-title":"J Parallel Distrib Comput"},{"issue":"4","key":"5040_CR48","doi-asserted-by":"publisher","first-page":"805","DOI":"10.1109\/TPDS.2021.3097283","volume":"33","author":"CR Trott","year":"2022","unstructured":"Trott CR, Lebrun-Grandi\u00e9 D, Arndt D, Ciesko J, Dang V, Ellingwood N, Gayatri R, Harvey E, Hollman DS, Ibanez D, Liber N, Madsen J, Miles J, Poliakoff D, Powell A, Rajamanickam S, Simberg M, Sunderland D, Turcksin B, Wilke J (2022) Kokkos 3: Programming model extensions for the exascale era. IEEE Trans Parallel Distrib Syst 33(4):805\u2013817. https:\/\/doi.org\/10.1109\/TPDS.2021.3097283","journal-title":"IEEE Trans Parallel Distrib Syst"},{"key":"5040_CR49","doi-asserted-by":"publisher","unstructured":"Ciesko J (2020) Distributed memory programming and multi-GPU Support with KOKKOS. Presented at SC\u201920. https:\/\/doi.org\/10.2172\/1829951. https:\/\/www.osti.gov\/biblio\/1829951","DOI":"10.2172\/1829951"},{"key":"5040_CR50","unstructured":"Khronos OpenCL working group (2020) SYCL 1.2.1 specification standard. (accessed February 1, 2022). https:\/\/www.khronos.org\/registry\/SYCL\/specs\/sycl-1.2.1.pdf"},{"key":"5040_CR51","doi-asserted-by":"publisher","unstructured":"Gorlatch S, Cole M (2011) In: Padua D (ed), Parallel Skeletons, pp 1417\u20131422. Springer, Boston. https:\/\/doi.org\/10.1007\/978-0-387-09766-4_24","DOI":"10.1007\/978-0-387-09766-4_24"},{"key":"5040_CR52","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1007\/s11227-021-03792-7","volume":"77","author":"G Rodriguez-Canal","year":"2021","unstructured":"Rodriguez-Canal G, Torres Y, Andujar FJ, Gonzalez-Escribano A (2021) Efficient heterogeneous programming with FPGAs using the Controller model. J Supercomput 77:1\u201316. https:\/\/doi.org\/10.1007\/s11227-021-03792-7","journal-title":"J Supercomput"}],"container-title":["The Journal of Supercomputing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11227-022-05040-y.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11227-022-05040-y\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11227-022-05040-y.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,4,24]],"date-time":"2023-04-24T22:03:11Z","timestamp":1682373791000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11227-022-05040-y"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,1,14]]},"references-count":52,"journal-issue":{"issue":"9","published-print":{"date-parts":[[2023,6]]}},"alternative-id":["5040"],"URL":"https:\/\/doi.org\/10.1007\/s11227-022-05040-y","relation":{},"ISSN":["0920-8542","1573-0484"],"issn-type":[{"value":"0920-8542","type":"print"},{"value":"1573-0484","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,1,14]]},"assertion":[{"value":"29 December 2022","order":1,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"14 January 2023","order":2,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors have no competing interests as defined by Springer, or other interests that might be perceived to influence the results and\/or discussion reported in this paper.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}},{"value":"Not applicable.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethical approval"}},{"value":"The source code of both modules is freely available on the repository of the Trasgo group: .","order":4,"name":"Ethics","group":{"name":"EthicsHeading","label":"Availability of data and materials"}}]}}