{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,25]],"date-time":"2026-06-25T17:26:45Z","timestamp":1782408405098,"version":"3.54.5"},"reference-count":21,"publisher":"Springer Science and Business Media LLC","issue":"5-6","license":[{"start":{"date-parts":[[2022,7,23]],"date-time":"2022-07-23T00:00:00Z","timestamp":1658534400000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2022,7,23]],"date-time":"2022-07-23T00:00:00Z","timestamp":1658534400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100004869","name":"Westf\u00e4lische Wilhelms-Universit\u00e4t M\u00fcnster","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100004869","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Int J Parallel Prog"],"published-print":{"date-parts":[[2022,12]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>The development of parallel applications is a difficult and error-prone task, especially for inexperienced programmers. Stencil operations are exceptionally complex for parallelization as synchronization and communication between the individual processes and threads are necessary. It gets even more difficult to efficiently distribute the computations and efficiently implement communication when heterogeneous computing environments are used. For using multiple nodes, each having multiple cores and accelerators such as GPUs, skills in combining frameworks such as MPI, OpenMP, and CUDA are required. The complexity of parallelizing the stencil operation increases the need for abstracting from the platform-specific details and simplify parallel programming. One way to abstract from details of parallel programming is to use algorithmic skeletons. This work introduces an implementation of the MapStencil skeleton that is able to generate parallel code for distributed memory environments, using multiple nodes with multicore CPUs and GPUs. Examples of practical applications of the MapStencil skeleton are the Jacobi Solver or the Canny Edge Detector. The main contribution of this paper is a discussion of the difficulties when implementing a universal Skeleton for MapStencil for heterogeneous computing environments and an outline of the identified best practices for communication intense skeletons.<\/jats:p>","DOI":"10.1007\/s10766-022-00735-4","type":"journal-article","created":{"date-parts":[[2022,7,23]],"date-time":"2022-07-23T08:05:02Z","timestamp":1658563502000},"page":"433-453","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":7,"title":["Stencil Calculations with Algorithmic Skeletons for Heterogeneous Computing Environments"],"prefix":"10.1007","volume":"50","author":[{"given":"Nina","family":"Herrmann","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Breno A.","family":"de Melo Menezes","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Herbert","family":"Kuchen","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2022,7,23]]},"reference":[{"key":"735_CR1","doi-asserted-by":"crossref","unstructured":"Aldinucci, M., Danelutto, M., Drocco, M., Kilpatrick, P., Pezzi, G.P., Torquati, M.: The loop-of-stencil-reduce paradigm. In: 2015 IEEE Trustcom\/BigDataSE\/ISPA, vol. 3, pp. 172\u2013177. IEEE (2015)","DOI":"10.1109\/Trustcom.2015.628"},{"key":"735_CR2","doi-asserted-by":"crossref","unstructured":"Aldinucci, M., Danelutto, M., Kilpatrick, P., Torquati, M.: Fastflow: high-level and efficient streaming on multi-core. In: Programming Multi-core and Many-core Computing Systems, Parallel and Distributed Computing (2017)","DOI":"10.1002\/9781119332015.ch13"},{"key":"735_CR3","doi-asserted-by":"crossref","unstructured":"Benoit, A., Cole, M., Gilmore, S., Hillston, J.: Flexible skeletal programming with eSkel. In: European Conference on Parallel Processing, pp. 761\u2013770. Springer (2005)","DOI":"10.1007\/11549468_83"},{"issue":"2","key":"735_CR4","doi-asserted-by":"publisher","first-page":"468","DOI":"10.1007\/s11227-015-1575-9","volume":"72","author":"TLB Cheikh","year":"2016","unstructured":"Cheikh, T.L.B., Aguiar, A., Tahar, S., Nicolescu, G.: Tuning framework for stencil computation in heterogeneous parallel platforms. J. Supercomput. 72(2), 468\u2013502 (2016)","journal-title":"J. Supercomput."},{"issue":"3","key":"735_CR5","doi-asserted-by":"publisher","first-page":"205","DOI":"10.1007\/s00450-011-0160-6","volume":"26","author":"M Christen","year":"2011","unstructured":"Christen, M., Schenk, O., Burkhart, H.: Automatic code generation and tuning for stencil kernels on modern shared memory architectures. Comput. Sci. Res. Dev. 26(3), 205\u2013210 (2011)","journal-title":"Comput. Sci. Res. Dev."},{"key":"735_CR6","volume-title":"Algorithmic Skeletons: Structured Management of Parallel Computation","author":"MI Cole","year":"1989","unstructured":"Cole, M.I.: Algorithmic Skeletons: Structured Management of Parallel Computation. Pitman, London (1989)"},{"key":"735_CR7","unstructured":"Corporation, N.: Cuda. https:\/\/developer.nvidia.com\/cuda-zone (2021). Accessed 10 May 2021"},{"key":"735_CR8","doi-asserted-by":"crossref","unstructured":"Crank, J., Nicolson, P.: A practical method for numerical evaluation of solutions of partial differential equations of the heat-conduction type. In: Mathematical Proceedings of the Cambridge Philosophical Society, vol. 43, pp. 50\u201367. Cambridge University Press (1947)","DOI":"10.1017\/S0305004100023197"},{"key":"735_CR9","doi-asserted-by":"publisher","first-page":"254","DOI":"10.1007\/978-3-642-28869-2_13","volume-title":"Programming Languages and Systems","author":"K Emoto","year":"2012","unstructured":"Emoto, K., Fischer, S., Hu, Z.: Generate, test, and aggregate. In: Seidl, H. (ed.) Programming Languages and Systems, pp. 254\u2013273. Springer, Berlin (2012)"},{"key":"735_CR10","doi-asserted-by":"crossref","unstructured":"Enmyren, J., Kessler, C.W.: Skepu: a multi-backend skeleton programming library for multi-gpu systems. In: Proceedings of the Fourth International Workshop on High-level Parallel Programming and Applications, pp. 5\u201314 (2010)","DOI":"10.1145\/1863482.1863487"},{"issue":"2","key":"735_CR11","doi-asserted-by":"publisher","first-page":"129","DOI":"10.1504\/IJHPCN.2012.046370","volume":"7","author":"S Ernsting","year":"2012","unstructured":"Ernsting, S., Kuchen, H.: Algorithmic skeletons for multi-core, multi-GPU systems and clusters. Int. J. High Perform. Comput. Netw. 7(2), 129\u2013138 (2012)","journal-title":"Int. J. High Perform. Comput. Netw."},{"issue":"2","key":"735_CR12","doi-asserted-by":"publisher","first-page":"283","DOI":"10.1007\/s10766-016-0416-7","volume":"45","author":"S Ernsting","year":"2017","unstructured":"Ernsting, S., Kuchen, H.: Data parallel algorithmic skeletons with accelerator support. Int. J. Parallel Prog. 45(2), 283\u2013299 (2017)","journal-title":"Int. J. Parallel Prog."},{"key":"735_CR13","unstructured":"Forum, M.: Mpi standard. https:\/\/www.mpi-forum.org\/docs\/ (2021). Accessed 10 May 2021"},{"key":"735_CR14","doi-asserted-by":"crossref","unstructured":"Hagedorn, B., Stoltzfus, L., Steuwer, M., Gorlatch, S., Dubach, C.: High performance stencil code generation with lift. In: Proceedings of the 2018 International Symposium on Code Generation and Optimization, pp. 100\u2013112 (2018)","DOI":"10.1145\/3168824"},{"issue":"1","key":"735_CR15","doi-asserted-by":"publisher","first-page":"72","DOI":"10.1109\/TPDS.2016.2549523","volume":"28","author":"X Mei","year":"2017","unstructured":"Mei, X., Chu, X.: Dissecting GPU memory hierarchy through microbenchmarking. IEEE Trans. Parallel Distrib. Syst. 28(1), 72\u201386 (2017). https:\/\/doi.org\/10.1109\/TPDS.2016.2549523","journal-title":"IEEE Trans. Parallel Distrib. Syst."},{"issue":"7","key":"735_CR16","doi-asserted-by":"publisher","first-page":"5038","DOI":"10.1007\/s11227-019-02824-7","volume":"76","author":"T \u00d6hberg","year":"2020","unstructured":"\u00d6hberg, T., Ernstsson, A., Kessler, C.: Hybrid CPU-GPU execution support in the skeleton programming framework SkePU. J. Supercomput. 76(7), 5038\u20135056 (2020)","journal-title":"J. Supercomput."},{"key":"735_CR17","unstructured":"OpenMP: OpenMP the openMP API specification for parallel programming. https:\/\/www.openmp.org\/ (2021). Accessed 10 May 2021"},{"key":"735_CR18","doi-asserted-by":"crossref","unstructured":"Tang, Y., Chowdhury, R.A., Kuszmaul, B.C., Luk, C.K., Leiserson, C.E.: The pochoir stencil compiler. In: Proceedings of the Twenty-Third Annual ACM Symposium on Parallelism in Algorithms and Architectures, pp. 117\u2013128 (2011)","DOI":"10.1145\/1989493.1989508"},{"key":"735_CR19","unstructured":"Van\u00a0Werkhoven, B., Maassen, J., Seinstra, F.J.: Optimizing convolution operations in cuda with adaptive tiling. In: 2nd Workshop on Applications for Multi and Many Core Processors (A4MMC 2011) (2011)"},{"issue":"7","key":"735_CR20","doi-asserted-by":"publisher","first-page":"5098","DOI":"10.1007\/s11227-019-02825-6","volume":"76","author":"F Wrede","year":"2020","unstructured":"Wrede, F., Rieger, C., Kuchen, H.: Generation of high-performance code based on a domain-specific language for algorithmic skeletons. J. Supercomput. 76(7), 5098\u20135116 (2020)","journal-title":"J. Supercomput."},{"key":"735_CR21","doi-asserted-by":"crossref","unstructured":"Zhang, Y., Mueller, F.: Auto-generation and auto-tuning of 3d stencil codes on GPU clusters. In: Proceedings of the Tenth International Symposium on Code Generation and Optimization, pp. 155\u2013164 (2012)","DOI":"10.1145\/2259016.2259037"}],"container-title":["International Journal of Parallel Programming"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10766-022-00735-4.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10766-022-00735-4\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10766-022-00735-4.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,11,21]],"date-time":"2022-11-21T21:38:51Z","timestamp":1669066731000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10766-022-00735-4"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,7,23]]},"references-count":21,"journal-issue":{"issue":"5-6","published-print":{"date-parts":[[2022,12]]}},"alternative-id":["735"],"URL":"https:\/\/doi.org\/10.1007\/s10766-022-00735-4","relation":{},"ISSN":["0885-7458","1573-7640"],"issn-type":[{"value":"0885-7458","type":"print"},{"value":"1573-7640","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,7,23]]},"assertion":[{"value":"12 October 2021","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"18 May 2022","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"23 July 2022","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}]}}