{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,2,21]],"date-time":"2025-02-21T07:44:38Z","timestamp":1740123878543,"version":"3.37.3"},"reference-count":16,"publisher":"Springer Science and Business Media LLC","issue":"2-3","license":[{"start":{"date-parts":[[2023,1,7]],"date-time":"2023-01-07T00:00:00Z","timestamp":1673049600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2023,1,7]],"date-time":"2023-01-07T00:00:00Z","timestamp":1673049600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100004869","name":"Westf\u00e4lische Wilhelms-Universit\u00e4t M\u00fcnster","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100004869","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Int J Parallel Prog"],"published-print":{"date-parts":[[2023,6]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Contemporary HPC hardware typically provides several levels of parallelism, e.g. multiple nodes, each having multiple cores (possibly with vectorization) and accelerators. Efficiently programming such systems usually requires skills in combining several low-level frameworks such as MPI, OpenMP, and CUDA. This overburdens programmers without substantial parallel programming skills. One way to overcome this problem and to abstract from details of parallel programming is to use algorithmic skeletons. In the present paper, we evaluate the multi-node, multi-CPU and multi-GPU implementation of the most essential skeletons Map, Reduce, and Zip. Our main contribution is a discussion of the efficiency of using multiple parallelization levels and the consideration of which fine-tune settings should be offered to the user.<\/jats:p>","DOI":"10.1007\/s10766-022-00742-5","type":"journal-article","created":{"date-parts":[[2023,1,7]],"date-time":"2023-01-07T12:05:35Z","timestamp":1673093135000},"page":"172-185","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":1,"title":["Distributed Calculations with Algorithmic Skeletons for Heterogeneous Computing Environments"],"prefix":"10.1007","volume":"51","author":[{"given":"Nina","family":"Herrmann","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Herbert","family":"Kuchen","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2023,1,7]]},"reference":[{"key":"742_CR1","unstructured":"MPI Forum. Mpi standard. https:\/\/www.mpi-forum.org\/docs\/ (2021). Accessed: 10.05.2021"},{"key":"742_CR2","unstructured":"OpenMP. Openmp the openmp api specification for parallel programming. https:\/\/www.openmp.org\/ (2021). Accessed: 10.05.2021"},{"key":"742_CR3","unstructured":"NVIDIA Corporation. Cuda. https:\/\/developer.nvidia.com\/cuda-zone (2021). Accessed: 10.05.2021"},{"key":"742_CR4","volume-title":"Algorithmic skeletons: structured management of parallel computation","author":"MI Cole","year":"1989","unstructured":"Cole, M.I.: Algorithmic skeletons: structured management of parallel computation. Pitman, London (1989)"},{"issue":"2","key":"742_CR5","doi-asserted-by":"publisher","first-page":"129","DOI":"10.1504\/IJHPCN.2012.046370","volume":"7","author":"S Ernsting","year":"2012","unstructured":"Ernsting, S., Kuchen, H.: Algorithmic skeletons for multi-core, multi-gpu systems and clusters. Int. J. High Perform. Comput. Netw. 7(2), 129\u2013138 (2012)","journal-title":"Int. J. High Perform. Comput. Netw."},{"issue":"2","key":"742_CR6","doi-asserted-by":"publisher","first-page":"283","DOI":"10.1007\/s10766-016-0416-7","volume":"45","author":"S Ernsting","year":"2017","unstructured":"Ernsting, S., Kuchen, H.: Data parallel algorithmic skeletons with accelerator support. Int. J. Parallel Program. 45(2), 283\u2013299 (2017)","journal-title":"Int. J. Parallel Program."},{"key":"742_CR7","doi-asserted-by":"crossref","unstructured":"Benoit, A., Cole, M., Gilmore, S., Hillston, J.: Flexible skeletal programming with eskel. In: European Conference on Parallel Processing, pp. 761\u2013770. Springer (2005)","DOI":"10.1007\/11549468_83"},{"issue":"7","key":"742_CR8","doi-asserted-by":"publisher","first-page":"5098","DOI":"10.1007\/s11227-019-02825-6","volume":"76","author":"F Wrede","year":"2020","unstructured":"Wrede, F., Rieger, C., Kuchen, H.: Generation of high-performance code based on a domain-specific language for algorithmic skeletons. J. Supercomput. 76(7), 5098\u20135116 (2020)","journal-title":"J. Supercomput."},{"issue":"6","key":"742_CR9","doi-asserted-by":"publisher","first-page":"846","DOI":"10.1007\/s10766-021-00704-3","volume":"49","author":"A Ernstsson","year":"2021","unstructured":"Ernstsson, A., Ahlqvist, J., Zouzoula, S., Kessler, C.: Skepu 3: Portable high-level programming of heterogeneous systems and hpc clusters. Int. J. Parallel Program. 49(6), 846\u2013866 (2021)","journal-title":"Int. J. Parallel Program."},{"key":"742_CR10","doi-asserted-by":"crossref","unstructured":"Aldinucci, M., Danelutto, M., Kilpatrick, P., Torquati, M.: Fastflow: high-level and efficient streaming on multi-core. Programming multi-core and many-core computing systems, parallel and distributed computing (2017)","DOI":"10.1002\/9781119332015.ch13"},{"issue":"7","key":"742_CR11","doi-asserted-by":"publisher","first-page":"5038","DOI":"10.1007\/s11227-019-02824-7","volume":"76","author":"T \u00d6hberg","year":"2020","unstructured":"\u00d6hberg, T., Ernstsson, A., Kessler, C.: Hybrid cpu-gpu execution support in the skeleton programming framework skepu. J. Supercomput. 76(7), 5038\u20135056 (2020)","journal-title":"J. Supercomput."},{"issue":"1","key":"742_CR12","doi-asserted-by":"publisher","first-page":"25","DOI":"10.1007\/s11227-014-1213-y","volume":"69","author":"M Steuwer","year":"2014","unstructured":"Steuwer, M., Gorlatch, S.: Skelcl: a high-level extension of opencl for multi-gpu systems. J. Supercomput. 69(1), 25\u201333 (2014)","journal-title":"J. Supercomput."},{"key":"742_CR13","doi-asserted-by":"crossref","unstructured":"Rieger, C., Wrede, F., Kuchen, H.: Musket: a domain-specific language for high-level parallel programming with algorithmic skeletons. In: Proceedings of the 34th ACM\/SIGAPP Symposium on Applied Computing, pp. 1534\u20131543 (2019)","DOI":"10.1145\/3297280.3297434"},{"key":"742_CR14","doi-asserted-by":"crossref","unstructured":"Soldado, F., Alexandre, F., Paulino, H.: Towards the transparent execution of compound opencl computations in multi-cpu\/multi-gpu environments. In: European Conference on Parallel Processing, pp. 177\u2013188. Springer (2014)","DOI":"10.1007\/978-3-319-14325-5_16"},{"key":"742_CR15","doi-asserted-by":"crossref","unstructured":"Luk, C. K., Hong, S., Kim, H.: Qilin: exploiting parallelism on heterogeneous multiprocessors with adaptive mapping. In: 2009 42nd Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO), pp. 45\u201355. IEEE (2009)","DOI":"10.1145\/1669112.1669121"},{"key":"742_CR16","doi-asserted-by":"publisher","first-page":"254","DOI":"10.1007\/978-3-642-28869-2_13","volume-title":"1Programming Languages and Systems","author":"K Emoto","year":"2012","unstructured":"Emoto, K., Fischer, S., Hu, Z.: Generate, test, and aggregate. In: Seidl, H. (ed.) 1Programming Languages and Systems, pp. 254\u2013273. Springer Berlin Heidelberg, Berlin, Heidelberg (2012)"}],"container-title":["International Journal of Parallel Programming"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10766-022-00742-5.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10766-022-00742-5\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10766-022-00742-5.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,3,30]],"date-time":"2023-03-30T11:13:47Z","timestamp":1680174827000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10766-022-00742-5"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,1,7]]},"references-count":16,"journal-issue":{"issue":"2-3","published-print":{"date-parts":[[2023,6]]}},"alternative-id":["742"],"URL":"https:\/\/doi.org\/10.1007\/s10766-022-00742-5","relation":{},"ISSN":["0885-7458","1573-7640"],"issn-type":[{"type":"print","value":"0885-7458"},{"type":"electronic","value":"1573-7640"}],"subject":[],"published":{"date-parts":[[2023,1,7]]},"assertion":[{"value":"25 September 2022","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"21 November 2022","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"7 January 2023","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare no competing interests.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}}]}}