{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,24]],"date-time":"2026-01-24T14:45:32Z","timestamp":1769265932116,"version":"3.49.0"},"reference-count":19,"publisher":"Springer Science and Business Media LLC","issue":"7","license":[{"start":{"date-parts":[[2022,12,2]],"date-time":"2022-12-02T00:00:00Z","timestamp":1669939200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2022,12,2]],"date-time":"2022-12-02T00:00:00Z","timestamp":1669939200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100004837","name":"Ministerio de Ciencia e Innovaci\u00f3n, Spain","doi-asserted-by":"crossref","award":["PID2020-113656RB-C21"],"award-info":[{"award-number":["PID2020-113656RB-C21"]}],"id":[{"id":"10.13039\/501100004837","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100004837","name":"Ministerio de Ciencia e Innovaci\u00f3n, Spain","doi-asserted-by":"crossref","award":["PID2019-106455GB-C21"],"award-info":[{"award-number":["PID2019-106455GB-C21"]}],"id":[{"id":"10.13039\/501100004837","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100003359","name":"Generalitat Valenciana","doi-asserted-by":"publisher","award":["PROMETEO\/2019\/109"],"award-info":[{"award-number":["PROMETEO\/2019\/109"]}],"id":[{"id":"10.13039\/501100003359","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Comunidad de Madrid, Spain","award":["MIMACUHSPACE-CM-UC3M"],"award-info":[{"award-number":["MIMACUHSPACE-CM-UC3M"]}]},{"name":"Comunidad de Madrid, Spain","award":["MIMACUHSPACE-CM-UC3M"],"award-info":[{"award-number":["MIMACUHSPACE-CM-UC3M"]}]},{"name":"Comunidad de Madrid, Spain","award":["MIMACUHSPACE-CM-UC3M"],"award-info":[{"award-number":["MIMACUHSPACE-CM-UC3M"]}]},{"DOI":"10.13039\/501100004834","name":"Universitat Jaume I","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100004834","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J Supercomput"],"published-print":{"date-parts":[[2023,5]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Achieving maximum parallel performance on multi-core CPUs and many-core GPUs is a challenging task depending on multiple factors. These include, for example, the number and granularity of the computations or the use of the memories of the devices. In this paper, we assess those factors by evaluating and comparing different parallelizations of the same problem on a multiprocessor containing a CPU with 40 cores and four P100 GPUs with Pascal architecture. We use, as study case, the convolutional operation behind a non-standard finite element mesh truncation technique in the context of open region electromagnetic wave propagation problems. A total of six parallel algorithms implemented using OpenMP and CUDA have been used to carry out the comparison by leveraging the same levels of parallelism on both types of platforms. Three of the algorithms are presented for the first time in this paper, including a multi-GPU method, and two others are improved versions of algorithms previously developed by some of the authors. This paper presents a thorough experimental evaluation of the parallel algorithms on a radar cross-sectional prediction problem. Results show that performance obtained on the GPU clearly overcomes those obtained in the CPU, much more so if we use multiple GPUs to distribute both data and computations. Accelerations close to 30 have been obtained on the CPU, while with the multi-GPU version accelerations larger than 250 have been achieved.<\/jats:p>","DOI":"10.1007\/s11227-022-04975-6","type":"journal-article","created":{"date-parts":[[2022,12,2]],"date-time":"2022-12-02T17:03:48Z","timestamp":1670000628000},"page":"7648-7664","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":2,"title":["Strategies to parallelize a finite element mesh truncation technique on multi-core and many-core architectures"],"prefix":"10.1007","volume":"79","author":[{"given":"Jose M.","family":"Badia","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Adrian","family":"Amor-Martin","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jose A.","family":"Belloch","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Luis Emilio","family":"Garcia-Castillo","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2022,12,2]]},"reference":[{"key":"4975_CR1","volume-title":"Finite element analysis of antennas and arrays","author":"JM Jin","year":"2009","unstructured":"Jin JM, Ryley DJ (2009) Finite element analysis of antennas and arrays. Wiley, Hoboken"},{"key":"4975_CR2","doi-asserted-by":"publisher","first-page":"88","DOI":"10.1016\/j.parco.2019.02.004","volume":"85","author":"P-H Tournier","year":"2019","unstructured":"Tournier P-H, Aliferis I, Bonazzoli M, De Buhan M, Darbas M, Dolean V, Hecht F, Jolivet P, El Kanfoud I, Migliaccio C et al (2019) Microwave tomographic imaging of cerebrovascular accidents by using high-performance computing. Parallel Comput 85:88\u201397","journal-title":"Parallel Comput"},{"issue":"1","key":"4975_CR3","doi-asserted-by":"publisher","first-page":"23","DOI":"10.1190\/geo2018-0451.1","volume":"84","author":"ES Um","year":"2019","unstructured":"Um ES, Kim J, Wilt MJ, Commer M, Kim S-S (2019) Finite-element analysis of top-casing electric source method for imaging hydraulically active fracture zones. Geophysics 84(1):23\u201335","journal-title":"Geophysics"},{"key":"4975_CR4","volume-title":"Computational finite element methods in nanotechnology","author":"SM Musa","year":"2012","unstructured":"Musa SM (2012) Computational finite element methods in nanotechnology. CRC Press, Boca Raton"},{"key":"4975_CR5","doi-asserted-by":"publisher","first-page":"637","DOI":"10.1016\/j.cma.2004.05.025","volume":"194\/2\u20135","author":"LE Garc\u00eda-Castillo","year":"2005","unstructured":"Garc\u00eda-Castillo LE, G\u00f3mez-Revuelto I, S\u00e1ez de Adana F, Salazar-Palma M (2005) A finite element method for the analysis of radiation and scattering of electromagnetic waves on complex environments. Comput Methods Appl Mech Eng 194\/2\u20135:637\u2013655","journal-title":"Comput Methods Appl Mech Eng"},{"issue":"2","key":"4975_CR6","doi-asserted-by":"publisher","first-page":"104","DOI":"10.1002\/mop.21094","volume":"47","author":"I G\u00f3mez-Revuelto","year":"2005","unstructured":"G\u00f3mez-Revuelto I, Garc\u00eda-Castillo LE, Salazar-Palma M, Sarkar TK (2005) Fully coupled hybrid-method FEM\/high-frequency technique for the analysis of 3D scattering and radiation problems. Microw Opt Technol Lett 47(2):104\u2013107","journal-title":"Microw Opt Technol Lett"},{"issue":"3","key":"4975_CR7","doi-asserted-by":"publisher","first-page":"774","DOI":"10.1109\/TAP.2008.916878","volume":"56","author":"R Fern\u00e1ndez-Recio","year":"2008","unstructured":"Fern\u00e1ndez-Recio R, Garc\u00eda-Castillo LE, G\u00f3mez-Revuelto I, Salazar-Palma M (2008) Fully coupled hybrid FEM-UTD method using NURBS for the analysis of radiation problems. IEEE Trans Antennas Propag 56(3):774\u2013783","journal-title":"IEEE Trans Antennas Propag"},{"issue":"12","key":"4975_CR8","first-page":"1743","volume":"118","author":"PP Silvester","year":"1971","unstructured":"Silvester PP, Hsieh MS (1971) Finite-element solution of 2-dimensional exterior-field problems. IEE Proc (Microw Antennas Propag) 118(12):1743\u20131747","journal-title":"IEE Proc (Microw Antennas Propag)"},{"key":"4975_CR9","unstructured":"Mairal J, Koniusz P, Harchaoui Z, Schmid C (2014) Convolutional kernel networks. arXiv preprint arXiv:1406.3332"},{"key":"4975_CR10","doi-asserted-by":"crossref","unstructured":"Sharma S, Soni A, Malviya V (2019) Face recognition based on convolution neural network (CNN) applications in image processing: a survey. In: Proceedings of Recent Advances in Interdisciplinary Trends in Engineering & Applications (RAITEA)","DOI":"10.2139\/ssrn.3372193"},{"key":"4975_CR11","doi-asserted-by":"crossref","unstructured":"Saidani T, Lacassagne L, Falcou J, Tadonki C, Bouaziz S (2011) Parallelization schemes for memory optimization on the cell processor: a case study on the Harris corner detector. In: Transactions on high-performance embedded architectures and compilers III. Springer, Berlin, pp 177\u2013200","DOI":"10.1007\/978-3-642-19448-1_10"},{"key":"4975_CR12","doi-asserted-by":"crossref","unstructured":"Shi T, Belkin M, Yu B (2009) Data spectroscopy: eigenspaces of convolution operators and clustering. Ann Stat 3960\u20133984","DOI":"10.1214\/09-AOS700"},{"key":"4975_CR13","doi-asserted-by":"publisher","first-page":"67","DOI":"10.1109\/ICPP.2016.15","volume-title":"2016 45th International Conference on Parallel Processing (ICPP)","author":"X Li","year":"2016","unstructured":"Li X, Zhang G, Huang HH, Wang Z, Zheng W (2016) Performance analysis of GPU-based convolutional neural networks. In: 2016 45th International Conference on Parallel Processing (ICPP). IEEE, pp 67\u201376"},{"key":"4975_CR14","doi-asserted-by":"publisher","first-page":"102041","DOI":"10.1016\/j.sysarc.2021.102041","volume":"115","author":"S Mittal","year":"2021","unstructured":"Mittal S et al (2021) A survey of accelerator architectures for 3D convolution neural networks. J Syst Archit 115:102041","journal-title":"J Syst Archit"},{"key":"4975_CR15","doi-asserted-by":"crossref","unstructured":"Garcia-Donoro D, Amor-Martin A, Garcia-Castillo LE (2017) Higher-order finite element electromagnetics code for HPC environments. In: International Conference on Computational Science (ICCS), Zurich, Switzerland, pp 819\u2013827","DOI":"10.1016\/j.procs.2017.05.239"},{"issue":"3","key":"4975_CR16","doi-asserted-by":"publisher","first-page":"1686","DOI":"10.1007\/s11227-018-02739-9","volume":"75","author":"JA Belloch","year":"2019","unstructured":"Belloch JA, Amor-Martin A, Garcia-Donoro D, Martinez-Zaldivar FS, Garcia-Castillo LE (2019) On the use of many-core machines for the acceleration of a mesh truncation technique for FEM. J Supercomput 75(3):1686\u20131696. https:\/\/doi.org\/10.1007\/s11227-018-02739-9","journal-title":"J Supercomput"},{"key":"4975_CR17","doi-asserted-by":"publisher","first-page":"94719","DOI":"10.1109\/ACCESS.2020.2993103","volume":"8","author":"JM Bad\u00eda","year":"2020","unstructured":"Bad\u00eda JM, Amor-Martin A, Belloch JA, Garc\u00eda-Castillo LE (2020) GPU acceleration of a non-standard finite element mesh truncation technique for electromagnetics. IEEE Access 8:94719\u201394730. https:\/\/doi.org\/10.1109\/ACCESS.2020.2993103","journal-title":"IEEE Access"},{"key":"4975_CR18","unstructured":"Nvidia (2016) NVIDIA Tesla P100 Whitepaper. WP-08019-001_v01.1"},{"key":"4975_CR19","unstructured":"Melendo A, Coll A, Pasenau M, Escolano E, Monros A (2021) GiD software. www.gidhome.com. Accessed Oct 2022"}],"container-title":["The Journal of Supercomputing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11227-022-04975-6.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11227-022-04975-6\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11227-022-04975-6.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,3,23]],"date-time":"2023-03-23T09:18:22Z","timestamp":1679563102000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11227-022-04975-6"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,12,2]]},"references-count":19,"journal-issue":{"issue":"7","published-print":{"date-parts":[[2023,5]]}},"alternative-id":["4975"],"URL":"https:\/\/doi.org\/10.1007\/s11227-022-04975-6","relation":{},"ISSN":["0920-8542","1573-0484"],"issn-type":[{"value":"0920-8542","type":"print"},{"value":"1573-0484","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,12,2]]},"assertion":[{"value":"22 November 2022","order":1,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"2 December 2022","order":2,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare that they have no known conflict of interest or personal relationships that could have appeared to influence the work reported in this paper.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}},{"value":"Not applicable.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethical approval"}}]}}