{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,1]],"date-time":"2026-05-01T22:46:11Z","timestamp":1777675571151,"version":"3.51.4"},"reference-count":46,"publisher":"SAGE Publications","issue":"1","license":[{"start":{"date-parts":[[2019,10,20]],"date-time":"2019-10-20T00:00:00Z","timestamp":1571529600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/journals.sagepub.com\/page\/policies\/text-and-data-mining-license"}],"funder":[{"name":"Spanish Ministerio de Economia y Competitividad (MINECO) and the European Regional Development Fund (ERDF\/FEDER).","award":["MTM2014-52056-P"],"award-info":[{"award-number":["MTM2014-52056-P"]}]}],"content-domain":{"domain":["journals.sagepub.com"],"crossmark-restriction":true},"short-container-title":["The International Journal of High Performance Computing Applications"],"published-print":{"date-parts":[[2020,1]]},"abstract":"<jats:p>The simulation of ultrashort two-dimensional double gate metal-oxide semiconductor field-effect transistors and similar semiconductor devices through a deterministic mesoscopic, hence accurate, model can be very useful for the industry: It can provide reference results for macroscopic solvers and properly describe weakly charged zones of the device. For the scope of this work, we use a Boltzmann\u2013Schr\u00f6dinger\u2013Poisson model. Its drawback is being particularly costly from the computational point of view, and a purely sequential code may take weeks to simulate high voltages. In this article, we develop a hybrid parallel solver for a graphics processing unit (GPU)-based platform. In order to accelerate the simulations, the Boltzmann transport equations are solved on GPU using the CUDA programing model, while the Schr\u00f6dinger\u2013Poisson block is performed on multicore CPUs using OpenMP. We have adapted the costliest computing phases to the GPU in an efficient manner, achieving high performance and drastically reducing the simulation time. We give details about the parallel-design strategy and show the performance results.<\/jats:p>","DOI":"10.1177\/1094342019879985","type":"journal-article","created":{"date-parts":[[2019,10,20]],"date-time":"2019-10-20T23:11:05Z","timestamp":1571613065000},"page":"81-102","update-policy":"https:\/\/doi.org\/10.1177\/sage-journals-update-policy","source":"Crossref","is-referenced-by-count":1,"title":["Hybrid OpenMP-CUDA parallel implementation of a deterministic solver for ultrashort DG-MOSFETs"],"prefix":"10.1177","volume":"34","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-4465-938X","authenticated-orcid":false,"given":"Jos\u00e9 M","family":"Mantas","sequence":"first","affiliation":[{"name":"Departamento de Lenguajes y Sistemas Inform\u00e1ticos, Universidad de Granada, Granada, Spain"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Francesco","family":"Vecil","sequence":"additional","affiliation":[{"name":"UFR de Math\u00e9matiques, Laboratoire de Math\u00e9matiques Blaise Pascal, Universit\u00e9 Clermont Auvergne, Clermont-Ferrand, France"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"179","published-online":{"date-parts":[[2019,10,20]]},"reference":[{"key":"bibr1-1094342019879985","first-page":"1","volume":"33","author":"Abdi DS","year":"2017","journal-title":"The International Journal of High Performance Computing Applications"},{"key":"bibr2-1094342019879985","doi-asserted-by":"publisher","DOI":"10.1016\/j.jcp.2016.05.053"},{"key":"bibr3-1094342019879985","doi-asserted-by":"publisher","DOI":"10.1137\/1.9780898719604"},{"key":"bibr4-1094342019879985","doi-asserted-by":"publisher","DOI":"10.1016\/j.jcp.2009.06.001"},{"key":"bibr5-1094342019879985","doi-asserted-by":"publisher","DOI":"10.1016\/j.crma.2004.09.025"},{"key":"bibr6-1094342019879985","doi-asserted-by":"publisher","DOI":"10.1016\/j.jpdc.2012.04.003"},{"key":"bibr7-1094342019879985","doi-asserted-by":"publisher","DOI":"10.1016\/j.mcm.2012.11.007"},{"key":"bibr8-1094342019879985","doi-asserted-by":"publisher","DOI":"10.1016\/S0021-9991(02)00032-3"},{"key":"bibr9-1094342019879985","doi-asserted-by":"publisher","DOI":"10.1016\/j.crme.2010.12.004"},{"key":"bibr10-1094342019879985","volume-title":"Using OpenMP: Portable Shared Memory Parallel Programming","author":"Chapman B","year":"2008"},{"key":"bibr11-1094342019879985","doi-asserted-by":"publisher","DOI":"10.1016\/j.compfluid.2012.01.012"},{"key":"bibr12-1094342019879985","doi-asserted-by":"publisher","DOI":"10.1007\/s11227-010-0406-2"},{"key":"bibr13-1094342019879985","doi-asserted-by":"publisher","DOI":"10.1016\/j.procs.2013.05.213"},{"key":"bibr14-1094342019879985","doi-asserted-by":"publisher","DOI":"10.1145\/3017994"},{"key":"bibr15-1094342019879985","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2013.123"},{"key":"bibr16-1094342019879985","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-7091-0778-2"},{"key":"bibr17-1094342019879985","unstructured":"Intel (2019) Intel Math Kernel Library. Developer Reference. Revision 024. Available at: https:\/\/software.intel.com\/sites\/default\/files\/mkl-2019-developer-reference-c_2.pdf (accessed 15 January 2019)."},{"key":"bibr18-1094342019879985","doi-asserted-by":"publisher","DOI":"10.1016\/j.procs.2013.05.401"},{"key":"bibr19-1094342019879985","doi-asserted-by":"publisher","DOI":"10.1109\/HPCASIA.2005.74"},{"key":"bibr21-1094342019879985","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2014.2308221"},{"key":"bibr22-1094342019879985","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2016.2516988"},{"key":"bibr23-1094342019879985","unstructured":"Lukarski D, Trost N (2016) PARALUTION \u2013 User Manual, Version 1.1.0. Available at: https:\/\/www.paralution.com\/downloads\/paralution-um.pdf (accessed 15 January 2019)."},{"key":"bibr24-1094342019879985","doi-asserted-by":"publisher","DOI":"10.1016\/j.cma.2008.10.003"},{"key":"bibr25-1094342019879985","doi-asserted-by":"publisher","DOI":"10.1016\/j.mcm.2011.09.026"},{"key":"bibr26-1094342019879985","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-12165-4_36"},{"key":"bibr27-1094342019879985","doi-asserted-by":"publisher","DOI":"10.1016\/j.powtec.2016.11.061"},{"key":"bibr28-1094342019879985","unstructured":"NVIDIA (2019) CUDA Toolkit Documentation. cuSOLVER. Available at: https:\/\/developer.nvidia.com\/cusolver (accessed 15 January 2019)."},{"key":"bibr29-1094342019879985","unstructured":"NVIDIA (nd a) CUDA Toolkit Documentation. CUDA C Programming Guide. Available at: http:\/\/docs.nvidia.com\/cuda\/cuda-c-programming-guide\/ (accessed March 2018)."},{"key":"bibr30-1094342019879985","unstructured":"NVIDIA (nd b) CUDA Toolkit Documentation. Profiler User\u2019s Guide. Available at: https:\/\/docs.nvidia.com\/cuda\/profiler-users-guide\/ (accessed March 2018)."},{"key":"bibr31-1094342019879985","unstructured":"NVIDIA (nd c) CUDA Zone. Available at: https:\/\/developer.nvidia.com\/cuda-zone (accessed March 2018)."},{"key":"bibr32-1094342019879985","unstructured":"NVIDIA (2012) NVIDIA\u2019s Next Generation CUDA Compute Architecture: Kepler GK110. Available at: http:\/\/www.ece.lsu.edu\/gp\/refs\/NVIDIA-Kepler-GK110-Architecture-Whitepaper.pdf (accessed March 2018)."},{"key":"bibr33-1094342019879985","doi-asserted-by":"publisher","DOI":"10.1109\/JPROC.2008.917757"},{"issue":"1","key":"bibr34-1094342019879985","volume":"5","author":"Prasher R","year":"2013","journal-title":"Journal of Nano- and Electronic Physics"},{"key":"bibr35-1094342019879985","doi-asserted-by":"publisher","DOI":"10.1137\/S0036142998335984"},{"key":"bibr36-1094342019879985","volume-title":"Electron devices meeting (IEDM), 2011 IEEE international, 2011","author":"Rupp K","year":"2011"},{"key":"bibr37-1094342019879985","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-30397-5_13"},{"key":"bibr38-1094342019879985","doi-asserted-by":"publisher","DOI":"10.1137\/15M1026419"},{"key":"bibr39-1094342019879985","doi-asserted-by":"publisher","DOI":"10.1137\/1.9780898718003"},{"key":"bibr40-1094342019879985","doi-asserted-by":"publisher","DOI":"10.1137\/070685804"},{"key":"bibr41-1094342019879985","doi-asserted-by":"publisher","DOI":"10.1016\/j.matpr.2017.06.390"},{"key":"bibr42-1094342019879985","doi-asserted-by":"publisher","DOI":"10.15748\/jasse.2.211"},{"key":"bibr43-1094342019879985","doi-asserted-by":"publisher","DOI":"10.1109\/HPCSim.2012.6266884"},{"key":"bibr44-1094342019879985","doi-asserted-by":"publisher","DOI":"10.1080\/13873951003679017"},{"key":"bibr45-1094342019879985","doi-asserted-by":"publisher","DOI":"10.1016\/j.camwa.2014.02.021"},{"key":"bibr46-1094342019879985","doi-asserted-by":"publisher","DOI":"10.1109\/TC.2014.2366731"},{"key":"bibr47-1094342019879985","doi-asserted-by":"publisher","DOI":"10.1016\/j.compfluid.2014.06.002"}],"container-title":["The International Journal of High Performance Computing Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/1094342019879985","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/full-xml\/10.1177\/1094342019879985","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/1094342019879985","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,29]],"date-time":"2026-04-29T08:15:53Z","timestamp":1777450553000},"score":1,"resource":{"primary":{"URL":"https:\/\/journals.sagepub.com\/doi\/10.1177\/1094342019879985"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019,10,20]]},"references-count":46,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2020,1]]}},"alternative-id":["10.1177\/1094342019879985"],"URL":"https:\/\/doi.org\/10.1177\/1094342019879985","relation":{},"ISSN":["1094-3420","1741-2846"],"issn-type":[{"value":"1094-3420","type":"print"},{"value":"1741-2846","type":"electronic"}],"subject":[],"published":{"date-parts":[[2019,10,20]]}}}