{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,27]],"date-time":"2026-06-27T14:47:51Z","timestamp":1782571671350,"version":"3.54.5"},"reference-count":59,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2026,6,27]],"date-time":"2026-06-27T00:00:00Z","timestamp":1782518400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Archit. Code Optim."],"published-print":{"date-parts":[[2026,6,30]]},"abstract":"<jats:p>Initially branded as dedicated graphics processing accelerators, GPUs now find applications in an ever-growing range of domains, including artificial intelligence, high-performance computing, self-driving vehicles, and bioinformatics. However, this diversity comes at the cost of reduced resource efficiency and micro-architectural affinity. Evidently, the homogeneity of the GPU hardware struggles to cope with the vast heterogeneity of GPU applications. Motivated by the aforementioned observations, this article introduces the concept of Single-ISA Heterogeneous GPU architectures. In order to explore the efficiency of the new GPU architectural paradigm, we extend Accel-Sim, the state-of-the-art, cycle-accurate GPU simulator to support single-ISA heterogeneous cores within the GPU chip. The proposed implementation, called AccelHSA, supports independently tuning the micro-architectural characteristics of the cores, unlocking a wide design space. The CUDA API is extended to allow control of the kernel-to-core-type mapping along with a newly developed kernel launching model that supports concurrent execution, aimed at, albeit not limited to, the context of the simulator. We showcase the impact of single-ISA heterogeneous GPU architectures via a case study targeting the collocation of resource sensitive and insensitive HPC kernels. Finally, the heterogeneous GPU architecture is evaluated against homogeneous GPU baselines, demonstrating a 27.07% average speedup with a marginal 0.47% area overhead.<\/jats:p>","DOI":"10.1145\/3811410","type":"journal-article","created":{"date-parts":[[2026,4,21]],"date-time":"2026-04-21T11:11:18Z","timestamp":1776769878000},"page":"1-25","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["AccelHSA: Modeling Single-ISA Heterogeneous GPU Architectures"],"prefix":"10.1145","volume":"23","author":[{"ORCID":"https:\/\/orcid.org\/0009-0001-3391-1009","authenticated-orcid":false,"given":"Alexandros","family":"Moiras","sequence":"first","affiliation":[{"name":"Electrical and Computer Engineering, National Technical University of Athens","place":["Zografou, Greece"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1403-6851","authenticated-orcid":false,"given":"Konstantinos","family":"Iliakis","sequence":"additional","affiliation":[{"name":"Electrical and Computer Engineering, National Technical University of Athens","place":["Zografou, Greece"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6930-6847","authenticated-orcid":false,"given":"Dimitrios","family":"Soudris","sequence":"additional","affiliation":[{"name":"Electrical and Computer Engineering, National Technical University of Athens","place":["Zografou, Greece"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3151-2730","authenticated-orcid":false,"given":"Sotirios","family":"Xydis","sequence":"additional","affiliation":[{"name":"Electrical and Computer Engineering, National Technical University of Athens","place":["Zografou, Greece"]}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,6,27]]},"reference":[{"key":"e_1_3_2_2_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2012.6168946"},{"key":"e_1_3_2_3_2","doi-asserted-by":"crossref","unstructured":"Nauman Ahmed Jonathan L\u00e9vy Shanshan Ren Hamid Mushtaq Koen Bertels and Zaid Al-Ars. 2019. GASAL2: A GPU accelerated sequence alignment library for high-throughput NGS data. BMC Bioinformatics 20 1 (2019) 520.","DOI":"10.1186\/s12859-019-3086-9"},{"key":"e_1_3_2_4_2","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2020.3013728"},{"key":"e_1_3_2_5_2","unstructured":"ARM limited. 2013. big.LITTLE Technology: The Future of Mobile: Making Very High Performance Available in a Mobile Envelope Without Sacrificing Energy Efficiency. ARM Limited. Retrieved from https:\/\/armkeil.blob.core.windows.net\/developer\/Files\/pdf\/white-paper\/big-little-technology-the-future-of-mobile.pdf"},{"issue":"1","key":"e_1_3_2_6_2","first-page":"1","article-title":"R-gpu: A reconfigurable gpu architecture","volume":"13","author":"Braak Gert-Jan Van Den","year":"2016","unstructured":"Gert-Jan Van Den Braak and Henk Corporaal. 2016. R-gpu: A reconfigurable gpu architecture. ACM Transactions on Architecture and Code Optimization (TACO) 13, 1 (2016), 1\u201324.","journal-title":"ACM Transactions on Architecture and Code Optimization (TACO)"},{"key":"e_1_3_2_7_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.jpdc.2012.04.003"},{"key":"e_1_3_2_8_2","doi-asserted-by":"crossref","unstructured":"Pablo Carvalho Esteban Clua Aline Paes Cristiana Bentes Bruno Lopes and L\u00facia Maria de A. Drummond. 2020. Using machine learning techniques to analyze the performance of concurrent kernel execution on GPUs. Future Generation Computer Systems 113 0167\u2013739X (2020) 528\u2013540.","DOI":"10.1016\/j.future.2020.07.038"},{"key":"e_1_3_2_9_2","doi-asserted-by":"publisher","unstructured":"R. D. Hornung J. A. Keasler and M. B. Gokhale. 2011. Hydrodynamics Challenge Problem. Lawrence Livermore National Lab. (LLNL) Livermore CA (United States). DOI:10.2172\/1117905","DOI":"10.2172\/1117905"},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA59077.2024.00011"},{"key":"e_1_3_2_11_2","doi-asserted-by":"publisher","DOI":"10.1109\/IISWC.2009.5306797"},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.1145\/3392717.3392738"},{"key":"e_1_3_2_13_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-0348-8534-8_21"},{"key":"e_1_3_2_14_2","doi-asserted-by":"publisher","DOI":"10.1145\/3731569.3764818"},{"key":"e_1_3_2_15_2","doi-asserted-by":"crossref","unstructured":"Rommel Cruz Lucia Drummond Esteban Clua and Cristiana Bentes. 2017. Analyzing and estimating the performance of concurrent kernels execution on GPUs. In XVIII Simp\u00f3sio em Sistemas Computacionais de Alto Desempenho. 136\u2013147.","DOI":"10.5753\/wscad.2017.245"},{"key":"e_1_3_2_16_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2018.00027"},{"key":"e_1_3_2_17_2","doi-asserted-by":"publisher","DOI":"10.1145\/3777466"},{"issue":"1","key":"e_1_3_2_18_2","first-page":"1","article-title":"GPU domain specialization via composable on-package architecture","volume":"19","author":"Fu Yaosheng","year":"2021","unstructured":"Yaosheng Fu, Evgeny Bolotin, Niladrish Chatterjee, David Nellans, and Stephen W. Keckler. 2021. GPU domain specialization via composable on-package architecture. ACM Transactions on Architecture and Code Optimization (TACO) 19, 1 (2021), 1\u201323.","journal-title":"ACM Transactions on Architecture and Code Optimization (TACO)"},{"key":"e_1_3_2_19_2","doi-asserted-by":"publisher","DOI":"10.1109\/InPar.2012.6339595"},{"key":"e_1_3_2_20_2","article-title":"Groq rocks neural networks","author":"Gwennap Linley","year":"2020","unstructured":"Linley Gwennap. 2020. Groq rocks neural networks. Microprocessor Report, Tech. Rep., jan (2020).","journal-title":"Microprocessor Report, Tech. Rep., jan"},{"key":"e_1_3_2_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2022.3192707"},{"key":"e_1_3_2_22_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2021.3093231"},{"key":"e_1_3_2_23_2","doi-asserted-by":"publisher","DOI":"10.1109\/CGO.2015.7054182"},{"key":"e_1_3_2_24_2","doi-asserted-by":"publisher","DOI":"10.1145\/2451116.2451158"},{"key":"e_1_3_2_25_2","doi-asserted-by":"publisher","DOI":"10.1145\/2818950.2818979"},{"key":"e_1_3_2_26_2","doi-asserted-by":"publisher","DOI":"10.1145\/3079856.3080246"},{"key":"e_1_3_2_27_2","doi-asserted-by":"publisher","DOI":"10.1145\/3676641.3715996"},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.1145\/3466752.3480063"},{"key":"e_1_3_2_29_2","doi-asserted-by":"publisher","DOI":"10.2172\/1059462"},{"key":"e_1_3_2_30_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA45697.2020.00047"},{"key":"e_1_3_2_31_2","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2003.1253185"},{"key":"e_1_3_2_32_2","doi-asserted-by":"publisher","DOI":"10.1145\/1028176.1006707"},{"key":"e_1_3_2_33_2","doi-asserted-by":"publisher","DOI":"10.1145\/2508148.2485964"},{"key":"e_1_3_2_34_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.jksuci.2023.101656"},{"key":"e_1_3_2_35_2","doi-asserted-by":"publisher","DOI":"10.1145\/1669112.1669172"},{"key":"e_1_3_2_36_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICPP.2011.88"},{"key":"e_1_3_2_37_2","doi-asserted-by":"publisher","DOI":"10.1145\/2925426.2926267"},{"key":"e_1_3_2_38_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISPASS57527.2023.00026"},{"key":"e_1_3_2_39_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISBI.2008.4541126"},{"key":"e_1_3_2_40_2","doi-asserted-by":"publisher","DOI":"10.1002\/widm.53"},{"key":"e_1_3_2_41_2","doi-asserted-by":"publisher","DOI":"10.5555\/2523721.2523743"},{"key":"e_1_3_2_42_2","unstructured":"NVIDIA Corporation. 2017. NVIDIA Tesla V100 GPU Architecture. NVIDIA Corporation. Retrieved April 30 2026 from https:\/\/images.nvidia.com\/content\/volta-architecture\/pdf\/volta-architecture-whitepaper.pdf"},{"key":"e_1_3_2_43_2","unstructured":"NVIDIA Corporation. 2022. NVIDIA Multi-Instance GPU User Guide. NVIDIA Corporation. Retrieved April 30 2026 from https:\/\/docs.nvidia.com\/datacenter\/tesla\/pdf\/NVIDIA_MIG_User_Guide.pdf"},{"key":"e_1_3_2_44_2","unstructured":"NVIDIA Corporation. 2024. NVIDIA Multi-Process Service. NVIDIA Corporation. Retrieved April 30 2026 from https:\/\/docs.nvidia.com\/deploy\/pdf\/CUDA_Multi_Process_Service_Overview.pdf"},{"key":"e_1_3_2_45_2","unstructured":"Justin Luitjens. 2014. CUDA streams: Best practices and common pitfalls. Retrieved April 30 2026 from https:\/\/www.hpcadmintech.com\/wp-content\/uploads\/2016\/03\/Carlo_Nardone_presentation.pdf"},{"key":"e_1_3_2_46_2","doi-asserted-by":"publisher","DOI":"10.1145\/2490301.2451160"},{"key":"e_1_3_2_47_2","doi-asserted-by":"crossref","unstructured":"German I. Parisi Ronald Kemker Jose L. Part Christopher Kanan and Stefan Wermter. 2019. Continual lifelong learning with neural networks: A review. Neural Networks 113 0893\u20136080 (2019) 54\u201371.","DOI":"10.1016\/j.neunet.2019.01.012"},{"key":"e_1_3_2_48_2","doi-asserted-by":"publisher","DOI":"10.1145\/3037697.3037707"},{"key":"e_1_3_2_49_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.parco.2014.09.011"},{"key":"e_1_3_2_50_2","doi-asserted-by":"publisher","DOI":"10.1145\/1996130.1996160"},{"key":"e_1_3_2_51_2","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2012.16"},{"key":"e_1_3_2_52_2","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2022.3164338"},{"key":"e_1_3_2_53_2","first-page":"241","article-title":"Propagation of strong shock waves","volume":"10","author":"Sedov Leonid Ivanovich","year":"1946","unstructured":"Leonid Ivanovich Sedov. 1946. Propagation of strong shock waves. Journal of Applied Mathematics and Mechanics 10 (1946), 241\u2013250.","journal-title":"Journal of Applied Mathematics and Mechanics"},{"key":"e_1_3_2_54_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2018.2890150"},{"key":"e_1_3_2_55_2","doi-asserted-by":"publisher","DOI":"10.1103\/PhysRevAccelBeams.26.114602"},{"key":"e_1_3_2_56_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCSim.2011.5999803"},{"key":"e_1_3_2_57_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2016.7446078"},{"key":"e_1_3_2_58_2","doi-asserted-by":"publisher","DOI":"10.1145\/3377138"},{"key":"e_1_3_2_59_2","doi-asserted-by":"publisher","DOI":"10.1145\/3007787.3001161"},{"key":"e_1_3_2_60_2","doi-asserted-by":"publisher","DOI":"10.1145\/3155284.3018754"}],"container-title":["ACM Transactions on Architecture and Code Optimization"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3811410","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,27]],"date-time":"2026-06-27T14:16:47Z","timestamp":1782569807000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3811410"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6,27]]},"references-count":59,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2026,6,30]]}},"alternative-id":["10.1145\/3811410"],"URL":"https:\/\/doi.org\/10.1145\/3811410","relation":{},"ISSN":["1544-3566","1544-3973"],"issn-type":[{"value":"1544-3566","type":"print"},{"value":"1544-3973","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,6,27]]},"assertion":[{"value":"2025-07-10","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-04-01","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-06-27","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}