{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,10]],"date-time":"2026-03-10T20:50:55Z","timestamp":1773175855006,"version":"3.50.1"},"reference-count":33,"publisher":"Association for Computing Machinery (ACM)","issue":"1","license":[{"start":{"date-parts":[[2024,3,11]],"date-time":"2024-03-11T00:00:00Z","timestamp":1710115200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"Key-Area R&D Program of Guangdong Province","award":["2021B0101190004"],"award-info":[{"award-number":["2021B0101190004"]}]},{"DOI":"10.13039\/501100001809","name":"Programs of National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62321003, U21A20461, 62172157, 92055213, 62227808"],"award-info":[{"award-number":["62321003, U21A20461, 62172157, 92055213, 62227808"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Major Projects of Xiangjiang Laboratory","award":["22xj01011"],"award-info":[{"award-number":["22xj01011"]}]},{"name":"Key R&D Program of Hunan Province","award":["2023GK2002"],"award-info":[{"award-number":["2023GK2002"]}]},{"name":"Shenzhen Science and Technology Program","award":["JCYJ20210324135409026"],"award-info":[{"award-number":["JCYJ20210324135409026"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Parallel Comput."],"published-print":{"date-parts":[[2024,3,31]]},"abstract":"<jats:p>Thanks to the recognition and promotion of chiplet-based High-Performance Computing (HPC) system design technology by semiconductor industry\/market leaders, chiplet-based multi-chip systems have gradually become the mainstream. Unfortunately, programming such systems to achieve efficient computing is a challenge, especially when considering dynamic task parallelism. This paper presents an Adaptive Batch-Stream Scheduling (ABSS) module for dynamic task parallelism on chiplet-based multi-chip systems. To this end, we propose an adaptive batch-stream scheduling method based on Graph Convolution Network (GCN) classifier to select the appropriate scheduling scheme. We further design a chiplet-based core-cluster binding mechanism, which establishes the affinity between threads and core-clusters on CPU-compute die. Moreover, to achieve dynamic workload balance, we propose a chiplet-based nearest task stealing method. We implement our ABSS module on the HiSilicon Kunpeng-920 chiplet-based multi-chip system. Experiments show that it outperforms state-of-the-art parallelism solutions, such as Intel Threading Building Blocks.<\/jats:p>","DOI":"10.1145\/3643597","type":"journal-article","created":{"date-parts":[[2024,1,29]],"date-time":"2024-01-29T12:38:54Z","timestamp":1706531934000},"page":"1-24","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":4,"title":["ABSS: An Adaptive Batch-Stream Scheduling Module for Dynamic Task Parallelism on Chiplet-based Multi-Chip Systems"],"prefix":"10.1145","volume":"11","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-7642-457X","authenticated-orcid":false,"given":"Qinyun","family":"Cai","sequence":"first","affiliation":[{"name":"Hunan University, Changsha, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5008-4829","authenticated-orcid":false,"given":"Guoqing","family":"Xiao","sequence":"additional","affiliation":[{"name":"Hunan University, Changsha, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3329-0924","authenticated-orcid":false,"given":"Shengle","family":"Lin","sequence":"additional","affiliation":[{"name":"Hunan University, Changsha, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2681-7898","authenticated-orcid":false,"given":"Wangdong","family":"Yang","sequence":"additional","affiliation":[{"name":"Hunan University, Changsha, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5224-4048","authenticated-orcid":false,"given":"Keqin","family":"Li","sequence":"additional","affiliation":[{"name":"State University of New York, NY, USA and Hunan University, Changsha, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2635-7716","authenticated-orcid":false,"given":"Kenli","family":"Li","sequence":"additional","affiliation":[{"name":"Hunan University, Changsha, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2024,3,11]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.1145\/2442516.2442538"},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1002\/(SICI)1099-1506(199607\/08)3:4<275::AID-NLA83>3.0.CO;2-7"},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2008.105"},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2008.105"},{"key":"e_1_3_1_6_2","doi-asserted-by":"publisher","DOI":"10.1145\/3399730"},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.1145\/209937.209958"},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10766-010-0136-3"},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1145\/3316781.3317771"},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.1145\/3399728"},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-09873-9_50"},{"key":"e_1_3_1_12_2","article-title":"Kunpeng Math Library","author":"CO. Huawei Technologies","year":"2023","unstructured":"Huawei Technologies CO.2023. Kunpeng Math Library. (2023). http:\/\/www.hikunpeng.com\/developer\/boostkit\/library\/math","journal-title":"("},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1145\/2499368.2451157"},{"issue":"1","key":"e_1_3_1_14_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/2049662.2049663","article-title":"The University of Florida sparse matrix collection","volume":"38","author":"Davis Timothy A.","year":"2011","unstructured":"Timothy A. Davis and Yifan Hu. 2011. The University of Florida sparse matrix collection. ACM Trans. Math. Software 38, 1 (2011), 1\u201325.","journal-title":"ACM Trans. Math. Software"},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11227-010-0405-3"},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.1145\/3458817.3476199"},{"key":"e_1_3_1_17_2","doi-asserted-by":"crossref","first-page":"308","DOI":"10.1145\/232973.233006","volume-title":"ISCA\u201996","author":"Lovett T.","year":"1996","unstructured":"T. Lovett and R. Clapp. 1996. STiNG: A CC-NUMA computer system for the commercial marketplace. In ISCA\u201996. 308\u2013317."},{"key":"e_1_3_1_18_2","article-title":"Introduction to Intel QuickPath Interconnect","author":"Maddox R.","year":"2009","unstructured":"R. Maddox and R. J. Safranek. 2009. Introduction to Intel QuickPath Interconnect. High Performance Multi-Core Processor Fabric (2009).","journal-title":"High Performance Multi-Core Processor Fabric"},{"key":"e_1_3_1_19_2","first-page":"342","volume-title":"ISC\u201919","author":"Popov M.","year":"2019","unstructured":"M. Popov and A. Jimborean. 2019. Efficient thread\/page\/parallelism autotuning for NUMA systems. In ISC\u201919. 342\u2013353."},{"key":"e_1_3_1_20_2","volume-title":"Intel Threading Building Blocks: Outfitting C++ for Multi-core Processor Parallelism","author":"Reinders J.","year":"2007","unstructured":"J. Reinders. 2007. Intel Threading Building Blocks: Outfitting C++ for Multi-core Processor Parallelism. O\u2019Reilly Media, Inc."},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.sysarc.2022.102393"},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.1145\/3453417.3453419"},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.1145\/2938389"},{"key":"e_1_3_1_24_2","first-page":"102","volume-title":"IWOMP\u201916","year":"2016","unstructured":"Christian Terboven, Jonas Hahnfeld, Xavier Teruel, Sergi Mateo, Alejandro Duran, Michael Klemm, Stephen L. Olivier, and Bronis R. de Supinski. 2016. Approaches for task affinity in OpenMP. In IWOMP\u201916. Springer, 102\u2013115."},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1145\/1837853.1693479"},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.1145\/3584373"},{"key":"e_1_3_1_27_2","first-page":"173","volume-title":"ISCA\u201920","year":"2020","unstructured":"MoyangWang, Tuan Ta, Lin Cheng, and Christopher Batten. 2020. Efficiently supporting dynamic task parallelism on heterogeneous cache-coherent systems. In ISCA\u201920. IEEE, 173\u2013186."},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA53966.2022.00091"},{"key":"e_1_3_1_29_2","article-title":"Automatically tuned linear algebra software","author":"Whaley R.","year":"1998","unstructured":"R. Whaley and J. Dontarra. 1998. Automatically tuned linear algebra software. IEEE (1998).","journal-title":"IEEE"},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.1145\/3145617.3145620"},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA53966.2022.00076"},{"issue":"99","key":"e_1_3_1_32_2","first-page":"1","article-title":"Kunpeng 920: The first 7nm chiplet-based 64-core ARM SoC for cloud services","year":"2021","unstructured":"Jing Xia, Chuanning Cheng, Xiping Zhou, Yuxing Hu, and Peter Chun. 2021. Kunpeng 920: The first 7nm chiplet-based 64-core ARM SoC for cloud services. IEEE Micro PP, 99 (2021), 1\u20131.","journal-title":"IEEE Micro"},{"key":"e_1_3_1_33_2","first-page":"1","article-title":"Parallel algorithm design and optimization of geodynamic numerical simulation application on the Tianhe new-generation high-performance computer","author":"Yang Jin","year":"2023","unstructured":"Jin Yang, Wangdong Yang, Ruixuan Qi, Qinyun Tsai, Shengle Lin, Fengkun Dong, Kenli Li, and Keqin Li. 2023. Parallel algorithm design and optimization of geodynamic numerical simulation application on the Tianhe new-generation high-performance computer. The Journal of Supercomputing (2023), 1\u201332.","journal-title":"The Journal of Supercomputing"},{"key":"e_1_3_1_34_2","first-page":"1","volume-title":"SC\u201920","year":"2020","unstructured":"Di Zhang, Dong Dai, Youbiao He, Forrest Sheng Bao, and Bing Xie. 2020. RLScheduler: An automated HPC batch job scheduler using reinforcement learning. In SC\u201920. IEEE, 1\u201315."}],"container-title":["ACM Transactions on Parallel Computing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3643597","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3643597","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T22:50:28Z","timestamp":1750287028000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3643597"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,3,11]]},"references-count":33,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2024,3,31]]}},"alternative-id":["10.1145\/3643597"],"URL":"https:\/\/doi.org\/10.1145\/3643597","relation":{},"ISSN":["2329-4949","2329-4957"],"issn-type":[{"value":"2329-4949","type":"print"},{"value":"2329-4957","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,3,11]]},"assertion":[{"value":"2023-07-25","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-01-23","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-03-11","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}