{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,3]],"date-time":"2026-06-03T16:03:34Z","timestamp":1780502614862,"version":"3.54.1"},"reference-count":36,"publisher":"Springer Science and Business Media LLC","issue":"5","license":[{"start":{"date-parts":[[2023,10,23]],"date-time":"2023-10-23T00:00:00Z","timestamp":1698019200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2023,10,23]],"date-time":"2023-10-23T00:00:00Z","timestamp":1698019200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["61202076"],"award-info":[{"award-number":["61202076"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J Supercomput"],"published-print":{"date-parts":[[2024,3]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>With the rapid development of heterogeneous network-on-chip (NoC), a vast amount of shared resources are integrated into NoC. Intense resource competition exists between CPUs and GPUs, leading to congestion and a decrease in overall network performance. Reasonable node placement can minimize network conflicts at the topology level. This paper first discusses the placement of shared last-level cache and memory controller, then selects a more rational placement method and optimizes the path. To solve the hot spots problem in center placement method, a task-based routing algorithm is designed to plan the path. Simulation results demonstrate that, compared to the traditional routing algorithm, the overall network latency is reduced by 9%, and the CPU performance is improved by 13.6%. Furthermore, a dynamic task-based routing algorithm is proposed. Compared to the static task routing algorithm, the overall network latency is reduced by 2.08%, and the CPU performance is improved by 4.09%.<\/jats:p>","DOI":"10.1007\/s11227-023-05700-7","type":"journal-article","created":{"date-parts":[[2023,10,23]],"date-time":"2023-10-23T17:01:36Z","timestamp":1698080496000},"page":"6311-6335","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":15,"title":["TB-TBP: a task-based adaptive routing algorithm for network-on-chip in heterogenous CPU-GPU architectures"],"prefix":"10.1007","volume":"80","author":[{"given":"Juan","family":"Fang","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Zhichao","family":"Wei","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yaqi","family":"Liu","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yumin","family":"Hou","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2023,10,23]]},"reference":[{"key":"5700_CR1","doi-asserted-by":"publisher","first-page":"4056","DOI":"10.1007\/s11227-015-1504-y","volume":"71","author":"J Fang","year":"2015","unstructured":"Fang J, Yu L, Liu S, Lu J, Chen T (2015) Kl_ga: an application mapping algorithm for mesh-of-tree (mot) architecture in network-on-chip design. J Supercomput 71:4056\u20134071","journal-title":"J Supercomput"},{"key":"5700_CR2","doi-asserted-by":"publisher","first-page":"20","DOI":"10.1109\/MM.2012.12","volume":"32","author":"E Rotem","year":"2012","unstructured":"Rotem E, Naveh A, Ananthakrishnan A, Weissmann E, Rajwan D (2012) Power-management architecture of the intel microarchitecture code-named sandy bridge. IEEE Micro 32:20\u201327","journal-title":"IEEE Micro"},{"key":"5700_CR3","doi-asserted-by":"publisher","first-page":"175","DOI":"10.1007\/s00450-012-0209-1","volume":"28","author":"KS Lee","year":"2013","unstructured":"Lee KS, Lin H, Feng W-c (2013) Performance characterization of data-intensive kernels on amd fusion architectures. Comput Sci Res Dev 28:175\u2013184","journal-title":"Comput Sci Res Dev"},{"key":"5700_CR4","doi-asserted-by":"crossref","unstructured":"Matoussi O (2021) Noc performance model for efficient network latency estimation. In: 2021 Design, Automation & Test in Europe Conference & Exhibition (DATE), pp 994\u2013999","DOI":"10.23919\/DATE51398.2021.9474101"},{"key":"5700_CR5","doi-asserted-by":"publisher","first-page":"763","DOI":"10.1109\/TC.2013.2295523","volume":"64","author":"S Ma","year":"2015","unstructured":"Ma S, Wang Z, Liu Z, Jerger NDE (2015) Leaving one slot empty: flit bubble flow control for torus cache-coherent NoCs. IEEE Trans Comput 64:763\u2013777","journal-title":"IEEE Trans Comput"},{"key":"5700_CR6","doi-asserted-by":"publisher","first-page":"3","DOI":"10.1109\/TCAD.2008.2010691","volume":"28","author":"R Marculescu","year":"2009","unstructured":"Marculescu R, Ogras \u00dcY, Peh L-S, Jerger NDE, Hoskote YV (2009) Outstanding research problems in NoC design: system, microarchitecture, and circuit perspectives. IEEE Trans Comput Aid Des Integr Circuits Syst 28:3\u201321","journal-title":"IEEE Trans Comput Aid Des Integr Circuits Syst"},{"key":"5700_CR7","doi-asserted-by":"crossref","unstructured":"Cong J, Gill M, Hao Y, Reinman GD, Yuan B (2015) On-chip interconnection network for accelerator-rich architectures. In: 2015 52nd ACM\/EDAC\/IEEE Design Automation Conference (DAC), pp 1\u20136","DOI":"10.1145\/2744769.2744879"},{"key":"5700_CR8","doi-asserted-by":"crossref","unstructured":"Zheng H, Wang K, Louri A (2020) A versatile and flexible chiplet-based system design for heterogeneous manycore architectures. In: 2020 57th ACM\/IEEE Design Automation Conference (DAC), pp 1\u20136","DOI":"10.1109\/DAC18072.2020.9218654"},{"key":"5700_CR9","doi-asserted-by":"crossref","unstructured":"Cheng X, Zhao Y, Zhao H, Xie Y (2018) Packet pump: overcoming network bottleneck in on-chip interconnects for gpgpus*. In: 2018 55th ACM\/ESDA\/IEEE Design Automation Conference (DAC), pp 1\u20136","DOI":"10.1109\/DAC.2018.8465889"},{"key":"5700_CR10","doi-asserted-by":"crossref","unstructured":"Kim HK, Kim J, Seo W, Cho Y-G, Ryu S (2012) Providing cost-effective on-chip network bandwidth in gpgpus. In: 2012 IEEE 30th International Conference on Computer Design (ICCD), pp 407\u2013412","DOI":"10.1109\/ICCD.2012.6378671"},{"key":"5700_CR11","doi-asserted-by":"publisher","first-page":"163","DOI":"10.1109\/TPDS.2022.3217296","volume":"34","author":"YS Lee","year":"2023","unstructured":"Lee YS, Kim YW, Han TH (2023) Mrcn: throughput-oriented multicast routing for customized network-on-chips. IEEE Trans Parallel Distrib Syst 34:163\u2013179","journal-title":"IEEE Trans Parallel Distrib Syst"},{"key":"5700_CR12","doi-asserted-by":"publisher","first-page":"523","DOI":"10.1007\/s11227-021-03906-1","volume":"78","author":"E Khodadadi","year":"2021","unstructured":"Khodadadi E, Barekatain B, Yaghoubi E, Mogharrabi-Rad Z (2021) Ft-pdc: an enhanced hybrid congestion-aware fault-tolerant routing technique based on path diversity for 3d NoC. J Supercomput 78:523\u2013558","journal-title":"J Supercomput"},{"issue":"1","key":"5700_CR13","doi-asserted-by":"publisher","first-page":"113","DOI":"10.52547\/joc.15.1.113","volume":"15","author":"M Alaei","year":"2021","unstructured":"Alaei M, Yazdanpanah F (2021) A high reliable multicast routing algorithm for 2d and 3d mesh-based NoCs with fuzzy-based load control. J Control 15(1):113\u2013125","journal-title":"J Control"},{"key":"5700_CR14","first-page":"1153","volume":"67","author":"R Salamat","year":"2018","unstructured":"Salamat R, Khayambashi M, Ebrahimi M, Bagherzadeh N (2018) Lead: an adaptive 3d-NoC routing algorithm with queuing-theory based analytical verification. IEEE Trans Comput 67:1153\u20131166","journal-title":"IEEE Trans Comput"},{"key":"5700_CR15","unstructured":"Salamat R (2018) Design and evaluation of high-performance and fault-tolerant routing algorithms for 3d-NoCs"},{"key":"5700_CR16","doi-asserted-by":"crossref","unstructured":"Charles S, Mishra P (2020) Lightweight and trust-aware routing in NoC-based SoCs. In: 2020 IEEE computer society annual symposium on VLSI (ISVLSI), pp 160\u2013167","DOI":"10.1109\/ISVLSI49217.2020.00038"},{"key":"5700_CR17","doi-asserted-by":"publisher","first-page":"704","DOI":"10.1109\/TC.2017.2775643","volume":"67","author":"M Tang","year":"2018","unstructured":"Tang M, Lin J, Palesi M (2018) The suboptimal routing algorithm for 2d mesh network. IEEE Trans Comput 67:704\u2013716","journal-title":"IEEE Trans Comput"},{"key":"5700_CR18","doi-asserted-by":"crossref","unstructured":"Kao S-C, Yang C-HH, Chen P-Y, Ma X, Krishna T (2019) Reinforcement learning based interconnection routing for adaptive traffic optimization. In: Proceedings of the 13th IEEE\/ACM international symposium on networks-on-chip","DOI":"10.1145\/3313231.3352369"},{"key":"5700_CR19","doi-asserted-by":"publisher","first-page":"1525","DOI":"10.1016\/j.jpdc.2013.07.014","volume":"73","author":"J Lee","year":"2013","unstructured":"Lee J, Li S, Kim H, Yalamanchili S (2013) Design space exploration of on-chip ring interconnection for a CPU-GPU heterogeneous architecture. J Parallel Distrib Comput 73:1525\u20131538","journal-title":"J Parallel Distrib Comput"},{"key":"5700_CR20","first-page":"1","volume":"18","author":"J Lee","year":"2013","unstructured":"Lee J, Li S, Kim H, Yalamanchili S (2013) Adaptive virtual channel partitioning for network-on-chip in heterogeneous architectures. ACM Trans Des Autom Electron Syst (TODAES) 18:1\u201328","journal-title":"ACM Trans Des Autom Electron Syst (TODAES)"},{"key":"5700_CR21","doi-asserted-by":"crossref","unstructured":"Cui Y-W, Prabhakar SM, Zhao H, Mohanty SP, Fang J (2020) A low-cost conflict-free NoC architecture for heterogeneous multicore systems. In: 2020 IEEE computer society annual symposium on VLSI (ISVLSI), pp 300\u2013305","DOI":"10.1109\/ISVLSI49217.2020.00062"},{"key":"5700_CR22","doi-asserted-by":"publisher","first-page":"274","DOI":"10.1109\/TSUSC.2020.2981340","volume":"6","author":"Y Li","year":"2020","unstructured":"Li Y, Louri A (2020) Alpha: a learning-enabled high-performance network-on-chip router design for heterogeneous manycore architectures. IEEE Trans Sustain Comput 6:274\u2013288","journal-title":"IEEE Trans Sustain Comput"},{"key":"5700_CR23","doi-asserted-by":"crossref","unstructured":"Yin J, Zhou P, Sapatnekar SS, Zhai A (2014) Energy-efficient time-division multiplexed hybrid-switched noc for heterogeneous multicore systems. In: 2014 IEEE 28th international parallel and distributed processing symposium, pp 293\u2013303","DOI":"10.1109\/IPDPS.2014.40"},{"key":"5700_CR24","doi-asserted-by":"crossref","unstructured":"Zhan J, Kayiran O, Loh GH, Das CR, Xie Y (2016) Oscar: orchestrating STT-RAM cache traffic for heterogeneous CPU-GPU architectures. In: 2016 49th annual IEEE\/ACM international symposium on microarchitecture (MICRO), pp 1\u201313","DOI":"10.1109\/MICRO.2016.7783731"},{"key":"5700_CR25","doi-asserted-by":"publisher","first-page":"74","DOI":"10.1007\/s11390-015-1505-6","volume":"30","author":"J Fang","year":"2015","unstructured":"Fang J, Leng Z, Liu S, Yao Z, Sui X (2015) Exploring heterogeneous NoC design space in heterogeneous GPU-CPU architectures. J Comput Sci Technol 30:74\u201383","journal-title":"J Comput Sci Technol"},{"key":"5700_CR26","doi-asserted-by":"crossref","unstructured":"Ma S, Lu H, Huang L, Shen L, Guo Y, Wang Z, Xue W (2018) Adaptive vc partitioning for NoCs in gpgpus. In: 2018 IEEE International Symposium on Circuits and Systems (ISCAS), pp 1\u20135","DOI":"10.1109\/ISCAS.2018.8350900"},{"key":"5700_CR27","unstructured":"Kim H, Lee J, Lakshminarayana NB, Sim J, Pho T (2012) Macsim: A CPU-GPU heterogeneous simulation framework user guide. Georgia Institute of Technology, pp 1\u201357"},{"key":"5700_CR28","doi-asserted-by":"publisher","first-page":"41","DOI":"10.1109\/LES.2015.2402197","volume":"7","author":"AB Kahng","year":"2015","unstructured":"Kahng AB, Lin B, Nath S (2015) Orion3.0: a comprehensive NoC router estimation tool. IEEE Embed Syst Lett 7:41\u201345","journal-title":"IEEE Embed Syst Lett"},{"key":"5700_CR29","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/1186736.1186737","volume":"34","author":"JL Henning","year":"2006","unstructured":"Henning JL (2006) Spec cpu2006 benchmark descriptions. SIGARCH Comput Archit News 34:1\u201317","journal-title":"SIGARCH Comput Archit News"},{"key":"5700_CR30","doi-asserted-by":"crossref","unstructured":"Che S, Boyer M, Meng J, Tarjan D, Sheaffer JW, Lee S-H, Skadron K (2009) Rodinia: a benchmark suite for heterogeneous computing. In: 2009 IEEE international symposium on workload characterization (IISWC), pp 44\u201354","DOI":"10.1109\/IISWC.2009.5306797"},{"key":"5700_CR31","doi-asserted-by":"crossref","unstructured":"Stratton JA, Rodrigues CI, Sung I-J, Obeid N, Chang L-W, Anssari N, Liu G, Hwu WW (2012) Parboil: a revised benchmark suite for scientific and commercial throughput computing, pp 1\u201311","DOI":"10.1109\/InPar.2012.6339605"},{"key":"5700_CR32","doi-asserted-by":"crossref","unstructured":"Patil H, Cohn RS, Charney MJ, Kapoor R, Sun A, Karunanidhi A (2004) Pinpointing representative portions of large intel; itanium; programs with dynamic instrumentation. In: 37th International Symposium on Microarchitecture (MICRO-37\u201904), pp 81\u201392","DOI":"10.1109\/MICRO.2004.28"},{"key":"5700_CR33","doi-asserted-by":"crossref","unstructured":"Farooqui N, Kerr A, Diamos GF, Yalamanchili S, Schwan K (2011) A framework for dynamically instrumenting gpu compute applications within gpu ocelot. In: GPGPU-4","DOI":"10.1145\/1964179.1964192"},{"key":"5700_CR34","doi-asserted-by":"publisher","first-page":"729","DOI":"10.1109\/71.877831","volume":"11","author":"G-M Chiu","year":"2000","unstructured":"Chiu G-M (2000) The odd-even turn model for adaptive routing. IEEE Trans Parallel Distrib Syst 11:729\u2013738","journal-title":"IEEE Trans Parallel Distrib Syst"},{"key":"5700_CR35","doi-asserted-by":"crossref","unstructured":"An J, You H, Sun J, Cao J (2021) Fault tolerant xy-yx routing algorithm supporting backtracking strategy for noc. 2021 IEEE Intl conf on parallel & distributed processing with applications, big data & cloud computing, sustainable computing & communications, social computing & networking (ISPA\/BDCloud\/SocialCom\/SustainCom), pp 632\u2013635","DOI":"10.1109\/ISPA-BDCloud-SocialCom-SustainCom52081.2021.00092"},{"key":"5700_CR36","first-page":"93","volume":"25","author":"W Chang-shan","year":"2009","unstructured":"Chang-shan W (2009) Sd: a routing algorithm for mesh based network-on-chip. Microcomput Inf 25:93\u201394","journal-title":"Microcomput Inf"}],"container-title":["The Journal of Supercomputing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11227-023-05700-7.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11227-023-05700-7\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11227-023-05700-7.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,10,31]],"date-time":"2024-10-31T17:28:26Z","timestamp":1730395706000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11227-023-05700-7"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,10,23]]},"references-count":36,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2024,3]]}},"alternative-id":["5700"],"URL":"https:\/\/doi.org\/10.1007\/s11227-023-05700-7","relation":{},"ISSN":["0920-8542","1573-0484"],"issn-type":[{"value":"0920-8542","type":"print"},{"value":"1573-0484","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,10,23]]},"assertion":[{"value":"30 September 2023","order":1,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"23 October 2023","order":2,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare that they have no competing interests.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}},{"value":"The authors readily consent to have this paper published.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent for publication"}}]}}