{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,2]],"date-time":"2026-07-02T04:55:43Z","timestamp":1782968143594,"version":"3.54.5"},"reference-count":71,"publisher":"Association for Computing Machinery (ACM)","issue":"1","license":[{"start":{"date-parts":[[2025,3,20]],"date-time":"2025-03-20T00:00:00Z","timestamp":1742428800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["60973010"],"award-info":[{"award-number":["60973010"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Archit. Code Optim."],"published-print":{"date-parts":[[2025,3,31]]},"abstract":"<jats:p>Innovative processor architecture designs are shifting towards Many-Core Architectures (MCAs) to meet the future demands of high-performance computing as the limits of Moore\u2019s Law have almost been reached. Many-core processors utilize shared memory hierarchies to achieve high-speed memory systems, improving memory access efficiency. However, as the number of cores multiplies, the scalability of this system is significantly constrained by the increased proportion of long-distance and Non-Uniform Memory Access (NUMA). Improving the scalability of MCAs is crucial for achieving large\/super-scale general-purpose many-core processors. This work proposes a high-scalability memory Network-on-Chip (NoC) for Triplet-Based Many-Core Architecture (TriBA), named TriBA-mNoC. TriBA-mNoC maintains a consistent core-to-core spacing as the network scale increases, effectively preventing increased long-distance memory access latency. Moreover, it leverages an inherent advantage of shared-inside hierarchical-groupings, alleviating common NUMA issues in the NoC design. Evaluations of static network characteristics show that TriBA-mNoC outperforms most classical NoCs in network diameter, average distance, and cost. TriBA-mNoC can be integrated with TriBA in the same silicon die with a tile-like floorplan, forming a novel NoC called TriBA-NoC, which can combine the strengths of both networks to maximize the architecture performance. We evaluated the memory access performance and scalability of TriBA-NoC using the mathematical evaluation models and actual simulations with real traffic (PARSEC 3.0 and SPLASH-2) at different network scales. The mathematical evaluation results indicate that TriBA-NoC achieves an aggregate speedup of approximately 3x compared with 2D-Mesh for a similar number of cores. Furthermore, TriBA-NoC\u2019s single-core speedup efficiency remains stable as the number of cores increases under the same cache hit ratio, whereas 2D-Mesh experiences a rapid decline, highlighting TriBA-NoC\u2019s exceptional scalability. Finally, the actual traffic simulation results show that TriBA-NoC achieves an average memory access latency and time reduction of 25.90% to 40.50% and 5.61% to 31.69%, respectively, compared with 2D-Mesh.<\/jats:p>","DOI":"10.1145\/3688610","type":"journal-article","created":{"date-parts":[[2024,11,2]],"date-time":"2024-11-02T08:46:28Z","timestamp":1730537188000},"page":"1-26","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":2,"title":["A High Scalability Memory NoC with Shared-Inside Hierarchical-Groupings for Triplet-Based Many-Core Architecture"],"prefix":"10.1145","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-8865-9744","authenticated-orcid":false,"given":"Chunfeng","family":"Li","sequence":"first","affiliation":[{"name":"School of Computer Science and Technology, Beijing Institute of Technology, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5175-9760","authenticated-orcid":false,"given":"Feng","family":"Shi","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Beijing Institute of Technology, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2207-6295","authenticated-orcid":false,"given":"Fei","family":"Yin","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Beijing Institute of Technology, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5296-9743","authenticated-orcid":false,"given":"Karim","family":"Soliman","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Beijing Institute of Technology, Alexandria, Egypt"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2665-278X","authenticated-orcid":false,"given":"Jin","family":"Wei","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Beijing Institute of Technology, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,3,20]]},"reference":[{"key":"e_1_3_2_2_2","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2007.4378779"},{"key":"e_1_3_2_3_2","doi-asserted-by":"publisher","DOI":"10.1049\/iet-cdt.2018.5220"},{"key":"e_1_3_2_4_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.vlsi.2022.06.014"},{"key":"e_1_3_2_5_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11227-022-04910-9"},{"key":"e_1_3_2_6_2","doi-asserted-by":"publisher","DOI":"10.5772\/intechopen.97262"},{"key":"e_1_3_2_7_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10836-023-06046-x"},{"key":"e_1_3_2_8_2","doi-asserted-by":"publisher","DOI":"10.2174\/2352096510666170425102503"},{"key":"e_1_3_2_9_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCITechn.2014.6997341"},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPSW.2010.5470784"},{"key":"e_1_3_2_11_2","doi-asserted-by":"publisher","DOI":"10.1145\/2591635.2667187"},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.1109\/TVLSI.2008.2003999"},{"key":"e_1_3_2_13_2","doi-asserted-by":"publisher","DOI":"10.1109\/IISWC.2009.5306792"},{"key":"e_1_3_2_14_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.proeng.2012.01.423"},{"key":"e_1_3_2_15_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCAD.2021.3091961"},{"key":"e_1_3_2_16_2","doi-asserted-by":"publisher","DOI":"10.1109\/APCCAS.2010.5774875"},{"key":"e_1_3_2_17_2","doi-asserted-by":"publisher","DOI":"10.1109\/ASICON.2009.5351597"},{"key":"e_1_3_2_18_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCC\/SmartCity\/DSS.2018.00103"},{"key":"e_1_3_2_19_2","unstructured":"Tilera Corporation. 2012. Tile Processor Architecture Overview for the Tile-GX Series. Release 0.20 Doc.No. UG130.https:\/\/cdn.manesht.ir\/17871___210769647-UG130-ArchOverview-TILE-Gx.pdf"},{"key":"e_1_3_2_20_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2009.4798252"},{"key":"e_1_3_2_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA56546.2023.10070981"},{"key":"e_1_3_2_22_2","doi-asserted-by":"publisher","DOI":"10.1109\/NEWCAS.2009.5290419"},{"key":"e_1_3_2_23_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-14313-2_21"},{"key":"e_1_3_2_24_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISSCC42614.2022.9731673"},{"key":"e_1_3_2_25_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2009.4798251"},{"key":"e_1_3_2_26_2","doi-asserted-by":"publisher","DOI":"10.1049\/iet-cdt.2016.0184"},{"key":"e_1_3_2_27_2","first-page":"70","volume-title":"Computer Architecture: A Quantitative Approach","author":"Hennessy John L.","year":"2011","unstructured":"John L. Hennessy and David A. Patterson. 2011. Computer Architecture: A Quantitative Approach. Elsevier, 70\u201376."},{"key":"e_1_3_2_28_2","first-page":"14","volume-title":"1st International Workshop on Network on Chip Architectures (NoCArc 2008)","author":"Hu Wen-Hsiang","year":"2008","unstructured":"Wen-Hsiang Hu, Seung Eun Lee, and Nader Bagherzadeh. 2008. DMesh: A diagonally-linked mesh network-on-chip architecture. In 1st International Workshop on Network on Chip Architectures (NoCArc 2008). 14\u201320."},{"key":"e_1_3_2_29_2","doi-asserted-by":"publisher","DOI":"10.1109\/DAC56929.2023.10248006"},{"key":"e_1_3_2_30_2","unstructured":"Intel. 2022. Intel\/microarchitectures\/skylake (client). Retrieved from https:\/\/en.wikichip.org\/wiki\/intel\/microarchitectures\/skylake_(client)#Memory_Hierarchy"},{"key":"e_1_3_2_31_2","doi-asserted-by":"publisher","DOI":"10.3390\/w15152810"},{"key":"e_1_3_2_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2007.29"},{"key":"e_1_3_2_33_2","doi-asserted-by":"publisher","DOI":"10.1145\/1273440.1250681"},{"key":"e_1_3_2_34_2","doi-asserted-by":"publisher","DOI":"10.1109\/71.113081"},{"key":"e_1_3_2_35_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA.2005.34"},{"key":"e_1_3_2_36_2","doi-asserted-by":"publisher","DOI":"10.1142\/S0218126623500767"},{"key":"e_1_3_2_37_2","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPSW.2013.154"},{"key":"e_1_3_2_38_2","doi-asserted-by":"publisher","DOI":"10.1145\/2508834.2513149"},{"key":"e_1_3_2_39_2","doi-asserted-by":"publisher","DOI":"10.1109\/NAS.2009.48"},{"key":"e_1_3_2_40_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2019.2895701"},{"key":"e_1_3_2_41_2","doi-asserted-by":"publisher","DOI":"10.1109\/MCSoC.2016.40"},{"key":"e_1_3_2_42_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11227-019-03072-5"},{"key":"e_1_3_2_43_2","doi-asserted-by":"publisher","DOI":"10.1038\/s41598-023-34297-3"},{"key":"e_1_3_2_44_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA52012.2021.00014"},{"key":"e_1_3_2_45_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.jpdc.2017.10.011"},{"key":"e_1_3_2_46_2","doi-asserted-by":"publisher","DOI":"10.1109\/HOTCHIPS.2011.7477491"},{"key":"e_1_3_2_47_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.micpro.2014.12.002"},{"key":"e_1_3_2_48_2","doi-asserted-by":"publisher","DOI":"10.1145\/1736065.1736069"},{"key":"e_1_3_2_49_2","unstructured":"Intel SDG. 2014. Intel\u00ae Xeon Phi\u2122 Coprocessor System Software Developers Guide. Retrieved from https:\/\/engineering.purdue.edu\/eigenman\/ECE563\/Handouts\/intel-xeon-phi-systemsoftwaredevelopersguide.pdf"},{"key":"e_1_3_2_50_2","doi-asserted-by":"publisher","DOI":"10.1360\/N112017-00065"},{"key":"e_1_3_2_51_2","unstructured":"Feng Shi ShengQiang Ruan and XiaoJun Wang. 2021. Deadlock Prevention Method for Base-Three Inter-core Networks Based on Cutting-edge Subgraph Classification of Transmitted Data. Patent CN112698960A. Beijing Institute of Technology China."},{"key":"e_1_3_2_52_2","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2016.25"},{"key":"e_1_3_2_53_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.jpdc.2009.05.002"},{"key":"e_1_3_2_54_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.comcom.2011.05.002"},{"key":"e_1_3_2_55_2","first-page":"139","volume-title":"International Conference on Networks (ICN\u201911)","author":"Thamarakuzhi Ajithkumar","year":"2011","unstructured":"Ajithkumar Thamarakuzhi and John A. Chandy. 2011. Adaptive load balanced routing for 2-dilated flattened butterfly switching network. In International Conference on Networks (ICN\u201911). 139\u2013244."},{"key":"e_1_3_2_56_2","volume-title":"CACTI 5.1","author":"Thoziyoor Shyamkumar","year":"2008","unstructured":"Shyamkumar Thoziyoor, Naveen Muralimanohar, Jung Ho Ahn, and Norman P. Jouppi. 2008. CACTI 5.1. Technical Report. Technical Report HPL-2008-20, HP Labs."},{"key":"e_1_3_2_57_2","doi-asserted-by":"publisher","DOI":"10.1145\/2370816.2370865"},{"key":"e_1_3_2_58_2","doi-asserted-by":"publisher","DOI":"10.36079\/lamintang.ijortas-0502.515"},{"key":"e_1_3_2_59_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10617-022-09266-0"},{"key":"e_1_3_2_60_2","doi-asserted-by":"publisher","DOI":"10.1002\/dac.5360"},{"key":"e_1_3_2_61_2","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2011.24"},{"key":"e_1_3_2_62_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2010.5416628"},{"key":"e_1_3_2_63_2","doi-asserted-by":"publisher","DOI":"10.1145\/225830.223990"},{"key":"e_1_3_2_64_2","doi-asserted-by":"publisher","DOI":"10.1145\/216585.216588"},{"key":"e_1_3_2_65_2","doi-asserted-by":"publisher","DOI":"10.5555\/2842825"},{"issue":"9","key":"e_1_3_2_66_2","doi-asserted-by":"crossref","first-page":"2194","DOI":"10.1360\/jos182194","article-title":"Xmesh: A mesh-like topology for network on chip","volume":"18","author":"Zhu X. J.","year":"2007","unstructured":"X. J. Zhu, W. W. Hu, K. Ma, and L. B. Zhang. 2007. Xmesh: A mesh-like topology for network on chip. Journal of Software 18, 9 (2007), 2194\u20132204. http:\/\/www.jos.org.cn\/1000-9825\/18\/2194.htm","journal-title":"Journal of Software"},{"key":"e_1_3_2_67_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.micpro.2011.04.001"},{"key":"e_1_3_2_68_2","doi-asserted-by":"publisher","DOI":"10.1145\/3053277.3053279"},{"key":"e_1_3_2_69_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-84522-3_7"},{"key":"e_1_3_2_70_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2023.3317296"},{"key":"e_1_3_2_71_2","doi-asserted-by":"publisher","DOI":"10.1520\/JTE20120212"},{"key":"e_1_3_2_72_2","doi-asserted-by":"publisher","DOI":"10.5555\/1558740.1558750"}],"container-title":["ACM Transactions on Architecture and Code Optimization"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3688610","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3688610","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T01:10:30Z","timestamp":1750295430000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3688610"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,3,20]]},"references-count":71,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2025,3,31]]}},"alternative-id":["10.1145\/3688610"],"URL":"https:\/\/doi.org\/10.1145\/3688610","relation":{},"ISSN":["1544-3566","1544-3973"],"issn-type":[{"value":"1544-3566","type":"print"},{"value":"1544-3973","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,3,20]]},"assertion":[{"value":"2023-12-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-07-29","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-03-20","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}