{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T05:05:16Z","timestamp":1750309516976,"version":"3.41.0"},"reference-count":38,"publisher":"Association for Computing Machinery (ACM)","issue":"1","license":[{"start":{"date-parts":[[2025,3,21]],"date-time":"2025-03-21T00:00:00Z","timestamp":1742515200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62025208, 62272476, 62172430, 62472435, 62272475 and U22B2005"],"award-info":[{"award-number":["62025208, 62272476, 62172430, 62472435, 62272475 and U22B2005"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"863 Program of China","award":["2023YFB3001504"],"award-info":[{"award-number":["2023YFB3001504"]}]},{"name":"STIP of Hunan Province","award":["2022RC3065"],"award-info":[{"award-number":["2022RC3065"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Archit. Code Optim."],"published-print":{"date-parts":[[2025,3,31]]},"abstract":"<jats:p>Deadlock-free adaptive routing is extensively adopted in both on-chip and off-chip interconnection networks to improve communication bandwidth and reduce latency. Introducing virtual channels (VCs), also known as virtual lanes (VLs). This is the mainstream technique to handle deadlocks incurred by adaptive routing and also provides VC preemption for higher priority traffic. However, existing deadlock-free flow control schemes either underutilize memory resources due to inefficient buffer management to simplify hardware implementation, or rely on complicated global coordination and synchronization with very high hardware complexity. Most hardware-friendly schemes use more VCs and memory resources to enable ease of implementation of deadlock-free flow control. In contrast, sophisticated schemes achieve deadlock freedom with minimum VC cost, even eliminating additional buffer requirement through the complicated control mechanisms. In this work, we rethink the root cause of the deadlock problem from a different perspective by considering it as a lack of credit, which makes us find an efficient solution to the deadlock problem. With minor modification of credit accumulation and return, our proposed bubble-swap flow control (BSFC) ensures atomic buffer swap between two adjacent routers only based on local credit status while making full use of the buffer space. BSFC achieves a better tradeoff between implementation complexity and memory overhead and can be easily integrated in the industrial router with no modification on buffer allocation or port arbitration. The simulation results demonstrate BSFC outperforms existing bubble-based deadlock-free methods by average 64% higher throughput. We further propose a credit reservation strategy to eliminate the escape virtual channel (VC) cost for fully adaptive routing implementation. The synthesizing results demonstrate that BSFC along with credit reservation (BSFC-CR) can reduce the area and power consumption by respectively 29% and 26% in contrast to the traditional critical bubble scheme (CBS).<\/jats:p>","DOI":"10.1145\/3705316","type":"journal-article","created":{"date-parts":[[2024,12,17]],"date-time":"2024-12-17T11:24:14Z","timestamp":1734434654000},"page":"1-26","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Bubble-Swap Flow Control"],"prefix":"10.1145","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-9929-6459","authenticated-orcid":false,"given":"Yi","family":"Dai","sequence":"first","affiliation":[{"name":"National University of Defense Technology, Changsha, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6378-7002","authenticated-orcid":false,"given":"Kai","family":"Lu","sequence":"additional","affiliation":[{"name":"National University of Defense Technology, Changsha, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1710-4060","authenticated-orcid":false,"given":"Sheng","family":"Ma","sequence":"additional","affiliation":[{"name":"National University of Defense Technology, Changsha, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9273-616X","authenticated-orcid":false,"given":"Jinshu","family":"Su","sequence":"additional","affiliation":[{"name":"National University of Defense Technology, Changsha, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9743-2034","authenticated-orcid":false,"given":"Dongsheng","family":"Li","sequence":"additional","affiliation":[{"name":"National University of Defense Technology, Changsha, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2025,3,21]]},"reference":[{"key":"e_1_3_1_2_2","unstructured":"InfiniBand Trade Association. Retrieved from www. infinibandta. org. InfiniBandTm architecture specification volume 1 release 1.0. Retrieved December 15 2023 from www.infinibandta.org"},{"key":"e_1_3_1_3_2","article-title":"Intel omni-path architecture technology overview","author":"Birrittella Mark S.","year":"2015","unstructured":"Mark S. Birrittella, Mark Debbage, Ram Huggahalli, James Kunz, Tom Lovett, Todd Rimmer, Keith D. Underwood, and Robert C. Zak. 2015. Intel omni-path architecture technology overview. Intel, Aug (2015), 1--19.","journal-title":"Intel, Aug"},{"key":"e_1_3_1_4_2","article-title":"The intel omni-path architecture (OPA) for machine learning","author":"Chari S.","year":"2017","unstructured":"S. Chari and M. R. Pamidi. Dec. 2017. The intel omni-path architecture (OPA) for machine learning. White Paper (2017), 1--11.","journal-title":"White Paper"},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2013.6522333"},{"key":"e_1_3_1_6_2","first-page":"592","volume-title":"Proceedings of the 2011 IEEE International Parallel & Distributed Processing Symposium","author":"Chen Lizhong","year":"2011","unstructured":"Lizhong Chen, Ruisheng Wang, and Timothy M. Pinkston. 2011. Critical bubble scheme: An efficient implementation of globally aware network flow control. In Proceedings of the 2011 IEEE International Parallel & Distributed Processing Symposium. IEEE, 592\u2013603."},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.1109\/MDAT.2023.3309742"},{"key":"e_1_3_1_8_2","doi-asserted-by":"crossref","first-page":"3","DOI":"10.1007\/978-3-030-78713-4_1","volume-title":"Proceedings of the High Performance Computing.","author":"Dai Yi","year":"2021","unstructured":"Yi Dai, Kai Lu, Junsheng Chang, Xingyun Qi, Jijun Cao, and Jianmin Zhang. 2021. Microarchitecture of a configurable high-radix router for the post-moore era. In Proceedings of the High Performance Computing.Bradford L. Chamberlain, Ana-Lucia Varbanescu, Hatem Ltaief, and Piotr Luszczek (Eds.), Springer International Publishing, Cham, 3\u201317."},{"key":"e_1_3_1_9_2","doi-asserted-by":"crossref","first-page":"1041","DOI":"10.23919\/DATE54114.2022.9774519","volume-title":"Proceedings of the 2022 Design, Automation & Test in Europe Conference & Exhibition","author":"Dai Yi","year":"2022","unstructured":"Yi Dai, Kai Lu, Sheng Ma, and Junsheng Chang. 2022. Full-credit flow control: A novel technique to implement deadlock-free adaptive routing. In Proceedings of the 2022 Design, Automation & Test in Europe Conference & Exhibition. IEEE, 1041\u20131046."},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2018.2873337"},{"key":"e_1_3_1_11_2","doi-asserted-by":"crossref","unstructured":"W. Dally. 1990. Virtual-channel flow control. In ACM SIGARCH Computer Architecture News 18 2SI (1990) 60\u201368.","DOI":"10.1145\/325096.325115"},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.1109\/TC.1987.1676939"},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.5555\/2821589"},{"key":"e_1_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.1109\/71.250114"},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.1109\/71.473515"},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.1109\/IPPS.1999.760469"},{"issue":"2","key":"e_1_3_1_17_2","doi-asserted-by":"crossref","first-page":"703","DOI":"10.1145\/3140659.3080253","article-title":"EbDa: A new theory on design and verification of deadlock-free interconnection networks","volume":"45","author":"Ebrahimi Masoumeh","year":"2017","unstructured":"Masoumeh Ebrahimi and Masoud Daneshtalab. 2017. EbDa: A new theory on design and verification of deadlock-free interconnection networks. ACM SIGARCH Computer Architecture News 45, 2 (2017), 703\u2013715.","journal-title":"ACM SIGARCH Computer Architecture News"},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1109\/71.707539"},{"issue":"2","key":"e_1_3_1_19_2","first-page":"265","article-title":"Blue Gene\/L torus interconnection network","volume":"49","author":"Gara A.","year":"2005","unstructured":"A. Gara, M. E. Giampapa, P. Heidelberger, S. Singh, B. D. Steinmacher-Burow, T. Takken, M. Tsao, and P. Vranas. 2005. Blue Gene\/L torus interconnection network. IBM Journal of Research and Development 49, 2.3 (2005), 265\u2013276.","journal-title":"IBM Journal of Research and Development"},{"key":"e_1_3_1_20_2","volume-title":"Proceedings of the International Conference on Parallel Processing, Volume I","author":"Glass Christopher J.","year":"1992","unstructured":"Christopher J. Glass and Lionel M. Ni. 1992. Maximally fully adaptive routing in 2D meshes. In Proceedings of the International Conference on Parallel Processing, Volume I. Citeseer."},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.1145\/146628.140384"},{"key":"e_1_3_1_22_2","first-page":"203","volume-title":"Proceedings of the High Performance Computer Architecture","author":"Gratz Paul.","year":"2008","unstructured":"Paul. Gratz, Boris Grot, and Stephen W. Keckler. 2008. Regional congestion awareness for load balance in networks-on-chip. In Proceedings of the High Performance Computer Architecture. 203\u2013214."},{"volume-title":"Dual Low Dropout Voltage Regulator","year":"2020","key":"e_1_3_1_23_2","unstructured":"InfiniBandSM Trade Association 2020. Dual Low Dropout Voltage Regulator. InfiniBandSM Trade Association. Rev. 1.4."},{"key":"e_1_3_1_24_2","first-page":"420","volume-title":"Proceedings of the ACM SIGARCH Computer Architecture News","author":"Kim John","year":"2005","unstructured":"John Kim, William J. Dally, Brian Towles, and Amit K. Gupta. 2005. Microarchitecture of a high-radix router. In Proceedings of the ACM SIGARCH Computer Architecture News. IEEE Computer Society, 420\u2013431."},{"key":"e_1_3_1_25_2","doi-asserted-by":"crossref","first-page":"2","DOI":"10.1109\/12.67315","article-title":"An adaptive and fault tolerant wormhole routing strategy for k-ary n-cubes","author":"Linder Daniel H.","year":"1991","unstructured":"Daniel H. Linder and James C. Harden. 1991. An adaptive and fault tolerant wormhole routing strategy for k-ary n-cubes. IEEE Transactions on Computers 40, 1 (1991), 2\u201312.","journal-title":"IEEE Transactions on Computers"},{"key":"e_1_3_1_26_2","volume-title":"Proceedings of the Symposium on Networked Systems Design and Implementation","author":"Lu Yuanwei","year":"2018","unstructured":"Yuanwei Lu, Guo Chen, Bojie Li, Kun Tan, Yongqiang Xiong, Peng Cheng, Jiansong Zhang, Enhong Chen, and Thomas Moscibroda. 2018. Multi-path transport for RDMA in datacenters. In Proceedings of the Symposium on Networked Systems Design and Implementation."},{"key":"e_1_3_1_27_2","first-page":"1","volume-title":"Proceedings of the IEEE International Symposium on High-Performance Comp Architecture","author":"Ma Sheng","year":"2012","unstructured":"Sheng Ma, Natalie Enright Jerger, and Zhiying Wang. 2012. Whole packet forwarding: Efficient design of fully adaptive routing algorithms for networks-on-chip. In Proceedings of the IEEE International Symposium on High-Performance Comp Architecture. IEEE, 1\u201312."},{"key":"e_1_3_1_28_2","volume-title":"Proceedings of the 2020 IEEE International Symposium on High Performance Computer Architecture","author":"Parasar Mayank","year":"2020","unstructured":"Mayank Parasar, Hossein Farrokhbakht, Natalie Enright Jerger, Paul V. Gratz, and Joshua San Miguel. 2020. DRAIN: Deadlock removal for arbitrary irregular networks. In Proceedings of the 2020 IEEE International Symposium on High Performance Computer Architecture."},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.1145\/3352460.3358255"},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.1006\/jpdc.2001.1746"},{"key":"e_1_3_1_31_2","first-page":"699","volume-title":"Proceedings of the 2018 ACM\/IEEE 45th Annual International Symposium on Computer Architecture.","author":"Ramrakhyani Aniruddh","year":"2018","unstructured":"Aniruddh Ramrakhyani, Paul V. Gratz, and Tushar Krishna. 2018. Synchronized progress in interconnection networks (spin): A new theory for deadlock freedom. In Proceedings of the 2018 ACM\/IEEE 45th Annual International Symposium on Computer Architecture.IEEE, 699\u2013711."},{"key":"e_1_3_1_32_2","volume-title":"Proceedings of the Tutorial at the International Symposium on Microarchitecture","author":"Research AMD","year":"2015","unstructured":"AMD Research. 2015. The AMD gem5 APU simulator: Modeling heterogeneous systems in gem5. In Proceedings of the Tutorial at the International Symposium on Microarchitecture."},{"key":"e_1_3_1_33_2","doi-asserted-by":"publisher","DOI":"10.1006\/jpdc.1996.0008"},{"issue":"3","key":"e_1_3_1_34_2","doi-asserted-by":"crossref","first-page":"763","DOI":"10.1109\/TC.2013.2295523","article-title":"Leaving one slot empty: Flit bubble flow control for torus cache-coherent NoCs","volume":"64","author":"Sheng Ma","year":"2015","unstructured":"Ma Sheng, Zhiying Wang, Zonglin Liu Liu, and Natalie Enright Jerger. 2015. Leaving one slot empty: Flit bubble flow control for torus cache-coherent NoCs. IEEE Transactions on Computers 64, 3 (2015), 763\u2013777.","journal-title":"IEEE Transactions on Computers"},{"key":"e_1_3_1_35_2","first-page":"194","volume-title":"Proceedings of the ACM SIGARCH Computer Architecture News","author":"Singh Arjun","year":"2003","unstructured":"Arjun Singh, William J. Dally, Amit K. Gupta, and Brian Towles. 2003. GOAL: A load-balanced adaptive routing algorithm for torus networks. In Proceedings of the ACM SIGARCH Computer Architecture News. IEEE, 194\u2013205."},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2003.1189584"},{"key":"e_1_3_1_37_2","doi-asserted-by":"crossref","first-page":"223","DOI":"10.1145\/371209.371542","volume-title":"Proceedings of the ACM SIGCPR Conference on Computer Personnel Research","author":"Vaidya Aniruddha S.","year":"2001","unstructured":"Aniruddha S. Vaidya, Anand Sivasubramaniam, and Chita R. Das. 2001. Impact of virtual channels and adaptive routing on application performance. In Proceedings of the ACM SIGCPR Conference on Computer Personnel Research. 223\u2013237."},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2011.60"},{"issue":"01","key":"e_1_3_1_39_2","doi-asserted-by":"crossref","first-page":"1450012","DOI":"10.1142\/S0218126614500121","article-title":"A fast and fair shared buffer for high-radix router","volume":"23","author":"Zhang Heying","year":"2014","unstructured":"Heying Zhang, Kefei Wang, Jianmin Zhang, Nan Wu, and Yi Dai. 2014. A fast and fair shared buffer for high-radix router. Journal of Circuits, Systems, and Computers 23, 01 (2014), 1450012.","journal-title":"Journal of Circuits, Systems, and Computers"}],"container-title":["ACM Transactions on Architecture and Code Optimization"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3705316","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3705316","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T01:18:02Z","timestamp":1750295882000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3705316"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,3,21]]},"references-count":38,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2025,3,31]]}},"alternative-id":["10.1145\/3705316"],"URL":"https:\/\/doi.org\/10.1145\/3705316","relation":{},"ISSN":["1544-3566","1544-3973"],"issn-type":[{"type":"print","value":"1544-3566"},{"type":"electronic","value":"1544-3973"}],"subject":[],"published":{"date-parts":[[2025,3,21]]},"assertion":[{"value":"2024-02-09","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-10-28","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-03-21","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}