{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,9,29]],"date-time":"2025-09-29T11:55:47Z","timestamp":1759146947869,"version":"3.41.0"},"reference-count":32,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2014,8,1]],"date-time":"2014-08-01T00:00:00Z","timestamp":1406851200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Des. Autom. Electron. Syst."],"published-print":{"date-parts":[[2014,8]]},"abstract":"<jats:p>WaveSync is a network-on-chip architecture for a globally asynchronous locally-synchronous (GALS) design. The WaveSync design facilitates low-latency communication leveraging the source-synchronous clock sent along with the data to time components in the datapath of a downstream router, reducing the number of synchronizations needed. WaveSync accomplishes this by partitioning the router components at each node into different clock domains, each synchronized with one of the orthogonal incoming source-synchronous clocks in a GALS 2D mesh network. The data and clock subsequently propagate through each node\/router synchronously until the destination is reached, regardless of the number of hops this may take. As long as the data travels in the path of clock propagation and no congestion is encountered, it will be propagated without latching as if in a long combinatorial path, with both the clock and the data accruing delay at the same rate. The result is that the need for synchronization between the mesochronous nodes and\/or the asynchronous control associated with the typical GALS network is completely eliminated. To further reduce the latency overhead of synchronization, for those occasions when synchronization is still required (when a flit takes a turn or arrives at the destination), we propose a novel less-than-one-cycle synchronizer. The proposed WaveSync network outperforms conventional GALS networks by 87--90% in average latency, synthesized using a 45nm CMOS library.<\/jats:p>","DOI":"10.1145\/2647950","type":"journal-article","created":{"date-parts":[[2014,8,26]],"date-time":"2014-08-26T12:08:55Z","timestamp":1409054935000},"page":"1-22","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["WaveSync"],"prefix":"10.1145","volume":"19","author":[{"given":"Yoon Seok","family":"Yang","sequence":"first","affiliation":[{"name":"Texas A&amp;M University, TX"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Reeshav","family":"Kumar","sequence":"additional","affiliation":[{"name":"Texas A&amp;M University, TX"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Gwan","family":"Choi","sequence":"additional","affiliation":[{"name":"Texas A&amp;M University, TX"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Paul V.","family":"Gratz","sequence":"additional","affiliation":[{"name":"Texas A&amp;M University, TX"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2014,8,29]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1049\/ip-cdt:20045093"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1109\/92.711317"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1109\/12.53599"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/12.83652"},{"key":"e_1_2_1_5_1","doi-asserted-by":"crossref","unstructured":"W. J. Dally and J. Poulton. 1998. Digital Systems Engineering. Cambridge University Press.   W. J. Dally and J. Poulton. 1998. Digital Systems Engineering. Cambridge University Press.","DOI":"10.1017\/CBO9781139166980"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/378239.379048"},{"key":"e_1_2_1_7_1","unstructured":"W. J. Dally and B. Towles. 2003. Principles and Practices of Interconnection Networks. Morgan Kaufmann San Fransisco.   W. J. Dally and B. Towles. 2003. Principles and Practices of Interconnection Networks. Morgan Kaufmann San Fransisco."},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/981066.981097"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/1999946.1999977"},{"volume-title":"Proceedings of the 4th Workshop on Chip Multiprocessor Memory Systems and Interconnects (CMP-MSI'10)","author":"Gratz P.","key":"e_1_2_1_10_1","unstructured":"P. Gratz and S. W. Keckler . 2010. Realistic workload characterization and analysis for networks-on-chip design . In Proceedings of the 4th Workshop on Chip Multiprocessor Memory Systems and Interconnects (CMP-MSI'10) . P. Gratz and S. W. Keckler. 2010. Realistic workload characterization and analysis for networks-on-chip design. In Proceedings of the 4th Workshop on Chip Multiprocessor Memory Systems and Interconnects (CMP-MSI'10)."},{"volume-title":"Proceedings of the IEEE International Conference on Computer Design (ICCD'06)","author":"Gratz P.","key":"e_1_2_1_11_1","unstructured":"P. Gratz , C. Kim , R. Mcdonald , S. W. Keckler , and D. Burger . 2006. Implementation and evaluation of on-chip network architectures . In Proceedings of the IEEE International Conference on Computer Design (ICCD'06) . P. Gratz, C. Kim, R. Mcdonald, S. W. Keckler, and D. Burger. 2006. Implementation and evaluation of on-chip network architectures. In Proceedings of the IEEE International Conference on Computer Design (ICCD'06)."},{"volume-title":"Proceedings of the International Symposium on High Performance Computer Architecture. 163--174","author":"Grot B.","key":"e_1_2_1_12_1","unstructured":"B. Grot , J. Hestness , S. W. Keckler , and O. Mutlu . 2009. Express cube topologies for on-chip interconnects . In Proceedings of the International Symposium on High Performance Computer Architecture. 163--174 . B. Grot, J. Hestness, S. W. Keckler, and O. Mutlu. 2009. Express cube topologies for on-chip interconnects. In Proceedings of the International Symposium on High Performance Computer Architecture. 163--174."},{"volume-title":"Proceedings of the International Solid-State Circuits Conference. 72--73","author":"Hatakeyama A.","key":"e_1_2_1_13_1","unstructured":"A. Hatakeyama , H. Mochizuki , T. Aikawa , M. Takiia , Y. Ishii , H. Tsuboi , S. Fujioka , S. Yamaguchi , M. Koga , Y. Serizawa , K. Nishimura , K. Kawabata , Y. Okajima , M. Kawano , H. Kojima , K. Mizutani , T. Anezaki , M. Hasegawa , and M. Taguchi . 1997. A 256 mb sdram using a register-controlled digital dll . In Proceedings of the International Solid-State Circuits Conference. 72--73 . A. Hatakeyama, H. Mochizuki, T. Aikawa, M. Takiia, Y. Ishii, H. Tsuboi, S. Fujioka, S. Yamaguchi, M. Koga, Y. Serizawa, K. Nishimura, K. Kawabata, Y. Okajima, M. Kawano, H. Kojima, K. Mizutani, T. Anezaki, M. Hasegawa, and M. Taguchi. 1997. A 256 mb sdram using a register-controlled digital dll. In Proceedings of the International Solid-State Circuits Conference. 72--73."},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/309847.310091"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISED.2010.12"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCAD.2011.2114970"},{"key":"e_1_2_1_17_1","unstructured":"International Technology Roadmap for Semiconductors. 2012. http:\/\/www.itrs.net\/Links\/2012ITRS\/Home2012.htm.  International Technology Roadmap for Semiconductors. 2012. http:\/\/www.itrs.net\/Links\/2012ITRS\/Home2012.htm."},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/NOCS.2010.15"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/1669112.1669145"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA.2005.35"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/1065579.1065726"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1109\/HOTI.2008.22"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1109\/VLSID.2011.73"},{"volume-title":"Proceedings of the 31st Annual International Symposium on Computer Architecture. 188--197","author":"Mullins R.","key":"e_1_2_1_24_1","unstructured":"R. Mullins , A. West , and S. Moore . 2004. Low-latency virtual-channel routers for on-chip networks . In Proceedings of the 31st Annual International Symposium on Computer Architecture. 188--197 . R. Mullins, A. West, and S. Moore. 2004. Low-latency virtual-channel routers for on-chip networks. In Proceedings of the 31st Annual International Symposium on Computer Architecture. 188--197."},{"volume-title":"Proceedings of the International Conference on Computer-Aided Design. 246--253","author":"Ogras U.","key":"e_1_2_1_25_1","unstructured":"U. Ogras and R. Marculescu . 2005. Application-specific network-on-chip architecture customization via long-range link insertion . In Proceedings of the International Conference on Computer-Aided Design. 246--253 . U. Ogras and R. Marculescu. 2005. Application-specific network-on-chip architecture customization via long-range link insertion. In Proceedings of the International Conference on Computer-Aided Design. 246--253."},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1109\/NOCS.2007.14"},{"volume-title":"Proceedings of the Custom Integrated Circuits Conference. 511--514","author":"Saeki T.","key":"e_1_2_1_27_1","unstructured":"T. Saeki , K. Minami , H. Yoshida , and H. Suzuki . 1998. The direct skew detect synchronous mirror delay (direct smd) for asics . In Proceedings of the Custom Integrated Circuits Conference. 511--514 . T. Saeki, K. Minami, H. Yoshida, and H. Suzuki. 1998. The direct skew detect synchronous mirror delay (direct smd) for asics. In Proceedings of the Custom Integrated Circuits Conference. 511--514."},{"key":"e_1_2_1_28_1","unstructured":"Synopsys. 2012. RTL synthesis and test. http:\/\/www.synopsys.com\/tools\/implementation\/rtlsynthesis\/pages\/default.aspx.  Synopsys. 2012. RTL synthesis and test. http:\/\/www.synopsys.com\/tools\/implementation\/rtlsynthesis\/pages\/default.aspx."},{"volume-title":"Proceedings of the International Symposium on Circuits and Systems. 996--999","author":"Tran A. T.","key":"e_1_2_1_29_1","unstructured":"A. T. Tran , D. N. Truong , and B. M. Baas . 2009. A low-cost high-speed source-synchronous interconnection technique for gals chip multiprocessors . In Proceedings of the International Symposium on Circuits and Systems. 996--999 . A. T. Tran, D. N. Truong, and B. M. Baas. 2009. A low-cost high-speed source-synchronous interconnection technique for gals chip multiprocessors. In Proceedings of the International Symposium on Circuits and Systems. 996--999."},{"volume-title":"Proceedings of the Solid-State Circuits Conference. 98--589","author":"Vangal S.","key":"e_1_2_1_30_1","unstructured":"S. Vangal , J. Howard , G. Ruhl , S. Dighe , H. Wilson , J. Tschanz , D. Finan , Iyer, P., A. Singh , A., T. Jacob , S. Jain , S. Venkataraman , Y. Hoskote , and N. Borkar . 2007. An 80-tile 1.28tflops network-on-chip in 65nm cmos . In Proceedings of the Solid-State Circuits Conference. 98--589 . S. Vangal, J. Howard, G. Ruhl, S. Dighe, H. Wilson, J. Tschanz, D. Finan, Iyer, P., A. Singh, A., T. Jacob, S. Jain, S. Venkataraman, Y. Hoskote, and N. Borkar. 2007. An 80-tile 1.28tflops network-on-chip in 65nm cmos. In Proceedings of the Solid-State Circuits Conference. 98--589."},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1145\/223982.223990"},{"key":"e_1_2_1_32_1","unstructured":"Xilinx. 2009. Power consumption at 40 and 45nm. http:\/\/www.xilinx.com\/support\/documentation\/white_papers\/wp298.pdf.  Xilinx. 2009. Power consumption at 40 and 45nm. http:\/\/www.xilinx.com\/support\/documentation\/white_papers\/wp298.pdf."}],"container-title":["ACM Transactions on Design Automation of Electronic Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2647950","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/2647950","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T07:19:40Z","timestamp":1750231180000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2647950"}},"subtitle":["Low-Latency Source-Synchronous Bypass Network-on-Chip Architecture"],"short-title":[],"issued":{"date-parts":[[2014,8]]},"references-count":32,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2014,8]]}},"alternative-id":["10.1145\/2647950"],"URL":"https:\/\/doi.org\/10.1145\/2647950","relation":{},"ISSN":["1084-4309","1557-7309"],"issn-type":[{"type":"print","value":"1084-4309"},{"type":"electronic","value":"1557-7309"}],"subject":[],"published":{"date-parts":[[2014,8]]},"assertion":[{"value":"2013-01-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2014-05-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2014-08-29","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}