{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T04:35:18Z","timestamp":1750307718336,"version":"3.41.0"},"reference-count":42,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2009,5,1]],"date-time":"2009-05-01T00:00:00Z","timestamp":1241136000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Des. Autom. Electron. Syst."],"published-print":{"date-parts":[[2009,5]]},"abstract":"<jats:p>\n            Behavior synthesis and optimization beyond the register-transfer level require an efficient utilization of the underlying platform features. This article presents a platform-based resource binding approach based on a\n            <jats:italic>Distributed Register-File Microarchitecture (DRFM)<\/jats:italic>\n            , which makes efficient use of distributed embedded memory blocks as register files in modern FPGAs. DRFM contains multiple islands, each having a local register file, a functional unit pool, and data-routing logic. Compared to the traditional discrete-register counterpart, a DRFM allows use of the platform-featured on-chip memory or register-file IP blocks to implement its local register files, and this results in a substantial saving of multiplexing logic and global interconnects. DRFM provides a useful architectural template and a direct optimization objective for minimizing interisland connections for synthesis algorithms. Given the scheduling solution and resource (functional units) constraints, two novel algorithms in the resource binding stage are developed based on DRFM: (i) a simultaneous DRFM clustering and binding algorithm, which decides the configuration of DRFM and the assignment of operations into islands with the focus on optimizing global connections; (ii) a data-forwarding scheduling algorithm, which takes advantage of the operation slacks to handle the read-port restriction of register files. On the Xilinx Virtex4 FPGA platform, experimental results with a set of real-life test cases show a 50% logic area reduction achieved by applying our approach, with a 14.6% performance improvement, compared to the traditional discrete-register-based approach. Also, experiments on small-size designs show that our algorithm produces the same number of total connections and at most one more maximum feeding-in connection compared to optimal solutions generated by ILP.\n          <\/jats:p>","DOI":"10.1145\/1529255.1529257","type":"journal-article","created":{"date-parts":[[2009,6,2]],"date-time":"2009-06-02T14:51:08Z","timestamp":1243954268000},"page":"1-31","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":17,"title":["Simultaneous resource binding and interconnection optimization based on a distributed register-file microarchitecture"],"prefix":"10.1145","volume":"14","author":[{"given":"Jason","family":"Cong","sequence":"first","affiliation":[{"name":"University of California, Los Angeles, Los Angeles, CA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yiping","family":"Fan","sequence":"additional","affiliation":[{"name":"AutoESL, Inc., Cupertino, CA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Junjuan","family":"Xu","sequence":"additional","affiliation":[{"name":"University of California, Los Angeles, Los Angeles, CA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2009,6,4]]},"reference":[{"key":"e_1_2_1_1_1","unstructured":"Altera. Altera Web site. http:\/\/www.altera.com.  Altera. Altera Web site. http:\/\/www.altera.com."},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1016\/0020-0190(79)90143-1"},{"volume-title":"School of Electrical and Computer Engineering","author":"Bunchua S.","key":"e_1_2_1_3_1"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1145\/217474.217502"},{"volume-title":"Proceedings of the Asian and South Pacific Design Automation Conference, 68--73","author":"Chen D.","key":"e_1_2_1_5_1"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/871506.871541"},{"volume-title":"Proceedings of IEEE International SOC Conference (invited paper), 199--202","author":"Cong J.","key":"e_1_2_1_7_1"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCAD.2004.825872"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/1233501.1233648"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/1403375.1403629"},{"volume-title":"Proceedings of the 20th Conference on Advanced Research in VLSI, 232--241","author":"Dally W. J.","key":"e_1_2_1_11_1"},{"volume-title":"Synthesis and Optimization of Digital Circuits","author":"De Micheli G.","key":"e_1_2_1_12_1"},{"volume-title":"Proceedings of the 30th International Symposium on Microarchitecture, 149--159","author":"Farkas K. I.","key":"e_1_2_1_13_1"},{"key":"e_1_2_1_14_1","unstructured":"FFT. FFT package. http:\/\/momonga.t.u-tokyo.ac.jp\/~ooura\/fft.html.  FFT. FFT package. http:\/\/momonga.t.u-tokyo.ac.jp\/~ooura\/fft.html."},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/28869.28874"},{"key":"e_1_2_1_16_1","doi-asserted-by":"crossref","unstructured":"Gajski D. Dutt N. Wu A. and Lin S. 1992. High-Level Synthesis C Introduction to Chip and System Design. Kulwer Academic.   Gajski D. Dutt N. Wu A. and Lin S. 1992. High-Level Synthesis C Introduction to Chip and System Design. Kulwer Academic.","DOI":"10.1007\/978-1-4615-3636-9"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/266021.266192"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/TVLSI.2007.904096"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/123186.123350"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/370155.370576"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1002\/j.1538-7305.1970.tb01770.x"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1109\/40.918001"},{"volume-title":"Proceedings of the International Conference on Computer Aided Design, 320--326","author":"Kim D.","key":"e_1_2_1_23_1"},{"volume-title":"Custom Integrated Circuits Conference, Proc. IEEE 1--4, 615--618","author":"Kim T.","key":"e_1_2_1_24_1"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1016\/0167-9260(95)00009-5"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1145\/280756.280873"},{"volume-title":"International Symposium on Microarchitecture, 330--335","author":"Lee C.","key":"e_1_2_1_27_1"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1145\/224818.224847"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/1278480.1278672"},{"volume-title":"Proceedings of the 21st International Conference on Computer Design, 140--145","author":"Luthra M.","key":"e_1_2_1_30_1"},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1016\/S0898-1221(98)00076-5"},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1109\/43.97625"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1109\/54.82037"},{"volume-title":"Proceedings of the 31st Annual ACM\/IEEE International Symposium on Microarchitecture.","author":"Rixner S.","key":"e_1_2_1_34_1"},{"volume-title":"Proceedings of the 6th International Symposium on High-Performance Computer Architecture, 375--386","author":"Rixner S.","key":"e_1_2_1_35_1"},{"volume-title":"Proceedings of the 35th Annual ACM\/IEEE International Symposium on Microarchitecture.","author":"Seznec A.","key":"e_1_2_1_36_1"},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1145\/605440.605448"},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1109\/92.365450"},{"volume-title":"Proceedings of the IEEE\/ACM International Conference on Computer-Aided Design.","year":"1992","author":"Stok L.","key":"e_1_2_1_39_1"},{"key":"e_1_2_1_40_1","first-page":"2862","article-title":"Module allocation and comparability graphs","volume":"5","author":"Stok L.","year":"1991","journal-title":"IEEE International Sympoisum on Circuits and Systems"},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1145\/513918.514141"},{"key":"e_1_2_1_42_1","unstructured":"Xilinx. Xilinx Web site. http:\/\/www.xilinx.com.  Xilinx. Xilinx Web site. http:\/\/www.xilinx.com."}],"container-title":["ACM Transactions on Design Automation of Electronic Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/1529255.1529257","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/1529255.1529257","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T13:30:26Z","timestamp":1750253426000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/1529255.1529257"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2009,5]]},"references-count":42,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2009,5]]}},"alternative-id":["10.1145\/1529255.1529257"],"URL":"https:\/\/doi.org\/10.1145\/1529255.1529257","relation":{},"ISSN":["1084-4309","1557-7309"],"issn-type":[{"type":"print","value":"1084-4309"},{"type":"electronic","value":"1557-7309"}],"subject":[],"published":{"date-parts":[[2009,5]]},"assertion":[{"value":"2008-07-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2009-02-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2009-06-04","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}