{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,8]],"date-time":"2026-06-08T11:47:55Z","timestamp":1780919275246,"version":"3.54.1"},"reference-count":29,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2016,8,22]],"date-time":"2016-08-22T00:00:00Z","timestamp":1471824000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"DARPA\/CMO","award":["HR0011-13-C-0005"],"award-info":[{"award-number":["HR0011-13-C-0005"]}]},{"name":"VIPER program at the University of Pennsylvania"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Reconfigurable Technol. Syst."],"published-print":{"date-parts":[[2016,9,20]]},"abstract":"<jats:p>The energy in FPGA computations is dominated by data communication energy, either in the form of memory references or data movement on interconnect. In this article, we explore how to use data placement and parallelism to reduce communication energy. We show that parallelism can reduce energy and that the optimal level of parallelism increases with the problem size. We further explore how FPGA memory architecture (memory block size(s), memory banking, and spacing between memory banks) can impact communication energy, and determine how to organize the memory architecture to guarantee that the energy overhead compared to the optimally matched architecture for the design is never more than 60%. We specifically show that an architecture with 32 bit wide, 16Kb internally banked memories placed every 8 columns of 10 4-LUT logic blocks is within 61% of the optimally matched architecture across the VTR 7 benchmark set and a set of parallelism-tunable benchmarks. Without internal banking, the worst-case overhead is 98%, achieved with an architecture with 32 bit wide, 8Kb memories placed every 9 columns, roughly comparable to the memory organization on the Cyclone V (where memories are placed about every 10 columns). Monolithic 32 bit wide, 16Kb memories placed every 10 columns (comparable to 18Kb and 20Kb memories used in Virtex 4 and Stratix V FPGAs) have a 180% worst-case energy overhead. Furthermore, we show practical cases where designs mapped for optimal parallelism use 4.7 \u00d7 less energy than designs using a single processing element.<\/jats:p>","DOI":"10.1145\/2857057","type":"journal-article","created":{"date-parts":[[2016,8,26]],"date-time":"2016-08-26T12:25:39Z","timestamp":1472214339000},"page":"1-23","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":7,"title":["Impact of Parallelism and Memory Architecture on FPGA Communication Energy"],"prefix":"10.1145","volume":"9","author":[{"given":"Edin","family":"Kadric","sequence":"first","affiliation":[{"name":"University of Pennsylvania, Philadelphia, PA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"David","family":"Lakata","sequence":"additional","affiliation":[{"name":"University of Pennsylvania, Philadelphia, PA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Andr\u00e9","family":"Dehon","sequence":"additional","affiliation":[{"name":"University of Pennsylvania, Philadelphia, PA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2016,8,22]]},"reference":[{"key":"e_1_2_1_1_1","volume-title":"PowerPlay Early Power Estimator","author":"Altera Corporation","unstructured":"Altera Corporation . 2013. PowerPlay Early Power Estimator . Altera Corporation , San Jose, CA . http:\/\/www.altera.com\/support\/devices\/estimator\/pow-powerplay.jsp. Altera Corporation. 2013. PowerPlay Early Power Estimator. Altera Corporation, San Jose, CA. http:\/\/www.altera.com\/support\/devices\/estimator\/pow-powerplay.jsp."},{"key":"e_1_2_1_2_1","volume-title":"Architecture and CAD for Deep-Submicron FPGAs","author":"Betz Vaughn","unstructured":"Vaughn Betz , Jonathan Rose , and Alexander Marquardt . 1999. Architecture and CAD for Deep-Submicron FPGAs . Kluwer , Norwell, MA . Vaughn Betz, Jonathan Rose, and Alexander Marquardt. 1999. Architecture and CAD for Deep-Submicron FPGAs. Kluwer, Norwell, MA."},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1016\/0022-0000(84)90071-0"},{"key":"e_1_2_1_4_1","unstructured":"Bluespec. 2012. Bluespec SystemVerilog 2012.01.A. Available at http:\/\/www.bluespec.com.  Bluespec. 2012. Bluespec SystemVerilog 2012.01.A. Available at http:\/\/www.bluespec.com."},{"key":"e_1_2_1_5_1","volume-title":"Proceedings of the International Conference on Field-Programmable Logic and Applications. 1--8. DOI:http:\/\/dx.doi.org\/10","author":"Chin S. Y. I.","year":"2006","unstructured":"S. Y. I. Chin , C. S. P. Lee , and Steven J. E. Wilton . 2006. Power implications of implementing logic using FPGA embedded memory arrays . In Proceedings of the International Conference on Field-Programmable Logic and Applications. 1--8. DOI:http:\/\/dx.doi.org\/10 .1109\/FPL. 2006 .311200 10.1109\/FPL.2006.311200 S. Y. I. Chin, C. S. P. Lee, and Steven J. E. Wilton. 2006. Power implications of implementing logic using FPGA embedded memory arrays. In Proceedings of the International Conference on Field-Programmable Logic and Applications. 1--8. DOI:http:\/\/dx.doi.org\/10.1109\/FPL.2006.311200"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/296399.296431"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1109\/JPROC.2014.2387696"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/2019583.2019584"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCS.1979.1084635"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/TVLSI.2013.2249295"},{"key":"e_1_2_1_11_1","volume-title":"Proceedings of the International Conference on Field-Programmable Technology. 229--234","author":"Goeders J. B.","year":"2012","unstructured":"J. B. Goeders and Steven J. E. Wilton . 2012. VersaPower: Power estimation for diverse FPGA architectures . In Proceedings of the International Conference on Field-Programmable Technology. 229--234 . DOI:http:\/\/dx.doi.org\/10.1109\/FPT. 2012 .6412139 10.1109\/FPT.2012.6412139 J. B. Goeders and Steven J. E. Wilton. 2012. VersaPower: Power estimation for diverse FPGA architectures. In Proceedings of the International Conference on Field-Programmable Technology. 229--234. DOI:http:\/\/dx.doi.org\/10.1109\/FPT.2012.6412139"},{"key":"e_1_2_1_12_1","volume-title":"The Thirteen Books of Euclid\u2019s Elements","author":"Heath Thomas L.","unstructured":"Thomas L. Heath and Euclid. 1956. The Thirteen Books of Euclid\u2019s Elements , Books I and II (2nd ed.). Dover Publications . Thomas L. Heath and Euclid. 1956. The Thirteen Books of Euclid\u2019s Elements, Books I and II (2nd ed.). Dover Publications."},{"key":"e_1_2_1_13_1","unstructured":"ITRS. 2012. International Technology Roadmap for Semiconductors. Available at http:\/\/www.itrs2.net\/itrs-reports.html.  ITRS. 2012. International Technology Roadmap for Semiconductors. Available at http:\/\/www.itrs2.net\/itrs-reports.html."},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/2684746.2689062"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.5555\/2650280.2650321"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/1950413.1950427"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCAD.2006.884574"},{"key":"e_1_2_1_18_1","volume-title":"Proceedings of the International Conference on Field-Programmable Logic and Applications. 1--8. DOI:http:\/\/dx.doi.org\/10","author":"Lamoureux J.","year":"2006","unstructured":"J. Lamoureux and Steven J. E. Wilton . 2006. Activity estimation for field-programmable gate arrays . In Proceedings of the International Conference on Field-Programmable Logic and Applications. 1--8. DOI:http:\/\/dx.doi.org\/10 .1109\/FPL. 2006 .311199 10.1109\/FPL.2006.311199 J. Lamoureux and Steven J. E. Wilton. 2006. Activity estimation for field-programmable gate arrays. In Proceedings of the International Conference on Field-Programmable Logic and Applications. 1--8. DOI:http:\/\/dx.doi.org\/10.1109\/FPL.2006.311199"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/T-C.1971.223159"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/1508128.1508135"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/2435264.2435292"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/1950413.1950457"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/2617593"},{"key":"e_1_2_1_24_1","first-page":"2009","article-title":"CACTI 6.0: A Tool to Model Large Caches","author":"Muralimanohar Naveen","year":"2009","unstructured":"Naveen Muralimanohar , Rajeev Balasubramonian , and Norman P. Jouppi . 2009 . CACTI 6.0: A Tool to Model Large Caches . HPL 2009 - 2085 . HP Labs, Palo Alto, CA. http:\/\/www.hpl.hp.com\/techreports\/2009\/HPL-2009-85.html. Naveen Muralimanohar, Rajeev Balasubramonian, and Norman P. Jouppi. 2009. CACTI 6.0: A Tool to Model Large Caches. HPL 2009-85. HP Labs, Palo Alto, CA. http:\/\/www.hpl.hp.com\/techreports\/2009\/HPL-2009-85.html.","journal-title":"HPL"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/1059876.1059881"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1145\/2145694.2145708"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCAD.2006.887924"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1145\/800135.804401"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/1950413.1950419"}],"container-title":["ACM Transactions on Reconfigurable Technology and Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2857057","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/2857057","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T04:39:08Z","timestamp":1750221548000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2857057"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2016,8,22]]},"references-count":29,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2016,9,20]]}},"alternative-id":["10.1145\/2857057"],"URL":"https:\/\/doi.org\/10.1145\/2857057","relation":{},"ISSN":["1936-7406","1936-7414"],"issn-type":[{"value":"1936-7406","type":"print"},{"value":"1936-7414","type":"electronic"}],"subject":[],"published":{"date-parts":[[2016,8,22]]},"assertion":[{"value":"2015-07-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2015-12-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2016-08-22","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}