{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,10]],"date-time":"2026-07-10T02:26:28Z","timestamp":1783650388364,"version":"3.55.0"},"reference-count":31,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2018,12,31]],"date-time":"2018-12-31T00:00:00Z","timestamp":1546214400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/100000028","name":"Semiconductor Research Corporation","doi-asserted-by":"crossref","id":[{"id":"10.13039\/100000028","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/100000001","name":"National Science Foundation","doi-asserted-by":"publisher","award":["CNS-1619558"],"award-info":[{"award-number":["CNS-1619558"]}],"id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Reconfigurable Technol. Syst."],"published-print":{"date-parts":[[2018,12,31]]},"abstract":"<jats:p>Field-programmable gate arrays (FPGAs) are used for a wide variety of computations in low-cost embedded systems. Although these systems often have modest performance constraints, their energy consumption must typically be limited. Many FPGA applications employ repetitive loops that cannot be straightforwardly split into parallel computations. Performing a loop sequentially generally requires high-speed clocks that consume considerable clock power and sometimes require clock generation using a phase-locked loop (PLL). Loop unrolling addresses the high-speed clock issue, but its use often leads to significant combinational glitch power.<\/jats:p>\n          <jats:p>In this work, a computer-aided design (CAD) approach that unrolls loops for designs targeted to low-cost FPGAs is described. Our approach considers latency constraints in an effort to minimize energy consumption for loop-based computation. To reduce glitch power, a glitch-filtering approach is introduced that provides a balance between glitch reduction and design performance. Glitch-filter enable signals are generated and routed to the filters using resources best suited to the target FPGA. Our approach automatically inserts glitch filters and associated control logic into a design prior to processing with FPGA synthesis, place, and route tools. Our energy-saving loop-unrolling approach has been evaluated using five benchmarks often used in low-cost FPGAs. The energy-saving capabilities of the approach have been evaluated for an Intel Cyclone IV and a Xilinx Artix-7 FPGA using board-level power measurement. The use of unrolling and glitch filtering is shown to reduce energy by at least 65% for an Artix-7 device and 50% for a Cyclone IV device while meeting design latency constraints.<\/jats:p>","DOI":"10.1145\/3289186","type":"journal-article","created":{"date-parts":[[2019,1,22]],"date-time":"2019-01-22T13:17:41Z","timestamp":1548163061000},"page":"1-23","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["Loop Unrolling for Energy Efficiency in Low-Cost Field-Programmable Gate Arrays"],"prefix":"10.1145","volume":"11","author":[{"given":"Naveen Kumar","family":"Dumpala","sequence":"first","affiliation":[{"name":"University of Massachusetts Amherst, MA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Shivukumar B.","family":"Patil","sequence":"additional","affiliation":[{"name":"University of Massachusetts Amherst, MA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Daniel","family":"Holcomb","sequence":"additional","affiliation":[{"name":"University of Massachusetts Amherst, MA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Russell","family":"Tessier","sequence":"additional","affiliation":[{"name":"University of Massachusetts Amherst, MA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2019,1,21]]},"reference":[{"key":"e_1_2_1_1_1","volume-title":"Altera Cyclone IV GX Development Board. Retrieved","year":"2018","unstructured":"Altera. 2017. Altera Cyclone IV GX Development Board. Retrieved December 9, 2018 from https:\/\/www.altera.com\/products\/boards_and_kits\/dev-kits\/altera\/kit-cyclone-iv-gx.html. Altera. 2017. Altera Cyclone IV GX Development Board. Retrieved December 9, 2018 from https:\/\/www.altera.com\/products\/boards_and_kits\/dev-kits\/altera\/kit-cyclone-iv-gx.html."},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1145\/275107.275139"},{"key":"e_1_2_1_3_1","volume-title":"Proceedings of the IEEE International Symposium on Field-Programmable Custom Computing Machines. IEEE. 70--81","author":"Babb J.","unstructured":"J. Babb , M. Renard , C. Andras Moritz , W. Lee , M. Frank , R. Barua , and S. Amarasinghe . 1999. Parallelizing applications to silicon . In Proceedings of the IEEE International Symposium on Field-Programmable Custom Computing Machines. IEEE. 70--81 . J. Babb, M. Renard, C. Andras Moritz, W. Lee, M. Frank, R. Barua, and S. Amarasinghe. 1999. Parallelizing applications to silicon. In Proceedings of the IEEE International Symposium on Field-Programmable Custom Computing Machines. IEEE. 70--81."},{"key":"e_1_2_1_4_1","volume-title":"Proceedings of Symposium on Hardware-Oriented Security and Trust. IEEE. 55--60","author":"Banik S.","unstructured":"S. Banik , A. Bogdanov , F. Regazzoni , T. Isobe , H. Hiwatari , and T. Akishita . 2016. Round gating for low energy block ciphers . In Proceedings of Symposium on Hardware-Oriented Security and Trust. IEEE. 55--60 . S. Banik, A. Bogdanov, F. Regazzoni, T. Isobe, H. Hiwatari, and T. Akishita. 2016. Round gating for low energy block ciphers. In Proceedings of Symposium on Hardware-Oriented Security and Trust. IEEE. 55--60."},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/2744769.2747946"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/2513683.2513692"},{"key":"e_1_2_1_7_1","volume-title":"WP398 (v1. 0) August 15","author":"Collins A.","year":"2011","unstructured":"A. Collins . 2011. Agile mixed signal addresses analog design challenges. White Paper , WP398 (v1. 0) August 15 ( 2011 ). A. Collins. 2011. Agile mixed signal addresses analog design challenges. White Paper, WP398 (v1. 0) August 15 (2011)."},{"key":"e_1_2_1_8_1","volume-title":"Device Handbook","year":"2010","unstructured":"Cyclone IV , Device Handbook . 2010 . Vol. 1 . Altera , Dec (2010). Cyclone IV, Device Handbook. 2010. Vol. 1. Altera, Dec (2010)."},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/1278480.1278563"},{"key":"e_1_2_1_10_1","volume-title":"Proceedings of Workshop on RFID Security and Privacy. Springer International Publishing.","author":"Dhanuskodi S. N.","unstructured":"S. N. Dhanuskodi and D. Holcomb . 2016. Energy optimization of unrolled block ciphers using combinational checkpointing . In Proceedings of Workshop on RFID Security and Privacy. Springer International Publishing. S. N. Dhanuskodi and D. Holcomb. 2016. Energy optimization of unrolled block ciphers using combinational checkpointing. In Proceedings of Workshop on RFID Security and Privacy. Springer International Publishing."},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCAD.2009.2035564"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/1575779.1575785"},{"key":"e_1_2_1_13_1","volume-title":"Proceedings of the IEEE Conference on Field-Programmable Custom Computing Machines","author":"Dumpala N. K.","unstructured":"N. K. Dumpala , S. B. Patil , D. E. Holcomb , and R. Tessier . 2017. Energy efficient loop unrolling for low-cost FPGAs . In Proceedings of the IEEE Conference on Field-Programmable Custom Computing Machines . Napa, CA, 17--20. N. K. Dumpala, S. B. Patil, D. E. Holcomb, and R. Tessier. 2017. Energy efficient loop unrolling for low-cost FPGAs. In Proceedings of the IEEE Conference on Field-Programmable Custom Computing Machines. Napa, CA, 17--20."},{"key":"e_1_2_1_14_1","volume-title":"International Solid State Circuits Conference. Mira Digital Publishing, 23--25","author":"Fick D.","unstructured":"D. Fick , N. Liu , Z. Foo , M. Fojtik , J. Seo , D. Sylvester , and D. Blaauw . 2010. In situ delay-slack monitor for high-performance processors using an all-digital self-calibrating 5ps resolution time-to-digital converter . In International Solid State Circuits Conference. Mira Digital Publishing, 23--25 . D. Fick, N. Liu, Z. Foo, M. Fojtik, J. Seo, D. Sylvester, and D. Blaauw. 2010. In situ delay-slack monitor for high-performance processors using an all-digital self-calibrating 5ps resolution time-to-digital converter. In International Solid State Circuits Conference. Mira Digital Publishing, 23--25."},{"key":"e_1_2_1_15_1","volume-title":"tiny_aes AES Core. Retrieved","author":"Hsing H.","year":"2018","unstructured":"H. Hsing . 2015. tiny_aes AES Core. Retrieved December 9, 2018 from http:\/\/opencores.org\/project,tiny_aes. H. Hsing. 2015. tiny_aes AES Core. Retrieved December 9, 2018 from http:\/\/opencores.org\/project,tiny_aes."},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/2847263.2847272"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.2478\/v10048-008-0021-z"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-33027-8_23"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/TVLSI.2008.2001237"},{"key":"e_1_2_1_20_1","volume-title":"IEEE\/ACM International Conference on Computer-Aided Design. IEEE Computer Society, 335--342","author":"Lim H.","unstructured":"H. Lim , K. Lee , Y. Cho , and N. Chang . 2005. Flip-flop insertion with shifted-phase clocks for FPGA power reduction . In IEEE\/ACM International Conference on Computer-Aided Design. IEEE Computer Society, 335--342 . H. Lim, K. Lee, Y. Cho, and N. Chang. 2005. Flip-flop insertion with shifted-phase clocks for FPGA power reduction. In IEEE\/ACM International Conference on Computer-Aided Design. IEEE Computer Society, 335--342."},{"key":"e_1_2_1_21_1","volume-title":"5th International Workshop on Power and Timing Modeling","author":"Musoll E.","unstructured":"E. Musoll and J. Cortadella . 1995. Low-power array multipliers with transition retaining barriers . In 5th International Workshop on Power and Timing Modeling . Oldenburg University, 227--235. E. Musoll and J. Cortadella. 1995. Low-power array multipliers with transition retaining barriers. In 5th International Workshop on Power and Timing Modeling. Oldenburg University, 227--235."},{"key":"e_1_2_1_22_1","unstructured":"National Institute of Standards and Technology. 2001. Advanced Encryption Standard (AES). Federal Information Processing Standards Publication FIPS-197.  National Institute of Standards and Technology. 2001. Advanced Encryption Standard (AES). Federal Information Processing Standards Publication FIPS-197."},{"key":"e_1_2_1_23_1","volume-title":"Southern Conference on Programmable Logic. IEEE Press, 1--5.","author":"Oliver J.","unstructured":"J. Oliver , J. P\u00e9rez , and E. Boemo . 2014. Power estimations versus power measurements in Spartan-6 devices . In Southern Conference on Programmable Logic. IEEE Press, 1--5. J. Oliver, J. P\u00e9rez, and E. Boemo. 2014. Power estimations versus power measurements in Spartan-6 devices. In Southern Conference on Programmable Logic. IEEE Press, 1--5."},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1109\/TC.2004.101"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCAD.2012.2192478"},{"key":"e_1_2_1_27_1","volume-title":"Proceedings of the IEEE\/ACM International Symposium on Low-Power Electronics and Design. IEEE Press, 27--32","author":"Shum W.","unstructured":"W. Shum and J. H. Anderson . 2011. FPGA glitch power analysis and reduction . In Proceedings of the IEEE\/ACM International Symposium on Low-Power Electronics and Design. IEEE Press, 27--32 . W. Shum and J. H. Anderson. 2011. FPGA glitch power analysis and reduction. In Proceedings of the IEEE\/ACM International Symposium on Low-Power Electronics and Design. IEEE Press, 27--32."},{"key":"e_1_2_1_28_1","volume-title":"Retrieved","author":"Usselmann R.","year":"2018","unstructured":"R. Usselmann . 2009. DES Core . Retrieved December 9, 2018 from http:\/\/opencores.org\/project,des. R. Usselmann. 2009. DES Core. Retrieved December 9, 2018 from http:\/\/opencores.org\/project,des."},{"key":"e_1_2_1_29_1","volume-title":"Proceedings of Conference on Field Programmable Logic and Application. Springer-Verlag Berlin Heidelberg, 719--728","author":"Wilton S.","unstructured":"S. Wilton , S. Ang , and W. Luk . 2004. The impact of pipelining on energy per operation in field-programmable gate arrays . In Proceedings of Conference on Field Programmable Logic and Application. Springer-Verlag Berlin Heidelberg, 719--728 . S. Wilton, S. Ang, and W. Luk. 2004. The impact of pipelining on energy per operation in field-programmable gate arrays. In Proceedings of Conference on Field Programmable Logic and Application. Springer-Verlag Berlin Heidelberg, 719--728."},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1109\/TNS.2010.2045901"},{"key":"e_1_2_1_31_1","volume-title":"Artix-7 35T Arty FPGA Evaluation Kit. Retrieved","year":"2018","unstructured":"Xilinx. 2017. Artix-7 35T Arty FPGA Evaluation Kit. Retrieved December 9, 2018 from http:\/\/www.xilinx.com\/products\/boards-and-kits\/arty.html#documentation. Xilinx. 2017. Artix-7 35T Arty FPGA Evaluation Kit. Retrieved December 9, 2018 from http:\/\/www.xilinx.com\/products\/boards-and-kits\/arty.html#documentation."},{"key":"e_1_2_1_32_1","volume-title":"Sorting Network IP Generator. Retrieved","author":"Zuluaga M.","year":"2018","unstructured":"M. Zuluaga . 2012. Sorting Network IP Generator. Retrieved December 9, 2018 from http:\/\/www.spiral.net\/hardware\/sort\/sort.html. M. Zuluaga. 2012. Sorting Network IP Generator. Retrieved December 9, 2018 from http:\/\/www.spiral.net\/hardware\/sort\/sort.html."}],"container-title":["ACM Transactions on Reconfigurable Technology and Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3289186","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3289186","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3289186","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T01:02:23Z","timestamp":1750208543000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3289186"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2018,12,31]]},"references-count":31,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2018,12,31]]}},"alternative-id":["10.1145\/3289186"],"URL":"https:\/\/doi.org\/10.1145\/3289186","relation":{},"ISSN":["1936-7406","1936-7414"],"issn-type":[{"value":"1936-7406","type":"print"},{"value":"1936-7414","type":"electronic"}],"subject":[],"published":{"date-parts":[[2018,12,31]]},"assertion":[{"value":"2018-04-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2018-10-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2019-01-21","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}