{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,11,18]],"date-time":"2025-11-18T12:17:34Z","timestamp":1763468254973,"version":"3.41.0"},"reference-count":29,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2015,1,23]],"date-time":"2015-01-23T00:00:00Z","timestamp":1421971200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"DARPA CRASH program through the United States Air Force Research Laboratory (AFRL) under Contract No. FA8650-10-C-7090"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Reconfigurable Technol. Syst."],"published-print":{"date-parts":[[2015,1,23]]},"abstract":"<jats:p>Associative memories can map sparsely used keys to values with low latency but can incur heavy area overheads. The lack of customized hardware for associative memories in today\u2019s mainstream FPGAs exacerbates the overhead cost of building these memories using the fixed address match BRAMs. In this article, we develop a new, FPGA-friendly, memory system architecture based on a multiple hash scheme that is able to achieve near-associative performance without the area-delay overheads of a fully associative memory on FPGAs. At the same time, we develop a novel memory management algorithm that allows us to statistically mimic an associative memory. Using the proposed architecture as a 64KB L1 data cache, we show that it is able to achieve near-associative miss rates while consuming 3--13 \u00d7 fewer FPGA memory resources for a set of benchmark programs from the SPEC CPU2006 suite than fully associative memories generated by the Xilinx Coregen tool. Benefits for our architecture increase with key width, allowing area reduction up to 100 \u00d7. Mapping delay is also reduced to 3.7ns for a 1,024-entry flat version or 6.1ns for an area-efficient version compared to 17.6ns for a fully associative memory for a 64-bit key on a Xilinx Virtex 6 device.<\/jats:p>","DOI":"10.1145\/2629471","type":"journal-article","created":{"date-parts":[[2015,1,28]],"date-time":"2015-01-28T14:05:51Z","timestamp":1422453951000},"page":"1-22","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":5,"title":["Area-Efficient Near-Associative Memories on FPGAs"],"prefix":"10.1145","volume":"7","author":[{"given":"Udit","family":"Dhawan","sequence":"first","affiliation":[{"name":"University of Pennsylvania, Philadelphia, PA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Andr\u00e9","family":"Dehon","sequence":"additional","affiliation":[{"name":"University of Pennsylvania, Philadelphia, PA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2015,1,23]]},"reference":[{"key":"e_1_2_2_1_1","doi-asserted-by":"publisher","DOI":"10.1145\/195058.195412"},{"key":"e_1_2_2_2_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2012.6169033"},{"key":"e_1_2_2_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/362686.362692"},{"key":"e_1_2_2_4_1","unstructured":"Bluespec Inc. 2012. Bluespec SystemVerilog 2012.01.A. Retrieved from http:\/\/www.bluespec.com.  Bluespec Inc. 2012. Bluespec SystemVerilog 2012.01.A. Retrieved from http:\/\/www.bluespec.com."},{"key":"e_1_2_2_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/129617.129622"},{"key":"e_1_2_2_6_1","volume-title":"Proceedings of ACM-SIAM Symposium on Discrete Algorithms (SODA\u201904)","author":"Chazelle Bernard","year":"2004","unstructured":"Bernard Chazelle , Joe Kilian , Ronitt Rubinfeld , and Ayellet Tal . 2004 . The Bloomier filter: An efficient data structure for static support lookup tables . In Proceedings of ACM-SIAM Symposium on Discrete Algorithms (SODA\u201904) . Society for Industrial and Applied Mathematics, Philadelphia, PA, 30--39. Bernard Chazelle, Joe Kilian, Ronitt Rubinfeld, and Ayellet Tal. 2004. The Bloomier filter: An efficient data structure for static support lookup tables. In Proceedings of ACM-SIAM Symposium on Discrete Algorithms (SODA\u201904). Society for Industrial and Applied Mathematics, Philadelphia, PA, 30--39."},{"key":"e_1_2_2_7_1","doi-asserted-by":"publisher","DOI":"10.1016\/0020-0190(92)90220-P"},{"key":"e_1_2_2_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/2435264.2435298"},{"key":"e_1_2_2_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/SASOW.2012.11"},{"key":"e_1_2_2_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/90.851975"},{"key":"e_1_2_2_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/1186736.1186737"},{"key":"e_1_2_2_12_1","volume-title":"Proceedings of the International Conference on Field-Programmable Technology. 73--80","author":"Ho J.","year":"2008","unstructured":"J. Ho and G. Lemieux . 2008. PERG: A scalable FPGA-based pattern-matching engine with consolidated Bloomier filters . In Proceedings of the International Conference on Field-Programmable Technology. 73--80 . DOI: http:\/\/dx.doi.org\/10.1109\/FPT. 2008 .4762368 10.1109\/FPT.2008.4762368 J. Ho and G. Lemieux. 2008. PERG: A scalable FPGA-based pattern-matching engine with consolidated Bloomier filters. In Proceedings of the International Conference on Field-Programmable Technology. 73--80. DOI: http:\/\/dx.doi.org\/10.1109\/FPT.2008.4762368"},{"key":"e_1_2_2_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/TNET.2010.2047868"},{"key":"e_1_2_2_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/2145694.2145731"},{"key":"e_1_2_2_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/1331897.1331901"},{"key":"e_1_2_2_16_1","doi-asserted-by":"publisher","DOI":"10.1017\/S0963548399003946"},{"key":"e_1_2_2_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/1477942.1477944"},{"key":"e_1_2_2_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2010.20"},{"key":"e_1_2_2_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/165123.165152"},{"key":"#cr-split#-e_1_2_2_20_1.1","doi-asserted-by":"crossref","unstructured":"Andr\u00e9 Seznec and Fran\u00e7ois Bodin. 1993. Skewed-associative caches. In Parallel Architectures and Languages Europe. 304--316. DOI: http:\/\/dx.doi.org\/10.1007\/3-540-56891-3_24 10.1007\/3-540-56891-3_24","DOI":"10.1007\/3-540-56891-3_24"},{"key":"#cr-split#-e_1_2_2_20_1.2","doi-asserted-by":"crossref","unstructured":"Andr\u00e9 Seznec and Fran\u00e7ois Bodin. 1993. Skewed-associative caches. In Parallel Architectures and Languages Europe. 304--316. DOI: http:\/\/dx.doi.org\/10.1007\/3-540-56891-3_24","DOI":"10.1007\/3-540-56891-3_24"},{"key":"e_1_2_2_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/1080091.1080114"},{"key":"e_1_2_2_22_1","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2007.39"},{"key":"e_1_2_2_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/1216919.1216936"},{"volume-title":"Proceedings of the International Conference on Computer Design. 288--294","author":"Wunderlich Rondald","key":"e_1_2_2_24_1","unstructured":"Rondald Wunderlich and James C. Hoe . 2004. In-system FPGA prototyping of an itanium microarchitecture . In Proceedings of the International Conference on Computer Design. 288--294 . Rondald Wunderlich and James C. Hoe. 2004. In-system FPGA prototyping of an itanium microarchitecture. In Proceedings of the International Conference on Computer Design. 288--294."},{"key":"e_1_2_2_25_1","unstructured":"Xilinx Inc. 2011a. Parameterizable Content-Addressable Memory. Xilinx Inc. 2100 Logic Drive San Jose CA 95124. XAPP 1151 http:\/\/www.xilinx.com\/support\/documentation\/application_notes\/xapp1 151_Param_CAM.pdf.  Xilinx Inc. 2011a. Parameterizable Content-Addressable Memory. Xilinx Inc. 2100 Logic Drive San Jose CA 95124. XAPP 1151 http:\/\/www.xilinx.com\/support\/documentation\/application_notes\/xapp1 151_Param_CAM.pdf."},{"key":"e_1_2_2_26_1","unstructured":"Xilinx Inc. 2011b. Virtex-6 FPGA Data Sheet: DC and Switching Characteristics. Xilinx Inc. 2100 Logic Drive San Jose CA 95124.  Xilinx Inc. 2011b. Virtex-6 FPGA Data Sheet: DC and Switching Characteristics. Xilinx Inc. 2100 Logic Drive San Jose CA 95124."},{"key":"e_1_2_2_27_1","doi-asserted-by":"publisher","DOI":"10.1109\/FPT.2003.1275768"},{"key":"e_1_2_2_28_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCAD.2006.887921"}],"container-title":["ACM Transactions on Reconfigurable Technology and Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2629471","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/2629471","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T07:01:17Z","timestamp":1750230077000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2629471"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2015,1,23]]},"references-count":29,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2015,1,23]]}},"alternative-id":["10.1145\/2629471"],"URL":"https:\/\/doi.org\/10.1145\/2629471","relation":{},"ISSN":["1936-7406","1936-7414"],"issn-type":[{"type":"print","value":"1936-7406"},{"type":"electronic","value":"1936-7414"}],"subject":[],"published":{"date-parts":[[2015,1,23]]},"assertion":[{"value":"2013-06-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2014-01-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2015-01-23","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}