{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T04:50:22Z","timestamp":1750308622274,"version":"3.41.0"},"reference-count":29,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2011,10,1]],"date-time":"2011-10-01T00:00:00Z","timestamp":1317427200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"Republic of Cyprus and the European Regional Development Fund","award":["T\u220fE\/\u220f\u2227HPO\/0609(BIE)\/09"],"award-info":[{"award-number":["T\u220fE\/\u220f\u2227HPO\/0609(BIE)\/09"]}]},{"DOI":"10.13039\/501100001810","name":"Research Promotion Foundation","doi-asserted-by":"publisher","award":["DESMI 2009-10"],"award-info":[{"award-number":["DESMI 2009-10"]}],"id":[{"id":"10.13039\/501100001810","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Archit. Code Optim."],"published-print":{"date-parts":[[2011,10]]},"abstract":"<jats:p>Cache-content-duplication (CCD) occurs when there is a miss for a block in a cache and the entire content of the missed block is already in the cache in a block with a different tag. Caches aware of content-duplication can have lower miss penalty by fetching, on a miss to a duplicate block, directly from the cache instead of accessing lower in the memory hierarchy, and can have lower miss rates by allowing only blocks with unique content to enter a cache.<\/jats:p>\n          <jats:p>This work examines the potential of CCD for instruction caches. We show that CCD is a frequent phenomenon and that an idealized duplication-detection mechanism for instruction caches has the potential to increase performance of an out-of-order processor, with a 16KB, 8-way, 8 instructions per block instruction cache, often by more than 10% and up to 36%.<\/jats:p>\n          <jats:p>This work also proposes CATCH, a hardware mechanism for dynamically detecting CCD for instruction caches. Experimental results for an out-of-order processor show that a duplication-detection mechanism with a 1.38KB cost captures on average 58% of the CCD's idealized potential.<\/jats:p>","DOI":"10.1145\/2019608.2019610","type":"journal-article","created":{"date-parts":[[2011,10,18]],"date-time":"2011-10-18T13:01:58Z","timestamp":1318942918000},"page":"1-27","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":2,"title":["CATCH"],"prefix":"10.1145","volume":"8","author":[{"given":"Marios","family":"Kleanthous","sequence":"first","affiliation":[{"name":"University of Cyprus, Nicosia, Cyprus"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yiannakis","family":"Sazeides","sequence":"additional","affiliation":[{"name":"University of Cyprus, Nicosia, Cyprus"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2011,10,18]]},"reference":[{"volume-title":"Proceedings of the 31st International Symposium on Computer Architecture. 212--223","author":"Alameldeen A. R.","key":"e_1_2_1_1_1","unstructured":"Alameldeen , A. R. and Wood , D. A . 2004. Adaptive cache compression for high-performance processors . In Proceedings of the 31st International Symposium on Computer Architecture. 212--223 . Alameldeen, A. R. and Wood, D. A. 2004. Adaptive cache compression for high-performance processors. In Proceedings of the 31st International Symposium on Computer Architecture. 212--223."},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1145\/125826.125932"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1109\/PACT.2005.38"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1145\/313817.313927"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/937503.937504"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/1555754.1555777"},{"key":"e_1_2_1_7_1","unstructured":"Compaq C. C. 1998. Alpha Architecture Handbook.  Compaq C. C. 1998. Alpha Architecture Handbook."},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/223982.224444"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/301618.301655"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/349214.349233"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/356571.356573"},{"key":"e_1_2_1_12_1","volume-title":"Proceedings of the USENIX Conference. 519--529","author":"Douglis F.","year":"1993","unstructured":"Douglis , F. 1993 . The Compression cache: Using on-line compression to extend physical memory . In Proceedings of the USENIX Conference. 519--529 . Douglis, F. 1993. The Compression cache: Using on-line compression to extend physical memory. In Proceedings of the USENIX Conference. 519--529."},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1145\/1054943.1054945"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA.2005.32"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/325164.325162"},{"volume-title":"Proceedings of the 22nd EUROMICRO Conference. 423--430","author":"Kjelso M.","key":"e_1_2_1_16_1","unstructured":"Kjelso , M. , Gooch , M. , and Jones , S . 1996. Design and performance of a main memory hardware data compressor . In Proceedings of the 22nd EUROMICRO Conference. 423--430 . Kjelso, M., Gooch, M., and Jones, S. 1996. Design and performance of a main memory hardware data compressor. In Proceedings of the 22nd EUROMICRO Conference. 423--430."},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/1403375.1403720"},{"volume-title":"Proceedings of the 8th International Symposium on Static Analysis. 40--56","author":"Komondoor R.","key":"e_1_2_1_18_1","unstructured":"Komondoor , R. and Horwitz , S . 2001. Using slicing to identify duplication in source code . In Proceedings of the 8th International Symposium on Static Analysis. 40--56 . Komondoor, R. and Horwitz, S. 2001. Using slicing to identify duplication in source code. In Proceedings of the 8th International Symposium on Static Analysis. 40--56."},{"volume-title":"Proceedings of the 30th Annual ACM\/IEEE International Symposium on Microarchitecture. 194--203","author":"Lefurgy C.","key":"e_1_2_1_19_1","unstructured":"Lefurgy , C. , Bird , P. , Chen , I.-C. , and Mudge , T . 1997. Improving code density using compression techniques . In Proceedings of the 30th Annual ACM\/IEEE International Symposium on Microarchitecture. 194--203 . Lefurgy, C., Bird, P., Chen, I.-C., and Mudge, T. 1997. Improving code density using compression techniques. In Proceedings of the 30th Annual ACM\/IEEE International Symposium on Microarchitecture. 194--203."},{"volume-title":"Proceedings of the 6th International Symposium on High Performance Computer Architecture. 218--228","author":"Lefurgy C.","key":"e_1_2_1_20_1","unstructured":"Lefurgy , C. , Piccininni , E. , and Mudge , T . 2000. Reducing code size with run-time decompression . In Proceedings of the 6th International Symposium on High Performance Computer Architecture. 218--228 . Lefurgy, C., Piccininni, E., and Mudge, T. 2000. Reducing code size with run-time decompression. In Proceedings of the 6th International Symposium on High Performance Computer Architecture. 218--228."},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/237090.237173"},{"key":"e_1_2_1_22_1","unstructured":"Malamy A. Patel R. N. and Hayes N. M. 1994. Methods and aparatus for implementing a pseudo-LRU cache memory replacement scheme with a locking feature. United States Patent 5353425.  Malamy A. Patel R. N. and Hayes N. M. 1994. Methods and aparatus for implementing a pseudo-LRU cache memory replacement scheme with a locking feature. United States Patent 5353425."},{"key":"e_1_2_1_23_1","unstructured":"Muralimanohar N. Balasubramonian R. and Jouppi N. P. 2009. CACTI 6.0: A tool to model large caches. Tech. rep. HPL-2009-85 HP Laboratories.  Muralimanohar N. Balasubramonian R. and Jouppi N. P. 2009. CACTI 6.0: A tool to model large caches. Tech. rep. HPL-2009-85 HP Laboratories."},{"key":"e_1_2_1_24_1","doi-asserted-by":"crossref","unstructured":"Postiff M. A. and Mudge T. 1999. Smart register file for high-performance microprocessors. Tech. rep. CSE-TR-403-99 University of Michigan.  Postiff M. A. and Mudge T. 1999. Smart register file for high-performance microprocessors. Tech. rep. CSE-TR-403-99 University of Michigan.","DOI":"10.21236\/ADA459519"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1109\/L-CA.2003.3"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1145\/605397.605403"},{"key":"e_1_2_1_27_1","volume-title":"Proceedings of the 22nd Annual Computer Measurement Group Conference. 819--828","author":"Tullsen D. M.","year":"1996","unstructured":"Tullsen , D. M. 1996 . Simulation and modeling of a simultaneous multithreading processor . In Proceedings of the 22nd Annual Computer Measurement Group Conference. 819--828 . Tullsen, D. M. 1996. Simulation and modeling of a simultaneous multithreading processor. In Proceedings of the 22nd Annual Computer Measurement Group Conference. 819--828."},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1145\/216585.216588"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/378993.379235"}],"container-title":["ACM Transactions on Architecture and Code Optimization"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2019608.2019610","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/2019608.2019610","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T19:07:42Z","timestamp":1750273662000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2019608.2019610"}},"subtitle":["A mechanism for dynamically detecting cache-content-duplication in instruction caches"],"short-title":[],"issued":{"date-parts":[[2011,10]]},"references-count":29,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2011,10]]}},"alternative-id":["10.1145\/2019608.2019610"],"URL":"https:\/\/doi.org\/10.1145\/2019608.2019610","relation":{},"ISSN":["1544-3566","1544-3973"],"issn-type":[{"type":"print","value":"1544-3566"},{"type":"electronic","value":"1544-3973"}],"subject":[],"published":{"date-parts":[[2011,10]]},"assertion":[{"value":"2010-03-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2011-06-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2011-10-18","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}