{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,12]],"date-time":"2026-02-12T17:32:18Z","timestamp":1770917538836,"version":"3.50.1"},"reference-count":38,"publisher":"Association for Computing Machinery (ACM)","issue":"1","license":[{"start":{"date-parts":[[2005,3,1]],"date-time":"2005-03-01T00:00:00Z","timestamp":1109635200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Archit. Code Optim."],"published-print":{"date-parts":[[2005,3]]},"abstract":"<jats:p>Caches contribute to much of a microprocessor system's power and energy consumption. Numerous new cache architectures, such as phased, pseudo-set-associative, way predicting, reactive-associative, way-shutdown, way-concatenating, and highly-associative, are intended to reduce power and\/or energy, but they all impose some performance overhead. We have developed a new cache architecture, called a way-halting cache, that reduces energy further than previously mentioned architectures, while imposing no performance overhead. Our way-halting cache is a four-way set-associative cache that stores the four lowest-order bits of all ways' tags into a fully associative memory, which we call the halt tag array. The lookup in the halt tag array is done in parallel with, and is no slower than, the set-index decoding. The halt tag array predetermines which tags cannot match due to their low-order 4 bits mismatching. Further accesses to ways with known mismatching tags are then halted, thus saving power. Our halt tag array has an additional feature of using static logic only, rather than dynamic logic used in highly associative caches, making our cache simpler to design with existing tools. We provide data from experiments on 29 benchmarks drawn from Powerstone, Mediabench, and Spec 2000, based on our layouts in 0.18 micron CMOS technology. On average, we obtained 55% savings of memory-access related energy over a conventional four-way set-associative cache. We show that savings are greater than previous methods, and nearly twice that of highly associative caches, while imposing no performance overhead and only 2% cache area overhead.<\/jats:p>","DOI":"10.1145\/1061267.1061270","type":"journal-article","created":{"date-parts":[[2005,8,1]],"date-time":"2005-08-01T17:31:42Z","timestamp":1122917502000},"page":"34-54","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":67,"title":["A way-halting cache for low-energy high-performance systems"],"prefix":"10.1145","volume":"2","author":[{"given":"Chuanjun","family":"Zhang","sequence":"first","affiliation":[{"name":"San Diego State University, San Diego, CA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Frank","family":"Vahid","sequence":"additional","affiliation":[{"name":"University of California, Riverside"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jun","family":"Yang","sequence":"additional","affiliation":[{"name":"University of California, Riverside"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Walid","family":"Najjar","sequence":"additional","affiliation":[{"name":"University of California, Riverside"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2005,3]]},"reference":[{"key":"e_1_2_1_1_1","unstructured":"Advanced Micro Devices. http:\/\/www.amd.com.  Advanced Micro Devices. http:\/\/www.amd.com."},{"key":"e_1_2_1_2_1","article-title":"Selective cache ways: On-demand cache resource allocation","author":"Albonesi D. H.","year":"2000","unstructured":"Albonesi , D. H. 2000 . Selective cache ways: On-demand cache resource allocation . Journal of Instruction Level Parallelism. Albonesi, D. H. 2000. Selective cache ways: On-demand cache resource allocation. Journal of Instruction Level Parallelism.","journal-title":"Journal of Instruction Level Parallelism."},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1109\/4.705359"},{"key":"e_1_2_1_4_1","volume-title":"International Conference on Parallel Architectures and Compilation Techniques.","author":"Batson B.","unstructured":"Batson , B. and Vijaykumar , T. N . 2000. Reactive-associative caches . In International Conference on Parallel Architectures and Compilation Techniques. Batson, B. and Vijaykumar, T. N. 2000. Reactive-associative caches. In International Conference on Parallel Architectures and Compilation Techniques."},{"key":"e_1_2_1_5_1","doi-asserted-by":"crossref","unstructured":"Burger D. and Austin T. M. 1997. The SimpleScalar tool set version 2.0. University of Wisconsin-Madison Computer Sciences Dept. Technical Report &num;1342.  Burger D. and Austin T. M. 1997. The SimpleScalar tool set version 2.0. University of Wisconsin-Madison Computer Sciences Dept. Technical Report &num;1342.","DOI":"10.1145\/268806.268810"},{"key":"e_1_2_1_6_1","unstructured":"Cadence. http:\/\/www.cadence.com.  Cadence. http:\/\/www.cadence.com."},{"key":"e_1_2_1_7_1","volume-title":"International Symposium on High Performance Computer Architecture.","author":"Calder B.","unstructured":"Calder , B. , Grunwall , D. , and Emer , J . 1996. Predictive sequential associative cache . In International Symposium on High Performance Computer Architecture. Calder, B., Grunwall, D., and Emer, J. 1996. Predictive sequential associative cache. In International Symposium on High Performance Computer Architecture."},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/40.621215"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.5555\/211554.211583"},{"key":"e_1_2_1_10_1","volume-title":"Proceedings of the International Symposium on Low Power Electronics and Design. 10","author":"Efthymiou A.","unstructured":"Efthymiou , A. and Garside , J. D . 2002. An adaptive serial-parallel CAM architecture for low-power cache blocks . In Proceedings of the International Symposium on Low Power Electronics and Design. 10 .1145\/566408.566445 Efthymiou, A. and Garside, J. D. 2002. An adaptive serial-parallel CAM architecture for low-power cache blocks. In Proceedings of the International Symposium on Low Power Electronics and Design. 10.1145\/566408.566445"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/54.914617"},{"key":"e_1_2_1_12_1","volume-title":"2nd International Symposium on Advanced Research in Asynchronous Circuits and Systems.","author":"Garside J. D.","unstructured":"Garside , J. D. , Temple , S. , and Mehra , R . 1996. The AMULET2e cache system . In 2nd International Symposium on Advanced Research in Asynchronous Circuits and Systems. Garside, J. D., Temple, S., and Mehra, R. 1996. The AMULET2e cache system. In 2nd International Symposium on Advanced Research in Asynchronous Circuits and Systems."},{"key":"e_1_2_1_13_1","doi-asserted-by":"crossref","unstructured":"Hasegawa A. Kawasaki I. Yamada K. Yoshioka S. Kawasaki S. and Biswas P. 1995. SH3: High code density low power. IEEE Micro Dec. 10.1109\/40.476254   Hasegawa A. Kawasaki I. Yamada K. Yoshioka S. Kawasaki S. and Biswas P. 1995. SH3: High code density low power. IEEE Micro Dec. 10.1109\/40.476254","DOI":"10.1109\/40.476254"},{"key":"e_1_2_1_14_1","volume-title":"Computer Architecture: A Quantitative Approach","author":"Hennessy J. L.","year":"2002","unstructured":"Hennessy , J. L. and Patterson , D. A . 2002 . Computer Architecture: A Quantitative Approach , 3 rd ed., International Student Edition. Morgan Kaufman , San Mateo, CA. Hennessy, J. L. and Patterson, D. A. 2002. Computer Architecture: A Quantitative Approach, 3rd ed., International Student Edition. Morgan Kaufman, San Mateo, CA.","edition":"3"},{"key":"e_1_2_1_15_1","volume-title":"International Symposium on Low Power Electronics and Design. 10","author":"Huang M.","unstructured":"Huang , M. , Renau , J. , Yoo , S. M. , and Torrellas , J . 2001. L1 data cache decomposition for energy efficiency . In International Symposium on Low Power Electronics and Design. 10 .1145\/383082.383086 Huang, M., Renau, J., Yoo, S. M., and Torrellas, J. 2001. L1 data cache decomposition for energy efficiency. In International Symposium on Low Power Electronics and Design. 10.1145\/383082.383086"},{"key":"e_1_2_1_16_1","unstructured":"IBM. http:\/\/www.ibm.com.  IBM. http:\/\/www.ibm.com."},{"key":"e_1_2_1_17_1","unstructured":"http:\/\/www.specbench.org\/osg\/cpu2000\/.  http:\/\/www.specbench.org\/osg\/cpu2000\/."},{"key":"e_1_2_1_18_1","volume-title":"International Symposium on Low Power Electronics and Design. 10","author":"Inoue K.","unstructured":"Inoue , K. , Ishihara , T. , and Murakami , K . 1999. Way-predictive set-associative cache for high performance and low energy consumption . In International Symposium on Low Power Electronics and Design. 10 .1145\/313817.313948 Inoue, K., Ishihara, T., and Murakami, K. 1999. Way-predictive set-associative cache for high performance and low energy consumption. In International Symposium on Low Power Electronics and Design. 10.1145\/313817.313948"},{"key":"e_1_2_1_19_1","volume-title":"Proceedings of the 27th Annual International Symposium on Computer Architecture. 10","author":"Juan","unstructured":"Juan , Lang, T., and Navarro , J . 1996. The difference-bit cache . In Proceedings of the 27th Annual International Symposium on Computer Architecture. 10 .1145\/232973.232986 Juan, Lang, T., and Navarro, J. 1996. The difference-bit cache. In Proceedings of the 27th Annual International Symposium on Computer Architecture. 10.1145\/232973.232986"},{"key":"e_1_2_1_20_1","volume-title":"International Symposium on Microarchitecture.","author":"Lee C.","unstructured":"Lee , C. , Potkonjak , M. , and Mangione-Smith , W . 1997. MediaBench: A tool for evaluating and synthesizing multimedia and communications systems . In International Symposium on Microarchitecture. Lee, C., Potkonjak, M., and Mangione-Smith, W. 1997. MediaBench: A tool for evaluating and synthesizing multimedia and communications systems. In International Symposium on Microarchitecture."},{"key":"e_1_2_1_21_1","volume-title":"Proceedings of the 27th Annual International Symposium on Microarchitecture. 10","author":"Liu L.","year":"1994","unstructured":"Liu , L. 1994 . Cache design with partial address matching . In Proceedings of the 27th Annual International Symposium on Microarchitecture. 10 .1145\/192724.192742 Liu, L. 1994. Cache design with partial address matching. In Proceedings of the 27th Annual International Symposium on Microarchitecture. 10.1145\/192724.192742"},{"key":"e_1_2_1_22_1","volume-title":"International Symposium on Low Power Electronics and Design. 10","author":"Malik A.","unstructured":"Malik , A. , Moyer , B. , and Cermak , D . 2000. A low power unified cache architecture providing power and poerformance flexibility . In International Symposium on Low Power Electronics and Design. 10 .1145\/344166.344610 Malik, A., Moyer, B., and Cermak, D. 2000. A low power unified cache architecture providing power and poerformance flexibility. In International Symposium on Low Power Electronics and Design. 10.1145\/344166.344610"},{"key":"e_1_2_1_23_1","unstructured":"MIPS Technologies Inc. http:\/\/www.mips.com.  MIPS Technologies Inc. http:\/\/www.mips.com."},{"key":"e_1_2_1_24_1","volume-title":"IEEE International Solid-State Circuits Conference.","author":"Montanaro J.","unstructured":"Montanaro , J. , Witek , R. T. , Anne , K. , Black , A. J. , Cooper , E. M. , Dobberpuhl , D. W. , Donahue , P. M. , Eno , J. , Farell , A. , Hoeppner , G. W. , Kruckemyer , D. , Lee , T. H. , Lin , P. , Madden , L. , Murray , D. , Pearce , M. , Santhanam , S. , Snyder , K. J. , Stephany , R. , and Thierauf , S. C . 1996. A 160 MHz 32 b 0.5 W CMOS RISC microprocessor . In IEEE International Solid-State Circuits Conference. Montanaro, J., Witek, R. T., Anne, K., Black, A. J., Cooper, E. M., Dobberpuhl, D. W., Donahue, P. M., Eno, J., Farell, A., Hoeppner, G. W., Kruckemyer, D., Lee, T. H., Lin, P., Madden, L., Murray, D., Pearce, M., Santhanam, S., Snyder, K. J., Stephany, R., and Thierauf, S. C. 1996. A 160 MHz 32 b 0.5 W CMOS RISC microprocessor. In IEEE International Solid-State Circuits Conference."},{"key":"e_1_2_1_25_1","unstructured":"The Mosis Service. http:\/\/www.mosis.org.  The Mosis Service. http:\/\/www.mosis.org."},{"key":"e_1_2_1_26_1","volume-title":"International Symposium on System Synthesis. 10","author":"Petrov P.","unstructured":"Petrov , P. and Orailoglu , A . 2001. Data cache energy minimizations through programmable tag size matching to the applications . In International Symposium on System Synthesis. 10 .1145\/500001.500028 Petrov, P. and Orailoglu, A. 2001. Data cache energy minimizations through programmable tag size matching to the applications. In International Symposium on System Synthesis. 10.1145\/500001.500028"},{"key":"e_1_2_1_27_1","volume-title":"International Symposium on Microarchitecture.","author":"Powell M.","unstructured":"Powell , M. , Agarwal , A. , Vijaykumar , T. N. , Falsafi , B. , and Roy , K . 2001. Reducing set-associative cache energy via way-prediction and selective direct-mapping . In International Symposium on Microarchitecture. Powell, M., Agarwal, A., Vijaykumar, T. N., Falsafi, B., and Roy, K. 2001. Reducing set-associative cache energy via way-prediction and selective direct-mapping. In International Symposium on Microarchitecture."},{"key":"e_1_2_1_28_1","first-page":"57","volume-title":"SLPE","author":"Panwar R.","unstructured":"Panwar , R. and Rennels , D . 1995. Reducing the frequency of tag compares for low power I-cache design . In SLPE , pp. 57 -- 62 . 10.1145\/224081.224092 Panwar, R. and Rennels, D. 1995. Reducing the frequency of tag compares for low power I-cache design. In SLPE, pp. 57--62. 10.1145\/224081.224092"},{"key":"e_1_2_1_29_1","unstructured":"Reinmann G. and Jouppi N. P. 1999. CACTI2.0: An Integrated Cache Timing and Power Model. COMPAQ Western Research Lab.  Reinmann G. and Jouppi N. P. 1999. CACTI2.0: An Integrated Cache Timing and Power Model. COMPAQ Western Research Lab."},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1109\/4.726584"},{"key":"e_1_2_1_31_1","volume-title":"International Solid-State Circuits Conference Tutorial.","author":"Segars S.","year":"2000","unstructured":"Segars , S. 2000 . Low power design techniques for microprocessors . In International Solid-State Circuits Conference Tutorial. Segars, S. 2000. Low power design techniques for microprocessors. In International Solid-State Circuits Conference Tutorial."},{"key":"e_1_2_1_32_1","volume-title":"Proceedings of the 17th Annual International Symposium on Computer Architecture. 10","author":"Taylor G.","unstructured":"Taylor , G. , Davis , P. , and Farmwald , M . 1990. The TLB slice---A low-cost high-speed address translation mechanisms . In Proceedings of the 17th Annual International Symposium on Computer Architecture. 10 .1145\/325164.325161 Taylor, G., Davis, P., and Farmwald, M. 1990. The TLB slice---A low-cost high-speed address translation mechanisms. In Proceedings of the 17th Annual International Symposium on Computer Architecture. 10.1145\/325164.325161"},{"key":"e_1_2_1_33_1","volume-title":"International Symposium on Microarchitecture.","author":"Witchel E.","unstructured":"Witchel , E. , Larsen , S. , Ananian , C. S. , and Asanovic , K . 2001. Direct addressed caches for reduced power consumption . In International Symposium on Microarchitecture. Witchel, E., Larsen, S., Ananian, C. S., and Asanovic, K. 2001. Direct addressed caches for reduced power consumption. In International Symposium on Microarchitecture."},{"key":"e_1_2_1_34_1","volume-title":"International Symposium on Microarchitecture.","author":"Yang J.","unstructured":"Yang , J. and Gupta , R . 2002. Energy efficient frequent value data cache design . In International Symposium on Microarchitecture. Yang, J. and Gupta, R. 2002. Energy efficient frequent value data cache design. In International Symposium on Microarchitecture."},{"key":"e_1_2_1_35_1","volume-title":"International Symposium on Low Power Electronics and Design. 10","author":"Zhang C.","unstructured":"Zhang , C. , Vahid , F. , Yang , J. , and Najjar , W . 2004. A way-halting cache for low-energy high performance systems . In International Symposium on Low Power Electronics and Design. 10 .1145\/1013235.1013272 Zhang, C., Vahid, F., Yang, J., and Najjar, W. 2004. A way-halting cache for low-energy high performance systems. In International Symposium on Low Power Electronics and Design. 10.1145\/1013235.1013272"},{"key":"e_1_2_1_36_1","volume-title":"International Symposium on Computer Architecture. 10","author":"Zhang C.","unstructured":"Zhang , C. , Vahid , F. , and Najjar , W . 2003. A highly-configurable cache architecture for embedded systems . In International Symposium on Computer Architecture. 10 .1145\/859618.859635 Zhang, C., Vahid, F., and Najjar, W. 2003. A highly-configurable cache architecture for embedded systems. In International Symposium on Computer Architecture. 10.1145\/859618.859635"},{"key":"e_1_2_1_37_1","unstructured":"Zhang C. Vahid F. and Najjar W. 2005. A highly-configurable cache architecture for embedded systems. ACM Transactions on Embedded Computing Systems. 10.1145\/1067915.1067921   Zhang C. Vahid F. and Najjar W. 2005. A highly-configurable cache architecture for embedded systems. ACM Transactions on Embedded Computing Systems. 10.1145\/1067915.1067921"},{"key":"e_1_2_1_38_1","volume-title":"Kool Chips Workshop, in conjunction with International Symposium on Microarchitecture.","author":"Zhang M.","unstructured":"Zhang , M. and Asanovic , K . 2000. Highly-associative caches for low-power processors . In Kool Chips Workshop, in conjunction with International Symposium on Microarchitecture. Zhang, M. and Asanovic, K. 2000. Highly-associative caches for low-power processors. In Kool Chips Workshop, in conjunction with International Symposium on Microarchitecture."}],"container-title":["ACM Transactions on Architecture and Code Optimization"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/1061267.1061270","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/1061267.1061270","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T21:36:56Z","timestamp":1750282616000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/1061267.1061270"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2005,3]]},"references-count":38,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2005,3]]}},"alternative-id":["10.1145\/1061267.1061270"],"URL":"https:\/\/doi.org\/10.1145\/1061267.1061270","relation":{},"ISSN":["1544-3566","1544-3973"],"issn-type":[{"value":"1544-3566","type":"print"},{"value":"1544-3973","type":"electronic"}],"subject":[],"published":{"date-parts":[[2005,3]]},"assertion":[{"value":"2005-03-01","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}