{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T04:52:31Z","timestamp":1750308751117,"version":"3.41.0"},"reference-count":31,"publisher":"Association for Computing Machinery (ACM)","issue":"1","license":[{"start":{"date-parts":[[2007,2,1]],"date-time":"2007-02-01T00:00:00Z","timestamp":1170288000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Embed. Comput. Syst."],"published-print":{"date-parts":[[2007,2]]},"abstract":"<jats:p>Because of stringent power constraints, aggressive latency-hiding approaches, such as prefetching, are absent in the state-of-the-art embedded processors. There are two main reasons that make prefetching power inefficient. First, compiler-inserted prefetch instructions increase code size and, therefore, could increase I-cache power. Second, inaccurate prefetching (especially for hardware prefetching) leads to high D-cache power consumption because of useless accesses. In this work, we show that it is possible to support power-efficient prefetching through bit-differential offset assignment. We target the prefetching of relocatable stack variables with a high degree of precision. By assigning the offsets of stack variables in such a way that most consecutive addresses differ by 1 bit, we can prefetch them with compact prefetch instructions to save I-cache power. The compiler first generates an access graph of consecutive memory references and then attempts a layout of the memory locations in the smallest hypercube. Each dimension of the hypercube represents a 1-bit differential addressing. The embedding is carried out in as compact a hypercube as possible in order to save memory space. Each load\/store instruction carries a hint regarding prefetching the next memory reference by encoding its differential address with respect to the current one. To reduce D-cache power cost, we further attempt to assign offsets so that most of the consecutive accesses map to the same cache line. Our prefetching is done using a one entry line buffer [Wilson et al. 1996]. Consequently, many look-ups in D-cache reduce to incremental ones. This results in D-cache activity reduction and power savings. Our prefetcher requires both compiler and hardware support. In this paper, we provide implementation on the processor model close to ARM with small modification to the ISA. We tackle issues such as out-of-order commit, predication, and speculation through simple modifications to the processor pipeline on noncritical paths. Our goal in this work is to boost performance while maintaining\/lowering power consumption. Our results show 12% speedup and slight power reduction. The runtime virtual space loss for stack and static data is about 11.8%.<\/jats:p>","DOI":"10.1145\/1210268.1210271","type":"journal-article","created":{"date-parts":[[2007,4,5]],"date-time":"2007-04-05T19:20:08Z","timestamp":1175800808000},"page":"3","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":7,"title":["Power-efficient prefetching for embedded processors"],"prefix":"10.1145","volume":"6","author":[{"given":"Xiaotong","family":"Zhuang","sequence":"first","affiliation":[{"name":"Georgia Institute of Technology, Atlanta, GA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Santosh","family":"Pande","sequence":"additional","affiliation":[{"name":"Georgia Institute of Technology, Atlanta, GA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2007,2]]},"reference":[{"key":"e_1_2_1_1_1","unstructured":"ARM Co.Ltd ARM 7TDMI Data Sheet. ARM Co.Ltd ARM 7TDMI Data Sheet."},{"key":"e_1_2_1_2_1","unstructured":"ARM Co.Ltd ARM7500FE Data Sheet. ARM Co.Ltd ARM7500FE Data Sheet."},{"key":"e_1_2_1_3_1","unstructured":"Aho A. V. Sethi R. and Ullman J. D. 1986. . . . Techniques Addison-Wesley Reading MA. Aho A. V. Sethi R. and Ullman J. D. 1986. Compilers Principles Techniques and Tools Addison-Wesley Reading MA."},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1145\/339647.339657"},{"key":"e_1_2_1_5_1","unstructured":"Basu K. Choudhary A. Pisharath J. and Kandemir M. 2002. . . . . . . . . . . . . . . . MICRO (Nov.). Basu K. Choudhary A. Pisharath J. and Kandemir M. 2002. Power protocol: Reducing power dissipation on off-chip data buses. MICRO (Nov.)."},{"volume-title":"Tech. Report 1342, Univ. of Wisconsin--Madison (May).","year":"1997","author":"Burger D.","key":"e_1_2_1_6_1"},{"volume-title":"Proceedings of Architectural Support for Programming Languages and Operating Systems. (Oct). 10","author":"Calder B.","key":"e_1_2_1_7_1"},{"volume-title":"Proceedings of the Seventh International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS VII), 222--233 (Oct.). 10","author":"Luk Chi-Keung","key":"e_1_2_1_8_1"},{"key":"e_1_2_1_9_1","unstructured":"Cho S. Yew P. C. and Lee G. 1999. . . . . . . . . . . . . . . . (May). 10.1145\/300979.300988 Cho S. Yew P. C. and Lee G. 1999. Decoupling local variable accesses in a wide-issue superscalar processor. ISCA (May). 10.1145\/300979.300988"},{"volume-title":"IEEE 4th Annual Workshop on Workload Characterization. 10","year":"2001","author":"Guthaus M. R.","key":"e_1_2_1_10_1"},{"volume-title":"Proceedings of International Symposium on Code Generation and Optimization.","author":"Haber G.","key":"e_1_2_1_11_1"},{"key":"e_1_2_1_12_1","unstructured":"Intel Corp. SA-110 Microprocessor Tech. Ref. Manual. Intel Corp. SA-110 Microprocessor Tech. Ref. Manual."},{"key":"e_1_2_1_13_1","unstructured":"Lee H. S. Smelyanskiy M. Newburn C. J. and Tyson G. S. 2001. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . (Jan). Lee H. S. Smelyanskiy M. Newburn C. J. and Tyson G. S. 2001. Stack value file: Custom microarchitecture for the stack. HPCA-7 (Jan)."},{"volume-title":"Proc. ICCAD.","author":"Leupers R.","key":"e_1_2_1_14_1"},{"key":"e_1_2_1_15_1","doi-asserted-by":"crossref","unstructured":"Liao S. Devadas S. Keutzer K. Tjiang S. and Wang A. 1996. . . . . . . . . . . . . . . . (May) 235--253. 10.1145\/229542.229543 Liao S. Devadas S. Keutzer K. Tjiang S. and Wang A. 1996. Storage assignment to decrease code size. ACM TOPLAS 18 3 (May) 235--253. 10.1145\/229542.229543","DOI":"10.1145\/229542.229543"},{"key":"e_1_2_1_16_1","doi-asserted-by":"crossref","unstructured":"Noth W. and Kolla R. 1990. . . . DATE 168--174. Noth W. and Kolla R. 1990. Spanning tree-based state encoding for low-power dissipation. DATE 168--174.","DOI":"10.1109\/DATE.1999.761114"},{"volume-title":"Proceedings of the 28th Annual IEEE\/ACM International Symposium on Microarchi-tecture (MICRO 28)","author":"Lipasti M. H.","key":"e_1_2_1_17_1"},{"volume-title":"Proceedings of the 28th Annual IEEE\/ACM International Symposium on Microarchitecture. (MICRO 28)","author":"Ozawa T.","key":"e_1_2_1_18_1"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/92.784092"},{"key":"e_1_2_1_20_1","unstructured":"Perez D. G. Mouchard G. and Temam O. 2004. . . . A case for the quantitative comparison of micro-architecture mechanisms. MICRO 37. 10.1109\/MICRO.2004.25 Perez D. G. Mouchard G. and Temam O. 2004. MicroLib: A case for the quantitative comparison of micro-architecture mechanisms. MICRO 37. 10.1109\/MICRO.2004.25"},{"key":"e_1_2_1_21_1","unstructured":"Pomerene J. Puzak T. Rechtschaffen R. and Sparacio F. 1989. . . . . . . . . . . . . . . . (Feb.). Pomerene J. Puzak T. Rechtschaffen R. and Sparacio F. 1989. Prefetching system for a cache having a second directory for sequentially accessed blocks. U. S. Patent number 4 807 110 (Feb.)."},{"key":"e_1_2_1_22_1","doi-asserted-by":"crossref","unstructured":"Rao A. and Pande S. 1999. . . . . . . . . . . . . . . . In ACM (PLDI). 128--138. 10.1145\/301618.301653 Rao A. and Pande S. 1999. Storage assignment optimizations to generate compact and efficient code on embedded DSPs. In ACM (PLDI). 128--138. 10.1145\/301618.301653","DOI":"10.1145\/301631.301653"},{"key":"e_1_2_1_23_1","unstructured":"Segars S. 2001. Low power design techniques for microprocessors. ISSCC (Feb.). Segars S. 2001. Low power design techniques for microprocessors. ISSCC (Feb.)."},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1145\/356887.356892"},{"key":"e_1_2_1_25_1","unstructured":"Udayanarayanan S. and Chakrabarti C. 2001. . . . . . . . . . . . . . . . DAC. Udayanarayanan S. and Chakrabarti C. 2001. Address code generation for DSPs. DAC."},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1137\/0219038"},{"volume-title":"MICRO'01 (Dec.).","author":"Witchel E.","key":"e_1_2_1_27_1"},{"key":"e_1_2_1_28_1","doi-asserted-by":"crossref","unstructured":"Wilson K. M. Olukotun K. and Rosenblum M. 1996. . . . ISCA. 10.1145\/232973.232989 Wilson K. M. Olukotun K. and Rosenblum M. 1996. Increasing cache port efficiency for dynamic superscalar microprocessors. ISCA. 10.1145\/232973.232989","DOI":"10.1145\/232973.232989"},{"volume-title":"Technical Report TN93\/5, Compaq Western Research Lab.","year":"1993","author":"Wilton S.","key":"e_1_2_1_29_1"},{"volume-title":"Proceedings of ACM SIGPLAN Conference on Languages, Compiler, and Tools for Embedded Systems (LCTES-03)","author":"Zhuang X.","key":"e_1_2_1_30_1"},{"volume-title":"Proceedings of ACM SIGPLAN Conference on Languages, Compiler, and Tools for Embedded Systems (LCTES-04)","author":"Zhuang X.","key":"e_1_2_1_31_1"}],"container-title":["ACM Transactions on Embedded Computing Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/1210268.1210271","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/1210268.1210271","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T20:22:21Z","timestamp":1750278141000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/1210268.1210271"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2007,2]]},"references-count":31,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2007,2]]}},"alternative-id":["10.1145\/1210268.1210271"],"URL":"https:\/\/doi.org\/10.1145\/1210268.1210271","relation":{},"ISSN":["1539-9087","1558-3465"],"issn-type":[{"type":"print","value":"1539-9087"},{"type":"electronic","value":"1558-3465"}],"subject":[],"published":{"date-parts":[[2007,2]]},"assertion":[{"value":"2007-02-01","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}