{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T04:51:52Z","timestamp":1750308712525,"version":"3.41.0"},"reference-count":61,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2012,12,1]],"date-time":"2012-12-01T00:00:00Z","timestamp":1354320000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/100000144","name":"Division of Computer and Network Systems","doi-asserted-by":"publisher","award":["CNS-0953447, CCR-0203829, and CCR-9876006"],"award-info":[{"award-number":["CNS-0953447, CCR-0203829, and CCR-9876006"]}],"id":[{"id":"10.13039\/100000144","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000001","name":"National Science Foundation","doi-asserted-by":"publisher","award":["CNS-0953447, CCR-0203829, and CCR-9876006"],"award-info":[{"award-number":["CNS-0953447, CCR-0203829, and CCR-9876006"]}],"id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Embed. Comput. Syst."],"published-print":{"date-parts":[[2012,12]]},"abstract":"<jats:p>The instruction cache is a popular optimization target due to the cache's high impact on system performance and power and because of the cache's predictable temporal and spatial locality. This article is an in depth study on the interaction of code reordering (a long-known technique) and cache configuration (a relatively new technique). Experimental results show that code reordering coupled with cache configuration reveals additional energy savings as high as 10--15% for several benchmarks with reduced cache area as high as 48%. To exploit these additional benefits, we architect and evaluate several design exploration heuristics for combining these two methods.<\/jats:p>","DOI":"10.1145\/2362336.2399177","type":"journal-article","created":{"date-parts":[[2013,1,11]],"date-time":"2013-01-11T15:42:48Z","timestamp":1357918968000},"page":"1-20","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["Combining code reordering and cache configuration"],"prefix":"10.1145","volume":"11","author":[{"given":"Ann","family":"Gordon-Ross","sequence":"first","affiliation":[{"name":"University of Florida"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Frank","family":"Vahid","sequence":"additional","affiliation":[{"name":"University of California, Riverside, CA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Nikil","family":"Dutt","sequence":"additional","affiliation":[{"name":"University of California, Irvine, CA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2013,1]]},"reference":[{"key":"e_1_2_1_1_1","unstructured":"Albonesi D. H. 2002. Selective cache ways: on demand cache resource allocation. J. Instruction Level Parallel.  Albonesi D. H. 2002. Selective cache ways: on demand cache resource allocation. J. Instruction Level Parallel."},{"key":"e_1_2_1_2_1","unstructured":"Altera. 2010. Nios embedded processor system development. http:\/\/www.altera.com\/corporate\/news_room\/releases\/products\/nr-nios_delivers_goods.html.  Altera. 2010. Nios embedded processor system development. http:\/\/www.altera.com\/corporate\/news_room\/releases\/products\/nr-nios_delivers_goods.html."},{"key":"e_1_2_1_3_1","unstructured":"Arc International 2010. www.arccores.com.  Arc International 2010. www.arccores.com."},{"key":"e_1_2_1_4_1","unstructured":"ARM. 2010. www.arm.com.  ARM. 2010. www.arm.com."},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/346023.346046"},{"volume-title":"Proceedings of the 3rd Workshop of Interaction Between Compilers and Computer Architecture.","author":"Bahar I.","key":"e_1_2_1_6_1"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/360128.360153"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/1113830.1113839"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/313817.313927"},{"key":"e_1_2_1_10_1","unstructured":"Burger D. Austin T. and Bennet S. 2000. Evaluating future microprocessors: The simplescalar toolset. Tech. rep. CS-TR-1308. Computer Science Department University of Wisconsin-Madison.  Burger D. Austin T. and Bennet S. 2000. Evaluating future microprocessors: The simplescalar toolset. Tech. rep. CS-TR-1308. Computer Science Department University of Wisconsin-Madison."},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/195473.195553"},{"volume-title":"Proceedings of the USENIX Windows NT Workshop.","author":"Chen J.","key":"e_1_2_1_12_1"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1145\/1250727.1250730"},{"volume-title":"Proceedings of the USENIX Windows NT Workshop.","author":"Cohn R.","key":"e_1_2_1_14_1"},{"key":"e_1_2_1_15_1","unstructured":"Cohn. R. and Lowney P. G. 2000. Design and analysis of profile-based optimization in Compaq's compilation tools for Alpha. J. Instruction Level Parallelism 2.  Cohn. R. and Lowney P. G. 2000. Design and analysis of profile-based optimization in Compaq's compilation tools for Alpha. J. Instruction Level Parallelism 2."},{"key":"e_1_2_1_16_1","unstructured":"Dinero I. 2010. http:\/\/www.cs.wisc.edu\/~markhill\/DineroIV\/.  Dinero I. 2010. http:\/\/www.cs.wisc.edu\/~markhill\/DineroIV\/."},{"key":"e_1_2_1_17_1","unstructured":"EEMBC. 2010. The Embedded Microprocessor Benchmark Consortium. www.eembc.org.  EEMBC. 2010. The Embedded Microprocessor Benchmark Consortium. www.eembc.org."},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.5555\/996070.1009912"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCAD.2002.804107"},{"volume-title":"Proceedings of the 30th Anual ACM\/IEEE International Symposium on Microarchitecture. 303--313","author":"Gloy N.","key":"e_1_2_1_20_1"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1109\/L-CA.2002.4"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/1366110.1366200"},{"volume-title":"Proceedings of the International Conference on Computer Design.","author":"Gordon-Ross A.","key":"e_1_2_1_23_1"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1109\/TVLSI.2008.2002459"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/258915.258931"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2007.18"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/1133956.1133980"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1145\/1134760.1134779"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/74925.74953"},{"key":"e_1_2_1_30_1","unstructured":"Kalmatianos J. and Kaeli D. 1999. Code reordering for multi-level cache hierarchies. Northeeastern University Computer Architecture Research Group. http:\/\/www.ece.neu.edu\/info\/architecture\/publications. html.  Kalmatianos J. and Kaeli D. 1999. Code reordering for multi-level cache hierarchies. Northeeastern University Computer Architecture Research Group. http:\/\/www.ece.neu.edu\/info\/architecture\/publications. html."},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.5555\/1153923.1154523"},{"volume-title":"Proceedings of the IEEE Micro.","author":"Kin J.","key":"e_1_2_1_32_1"},{"volume-title":"Proceedings of the Windows NT Symposium.","author":"Lee D.","key":"e_1_2_1_33_1"},{"key":"e_1_2_1_34_1","unstructured":"Lee L. H. Moyer W. and Arends J. 1999b. Low cost Embedded Program Loop Caching -- Revisited. Tech. rep. N CSE-TR-411-99 University of Michigan.  Lee L. H. Moyer W. and Arends J. 1999b. Low cost Embedded Program Loop Caching -- Revisited. Tech. rep. N CSE-TR-411-99 University of Michigan."},{"volume-title":"Proceedings of the 30th Annual International Symposium on Microarchitecture.","author":"Lee C.","key":"e_1_2_1_35_1"},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1145\/344166.344610"},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1145\/70082.68200"},{"key":"e_1_2_1_38_1","unstructured":"MIPS Technologies. 2010. www.mips.com.  MIPS Technologies. 2010. www.mips.com."},{"volume-title":"Proceedings of the 3rd IEEE International Workshop of Source Code Analysis and Manipulation.","author":"Moseley P.","key":"e_1_2_1_39_1"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1002\/1097-024X(200101)31:1%3C67::AID-SPE357%3E3.0.CO;2-A"},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1145\/774789.774804"},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1145\/93542.93550"},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1145\/1065010.1065025"},{"key":"e_1_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1023\/A:1019992713965"},{"volume-title":"Proceedings of the International Conference on Parallel Architectures and Compilation Techniques (PACT).","author":"Ramirez A.","key":"e_1_2_1_45_1"},{"key":"e_1_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.1109\/5.964440"},{"key":"e_1_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.1109\/TC.2005.13"},{"key":"e_1_2_1_48_1","unstructured":"Reinman G. and Jouppi N. P. 1999. Cacti2.0: An integraded cache timing and power model. Tech rep. COMPAQ Western Research Lab.  Reinman G. and Jouppi N. P. 1999. Cacti2.0: An integraded cache timing and power model. Tech rep. COMPAQ Western Research Lab."},{"key":"e_1_2_1_49_1","unstructured":"Samples A. D. and Hilfinger P. N. 1988. Code reorganization for instruction caches. Techn. rep. UCB\/CSD 88\/447 University of California Berkeley.   Samples A. D. and Hilfinger P. N. 1988. Code reorganization for instruction caches. Techn. rep. UCB\/CSD 88\/447 University of California Berkeley."},{"volume-title":"Proceedings of the Workshop on Optimizations for DSP and Embedded Systems (ODES).","author":"Sanghai K.","key":"e_1_2_1_50_1"},{"key":"e_1_2_1_51_1","unstructured":"Scales D. 1998. Efficient dynamic procedure placement. Tech. rep. WRL-98\/5 Compaq WRL Research Lab.  Scales D. 1998. Efficient dynamic procedure placement. Tech. rep. WRL-98\/5 Compaq WRL Research Lab."},{"volume-title":"Proceedings of the Workshop on Binary Translation (WBT).","author":"Scharz B.","key":"e_1_2_1_52_1"},{"key":"e_1_2_1_53_1","doi-asserted-by":"publisher","DOI":"10.1147\/sj.372.0270"},{"key":"e_1_2_1_54_1","first-page":"1","article-title":"A practical system of intermodule code optimization at link-time","volume":"11","author":"Srivastava A.","year":"1992","journal-title":"J. Program. Lang."},{"key":"e_1_2_1_55_1","doi-asserted-by":"publisher","DOI":"10.1145\/224081.224093"},{"key":"e_1_2_1_56_1","unstructured":"Tensilica. 2010. Xtensa processor generator. http:\/\/www.tensilica.com\/.  Tensilica. 2010. Xtensa processor generator. http:\/\/www.tensilica.com\/."},{"key":"e_1_2_1_57_1","unstructured":"Villarreal J. Lysecky R. Cotterell S. and Vahid F. 2001. Loop analysis of embedded applications. Tech. rep. UCR-CSR-01-03 University of California Riverside.  Villarreal J. Lysecky R. Cotterell S. and Vahid F. 2001. Loop analysis of embedded applications. Tech. rep. UCR-CSR-01-03 University of California Riverside."},{"volume-title":"Proceedings of the 14th IEEE International Workshop on Rapid System Prototyping (RSP- 03)","author":"Zhang C.","key":"e_1_2_1_58_1"},{"key":"e_1_2_1_59_1","doi-asserted-by":"publisher","DOI":"10.1145\/859618.859635"},{"volume-title":"Proceedings of the Design, Automation and Test (DATE) Conference in Europe.","author":"Zhang C.","key":"e_1_2_1_60_1"},{"volume-title":"Proceedings of the Design, Automation and Test (DATE) Conference in Europe.","author":"Zhang C.","key":"e_1_2_1_61_1"}],"container-title":["ACM Transactions on Embedded Computing Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2362336.2399177","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/2362336.2399177","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T20:14:16Z","timestamp":1750277656000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2362336.2399177"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2012,12]]},"references-count":61,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2012,12]]}},"alternative-id":["10.1145\/2362336.2399177"],"URL":"https:\/\/doi.org\/10.1145\/2362336.2399177","relation":{},"ISSN":["1539-9087","1558-3465"],"issn-type":[{"type":"print","value":"1539-9087"},{"type":"electronic","value":"1558-3465"}],"subject":[],"published":{"date-parts":[[2012,12]]},"assertion":[{"value":"2009-09-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2010-06-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2013-01-01","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}