{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T04:36:35Z","timestamp":1750307795330,"version":"3.41.0"},"reference-count":19,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2007,9,1]],"date-time":"2007-09-01T00:00:00Z","timestamp":1188604800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100003176","name":"Ministerio de Educaci\u00f3n, Cultura y Deporte","doi-asserted-by":"publisher","award":["TIN-2004-07739-C02-01AP2003-3682"],"award-info":[{"award-number":["TIN-2004-07739-C02-01AP2003-3682"]}],"id":[{"id":"10.13039\/501100003176","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["SIGARCH Comput. Archit. News"],"published-print":{"date-parts":[[2007,9]]},"abstract":"<jats:p>To alleviate the memory wall problem, current architectural trends suggest implementing large instruction windows able to maintain a high number of in-flight instructions. However, the benefits achieved by these recent proposals may be limited because more instructions are executed down the wrong path of a mispredicted branch. The larger number of misspeculated instructions involves increasing the energy consumed compared to traditional designs with smaller instruction windows. Our analysis shows that, for some SPEC2000 integer benchmarks, up to 2, 5X wrong-path load instructions are executed when the instruction window of a 4-way superscalar processor is increased from 256 to 1024 entries.<\/jats:p>\n          <jats:p>This paper describes a simple speculative control technique to prevent wrong-path load instructions from being executed. Our technique extends the functionality of the load-store queue to block those load instructions that depend on a hard-to-predict conditional branch until it is resolved. If the branch is actually mispredicted, unnecessary cache misses can be avoided, saving energy down the wrong path. Furthermore, instructions that depend on a blocked load are not issued because their source values are not available, which also saves dynamic energy. Our results show that the proposed mechanism reduces, on average, up to 26% misspeculated load instructions and 18% wrong-path instructions without any performance loss.<\/jats:p>","DOI":"10.1145\/1327312.1327318","type":"journal-article","created":{"date-parts":[[2007,12,21]],"date-time":"2007-12-21T14:52:36Z","timestamp":1198248756000},"page":"29-36","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["Energy saving through a simple load control mechanism"],"prefix":"10.1145","volume":"35","author":[{"given":"Tanaus\u00fa","family":"Ram\u00edrez","sequence":"first","affiliation":[{"name":"DAC - UPC, Campus Nord, Barcelona, Spain"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Alex","family":"Pajuelo","sequence":"additional","affiliation":[{"name":"DAC - UPC, Campus Nord, Barcelona, Spain"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Oliverio J.","family":"Santana","sequence":"additional","affiliation":[{"name":"DIS - ULPGC, Campus Tafira, Las Palmas de GC, Spain"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Mateo","family":"Valero","sequence":"additional","affiliation":[{"name":"DAC - UPC, Campus Nord, Barcelona, Spain"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2007,9]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.5555\/956417.956554"},{"key":"e_1_2_1_2_1","volume-title":"Workshop on Performance Analysis and its Impact on Design, ISCA'98","author":"Bahar R.","year":"1998","unstructured":"R. Bahar and G. Albera . Performance analysis of wrong-path data cache accesses . In Workshop on Performance Analysis and its Impact on Design, ISCA'98 ., 1998 . R. Bahar and G. Albera. Performance analysis of wrong-path data cache accesses. In Workshop on Performance Analysis and its Impact on Design, ISCA'98., 1998."},{"key":"e_1_2_1_3_1","volume-title":"Proceedings of the 34th annual Intl. symposium on Microarchitecture","author":"Cher C.-Y.","year":"2001","unstructured":"C.-Y. Cher and T. N. Vijaykumar . Skipper: a microarchitecture for exploiting control-flow independence . In Proceedings of the 34th annual Intl. symposium on Microarchitecture , Washington, DC, USA , 2001 . C.-Y. Cher and T. N. Vijaykumar. Skipper: a microarchitecture for exploiting control-flow independence. In Proceedings of the 34th annual Intl. symposium on Microarchitecture, Washington, DC, USA, 2001."},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.5555\/646664.700879"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2004.10008"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2005.53"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/279358.279376"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.5555\/243846.243880"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.5555\/580550.876441"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/279358.279377"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/SBAC-PAD.2004.11"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/TC.2005.190"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.5555\/520549.822775"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.5555\/646667.700033"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.5555\/645988.674158"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.5555\/800052.801871"},{"key":"e_1_2_1_18_1","unstructured":"SPEC. Standard performance evaluation corporation (spec) 2000 benchmark suite. http:\/\/www.spec.org\/cpu2000\/.  SPEC. Standard performance evaluation corporation (spec) 2000 benchmark suite. http:\/\/www.spec.org\/cpu2000\/."},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/1024393.1024407"},{"key":"e_1_2_1_20_1","volume-title":"Int. Annual Computer Measurement Group Conference","author":"Tullsen D. M.","year":"1996","unstructured":"D. M. Tullsen . Simulation and modeling of a simultaneous multithreading processor . In Int. Annual Computer Measurement Group Conference , 1996 . D. M. Tullsen. Simulation and modeling of a simultaneous multithreading processor. In Int. Annual Computer Measurement Group Conference, 1996."}],"container-title":["ACM SIGARCH Computer Architecture News"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/1327312.1327318","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/1327312.1327318","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T13:56:25Z","timestamp":1750254985000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/1327312.1327318"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2007,9]]},"references-count":19,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2007,9]]}},"alternative-id":["10.1145\/1327312.1327318"],"URL":"https:\/\/doi.org\/10.1145\/1327312.1327318","relation":{},"ISSN":["0163-5964"],"issn-type":[{"type":"print","value":"0163-5964"}],"subject":[],"published":{"date-parts":[[2007,9]]},"assertion":[{"value":"2007-09-01","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}