{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,19]],"date-time":"2026-06-19T02:47:12Z","timestamp":1781837232283,"version":"3.54.5"},"reference-count":69,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2018,12,7]],"date-time":"2018-12-07T00:00:00Z","timestamp":1544140800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"FAPESP fellowships","award":["2013\/08293-7, 2014\/03840-2 and 2016\/18929-4"],"award-info":[{"award-number":["2013\/08293-7, 2014\/03840-2 and 2016\/18929-4"]}]},{"name":"CNPq, and CAPES"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Archit. Code Optim."],"published-print":{"date-parts":[[2018,12,31]]},"abstract":"<jats:p>\n            Value prediction improves instruction level parallelism in superscalar processors by breaking true data dependencies. Although this technique can significantly improve overall performance, most of the state-of-the-art value prediction approaches require high hardware cost, which is the main obstacle for its wide adoption in current processors. To tackle this issue, we revisit\n            <jats:italic>load<\/jats:italic>\n            value prediction as an efficient alternative to the classical approaches that predict\n            <jats:italic>all<\/jats:italic>\n            instructions. By speculating only on loads, the pressure over shared resources (e.g., the Physical Register File) and the predictor size can be substantially reduced (e.g., more than 90% reduction compared to recent works). We observe that existing value predictors cannot achieve very high performance when speculating only on load instructions. To solve this problem, we propose a new, accurate and low-cost mechanism for predicting the values of load instructions: the Address-first Value-next Predictor with Value Prefetching (AVPP). The key idea of our predictor is to predict the load address first (which, we find, is much more predictable than the value) and to use a small non-speculative Value Table (VT)\u2014indexed by the predicted address\u2014to predict the value next. To increase the coverage of AVPP, we aim to increase the hit rate of the VT by predicting also the load address of a\n            <jats:italic>future instance<\/jats:italic>\n            of the same load instruction and prefetching its value in the VT. We show that AVPP is relatively easy to implement, requiring only 2.5% of the area of a 32KB L1 data cache. We compare our mechanism with five state-of-the-art value prediction techniques, evaluated within the context of load value prediction, in a relatively narrow out-of-order processor. On average, our AVPP predictor achieves 11.2% speedup and 3.7% of energy savings over the baseline processor, outperforming\n            <jats:italic>all<\/jats:italic>\n            the state-of-the-art predictors in 16 of the 23 benchmarks we evaluate. We evaluate AVPP implemented together with different prefetching techniques, showing additive performance gains (20% average speedup). In addition, we propose a new taxonomy to classify different value predictor\n            <jats:italic>policies<\/jats:italic>\n            regarding predictor update, predictor availability, and in-flight pending updates. We evaluate these policies in detail.\n          <\/jats:p>","DOI":"10.1145\/3239567","type":"journal-article","created":{"date-parts":[[2018,12,7]],"date-time":"2018-12-07T13:17:29Z","timestamp":1544188649000},"page":"1-30","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":12,"title":["AVPP"],"prefix":"10.1145","volume":"15","author":[{"given":"Lois","family":"Orosa","sequence":"first","affiliation":[{"name":"University of Campinas (UNICAMP) and ETH Z\u00fcrich, Z\u00fcrich, Switzerland"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Rodolfo","family":"Azevedo","sequence":"additional","affiliation":[{"name":"University of Campinas (UNICAMP), Campinas, Brazil"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Onur","family":"Mutlu","sequence":"additional","affiliation":[{"name":"ETH Z\u00fcrich, Z\u00fcrich, Switzerland"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2018,12,7]]},"reference":[{"key":"e_1_2_1_1_1","volume-title":"Proceedings of the 25th International Conference on Very Large Data Bases (VLDB\u201999)","author":"Ailamaki Anastassia"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1145\/1787275.1787321"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/1465482.1465560"},{"key":"e_1_2_1_4_1","unstructured":"ARM. 2016. ARM Cortex-A72 MPCore Processor Technical Reference Manual. Retrieved from http:\/\/infocenter.arm.com\/help\/index.jsp?topic&equals;\/com.arm.doc.100095_0003_06_en\/index.html.  ARM. 2016. ARM Cortex-A72 MPCore Processor Technical Reference Manual. Retrieved from http:\/\/infocenter.arm.com\/help\/index.jsp?topic&equals;\/com.arm.doc.100095_0003_06_en\/index.html."},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.5555\/225160.225176"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/300979.300984"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/1454115.1454128"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.5555\/1924943.1924944"},{"key":"e_1_2_1_9_1","volume-title":"Proceedings of the 1999 International Conference on Parallel Architectures and Compilation Techniques (PACT\u201999)","author":"Burtscher Martin"},{"key":"e_1_2_1_10_1","first-page":"1","article-title":"A comparative survey of load speculation architectures","volume":"2","author":"Calder Brad","year":"2000","journal-title":"J. Instruct.-Level Parallel."},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/300979.300985"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/1138035.1138038"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/12.381947"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/279358.279378"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1147\/rd.374.0547"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/3090634"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2008.44"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/379240.379253"},{"key":"e_1_2_1_19_1","unstructured":"Agner Fog. 2017. Lists of Instruction Latencies Throughputs and Micro-Operation Breakdowns for Intel AMD and VIA CPUs. Retrieved from http:\/\/www.agner.org\/optimize\/.  Agner Fog. 2017. Lists of Instruction Latencies Throughputs and Micro-Operation Breakdowns for Intel AMD and VIA CPUs. Retrieved from http:\/\/www.agner.org\/optimize\/."},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/291069.291058"},{"key":"e_1_2_1_21_1","volume-title":"Proceedings of the 1991 International Conference on Parallel Processing. 355--364","author":"Gharachorloo Kourosh","year":"1991"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.5555\/580550.876442"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/263580.263631"},{"key":"e_1_2_1_25_1","volume-title":"Proceedings of the 49th Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO\u201916)","author":"Hashemi Milad"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1145\/2830772.2830812"},{"key":"e_1_2_1_27_1","volume-title":"Proceedings of the 32nd Annual ACM\/IEEE International Symposium on Microarchitecture (MICRO\u201999)","author":"Heil Timothy H."},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1145\/1542275.1542349"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/2150976.2151001"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1145\/2485922.2485936"},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2014.29"},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2005.9"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1145\/1669112.1669172"},{"key":"e_1_2_1_34_1","volume-title":"Proceedings of the 29th Annual ACM\/IEEE International Symposium on Microarchitecture (MICRO\u201996)","author":"Mikko"},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1145\/237090.237173"},{"key":"e_1_2_1_36_1","unstructured":"Paul E. McKenney. 2017. Is parallel programming hard and if so what can you do about it?(v2017. 01.02 a). arXiv Preprint arXiv:1701.00854 (2017).  Paul E. McKenney. 2017. Is parallel programming hard and if so what can you do about it?(v2017. 01.02 a). arXiv Preprint arXiv:1701.00854 (2017)."},{"key":"e_1_2_1_37_1","unstructured":"A. Mendelson and F. Gabbay. 1996. Speculative Execution based on Value Prediction. Technical Report. EE Department TR 1080 Technion Israel Institue of Technology.  A. Mendelson and F. Gabbay. 1996. Speculative Execution based on Value Prediction. Technical Report. EE Department TR 1080 Technion Israel Institue of Technology."},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2014.22"},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1145\/264107.264189"},{"key":"e_1_2_1_40_1","volume-title":"Proceedings of the 30th Annual ACM\/IEEE International Symposium on Microarchitecture (MICRO\u201997)","author":"Moshovos Andreas"},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA.2005.49"},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.5555\/1116644.1116665"},{"key":"e_1_2_1_43_1","volume-title":"Proceedings of the 9th International Symposium on High-Performance Computer Architecture (HPCA\u201903)","author":"Mutlu Onur"},{"key":"e_1_2_1_44_1","volume-title":"Proceedings of the 5th International Symposium on High Performance Computer Architecture (HPCA\u201999)","author":"Nakra Tarun"},{"key":"e_1_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.1145\/1772954.1772958"},{"key":"e_1_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.5555\/3195638.3195643"},{"key":"e_1_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.5555\/2665671.2665742"},{"key":"e_1_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2014.6835952"},{"key":"e_1_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2015.7056018"},{"key":"e_1_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.5555\/942806.943854"},{"key":"e_1_2_1_51_1","doi-asserted-by":"publisher","DOI":"10.1145\/2485922.2485963"},{"key":"e_1_2_1_52_1","doi-asserted-by":"publisher","DOI":"10.1007\/3-540-47847-7_11"},{"key":"e_1_2_1_53_1","doi-asserted-by":"publisher","DOI":"10.5555\/874076.876459"},{"key":"e_1_2_1_55_1","volume-title":"Proceedings of the 30th Annual ACM\/IEEE International Symposium on Microarchitecture (MICRO\u201997)","author":"Sazeides Yiannakis"},{"key":"e_1_2_1_56_1","volume-title":"Proceedings of the 29th Annual ACM\/IEEE International Symposium on Microarchitecture (MICRO\u201996)","author":"Sazeides Yiannakis"},{"key":"e_1_2_1_57_1","first-page":"1","article-title":"A case for (partially) TAgged GEometric history length branch prediction","volume":"8","author":"Seznec Andr\u00e9","year":"2006","journal-title":"J. Instruct. Level Parallel."},{"key":"e_1_2_1_58_1","doi-asserted-by":"publisher","DOI":"10.1145\/3123939.3123951"},{"key":"e_1_2_1_59_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2007.346185"},{"key":"e_1_2_1_60_1","doi-asserted-by":"publisher","DOI":"10.1145\/1508244.1508274"},{"key":"e_1_2_1_61_1","doi-asserted-by":"publisher","DOI":"10.1145\/2628071.2628110"},{"key":"e_1_2_1_62_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2005.22"},{"key":"e_1_2_1_63_1","doi-asserted-by":"publisher","DOI":"10.1145\/300979.301002"},{"key":"e_1_2_1_64_1","doi-asserted-by":"publisher","DOI":"10.5555\/645989.674310"},{"key":"e_1_2_1_65_1","volume-title":"Proceedings of the 30th Annual ACM\/IEEE International Symposium on Microarchitecture (MICRO\u201997)","author":"Gary"},{"key":"e_1_2_1_66_1","doi-asserted-by":"publisher","DOI":"10.5555\/266800.266827"},{"key":"e_1_2_1_67_1","doi-asserted-by":"publisher","DOI":"10.1145\/223982.223990"},{"key":"e_1_2_1_68_1","doi-asserted-by":"publisher","DOI":"10.1145\/2836168"},{"key":"e_1_2_1_69_1","doi-asserted-by":"publisher","DOI":"10.1109\/MDAT.2015.2504899"},{"key":"e_1_2_1_70_1","doi-asserted-by":"publisher","DOI":"10.1109\/TC.2005.117"},{"key":"e_1_2_1_71_1","doi-asserted-by":"publisher","DOI":"10.1145\/280756.280943"}],"container-title":["ACM Transactions on Architecture and Code Optimization"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3239567","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3239567","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T01:08:20Z","timestamp":1750208900000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3239567"}},"subtitle":["Address-first Value-next Predictor with Value Prefetching for Improving the Efficiency of Load Value Prediction"],"short-title":[],"issued":{"date-parts":[[2018,12,7]]},"references-count":69,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2018,12,31]]}},"alternative-id":["10.1145\/3239567"],"URL":"https:\/\/doi.org\/10.1145\/3239567","relation":{},"ISSN":["1544-3566","1544-3973"],"issn-type":[{"value":"1544-3566","type":"print"},{"value":"1544-3973","type":"electronic"}],"subject":[],"published":{"date-parts":[[2018,12,7]]},"assertion":[{"value":"2017-05-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2018-07-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2018-12-07","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}