{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,2]],"date-time":"2026-07-02T15:58:27Z","timestamp":1783007907863,"version":"3.54.5"},"reference-count":56,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2016,1,6]],"date-time":"2016-01-06T00:00:00Z","timestamp":1452038400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Archit. Code Optim."],"published-print":{"date-parts":[[2016,1,7]]},"abstract":"<jats:p>This article aims to tackle two fundamental memory bottlenecks: limited off-chip bandwidth (bandwidth wall) and long access latency (memory wall). To achieve this goal, our approach exploits the inherent error resilience of a wide range of applications. We introduce an approximation technique, called Rollback-Free Value Prediction (RFVP). When certain safe-to-approximate load operations miss in the cache, RFVP predicts the requested values. However, RFVP does not check for or recover from load-value mispredictions, hence, avoiding the high cost of pipeline flushes and re-executions. RFVP mitigates the memory wall by enabling the execution to continue without stalling for long-latency memory accesses. To mitigate the bandwidth wall, RFVP drops a fraction of load requests that miss in the cache after predicting their values. Dropping requests reduces memory bandwidth contention by removing them from the system. The drop rate is a knob to control the trade-off between performance\/energy efficiency and output quality. Our extensive evaluations show that RFVP, when used in GPUs, yields significant performance improvement and energy reduction for a wide range of quality-loss levels. We also evaluate RFVP\u2019s latency benefits for a single core CPU. The results show performance improvement and energy reduction for a wide variety of applications with less than 1% loss in quality.<\/jats:p>","DOI":"10.1145\/2836168","type":"journal-article","created":{"date-parts":[[2016,1,7]],"date-time":"2016-01-07T14:04:54Z","timestamp":1452175494000},"page":"1-26","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":61,"title":["RFVP"],"prefix":"10.1145","volume":"12","author":[{"given":"Amir","family":"Yazdanbakhsh","sequence":"first","affiliation":[{"name":"Georgia Institute of Technology"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Gennady","family":"Pekhimenko","sequence":"additional","affiliation":[{"name":"Carnegie Mellon University"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Bradley","family":"Thwaites","sequence":"additional","affiliation":[{"name":"Georgia Institute of Technology"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Hadi","family":"Esmaeilzadeh","sequence":"additional","affiliation":[{"name":"Georgia Institute of Technology"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Onur","family":"Mutlu","sequence":"additional","affiliation":[{"name":"Carnegie Mellon University"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Todd C.","family":"Mowry","sequence":"additional","affiliation":[{"name":"Carnegie Mellon University"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2016,1,6]]},"reference":[{"key":"e_1_2_2_1_1","doi-asserted-by":"publisher","DOI":"10.1145\/567067.567085"},{"key":"e_1_2_2_2_1","doi-asserted-by":"publisher","DOI":"10.1109\/TC.2005.119"},{"key":"e_1_2_2_3_1","doi-asserted-by":"crossref","unstructured":"Ren\u00e9e St. Amant Amir Yazdanbakhsh Jongse Park Bradley Thwaites Hadi Esmaeilzadeh Arjang Hassibi Luis Ceze and Doug Burger. 2014. General-purpose code acceleration with limited-precision analog computation. In ISCA.   Ren\u00e9e St. Amant Amir Yazdanbakhsh Jongse Park Bradley Thwaites Hadi Esmaeilzadeh Arjang Hassibi Luis Ceze and Doug Burger. 2014. General-purpose code acceleration with limited-precision analog computation. In ISCA.","DOI":"10.1109\/ISCA.2014.6853213"},{"key":"e_1_2_2_4_1","doi-asserted-by":"crossref","unstructured":"Jose-Maria Arnau Joan-Manuel Parcerisa and Polychronis Xekalakis. 2014. Eliminating redundant fragment shader executions on a mobile GPU via hardware memoization. In ISCA.   Jose-Maria Arnau Joan-Manuel Parcerisa and Polychronis Xekalakis. 2014. Eliminating redundant fragment shader executions on a mobile GPU via hardware memoization. In ISCA.","DOI":"10.1109\/ISCA.2014.6853207"},{"key":"e_1_2_2_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/1806596.1806620"},{"key":"e_1_2_2_6_1","doi-asserted-by":"crossref","unstructured":"A Bakhoda G. L. Yuan W. W. L. Fung H. Wong and T. M. Aamodt. 2009. Analyzing CUDA workloads using a detailed GPU simulator. In ISPASS.  A Bakhoda G. L. Yuan W. W. L. Fung H. Wong and T. M. Aamodt. 2009. Analyzing CUDA workloads using a detailed GPU simulator. In ISPASS.","DOI":"10.1109\/ISPASS.2009.4919648"},{"key":"e_1_2_2_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/2509136.2509546"},{"key":"e_1_2_2_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/1138035.1138038"},{"key":"e_1_2_2_9_1","doi-asserted-by":"crossref","unstructured":"Lakshmi N. Chakrapani Bilge E. S. Akgul Suresh Cheemalavagu Pinar Korkmaz Krishna V. Palem and Balasubramanian Seshasayee. 2006. Ultra-efficient (Embedded) SoC architectures based on probabilistic CMOS (PCMOS) technology. In DATE.   Lakshmi N. Chakrapani Bilge E. S. Akgul Suresh Cheemalavagu Pinar Korkmaz Krishna V. Palem and Balasubramanian Seshasayee. 2006. Ultra-efficient (Embedded) SoC architectures based on probabilistic CMOS (PCMOS) technology. In DATE.","DOI":"10.1109\/DATE.2006.243978"},{"key":"e_1_2_2_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/IISWC.2009.5306797"},{"key":"e_1_2_2_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2010.36"},{"key":"e_1_2_2_12_1","doi-asserted-by":"publisher","DOI":"10.5555\/1884795.1884804"},{"key":"e_1_2_2_13_1","doi-asserted-by":"publisher","DOI":"10.1145\/379240.379248"},{"key":"e_1_2_2_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/1815961.1816026"},{"key":"e_1_2_2_15_1","doi-asserted-by":"publisher","DOI":"10.1147\/rd.374.0547"},{"key":"e_1_2_2_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/2150976.2151008"},{"key":"e_1_2_2_17_1","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2012.48"},{"key":"e_1_2_2_18_1","unstructured":"Bart Goeman Hans Vandierendonck and Koenraad De Bosschere. 2001. Differential FCM: Increasing value prediction accuracy by improving table usage efficiency. In HPCA.   Bart Goeman Hans Vandierendonck and Koenraad De Bosschere. 2001. Differential FCM: Increasing value prediction accuracy by improving table usage efficiency. In HPCA."},{"key":"e_1_2_2_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/1054907.1054913"},{"key":"e_1_2_2_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/1454115.1454152"},{"key":"e_1_2_2_21_1","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2011.89"},{"key":"e_1_2_2_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/2485922.2485964"},{"key":"e_1_2_2_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/1669112.1669172"},{"key":"e_1_2_2_24_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2007.346196"},{"key":"e_1_2_2_25_1","volume-title":"Lipasti and John Paul Shen","author":"Mikko","year":"1996","unstructured":"Mikko H. Lipasti and John Paul Shen . 1996 . Exceeding the dataflow limit via value prediction. In MICRO. Mikko H. Lipasti and John Paul Shen. 1996. Exceeding the dataflow limit via value prediction. In MICRO."},{"key":"e_1_2_2_26_1","doi-asserted-by":"publisher","DOI":"10.1145\/237090.237173"},{"key":"e_1_2_2_27_1","volume-title":"Zorn","author":"Liu Song","year":"2011","unstructured":"Song Liu , Karthik Pattabiraman , Thomas Moscibroda , and Benjamin G . Zorn . 2011 . Flikker : Saving refresh-power in mobile devices through critical data partitioning. In ASPLOS. Song Liu, Karthik Pattabiraman, Thomas Moscibroda, and Benjamin G. Zorn. 2011. Flikker: Saving refresh-power in mobile devices through critical data partitioning. In ASPLOS."},{"key":"e_1_2_2_28_1","doi-asserted-by":"publisher","DOI":"10.1109\/DSN.2014.50"},{"key":"e_1_2_2_29_1","volume-title":"Axilog: Abstractions for approximate hardware design and reuse","author":"Mahajan Divya","year":"2015","unstructured":"Divya Mahajan , Kartik Ramkrishnan , Rudra Jariwala , Amir Yazdanbakhsh , Jongse Park , Bradley Thwaites , Anandhavel Nagendrakumar , Abbas Rahimi , Hadi Esmaeilzadeh , and Kia Bazargan . 2015 . Axilog: Abstractions for approximate hardware design and reuse . In IEEE Micro . Divya Mahajan, Kartik Ramkrishnan, Rudra Jariwala, Amir Yazdanbakhsh, Jongse Park, Bradley Thwaites, Anandhavel Nagendrakumar, Abbas Rahimi, Hadi Esmaeilzadeh, and Kia Bazargan. 2015. Axilog: Abstractions for approximate hardware design and reuse. In IEEE Micro."},{"key":"e_1_2_2_30_1","doi-asserted-by":"publisher","DOI":"10.1109\/PACT.2009.22"},{"key":"e_1_2_2_31_1","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2007.30"},{"key":"e_1_2_2_32_1","unstructured":"Makoto Murase. 1992. Linear feedback shift register. US Patent.  Makoto Murase. 1992. Linear feedback shift register. US Patent."},{"key":"e_1_2_2_33_1","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2005.11"},{"key":"e_1_2_2_34_1","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2003.1261383"},{"key":"e_1_2_2_35_1","doi-asserted-by":"publisher","DOI":"10.1145\/1250734.1250746"},{"key":"e_1_2_2_36_1","doi-asserted-by":"publisher","DOI":"10.1145\/2786805.2786807"},{"key":"e_1_2_2_37_1","doi-asserted-by":"publisher","DOI":"10.1145\/2024724.2024954"},{"key":"e_1_2_2_38_1","doi-asserted-by":"crossref","unstructured":"Gennady Pekhimenko Evgeny Bolotin Mike O\u2019Connor Onur Mutlu Todd Mowry and Stephen Keckler. 2015. Toggle-aware compression for GPUs. Computer Architecture Letters.  Gennady Pekhimenko Evgeny Bolotin Mike O\u2019Connor Onur Mutlu Todd Mowry and Stephen Keckler. 2015. Toggle-aware compression for GPUs. Computer Architecture Letters.","DOI":"10.1109\/HPCA.2016.7446064"},{"key":"e_1_2_2_39_1","doi-asserted-by":"publisher","DOI":"10.1145\/2370816.2370870"},{"key":"e_1_2_2_40_1","doi-asserted-by":"crossref","unstructured":"A. Perais and A. Seznec. 2014. Practical data value speculation for future high-end processors. In HPCA.  A. Perais and A. Seznec. 2014. Practical data value speculation for future high-end processors. In HPCA.","DOI":"10.1109\/HPCA.2014.6835952"},{"key":"e_1_2_2_41_1","doi-asserted-by":"publisher","DOI":"10.1145\/1555754.1555801"},{"key":"e_1_2_2_42_1","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2012.16"},{"key":"e_1_2_2_43_1","doi-asserted-by":"publisher","DOI":"10.1145\/2541940.2541948"},{"key":"e_1_2_2_44_1","doi-asserted-by":"publisher","DOI":"10.1145\/2540708.2540711"},{"key":"e_1_2_2_45_1","doi-asserted-by":"publisher","DOI":"10.1145\/1993498.1993518"},{"key":"e_1_2_2_46_1","doi-asserted-by":"publisher","DOI":"10.1145\/2540708.2540712"},{"key":"e_1_2_2_47_1","doi-asserted-by":"crossref","unstructured":"Joshua San Miguel Mario Badr and Natalie Enright Jerger. 2014. Load value approximation. In MICRO.  Joshua San Miguel Mario Badr and Natalie Enright Jerger. 2014. Load value approximation. In MICRO.","DOI":"10.1109\/MICRO.2014.22"},{"key":"e_1_2_2_48_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2012.2232647"},{"key":"e_1_2_2_49_1","volume-title":"Smith","author":"Sazeides Yiannakis","year":"1997","unstructured":"Yiannakis Sazeides and James E . Smith . 1997 . The predictability of data values. In MICRO. Yiannakis Sazeides and James E. Smith. 1997. The predictability of data values. In MICRO."},{"key":"e_1_2_2_50_1","doi-asserted-by":"publisher","DOI":"10.1109\/L-CA.2011.22"},{"key":"e_1_2_2_51_1","doi-asserted-by":"publisher","DOI":"10.1145\/2025113.2025133"},{"key":"e_1_2_2_52_1","doi-asserted-by":"crossref","unstructured":"R. Thomas and M. Franklin. 2001. Using dataflow based context for accurate value prediction. In PACT.   R. Thomas and M. Franklin. 2001. Using dataflow based context for accurate value prediction. In PACT.","DOI":"10.1007\/3-540-36265-7_55"},{"key":"e_1_2_2_53_1","doi-asserted-by":"publisher","DOI":"10.1145\/2628071.2628110"},{"key":"e_1_2_2_54_1","doi-asserted-by":"publisher","DOI":"10.1145\/2749469.2750399"},{"key":"e_1_2_2_55_1","volume-title":"Axilog: Language support for approximate hardware design. In DATE.","author":"Yazdanbakhsh Amir","year":"2015","unstructured":"Amir Yazdanbakhsh , Divya Mahajan , Bradley Thwaites , Jongse Park , Anandhavel Nagendrakumar , Sindhuja Sethuraman , Kartik Ramkrishnan , Nishanthi Ravindran , Rudra Jariwala , Abbas Rahimi , Hadi Esmaeilzadeh , and Kia Bazargan . 2015 . Axilog: Language support for approximate hardware design. In DATE. Amir Yazdanbakhsh, Divya Mahajan, Bradley Thwaites, Jongse Park, Anandhavel Nagendrakumar, Sindhuja Sethuraman, Kartik Ramkrishnan, Nishanthi Ravindran, Rudra Jariwala, Abbas Rahimi, Hadi Esmaeilzadeh, and Kia Bazargan. 2015. Axilog: Language support for approximate hardware design. In DATE."},{"key":"e_1_2_2_56_1","doi-asserted-by":"publisher","DOI":"10.1109\/TC.2005.117"}],"container-title":["ACM Transactions on Architecture and Code Optimization"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2836168","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/2836168","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T05:48:41Z","timestamp":1750225721000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2836168"}},"subtitle":["Rollback-Free Value Prediction with Safe-to-Approximate Loads"],"short-title":[],"issued":{"date-parts":[[2016,1,6]]},"references-count":56,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2016,1,7]]}},"alternative-id":["10.1145\/2836168"],"URL":"https:\/\/doi.org\/10.1145\/2836168","relation":{},"ISSN":["1544-3566","1544-3973"],"issn-type":[{"value":"1544-3566","type":"print"},{"value":"1544-3973","type":"electronic"}],"subject":[],"published":{"date-parts":[[2016,1,6]]},"assertion":[{"value":"2015-06-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2015-10-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2016-01-06","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}