{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,11,18]],"date-time":"2025-11-18T12:16:46Z","timestamp":1763468206799,"version":"3.41.0"},"reference-count":33,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2014,8,29]],"date-time":"2014-08-29T00:00:00Z","timestamp":1409270400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/100000143","name":"Division of Computing and Communication Foundations","doi-asserted-by":"publisher","award":["CNS-0964478 and CCF-0916689"],"award-info":[{"award-number":["CNS-0964478 and CCF-0916689"]}],"id":[{"id":"10.13039\/100000143","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000144","name":"Division of Computer and Network Systems","doi-asserted-by":"publisher","award":["CNS-0964478 and CCF-0916689"],"award-info":[{"award-number":["CNS-0964478 and CCF-0916689"]}],"id":[{"id":"10.13039\/100000144","id-type":"DOI","asserted-by":"publisher"}]},{"name":"ARM Ltd"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Comput. Syst."],"published-print":{"date-parts":[[2014,9,23]]},"abstract":"<jats:p>Approximate computing, where computation accuracy is traded off for better performance or higher data throughput, is one solution that can help data processing keep pace with the current and growing abundance of information. For particular domains, such as multimedia and learning algorithms, approximation is commonly used today. We consider automation to be essential to provide transparent approximation, and we show that larger benefits can be achieved by constructing the approximation techniques to fit the underlying hardware. Our target platform is the GPU because of its high performance capabilities and difficult programming challenges that can be alleviated with proper automation. Our approach\u2014SAGE\u2014combines a static compiler that automatically generates a set of CUDA kernels with varying levels of approximation with a runtime system that iteratively selects among the available kernels to achieve speedup while adhering to a target output quality set by the user. The SAGE compiler employs three optimization techniques to generate approximate kernels that exploit the GPU microarchitecture: selective discarding of atomic operations, data packing, and thread fusion. Across a set of machine learning and image processing kernels, SAGE's approximation yields an average of 2.5\u00d7 speedup with less than 10% quality loss compared to the accurate execution on a NVIDIA GTX 560 GPU.<\/jats:p>","DOI":"10.1145\/2631913","type":"journal-article","created":{"date-parts":[[2014,9,2]],"date-time":"2014-09-02T12:48:26Z","timestamp":1409662106000},"page":"1-29","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":3,"title":["Scaling Performance via Self-Tuning Approximation for Graphics Engines"],"prefix":"10.1145","volume":"32","author":[{"given":"Mehrzad","family":"Samadi","sequence":"first","affiliation":[{"name":"University of Michigan, MI"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Janghaeng","family":"Lee","sequence":"additional","affiliation":[{"name":"University of Michigan, MI"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"D. Anoushe","family":"Jamshidi","sequence":"additional","affiliation":[{"name":"University of Michigan, MI"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Scott","family":"Mahlke","sequence":"additional","affiliation":[{"name":"University of Michigan, MI"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Amir","family":"Hormati","sequence":"additional","affiliation":[{"name":"Google Inc., Seattle, WA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2014,8,29]]},"reference":[{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1145\/1542476.1542481"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.5555\/2190025.2190056"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1145\/1806596.1806620"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/1390156.1390170"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/1815961.1816026"},{"key":"e_1_2_1_7_1","unstructured":"EMC Corporation. 2011. Extracting Value from Chaos. http:\/\/www.emc.com\/collateral\/analyst-reports\/idc-extracting-value-from-chaos-ar.pdf.  EMC Corporation. 2011. Extracting Value from Chaos. http:\/\/www.emc.com\/collateral\/analyst-reports\/idc-extracting-value-from-chaos-ar.pdf."},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/2150976.2151008"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2012.48"},{"volume-title":"Retrieved","year":"2010","author":"Frank Andrew","key":"e_1_2_1_10_1"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/1950365.1950390"},{"key":"e_1_2_1_12_1","unstructured":"Alex Kulesza and Fernando Pereira. 2008. Structured learning with approximate inference. In Advances in Neural Information Processing Systems. 785--792.  Alex Kulesza and Fernando Pereira. 2008. Structured learning with approximate inference. In Advances in Neural Information Processing Systems. 785--792."},{"volume-title":"Proceedings of the 16th Workshop on Languages and Compilers for Parallel Computing. 539--553","year":"2003","author":"Lee Sang Ik","key":"e_1_2_1_13_1"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2007.346196"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/CIT.2010.60"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/1806799.1806808"},{"volume-title":"Retrieved","year":"2013","author":"NVIDIA.","key":"e_1_2_1_17_1"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/1183401.1183447"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/1297027.1297055"},{"volume-title":"Artificial Intelligence: A Modern Approach","year":"2009","author":"Russell Stuart","key":"e_1_2_1_20_1"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/2254064.2254067"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/2541940.2541948"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/2540708.2540711"},{"volume-title":"Proceedings of the 1st Workshop on Approximate Computing across the System Stack. 1--3.","year":"2014","author":"Samadi Mehrzad","key":"e_1_2_1_24_1"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/1993316.1993518"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1145\/2540708.2540712"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/2370816.2370879"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2006.881959"},{"volume-title":"Meyerson","year":"2011","author":"Shindler Michael","key":"e_1_2_1_29_1"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1145\/2025113.2025133"},{"volume-title":"Proceedings of the 13th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming. 16--30","author":"Stratton John A.","key":"e_1_2_1_31_1"},{"volume-title":"Proceedings of the 28th International Conference on Machine Learning. 609--616","author":"Sujeeth Arvind K.","key":"e_1_2_1_32_1"},{"volume-title":"Dunlop","year":"2000","author":"Tamhane Ajit C.","key":"e_1_2_1_33_1"},{"volume-title":"Bovik","year":"2006","author":"Sheikh Hamid R.","key":"e_1_2_1_34_1"}],"container-title":["ACM Transactions on Computer Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2631913","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/2631913","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T07:19:13Z","timestamp":1750231153000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2631913"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2014,8,29]]},"references-count":33,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2014,9,23]]}},"alternative-id":["10.1145\/2631913"],"URL":"https:\/\/doi.org\/10.1145\/2631913","relation":{},"ISSN":["0734-2071","1557-7333"],"issn-type":[{"type":"print","value":"0734-2071"},{"type":"electronic","value":"1557-7333"}],"subject":[],"published":{"date-parts":[[2014,8,29]]},"assertion":[{"value":"2014-04-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2014-05-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2014-08-29","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}