{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,10]],"date-time":"2026-01-10T07:17:28Z","timestamp":1768029448585,"version":"3.49.0"},"reference-count":67,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2018,12,21]],"date-time":"2018-12-21T00:00:00Z","timestamp":1545350400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/100000001","name":"National Science Foundation","doi-asserted-by":"publisher","award":["1526750, 1763681, 1439057, 1439021, 1629129, 1409095, 1626251, 1629915"],"award-info":[{"award-number":["1526750, 1763681, 1439057, 1439021, 1629129, 1409095, 1626251, 1629915"]}],"id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["Proc. ACM Meas. Anal. Comput. Syst."],"published-print":{"date-parts":[[2018,12,21]]},"abstract":"<jats:p>One cost that plays a significant role in shaping the overall performance of both single-threaded and multi-thread applications in modern computing systems is the cost of moving data between compute elements and storage elements. Traditional approaches to address this cost are code and data layout reorganizations and various hardware enhancements. More recently, an alternative paradigm, called Near Data Computing (NDC) or Near Data Processing (NDP), has been shown to be effective in reducing the data movements costs, by moving computation to data, instead of the traditional approach of moving data to computation. Unfortunately, the existing Near Data Computing proposals require significant modifications to hardware and are yet to be widely adopted.<\/jats:p>\n          <jats:p>In this paper, we present a software-only (compiler-driven) approach to reducing data movement costs in both single-threaded and multi-threaded applications. Our approach, referred to as Computing with Near Data (CND), is built upon a concept called \"recomputation,\" in which a costly data access is replaced by a few less costly data accesses plus some extra computation, if the cumulative cost of the latter is less than that of the costly data access. If implemented carefully, CND can successfully trade off data access with computation, and considering the continuously increasing latency gap between the two, doing so can significantly reduce the execution latencies of both sequential and parallel application programs.<\/jats:p>\n          <jats:p>We i) quantify the intrinsic recomputability of a set of single-threaded and multi-threaded applications, ii) propose a practical, compiler-driven approach that automatically transforms a given application code fragment to a version that employs recomputation, iii) discuss an optimization strategy that increases recomputability; and iv) compare CND, both qualitatively and quantitatively, against NDC. Our experimental analysis of CND reveals that i) the average recomputability across our benchmarks is 51.1%, ii) our compiler-driven strategy is able to exploit 79.3% of the recomputation opportunities presented by our workloads, and iii) our enhancements increase the value of the recomputability metric significantly. As a result, our compiler-driven approach with the proposed enhancements brings an average execution time improvement of 40.1%.<\/jats:p>","DOI":"10.1145\/3287321","type":"journal-article","created":{"date-parts":[[2018,12,26]],"date-time":"2018-12-26T12:39:28Z","timestamp":1545827968000},"page":"1-30","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":5,"title":["Computing with Near Data"],"prefix":"10.1145","volume":"2","author":[{"given":"Xulong","family":"Tang","sequence":"first","affiliation":[{"name":"Pennsylvania State University, State College, PA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Mahmut Taylan","family":"Kandemir","sequence":"additional","affiliation":[{"name":"Pennsylvania State University, State College, PA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hui","family":"Zhao","sequence":"additional","affiliation":[{"name":"University of North Texas, Denton, TX, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Myoungsoo","family":"Jung","sequence":"additional","affiliation":[{"name":"Yonsei University, Seoul, South Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Mustafa","family":"Karakoy","sequence":"additional","affiliation":[{"name":"TOBB University of Economics and Technology, Ankara, Turkey"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2018,12,21]]},"reference":[{"key":"e_1_2_1_1_1","volume-title":"Proceedings of the 2009 Conference on Hot Topics in Cloud Computing.","author":"Adams Ian F.","unstructured":"Ian F. Adams , Darrell D. E. Long , Ethan L. Miller , Shankar Pasupathy , and Mark W. Storer . 2009. Maximizing Efficiency by Trading Storage for Computation . In Proceedings of the 2009 Conference on Hot Topics in Cloud Computing. Ian F. Adams, Darrell D. E. Long, Ethan L. Miller, Shankar Pasupathy, and Mark W. Storer. 2009. Maximizing Efficiency by Trading Storage for Computation. In Proceedings of the 2009 Conference on Hot Topics in Cloud Computing."},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1145\/2749469.2750386"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/2749469.2750385"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1145\/2749469.2750397"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/3037697.3037741"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/155090.155101"},{"key":"e_1_2_1_7_1","volume-title":"Near-data processing: Insights from a MICRO-46 Workshop. Micro","author":"Balasubramonian Rajeev","year":"2014","unstructured":"Rajeev Balasubramonian , Jichuan Chang , Troy Manning , Jaime H Moreno , Richard Murphy , Ravi Nair , and Steven Swanson . 2014. Near-data processing: Insights from a MICRO-46 Workshop. Micro , IEEE ( 2014 ). Rajeev Balasubramonian, Jichuan Chang, Troy Manning, Jaime H Moreno, Richard Murphy, Ravi Nair, and Steven Swanson. 2014. Near-data processing: Insights from a MICRO-46 Workshop. Micro, IEEE (2014)."},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2015.7056061"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/2024716.2024718"},{"key":"e_1_2_1_10_1","volume-title":"Proceedings of Programming Language Design And Implementation (PLDI).","author":"Bondhugula Uday","unstructured":"Uday Bondhugula , J. Ramanujam, and et al. 2008. PLuTo: A practical and fully automatic polyhedral program optimization system . In Proceedings of Programming Language Design And Implementation (PLDI). Uday Bondhugula, J. Ramanujam, and et al. 2008. PLuTo: A practical and fully automatic polyhedral program optimization system. In Proceedings of Programming Language Design And Implementation (PLDI)."},{"key":"e_1_2_1_11_1","volume-title":"LazyPIM: An Efficient Cache Coherence Mechanism for Processing-in-Memory","author":"Boroumand Amirali","year":"2016","unstructured":"Amirali Boroumand , Saugata Ghose , Brandon Lucia , Kevin Hsieh , Krishna Malladi , Hongzhong Zheng , and Onur Mutlu . 2016. LazyPIM: An Efficient Cache Coherence Mechanism for Processing-in-Memory . IEEE Computer Architecture Letters ( 2016 ). Amirali Boroumand, Saugata Ghose, Brandon Lucia, Kevin Hsieh, Krishna Malladi, Hongzhong Zheng, and Onur Mutlu. 2016. LazyPIM: An Efficient Cache Coherence Mechanism for Processing-in-Memory. IEEE Computer Architecture Letters (2016)."},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/195473.195557"},{"key":"e_1_2_1_13_1","volume-title":"Proceedings of the 5th International Symposium on High Performance Computer Architecture (ISCA).","author":"Carter J.","unstructured":"J. Carter , W. Hsieh , L. Stoller , M. Swanson , Lixin Zhang , E. Brunvand , A. Davis , Chen-Chi Kuo , R. Kuramkote , M. Parker , L. Schaelicke , and T. Tateyama . 1999. Impulse: Building a Smarter Memory Controller . In Proceedings of the 5th International Symposium on High Performance Computer Architecture (ISCA). J. Carter, W. Hsieh, L. Stoller, M. Swanson, Lixin Zhang, E. Brunvand, A. Davis, Chen-Chi Kuo, R. Kuramkote, M. Parker, L. Schaelicke, and T. Tateyama. 1999. Impulse: Building a Smarter Memory Controller. In Proceedings of the 5th International Symposium on High Performance Computer Architecture (ISCA)."},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA.2016.13"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/207110.207145"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/2370816.2370893"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/2737924.2737989"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/291069.291051"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/2.375174"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/1669112.1669149"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/224170.224337"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA.2016.46"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2016.7446096"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA.2016.27"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1109\/VLSIT.2012.6242474"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1006\/jpdc.1999.1552"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1109\/TC.2001.970571"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1145\/2745844.2745867"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/3192366.3192386"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1109\/PACT.2017.20"},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1145\/258915.258946"},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1145\/106972.106981"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1145\/1669112.1669155"},{"key":"e_1_2_1_34_1","volume-title":"Optimizing data locality by array restructuring.Department of Computer Science and Engineering","author":"Leung Shun-Tak","unstructured":"Shun-Tak Leung and John Zahorjan . 1995. Optimizing data locality by array restructuring.Department of Computer Science and Engineering , University of Washington. Shun-Tak Leung and John Zahorjan. 1995. Optimizing data locality by array restructuring.Department of Computer Science and Engineering, University of Washington."},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1145\/3123939.3123977"},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1145\/305138.305197"},{"key":"e_1_2_1_37_1","unstructured":"G.J. Lipovski and C. Yu. 1999. The Dynamic Associative Access Memory Chip and Its Application to SIMD Processing and Full-Text Database Retrieval. In MTDT.   G.J. Lipovski and C. Yu. 1999. The Dynamic Associative Access Memory Chip and Its Application to SIMD Processing and Full-Text Database Retrieval. In MTDT."},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1145\/2370816.2370869"},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA.2008.15"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1109\/PACT.2009.36"},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1145\/2463209.2488779"},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1145\/2000064.2000111"},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1147\/JRD.2015.2409732"},{"key":"e_1_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1006\/jpdc.2001.1815"},{"key":"e_1_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.1145\/2967938.2967940"},{"key":"e_1_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.1145\/277650.277661"},{"key":"e_1_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1109\/MASCOTS.2018.00022"},{"key":"e_1_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.1109\/MASCOTS.2017.16"},{"key":"e_1_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2016.25"},{"key":"e_1_2_1_51_1","doi-asserted-by":"publisher","DOI":"10.1145\/301618.301668"},{"key":"e_1_2_1_52_1","volume-title":"IEEE Transaction Computing","author":"Stone Harold S.","year":"1970","unstructured":"Harold S. Stone . 1970. IEEE Transaction Computing ( 1970 ). Harold S. Stone. 1970. IEEE Transaction Computing (1970)."},{"key":"e_1_2_1_53_1","doi-asserted-by":"publisher","DOI":"10.5555\/3195638.3195708"},{"key":"e_1_2_1_54_1","doi-asserted-by":"publisher","DOI":"10.1145\/3123939.3123954"},{"key":"e_1_2_1_55_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2017.14"},{"key":"e_1_2_1_56_1","doi-asserted-by":"publisher","DOI":"10.1145\/3287318"},{"key":"e_1_2_1_57_1","unstructured":"S. Verdoolaege M. Bruynooghe G. Janssens and P. Catthoor. 2003. Multi-dimensional incremental loop fusion for data locality. In ASAP.  S. Verdoolaege M. Bruynooghe G. Janssens and P. Catthoor. 2003. Multi-dimensional incremental loop fusion for data locality. In ASAP."},{"key":"e_1_2_1_58_1","doi-asserted-by":"publisher","DOI":"10.1145\/237090.237205"},{"key":"e_1_2_1_59_1","doi-asserted-by":"publisher","DOI":"10.1145\/113445.113449"},{"key":"e_1_2_1_60_1","doi-asserted-by":"publisher","DOI":"10.1109\/71.97902"},{"key":"e_1_2_1_61_1","doi-asserted-by":"publisher","DOI":"10.1145\/223982.223990"},{"key":"e_1_2_1_62_1","volume-title":"Proceedings of the 3rd Workshop on Near-Data Processing.","author":"Xu Lifan","year":"2015","unstructured":"Lifan Xu , Dong Ping Zhang , and Nuwan Jayasena . 2015 . Scaling Deep Learning on Multiple In-Memory Processors . In Proceedings of the 3rd Workshop on Near-Data Processing. Lifan Xu, Dong Ping Zhang, and Nuwan Jayasena. 2015. Scaling Deep Learning on Multiple In-Memory Processors. In Proceedings of the 3rd Workshop on Near-Data Processing."},{"key":"e_1_2_1_63_1","doi-asserted-by":"publisher","DOI":"10.1145\/2600212.2600213"},{"key":"e_1_2_1_64_1","doi-asserted-by":"publisher","DOI":"10.1145\/3195970.3196052"},{"key":"e_1_2_1_65_1","volume-title":"Shulin Zhao, Nachiappan Chidambaram Nachiappan, Anand Sivasubramaniam, Mahmut T. Kandemir, Ravi Iyer, and Chita R. Das.","author":"Zhang Haibo","year":"2017","unstructured":"Haibo Zhang , Prasanna Venkatesh Rengasamy , Shulin Zhao, Nachiappan Chidambaram Nachiappan, Anand Sivasubramaniam, Mahmut T. Kandemir, Ravi Iyer, and Chita R. Das. 2017 . Race-to-sleep Haibo Zhang, Prasanna Venkatesh Rengasamy, Shulin Zhao, Nachiappan Chidambaram Nachiappan, Anand Sivasubramaniam, Mahmut T. Kandemir, Ravi Iyer, and Chita R. Das. 2017. Race-to-sleep"},{"key":"e_1_2_1_66_1","unstructured":"Content Caching  Content Caching"},{"key":"e_1_2_1_67_1","volume-title":"Caching: A Recipe for Energy-efficient Video Streaming on Handhelds. In Proceedings of the 50th Annual IEEE\/ACM International Symposium on Microarchitecture.","author":"Display","unstructured":"Display Caching: A Recipe for Energy-efficient Video Streaming on Handhelds. In Proceedings of the 50th Annual IEEE\/ACM International Symposium on Microarchitecture. Display Caching: A Recipe for Energy-efficient Video Streaming on Handhelds. In Proceedings of the 50th Annual IEEE\/ACM International Symposium on Microarchitecture."},{"key":"e_1_2_1_68_1","unstructured":"Zhao Zhang Zhichun Zhu and Xiaodong Zhang. 2002. Breaking Address Mapping Symmetry at Multi-levels of Memory Hierarchy to Reduce DRAM Row-buffer Conflicts. In The Journal of Instruction-Level Parallelism (JILP).  Zhao Zhang Zhichun Zhu and Xiaodong Zhang. 2002. Breaking Address Mapping Symmetry at Multi-levels of Memory Hierarchy to Reduce DRAM Row-buffer Conflicts. In The Journal of Instruction-Level Parallelism (JILP)."}],"container-title":["Proceedings of the ACM on Measurement and Analysis of Computing Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3287321","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3287321","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3287321","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T01:01:53Z","timestamp":1750208513000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3287321"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2018,12,21]]},"references-count":67,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2018,12,21]]}},"alternative-id":["10.1145\/3287321"],"URL":"https:\/\/doi.org\/10.1145\/3287321","relation":{},"ISSN":["2476-1249"],"issn-type":[{"value":"2476-1249","type":"electronic"}],"subject":[],"published":{"date-parts":[[2018,12,21]]},"assertion":[{"value":"2018-12-21","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}