{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,31]],"date-time":"2026-01-31T07:30:40Z","timestamp":1769844640472,"version":"3.49.0"},"reference-count":65,"publisher":"Association for Computing Machinery (ACM)","issue":"1","license":[{"start":{"date-parts":[[2021,12,6]],"date-time":"2021-12-06T00:00:00Z","timestamp":1638748800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100004410","name":"Scientific and Technological Research Council of Turkey","doi-asserted-by":"crossref","award":["120E492"],"award-info":[{"award-number":["120E492"]}],"id":[{"id":"10.13039\/501100004410","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Royal Society-Newton Advanced Fellowship"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Archit. Code Optim."],"published-print":{"date-parts":[[2022,3,31]]},"abstract":"<jats:p>\n            One widely used metric that measures data locality is\n            <jats:italic>reuse distance<\/jats:italic>\n            \u2014the number of unique memory locations that are accessed between two consecutive accesses to a particular memory location. State-of-the-art techniques that measure reuse distance in parallel applications rely on simulators or binary instrumentation tools that incur large performance and memory overheads. Moreover, the existing sampling-based tools are limited to measuring reuse distances of a single thread and discard interactions among threads in multi-threaded programs. In this work, we propose\n            <jats:sc>ReuseTracker<\/jats:sc>\n            \u2014a fast and accurate reuse distance analyzer that leverages existing hardware features in commodity CPUs.\n            <jats:sc>ReuseTracker<\/jats:sc>\n            is designed for multi-threaded programs and takes cache-coherence effects into account. By utilizing hardware features like performance monitoring units and debug registers,\n            <jats:sc>ReuseTracker<\/jats:sc>\n            can accurately profile reuse distance in parallel applications with much lower overheads than existing tools. It introduces only 2.9\u00d7 runtime and 2.8\u00d7 memory overheads. Our tool achieves 92% accuracy when verified against a newly developed configurable benchmark that can generate a variety of different reuse distance patterns. We demonstrate the tool\u2019s functionality with two use-case scenarios using PARSEC, Rodinia, and Synchrobench benchmark suites where\n            <jats:sc>ReuseTracker<\/jats:sc>\n            guides code refactoring in these benchmarks by detecting spatial reuses in shared caches that are also false sharing and successfully predicts whether some benchmarks in these suites can benefit from adjacent cache line prefetch optimization.\n          <\/jats:p>","DOI":"10.1145\/3484199","type":"journal-article","created":{"date-parts":[[2021,12,6]],"date-time":"2021-12-06T21:29:34Z","timestamp":1638826174000},"page":"1-25","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":14,"title":["<scp>ReuseTracker<\/scp>\n            : Fast Yet Accurate Multicore Reuse Distance Analyzer"],"prefix":"10.1145","volume":"19","author":[{"given":"Muhammad Aditya","family":"Sasongko","sequence":"first","affiliation":[{"name":"Ko\u00e7 University, Istanbul, Turkey"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Milind","family":"Chabbi","sequence":"additional","affiliation":[{"name":"Scalable Machines Research, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Mandana Bagheri","family":"Marzijarani","sequence":"additional","affiliation":[{"name":"Ko\u00e7 University, Istanbul, Turkey"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Didem","family":"Unat","sequence":"additional","affiliation":[{"name":"Ko\u00e7 University, Istanbul, Turkey"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2021,12,6]]},"reference":[{"key":"e_1_3_2_2_2","article-title":"dcompiler\/loca: Program Locality Analysis Tools","author":"GitHub.","unstructured":"GitHub. [n.d.]. dcompiler\/loca: Program Locality Analysis Tools. Retrieved on 20 July, 2020 from https:\/\/github.com\/dcompiler\/loca.","journal-title":"Retrieved on 20 July, 2020 from https:\/\/github.com\/dcompiler\/loca"},{"key":"e_1_3_2_3_2","article-title":"Harmonic Progression","author":"Wikipedia.","unstructured":"Wikipedia. [n.d.]. Harmonic Progression. Retrieved on 12 January, 2021 from https:\/\/en.wikipedia.org\/wiki\/Harmonic_progression_(mathematics).","journal-title":"Retrieved on 12 January, 2021 from https:\/\/en.wikipedia.org\/wiki\/Harmonic_progression_(mathematics)"},{"key":"e_1_3_2_4_2","article-title":"Thread Affinity Interface (Linux* and Windows*)","author":"Intel.","unstructured":"Intel. [n.d.]. Thread Affinity Interface (Linux* and Windows*). Retrieved on 1 February, 2021 from https:\/\/software.intel.com\/content\/www\/us\/en\/develop\/documentation\/cpp-compiler-developer-guide-and-reference\/top\/optimization-and-programming-guide\/openmp-support\/openmp-library-support\/thread-affinity-interface-linux-and-windows.","journal-title":"R"},{"key":"e_1_3_2_5_2","doi-asserted-by":"publisher","DOI":"10.5555\/1753228.1753233"},{"key":"e_1_3_2_6_2","doi-asserted-by":"publisher","DOI":"10.1145\/3422575.3422806"},{"key":"e_1_3_2_7_2","doi-asserted-by":"publisher","DOI":"10.5555\/1153925.1154584"},{"key":"e_1_3_2_8_2","doi-asserted-by":"publisher","DOI":"10.1145\/1071690.1064232"},{"key":"e_1_3_2_9_2","first-page":"617","volume-title":"Proceedings of the International Conference on Parallel and Distributed Computing and Systems (IASTED\u201901)","author":"Beyls Kristof","year":"2001","unstructured":"Kristof Beyls and Erik D\u2019Hollander. 2001. Reuse distance as a metric for cache behavior. In Proceedings of the International Conference on Parallel and Distributed Computing and Systems (IASTED\u201901). 617\u2013622."},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","DOI":"10.1145\/1454115.1454128"},{"key":"e_1_3_2_11_2","doi-asserted-by":"publisher","DOI":"10.1145\/782814.782836"},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-49956-7_4"},{"key":"e_1_3_2_13_2","doi-asserted-by":"publisher","DOI":"10.1109\/IISWC.2009.5306797"},{"key":"e_1_3_2_14_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2005.27"},{"key":"e_1_3_2_15_2","volume-title":"A Composable Model for Analyzing Locality of Multi-threaded Programs","author":"Ding Chen","year":"2009","unstructured":"Chen Ding and Trishul Chilimbi. 2009. A Composable Model for Analyzing Locality of Multi-threaded Programs. Technical Report MSR-TR-2009-107. Retrieved from https:\/\/www.microsoft.com\/en-us\/research\/publication\/a-composable-model-for-analyzing-locality-of-multi-threaded-programs\/."},{"key":"e_1_3_2_16_2","doi-asserted-by":"publisher","DOI":"10.5555\/898449"},{"key":"e_1_3_2_17_2","doi-asserted-by":"publisher","DOI":"10.1145\/781131.781159"},{"key":"e_1_3_2_18_2","article-title":"Instruction-based Sampling: A New Performance Analysis Technique for AMD Family 10h Processors","author":"Drongowski Paul J.","year":"2007","unstructured":"Paul J. Drongowski. 2007. Instruction-based Sampling: A New Performance Analysis Technique for AMD Family 10h Processors. Retrieved from https:\/\/pdfs.semanticscholar.org\/5219\/4b43b8385ce39b2b08ecd409c753e0efafe5.pdf.","journal-title":"Retrieved from https:\/\/pdfs.semanticscholar.org\/5219\/4b43b8385ce39b2b08ecd409c753e0efafe5.pdf"},{"key":"e_1_3_2_19_2","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2012.43"},{"key":"e_1_3_2_20_2","doi-asserted-by":"publisher","DOI":"10.1145\/1854273.1854347"},{"key":"e_1_3_2_21_2","doi-asserted-by":"publisher","DOI":"10.1145\/1944862.1944885"},{"key":"e_1_3_2_22_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISPASS.2010.5452069"},{"key":"e_1_3_2_23_2","doi-asserted-by":"publisher","DOI":"10.1145\/2858788.2688501"},{"key":"e_1_3_2_24_2","article-title":"Optimizing Application Performance on Intel Core Microarchitecture Using Hardware-Implemented Prefetchers","author":"Hegde Ravi","year":"2015","unstructured":"Ravi Hegde. 2015. Optimizing Application Performance on Intel Core Microarchitecture Using Hardware-Implemented Prefetchers. Retrieved from https:\/\/software.intel.com\/content\/www\/us\/en\/develop\/articles\/optimizing-application-performance-on-intel-coret-microarchitecture-using-hardware-implemented-prefetchers.html.","journal-title":"R"},{"key":"e_1_3_2_25_2","doi-asserted-by":"publisher","DOI":"10.1145\/3185751"},{"key":"e_1_3_2_26_2","article-title":"Intel Microarchitecture Codename Nehalem Performance Monitoring Unit Programming Guide","year":"2010","unstructured":"Intel. 2010. Intel Microarchitecture Codename Nehalem Performance Monitoring Unit Programming Guide. Retrieved from https:\/\/software.intel.com\/sites\/default\/files\/m\/5\/2\/c\/f\/1\/30320-Nehalem-PMU-Programming-Guide-Core.pdf.","journal-title":"Retrieved from https:\/\/software.intel.com\/sites\/default\/files\/m\/5\/2\/c\/f\/1\/30320-Nehalem-PMU-Programming-Guide-Core.pdf"},{"key":"e_1_3_2_27_2","first-page":"329","volume-title":"Proceedings of the International Conference on Computer Science, Electronics and Communication Engineering (CSECE\u201918)","author":"Ji Kecheng","unstructured":"Kecheng Ji, Ming Ling, and Li Liu. 2018\/02. A probability model of calculating L2 cache misses. In Proceedings of the International Conference on Computer Science, Electronics and Communication Engineering (CSECE\u201918). Atlantis Press, 329\u2013332. DOI: https:\/\/doi.org\/10.2991\/csece-18.2018.71"},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.micpro.2017.10.001"},{"key":"e_1_3_2_29_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.micpro.2017.02.005"},{"key":"e_1_3_2_30_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-11970-5_15"},{"key":"e_1_3_2_31_2","doi-asserted-by":"publisher","DOI":"10.1145\/964750.801837"},{"key":"e_1_3_2_32_2","doi-asserted-by":"publisher","DOI":"10.1145\/359863.359878"},{"key":"e_1_3_2_33_2","article-title":"Performance Analysis Guide for Intel Core i7 Processor and Intel Xeon 5500 Processors","author":"Levinthal David","year":"2009","unstructured":"David Levinthal. 2009. Performance Analysis Guide for Intel Core i7 Processor and Intel Xeon 5500 Processors. Retrieved from https:\/\/www.amd.com\/system\/files\/TechDocs\/24594.pdf.","journal-title":"Retrieved from https:\/\/www.amd.com\/system\/files\/TechDocs\/24594.pdf"},{"key":"e_1_3_2_34_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.sysarc.2020.101745"},{"key":"e_1_3_2_35_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2021.3053350"},{"key":"e_1_3_2_36_2","article-title":"perf_event_open\u2014Linux Man Page","year":"2012","unstructured":"Linux. 2012. perf_event_open\u2014Linux Man Page. Retrieved from https:\/\/linux.die.net\/man\/2\/perf_event_open.","journal-title":"Retrieved from https:\/\/linux.die.net\/man\/2\/perf_event_open"},{"key":"e_1_3_2_37_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2017.11"},{"key":"e_1_3_2_38_2","doi-asserted-by":"publisher","DOI":"10.1147\/sj.92.0078"},{"key":"e_1_3_2_39_2","doi-asserted-by":"publisher","DOI":"10.1145\/964750.801833"},{"key":"e_1_3_2_40_2","article-title":"Debugging and Performance Tuning with Library Interposers","author":"Nakhimovsky Greg","year":"2001","unstructured":"Greg Nakhimovsky. 2001. Debugging and Performance Tuning with Library Interposers. Retrieved from http:\/\/dsc.sun.com\/solaris\/articles\/lib_interposers.html.","journal-title":"Retrieved from http:\/\/dsc.sun.com\/solaris\/articles\/lib_interposers.html"},{"key":"e_1_3_2_41_2","doi-asserted-by":"publisher","DOI":"10.1145\/1772954.1772958"},{"key":"e_1_3_2_42_2","doi-asserted-by":"publisher","DOI":"10.1145\/2597652.2597674"},{"key":"e_1_3_2_43_2","doi-asserted-by":"publisher","DOI":"10.1145\/1275571.1275600"},{"key":"e_1_3_2_44_2","doi-asserted-by":"publisher","DOI":"10.1145\/1275571.1275600"},{"key":"e_1_3_2_45_2","doi-asserted-by":"publisher","DOI":"10.1145\/800015.808203"},{"key":"e_1_3_2_46_2","doi-asserted-by":"publisher","DOI":"10.1109\/TC.2017.2723878"},{"key":"e_1_3_2_47_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2019.2896633"},{"key":"e_1_3_2_48_2","doi-asserted-by":"publisher","DOI":"10.1145\/1854273.1854286"},{"key":"e_1_3_2_49_2","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPSW.2010.5470780"},{"key":"e_1_3_2_50_2","doi-asserted-by":"publisher","DOI":"10.1145\/2494232.2465756"},{"key":"e_1_3_2_51_2","volume-title":"Accurate Approximation of Locality from Time Distance Histograms","author":"Shen Xipeng","year":"2006","unstructured":"Xipeng Shen, Jonathan Shaw, and Brian Meeker. 2006. Accurate Approximation of Locality from Time Distance Histograms. Technical Report."},{"key":"e_1_3_2_52_2","doi-asserted-by":"publisher","DOI":"10.1145\/1190216.1190227"},{"issue":"3","key":"e_1_3_2_53_2","first-page":"4:1\u20134:19","article-title":"IBM POWER7 performance modeling, verification, and evaluation","volume":"55","author":"Srinivas M.","year":"2011","unstructured":"M. Srinivas, B. Sinharoy, R. J. Eickemeyer, R. Raghavan, S. Kunkel, T. Chen, W. Maron, D. Flemming, A. Blanchard, P. Seshadri, J. W. Kellington, A. Mericas, A. E. Petruski, V. R. Indukuru, and S. Reyes. 2011. IBM POWER7 performance modeling, verification, and evaluation. IBM JRD 55, 3 (May-June 2011), 4:1\u20134:19.","journal-title":"IBM JRD"},{"key":"e_1_3_2_54_2","doi-asserted-by":"publisher","DOI":"10.1145\/225830.224449"},{"key":"e_1_3_2_55_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2017.2703149"},{"key":"e_1_3_2_56_2","doi-asserted-by":"crossref","first-page":"116","DOI":"10.1007\/978-3-319-41321-1_7","volume-title":"High Performance Computing","author":"Unat Didem","year":"2016","unstructured":"Didem Unat, Tan Nguyen, Weiqun Zhang, Muhammed Nufail Farooqi, Burak Bastem, George Michelogiannakis, Ann Almgren, and John Shalf. 2016. TiDA: High-level programming abstractions for data locality management. In High Performance Computing, Julian M. Kunkel, Pavan Balaji, and Jack Dongarra (Eds.). Springer International Publishing, Cham, 116\u2013135."},{"key":"e_1_3_2_57_2","doi-asserted-by":"publisher","DOI":"10.1109\/LCA.2017.2701370"},{"key":"e_1_3_2_58_2","article-title":"Intel Memory Latency Checker v3.9","author":"Viswanathan Krishnaswamy","year":"2013","unstructured":"Krishnaswamy Viswanathan. 2013. Intel Memory Latency Checker v3.9. Retrieved from https:\/\/software.intel.com\/content\/www\/us\/en\/develop\/articles\/intelr-memory-latency-checker.html.","journal-title":"Retrieved from https:\/\/software.intel.com\/content\/www\/us\/en\/develop\/articles\/intelr-memory-latency-checker.html"},{"key":"e_1_3_2_59_2","doi-asserted-by":"publisher","DOI":"10.1145\/3453165"},{"key":"e_1_3_2_60_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICPADS47876.2019.00046"},{"key":"e_1_3_2_61_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2019.00056"},{"key":"e_1_3_2_62_2","doi-asserted-by":"publisher","DOI":"10.1145\/3296957.3177159"},{"key":"e_1_3_2_63_2","doi-asserted-by":"publisher","DOI":"10.1145\/3296957.3177159"},{"key":"e_1_3_2_64_2","doi-asserted-by":"publisher","DOI":"10.1109\/PACT.2011.58"},{"key":"e_1_3_2_65_2","doi-asserted-by":"publisher","DOI":"10.1145\/2499368.2451153"},{"key":"e_1_3_2_66_2","doi-asserted-by":"publisher","DOI":"10.1145\/1552309.1552310"}],"container-title":["ACM Transactions on Architecture and Code Optimization"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3484199","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3484199","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T20:17:13Z","timestamp":1750191433000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3484199"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,12,6]]},"references-count":65,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2022,3,31]]}},"alternative-id":["10.1145\/3484199"],"URL":"https:\/\/doi.org\/10.1145\/3484199","relation":{},"ISSN":["1544-3566","1544-3973"],"issn-type":[{"value":"1544-3566","type":"print"},{"value":"1544-3973","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,12,6]]},"assertion":[{"value":"2021-05-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2021-08-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2021-12-06","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}