{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,25]],"date-time":"2026-06-25T16:46:25Z","timestamp":1782405985445,"version":"3.54.5"},"reference-count":67,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2026,6,25]],"date-time":"2026-06-25T00:00:00Z","timestamp":1782345600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Archit. Code Optim."],"published-print":{"date-parts":[[2026,6,30]]},"abstract":"<jats:p>Sparse data structures are ubiquitous in graph analytics, machine learning, and high-performance computing. Algorithms operating on these structures typically exhibit highly irregular data-dependent memory access (DDMA) patterns, leading to frequent cache misses and degraded memory performance. Prior work on hardware prefetching to mitigate DDMA-induced misses falls into two categories: address-based methods that sample correlated sequences of load data and addresses from cache miss streams, and instruction-based methods that record instruction-level dependency chains. Although both learn single relations effectively, they struggle with multi-level range relations prevalent in DDMA-intensive workloads, leaving substantial prefetching opportunities unexploited. In address-based schemes, misses from deeper-level consumers are often miscorrelated with the producer. Moreover, out-of-order execution and the range relations themselves perturb sampling, yielding mismatched load instances. In instruction-based schemes, chain-structured representations and suboptimal learning strategies prevent the construction of complete dependency chains for these relations.<\/jats:p>\n                  <jats:p>To overcome these limitations, we present Thoth, a hardware prefetcher that operates at the granularity of explicit producer-consumer load pairs rather than constructing dependency chains. Thoth detects such pairs via register-level dependency tracking. It adopts an annotation-directed load sampling strategy that annotates matched producer-consumer load instances and samples only those annotated instances, thereby robustly uncovering DDMA patterns\u2014including multi-level range relations\u2014while avoiding mismatches. To maintain annotation correctness across pipeline flushes, Thoth employs precise load annotation, which leverages reorder identifiers to resume or terminate annotation precisely. On a suite of DDMA-intensive benchmarks, Thoth delivers a 51.1% speedup over a no-prefetching baseline and outperforms two state-of-the-art DDMA prefetchers by 14.7% and 8.2%, respectively.<\/jats:p>","DOI":"10.1145\/3806835","type":"journal-article","created":{"date-parts":[[2026,4,3]],"date-time":"2026-04-03T10:10:06Z","timestamp":1775211006000},"page":"1-26","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Thoth: Uncovering Data-Dependent Memory Access Patterns via Annotation-Directed Load Sampling"],"prefix":"10.1145","volume":"23","author":[{"ORCID":"https:\/\/orcid.org\/0009-0006-8220-5238","authenticated-orcid":false,"given":"Kanheng","family":"Jiang","sequence":"first","affiliation":[{"name":"State Key Laboratory of Integrated Chips and Systems, Fudan University","place":["Shanghai, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-2997-5285","authenticated-orcid":false,"given":"Yongxin","family":"Lyu","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Integrated Chips and Systems, Fudan University","place":["Shanghai, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-2072-075X","authenticated-orcid":false,"given":"Zhiyuan","family":"Zhang","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Integrated Chips and Systems, Fudan University","place":["Shanghai, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-5575-8440","authenticated-orcid":false,"given":"Zengshi","family":"Wang","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Integrated Chips and Systems, Fudan University","place":["Shanghai, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9286-0275","authenticated-orcid":false,"given":"Chao","family":"Fu","sequence":"additional","affiliation":[{"name":"Shaoxin Laboratory, Fudan University","place":["Shaoxing, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5245-0754","authenticated-orcid":false,"given":"Jun","family":"Han","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Integrated Chips and Systems, Fudan University","place":["Shanghai, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,6,25]]},"reference":[{"key":"e_1_3_2_2_2","doi-asserted-by":"publisher","DOI":"10.1145\/3725843.3756081"},{"key":"e_1_3_2_3_2","doi-asserted-by":"publisher","DOI":"10.1145\/2925426.2926254"},{"key":"e_1_3_2_4_2","doi-asserted-by":"publisher","DOI":"10.5555\/3049832.3049865"},{"key":"e_1_3_2_5_2","doi-asserted-by":"publisher","DOI":"10.1145\/3296957.3173189"},{"key":"e_1_3_2_6_2","doi-asserted-by":"publisher","DOI":"10.1145\/3319393"},{"key":"e_1_3_2_7_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA59077.2024.00090"},{"key":"e_1_3_2_8_2","doi-asserted-by":"publisher","DOI":"10.1145\/384285.379251"},{"key":"e_1_3_2_9_2","doi-asserted-by":"publisher","DOI":"10.1145\/2228360.2228584"},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","DOI":"10.1145\/125826.125925"},{"key":"e_1_3_2_11_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE.2013.6544839"},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2019.00051"},{"key":"e_1_3_2_13_2","unstructured":"Scott Beamer Krste Asanovi\u0107 and David Patterson. 2015. The GAP benchmark suite. arXiv:1508.03619. Retrieved from https:\/\/arxiv.org\/abs\/1508.03619"},{"key":"e_1_3_2_14_2","doi-asserted-by":"publisher","DOI":"10.1145\/3352460.3358325"},{"key":"e_1_3_2_15_2","doi-asserted-by":"publisher","DOI":"10.1145\/2024716.2024718"},{"key":"e_1_3_2_16_2","doi-asserted-by":"publisher","DOI":"10.1109\/HCS59251.2023.10254718"},{"key":"e_1_3_2_17_2","doi-asserted-by":"publisher","DOI":"10.1145\/3185768.3185771"},{"key":"e_1_3_2_18_2","first-page":"578","volume-title":"Proceedings of the 13th USENIX Symposium on Operating Systems Design and Implementation18)","author":"Chen Tianqi","year":"2018","unstructured":"Tianqi Chen, Thierry Moreau, Ziheng Jiang, Lianmin Zheng, Eddie Yan, Haichen Shen, Meghan Cowan, Leyuan Wang, Yuwei Hu, Luis Ceze, et\u00a0al. 2018. \\(\\lbrace\\) TVM \\(\\rbrace\\) : An automated \\(\\lbrace\\) End-to-End \\(\\rbrace\\) optimizing compiler for deep learning. In Proceedings of the 13th USENIX Symposium on Operating Systems Design and Implementation18). 578\u2013594."},{"key":"e_1_3_2_19_2","doi-asserted-by":"publisher","DOI":"10.1109\/12.381947"},{"key":"e_1_3_2_20_2","doi-asserted-by":"publisher","DOI":"10.1145\/605432.605427"},{"key":"e_1_3_2_21_2","unstructured":"Ganesh Suryanarayan Dasika and Rune Holm. 2015. Data processing method and apparatus for prefetching. US Patent 9 037 835."},{"key":"e_1_3_2_22_2","first-page":"150","article-title":"Toward a new metric for ranking high performance computing systems","volume":"312","author":"Dongarra Jack","year":"2013","unstructured":"Jack Dongarra and Michael A. Heroux. 2013. Toward a new metric for ranking high performance computing systems. Sandia Report, SAND2013-4744 312 (2013), 150.","journal-title":"Sandia Report, SAND2013-4744"},{"key":"e_1_3_2_23_2","doi-asserted-by":"publisher","DOI":"10.1109\/LCA.2024.3361925"},{"key":"e_1_3_2_24_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA57654.2024.00040"},{"key":"e_1_3_2_25_2","doi-asserted-by":"publisher","DOI":"10.1145\/3695053.3731054"},{"key":"e_1_3_2_26_2","doi-asserted-by":"publisher","DOI":"10.1145\/3725843.3756133"},{"key":"e_1_3_2_27_2","doi-asserted-by":"publisher","DOI":"10.1145\/1186736.1186737"},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO56248.2022.00071"},{"key":"e_1_3_2_29_2","first-page":"206","volume-title":"Proceedings of the 6th International Symposium on High-Performance Computer Architecture. HPCA-6 (Cat. No. PR00550)","author":"Karlsson Magnus","year":"2000","unstructured":"Magnus Karlsson, Fredrik Dahlgren, and Per Stenstrom. 2000. A prefetching technique for irregular accesses to linked data structures. In Proceedings of the 6th International Symposium on High-Performance Computer Architecture. HPCA-6 (Cat. No. PR00550). IEEE, 206\u2013217."},{"key":"e_1_3_2_30_2","doi-asserted-by":"publisher","DOI":"10.1145\/3439803"},{"key":"e_1_3_2_31_2","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2016.7783763"},{"key":"e_1_3_2_32_2","article-title":"SNAP Datasets: Stanford Large Network Dataset Collection","author":"Leskovec Jure","year":"2014","unstructured":"Jure Leskovec and Andrej Krevl. 2014. SNAP Datasets: Stanford Large Network Dataset Collection. Retrieved April 10, 2026 from http:\/\/snap.stanford.edu\/data","journal-title":"Retrieved April 10, 2026 from"},{"key":"e_1_3_2_33_2","doi-asserted-by":"publisher","DOI":"10.1145\/1669112.1669172"},{"key":"e_1_3_2_34_2","article-title":"Arm\u00ae Neoverse\u2122 V2 Core Technical Reference Manual","author":"Limited Arm","year":"2022","unstructured":"Arm Limited. 2022. Arm\u00ae Neoverse\u2122 V2 Core Technical Reference Manual. Retrieved April 10, 2026 from https:\/\/developer.arm.com\/documentation\/102375\/0002","journal-title":"R"},{"key":"e_1_3_2_35_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCAD.2024.3442981"},{"key":"e_1_3_2_36_2","doi-asserted-by":"publisher","DOI":"10.1145\/237090.237190"},{"key":"e_1_3_2_37_2","doi-asserted-by":"publisher","DOI":"10.1142\/S0129626407002843"},{"key":"e_1_3_2_38_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2016.7446087"},{"key":"e_1_3_2_39_2","doi-asserted-by":"publisher","DOI":"10.1145\/1298306.1298311"},{"key":"e_1_3_2_40_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA52012.2021.00024"},{"key":"e_1_3_2_41_2","doi-asserted-by":"publisher","DOI":"10.1145\/3613424.3614255"},{"key":"e_1_3_2_42_2","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO56248.2022.00072"},{"key":"e_1_3_2_43_2","volume-title":"The PageRank Citation Ranking: Bringing Order to the Web.","author":"Page Lawrence","year":"1999","unstructured":"Lawrence Page, Sergey Brin, Rajeev Motwani, and Terry Winograd. 1999. The PageRank Citation Ranking: Bringing Order to the Web.Technical Report. Stanford infolab."},{"key":"e_1_3_2_44_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA45697.2020.00021"},{"key":"e_1_3_2_45_2","doi-asserted-by":"publisher","DOI":"10.1109\/SC.2014.82"},{"key":"e_1_3_2_46_2","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO61859.2024.00101"},{"key":"e_1_3_2_47_2","doi-asserted-by":"publisher","DOI":"10.1145\/291069.291034"},{"key":"e_1_3_2_48_2","doi-asserted-by":"publisher","DOI":"10.1145\/300979.300989"},{"key":"e_1_3_2_49_2","doi-asserted-by":"publisher","DOI":"10.1145\/605432.605403"},{"key":"e_1_3_2_50_2","unstructured":"Alexander Cole Shulyak Joseph Michael Pusdesris Adrian Montero and Balaji Vijayan. 2022. Determining prefetch patterns with discontinuous strides. US Patent 11 385 896."},{"key":"e_1_3_2_51_2","doi-asserted-by":"publisher","DOI":"10.1145\/2442516.2442530"},{"key":"e_1_3_2_52_2","doi-asserted-by":"publisher","DOI":"10.1145\/327070.327125"},{"key":"e_1_3_2_53_2","article-title":"Synopsys Design Compiler","unstructured":"synopsys. [n. d.]. Synopsys Design Compiler. Retrieved April 10, 2026 from https:\/\/www.synopsys.com\/implementation-and-signoff\/rtl-synthesis-test\/dc-ultra.html","journal-title":"R"},{"key":"e_1_3_2_54_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA51647.2021.00061"},{"key":"e_1_3_2_55_2","article-title":"Intel Core i7-3770 Specifications","unstructured":"TechPowerUp. [n. d.]. Intel Core i7-3770 Specifications. Retrieved April 10, 2026 from https:\/\/www.techpowerup.com\/cpu-specs\/core-i7-3770.c1002","journal-title":"R"},{"key":"e_1_3_2_56_2","article-title":"Intel Xeon E7-8860 Specifications","unstructured":"TechPowerUp. [n. d.]. Intel Xeon E7-8860 Specifications. Retrieved April 10, 2026 from https:\/\/www.techpowerup.com\/cpu-specs\/xeon-e7-8860.c1467","journal-title":"R"},{"key":"e_1_3_2_57_2","article-title":"TSMC 28nm Technology","unstructured":"TSMC. [n. d.]. TSMC 28nm Technology. Retrieved April 10, 2026 from https:\/\/www.tsmc.com\/english\/dedicatedFoundry\/technology\/logic\/l_28nm","journal-title":"R"},{"key":"e_1_3_2_58_2","doi-asserted-by":"publisher","DOI":"10.1145\/2465351.2465371"},{"key":"e_1_3_2_59_2","article-title":"The RISC-V instruction set manual volume II: Privileged architecture version 1.7","author":"Waterman Andrew","year":"2016","unstructured":"Andrew Waterman, Yunsup Lee, Rimas Avizienis, David A. Patterson, and Krste Asanovic. 2016. The RISC-V instruction set manual volume II: Privileged architecture version 1.7. EECS Department, University of California, Berkeley, Tech. Rep. UCB\/EECS-2016-129 (2016).","journal-title":"EECS Department, University of California, Berkeley, Tech. Rep. UCB\/EECS-2016-129"},{"key":"e_1_3_2_60_2","first-page":"4","article-title":"The RISC-V instruction set manual, volume I: User-level ISA, version 2.0","author":"Waterman Andrew","year":"2014","unstructured":"Andrew Waterman, Yunsup Lee, David A. Patterson, and Krste Asanovic. 2014. The RISC-V instruction set manual, volume I: User-level ISA, version 2.0. EECS Department, University of California, Berkeley, Tech. Rep. UCB\/EECS-2014-54 (2014), 4.","journal-title":"EECS Department, University of California, Berkeley, Tech. Rep. UCB\/EECS-2014-54"},{"key":"e_1_3_2_61_2","doi-asserted-by":"publisher","DOI":"10.1109\/TC.2021.3065909"},{"key":"e_1_3_2_62_2","doi-asserted-by":"publisher","DOI":"10.1145\/3352460.3358300"},{"key":"e_1_3_2_63_2","doi-asserted-by":"publisher","DOI":"10.1145\/3307650.3322225"},{"key":"e_1_3_2_64_2","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO56248.2022.00080"},{"key":"e_1_3_2_65_2","doi-asserted-by":"publisher","DOI":"10.1145\/3641853"},{"key":"e_1_3_2_66_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISPASS.2014.6844459"},{"key":"e_1_3_2_67_2","doi-asserted-by":"publisher","DOI":"10.1145\/2830772.2830807"},{"key":"e_1_3_2_68_2","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO50266.2020.00057"}],"container-title":["ACM Transactions on Architecture and Code Optimization"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3806835","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,25]],"date-time":"2026-06-25T15:54:35Z","timestamp":1782402875000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3806835"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6,25]]},"references-count":67,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2026,6,30]]}},"alternative-id":["10.1145\/3806835"],"URL":"https:\/\/doi.org\/10.1145\/3806835","relation":{},"ISSN":["1544-3566","1544-3973"],"issn-type":[{"value":"1544-3566","type":"print"},{"value":"1544-3973","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,6,25]]},"assertion":[{"value":"2025-09-16","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-03-30","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-06-25","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}