{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,25]],"date-time":"2026-06-25T16:46:27Z","timestamp":1782405987243,"version":"3.54.5"},"reference-count":49,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2026,6,25]],"date-time":"2026-06-25T00:00:00Z","timestamp":1782345600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by-nc-nd\/4.0\/legalcode"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Archit. Code Optim."],"published-print":{"date-parts":[[2026,6,30]]},"abstract":"<jats:p>Accelerating load requests is an effective way to improve processor performance through reducing load request latency. Currently, state-of-the-art methods, including Hermes and TLP, are limited to accelerating off-chip load requests that are correctly predicted and cannot accelerate mispredicted off-chip load requests and any on-chip load requests. To overcome this limitation, we propose a new technique called Apollo that integrates TLP. The key innovation of Apollo lies in its transformation of the perceptron-based off-chip prediction paradigm, pioneered by Hermes, into a comprehensive multi-level cache miss prediction technique. The workflow of Apollo is as follows: (1) predicting whether a load request will miss the L1D or L2, and (2) performing arbitration to decide whether to issue a speculative load request and to which cache level the request is issued, and (3) issuing a speculative load request to the lower-level cache (either L2 or LLC) after arbitration for those predicted to miss the L1D or L2, while allowing the regular load request to concurrently access the cache hierarchy. If the prediction is correct, the regular load request eventually misses the L1D or L2 and waits for the speculative load request to finish. Therefore, Apollo can hide the L1D access latency for correctly predicted L1D miss load requests, and both L1D and L2 access latency for correctly predicted L2 miss load requests.<\/jats:p>\n                  <jats:p>To enable Apollo, we propose a lightweight L1D miss load predictor (L1MP), a lightweight L2 miss load predictor (L2MP), and an arbiter called Athena. L1MP and L2MP predict whether load requests will miss in the L1D and L2, respectively, while Athena performs arbitration to control the issuance of speculative load requests. Our evaluation using a diverse set of workloads shows that Apollo provides a geometric mean (geomean) performance improvement of 15.2% for the single-core processor, outperforming Hermes by 9.6% and TLP by 6.4%. For the multi-core processor, Apollo provides a geomean performance improvement of 18.5%, which is 16.8% higher than Hermes and 5.8% higher than TLP.<\/jats:p>","DOI":"10.1145\/3818683","type":"journal-article","created":{"date-parts":[[2026,5,28]],"date-time":"2026-05-28T11:25:31Z","timestamp":1779967531000},"page":"1-25","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Apollo: Accelerating Load Requests via Multi-Level Cache Miss Load Prediction"],"prefix":"10.1145","volume":"23","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-9914-1254","authenticated-orcid":false,"given":"Zhengwei","family":"Huang","sequence":"first","affiliation":[{"name":"College of Computer Science and Technology, National University of Defense Technology","place":["Changsha, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-6704-0966","authenticated-orcid":false,"given":"Wei","family":"Guo","sequence":"additional","affiliation":[{"name":"College of Computer Science and Technology, National University of Defense Technology","place":["Changsha, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-2514-2052","authenticated-orcid":false,"given":"Yongwen","family":"Wang","sequence":"additional","affiliation":[{"name":"College of Computer Science and Technology, National University of Defense Technology","place":["Changsha, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,6,25]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA59077.2024.00090"},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2019.00053"},{"key":"e_1_3_1_4_2","unstructured":"Scott Beamer Krste Asanovi\u0107 and David Patterson. 2015. The GAP benchmark suite. arXiv:1508.03619. Retrieved from https:\/\/arxiv.org\/abs\/1508.03619"},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO56248.2022.00015"},{"key":"e_1_3_1_6_2","doi-asserted-by":"publisher","DOI":"10.1145\/3466752.3480114"},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","unstructured":"Rahul Bera Zhenrong Lang Caroline Hengartner Konstantinos Kanellopoulos Rakesh Kumar Mohammad Sadrosadati and Onur Mutlu. 2026. Athena: Synergizing data prefetching and off-chip prediction via online reinforcement learning. In 2026 IEEE International Symposium on High Performance Computer Architecture (HPCA). 1\u201319. DOI:10.1109\/HPCA68181.2026.11408449","DOI":"10.1109\/HPCA68181.2026.11408449"},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.1145\/3352460.3358325"},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1145\/3307650.3322207"},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA61900.2025.00024"},{"key":"e_1_3_1_11_2","unstructured":"ChampSim Contributors. 2021. ChampSim: An open-source trace based simulator. Retrieved October 6 2025 from https:\/\/github.com\/ChampSim\/ChampSim"},{"key":"e_1_3_1_12_2","unstructured":"The Standard Performance Evaluation Corporation. 2006. SPEC CPU 2006. Retrieved October 6 2025 from https:\/\/www.spec.org\/cpu2006\/"},{"key":"e_1_3_1_13_2","unstructured":"The Standard Performance Evaluation Corporation. 2017. SPEC CPU 2017. Retrieved October 6 2025 from https:\/\/www.spec.org\/cpu2017\/"},{"key":"e_1_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.1145\/3675398"},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA59077.2024.00088"},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA57654.2024.00040"},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1145\/3307650.3322217"},{"key":"e_1_3_1_18_2","unstructured":"Nathan Gober Gino Chacon Lei Wang Paul V. Gratz Daniel A. Jimenez Elvira Teran Seth Pugsley and Jinchun Kim. 2022. The championship simulator: Architectural simulation for education and competition. arXiv:2210.14324. Retrieved from https:\/\/arxiv.org\/abs\/2210.14324"},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","DOI":"10.1145\/3656019.3689613"},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA53966.2022.00054"},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA57654.2024.00046"},{"key":"e_1_3_1_22_2","doi-asserted-by":"crossref","unstructured":"He Jiang Liuwei Fu Dong Liu Zhilei Ren Yuting Chen and Lei Qiao. 2025. TRACED: A Temporal Graph Neural Networks-based Model for Data Prefetching. ACM Transactions on Architecture and Code Optimization 22 3 (2025) 1\u201325.","DOI":"10.1145\/3747843"},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO56248.2022.00071"},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2003.1253199"},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1145\/571637.571639"},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.1145\/3123939.3123942"},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2016.7783763"},{"key":"e_1_3_1_28_2","unstructured":"Chester Lam. 2022. A Preview of Raptor Lake\u2019s Improved L2 Caches. Retrieved October 5 2025 from https:\/\/chipsandcheese.com\/p\/a-preview-of-raptor-lakes-improved-l2-caches"},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA.2005.49"},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2003.1183532"},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO56248.2022.00072"},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA45697.2020.00021"},{"key":"e_1_3_1_33_2","doi-asserted-by":"publisher","DOI":"10.1145\/3613424.3614245"},{"key":"e_1_3_1_34_2","doi-asserted-by":"publisher","DOI":"10.1145\/3345000"},{"key":"e_1_3_1_35_2","doi-asserted-by":"publisher","DOI":"10.1145\/2678373.2665694"},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2017.25"},{"issue":"2019","key":"e_1_3_1_37_2","article-title":"Multi-lookahead offset prefetching","author":"Shakerinava Mehran","year":"2019","unstructured":"Mehran Shakerinava, Mohammad Bakhshalipour, Pejman Lotfi-Kamran, and Hamid Sarbazi-Azad. 2019. Multi-lookahead offset prefetching. The Third Data Prefetching Championship2019 (2019).","journal-title":"The Third Data Prefetching Championship"},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","DOI":"10.1145\/3445814.3446752"},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","DOI":"10.1145\/2442516.2442530"},{"key":"e_1_3_1_40_2","doi-asserted-by":"publisher","DOI":"10.1145\/1089008.1089011"},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2016.7783705"},{"key":"e_1_3_1_42_2","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO56248.2022.00070"},{"key":"e_1_3_1_43_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA61900.2025.00025"},{"key":"e_1_3_1_44_2","unstructured":"Wikipedia. 2019. Cascade Lake - Microarchitectures - Intel. Retrieved October 6 2025 from https:\/\/en.wikipedia.org\/wiki\/Cascade_Lake"},{"key":"e_1_3_1_45_2","unstructured":"Wikipedia. 2025. Sunny Cove (microarchitecture). Retrieved October 5 2025 from https:\/\/en.wikipedia.org\/wiki\/Sunny_Cove_%28microarchitecture%29"},{"key":"e_1_3_1_46_2","unstructured":"Wikipedia. 2025. Zen 5. Retrieved October 5 2025 from https:\/\/en.wikipedia.org\/wiki\/Zen_5"},{"key":"e_1_3_1_47_2","first-page":"1430","volume-title":"Proceedings of the 31st ACM International Conference on Architectural Support for Programming Languages and Operating Systems","author":"Xu Ceyu","year":"2026","unstructured":"Ceyu Xu, Xiangfeng Sun, Weihang Li, Chen Bai, Bangyan Wang, Mengming Li, Zhiyao Xie, and Yuan Xie. 2026. PF-LLM: Large language model hinted hardware prefetching. In Proceedings of the 31st ACM International Conference on Architectural Support for Programming Languages and Operating Systems. 1430\u20131444."},{"key":"e_1_3_1_48_2","doi-asserted-by":"publisher","DOI":"10.1145\/3641853"},{"key":"e_1_3_1_49_2","doi-asserted-by":"publisher","DOI":"10.1145\/3762997"},{"key":"e_1_3_1_50_2","doi-asserted-by":"publisher","DOI":"10.1145\/300979.300983"}],"container-title":["ACM Transactions on Architecture and Code Optimization"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3818683","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,25]],"date-time":"2026-06-25T15:53:50Z","timestamp":1782402830000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3818683"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6,25]]},"references-count":49,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2026,6,30]]}},"alternative-id":["10.1145\/3818683"],"URL":"https:\/\/doi.org\/10.1145\/3818683","relation":{},"ISSN":["1544-3566","1544-3973"],"issn-type":[{"value":"1544-3566","type":"print"},{"value":"1544-3973","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,6,25]]},"assertion":[{"value":"2025-10-14","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-05-18","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-06-25","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}