{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,14]],"date-time":"2026-07-14T12:35:05Z","timestamp":1784032505149,"version":"3.55.0"},"reference-count":74,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2024,9,14]],"date-time":"2024-09-14T00:00:00Z","timestamp":1726272000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"National Key Research and Development Program of China NKRDPC","award":["2022YFB4500403 NKRDPC"],"award-info":[{"award-number":["2022YFB4500403 NKRDPC"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Archit. Code Optim."],"published-print":{"date-parts":[[2024,9,30]]},"abstract":"<jats:p>The growing memory demands of modern applications have driven the adoption of far memory technologies in data centers to provide cost-effective, high-capacity memory solutions. However, far memory presents new performance challenges because its access latencies are significantly longer and more variable than local DRAM. For applications to achieve acceptable performance on far memory, a high degree of memory-level parallelism (MLP) is needed to tolerate the long access latency.<\/jats:p>\n          <jats:p>While modern out-of-order processors are capable of exploiting a certain degree of MLP, they are constrained by resource limitations and hardware complexity. The key obstacle is the synchronous memory access semantics of traditional load\/store instructions, which occupy critical hardware resources for a long time. The longer far memory latencies exacerbate this limitation.<\/jats:p>\n          <jats:p>This article proposes a set of Asynchronous Memory Access Instructions (AMI) and its supporting function unit, Asynchronous Memory Access Unit (AMU), inside contemporary Out-of-Order Core. AMI separates memory request issuing from response handling to reduce resource occupation. Additionally, AMU architecture supports up to several hundreds of asynchronous memory requests through re-purposing a portion of L2 Cache as scratchpad memory (SPM) to provide sufficient temporal storage. Together with a coroutine-based programming framework, this scheme can achieve significantly higher MLP for hiding far memory latencies.<\/jats:p>\n          <jats:p>Evaluation with a cycle-accurate simulation shows AMI achieves 2.42\u00d7 speedup on average for memory-bound benchmarks with 1\u03bcs additional far memory latency. Over 130 outstanding requests are supported with 26.86\u00d7 speedup for GUPS (random access) with 5 \u03bcs latency. These demonstrate how the techniques tackle far memory performance impacts through explicit MLP expression and latency adaptation.<\/jats:p>","DOI":"10.1145\/3663479","type":"journal-article","created":{"date-parts":[[2024,5,9]],"date-time":"2024-05-09T08:32:53Z","timestamp":1715243573000},"page":"1-28","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":8,"title":["Asynchronous Memory Access Unit: Exploiting Massive Parallelism for Far Memory Access"],"prefix":"10.1145","volume":"21","author":[{"ORCID":"https:\/\/orcid.org\/0009-0006-2508-7175","authenticated-orcid":false,"given":"Luming","family":"Wang","sequence":"first","affiliation":[{"name":"State Key Lab of Processors, Institute of Computing Technology Chinese Academy of Sciences, Beijing, China and University of Chinese Academy of Sciences, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0009-0027-9818","authenticated-orcid":false,"given":"Xu","family":"Zhang","sequence":"additional","affiliation":[{"name":"State Key Lab of Processors, Institute of Computing Technology Chinese Academy of Sciences, Beijing China and University of Chinese Academy of Sciences, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0003-2732-7820","authenticated-orcid":false,"given":"Songyue","family":"Wang","sequence":"additional","affiliation":[{"name":"State Key Lab of Processors, Institute of Computing Technology Chinese Academy of Sciences, Beijing China and University of Chinese Academy of Sciences, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-9434-7399","authenticated-orcid":false,"given":"Zhuolun","family":"Jiang","sequence":"additional","affiliation":[{"name":"State Key Lab of Processors, Institute of Computing Technology Chinese Academy of Sciences, Beijing China and University of Chinese Academy of Sciences, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0001-8808-0935","authenticated-orcid":false,"given":"Tianyue","family":"Lu","sequence":"additional","affiliation":[{"name":"State Key Lab of Processors, Institute of Computing Technology Chinese Academy of Sciences, Beijing China and University of Chinese Academy of Sciences, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4469-1037","authenticated-orcid":false,"given":"Mingyu","family":"Chen","sequence":"additional","affiliation":[{"name":"State Key Lab of Processors, Institute of Computing Technology Chinese Academy of Sciences, Beijing China and University of Chinese Academy of Sciences, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-8120-7343","authenticated-orcid":false,"given":"Siwei","family":"Luo","sequence":"additional","affiliation":[{"name":"Huawei Technologies Co Ltd, Shenzhen, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0009-5496-5522","authenticated-orcid":false,"given":"Keji","family":"Huang","sequence":"additional","affiliation":[{"name":"Huawei Technologies Co Ltd, Shenzhen, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2024,9,14]]},"reference":[{"key":"e_1_3_2_2_2","unstructured":"2021. NVIDIA CONNECTX-6 Datasheet. https:\/\/www.nvidia.com\/content\/dam\/en-zz\/Solutions\/networking\/ethernet-adapters\/connectX-6-dx-datasheet.pdf [Online; accessed: 2024]."},{"key":"e_1_3_2_3_2","unstructured":"2017. Cray MTA-2 System. (2017). Retrieved from ttp:\/\/www.cray.com\/products\/programs\/mta_2\/[Online]."},{"key":"e_1_3_2_4_2","unstructured":"2017. openCAPI Specification. (2017). Retrieved from http:\/\/opencapi.org[Online; accessed: Febrary 2022]."},{"key":"e_1_3_2_5_2","unstructured":"2018. Gen-Z Specification. (2018). Retrieved from https:\/\/genzconsortium.org\/specifications[Online; accessed: Febrary 2022]."},{"key":"e_1_3_2_6_2","unstructured":"2020. IBM Reveals Next-Generation IBM POWER10 Processor. (2020). Retrieved from https:\/\/newsroom.ibm.com\/2020-08-17-IBM-Reveals-Next-Generation-IBM-POWER10-Processor[Online; accessed: Febrary 2022]."},{"key":"e_1_3_2_7_2","unstructured":"2022. Intel Optane Persistent Memory. (2022). Retrieved from https:\/\/www.intel.com\/content\/www\/us\/en\/architecture-and-technology\/optane-dc-persistent-memory.html[Online; accessed: Febrary 2022]."},{"key":"e_1_3_2_8_2","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2023.3295848"},{"key":"e_1_3_2_9_2","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2023.3295848"},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","DOI":"10.23919\/DATE.2019.8715034"},{"key":"e_1_3_2_11_2","volume-title":"The Rocket Chip Generator","author":"Asanovi\u0107 Krste","year":"2016","unstructured":"Krste Asanovi\u0107, Rimas Avizienis, Jonathan Bachrach, Scott Beamer, David Biancolin, Christopher Celio, Henry Cook, Daniel Dabbelt, John Hauser, Adam Izraelevitz, Sagar Karandikar, Ben Keller, Donggyu Kim, John Koenig, Yunsup Lee, Eric Love, Martin Maas, Albert Magyar, Howard Mao, Miquel Moreto, Albert Ou, David A. Patterson, Brian Richards, Colin Schmidt, Stephen Twigg, Huy Vo, and Andrew Waterman. 2016. The Rocket Chip Generator. Technical Report UCB\/EECS-2016-17. EECS Department, University of California, Berkeley. Retrieved from http:\/\/www2.eecs.berkeley.edu\/Pubs\/TechRpts\/2016\/EECS-2016-17.html"},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.1145\/3289602.3293901"},{"key":"e_1_3_2_13_2","doi-asserted-by":"publisher","DOI":"10.1145\/125826.125925"},{"key":"e_1_3_2_14_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2018.00021"},{"key":"e_1_3_2_15_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2019.00053"},{"key":"e_1_3_2_16_2","doi-asserted-by":"crossref","first-page":"362","DOI":"10.1109\/ICDE.2013.6544839","volume-title":"2013 IEEE 29th International Conference on Data Engineering (ICDE\u201913)","author":"Balkesen Cagri","year":"2013","unstructured":"Cagri Balkesen, Jens Teubner, Gustavo Alonso, and M. Tamer \u00d6zsu. 2013. Main-memory hash joins on multi-core CPUs: Tuning to the underlying hardware. In 2013 IEEE 29th International Conference on Data Engineering (ICDE\u201913). IEEE, 362\u2013373."},{"issue":"3","key":"e_1_3_2_17_2","doi-asserted-by":"crossref","first-page":"17\u2013es","DOI":"10.1145\/1272743.1272747","article-title":"Improving hash join performance through prefetching","volume":"32","author":"Chen Shimin","year":"2007","unstructured":"Shimin Chen, Anastassia Ailamaki, Phillip B. Gibbons, and Todd C. Mowry. 2007. Improving hash join performance through prefetching. ACM Transactions on Database Systems (TODS) 32, 3 (2007), 17\u2013es.","journal-title":"ACM Transactions on Database Systems (TODS)"},{"key":"e_1_3_2_18_2","doi-asserted-by":"publisher","DOI":"10.1109\/PACT.1998.727154"},{"key":"e_1_3_2_19_2","doi-asserted-by":"crossref","first-page":"631","DOI":"10.1145\/2694344.2694359","volume-title":"Proceedings of the Twentieth International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS\u201915)","author":"David Tudor","year":"2015","unstructured":"Tudor David, Rachid Guerraoui, and Vasileios Trigonakis. 2015. Asynchronized concurrency: The secret to scaling concurrent search data structures. In Proceedings of the Twentieth International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS\u201915). Association for Computing Machinery, New York, NY, 631\u2013644. DOI:10.1145\/2694344.2694359"},{"key":"e_1_3_2_20_2","doi-asserted-by":"publisher","DOI":"10.1145\/3548681"},{"key":"e_1_3_2_21_2","doi-asserted-by":"publisher","DOI":"10.2200\/S00581ED1V01Y201405CAC028"},{"key":"e_1_3_2_22_2","first-page":"596","volume-title":"2018 IEEE International Symposium on High Performance Computer Architecture (HPCA\u201918)","author":"Fan Dongrui","year":"2018","unstructured":"Dongrui Fan, Wenming Li, Xiaochun Ye, Da Wang, Hao Zhang, Zhimin Tang, and Ninghui Sun. 2018. Smarco: An efficient many-core processor for high-throughput applications in datacenters. In 2018 IEEE International Symposium on High Performance Computer Architecture (HPCA\u201918). IEEE, 596\u2013607."},{"key":"e_1_3_2_23_2","doi-asserted-by":"publisher","DOI":"10.1145\/3132402.3132409"},{"key":"e_1_3_2_24_2","doi-asserted-by":"publisher","DOI":"10.1109\/TC.2021.3086069"},{"key":"e_1_3_2_25_2","first-page":"28","volume-title":"Proceedings of the 2nd Conference on Computing Frontiers (CF\u201905)","author":"Feo John","year":"2005","unstructured":"John Feo, David Harper, Simon Kahan, and Petr Konecny. 2005. ELDORADO. In Proceedings of the 2nd Conference on Computing Frontiers (CF\u201905). Association for Computing Machinery, New York, NY, 28\u201334. DOI:10.1145\/1062261.1062268"},{"key":"e_1_3_2_26_2","doi-asserted-by":"crossref","unstructured":"Haohuan Fu Junfeng Liao Jinzhe Yang Lanning Wang Zhenya Song Xiaomeng Huang Chao Yang Wei Xue Fangfang Liu Fangli Qiao Wei Zhao Xunqiang Yin Chaofeng Hou Chenglong Zhang Wei Ge Jian Zhang Yangang Wang Chunbo Zhou and Yang Guangwen. 2016. The Sunway TaihuLight supercomputer: system and applications. Science China Information Sciences 59 7 (2016) 1\u201316.","DOI":"10.1007\/s11432-016-5588-7"},{"key":"e_1_3_2_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA45697.2020.00089"},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.1145\/2830772.2830812"},{"key":"e_1_3_2_29_2","doi-asserted-by":"publisher","DOI":"10.5555\/2385452"},{"key":"e_1_3_2_30_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2006.1598133"},{"key":"e_1_3_2_31_2","doi-asserted-by":"publisher","DOI":"10.1145\/1542275.1542349"},{"key":"e_1_3_2_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2005.9"},{"key":"e_1_3_2_33_2","doi-asserted-by":"publisher","DOI":"10.14778\/2856318.2856321"},{"key":"e_1_3_2_34_2","volume-title":"USENIX Annual Technical Conference","author":"Koh S.","year":"2018","unstructured":"S. Koh, C. Lee, M. Kwon, and M. Jung. 2018. Exploring system challenges of ultra-low latency solid state drives. In USENIX Annual Technical Conference."},{"key":"e_1_3_2_35_2","doi-asserted-by":"publisher","DOI":"10.1145\/3599691.3603406"},{"key":"e_1_3_2_36_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISPASS.2013.6557176"},{"key":"e_1_3_2_37_2","doi-asserted-by":"publisher","unstructured":"Andres Lagar-Cavilla Junwhan Ahn Suleiman Souhlal Neha Agarwal Radoslaw Burny Shakeel Butt Jichuan Chang Ashwin Chaugule Nan Deng Junaid Shahid Greg Thelen Kamil Adam Yurtsever Yu Zhao and Parthasarathy Ranganathan. 2019. Software-defined far memory in warehouse-scale computers. In Proceedings of the Twenty-Fourth International Conference on Architectural Support for Programming Languages and Operating Systems (Providence RI USA) (ASPLOS\u201919). Association for Computing Machinery New York NY USA 317\u2013330. 10.1145\/3297858.330405","DOI":"10.1145\/3297858.330405"},{"key":"e_1_3_2_38_2","doi-asserted-by":"publisher","DOI":"10.1145\/2133382.2133384"},{"key":"e_1_3_2_39_2","doi-asserted-by":"publisher","DOI":"10.1109\/IEDM.2016.7838026"},{"key":"e_1_3_2_40_2","doi-asserted-by":"publisher","DOI":"10.1145\/3445814.3446717"},{"key":"e_1_3_2_41_2","doi-asserted-by":"publisher","unstructured":"Huaicheng Li Daniel S. Berger Stanko Novakovic Lisa Hsu Dan Ernst Pantea Zardoshti Monish Shah Samir Rajadnya Scott Lee Ishwar Agarwal Mark D. Hill Marcus Fontoura and Ricardo Bianchini. 2022. Pond: CXL-Based Memory Pooling Systems for Cloud Platforms. (Oct.2022). DOI:10.48550\/arXiv.2203.00241arXiv:2203.00241 [cs].","DOI":"10.48550\/arXiv.2203.00241"},{"key":"e_1_3_2_42_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA56546.2023.10070986"},{"key":"e_1_3_2_43_2","doi-asserted-by":"publisher","DOI":"10.1145\/1669112.1669172"},{"key":"e_1_3_2_44_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISPASS.2018.00028"},{"key":"e_1_3_2_45_2","doi-asserted-by":"publisher","DOI":"10.1145\/3503222.3507745"},{"key":"e_1_3_2_46_2","doi-asserted-by":"publisher","DOI":"10.1109\/TC.2020.2973134"},{"key":"e_1_3_2_47_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11390-020-0780-z"},{"key":"e_1_3_2_48_2","unstructured":"Jason Lowe-Power Abdul Mutaal Ahmad Ayaz Akram Mohammad Alian Rico Amslinger Matteo Andreozzi Adri\u00e0 Armejach Nils Asmussen Brad Beckmann Srikant Bharadwaj Gabe Black Gedare Bloom Bobby R. Bruce Daniel Rodrigues Carvalho Jeronimo Castrillon Lizhong Chen Nicolas Derumigny Stephan Diestelhorst Wendy Elsasser Carlos Escuin Marjan Fariborz Amin Farmahini-Farahani Pouya Fotouhi Ryan Gambord Jayneel Gandhi Dibakar Gope Thomas Grass Anthony Gutierrez Bagus Hanindhito Andreas Hansson Swapnil Haria Austin Harris Timothy Hayes Adrian Herrera Matthew Horsnell Syed Ali Raza Jafri Radhika Jagtap Hanhwi Jang Reiley Jeyapaul Timothy M. Jones Matthias Jung Subash Kannoth Hamidreza Khaleghzadeh Yuetsu Kodama Tushar Krishna Tommaso Marinelli Christian Menard Andrea Mondelli Miquel Moreto Tiago M\u00fcck Omar Naji Krishnendra Nathella Hoa Nguyen Nikos Nikoleris Lena E. Olson Marc Orr Binh Pham Pablo Prieto Trivikram Reddy Alec Roelke Mahyar Samani Andreas Sandberg Javier Setoain Boris Shingarov Matthew D. Sinclair Tuan Ta Rahul Thakur Giacomo Travaglini Michael Upton Nilay Vaish Ilias Vougioukas William Wang Zhengrong Wang Norbert Wehn Christian Weis David A. Wood Hongil Yoon and \u00c9der F. Zulian. 2020. The gem5 Simulator: Version 20.0+. arXiv:2007.0315."},{"key":"e_1_3_2_49_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2016.7446087"},{"key":"e_1_3_2_50_2","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2014.117"},{"key":"e_1_3_2_51_2","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2003.1261383"},{"key":"e_1_3_2_52_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA52012.2021.00024"},{"key":"e_1_3_2_53_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA47549.2020.00040"},{"key":"e_1_3_2_54_2","volume-title":"2023 56rd Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO\u201923)","author":"Naithani Ajeya","year":"2023","unstructured":"Ajeya Naithani, Jaime Roelandts, Sam Ainsworth, Timothy M. Jones, and Lieven Eeckhout. 2023. Decoupled vector runahead. In 2023 56rd Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO\u201923)."},{"key":"e_1_3_2_55_2","unstructured":"Jacob Nelson Brandon Holt Brandon Myers Preston Briggs Luis Ceze Simon Kahan and Mark Oskin. 2015. {Latency-Tolerant} software distributed shared memory. 291\u2013305. Retrieved fromhttps:\/\/www.usenix.org\/conference\/atc15\/technical-session\/presentation\/nelson"},{"key":"e_1_3_2_56_2","doi-asserted-by":"publisher","DOI":"10.1145\/3470496.3527400"},{"key":"e_1_3_2_57_2","volume-title":"Design and programming of a coprocessor for a RISC-V architecture","author":"Pala Davide","year":"2016","unstructured":"Davide Pala. 2016-2017. Design and programming of a coprocessor for a RISC-V architecture. Master\u2019s Thesis. POLITECNICO DI TORINO."},{"key":"e_1_3_2_58_2","first-page":"501","volume-title":"2021 Design, Automation & Test in Europe Conference & Exhibition (DATE\u201921)","author":"Pan Haiyang","year":"2021","unstructured":"Haiyang Pan, Yuhang Liu, Tianyue Lu, and Mingyu Chen. 2021. LSP: Collective cross-page prefetching for NVM. In 2021 Design, Automation & Test in Europe Conference & Exhibition (DATE\u201921). 501\u2013506. DOI:10.23919\/DATE51398.2021.9474127"},{"key":"e_1_3_2_59_2","doi-asserted-by":"publisher","DOI":"10.1145\/1555815.1555760"},{"key":"e_1_3_2_60_2","doi-asserted-by":"publisher","DOI":"10.5555\/2971808.2972024"},{"key":"e_1_3_2_61_2","doi-asserted-by":"publisher","DOI":"10.1109\/HOTI55740.2022.00017"},{"key":"e_1_3_2_62_2","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2004.71"},{"key":"e_1_3_2_63_2","doi-asserted-by":"publisher","DOI":"10.1109\/CGO.2017.7863738"},{"key":"e_1_3_2_64_2","doi-asserted-by":"publisher","DOI":"10.1145\/3192366.3192393"},{"key":"e_1_3_2_65_2","doi-asserted-by":"crossref","first-page":"409","DOI":"10.1109\/MICRO.2006.44","volume-title":"2006 39th Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO\u201906)","author":"Tuck James","year":"2006","unstructured":"James Tuck, Luis Ceze, and Josep Torrellas. 2006. Scalable cache miss handling for high memory-level parallelism. In 2006 39th Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO\u201906). IEEE, 409\u2013422."},{"key":"e_1_3_2_66_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.tbench.2022.100061"},{"key":"e_1_3_2_67_2","unstructured":"Songyue Wang Luming Wang Tianyue Lu and Mingyu Chen. 2022. Architecture and RISC-V ISA extension supporting asynchronous and flexible parallel far memory access. In Sixth Workshop on Computer Architecture Research with RISC-V (CARRV\u201922)."},{"key":"e_1_3_2_68_2","doi-asserted-by":"publisher","DOI":"10.1109\/DAC18074.2021.9586186"},{"key":"e_1_3_2_69_2","doi-asserted-by":"publisher","DOI":"10.1145\/3494536"},{"key":"e_1_3_2_70_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCAD.2020.3012213"},{"key":"e_1_3_2_71_2","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO56248.2022.00080"},{"key":"e_1_3_2_72_2","first-page":"178","volume-title":"Proceedings of the 48th International Symposium on Microarchitecture (MICRO-48)","author":"Yu Xiangyao","year":"2015","unstructured":"Xiangyao Yu, Christopher J. Hughes, Nadathur Satish, and Srinivas Devadas. 2015. IMP: Indirect memory prefetcher. In Proceedings of the 48th International Symposium on Microarchitecture (MICRO-48). ACM, New York, NY, 178\u2013190."},{"key":"e_1_3_2_73_2","doi-asserted-by":"publisher","DOI":"10.1109\/BigData.2017.8257937"},{"key":"e_1_3_2_74_2","doi-asserted-by":"publisher","DOI":"10.1145\/1250662.1250668"},{"key":"e_1_3_2_75_2","doi-asserted-by":"publisher","unstructured":"Li-Cheng Chen Ming-Yu Chen Yuan Ruan Yong-Bing Huang Ze-Han Cui Tian-Yue Lu and Yun-Gang Bao. 2014. MIMS: Towards a message interface based memory system. Journal of Computer Science and Technology 29 2 (March 2014) 255\u2013272. 10.1007\/s11390-014-1428-7","DOI":"10.1007\/s11390-014-1428-7"}],"container-title":["ACM Transactions on Architecture and Code Optimization"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3663479","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3663479","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T23:56:46Z","timestamp":1750291006000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3663479"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,9,14]]},"references-count":74,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2024,9,30]]}},"alternative-id":["10.1145\/3663479"],"URL":"https:\/\/doi.org\/10.1145\/3663479","relation":{},"ISSN":["1544-3566","1544-3973"],"issn-type":[{"value":"1544-3566","type":"print"},{"value":"1544-3973","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,9,14]]},"assertion":[{"value":"2023-12-04","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-04-15","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-09-14","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}