{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,21]],"date-time":"2026-05-21T22:28:29Z","timestamp":1779402509953,"version":"3.53.1"},"reference-count":177,"publisher":"Association for Computing Machinery (ACM)","issue":"12","license":[{"start":{"date-parts":[[2026,5,15]],"date-time":"2026-05-15T00:00:00Z","timestamp":1778803200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"funder":[{"DOI":"10.13039\/501100012166","name":"National Key R&D Program of China","doi-asserted-by":"crossref","award":["2024YFB4504300"],"award-info":[{"award-number":["2024YFB4504300"]}],"id":[{"id":"10.13039\/501100012166","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["U24A20234"],"award-info":[{"award-number":["U24A20234"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Comput. Surv."],"published-print":{"date-parts":[[2026,9,30]]},"abstract":"<jats:p>The growing scale of data requires efficient memory subsystems with large memory capacity and high memory performance. Disaggregated architecture has become a promising solution for today\u2019s cloud and edge computing for its scalability and elasticity. As a critical part of disaggregation, disaggregated memory faces many design challenges in different dimensions, including hardware scalability, architecture structure, software system design, application programmability, resource allocation, power management, and so on. These challenges inspire a number of novel solutions at different system levels to improve overall efficiency. In this article, we provide a comprehensive review of disaggregated memory, including the methodology and technologies of disaggregated memory system foundation, optimization, and management. We study the technical essentials of disaggregated memory systems and analyze them from a bottom-up perspective, covering hardware, architecture, system, and application levels. Then, we compare the design details of typical cross-layer designs on disaggregated memory. Finally, we discuss the challenges and opportunities of future disaggregated memory works that serve better for next-generation elastic and efficient datacenters.<\/jats:p>","DOI":"10.1145\/3807443","type":"journal-article","created":{"date-parts":[[2026,4,7]],"date-time":"2026-04-07T11:35:45Z","timestamp":1775561745000},"page":"1-38","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["Survey of Disaggregated Memory: Cross-Layer Technique Insights for Next-Generation Datacenters"],"prefix":"10.1145","volume":"58","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-7260-0521","authenticated-orcid":false,"given":"Jing","family":"Wang","sequence":"first","affiliation":[{"name":"East China Normal University","place":["Shanghai, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6218-4659","authenticated-orcid":false,"given":"Chao","family":"Li","sequence":"additional","affiliation":[{"name":"Shanghai Jiao Tong University","place":["Shanghai, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0166-1447","authenticated-orcid":false,"given":"Taolei","family":"Wang","sequence":"additional","affiliation":[{"name":"Shanghai Jiao Tong University","place":["Shanghai, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4405-0424","authenticated-orcid":false,"given":"Jinyang","family":"Guo","sequence":"additional","affiliation":[{"name":"Shanghai Jiao Tong University","place":["Shanghai, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-6097-6952","authenticated-orcid":false,"given":"Hanzhang","family":"Yang","sequence":"additional","affiliation":[{"name":"University of Toronto","place":["Toronto, Canada"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0009-9057-6452","authenticated-orcid":false,"given":"Yiming","family":"Zhuansun","sequence":"additional","affiliation":[{"name":"Shanghai Jiao Tong University","place":["Shanghai, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0034-2302","authenticated-orcid":false,"given":"Minyi","family":"Guo","sequence":"additional","affiliation":[{"name":"Shanghai Jiao Tong University","place":["Shanghai, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,5,15]]},"reference":[{"key":"e_1_3_2_2_2","unstructured":"2025. Compute Express Link. Retrieved April 2026 from https:\/\/www.computeexpresslink.org\/"},{"key":"e_1_3_2_3_2","unstructured":"2025. Control Group v2. Retrieved April 2026 from https:\/\/www.kernel.org\/doc\/Documentation\/cgroup-v2.txt"},{"key":"e_1_3_2_4_2","unstructured":"2025. High-performance remote procedure call framework. Retrieved April 2026 from https:\/\/github.com\/grpc"},{"key":"e_1_3_2_5_2","unstructured":"2025. Instruction-Based Sampling: A New Performance Analysis Technique for AMD Family 10h Processors. Retrieved April 2026 from http:\/\/developer.amd.com\/wordpress\/media\/2012\/10\/AMD_IBS_paper_EN.pdf"},{"key":"e_1_3_2_6_2","unstructured":"2025. KVM (for Kernel-based Virtual Machine). Retrieved April 2026 from https:\/\/www.linux-kvm.org\/page\/Main_Page"},{"key":"e_1_3_2_7_2","unstructured":"2025. MGLRU. Retrieved April 2026 from https:\/\/lore.kernel.org\/all\/20230714194610.828210-1-hannes@cmpxchg.org\/"},{"key":"e_1_3_2_8_2","unstructured":"2025. Nvidia DPU. Retrieved April 2026 from https:\/\/www.nvidia.cn\/networking\/products\/data-processing-unit"},{"key":"e_1_3_2_9_2","unstructured":"2025. NVM-Express (NVMe) Technology. Retrieved April 2026 from https:\/\/nvmexpress.org\/"},{"key":"e_1_3_2_10_2","unstructured":"2025. OFED from OpenFabrics. Retrieved April 2026 from https:\/\/network.nvidia.com\/products\/infiniband-drivers\/linux\/mlnx_ofed\/"},{"key":"e_1_3_2_11_2","unstructured":"2025. OpenCAPI Specification. Retrieved April 2026 from https:\/\/opencapi.org\/"},{"key":"e_1_3_2_12_2","unstructured":"2025. PCIe 6.0 is coming. Retrieved April 2026 from https:\/\/www.theverge.com\/2022\/1\/12\/22879732\/pcie-6-0-final-specification-bandwidth-speeds"},{"key":"e_1_3_2_13_2","unstructured":"2025. QEMU. Retrieved April 2026 from https:\/\/github.com\/qemu\/qemu"},{"key":"e_1_3_2_14_2","unstructured":"2025. Sumsung Solid State Drives (SSD). Retrieved April 2026 from https:\/\/www.samsung.com\/us\/computing\/memory-storage\/solid-state-drives"},{"key":"e_1_3_2_15_2","unstructured":"2025. VMware. Retrieved April 2026 from https:\/\/www.vmware.com"},{"key":"e_1_3_2_16_2","unstructured":"Marcos K. Aguilera Nadav Amit Irina Calciu Xavier Deguillard Jayneel Gandhi Stanko Novakovic Arun Ramanathan Pratap Subrahmanyam Lalith Suresh et\u00a0al. 2018. Remote regions: A simple abstraction for remote memory. In USENIX Annual Technical Conference (USENIX ATC). 1\u201314."},{"key":"e_1_3_2_17_2","doi-asserted-by":"crossref","unstructured":"Marcos K. Aguilera Kimberly Keeton Stanko Novakovic and Sharad Singhal. 2019. Designing far memory data structures: Think outside the box. In Workshop on Hot Topics in Operating Systems (HotOS). 120\u2013126.","DOI":"10.1145\/3317550.3321433"},{"key":"e_1_3_2_18_2","volume-title":"Proceedings of the USENIX Conference on Usenix Annual Technical Conference (USENIX ATC)","author":"Maruf Hasan Al","year":"2020","unstructured":"Hasan Al Maruf and Mosharaf Chowdhury. 2020. Effectively prefetching remote memory with leap. In Proceedings of the USENIX Conference on Usenix Annual Technical Conference (USENIX ATC). 843\u2013857."},{"key":"e_1_3_2_19_2","doi-asserted-by":"crossref","unstructured":"Emmanuel Amaro Christopher Branner-Augmon Zhihong Luo Amy Ousterhout Marcos K. Aguilera Aurojit Panda Sylvia Ratnasamy and Scott Shenker. 2020. Can far memory improve job throughput? In European Conference on Computer Systems (Eurosys). 1\u201315.","DOI":"10.1145\/3342195.3387522"},{"key":"e_1_3_2_20_2","doi-asserted-by":"crossref","unstructured":"Nadav Amit Dan Tsafrir and Assaf Schuster. 2014. Vswapper: A memory swapper for virtualized environments. Sigplan Notices (2014).","DOI":"10.1145\/2541940.2541969"},{"key":"e_1_3_2_21_2","unstructured":"Thomas E. Anderson Marco Canini Jongyul Kim Dejan Kosti\u0107 Youngjin Kwon Simon Peter Waleed Reda Henry N. Schuh and Emmett Witchel. 2020. Assise: Performance and availability via client-local nvm in a distributed file system. In USENIX Symposium on Operating Systems Design and Implementation (OSDI). 1\u201318."},{"key":"e_1_3_2_22_2","doi-asserted-by":"crossref","unstructured":"Trinayan Baruah Yifan Sun Ali Tolga Din\u00e7er Saiful A. Mojumder Jos\u00e9 L. Abell\u00e1n Yash Ukidave Ajay Joshi Norman Rubin John Kim and David Kaeli. 2020. Griffin: Hardware-software support for efficient page migration in multi-GPU systems. In IEEE International Symposium on High Performance Computer Architecture (HPCA). 596\u2013609.","DOI":"10.1109\/HPCA47549.2020.00055"},{"key":"e_1_3_2_23_2","doi-asserted-by":"crossref","unstructured":"Soroush Bateni Zhendong Wang et\u00a0al. 2020. Co-optimizing performance and memory footprint via integrated CPU\/GPU memory management. In Proceedings of the Real-Time and Embedded Technology and Applications Symposium (RTAS). 1\u201314.","DOI":"10.1109\/RTAS48715.2020.00007"},{"key":"e_1_3_2_24_2","doi-asserted-by":"crossref","unstructured":"Maciej Bielski Christian Pinto Daniel Raho and Renaud Pacalet. 2016. Survey on memory and devices disaggregation solutions for HPC systems. In IEEE Intl Conference on Computational Science and Engineering (CSE) and IEEE Intl Conference on Embedded and Ubiquitous Computing (EUC) and Intl Symposium on Distributed Computing and Applications for Business Engineering (DCABES). 197\u2013204.","DOI":"10.1109\/CSE-EUC-DCABES.2016.185"},{"key":"e_1_3_2_25_2","doi-asserted-by":"crossref","unstructured":"Rajarshi Biswas Xiaoyi Lu and Dhabaleswar K. Panda. 2018. Accelerating tensorflow with adaptive RDMA-based gRPC. In IEEE International Conference on High Performance Computing (HiPC). 2\u201311.","DOI":"10.1109\/HiPC.2018.00010"},{"key":"e_1_3_2_26_2","doi-asserted-by":"crossref","unstructured":"Andreas Blenk Arsany Basta Martin Reisslein and Wolfgang Kellerer. 2015. Survey on network virtualization hypervisors for software defined networking. IEEE Communications Surveys & Tutorials (2015).","DOI":"10.1109\/COMST.2015.2489183"},{"key":"e_1_3_2_27_2","unstructured":"Eric Boutin Jaliya Ekanayake Wei Lin Bing Shi Jingren Zhou Zhengping Qian Ming Wu and Lidong Zhou. 2014. Apollo: Scalable and coordinated scheduling for cloud-scale computing. In USENIX Symposium on Operating Systems Design and Implementation (OSDI). 285\u2013300."},{"key":"e_1_3_2_28_2","doi-asserted-by":"crossref","unstructured":"Irina Calciu M. Talha Imran Ivan Puddu Sanidhya Kashyap Hasan Al Maruf Onur Mutlu and Aasheesh Kolli. 2021. Rethinking software runtimes for disaggregated memory. In Architectural Support for Programming Languages and Operating Systems (ASPLOS). 79\u201392.","DOI":"10.1145\/3445814.3446713"},{"key":"e_1_3_2_29_2","unstructured":"Wenqi Cao. 2024. FastSwap. Retrieved from https:\/\/github.com\/git-disl\/FastSwap"},{"key":"e_1_3_2_30_2","doi-asserted-by":"publisher","DOI":"10.1109\/TC.2020.2968525"},{"key":"e_1_3_2_31_2","doi-asserted-by":"crossref","unstructured":"Yuhai Cao Chao Li Quan Chen Jingwen Leng Minyi Guo Jing Wang and Weigong Zhang. 2018. DR DRAM: Accelerating memory-readintensive applications. In International Conference on Computer Design (ICCD). 301\u2013309.","DOI":"10.1109\/ICCD.2018.00053"},{"key":"e_1_3_2_32_2","doi-asserted-by":"crossref","unstructured":"Chia-Hao Chang Jihoon Han Anand Sivasubramaniam Vikram Sharma Mailthody Zaid Qureshi and Wen-Mei Hwu. 2024. GMT: GPU orchestrated memory tiering for the big data era. In Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems 3 (2024) 464\u2013478.","DOI":"10.1145\/3620666.3651353"},{"key":"e_1_3_2_33_2","doi-asserted-by":"crossref","unstructured":"Rong Chen Jiaxin Shi Yanzhe Chen Binyu Zang Haibing Guan and Haibo Chen. 2019. Powerlyra: Differentiated graph computation and partitioning on skewed graphs. ACM Transactions on Parallel Computing (TPC) 5 3 (2019) 1\u201339.","DOI":"10.1145\/3298989"},{"key":"e_1_3_2_34_2","doi-asserted-by":"crossref","unstructured":"Tingting Chen Haikun Liu Xiaofei Liao and Hai Jin. 2021. Resource abstraction and data placement for distributed hybrid memory pool. Frontiers of Computer Science (FCS) 15 1 (2021).","DOI":"10.1007\/s11704-020-9448-7"},{"key":"e_1_3_2_35_2","doi-asserted-by":"crossref","unstructured":"Esha Choukse Michael B. Sullivan Mike O\u2019Connor Mattan Erez Jeff Pool David Nellans and StephenW. Keckler. 2020. Buddy compression: Enabling larger memory for deep learning and hpc workloads on GPUs. In International Symposium on Computer Architecture (ISCA). 926\u2013939.","DOI":"10.1109\/ISCA45697.2020.00080"},{"key":"e_1_3_2_36_2","unstructured":"CCIX Consortium. 2023. Cache Coherent Interconnect for Accelerators. Retrieved 2026 from https:\/\/genzconsortium.org"},{"key":"e_1_3_2_37_2","doi-asserted-by":"crossref","unstructured":"Eli Cortez Anand Bonde Alexandre Muzio Mark Russinovich Marcus Fontoura and Ricardo Bianchini. 2017. Resource central: Understanding and predicting workloads for improved resource management in large cloud platforms. In Symposium on Operating Systems Principles (SOSP). 153\u2013167.","DOI":"10.1145\/3132747.3132772"},{"key":"e_1_3_2_38_2","doi-asserted-by":"crossref","unstructured":"Ala Darabseh Mahmoud Al-Ayyoub Yaser Jararweh Elhadj Benkhelifa Mladen Vouk and Andy Rindos. 2015. Sdstorage: A software defined storage experimental framework. In 2015 IEEE International Conference on Cloud Engineering. 341\u2013346.","DOI":"10.1109\/IC2E.2015.60"},{"key":"e_1_3_2_39_2","doi-asserted-by":"publisher","DOI":"10.1109\/TSE.1980.230464"},{"key":"e_1_3_2_40_2","unstructured":"Data Dog. 2023. The state of Serverless. Retrieved April 2026 from https:\/\/www.datadoghq.com\/state-of-serverless\/"},{"key":"e_1_3_2_41_2","doi-asserted-by":"crossref","unstructured":"Thaleia Dimitra Doudali Sergey Blagodurov Abhinav Vishnu Sudhanva Gurumurthi and Ada Gavrilovska. 2019. Kleio: A hybrid memory page scheduler with machine intelligence. In International Symposium on High-Performance Parallel and Distributed Computing. 37\u201348.","DOI":"10.1145\/3307681.3325398"},{"key":"e_1_3_2_42_2","unstructured":"Aleksandar Dragojevic Dushyanth Narayanan Miguel Castro and Orion Hodson. 2014. FaRM: Fast remote memory. In USENIX Symposium on Networked Systems Design and Implementation (NSDI). 401\u2013414."},{"key":"e_1_3_2_43_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2023.3250407"},{"key":"e_1_3_2_44_2","unstructured":"FFmpeg. 2013. A complete cross-platform solution to record convert and stream audio and video. Retrieved 2013 fromhttp:\/\/ffmpeg.org"},{"key":"e_1_3_2_45_2","unstructured":"MPI Forum.2025. Working set. Retrieved April 2026 from https:\/\/en.wikipedia.org\/wiki\/Working_set"},{"key":"e_1_3_2_46_2","unstructured":"Peter X. Gao Akshay Narayan Sagar Karandikar Joao Carreira Sangjin Han Rachit Agarwal Sylvia Ratnasamy and Scott Shenker. 2016. Network requirements for resource disaggregation. In USENIX Symposium on Operating Systems Design and Implementation (OSDI). 249\u2013264."},{"key":"e_1_3_2_47_2","doi-asserted-by":"crossref","unstructured":"Jorge Gonzalez Alexander Gazman Maarten Hattink Mauricio G. Palma Meisam Bahadori Ruth Rubio-Noriega Lois Orosa Madeleine Glick Onur Mutlu Keren Bergman and Rodolfo Azevedo. 2022. Optically connected memory for disaggregated data centers. 43\u201350.","DOI":"10.1109\/SBAC-PAD49847.2020.00017"},{"key":"e_1_3_2_48_2","unstructured":"Donghyun Gouk Sangwon Lee Miryeong Kwon and Myoungsoo Jung. 2022. Direct access High-Performance memory disaggregation with DirectCXL. In USENIX Annual Technical Conference (USENIX ATC). 287\u2013294."},{"key":"e_1_3_2_49_2","unstructured":"Robert Grandl Srikanth Kandula Sriram Rao Aditya Akella and Janardhan Kulkarni. 2016. GRAPHENE: Packing and dependency-aware scheduling for data-parallel clusters. In USENIX Symposium on Operating Systems Design and Implementation (OSDI). 81\u201397."},{"key":"e_1_3_2_50_2","unstructured":"Juncheng Gu Youngmoon Lee Yiwen Zhang Mosharaf Chowdhury and Kang G. Shin. 2017. Efficient memory disaggregation with infiniswap. In USENIX Symposium on Networked Systems Design and Implementation (NSDI). 649\u2013667."},{"key":"e_1_3_2_51_2","doi-asserted-by":"crossref","unstructured":"Zhiyuan Guo Yizhou Shan Xuhao Luo Yutong Huang and Yiying Zhang. 2022. Clio: A hardware-software co-designed disaggregated memory system. In ACM International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS). 417\u2013433.","DOI":"10.1145\/3503222.3507762"},{"key":"e_1_3_2_52_2","doi-asserted-by":"crossref","unstructured":"Cunchen Hu Chenxi Wang Sa Wang Ninghui Sun Yungang Bao Jieru Zhao Sanidhya Kashyap Pengfei Zuo Xusheng Chen Liangliang Xu et\u00a0al. 2023. Skadi: Building a distributed runtime for data systems in disaggregated data centers. In Workshop on Hot Topics in Operating Systems (HotOS). 94\u2013102.","DOI":"10.1145\/3593856.3595897"},{"key":"e_1_3_2_53_2","doi-asserted-by":"publisher","DOI":"10.1145\/3373376.3378530"},{"key":"e_1_3_2_54_2","doi-asserted-by":"crossref","unstructured":"Wenqin Huangfu Krishna T. Malladi Andrew Chang and Yuan Xie. 2022. BEACON: Scalable near-data-processing accelerators for genome analysis near memory pool with the CXL support. In IEEE\/ACM International Symposium on Microarchitecture (MICRO). 727\u2013743.","DOI":"10.1109\/MICRO56248.2022.00057"},{"key":"e_1_3_2_55_2","unstructured":"IBM. 2025. IBM Power 9 CPU. Retrieved April 2026 from https:\/\/www.ibm.com\/it-infrastructure\/power\/power9"},{"key":"e_1_3_2_56_2","unstructured":"Intel. 2025. Intel Performance Counter Monitor A Better Way to Measure CPU Utilization. Retrieved April 2026 from https:\/\/www.intel.com\/content\/www\/us\/en\/developer\/articles\/technical\/performance-counter-monitor.html"},{"key":"e_1_3_2_57_2","volume-title":"Memory Systems: Cache, DRAM, Disk","author":"Jacob Bruce","year":"2010","unstructured":"Bruce Jacob, David Wang, and Spencer Ng. 2010. Memory Systems: Cache, DRAM, Disk. Morgan Kaufmann."},{"key":"e_1_3_2_58_2","unstructured":"Junhyeok Jang Hanjin Choi Hanyeoreum Bae Seungjun Lee Miryeong Kwon and Myoungsoo Jung. 2023. CXL-ANNS: Software-Hardware collaborative memory disaggregation and computation for Billion-Scale approximate nearest neighbor search. In USENIX Annual Technical Conference (USENIX ATC). 585\u2013600."},{"key":"e_1_3_2_59_2","doi-asserted-by":"crossref","unstructured":"Tatiana Jin Zhenkun Cai Boyang Li Chengguang Zheng Guanxian Jiang and James Cheng. 2020. Improving resource utilization by timely fine-grained scheduling. In European Conference on Computer Systems (EuroSys). 1\u201316.","DOI":"10.1145\/3342195.3387551"},{"key":"e_1_3_2_60_2","doi-asserted-by":"crossref","unstructured":"Junyi Mei Shixuan Sun Chao Li Cheng Xu Cheng Chen Yibo Liu JingWang Cheng Zhao Xiaofeng Hou Minyi Guo et\u00a0al. 2024. FlowWalker: A memory-efficient and high-performance GPU-based dynamic graph random walk framework. In Proceedings of the VLDB Endow. 1788\u20131801.","DOI":"10.14778\/3659437.3659438"},{"key":"e_1_3_2_61_2","doi-asserted-by":"crossref","unstructured":"G. Kandiraju H. Franke M. D. Williams M. Steinder and S. M. Black. 2014. Software defined infrastructures. 58 (2014) 2:1\u20132:13.","DOI":"10.1147\/JRD.2014.2298133"},{"key":"e_1_3_2_62_2","unstructured":"Hiwot Tadese Kassa Jason Akers Mrinmoy Ghosh Zhichao Cao Vaibhav Gogte and Ronald Dreslinski. 2021. Improving performance of flash based key-value stores using storage class memory as a volatile memory extension. In USENIX Annual Technical Conference (USENIX ATC). 821\u2013837."},{"key":"e_1_3_2_63_2","doi-asserted-by":"crossref","unstructured":"Anurag Khandelwal Yupeng Tang Rachit Agarwal Aditya Akella and Ion Stoica. Jiffy: Elastic far-memory for stateful serverless analytics. In European Conference on Computer Systems (EuroSys). 697\u2013713.","DOI":"10.1145\/3492321.3527539"},{"key":"e_1_3_2_64_2","unstructured":"Daehyeok Kim Tianlong Yu Hongqiang Harry Liu Yibo Zhu Jitu Padhye Shachar Raindel Chuanxiong Guo Vyas Sekar and Srinivasan Seshan. 2019. FreeFlow: Software-based virtual RDMA networking for containerized clouds. In USENIX Symposium on Networked Systems Design and Implementation (NSDI). 113\u2013125."},{"key":"e_1_3_2_65_2","doi-asserted-by":"crossref","unstructured":"Gwangsun Kim John Kim Jung Ho Ahn and Jaeha Kim. 2013. Memory-centric system interconnect design with hybrid memory cubes. In International conference on Parallel architectures and compilation techniques (PACT). 145\u2013155.","DOI":"10.1109\/PACT.2013.6618812"},{"key":"e_1_3_2_66_2","doi-asserted-by":"crossref","unstructured":"Atsushi Koshiba Felix Gust Julian Pritzi Anjo Vahldiek-Oberwagner Nuno Santos and Pramod Bhatotia. 2023. Trusted heterogeneous disaggregated architectures. In ACM SIGOPS Asia-Pacific Workshop on Systems (APSys). 72\u201379.","DOI":"10.1145\/3609510.3609812"},{"key":"e_1_3_2_67_2","doi-asserted-by":"crossref","unstructured":"Anthony Kougkas Hariharan Devarajan and Xian-He Sun. 2018. Hermes: A heterogeneous-aware multi-tiered distributed I\/O buffering system. In International Symposium on High-Performance Parallel and Distributed Computing (HPDC). 219\u2013230.","DOI":"10.1145\/3208040.3208059"},{"key":"e_1_3_2_68_2","doi-asserted-by":"crossref","unstructured":"Panos Koutsovasilis Michele Gazzetti and Christian Pinto. 2021. A holistic system software integration of disaggregated memory for next-generation cloud infrastructures. In IEEE\/ACM International Symposium on Cluster Cloud and Internet Computing (CCGrid). 576\u2013585.","DOI":"10.1109\/CCGrid51090.2021.00067"},{"key":"e_1_3_2_69_2","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2018.00021"},{"key":"e_1_3_2_70_2","doi-asserted-by":"crossref","unstructured":"Andres Lagar-Cavilla Junwhan Ahn Suleiman Souhlal Neha Agarwal Radoslaw Burny Shakeel Butt Jichuan Chang Ashwin Chaugule Nan Deng Junaid Shahid et\u00a0al. 2019. Software-defined far memory in warehouse-scale computers. In Architectural Support for Programming Languages and Operating Systems (ASPLOS). 317\u2013330.","DOI":"10.1145\/3297858.3304053"},{"key":"e_1_3_2_71_2","unstructured":"Christopher Lameter. 2024. Flavors of Memory supported by Linux Their Use and Benefit. Retrieved April 2026 from https:\/\/events.linuxfoundation.org\/"},{"key":"e_1_3_2_72_2","doi-asserted-by":"crossref","unstructured":"Jong Chern Lee Jihwan Kim Kyung Whan Kim Young Jun Ku Dae Suk Kim Chunseok Jeong Tae Sik Yun Hongjung Kim Ho Sung Cho Sangmuk Oh et\u00a0al. 2016. High bandwidth memory (HBM) with TSV technique. In International SoC Design Conference (ISOCC). 181\u2013182.","DOI":"10.1109\/ISOCC.2016.7799847"},{"key":"e_1_3_2_73_2","doi-asserted-by":"crossref","unstructured":"Sekwon Lee Soujanya Ponnapalli Sharad Singhal Marcos K. Aguilera Kimberly Keeton and Vijay Chidambaram. 2022. DINOMO: An elastic scalable high-performance key-value store for disaggregated persistent memory. In Proceedings of the VLDB Endowment 15 13 (2022) 4023\u20134037.","DOI":"10.14778\/3565838.3565854"},{"key":"e_1_3_2_74_2","doi-asserted-by":"crossref","unstructured":"Seung-seob Lee Yanpeng Yu Yupeng Tang Anurag Khandelwal Lin Zhong and Abhishek Bhattacharjee. 2021. Mind: In-network memory management for disaggregated data centers. In Symposium on Operating Systems Principles (SOSP). 488\u2013504.","DOI":"10.1145\/3477132.3483561"},{"key":"e_1_3_2_75_2","unstructured":"Youngmoon Lee Hasan Al Maruf Mosharaf Chowdhury Asaf Cidon and Kang G. Shin. 2022. Hydra: Resilient and highly available remote memory. In Proceedings of the 20th USENIX Conference on File and Storage Technologies (FAST). 1\u201315."},{"key":"e_1_3_2_76_2","unstructured":"Youngmoon Lee Hassan Al Maruf Mosharaf Chowdhury and Kang G. Shin. 2019. Mitigating the performance-efficiency tradeoff in resilient memory disaggregation. Arxiv abs\/1910.09727. (2019) 1\u201315."},{"key":"e_1_3_2_77_2","doi-asserted-by":"crossref","unstructured":"Baolin Li Rohin Arora Siddharth Samsi Tirthak Patel William Arcand David Bestor Chansup Byun Rohan Basu Roy Bill Bergeron John Holodnak et\u00a0al. 2022. AI-enabling workloads on large-scale GPU-accelerated system: Characterization opportunities and implications. In IEEE International Symposium on High-Performance Computer Architecture (HPCA). 1224\u20131237.","DOI":"10.1109\/HPCA53966.2022.00093"},{"key":"e_1_3_2_78_2","doi-asserted-by":"crossref","unstructured":"Chao Li Yushu Xue Jing Wang Weigong Zhang and Tao Li. 2018. Edge-oriented computing paradigms: A survey on architecture design and system management. ACM Computer Survey 51 2 (2018) 1\u201334.","DOI":"10.1145\/3154815"},{"key":"e_1_3_2_79_2","doi-asserted-by":"publisher","DOI":"10.14778\/3554821.3554893"},{"key":"e_1_3_2_80_2","doi-asserted-by":"publisher","DOI":"10.1145\/3575693.3578835"},{"key":"e_1_3_2_81_2","doi-asserted-by":"crossref","unstructured":"Haifeng Li Ke Liu Ting Liang Zuojun Li Tianyue Lu Hui Yuan Yinben Xia Yungang Bao Mingyu Chen and Yizhou Shan. 2023. Hopp: Hardware-software co-designed page prefetching for disaggregated memory. In IEEE International Symposium on High-Performance Computer Architecture (HPCA). 1168\u20131181.","DOI":"10.1109\/HPCA56546.2023.10070986"},{"key":"e_1_3_2_82_2","doi-asserted-by":"publisher","DOI":"10.1145\/75104.75105"},{"key":"e_1_3_2_83_2","unstructured":"Pengfei Li Yu Hua Pengfei Zuo Zhangyu Chen and Jiajie Sheng. 2023. Rolex: A scalable RDMA-oriented learned key-value store for disaggregated memory systems. In USENIX Conference on File and Storage Technologies (FAST). 99\u2013114."},{"key":"e_1_3_2_84_2","doi-asserted-by":"crossref","unstructured":"Yang Li Saugata Ghose Jongmoo Choi Jin Sun Hui Wang and Onur Mutlu. 2017. Utility-based hybrid memory management. In IEEE International Conference on Cluster Computing (CLUSTER). 152\u2013165.","DOI":"10.1109\/CLUSTER.2017.130"},{"key":"e_1_3_2_85_2","doi-asserted-by":"publisher","DOI":"10.1145\/3342280.3342296"},{"key":"e_1_3_2_86_2","doi-asserted-by":"publisher","DOI":"10.1109\/CLUSTR.2005.347050"},{"key":"e_1_3_2_87_2","doi-asserted-by":"crossref","unstructured":"Kevin Lim Jichuan Chang Trevor Mudge Parthasarathy Ranganathan Steven K. Reinhardt and Thomas F. Wenisch. 2009. Disaggregated memory for expansion and sharing in blade servers. In International Symposium on Computer Architecture (ISCA). 267\u2013278.","DOI":"10.1145\/1555754.1555789"},{"key":"e_1_3_2_88_2","unstructured":"Linux. 2025. Model specific registers (MSRs). Retrieved April 2026 from https:\/\/man7.org\/linux\/man-pages\/man4\/msr.4.html"},{"key":"e_1_3_2_89_2","unstructured":"Linux. 2025. Zswap kernel. Retrieved April 2026 from https:\/\/www.kernel.org\/doc\/html\/latest\/admin-guide\/mm\/zswap.html"},{"key":"e_1_3_2_90_2","doi-asserted-by":"crossref","unstructured":"Haifeng Liu Long Zheng Yu Huang Jingyi Zhou Chaoqiang Liu RunzeWang Xiaofei Liao Hai Jin and Jingling Xue. 2024. Enabling efficient large recommendation model training with near CXL memory processing. In 2024 ACM\/IEEE 51st Annual International Symposium on Computer Architecture (ISCA). 382\u2013395.","DOI":"10.1109\/ISCA59077.2024.00036"},{"key":"e_1_3_2_91_2","doi-asserted-by":"crossref","unstructured":"Ling Liu Wenqi Cao Semih Sahin Qi Zhang Juhyun Bae and Yanzhao Wu. 2019. Memory disaggregation: Research problems and opportunities. In IEEE International Conference on Distributed Computing Systems (ICDCS). 1664\u20131673.","DOI":"10.1109\/ICDCS.2019.00165"},{"key":"e_1_3_2_92_2","doi-asserted-by":"crossref","unstructured":"Lei Liu Shengjie Yang Lu Peng and Xinyu Li. 2019. Hierarchical hybrid memory management in OS for tiered memory systems. IEEE Transactions on Parallel and Distributed Systems (TPDS) 30 10 (2019) 2223\u20132236.","DOI":"10.1109\/TPDS.2019.2908175"},{"key":"e_1_3_2_93_2","doi-asserted-by":"crossref","unstructured":"Baotong Lu Kaisong Huang Chieh-Jan Mike Liang Tianzheng Wang and Eric Lo. 2024. DEX: Scalable range indexing on disaggregated memory. In Proceedings of the VLDB Endow. 17 10 (June 2024) 2603\u20132616.","DOI":"10.14778\/3675034.3675050"},{"key":"e_1_3_2_94_2","unstructured":"Xuchuan Luo Pengfei Zuo Jiacheng Shen Jiazhen Gu Xin Wang Michael R. Lyu and Yangfan Zhou. 2023. SMART: A high-performance adaptive radix tree for disaggregated memory. In USENIX Symposium on Operating Systems Design and Implementation (OSDI). 553\u2013571."},{"key":"e_1_3_2_95_2","doi-asserted-by":"crossref","unstructured":"Steffen Maass Changwoo Min Sanidhya Kashyap Woonhak Kang Mohan Kumar and Taesoo Kim. 2017. Mosaic: Processing a trillion-edge graph on a single machine. In European Conference on Computer Systems (EuroSys). 527\u2013543.","DOI":"10.1145\/3064176.3064191"},{"key":"e_1_3_2_96_2","unstructured":"The Machine. 2025. The Machine: A new kind of computer. Retrieved April 2026 from https:\/\/www.labs.hpe.com\/the-machine"},{"key":"e_1_3_2_97_2","first-page":"414","volume-title":"Proceedings of the SC18: International Conference for High Performance Computing, Networking, Storage and Analysis","year":"2018","unstructured":"Pak Markthub, Mehmet E. Belviranli, Seyong Lee, Jeffrey S. Vetter, and Satoshi Matsuoka. 2018. DRAGON: Breaking GPU memory capacity limits with direct NVM access. In Proceedings of the SC18: International Conference for High Performance Computing, Networking, Storage and Analysis. IEEE, 414\u2013426."},{"key":"e_1_3_2_98_2","unstructured":"Hasan Al Maruf Yuhong Zhong et\u00a0al. 2021. Memtrade: A disaggregated-memory marketplace for public clouds. International Conference on Measurement and Modeling of Computer Systems (SIGMETRICS)."},{"key":"e_1_3_2_99_2","doi-asserted-by":"publisher","DOI":"10.1109\/MC.2022.3187847"},{"key":"e_1_3_2_100_2","unstructured":"Mellanox. [n.d.]. Mellanox Interconnect Community. Retrieved from https:\/\/forums.developer.nvidia.com\/c\/infrastructure\/369"},{"key":"e_1_3_2_101_2","doi-asserted-by":"crossref","unstructured":"Xupeng Miao Hailin Zhang Yining Shi Xiaonan Nie Zhi Yang Yangyu Tao and Bin Cui. 2021. HET: Scaling out huge embedding model training via cache-enabled distributed framework. (2021) 312\u2013320.","DOI":"10.14778\/3489496.3489511"},{"key":"e_1_3_2_102_2","unstructured":"Dave Montgomery. 2019. The Future of Data Infrastructure: CDI. Retrieved April 2026 from https:\/\/www.datacenterknowledge.com\/industry-perspectives\/future-data-infrastructure"},{"key":"e_1_3_2_103_2","doi-asserted-by":"crossref","unstructured":"Xiaonan Nie Yi Liu Fangcheng Fu Jinbao Xue Dian Jiao Xupeng Miao Yangyu Tao and Bin Cui. 2023. Angel-PTM: A scalable and economical large-scale pre-training system in tencent. 16 12 (2023) 3781\u20133794.","DOI":"10.14778\/3611540.3611564"},{"key":"e_1_3_2_104_2","doi-asserted-by":"crossref","unstructured":"Vlad Nitu Boris Teabe Alain Tchana Canturk Isci and Daniel Hagimont. 2018. Welcome to zombieland: Practical and energy-efficient memory disaggregation in a datacenter. In European Conference on Computer Systems (Eurosys). 1\u201312.","DOI":"10.1145\/3190508.3190537"},{"key":"e_1_3_2_105_2","unstructured":"NVIDIA. 2024. NvLink Interconnect. Retrieved April 2026 from http:\/\/www.nvidia.com\/object\/nvlink.html"},{"key":"e_1_3_2_106_2","unstructured":"NVIDIA. 2025. GPUDirect RDMA. Retrieved April 2026 from https:\/\/docs.nvidia.com\/cuda\/gpudirect-rdma\/index.html"},{"key":"e_1_3_2_107_2","unstructured":"NVIDIA. 2025. InfiniBand Architecture. Retrieved April 2026 from https:\/\/www.nvidia.com\/en-us\/networking\/infiniband-adapters\/"},{"key":"e_1_3_2_108_2","unstructured":"NVIDIA. 2025. Magnum IO GPUDirect Storage. Retrieved April 2026 from https:\/\/developer.nvidia.com\/gpudirect-storage"},{"key":"e_1_3_2_109_2","doi-asserted-by":"crossref","unstructured":"Matheus Ogleari Ye Yu Chen Qian Ethan Miller and Jishen Zhao. 2019. String figure: A scalable and elastic memory network architecture. In IEEE International Symposium on High Performance Computer Architecture (HPCA). 647\u2013660.","DOI":"10.1109\/HPCA.2019.00016"},{"key":"e_1_3_2_110_2","doi-asserted-by":"crossref","unstructured":"Chang Hyun Park Taekyung Heo Jungi Jeong and Jaehyuk Huh. 2017. Hybrid TLB coalescing: Improving TLB translation coverage under diverse fragmented memory allocations. In International Symposium on Computer Architecture (ISCA). 444\u2013456.","DOI":"10.1145\/3079856.3080217"},{"key":"e_1_3_2_111_2","doi-asserted-by":"publisher","DOI":"10.14778\/3648160.3648166"},{"key":"e_1_3_2_112_2","doi-asserted-by":"crossref","unstructured":"Adarsh Patil Vijay Nagarajan Nikos Nikoleris and Nicolai Oswald. 2023. Apta: Fault-tolerant object-granular CXL disaggregated memory for accelerating faas. In IEEE\/IFIP International Conference on Dependable Systems and Networks (DSN). 201\u2013215.","DOI":"10.1109\/DSN58367.2023.00030"},{"key":"e_1_3_2_113_2","doi-asserted-by":"crossref","unstructured":"Christian Pinto Dimitris Syrivelis Michele Gazzetti Panos Koutsovasilis Andrea Reale Kostas Katrinis and H. Peter Hofstee. 2020. Thymesisflow: A software-defined HW\/SW co-designed interconnect stack for rack-scale memory disaggregation. In Proceedings of the 2020 53rd Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO). 868\u2013880.","DOI":"10.1109\/MICRO50266.2020.00075"},{"key":"e_1_3_2_114_2","doi-asserted-by":"crossref","unstructured":"Matthew Poremba Itir Akgun Jieming Yin Onur Kayiran Yuan Xie and Gabriel H. Loh. 2017. There and back again: Optimizing the interconnect in networks of memory cubes. ACM SIGARCH Computer Architecture News (CAN). 678\u2013690.","DOI":"10.1145\/3079856.3080251"},{"key":"e_1_3_2_115_2","unstructured":"Open Compute Project. 2023. Open Accelerator Infrastructure (OAI) - OCP Accelerator Module (OAM) Base Specification. Retrieved April 2026 from https:\/\/www.opencompute.org\/documents\/oai-oam-base-specification-r2-0-v1-0-20230919-pdf"},{"key":"e_1_3_2_116_2","doi-asserted-by":"crossref","unstructured":"Zaid Qureshi Vikram Sharma Mailthody Isaac Gelado Seungwon Min Amna Masood Jeongmin Park Jinjun Xiong C. J. Newburn Dmitri Vainbrand et\u00a0al. 2023. GPU-initiated on-demand high-throughput storage access in the bam system architecture. 325\u2013339.","DOI":"10.1145\/3575693.3575748"},{"key":"e_1_3_2_117_2","doi-asserted-by":"publisher","DOI":"10.1145\/3458817.3476205"},{"key":"e_1_3_2_118_2","doi-asserted-by":"crossref","unstructured":"Jie Ren Jiaolin Luo Kai Wu Minjia Zhang Hyeran Jeon and Dong Li. 2021. Sentinel: Efficient tensor migration and allocation on heterogeneous memory systems for deep learning. In IEEE International Symposium on High-Performance Computer Architecture (HPCA). 598\u2013611.","DOI":"10.1109\/HPCA51647.2021.00057"},{"key":"e_1_3_2_119_2","doi-asserted-by":"crossref","unstructured":"Amir Roozbeh Jo\u00e3o Soares Gerald Q. Maguire Fetahi Wuhib Chakri Padala Mozhgan Mahloo Daniel Turull Vinay Yadhav and Dejan Kosti\u0107. 2028. Software-defined \u201chardware\u201d infrastructures: A survey on enabling technologies and open research directions. IEEE Communications Surveys & Tutorials 20 3 (2018) 2454\u20132485.","DOI":"10.1109\/COMST.2018.2834731"},{"key":"e_1_3_2_120_2","doi-asserted-by":"crossref","unstructured":"Amitabha Roy Laurent Bindschaedler Jasmina Malicevic and Willy Zwaenepoel. 2015. Chaos: Scale-out graph processing from secondary storage. In Symposium on Operating Systems Principles (SOSP). 410\u2013424.","DOI":"10.1145\/2815400.2815408"},{"key":"e_1_3_2_121_2","doi-asserted-by":"publisher","DOI":"10.1109\/CLOUD.2011.42"},{"key":"e_1_3_2_122_2","unstructured":"Zhenyuan Ruan Malte Schwarzkopf Marcos K. Aguilera and Adam Belay. AIFM: High-performance application-integrated far memory. In USENIX Symposium on Operating Systems Design and Implementation (OSDI). 315\u2013332."},{"key":"e_1_3_2_123_2","doi-asserted-by":"publisher","DOI":"10.1145\/3492321.3524272"},{"key":"e_1_3_2_124_2","doi-asserted-by":"crossref","unstructured":"Malte Schwarzkopf Andy Konwinski Michael Abd-El-Malek and John Wilkes. 2013. Omega: Flexible scalable schedulers for large compute clusters. In ACM European Conference on Computer Systems (EuroSys). 351\u2013364.","DOI":"10.1145\/2465351.2465386"},{"key":"e_1_3_2_125_2","unstructured":"Yizhou Shan Yutong Huang Yilun Chen and Yiying Zhang. 2018. Legoos: A disseminated distributed OS for hardware resource disaggregation. In USENIX Symposium on Operating Systems Design and Implementation (OSDI). 69\u201387."},{"key":"e_1_3_2_126_2","doi-asserted-by":"publisher","DOI":"10.1145\/3127479.3128610"},{"key":"e_1_3_2_127_2","doi-asserted-by":"crossref","unstructured":"Chuanming Shao Jinyang Guo Pengyu Wang Jing Wang Chao Li and Minyi Guo. 2022. Oversubscribing gpu unified virtual memory: Implications and suggestions. In ACM\/SPEC International Conference on Performance Engineering (ICPE). 67\u201375.","DOI":"10.1145\/3489525.3511691"},{"key":"e_1_3_2_128_2","doi-asserted-by":"publisher","DOI":"10.1145\/2442516.2442530"},{"key":"e_1_3_2_129_2","doi-asserted-by":"publisher","DOI":"10.1145\/3381898.3397215"},{"key":"e_1_3_2_130_2","doi-asserted-by":"crossref","unstructured":"Jie Sun Zuocheng Shi Li Su Wenting Shen Zeke Wang Yong Li Wenyuan Yu Wei Lin Fei Wu Bingsheng He and Jingren Zhou. 2025. Helios: Efficient distributed dynamic graph sampling for online GNN inference. In ACM SIGPLAN Annual Symposium on Principles and Practice of Parallel Programming (PPoPP). 2\u201315.","DOI":"10.1145\/3710848.3710854"},{"key":"e_1_3_2_131_2","doi-asserted-by":"crossref","unstructured":"Xuan Sun Hu Wan Qiao Li Chia-Lin Yang Tei-Wei Kuo and Chun Jason Xue. 2022. RM-SSD: In-storage computing for large-scale recommendation inference. In IEEE International Symposium on High-Performance Computer Architecture (HPCA). 1056\u20131070.","DOI":"10.1109\/HPCA53966.2022.00081"},{"key":"e_1_3_2_132_2","unstructured":"Shin-Yeh Tsai Yizhou Shan and Yiying Zhang. 2020. Disaggregating persistent memory and controlling them remotely: An exploration of passive disaggregated key-valuestores. In USENIX Annual Technical Conference (USENIX ATC). 1\u201316."},{"key":"e_1_3_2_133_2","doi-asserted-by":"publisher","DOI":"10.1145\/3132747.3132762"},{"key":"e_1_3_2_134_2","doi-asserted-by":"crossref","unstructured":"Evangelos Vasilakis Vassilis Papaefstathiou Pedro Trancoso and Ioannis Sourdis. 2020. Hybrid \\(^2\\) : Combining caching and migration in hybrid memory systems. In 2020 IEEE International Symposium on High Performance Computer Architecture (HPCA). IEEE 649\u2013662.","DOI":"10.1109\/HPCA47549.2020.00059"},{"key":"e_1_3_2_135_2","doi-asserted-by":"crossref","unstructured":"ChenxiWang Huimin Cui Ting Cao John Zigman Haris Volos Onur Mutlu Fang Lv Xiaobing Feng and Guoqing Harry Xu. 2019. Panthera: Holistic memory management for big data processing over hybrid memories. In ACM SIGPLAN Conference on Programming Language Design and Implementation. 347\u2013362.","DOI":"10.1145\/3314221.3314650"},{"key":"e_1_3_2_136_2","unstructured":"Chenxi Wang Haoran Ma Shi Liu Yuanqi Li Zhenyuan Ruan Khanh Nguyen Michael D. Bond Ravi Netravali Miryung Kim and Guoqing Harry Xu. 2020. Semeru: A memory-disaggregated managed runtime. In USENIX Symposium on Operating Systems Design and Implementation (OSDI). 261\u2013280."},{"key":"e_1_3_2_137_2","unstructured":"Chenxi Wang Haoran Ma Shi Liu Yifan Qiao Jonathan Eyolfson Christian Navasca Shan Lu and Guoqing Harry Xu. 2022. MemLiner: Lining up tracing and application for a far-memory-friendly runtime. In USENIX Symposium on Operating Systems Design and Implementation (OSDI). 35\u201353."},{"key":"e_1_3_2_138_2","unstructured":"Chenxi Wang Yifan Qiao Haoran Ma Shi Liu Wenguang Chen Ravi Netravali Miryung Kim and Guoqing Harry Xu. 2023. Canvas: Isolated and adaptive swapping for multi-applications on remote memory. In USENIX Symposium on Networked Systems Design and Implementation (NSDI 23). 161\u2013179."},{"key":"e_1_3_2_139_2","doi-asserted-by":"crossref","unstructured":"Jing Wang Chao Li Taolei Wang Lu Zhang Pengyu Wang Junyi Mei and Minyi Guo. 2022. Excavating the potential of graph workload on RDMA-based far memory architecture (ipdps). In 2022 IEEE International Parallel and Distributed Processing Symposium (IPDPS). 1029\u20131039.","DOI":"10.1109\/IPDPS53621.2022.00104"},{"key":"e_1_3_2_140_2","doi-asserted-by":"crossref","unstructured":"Jing Wang Chao Li Yibo Liu Taolei Wang Lu Zhang Pengyu Wang Junyi Mei and Minyi Guo. 2023. Fargraph+: Excavating the parallelism of graph processing workload on RDMA-based far memory system. In Journal of Parallel and Distributed Computing (JPDC) 177 (2023) 144\u2013159.","DOI":"10.1016\/j.jpdc.2023.02.015"},{"key":"e_1_3_2_141_2","doi-asserted-by":"crossref","unstructured":"Jing Wang Chao Li Junyi Mei Hao He Taolei Wang Pengyu Wang Lu Zhang Minyi Guo Hanqing Wu Dongbai Chen and Xiangwen Liu. 2022. HyFarM: Task orchestration on hybrid far memory for high performance per bit. In International Conference on Computer Design (ICCD). 33\u201341.","DOI":"10.1109\/ICCD56317.2022.00016"},{"key":"e_1_3_2_142_2","doi-asserted-by":"crossref","unstructured":"Jing Wang Qing Wang Yuhao Zhang and Jiwu Shu. 2025. Deft: A scalable tree index for disaggregated memory. In Proceedings of the Twentieth European Conference on Computer Systems (Eurosys). 886\u2013901.","DOI":"10.1145\/3689031.3696062"},{"key":"e_1_3_2_143_2","doi-asserted-by":"crossref","unstructured":"Jing Wang Hanzhang Yang Chao Li Yiming Zhuansun Wang Yuan Cheng Xu Xiaofeng Hou Minyi Guo Yang Hu and Yaqian Zhao. 2024. Boosting data center performance via intelligently managed multi-backend disaggregated memory. In International Conference for High Performance Computing Networking Storage and Analysis (SC). 1\u201318.","DOI":"10.1109\/SC41406.2024.00043"},{"key":"e_1_3_2_144_2","doi-asserted-by":"crossref","unstructured":"Linnan Wang Jinmian Ye Yiyang Zhao Wei Wu Ang Li Shuaiwen Leon Song Zenglin Xu and Tim Kraska. 2018. Superneurons: Dynamic GPU memory management for training deep neural networks. In Principles and Practice of Parallel Programming (PPoPP). 41\u201353.","DOI":"10.1145\/3178487.3178491"},{"key":"e_1_3_2_145_2","doi-asserted-by":"crossref","unstructured":"Pengyu Wang Chao Li Jing Wang Taolei Wang Lu Zhang Jingwen Leng Quan Chen and Minyi Guo. 2021. Skywalker: Efficient aliasmethod-based graph sampling and random walk on GPUs. In Parallel Architectures and Compilation Techniques (PACT). 304\u2013317.","DOI":"10.1109\/PACT52795.2021.00029"},{"key":"e_1_3_2_146_2","doi-asserted-by":"crossref","unstructured":"Pengyu Wang Jing Wang Chao Li Jianzong Wang Haojin Zhu and Minyi Guo. 2021. Grus: Toward unified-memory-efficient highperformance graph processing on GPU. ACM Transactions on Architecture and Code Optimization (TACO) 18 2 (2021) 1\u201325.","DOI":"10.1145\/3444844"},{"key":"e_1_3_2_147_2","doi-asserted-by":"crossref","unstructured":"Qing Wang Youyou Lu and Jiwu Shu. 2022. Sherman: A write-optimized distributed B+tree index on disaggregated memory. In International Conference on Management of Data (SIGMOD). 1\u201316.","DOI":"10.1145\/3604437.3604448"},{"key":"e_1_3_2_148_2","doi-asserted-by":"crossref","unstructured":"Ruihong Wang Chuqing Gao Jianguo Wang Prishita Kadam M. Tamer \u00d6zsu and Walid G. Aref. 2024. Optimizing LSM-based indexes for disaggregated memory. The VLDB Journal. 1\u201324.","DOI":"10.1007\/s00778-024-00863-y"},{"key":"e_1_3_2_149_2","doi-asserted-by":"crossref","unstructured":"Ruihong Wang Jianguo Wang Stratos Idreos M. Tamer Ozsu and Walid G. Aref. 2022. The case for distributed shared-memory databases with RDMA-enabled memory disaggregation. VLDB. 15\u201322.","DOI":"10.14778\/3561261.3561263"},{"key":"e_1_3_2_150_2","doi-asserted-by":"crossref","unstructured":"Ruihong Wang Jianguo Wang Prishita Kadam M. Tamer \u00d6zsu and Walid G. Aref. 2023. dLSM: An lsm-based index for memory disaggregation. In IEEE International Conference on Data Engineering (ICDE). 2835\u20132849.","DOI":"10.1109\/ICDE55515.2023.00217"},{"key":"e_1_3_2_151_2","doi-asserted-by":"crossref","unstructured":"Xinkai Wang Hao He Yuancheng Li Chao Li Xiaofeng Hou Jing Wang Quan Chen Jingwen Leng Minyi Guo and Leibo Wang. 2023. Not all resources are visible: Exploiting fragmented shadow resources in shared-state scheduler architecture. In ACM Symposium on Cloud Computing (SoCC). 109\u2013124.","DOI":"10.1145\/3620678.3624650"},{"key":"e_1_3_2_152_2","doi-asserted-by":"crossref","unstructured":"Zixuan Wang Joonseop Sim Euicheol Lim and Jishen Zhao. 2022. Enabling efficient large-scale deep learning training with cache coherent disaggregated memory systems. In International Symposium on High-Performance Computer Architecture (HPCA). 126\u2013140.","DOI":"10.1109\/HPCA53966.2022.00018"},{"key":"e_1_3_2_153_2","unstructured":"Xingda Wei Rongxin Cheng Yuhan Yang Rong Chen and Haibo Chen. 2023. Characterizing Off-path SmartNIC for accelerating distributed systems. In USENIX Symposium on Operating Systems Design and Implementation (OSDI). 987\u20131004."},{"key":"e_1_3_2_154_2","unstructured":"Xingda Wei Fangming Lu Rong Chen and Haibo Chen. 2022. KRCORE: A microsecond-scale RDMA control plane for elastic computing. In USENIX Annual Technical Conference (USENIX ATC). 121\u2013136."},{"key":"e_1_3_2_155_2","unstructured":"Xingda Wei Fangming Lu Tianxia Wang Jinyu Gu Yuhan Yang Rong Chen and Haibo Chen. 2023. No provisioned concurrency: Fast RDMA-codesigned remote fork for serverless computing. In USENIX Symposium on Operating Systems Design and Implementation (OSDI). 497\u2013517."},{"key":"e_1_3_2_156_2","doi-asserted-by":"crossref","unstructured":"Michele Weiland Holger Brunst Tiago Quintino Nick Johnson Olivier Iffrig Simon Smart Christian Herold Antonino Bonanni Adrian Jackson and Mark Parsons. 2019. An early evaluation of intel\u2019s optane DC persistent memory module and its impact on high-performance scientific applications. In International Conference for High Performance Computing Networking Storage and Analysis (SC). 1\u201319.","DOI":"10.1145\/3295500.3356159"},{"key":"e_1_3_2_157_2","doi-asserted-by":"crossref","unstructured":"Johannes Weiner Niket Agarwal Dan Schatzberg Leon Yang Hao Wang Blaise Sanouillet Bikash Sharma Tejun Heo Mayank Jain Chunqiang Tang and Dimitrios Skarlatos. 2022. Tmo: Transparent memory offloading in datacenters. In ACM International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS). 609\u2013621.","DOI":"10.1145\/3503222.3507731"},{"key":"e_1_3_2_158_2","doi-asserted-by":"crossref","unstructured":"Du Xiang Tao Liu Jilian Xu Jun Y. Tan Zehua Hu Bo Lei Yue Zheng Jing Wu A. H. Castro Neto Lei Liu and Wei Chen. 2018. Two dimensional multibit optoelectronic memory with broadband spectrum distinction. Nature Communications 9 2966 (2018) 1\u201315.","DOI":"10.1038\/s41467-018-05397-w"},{"key":"e_1_3_2_159_2","unstructured":"Shao-Peng Yang Minjae Kim Sanghyun Nam Juhyung Park Jin yong Choi Eyee Hyun Nam Eunji Lee Sungjin Lee and Bryan S. Kim. 2023. Overcoming the memory wall with CXL-enabled SSDs. In Proceedings of the USENIX Annual Technical Conference (USENIX ATC). 601\u2013617."},{"key":"e_1_3_2_160_2","doi-asserted-by":"crossref","unstructured":"Daniel Zahka and Ada Gavrilovska. 2022. FAM-graph: Graph analytics on disaggregated memory. In Proceedings of the IEEE International Parallel and Distributed Processing Symposium (IPDPS). 1\u201314.","DOI":"10.1109\/IPDPS53621.2022.00017"},{"key":"e_1_3_2_161_2","doi-asserted-by":"crossref","unstructured":"Jia Zhan Itir Akgun Jishen Zhao Al Davis Paolo Faraboschi Yuangang Wang and Yuan Xie. 2016. A unified memory network architecture for in-memory computing in commodity servers. In IEEE\/ACM International Symposium on Microarchitecture (MICRO). 1\u201314.","DOI":"10.1109\/MICRO.2016.7783732"},{"key":"e_1_3_2_162_2","unstructured":"Haoran Zhang Adney Cardoza Peter Baile Chen Sebastian Angel and Vincent Liu. 2020. Fault-Tolerant and transactional stateful serverless workflows. In USENIX Symposium on Operating Systems Design and Implementation (OSDI). 1187\u20131204."},{"key":"e_1_3_2_163_2","doi-asserted-by":"crossref","unstructured":"Haoyang Zhang Yirui Zhou et\u00a0al. 2023. G10: Enabling an efficient unified GPU memory and storage architecture with smart tensor migrations. In Proceedings of the 56th Annual IEEE\/ACM International Symposium on Microarchitecture. 395\u2013410.","DOI":"10.1145\/3613424.3614309"},{"key":"e_1_3_2_164_2","doi-asserted-by":"publisher","DOI":"10.1145\/2688500.2688507"},{"key":"e_1_3_2_165_2","unstructured":"Ming Zhang Yu Hua Pengfei Zuo and Lurong Liu. 2022. FORD: Fast one-sided RDMA-based distributed transactions for disaggregated persistent memory. In USENIX Conference on File and Storage Technologies (FAST). 51\u201368."},{"key":"e_1_3_2_166_2","doi-asserted-by":"crossref","unstructured":"Mingxing Zhang Yongwei Wu Youwei Zhuo Xuehai Qian Chengying Huan and Kang Chen. 2018. Wonderland: A novel abstraction-based out-of-core graph processing system. Architectural Support for Programming Languages and Operating Systems (ASPLOS). 608\u2013621.","DOI":"10.1145\/3173162.3173208"},{"key":"e_1_3_2_167_2","doi-asserted-by":"crossref","unstructured":"Pengfei Zhang Xi Li Rui Chu and Huaimin Wang. 2015. HybridSwap: A scalable and synthetic framework for guest swapping on virtualization platform. In IEEE Conference on Computer Communications (INFOCOMM). 864\u2013872.","DOI":"10.1109\/INFOCOM.2015.7218457"},{"key":"e_1_3_2_168_2","doi-asserted-by":"crossref","unstructured":"Qizhen Zhang Xinyi Chen Sidharth Sankhe Zhilei Zheng Ke Zhong Sebastian Angel Ang Chen Vincent Liu and Boon Thau Loo. 2022. Optimizing data-intensive systems in disaggregated data centers with teleport. In International Conference on Management of Data (SIGMOD). 1345\u20131359.","DOI":"10.1145\/3514221.3517856"},{"key":"e_1_3_2_169_2","doi-asserted-by":"crossref","unstructured":"Yunming Zhang Ajay Brahmakshatriya Xinyi Chen Laxman Dhulipala Shoaib Kamil Saman Amarasinghe and Julian Shun. 2020. Optimizing ordered graph algorithms with graphit. In The International Symposium on Code Generation and Optimization (CGO). 158\u2013170.","DOI":"10.1145\/3368826.3377909"},{"key":"e_1_3_2_170_2","doi-asserted-by":"crossref","unstructured":"Yunming Zhang Vladimir Kiriansky Charith Mendis Saman Amarasinghe and Matei Zaharia. 2017. Making caches work for graph analytics. In IEEE International Conference on Big Data (Big Data). 293\u2013302.","DOI":"10.1109\/BigData.2017.8257937"},{"key":"e_1_3_2_171_2","doi-asserted-by":"crossref","unstructured":"Yuhong Zhong Daniel S. Berger et\u00a0al. 2025. Oasis: Pooling PCIe devices over CXL to boost utilization. In Proceedings of the ACM SIGOPS 31st Symposium on Operating Systems Principles (SOSP). 101\u2013119.","DOI":"10.1145\/3731569.3764812"},{"key":"e_1_3_2_172_2","unstructured":"Yijie Zhong Minqiang Zhou Zhirong Shen and Jiwu Shu. 2024. UniMem: Redesigning disaggregated memory within a unified local-remote memory hierarchy. In USENIX Annual Technical Conference (USENIX ATC). 463\u2013477."},{"key":"e_1_3_2_173_2","unstructured":"Yang Zhou Hassan M. G. Wassel Sihang Liu Jiaqi Gao James Mickens Minlan Yu Chris Kennelly Paul Turner David E. Culler Henry M. Levy and Amin Vahdat. 2022. Carbink: Fault-tolerant far memory. In USENIX Symposium on Operating Systems Design and Implementation (OSDI). 55\u201371."},{"key":"e_1_3_2_174_2","doi-asserted-by":"crossref","unstructured":"Bohong Zhu Youmin Chen Qing Wang Youyou Lu and Jiwu Shu. 2021. Octopus+: An RDMA-enabled distributed persistent memory file system. ACM Transactions on Storage (TOS) 17 3 (2021) 1\u201325.","DOI":"10.1145\/3448418"},{"key":"e_1_3_2_175_2","unstructured":"Xiaowei Zhu Wentao Han andWenguang Chen. 2015. GridGraph: Large-scale graph processing on a single machine using 2-level hierarchical partitioning. In USENIX Annual Technical Conference (USENIX ATC). 1\u201315."},{"key":"e_1_3_2_176_2","doi-asserted-by":"crossref","unstructured":"Ziyi Zhu Yiwen Shen Yishen Huang Alexander Gazman Maarten Hattink and Keren Bergman. 2019. Flexible resource allocation using photonic switched interconnects for disaggregated system architectures. In Proceedings fo the Optical Fiber Communication Conference (OFC). 1\u201313.","DOI":"10.1364\/OFC.2019.M3F.3"},{"key":"e_1_3_2_177_2","doi-asserted-by":"crossref","unstructured":"Tobias Ziegler Sumukha Tumkur Vani et\u00a0al. 2019. Designing distributed tree-based index structures for fast RDMA-capable networks. In International Conference on Management of Data (SIGMOD). Association for Computing Machinery 741\u2013758.","DOI":"10.1145\/3299869.3300081"},{"key":"e_1_3_2_178_2","unstructured":"Pengfei Zuo Jiazhao Sun Liu Yang Shuangwu Zhang and Yu Hua. 2021. One-sided RDMA-conscious extendible hashing for disaggregated memory. In USENIX Annual Technical Conference (USENIX ATC). 15\u201329."}],"container-title":["ACM Computing Surveys"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3807443","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,5,15]],"date-time":"2026-05-15T16:10:37Z","timestamp":1778861437000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3807443"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,5,15]]},"references-count":177,"journal-issue":{"issue":"12","published-print":{"date-parts":[[2026,9,30]]}},"alternative-id":["10.1145\/3807443"],"URL":"https:\/\/doi.org\/10.1145\/3807443","relation":{},"ISSN":["0360-0300","1557-7341"],"issn-type":[{"value":"0360-0300","type":"print"},{"value":"1557-7341","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,5,15]]},"assertion":[{"value":"2024-08-05","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-03-24","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-05-15","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}