{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,12,8]],"date-time":"2025-12-08T22:37:58Z","timestamp":1765233478049,"version":"3.41.0"},"reference-count":114,"publisher":"Association for Computing Machinery (ACM)","issue":"1","license":[{"start":{"date-parts":[[2023,2,27]],"date-time":"2023-02-27T00:00:00Z","timestamp":1677456000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["Proc. ACM Meas. Anal. Comput. Syst."],"published-print":{"date-parts":[[2023,2,27]]},"abstract":"<jats:p>Resource disaggregation offers a cost effective solution to resource scaling, utilization, and failure-handling in data centers by physically separating hardware devices in a server. Servers are architected as pools of processor, memory, and storage devices, organized as independent failure-isolated components interconnected by a high-bandwidth network. A critical challenge, however, is the high performance penalty of accessing data from a remote memory module over the network. Addressing this challenge is difficult as disaggregated systems have high runtime variability in network latencies\/bandwidth, and page migration can significantly delay critical path cache line accesses in other pages.<\/jats:p>\n          <jats:p>This paper conducts a characterization analysis on different data movement strategies in fully disaggregated systems, evaluates their performance overheads in a variety of workloads, and introduces DaeMon, the first software-transparent mechanism to significantly alleviate data movement overheads in fully disaggregated systems. First, to enable scalability to multiple hardware components in the system, we enhance each compute and memory unit with specialized engines that transparently handle data migrations. Second, to achieve high performance and provide robustness across various network, architecture and application characteristics, we implement a synergistic approach of bandwidth partitioning, link compression, decoupled data movement of multiple granularities, and adaptive granularity selection in data movements. We evaluate DaeMon in a wide variety of workloads at different network and architecture configurations using a state-of-the-art simulator. DaeMon improves system performance and data access costs by 2.39\u00d7 and 3.06\u00d7, respectively, over the widely-adopted approach of moving data at page granularity.<\/jats:p>","DOI":"10.1145\/3579445","type":"journal-article","created":{"date-parts":[[2023,3,2]],"date-time":"2023-03-02T23:50:57Z","timestamp":1677801057000},"page":"1-36","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":10,"title":["DaeMon: Architectural Support for Efficient Data Movement in Fully Disaggregated Systems"],"prefix":"10.1145","volume":"7","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-0162-4547","authenticated-orcid":false,"given":"Christina","family":"Giannoula","sequence":"first","affiliation":[{"name":"University of Toronto &amp; National Technical University of Athens, Toronto, ON, Canada"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8289-5597","authenticated-orcid":false,"given":"Kailong","family":"Huang","sequence":"additional","affiliation":[{"name":"University of Toronto, Toronto, ON, Canada"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6893-2598","authenticated-orcid":false,"given":"Jonathan","family":"Tang","sequence":"additional","affiliation":[{"name":"University of Toronto, Toronto, ON, Canada"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4890-8427","authenticated-orcid":false,"given":"Nectarios","family":"Koziris","sequence":"additional","affiliation":[{"name":"National Technical University of Athens, Athens, Greece"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7811-4831","authenticated-orcid":false,"given":"Georgios","family":"Goumas","sequence":"additional","affiliation":[{"name":"National Technical University of Athens, Athens, Greece"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1455-4843","authenticated-orcid":false,"given":"Zeshan","family":"Chishti","sequence":"additional","affiliation":[{"name":"Intel Corporation, Portland, OR, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3315-9336","authenticated-orcid":false,"given":"Nandita","family":"Vijaykumar","sequence":"additional","affiliation":[{"name":"University of Toronto, Toronto, ON, Canada"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2023,3,2]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"crossref","unstructured":"B. Abali H. Franke D. E. Poff R. A. Saccone C. O. Schulz L. M. Herger and T. B. Smith. 2001. Memory Expansion Technology (MXT): Software Support and Performance. IBM Journal of Research and Development (2001).","DOI":"10.1147\/rd.452.0287"},{"key":"e_1_2_1_2_1","doi-asserted-by":"crossref","unstructured":"Atul Adya Robert Grandl Daniel Myers and Henry Qin. 2019. Fast Key-Value Stores: An Idea Whose Time Has Come and Gone. In HotOS.","DOI":"10.1145\/3317550.3321434"},{"key":"e_1_2_1_3_1","volume-title":"Keckler","author":"Agarwal Neha","year":"2015","unstructured":"Neha Agarwal, David Nellans, Mark Stephenson, Mike O'Connor, and Stephen W. Keckler. 2015. Page Placement Strategies for GPUs within Heterogeneous Memory Systems. In ASPLOS."},{"key":"e_1_2_1_4_1","volume-title":"Wenisch","author":"Agarwal Neha","year":"2017","unstructured":"Neha Agarwal and Thomas F. Wenisch. 2017. Thermostat: Application-Transparent Page Management for Two-Tiered Main Memory. In ASPLOS."},{"key":"e_1_2_1_5_1","volume-title":"Remote Regions: A Simple Abstraction for Remote Memory. In ATC.","author":"Aguilera Marcos K.","year":"2018","unstructured":"Marcos K. Aguilera, Nadav Amit, Irina Calciu, Xavier Deguillard, Jayneel Gandhi, Stanko Novakovi\u0107, Arun Ramanathan, Pratap Subrahmanyam, Lalith Suresh, Kiran Tati, Rajesh Venkatasubramanian, and Michael Wei. 2018. Remote Regions: A Simple Abstraction for Remote Memory. In ATC."},{"key":"e_1_2_1_6_1","doi-asserted-by":"crossref","unstructured":"Marcos K. Aguilera Nadav Amit Irina Calciu Xavier Deguillard Jayneel Gandhi Pratap Subrahmanyam Lalith Suresh Kiran Tati Rajesh Venkatasubramanian and Michael Wei. 2017. Remote Memory in the Age of Fast Networks. In SoCC.","DOI":"10.1145\/3127479.3131612"},{"key":"e_1_2_1_7_1","doi-asserted-by":"crossref","unstructured":"A.R. Alameldeen and D.A. Wood. 2004 a. Adaptive Cache Compression for High-Performance Processors. In ISCA.","DOI":"10.1145\/1028176.1006719"},{"key":"e_1_2_1_8_1","unstructured":"Alaa Alameldeen and David Wood. 2004 b. Frequent Pattern Compression: A Significance-Based Compression Scheme for L2 Caches. (2004)."},{"key":"e_1_2_1_9_1","volume-title":"Wood","author":"Alameldeen Alaa R.","year":"2007","unstructured":"Alaa R. Alameldeen and David A. Wood. 2007. Interactions Between Compression and Prefetching in Chip Multiprocessors. In HPCA."},{"key":"e_1_2_1_10_1","unstructured":"Sebastian Angel Mihir Nanavati and Siddhartha Sen. 2020. Disaggregation and the Application. In HotCloud."},{"key":"e_1_2_1_11_1","doi-asserted-by":"crossref","unstructured":"Angelos Arelakis Fredrik Dahlgren and Per Stenstrom. 2015. HyComp: A Hybrid Cache Compression Method for Selection of Data-Type-Specific Compression Methods. In MICRO.","DOI":"10.1145\/2830772.2830823"},{"key":"e_1_2_1_12_1","doi-asserted-by":"crossref","unstructured":"Angelos Arelakis and Per Stenstrom. 2014. SC2: A Statistical Compression Cache Scheme. In ISCA.","DOI":"10.1145\/2678373.2665696"},{"key":"e_1_2_1_13_1","first-page":"D79","volume":"201","unstructured":"JEDEC Solid State Technology Assn. 2017. JESD79--4B: DDR4 SDRAM Standard. http:\/\/www.softnology.biz\/pdf\/JESD79--4B.pdf","journal-title":"JEDEC Solid State Technology Assn."},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/3373376.3378504"},{"key":"e_1_2_1_15_1","doi-asserted-by":"crossref","unstructured":"Dhantu Buragohain Abhishek Ghogare Trishal Patel Mythili Vutukuru and Purushottam Kulkarni. 2017. DiME: A Performance Emulator for Disaggregated Memory Architectures. In APSys.","DOI":"10.1145\/3124680.3124731"},{"key":"e_1_2_1_16_1","volume-title":"Onur Mutlu, and Aasheesh Kolli.","author":"Calciu Irina","year":"2021","unstructured":"Irina Calciu, M. Talha Imran, Ivan Puddu, Sanidhya Kashyap, Hasan Al Maruf, Onur Mutlu, and Aasheesh Kolli. 2021. Rethinking Software Runtimes for Disaggregated Memory. In ASPLOS."},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/2063384.2063454"},{"key":"e_1_2_1_18_1","volume-title":"An Evaluation of High-Level Mechanistic Core Models. TACO","author":"Carlson Trevor E.","year":"2014","unstructured":"Trevor E. Carlson, Wim Heirman, Stijn Eyerman, Ibrahim Hur, and Lieven Eeckhout. 2014. An Evaluation of High-Level Mechanistic Core Models. TACO (2014)."},{"key":"e_1_2_1_19_1","doi-asserted-by":"crossref","unstructured":"Chia-Hao Chang Adithya Kumar and Anand Sivasubramaniam. 2021. To Move or Not to Move? Page Migration for Irregular Applications in over-Subscribed GPU Memory Systems with DynaMap. In SYSTOR.","DOI":"10.1145\/3456727.3463766"},{"key":"e_1_2_1_20_1","volume-title":"Rodinia: A Benchmark Suite for Heterogeneous Computing. In IISWC.","author":"Che Shuai","year":"2009","unstructured":"Shuai Che, Michael Boyer, Jiayuan Meng, David Tarjan, Jeremy W. Sheaffer, Sang-Ha Lee, and Kevin Skadron. 2009. Rodinia: A Benchmark Suite for Heterogeneous Computing. In IISWC."},{"key":"e_1_2_1_21_1","volume-title":"C-Pack: A High-Performance Microprocessor Cache Compression Algorithm. VLSI","author":"Chen Xi","year":"2010","unstructured":"Xi Chen, Lei Yang, Robert P. Dick, Li Shang, and Haris Lekatsas. 2010. C-Pack: A High-Performance Microprocessor Cache Compression Algorithm. VLSI (2010)."},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/3132402.3132404"},{"key":"e_1_2_1_23_1","volume-title":"Qureshi","author":"Chou Chia Chen","year":"2014","unstructured":"Chia Chen Chou, Aamer Jaleel, and Moinuddin K. Qureshi. 2014. CAMEO: A Two-Level Memory Organization with Capacity of Main Memory and Flexibility of Hardware-Managed Cache. In MICRO."},{"key":"e_1_2_1_24_1","volume-title":"Alameldeen","author":"Choukse Esha","year":"2018","unstructured":"Esha Choukse, Mattan Erez, and Alaa R. Alameldeen. 2018. Compresso: Pragmatic Main Memory Compression. In MICRO."},{"key":"e_1_2_1_25_1","volume-title":"Enzian: An Open, General, CPU\/FPGA Platform for Systems Software Research. In ASPLOS.","author":"Cock David","year":"2022","unstructured":"David Cock, Abishek Ramdas, Daniel Schwyn, Michael Giardino, Adam Turowski, Zhenhao He, Nora Hossle, Dario Korolija, Melissa Licciardello, Kristina Martsenko, Reto Achermann, Gustavo Alonso, and Timothy Roscoe. 2022. Enzian: An Open, General, CPU\/FPGA Platform for Systems Software Research. In ASPLOS."},{"key":"e_1_2_1_26_1","volume-title":"Jouppi","author":"Dong Xiangyu","year":"2010","unstructured":"Xiangyu Dong, Yuan Xie, Naveen Muralimanohar, and Norman P. Jouppi. 2010. Simple but Effective Heterogeneous Main Memory with On-Chip Memory Controller Support. In SC."},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/3307681.3325398"},{"key":"e_1_2_1_28_1","volume-title":"Cori: Dancing to the Right Beat of Periodic Data Movements over Hybrid Memory Systems. In IPDPS.","author":"Doudali Thaleia Dimitra","year":"2021","unstructured":"Thaleia Dimitra Doudali, Daniel Zahka, and Ada Gavrilovska. 2021. Cori: Dancing to the Right Beat of Periodic Data Movements over Hybrid Memory Systems. In IPDPS."},{"key":"e_1_2_1_29_1","doi-asserted-by":"crossref","unstructured":"Subramanya R. Dulloor Amitabha Roy Zheguang Zhao Narayanan Sundaram Nadathur Satish Rajesh Sankaran Jeff Jackson and Karsten Schwan. 2016. Data Tiering in Heterogeneous Memory Systems. In EuroSys.","DOI":"10.1145\/2901318.2901344"},{"volume-title":"2005 a","author":"Ekman Magnus","key":"e_1_2_1_30_1","unstructured":"Magnus Ekman and Per Stenstrom. 2005 a. A Cost-Effective Main Memory Organization for Future Servers. In IPDPS."},{"key":"e_1_2_1_31_1","doi-asserted-by":"crossref","unstructured":"M. Ekman and P. Stenstrom. 2005 b. A Robust Main-Memory Compression Scheme. In ISCA.","DOI":"10.1109\/ISCA.2005.6"},{"key":"e_1_2_1_32_1","doi-asserted-by":"crossref","unstructured":"M. J. Feeley W. E. Morgan E. P. Pighin A. R. Karlin H. M. Levy and C. A. Thekkath. 1995. Implementing Global Memory Management in a Workstation Cluster. In SOSP.","DOI":"10.1145\/224056.224072"},{"key":"e_1_2_1_33_1","unstructured":"Peter X. Gao Akshay Narayan Sagar Karandikar Joao Carreira Sangjin Han Rachit Agarwal Sylvia Ratnasamy and Scott Shenker. 2016. Network Requirements for Resource Disaggregation. In OSDI."},{"key":"e_1_2_1_34_1","unstructured":"Christina Giannoula. 2022. Accelerating Irregular Applications via Efficient Synchronization and Data Access Techniques. https:\/\/arxiv.org\/abs\/2211.05908"},{"key":"e_1_2_1_35_1","doi-asserted-by":"crossref","unstructured":"Christina Giannoula Kailong Huang Jonathan Tang Nectarios Koziris Georgios Goumas Zeshan Chishti and Nandita Vijaykumar. 2023. DaeMon: Architectural Support for Efficient Data Movement in Disaggregated Systems. In CoRR. https:\/\/arxiv.org\/abs\/2301.00414","DOI":"10.1145\/3578338.3593533"},{"key":"e_1_2_1_36_1","unstructured":"Donghyun Gouk Sangwon Lee Miryeong Kwon and Myoungsoo Jung. 2022. Direct Access High-Performance Memory Disaggregation with DirectCXL. In ATC."},{"key":"e_1_2_1_37_1","volume-title":"Shin","author":"Gu Juncheng","year":"2017","unstructured":"Juncheng Gu, Youngmoon Lee, Yiwen Zhang, Mosharaf Chowdhury, and Kang G. Shin. 2017. Efficient Memory Disaggregation with Infiniswap. In NSDI."},{"key":"e_1_2_1_38_1","unstructured":"Chuanxiong Guo Haitao Wu Zhong Deng Gaurav Soni Jianxi Ye Jitu Padhye and Marina Lipshteyn. 2016. RDMA over Commodity Ethernet at Scale. In SIGCOMM."},{"key":"e_1_2_1_39_1","volume-title":"Clio: A Hardware-Software Co-Designed Disaggregated Memory System. In ASPLOS.","author":"Guo Zhiyuan","year":"2022","unstructured":"Zhiyuan Guo, Yizhou Shan, Xuhao Luo, Yutong Huang, and Yiying Zhang. 2022. Clio: A Hardware-Software Co-Designed Disaggregated Memory System. In ASPLOS."},{"key":"e_1_2_1_40_1","unstructured":"Sangjin Han Norbert Egi Aurojit Panda Sylvia Ratnasamy Guangyu Shi and Scott Shenker. 2013. Network Support for Resource Disaggregation in Next-Generation Datacenters. In HotNets."},{"key":"e_1_2_1_41_1","unstructured":"HPCG. 2019. High Performance Conjugate Gradient Benchmark. https:\/\/github.com\/hpcg-benchmark\/hpcg"},{"key":"e_1_2_1_42_1","volume-title":"Centaur: A Chiplet-Based, Hybrid Sparse-Dense Accelerator for Personalized Recommendations. In ISCA.","author":"Hwang Ranggi","year":"2020","unstructured":"Ranggi Hwang, Taehun Kim, Youngeun Kwon, and Minsoo Rhu. 2020. Centaur: A Chiplet-Based, Hybrid Sparse-Dense Accelerator for Personalized Recommendations. In ISCA."},{"key":"e_1_2_1_43_1","unstructured":"Intel. 2021. Intel Omni-Path Architecture. https:\/\/www.intel.com\/content\/www\/us\/en\/high-performance-computing-fabrics\/omni-path-driving-exascale-computing.html"},{"key":"e_1_2_1_44_1","doi-asserted-by":"crossref","unstructured":"Djordje Jevdjic Stavros Volos and Babak Falsafi. 2013. Die-Stacked DRAM Caches for Servers: Hit Ratio Latency or Bandwidth? Have It All with Footprint Cache. In ISCA.","DOI":"10.1145\/2485922.2485957"},{"key":"e_1_2_1_45_1","volume-title":"CHOP: Adaptive Filter-Based DRAM Caching for CMP Server Platforms. In HPCA.","author":"Jiang Xiaowei","year":"2010","unstructured":"Xiaowei Jiang, Niti Madan, Li Zhao, Mike Upton, Ravishankar Iyer, Srihari Makineni, Donald Newell, Yan Solihin, and Rajeev Balasubramonian. 2010. CHOP: Adaptive Filter-Based DRAM Caching for CMP Server Platforms. In HPCA."},{"key":"e_1_2_1_46_1","unstructured":"Hongshin Jun Jinhee Cho Kangseol Lee Ho-Young Son Kwiwook Kim Hanho Jin and Keith Kim. 2017. HBM DRAM Technology and Architecture. In IMW."},{"key":"e_1_2_1_47_1","doi-asserted-by":"crossref","unstructured":"Sudarsun Kannan Ada Gavrilovska Vishal Gupta and Karsten Schwan. 2017. HeteroOS: OS Design for Heterogeneous Memory Management in Datacenter. In ISCA.","DOI":"10.1145\/3079856.3080245"},{"key":"e_1_2_1_48_1","doi-asserted-by":"crossref","unstructured":"K. Katrinis D. Syrivelis D. Pnevmatikatos G. Zervas D. Theodoropoulos I. Koutsopoulos K. Hasharoni D. Raho C. Pinto F. Espina S. Lopez-Buedo Q. Chen M. Nemirovsky D. Roca H. Klos and T. Berends. 2016. Rack-Scale Disaggregated Cloud Data Centers: The dReDBox Project Vision. In DATE.","DOI":"10.3850\/9783981537079_1014"},{"key":"e_1_2_1_49_1","unstructured":"Jonghyeon Kim Wonkyo Choe and Jeongseob Ahn. 2021. Exploring the Design Space of Page Management for Multi-Tiered Memory Systems. In ATC."},{"key":"e_1_2_1_50_1","unstructured":"Jungrae Kim Michael Sullivan Esha Choukse and Mattan Erez. 2016. Bit-Plane Compression: Transforming Data for Better Compression in Many-Core Architectures. In ISCA."},{"key":"e_1_2_1_51_1","unstructured":"Seikwon Kim Seonyoung Lee Taehoon Kim and Jaehyuk Huh. 2017. Transparent Dual Memory Compression Architecture. In PACT."},{"key":"e_1_2_1_52_1","unstructured":"M. Kjelso M. Gooch and S. Jones. 1996. Design and Performance of a Main Memory Hardware Data Compressor. In EUROMICRO."},{"key":"e_1_2_1_53_1","volume-title":"Taco: A Tool to Generate Tensor Algebra Kernels. In ASE.","author":"Kjolstad Fredrik","year":"2017","unstructured":"Fredrik Kjolstad, Stephen Chou, David Lugato, Shoaib Kamil, and Saman Amarasinghe. 2017. Taco: A Tool to Generate Tensor Algebra Kernels. In ASE."},{"key":"e_1_2_1_54_1","volume-title":"Kandemir","author":"Kotra Jagadish B.","year":"2018","unstructured":"Jagadish B. Kotra, Haibo Zhang, Alaa R. Alameldeen, Chris Wilkerson, and Mahmut T. Kandemir. 2018. CHAMELEON: A Dynamically Reconfigurable Heterogeneous Memory System. In MICRO."},{"key":"e_1_2_1_55_1","volume-title":"Yu Zhao, and Parthasarathy Ranganathan.","author":"Lagar-Cavilla Andres","year":"2019","unstructured":"Andres Lagar-Cavilla, Junwhan Ahn, Suleiman Souhlal, Neha Agarwal, Radoslaw Burny, Shakeel Butt, Jichuan Chang, Ashwin Chaugule, Nan Deng, Junaid Shahid, Greg Thelen, Kamil Adam Yurtsever, Yu Zhao, and Parthasarathy Ranganathan. 2019. Software-Defined Far Memory in Warehouse-Scale Computers. In ASPLOS."},{"key":"e_1_2_1_56_1","volume-title":"MIND: In-Network Memory Management for Disaggregated Data Centers. In SOSP.","author":"Yu Yanpeng","year":"2021","unstructured":"Seung-seob Lee, Yanpeng Yu, Yupeng Tang, Anurag Khandelwal, Lin Zhong, and Abhishek Bhattacharjee. 2021. MIND: In-Network Memory Management for Disaggregated Data Centers. In SOSP."},{"key":"e_1_2_1_57_1","unstructured":"Yang Li Saugata Ghose Jongmoo Choi Jin Sun Hui Wang and Onur Mutlu. 2017. Utility-Based Hybrid Memory Management. In CLUSTER."},{"key":"e_1_2_1_58_1","volume-title":"Wenisch","author":"Lim Kevin","year":"2009","unstructured":"Kevin Lim, Jichuan Chang, Trevor Mudge, Parthasarathy Ranganathan, Steven K. Reinhardt, and Thomas F. Wenisch. 2009. Disaggregated Memory for Expansion and Sharing in Blade Servers. In ISCA."},{"key":"e_1_2_1_59_1","volume-title":"Alvin AuYoung, Jichuan Chang, Parthasarathy Ranganathan, and Thomas F. Wenisch.","author":"Lim Kevin","year":"2012","unstructured":"Kevin Lim, Yoshio Turner, Jose Renato Santos, Alvin AuYoung, Jichuan Chang, Parthasarathy Ranganathan, and Thomas F. Wenisch. 2012. System-Level Implications of Disaggregated Memory. In HPCA."},{"key":"e_1_2_1_60_1","unstructured":"Haikun Liu Yujie Chen Xiaofei Liao Hai Jin Bingsheng He Long Zheng and Rentong Guo. 2017. Hardware\/Software Cooperative Caching for Hybrid DRAM\/NVM Memory Architectures. In ICS."},{"key":"e_1_2_1_61_1","volume-title":"Hierarchical Hybrid Memory Management in OS for Tiered Memory Systems. TPDS","author":"Liu Lei","year":"2019","unstructured":"Lei Liu, Shengjie Yang, Lu Peng, and Xinyu Li. 2019. Hierarchical Hybrid Memory Management in OS for Tiered Memory Systems. TPDS (2019)."},{"key":"e_1_2_1_62_1","volume-title":"Hill","author":"Loh Gabriel","year":"2012","unstructured":"Gabriel Loh and Mark D. Hill. 2012. Supporting Very Large DRAM Caches with Compound-Access Scheduling and MissMap. IEEE Micro (2012)."},{"volume-title":"Efficiently Enabling Conventional Block Sizes for Very Large Die-Stacked DRAM Caches. In MICRO.","author":"Gabriel","key":"e_1_2_1_63_1","unstructured":"Gabriel H. Loh and Mark D. Hill. 2011. Efficiently Enabling Conventional Block Sizes for Very Large Die-Stacked DRAM Caches. In MICRO."},{"key":"e_1_2_1_64_1","unstructured":"Hasan Al Maruf and Mosharaf Chowdhury. 2020. Effectively Prefetching Remote Memory with Leap. In ATC."},{"key":"e_1_2_1_65_1","unstructured":"Mellanox. 2020. Mellanox Innova Adapters. https:\/\/www.nvidia.com\/en-us\/networking\/products\/data-processing-unit\/?mtag=programmable_adapter_cards"},{"key":"e_1_2_1_66_1","volume-title":"Loh","author":"Meswani Mitesh R.","year":"2015","unstructured":"Mitesh R. Meswani, Sergey Blagodurov, David Roberts, John Slice, Mike Ignatowski, and Gabriel H. Loh. 2015. Heterogeneous Memory Architectures: A HW\/SW Approach for Mixing Die-Stacked and Off-Package Memories. In HPCA."},{"key":"e_1_2_1_67_1","volume-title":"Vetter","author":"Mittal Sparsh","year":"2016","unstructured":"Sparsh Mittal and Jeffrey S. Vetter. 2016. A Survey Of Architectural Approaches for Data Compression in Cache and Main Memory Systems. TPDS (2016)."},{"key":"e_1_2_1_68_1","doi-asserted-by":"crossref","unstructured":"Naveen Muralimanohar Rajeev Balasubramonian and Norm Jouppi. 2007. Optimizing NUCA Organizations and Wiring Alternatives for Large Caches with CACTI 6.0. In MICRO.","DOI":"10.1109\/MICRO.2007.33"},{"key":"e_1_2_1_69_1","unstructured":"Maxim Naumov Dheevatsa Mudigere Hao-Jun Michael Shi Jianyu Huang Narayanan Sundaraman Jongsoo Park Xiaodong Wang Udit Gupta Carole-Jean Wu Alisson G. Azzolini Dmytro Dzhulgakov Andrey Mallevich Ilia Cherniavskii Yinghai Lu Raghuraman Krishnamoorthi Ansha Yu Volodymyr Kondratenko Stephanie Pereira Xianjie Chen Wenlin Chen Vijay Rao Bill Jia Liang Xiong and Misha Smelyanskiy. 2019. Deep Learning Recommendation Model for Personalization and Recommendation Systems. In arXiv."},{"key":"e_1_2_1_70_1","volume-title":"CABLE: A CAche-Based Link Encoder for Bandwidth-Starved Manycores. In MICRO.","author":"Nguyen Tri M.","year":"2018","unstructured":"Tri M. Nguyen, Adi Fuchs, and David Wentzlaff. 2018. CABLE: A CAche-Based Link Encoder for Bandwidth-Starved Manycores. In MICRO."},{"key":"e_1_2_1_71_1","volume-title":"Nguyen and David Wentzlaff","author":"Tri","year":"2015","unstructured":"Tri M. Nguyen and David Wentzlaff. 2015. MORC: A Manycore-Oriented Compressed Cache. In MICRO."},{"key":"e_1_2_1_72_1","doi-asserted-by":"crossref","unstructured":"Vlad Nitu Boris Teabe Alain Tchana Canturk Isci and Daniel Hagimont. 2018. Welcome to Zombieland: Practical and Energy-Efficient Memory Disaggregation in a Datacenter. In EuroSys.","DOI":"10.1145\/3190508.3190537"},{"key":"e_1_2_1_73_1","volume-title":"Jung Ho Ahn, and G. Edward Suh","author":"Park Sungbo","year":"2021","unstructured":"Sungbo Park, Ingab Kang, Yaebin Moon, Jung Ho Ahn, and G. Edward Suh. 2021. BCD Deduplication: Effective Memory Compression Using Partial Cache-Line Deduplication. In ASPLOS."},{"key":"e_1_2_1_74_1","volume-title":"Mowry","author":"Pekhimenko Gennady","year":"2013","unstructured":"Gennady Pekhimenko, Vivek Seshadri, Yoongu Kim, Hongyi Xin, Onur Mutlu, Phillip B. Gibbons, Michael A. Kozuch, and Todd C. Mowry. 2013. Linearly Compressed Pages: A Low-Complexity, Low-Latency Main Memory Compression Framework. In MICRO."},{"key":"e_1_2_1_75_1","volume-title":"Mowry","author":"Pekhimenko Gennady","year":"2012","unstructured":"Gennady Pekhimenko, Vivek Seshadri, Onur Mutlu, Phillip B. Gibbons, Michael A. Kozuch, and Todd C. Mowry. 2012. Base-Delta-Immediate Compression: Practical Data Compression for on-Chip Caches. In PACT."},{"key":"e_1_2_1_76_1","doi-asserted-by":"crossref","unstructured":"Ivy Peng Roger Pearce and Maya Gokhale. 2020. On the Memory Underutilization: Exploring Disaggregated Memory on HPC Systems. In SBAC-PAD.","DOI":"10.1109\/SBAC-PAD49847.2020.00034"},{"key":"e_1_2_1_77_1","doi-asserted-by":"crossref","unstructured":"Christian Pinto Dimitris Syrivelis Michele Gazzetti Panos Koutsovasilis Andrea Reale Kostas Katrinis and H. Peter Hofstee. 2020. ThymesisFlow: A Software-Defined HW\/SW co-Designed Interconnect Stack for Rack-Scale Memory Disaggregation. In MICRO.","DOI":"10.1109\/MICRO50266.2020.00075"},{"key":"e_1_2_1_78_1","volume-title":"Tullsen","author":"Prodromou Andreas","year":"2017","unstructured":"Andreas Prodromou, Mitesh Meswani, Nuwan Jayasena, Gabriel Loh, and Dean M. Tullsen. 2017. MemPod: A Clustered Architecture for Efficient and Scalable Migration in Flat Address Space Multi-level Memories. In HPCA."},{"key":"e_1_2_1_79_1","doi-asserted-by":"publisher","DOI":"10.1145\/3203217.3203235"},{"key":"e_1_2_1_80_1","unstructured":"Pramod Subba Rao and George Porter. 2016. Is Memory Disaggregation Feasible? A Case Study with Spark SQL. In ANCS."},{"key":"e_1_2_1_81_1","unstructured":"RDMA. 2019. RDMA Consortium. http:\/\/www.rdmaconsortium.org\/"},{"key":"e_1_2_1_82_1","unstructured":"RDMA. 2022. Gen-Z Core Specification. https:\/\/genzconsortium.org\/"},{"key":"e_1_2_1_83_1","unstructured":"Joseph Redmon. 2013--2016. Darknet: Open Source Neural Networks in C. http:\/\/pjreddie.com\/darknet\/"},{"key":"e_1_2_1_84_1","volume-title":"AIFM: High-Performance, Application-Integrated Far Memory. In OSDI.","author":"Ruan Zhenyuan","year":"2020","unstructured":"Zhenyuan Ruan, Malte Schwarzkopf, Marcos K. Aguilera, and Adam Belay. 2020. AIFM: High-Performance, Application-Integrated Far Memory. In OSDI."},{"key":"e_1_2_1_85_1","volume-title":"John","author":"Ryoo Jee Ho","year":"2017","unstructured":"Jee Ho Ryoo, Mitesh R. Meswani, Andreas Prodromou, and Lizy K. John. 2017. SILC-FM: Subblocked InterLeaved Cache-Like Flat Memory Organization. In HPCA."},{"key":"e_1_2_1_86_1","doi-asserted-by":"crossref","unstructured":"Amedeo Sapio Ibrahim Abdelaziz Abdulla Aldilaijan Marco Canini and Panos Kalnis. 2017. In-Network Computation is a Dumb Idea Whose Time Has Come. In HotNets.","DOI":"10.1145\/3152434.3152461"},{"key":"e_1_2_1_87_1","doi-asserted-by":"crossref","unstructured":"Vijay Sathish Michael J. Schulte and Nam Sung Kim. 2012. Lossless and Lossy Memory I\/O Link Compression for Improving Performance of GPGPU Workloads. In PACT.","DOI":"10.1145\/2370816.2370864"},{"key":"e_1_2_1_88_1","doi-asserted-by":"crossref","unstructured":"Ali Shafiee Meysam Taassori Rajeev Balasubramonian and Al Davis. 2014. MemZip: Exploring Unconventional Benefits from Memory Compression. In HPCA.","DOI":"10.1109\/HPCA.2014.6835972"},{"key":"e_1_2_1_89_1","unstructured":"Yizhou Shan Yutong Huang Yilun Chen and Yiying Zhang. 2018. LegoOS: A Disseminated Distributed OS for Hardware Resource Disaggregation. In OSDI."},{"key":"e_1_2_1_90_1","volume-title":"Blelloch","author":"Shun Julian","year":"2013","unstructured":"Julian Shun and Guy E. Blelloch. 2013. Ligra: A Lightweight Graph Processing Framework for Shared Memory. In PpopP."},{"key":"e_1_2_1_91_1","volume-title":"Sibyl: Adaptive and Extensible Data Placement in Hybrid Storage Systems Using Online Reinforcement Learning. In ISCA.","author":"Singh Gagandeep","year":"2022","unstructured":"Gagandeep Singh, Rakesh Nadig, Jisung Park, Rahul Bera, Nastaran Hajinazar, David Novo, Juan G\u00f3mez-Luna, Sander Stuijk, Henk Corporaal, and Onur Mutlu. 2022. Sibyl: Adaptive and Extensible Data Placement in Hybrid Storage Systems Using Online Reinforcement Learning. In ISCA."},{"key":"e_1_2_1_92_1","volume-title":"Memory-Link Compression Schemes: A Value Locality Perspective","author":"Thuresson Martin","year":"2008","unstructured":"Martin Thuresson, Lawrence Spracklen, and Per Stenstrom. 2008. Memory-Link Compression Schemes: A Value Locality Perspective. IEEE Trans. Comput. (2008)."},{"key":"e_1_2_1_93_1","doi-asserted-by":"crossref","unstructured":"Martin Thuresson and Per Stenstr\u00f6 m. 2008. Accommodation of the Bandwidth of Large Cache Blocks Using Cache\/Memory Link Compression. In ICPP.","DOI":"10.1109\/ICPP.2008.47"},{"key":"e_1_2_1_94_1","volume-title":"Loh","author":"Tian Yingying","year":"2014","unstructured":"Yingying Tian, Samira M. Khan, Daniel A. Jim\u00e9nez, and Gabriel H. Loh. 2014. Last-Level Cache Deduplication. In ICS."},{"key":"e_1_2_1_95_1","volume-title":"Pinnacle: IBM MXT in a Memory Controller Chip","author":"Tremaine R.B.","year":"2001","unstructured":"R.B. Tremaine, T.B. Smith, M. Wazlowski, D. Har, Kwok-Ken Mak, and S. Arramreddy. 2001. Pinnacle: IBM MXT in a Memory Controller Chip. IEEE Micro (2001)."},{"key":"e_1_2_1_96_1","doi-asserted-by":"crossref","unstructured":"Shin-Yeh Tsai and Yiying Zhang. 2017. LITE Kernel RDMA Support for Datacenter Applications. In SOSP.","DOI":"10.1145\/3132747.3132762"},{"key":"e_1_2_1_97_1","unstructured":"Irina Chihaia Tuduce and Thomas Gross. 2005. Adaptive Main Memory Compression. In ATC."},{"key":"e_1_2_1_98_1","doi-asserted-by":"crossref","unstructured":"Evangelos Vasilakis Vassilis Papaefstathiou Pedro Trancoso and Ioannis Sourdis. 2019. LLC-Guided Data Migration in Hybrid Memory Systems. In IPDPS.","DOI":"10.1109\/IPDPS.2019.00101"},{"key":"e_1_2_1_99_1","doi-asserted-by":"crossref","unstructured":"Vasilakis Evangelos and Papaefstathiou Vassilis and Trancoso Pedro and Sourdis Ioannis. 2020. Hybrid2: Combining Caching and Migration in Hybrid Memory Systems. In HPCA.","DOI":"10.1109\/HPCA47549.2020.00059"},{"key":"e_1_2_1_100_1","doi-asserted-by":"crossref","unstructured":"Nandita Vijaykumar Gennady Pekhimenko Adwait Jog Abhishek Bhowmick Rachata Ausavarungnirun Chita Das Mahmut Kandemir Todd C. Mowry and Onur Mutlu. 2015. A Case for Core-Assisted Bottleneck Acceleration in GPUs: Enabling Flexible Data Compression with Assist Warps. In ISCA.","DOI":"10.1145\/2749469.2750399"},{"key":"e_1_2_1_101_1","volume-title":"Semeru: A Memory-Disaggregated Managed Runtime. In OSDI.","author":"Wang Chenxi","year":"2020","unstructured":"Chenxi Wang, Haoran Ma, Shi Liu, Yuanqi Li, Zhenyuan Ruan, Khanh Nguyen, Michael D. Bond, Ravi Netravali, Miryung Kim, and Guoqing Harry Xu. 2020. Semeru: A Memory-Disaggregated Managed Runtime. In OSDI."},{"key":"e_1_2_1_102_1","volume-title":"TMO: Transparent Memory Offloading in Datacenters. In ASPLOS.","author":"Weiner Johannes","year":"2022","unstructured":"Johannes Weiner, Niket Agarwal, Dan Schatzberg, Leon Yang, Hao Wang, Blaise Sanouillet, Bikash Sharma, Tejun Heo, Mayank Jain, Chunqiang Tang, and Dimitrios Skarlatos. 2022. TMO: Transparent Memory Offloading in Datacenters. In ASPLOS."},{"key":"e_1_2_1_103_1","unstructured":"P. Wilson Scott F. Kaplan and Y. Smaragdakis. 1999. The Case for Compressed Caching in Virtual Memory Systems. In ATC."},{"key":"e_1_2_1_104_1","volume-title":"Dean L. Lewis, and Hsien-Hsin Sean Lee.","author":"Woo Dong Hyuk","year":"2010","unstructured":"Dong Hyuk Woo, Nak Hee Seong, Dean L. Lewis, and Hsien-Hsin Sean Lee. 2010. An Optimized 3D-stacked Memory Architecture by Exploiting Excessive, High-Density TSV Bandwidth. HPCA (2010)."},{"key":"e_1_2_1_105_1","doi-asserted-by":"crossref","unstructured":"Zi Yan Daniel Lustig David Nellans and Abhishek Bhattacharjee. 2019. Nimble Page Management for Tiered Memory Systems. In ASPLOS.","DOI":"10.1145\/3297858.3304024"},{"key":"e_1_2_1_106_1","volume-title":"Frequent Value Encoding for Low Power Data Buses. TODAES","author":"Yang Jun","year":"2004","unstructured":"Jun Yang, Rajiv Gupta, and Chuanjun Zhang. 2004. Frequent Value Encoding for Low Power Data Buses. TODAES (2004)."},{"key":"e_1_2_1_107_1","doi-asserted-by":"crossref","unstructured":"Jun Yang Youtao Zhang and R. Gupta. 2000. Frequent Value Compression in Data Caches. In MICRO.","DOI":"10.1145\/360128.360154"},{"key":"e_1_2_1_108_1","volume-title":"Diego Furtado Silva, Abdullah Mueen, and Eamonn Keogh.","author":"Michael Yeh Chin-Chia","year":"2016","unstructured":"Chin-Chia Michael Yeh, Yan Zhu, Liudmila Ulanova, Nurjahan Begum, Yifei Ding, Hoang Anh Dau, Diego Furtado Silva, Abdullah Mueen, and Eamonn Keogh. 2016. Matrix Profile I: All Pairs Similarity Joins for Time Series: A Unifying View that Includes Motifs, Discords and Shapelets. In ICDM."},{"key":"e_1_2_1_109_1","volume-title":"Qureshi","author":"Young Vinson","year":"2019","unstructured":"Vinson Young, Sanjay Kariyappa, and Moinuddin K. Qureshi. 2019. Enabling Transparent Memory-Compression for Commodity Memory Systems. In HPCA."},{"key":"e_1_2_1_110_1","volume-title":"Optically Disaggregated Data Centers with Minimal Remote Memory Latency: Technologies, Architectures, and Resource Allocation. JOCN","author":"Zervas Georgios","year":"2018","unstructured":"Georgios Zervas, Hui Yuan, Arsalan Saljoghei, Qianqiao Chen, and Vaibhawa Mishra. 2018. Optically Disaggregated Data Centers with Minimal Remote Memory Latency: Technologies, Architectures, and Resource Allocation. JOCN (2018)."},{"key":"e_1_2_1_111_1","unstructured":"Qizhen Zhang Yifan Cai Sebastian G. Angel Vincent Liu Ang Chen and B. T. Loo. 2020. Rethinking Data Management Systems for Disaggregated Data Centers. In CIDR."},{"key":"e_1_2_1_112_1","volume-title":"Carbink: Fault-Tolerant Far Memory. In OSDI.","author":"Zhou Yang","year":"2022","unstructured":"Yang Zhou, Hassan M. G. Wassel, Sihang Liu, Jiaqi Gao, James Mickens, Minlan Yu, Chris Kennelly, Paul Turner, David E. Culler, Henry M. Levy, and Amin Vahdat. 2022. Carbink: Fault-Tolerant Far Memory. In OSDI."},{"key":"e_1_2_1_113_1","doi-asserted-by":"crossref","unstructured":"J. Ziv and A. Lempel. 1977. A Universal Algorithm for Sequential Data Compression. IEEE Transactions on Information Theory (1977).","DOI":"10.1109\/TIT.1977.1055714"},{"key":"e_1_2_1_114_1","unstructured":"Pengfei Zuo Jiazhao Sun Liu Yang Shuangwu Zhang and Yu Hua. 2021. One-sided RDMA-Conscious Extendible Hashing for Disaggregated Memory. In ATC."}],"container-title":["Proceedings of the ACM on Measurement and Analysis of Computing Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3579445","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3579445","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T16:46:41Z","timestamp":1750178801000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3579445"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,2,27]]},"references-count":114,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2023,2,27]]}},"alternative-id":["10.1145\/3579445"],"URL":"https:\/\/doi.org\/10.1145\/3579445","relation":{},"ISSN":["2476-1249"],"issn-type":[{"type":"electronic","value":"2476-1249"}],"subject":[],"published":{"date-parts":[[2023,2,27]]},"assertion":[{"value":"2023-03-02","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}