{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,16]],"date-time":"2026-06-16T15:10:00Z","timestamp":1781622600813,"version":"3.54.5"},"publisher-location":"New York, NY, USA","reference-count":60,"publisher":"ACM","license":[{"start":{"date-parts":[[2025,5,14]],"date-time":"2025-05-14T00:00:00Z","timestamp":1747180800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by-nc-nd\/4.0\/"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2025,5,14]]},"DOI":"10.1145\/3713082.3730381","type":"proceedings-article","created":{"date-parts":[[2025,6,6]],"date-time":"2025-06-06T09:53:51Z","timestamp":1749203631000},"page":"111-118","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":3,"title":["Storage Class Memory is Dead, All Hail Managed-Retention Memory: Rethinking Memory for the AI Era"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0009-0001-5596-8962","authenticated-orcid":false,"given":"Sergey","family":"Legtchenko","sequence":"first","affiliation":[{"name":"Microsoft Research"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-2347-2715","authenticated-orcid":false,"given":"Ioan","family":"Stefanovici","sequence":"additional","affiliation":[{"name":"Microsoft Research"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0003-7032-8458","authenticated-orcid":false,"given":"Richard","family":"Black","sequence":"additional","affiliation":[{"name":"Microsoft Research"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0009-5936-6895","authenticated-orcid":false,"given":"Antony","family":"Rowstron","sequence":"additional","affiliation":[{"name":"Microsoft Research"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4277-1802","authenticated-orcid":false,"given":"Junyi","family":"Liu","sequence":"additional","affiliation":[{"name":"Microsoft Research"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1939-5690","authenticated-orcid":false,"given":"Paolo","family":"Costa","sequence":"additional","affiliation":[{"name":"Microsoft Research"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8406-7768","authenticated-orcid":false,"given":"Burcu","family":"Canakci","sequence":"additional","affiliation":[{"name":"Microsoft Research"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0000-1194-2958","authenticated-orcid":false,"given":"Dushyanth","family":"Narayanan","sequence":"additional","affiliation":[{"name":"Microsoft Research"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1649-2612","authenticated-orcid":false,"given":"Xingbo","family":"Wu","sequence":"additional","affiliation":[{"name":"Microsoft Research"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,6,6]]},"reference":[{"key":"e_1_3_2_1_1_1","unstructured":"2025. Next-generation memory for computers. https:\/\/www.intrinsicsemi.com\/."},{"key":"e_1_3_2_1_2_1","unstructured":"2025. The ReRAM Market Opportunity. https:\/\/www.weebit-nano.com\/market\/market-overview\/."},{"key":"e_1_3_2_1_3_1","volume-title":"SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills. arXiv:2308.16369 [cs.LG] https:\/\/arxiv.org\/abs\/2308.16369","author":"Agrawal Amey","year":"2023","unstructured":"Amey Agrawal, Ashish Panwar, Jayashree Mohan, Nipun Kwatra, Bhargav S. Gulavani, and Ramachandran Ramjee. 2023. SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills. arXiv:2308.16369 [cs.LG] https:\/\/arxiv.org\/abs\/2308.16369"},{"key":"e_1_3_2_1_4_1","unstructured":"Artificial-Fintelligence 2023. Transformer inference tricks. https:\/\/www.artfintel.com\/p\/transformer-inference-tricks."},{"key":"e_1_3_2_1_5_1","unstructured":"blocksandfiles.com 2019. Is Optane DIMM endurance good enough? Quick answer ... Yes Intel has delivered. https:\/\/blocksandfiles.com\/2019\/04\/04\/enduring-optane-dimm-question-is-its-endurance-good-enough-yes-intel-has-delivered\/."},{"key":"e_1_3_2_1_6_1","unstructured":"blocksandfiles.com 2022. CrossBar tries to secure embedded ReRAM IoT market. https:\/\/blocksandfiles.com\/2022\/04\/21\/no-sniffing-crossbar-tries-to-secure-its-embedded-reram-iot-market-niche\/."},{"key":"e_1_3_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/1278480.1278533"},{"key":"e_1_3_2_1_8_1","unstructured":"S. Dolinar Dariush Divsalar and F. Pollara. 1998. Code Performance as a Function of Block Size. Telecommunications and Mission Operations Progress Report (01 1998)."},{"key":"e_1_3_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISCAS.1995.519949"},{"key":"e_1_3_2_1_10_1","unstructured":"Keming Fan Wei-Chen Chen Sumukh Pinge H. S. Philip Wong and Tajana Rosing. 2024. Efficient Open Modification Spectral Library Searching in High-Dimensional Space with Multi-Level-Cell Memory. arXiv:2405.02756 [cs.AR] https:\/\/arxiv.org\/abs\/2405.02756"},{"key":"e_1_3_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/IEDM45625.2022.10019367"},{"key":"e_1_3_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2021.3076445"},{"key":"e_1_3_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1145\/2155620.2155642"},{"key":"e_1_3_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1109\/MUE.2008.61"},{"key":"e_1_3_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/IRPS.2010.5488761"},{"key":"e_1_3_2_1_16_1","unstructured":"Intel 2019. Intel Optane Memory - Responsive Memory Accelerated Performance. https:\/\/www.intel.com\/content\/www\/us\/en\/products\/details\/memory-storage\/optane-memory.html."},{"key":"e_1_3_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/1736020.1736023"},{"key":"e_1_3_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/2228360.2228406"},{"key":"e_1_3_2_1_19_1","unstructured":"Myoungsoo Jung Youngbin Jin and Mustafa Shihab. 2014. Area Power and Latency Considerations of STT-MRAM to Substitute for Main Memory."},{"key":"e_1_3_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2024.3375352"},{"key":"e_1_3_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1109\/HCS59251.2023.10254711"},{"key":"e_1_3_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/3600006.3613165"},{"key":"e_1_3_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1088\/1361-6641\/abf29d"},{"key":"e_1_3_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1145\/1555815.1555758"},{"key":"e_1_3_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1109\/IEDM.2010.5703395"},{"key":"e_1_3_2_1_26_1","doi-asserted-by":"publisher","unstructured":"Shuhan Liu Shengjun Qin Koustav Jana Jian Chen Kasidit Toprasertpong and H.-S. Philip Wong. 2024. First Experimental Demonstration of Hybrid Gain Cell Memory with Si PMOS and ITO FET for High-speed On-chip Memory. In 2024 IEEE Symposium on VLSI Technology and Circuits (VLSI Technology and Circuits). 1--2. https:\/\/doi.org\/10.1109\/VLSITechnologyandCir46783.2024.10631344","DOI":"10.1109\/VLSITechnologyandCir46783.2024.10631344"},{"key":"e_1_3_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/3651890.3672274"},{"key":"e_1_3_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1145\/3490391"},{"key":"e_1_3_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1109\/IEDM.1987.191485"},{"key":"e_1_3_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1186\/1556-276X-9-526"},{"key":"e_1_3_2_1_31_1","volume-title":"Alan Zhu, Lijie Yang, Xiaoxiang Shi, et al.","author":"Miao Xupeng","year":"2023","unstructured":"Xupeng Miao, Gabriele Oliaro, Zhihao Zhang, Xinhao Cheng, Zeyu Wang, Zhengxin Zhang, Rae Ying Yee Wong, Alan Zhu, Lijie Yang, Xiaoxiang Shi, et al. 2023. SpecInfer: Accelerating Generative Large Language Model Serving with Tree-based Speculative Inference and Verification. arXiv preprint arXiv:2305.09781 (2023)."},{"key":"e_1_3_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1109\/IMW52921.2022.9779293"},{"key":"e_1_3_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1109\/T-C.1974.223953"},{"key":"e_1_3_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1109\/IEDM.2016.7838346"},{"key":"e_1_3_2_1_35_1","unstructured":"nvidia.com 2024. The NVIDIA Blackwell Architecture. https:\/\/resources.nvidia.com\/en-us-blackwell-architecture?ncid=no-ncid."},{"key":"e_1_3_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1145\/1629911.1630085"},{"key":"e_1_3_2_1_37_1","volume-title":"Splitwise: Efficient generative LLM inference using phase splitting. In ISCA. https:\/\/www.microsoft.com\/en-us\/research\/publication\/splitwise-efficient-generative-llm-inference-using-phase-splitting\/","author":"Patel Pratyush","year":"2024","unstructured":"Pratyush Patel, Esha Choukse, Chaojie Zhang, Aashaka Shah, \u00cd\u00f1igo Goiri, Saeed Maleki, and Ricardo Bianchini. 2024. Splitwise: Efficient generative LLM inference using phase splitting. In ISCA. https:\/\/www.microsoft.com\/en-us\/research\/publication\/splitwise-efficient-generative-llm-inference-using-phase-splitting\/"},{"key":"e_1_3_2_1_38_1","volume-title":"Microsoft plans to invest $80 billion on AI-enabled data centers in fiscal","year":"2025","unstructured":"Reuters.com 2025. Microsoft plans to invest $80 billion on AI-enabled data centers in fiscal 2025. https:\/\/www.reuters.com\/technology\/artificial-intelligence\/microsoft-plans-spend-80-bln-ai-enabled-data-centers-fiscal-2025-cnbc-reports-2025-01-03\/."},{"key":"e_1_3_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.23919\/VLSIT.2017.7998174"},{"key":"e_1_3_2_1_40_1","unstructured":"SiliconMatter 2024. The Memory Wall and Its Implications. https:\/\/siliconmatter.substack.com\/p\/the-memory-wall-and-its-implications."},{"key":"e_1_3_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1145\/3656019.3676890"},{"key":"e_1_3_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA59077.2024.00068"},{"key":"e_1_3_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2011.5749716"},{"key":"e_1_3_2_1_44_1","unstructured":"Spheron 2024. How Much GPU Memory is Required to Run a Large Language Model? Find Out Here! https:\/\/blog.spheron.network\/how-much-gpu-memory-is-required-to-run-a-large-language-model-find-out-here."},{"key":"e_1_3_2_1_45_1","unstructured":"Stanford 2024. DAM: Differentiated Access Memory Systems and Applications. https:\/\/dam.stanford.edu\/assets\/Stanford_DAM_2_Pages_2024.pdfl."},{"key":"e_1_3_2_1_46_1","volume-title":"TAPAS: Thermal- and Power-Aware Scheduling for LLM Inference in Cloud Platforms. arXiv:2501.02600 [cs.DC] https:\/\/arxiv.org\/abs\/2501.02600","author":"Stojkovic Jovan","year":"2025","unstructured":"Jovan Stojkovic, Chaojie Zhang, \u00cd\u00f1igo Goiri, Esha Choukse, Haoran Qiu, Rodrigo Fonseca, Josep Torrellas, and Ricardo Bianchini. 2025. TAPAS: Thermal- and Power-Aware Scheduling for LLM Inference in Cloud Platforms. arXiv:2501.02600 [cs.DC] https:\/\/arxiv.org\/abs\/2501.02600"},{"key":"e_1_3_2_1_47_1","volume-title":"Exploring Memory Hierarchy Design with Emerging Memory Technologies","author":"Sun Guangyu","unstructured":"Guangyu Sun. 2013. Exploring Memory Hierarchy Design with Emerging Memory Technologies. Springer Publishing Company, Incorporated."},{"key":"e_1_3_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1145\/2155620.2155659"},{"key":"e_1_3_2_1_49_1","volume-title":"Exploring CXL-based KV Cache Storage for LLM Serving. Workshop on ML for Systems at NeurIPS 2024 (2024","author":"Tang Yupeng","year":"2024","unstructured":"Yupeng Tang, Runxiang Cheng, Ping Zhou, Tongping Liu, Fei Liu, Wei Tang, Kyoungryun Bae, Jianjun Chen, Wu Xiang, and Rui Shi. 2024. Exploring CXL-based KV Cache Storage for LLM Serving. Workshop on ML for Systems at NeurIPS 2024 (2024). https:\/\/mlforsystems.org\/assets\/papers\/neurips2024\/paper17.pdf"},{"key":"e_1_3_2_1_50_1","volume-title":"Micron Plans HBM4E","author":"TomsHardware","year":"2028","unstructured":"TomsHardware 2023. Micron Plans HBM4E in 2028. https:\/\/www.tomshardware.com\/pc-components\/ddr5\/micron-plans-hbm4e-in-2028-256gb-ddr5-12800-ram-sticks-in-2026."},{"key":"e_1_3_2_1_51_1","unstructured":"TomsHardware.com 2024. Nvidia's next-gen AI GPU is 4X faster than Hopper: Blackwell B200 GPU delivers up to 20 petaflops of compute and other massive improvements. https:\/\/www.tomshardware.com\/pc-components\/gpus\/nvidias-next-gen-ai-gpu-revealed-blackwell-b200-gpu-delivers-up-to-20-petaflops-of-compute-and-massive-improvements-over-hopper-h100."},{"key":"e_1_3_2_1_52_1","volume-title":"\u0141 ukasz Kaiser, and Illia Polosukhin","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, \u0141 ukasz Kaiser, and Illia Polosukhin. 2017. Attention is All you Need. In Advances in Neural Information Processing Systems, I. Guyon, U. Von Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett (Eds.), Vol. 30. Curran Associates, Inc. https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2017\/file\/3f5ee243547dee91fbd053c1c4a845aa-Paper.pdf"},{"key":"e_1_3_2_1_53_1","doi-asserted-by":"publisher","DOI":"10.1109\/MSSC.2019.2922889"},{"key":"e_1_3_2_1_54_1","unstructured":"vLLM.ai 2024. Automatic Prefix Caching. https:\/\/docs.vllm.ai\/en\/latest\/features\/automatic_prefix_caching.html."},{"key":"e_1_3_2_1_55_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICSICT55466.2022.9963306"},{"key":"e_1_3_2_1_56_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2015.7056056"},{"key":"e_1_3_2_1_57_1","volume-title":"ProTrain: Efficient LLM Training via Memory-Aware Techniques. arXiv preprint arXiv:2406.08334","author":"Yang Hanmei","year":"2024","unstructured":"Hanmei Yang, Jin Zhou, Yao Fu, Xiaoqun Wang, Ramine Roane, Hui Guan, and Tongping Liu. 2024. ProTrain: Efficient LLM Training via Memory-Aware Techniques. arXiv preprint arXiv:2406.08334 (2024)."},{"key":"e_1_3_2_1_58_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCSI.2021.3072200"},{"key":"e_1_3_2_1_59_1","unstructured":"Penghao Zhao Hailin Zhang Qinhan Yu Zhengren Wang Yunteng Geng Fangcheng Fu Ling Yang Wentao Zhang Jie Jiang and Bin Cui. 2024. Retrieval-Augmented Generation for AI-Generated Content: A Survey. arXiv:2402.19473 [cs.CV] https:\/\/arxiv.org\/abs\/2402.19473"},{"key":"e_1_3_2_1_60_1","unstructured":"zonedstorage.io 2019. SSDs with NVMe Zoned Namespace (ZNS) Support. https:\/\/zonedstorage.io\/docs\/introduction\/zns."}],"event":{"name":"HOTOS '25: Workshop on Hot Topics in Operating Systems","location":"Banff AB Canada","acronym":"HOTOS '25","sponsor":["SIGOPS ACM Special Interest Group on Operating Systems"]},"container-title":["Proceedings of the Workshop on Hot Topics in Operating Systems"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3713082.3730381","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3713082.3730381","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,8,29]],"date-time":"2025-08-29T16:49:15Z","timestamp":1756486155000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3713082.3730381"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,5,14]]},"references-count":60,"alternative-id":["10.1145\/3713082.3730381","10.1145\/3713082"],"URL":"https:\/\/doi.org\/10.1145\/3713082.3730381","relation":{},"subject":[],"published":{"date-parts":[[2025,5,14]]},"assertion":[{"value":"2025-06-06","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}