{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,11,3]],"date-time":"2025-11-03T14:21:07Z","timestamp":1762179667673,"version":"build-2065373602"},"reference-count":56,"publisher":"Association for Computing Machinery (ACM)","issue":"4","funder":[{"DOI":"10.13039\/100004358","name":"Samsung Electronics Co., Ltd","doi-asserted-by":"crossref","award":["IO230419-05997-01"],"award-info":[{"award-number":["IO230419-05997-01"]}],"id":[{"id":"10.13039\/100004358","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Storage"],"published-print":{"date-parts":[[2025,11,30]]},"abstract":"<jats:p>DRAM accounts for a large fraction of the total cost of ownership of memory systems in deep learning acceleration systems. To achieve sustainable scalability, tiered memory systems with denser technologies become critical. Prior work has proposed various tiered memory systems, but no matter how a system is designed, data movements between tiers consume substantial energy. In particular, as model sizes and memory capacity demands grow, the data movements between memory tiers become more frequent, posing a challenge that tiered memory systems may reduce deployment costs but suffer from low energy efficiency. If a memory system proactively places data into a tier and timely fetches it from the tier, then the excessive data movement between tiers can be mitigated. We find that a system can statically anticipate such behaviors for all pages and localities. With this insight, we propose a new DNN acceleration system called TM-Training using flash memory. TM-Training capitalizes on the repetitive nature of the same computational patterns during execution, enabling the static establishment of optimal data placement for subsequent operations. Moreover, TM-Training employs a new data-splitting scheme to enable precise memory management. Our evaluation demonstrates that TM-Training reduces inter-tier data traffic by 64% and achieves a 55% higher throughput per watt in training than prior work.<\/jats:p>","DOI":"10.1145\/3721484","type":"journal-article","created":{"date-parts":[[2025,3,5]],"date-time":"2025-03-05T05:11:38Z","timestamp":1741151498000},"page":"1-26","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["TM-Training: An Energy-Efficient Tiered Memory System for Deep Learning Training in NPUs"],"prefix":"10.1145","volume":"21","author":[{"ORCID":"https:\/\/orcid.org\/0009-0001-4373-3963","authenticated-orcid":false,"given":"Jaeyong","family":"Park","sequence":"first","affiliation":[{"name":"School of Electrical Engineering, Korea University","place":["Seongbuk-gu, Korea (the Republic of)"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0001-5803-361X","authenticated-orcid":false,"given":"Sangun","family":"Choi","sequence":"additional","affiliation":[{"name":"School of Electrical Engineering, Korea University","place":["Seoul, Korea (the Republic of)"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-0486-5493","authenticated-orcid":false,"given":"Jongmin","family":"Kim","sequence":"additional","affiliation":[{"name":"School of Electrical Engineering, Korea University","place":["Seoul, Korea (the Republic of)"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1706-6850","authenticated-orcid":false,"given":"GunJae","family":"Koo","sequence":"additional","affiliation":[{"name":"Computer Science and Engineering, Korea University","place":["Seongbuk-gu, Korea (the Republic of)"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9332-0251","authenticated-orcid":false,"given":"Myung Kuk","family":"Yoon","sequence":"additional","affiliation":[{"name":"Department of Computer Science and Engineering, Ewha Womans University","place":["Seoul, Korea (the Republic of)"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6442-3705","authenticated-orcid":false,"given":"Yunho","family":"Oh","sequence":"additional","affiliation":[{"name":"School of Electrical Engineering, Korea University","place":["Seoul, Korea (the Republic of)"]}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2025,11,3]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.1145\/3297858.3304061"},{"key":"e_1_3_1_3_2","volume-title":"8th USENIX Symposium on Networked Systems Design and Implementation (NSDI\u201911)","author":"Badam Anirudh","year":"2011","unstructured":"Anirudh Badam and Vivek S. Pai. 2011. SSDAlloc: Hybrid SSD\/RAM memory management made easy. In 8th USENIX Symposium on Networked Systems Design and Implementation (NSDI\u201911)."},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA.2018.00043"},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.1145\/3442188.3445922"},{"key":"e_1_3_1_6_2","first-page":"1877","article-title":"Language models are few-shot learners","volume":"33","author":"Brown Tom","year":"2020","unstructured":"Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D. Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell et\u00a0al. 2020. Language models are few-shot learners. Advan. Neural Inf. Process. Syst. 33 (2020), 1877\u20131901.","journal-title":"Advan. Neural Inf. Process. Syst."},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISPASS.2013.6557141"},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISPASS.2014.6844456"},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1145\/3132402.3132404"},{"key":"e_1_3_1_10_2","article-title":"PaLM: Scaling language modeling with pathways","author":"Chowdhery Aakanksha","year":"2022","unstructured":"Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann et\u00a0al. 2022. PaLM: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311 (2022).","journal-title":"arXiv preprint arXiv:2204.02311"},{"key":"e_1_3_1_11_2","article-title":"BERT: Pre-training of deep bidirectional transformers for language understanding","author":"Devlin Jacob","year":"2018","unstructured":"Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. BERT: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805 (2018).","journal-title":"arXiv preprint arXiv:1810.04805"},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPSW.2019.00080"},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1145\/3190508.3190524"},{"key":"e_1_3_1_14_2","first-page":"40","article-title":"Bandana: Using non-volatile memory for storing deep learning models","volume":"1","author":"Eisenman Assaf","year":"2019","unstructured":"Assaf Eisenman, Maxim Naumov, Darryl Gardner, Misha Smelyanskiy, Sergey Pupyrev, Kim Hazelwood, Asaf Cidon, and Sachin Katti. 2019. Bandana: Using non-volatile memory for storing deep learning models. Proc. Mach. Learn. Syst. 1 (2019), 40\u201352.","journal-title":"Proc. Mach. Learn. Syst."},{"key":"e_1_3_1_15_2","article-title":"Samsung V-NAND SSD 990 PRO","author":"Electronics Samsung","year":"2022","unstructured":"Samsung Electronics. 2022. Samsung V-NAND SSD 990 PRO. Retrieved from https:\/\/download.semiconductor.samsung.com\/resources\/data-sheet\/Samsung_NVMe_SSD_990_PRO_Datasheet_Rev.1.0.pdf","journal-title":"R"},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA56546.2023.10070955"},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1109\/TC.2019.2949408"},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2017.19"},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1145\/3373376.3378465"},{"key":"e_1_3_1_21_2","article-title":"Training compute-optimal large language models","author":"Hoffmann Jordan","year":"2022","unstructured":"Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark et\u00a0al. 2022. Training compute-optimal large language models. arXiv preprint arXiv:2203.15556 (2022).","journal-title":"arXiv preprint arXiv:2203.15556"},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.1145\/3373376.3378494"},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.1145\/3357526.3357561"},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2014.51"},{"key":"e_1_3_1_25_2","first-page":"1","volume-title":"International Symposium on Computer Architecture (ISCA\u201914)","author":"Jin Youngbin","year":"2014","unstructured":"Youngbin Jin, Mustafa Shihab, and Myoungsoo Jung. 2014. Area, power, and latency considerations of STT-MRAM to substitute for main memory. In International Symposium on Computer Architecture (ISCA\u201914). 1\u20134."},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.1145\/3579371.3589350"},{"key":"e_1_3_1_27_2","first-page":"1","volume-title":"ACM\/IEEE 48th Annual International Symposium on Computer Architecture (ISCA\u201921)","author":"Jouppi Norman P.","year":"2021","unstructured":"Norman P. Jouppi, Doe Hyun Yoon, Matthew Ashcraft, Mark Gottscho, Thomas B. Jablin, George Kurian, James Laudon, Sheng Li, Peter Ma, Xiaoyu Ma et\u00a0al. 2021. Ten lessons from three generations shaped Google\u2019s TPUv4i: Industrial product. In ACM\/IEEE 48th Annual International Symposium on Computer Architecture (ISCA\u201921). IEEE, 1\u201314."},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA52012.2021.00059"},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.1145\/3079856.3080245"},{"key":"e_1_3_1_30_2","article-title":"MTrainS: Improving DLRM training efficiency using heterogeneous memories","author":"Kassa Hiwot Tadese","year":"2023","unstructured":"Hiwot Tadese Kassa, Paul Johnson, Jason Akers, Mrinmoy Ghosh, Andrew Tulloch, Dheevatsa Mudigere, Jongsoo Park, Xing Liu, Ronald Dreslinski, and Ehsan K. Ardestani. 2023. MTrainS: Improving DLRM training efficiency using heterogeneous memories. arXiv preprint arXiv:2305.01515 (2023).","journal-title":"arXiv preprint arXiv:2305.01515"},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.1145\/3564695.3564777"},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISPASS.2013.6557176"},{"key":"e_1_3_1_33_2","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2020.2985963"},{"key":"e_1_3_1_34_2","doi-asserted-by":"publisher","DOI":"10.1145\/3296957.3173176"},{"key":"e_1_3_1_35_2","doi-asserted-by":"publisher","DOI":"10.1109\/TC.2020.2973134"},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2019.2908175"},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.1145\/2383276.2383284"},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","DOI":"10.23919\/DATE51398.2021.9474252"},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE53745.2022.00241"},{"key":"e_1_3_1_40_2","doi-asserted-by":"publisher","DOI":"10.1145\/3123939.3124545"},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","DOI":"10.1145\/1555754.1555760"},{"key":"e_1_3_1_42_2","doi-asserted-by":"publisher","DOI":"10.5555\/3455716.3455856"},{"key":"e_1_3_1_43_2","doi-asserted-by":"publisher","DOI":"10.1145\/1995896.1995911"},{"key":"e_1_3_1_44_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA51647.2021.00057"},{"key":"e_1_3_1_45_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISPASS48437.2020.00016"},{"key":"e_1_3_1_46_2","article-title":"990 PRO PCIe 4.0 NVMe SSD 1TB","year":"2024","unstructured":"Samsung. 2024. 990 PRO PCIe 4.0 NVMe SSD 1TB. Retrieved from https:\/\/www.samsung.com\/us\/computing\/memory-storage\/solid-state-drives\/990-pro-pcie--4-0-nvme--ssd-1tb-mz-v9p1t0b-am\/","journal-title":"R"},{"key":"e_1_3_1_47_2","article-title":"Report: DDR5 RDIMM production impacted by PMIC compatibility issues","author":"Shilov Anton","year":"2023","unstructured":"Anton Shilov. 2023. Report: DDR5 RDIMM production impacted by PMIC compatibility issues. Retrieved from https:\/\/www.anandtech.com\/show\/18830\/report-ddr5-rdimm-production-impacted-by-pmic-compatibility-issues","journal-title":"R"},{"key":"e_1_3_1_48_2","article-title":"Explosive HBM demand fueling an expected 20% increase in DDR5 memory pricing - Demand for AI GPUs drives production cuts for standard PC memory","author":"Shilov Anton","year":"2024","unstructured":"Anton Shilov. 2024. Explosive HBM demand fueling an expected 20% increase in DDR5 memory pricing - Demand for AI GPUs drives production cuts for standard PC memory. Retrieved from https:\/\/www.tomshardware.com\/pc-components\/gpus\/explosive-hbm-demand-fueling-an-expected-20-increase-in-ddr5-memory-pricing-demand-for-ai-gpus-drives-production-cuts-for-standard-pc-memory","journal-title":"R"},{"key":"e_1_3_1_49_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298594"},{"key":"e_1_3_1_50_2","article-title":"Attention is all you need","volume":"30","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advan. Neural Inf. Process. Syst. 30 (2017), 6000\u20136010.","journal-title":"Advan. Neural Inf. Process. Syst."},{"key":"e_1_3_1_51_2","doi-asserted-by":"publisher","DOI":"10.1145\/3643641"},{"key":"e_1_3_1_52_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISPASS.2006.1620785"},{"key":"e_1_3_1_53_2","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2006.79"},{"key":"e_1_3_1_54_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCAD45719.2019.8942149"},{"key":"e_1_3_1_55_2","doi-asserted-by":"publisher","DOI":"10.1145\/859618.859629"},{"key":"e_1_3_1_56_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA.2018.00036"},{"key":"e_1_3_1_57_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2024.3431910"}],"container-title":["ACM Transactions on Storage"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3721484","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,11,3]],"date-time":"2025-11-03T13:32:29Z","timestamp":1762176749000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3721484"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,11,3]]},"references-count":56,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2025,11,30]]}},"alternative-id":["10.1145\/3721484"],"URL":"https:\/\/doi.org\/10.1145\/3721484","relation":{},"ISSN":["1553-3077","1553-3093"],"issn-type":[{"type":"print","value":"1553-3077"},{"type":"electronic","value":"1553-3093"}],"subject":[],"published":{"date-parts":[[2025,11,3]]},"assertion":[{"value":"2024-07-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-01-31","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-11-03","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}