{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,6]],"date-time":"2026-07-06T18:41:15Z","timestamp":1783363275184,"version":"3.54.6"},"reference-count":34,"publisher":"MDPI AG","issue":"5","license":[{"start":{"date-parts":[[2022,5,10]],"date-time":"2022-05-10T00:00:00Z","timestamp":1652140800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"NSERC"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Future Internet"],"abstract":"<jats:p>From self-driving cars to detecting cancer, the applications of modern artificial intelligence (AI) rely primarily on deep neural networks (DNNs). Given raw sensory data, DNNs are able to extract high-level features after the network has been trained using statistical learning. However, due to the massive amounts of parallel processing in computations, the memory wall largely affects the performance. Thus, a review of the different memory architectures applied in DNN accelerators would prove beneficial. While the existing surveys only address DNN accelerators in general, this paper investigates novel advancements in efficient memory organizations and design methodologies in the DNN accelerator. First, an overview of the various memory architectures used in DNN accelerators will be provided, followed by a discussion of memory organizations on non-ASIC DNN accelerators. Furthermore, flexible memory systems incorporating an adaptable DNN computation will be explored. Lastly, an analysis of emerging memory technologies will be conducted. The reader, through this article, will: 1\u2014gain the ability to analyze various proposed memory architectures; 2\u2014discern various DNN accelerators with different memory designs; 3\u2014become familiar with the trade-offs associated with memory organizations; and 4\u2014become familiar with proposed new memory systems for modern DNN accelerators to solve the memory wall and other mentioned current issues.<\/jats:p>","DOI":"10.3390\/fi14050146","type":"journal-article","created":{"date-parts":[[2022,5,10]],"date-time":"2022-05-10T08:31:55Z","timestamp":1652171515000},"page":"146","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":18,"title":["A Survey on Memory Subsystems for Deep Neural Network Accelerators"],"prefix":"10.3390","volume":"14","author":[{"given":"Arghavan","family":"Asad","sequence":"first","affiliation":[{"name":"Electrical and Computer Engineering Department, Toronto Metropolitan University, 350 Victoria St, Toronto, ON M5B 2K3, Canada"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0065-6280","authenticated-orcid":false,"given":"Rupinder","family":"Kaur","sequence":"additional","affiliation":[{"name":"Electrical and Computer Engineering Department, Toronto Metropolitan University, 350 Victoria St, Toronto, ON M5B 2K3, Canada"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Farah","family":"Mohammadi","sequence":"additional","affiliation":[{"name":"Electrical and Computer Engineering Department, Toronto Metropolitan University, 350 Victoria St, Toronto, ON M5B 2K3, Canada"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2022,5,10]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Sze, V., Chen, Y.H., Yang, T.J., and Emer, J.S. (2017). Efficient processing of deep neural networks: A Tutorial and Survey. arXiv.","DOI":"10.1109\/JPROC.2017.2761740"},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"703","DOI":"10.1557\/mrs.2014.139","article-title":"Phase change materials and phase change memory","volume":"39","author":"Raoux","year":"2014","journal-title":"MRS Bull."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"76","DOI":"10.1016\/j.micpro.2017.03.011","article-title":"Optimization-based power and thermal management for dark silicon aware 3D chip multiprocessors using heterogeneous cache hierarchy","volume":"51","author":"Asad","year":"2017","journal-title":"Microprocess. Microsyst."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"380","DOI":"10.1145\/3007787.3001178","article-title":"Neurocube: A programmable digital neuromorphic architecture with high-density 3D memory","volume":"44","author":"Kim","year":"2016","journal-title":"ACM SIGARCH Comput. Archit. News"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Gao, M., Pu, J., Yang, X., Horowitz, M., and Kozyrakis, C. (2017, January 8\u201312). TETRIS: Scalable and Efficient Neural Network Acceleration with 3D Memory. Proceedings of the Twenty-Second International Conference on Architectural Support for Programming Languages and Operating Systems, Xi\u2019an, China.","DOI":"10.1145\/3037697.3037702"},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"852","DOI":"10.1109\/TC.2018.2889053","article-title":"Learning-based application-agnostic 3D NoC design for heterogeneous manycore systems","volume":"68","author":"Joardar","year":"2018","journal-title":"IEEE Trans. Comput."},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Firuzan, A., Modarressi, M., Daneshtalab, M., and Reshadi, M. (2018, January 4\u20135). Reconfigurable network-on-chip for 3D neural network accelerators. Proceedings of the 2018 Twelfth IEEE\/ACM International Symposium on Networks-on-Chip (NOCS), Torino, Italy.","DOI":"10.1109\/NOCS.2018.8512170"},{"key":"ref_8","unstructured":"Mohsen, I., Samragh, M., Gupta, S., Koushanfar, F., and Rosing, T. (2018). RAPIDNN: In-memory Deep Neural Network Acceleration Framework. arXiv."},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Kim, J.S., and Yang, J.S. (2019, January 2\u20136). DRIS-3: Deep neural network reliability improvement scheme in 3D die-stacked memory based on fault analysis. Proceedings of the 2019 56th ACM\/IEEE Design Automation Conference (DAC), New York, NY, USA.","DOI":"10.1145\/3316781.3317805"},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"101689","DOI":"10.1016\/j.sysarc.2019.101689","article-title":"A survey on modeling and improving reliability of DNN algorithms and accelerators","volume":"104","author":"Mittal","year":"2020","journal-title":"J. Syst. Archit."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"2360","DOI":"10.1109\/TCAD.2018.2858358","article-title":"DeepTrain: A programmable embedded platform for training deep neural networks","volume":"37","author":"Kim","year":"2018","journal-title":"IEEE Trans. Comput. Aided Des. Integr. Circuits Syst."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"186","DOI":"10.1109\/JSSC.2018.2871623","article-title":"QUEST: Multi-Purpose Log-Quantized DNN Inference Engine Stacked on 96-MB 3-D SRAM Using Inductive Coupling Technology in 40-nm CMOS","volume":"54","author":"Ueyoshi","year":"2019","journal-title":"IEEE J. Solid-State Circuits"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Imani, M., Gupta, S., Kim, Y., and Rosing, T. (2019, January 22\u201326). FloatPIM: In-Memory Acceleration of Deep Neural Network Training with High Precision. Proceedings of the 2019 ACM\/IEEE 46th Annual International Symposium on Computer Architecture (ISCA), Phoenix, AZ, USA.","DOI":"10.1145\/3307650.3322237"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Angizi, S., He, Z., and Fan, D. (2019, January 23). ParaPIM: A Parallel Processing-in-Memory Accelerator for Binary Weight Deep Neural Networks. Proceedings of the 24th Asia and South Pacific Design Automation Conference, Tokyo, Japan.","DOI":"10.1145\/3287624.3287644"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Angizi, S., He, Z., Fan, D., and Rakin, A.S. (2018, January 24\u201329). CMP-PIM: An Energy Efficient Comparator-based Processing-in-Memory Neural Network Accelerator. Proceedings of the 55th Annual Design Automation Conference, San Francisco, CA, USA.","DOI":"10.1145\/3195970.3196009"},{"key":"ref_16","unstructured":"Li, T., Zhong, J., Ji, L., Wu, W., and Zhang, C. (2018, January 27\u201331). Ease.ml: Towards multi-tenant resource sharing for machine learning workloads. Proceedings of the 44th International Conference on Very Large Data Bases Endowment, Rio de Janeiro, Brazil."},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Liu, J., Zhao, H., Ogleari, M.A., Li, D., and Zhao, J. (2018, January 20\u201324). Processing-in-memory for energy-efficient neural network training: A heterogeneous approach. Proceedings of the 2018 51st Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO), Fukuoka, Japan.","DOI":"10.1109\/MICRO.2018.00059"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Kim, H., Qian, C., Yoo, T., Kim, T.T., and Kim, B. (2019, January 6). A Bit-Precision Reconfigurable Digital In-Memory Computing Macro for Energy-Efficient Processing of Artificial Neural Networks. Proceedings of the 2019 International SoC Design Conference (ISOCC), Jeju Island, Korea.","DOI":"10.1109\/ISOCC47750.2019.9027679"},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"171","DOI":"10.1109\/LCA.2021.3126450","article-title":"Near-Data Processing in Memory Expander for DNN Acceleration on GPUs","volume":"20","author":"Ham","year":"2021","journal-title":"IEEE Comput. Archit. Lett."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Inci, A.F., Isgenc, M.M., and Marculescu, D. (2020, January 9\u201313). DeepNVM: A framework for modeling and analysis of non-volatile memory technologies for deep learning applications. Proceedings of the IEEE 2020 Design, Automation & Test in Europe Conference & Exhibition (DATE), Grenoble, France.","DOI":"10.23919\/DATE48585.2020.9116263"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Qiao, X., Cao, X., Yang, H., Song, L., and Li, H. (2018, January 24\u201329). AtomLayer: A Universal ReRAM-based CNN Accelerator with Atomic Layer Computation. Proceedings of the 55th Annual Design Automation Conference, San Francisco, CA, USA.","DOI":"10.1145\/3195970.3195998"},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"27","DOI":"10.1145\/3007787.3001140","article-title":"Prime: A novel processing-in-memory architecture for neural network computation in reram-based main memory","volume":"44","author":"Chi","year":"2016","journal-title":"ACM SIGARCH Comput. Archit. News"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Dai, G., Huang, T., Wang, Y., Yang, H., and Wawrzynek, J. (2019, January 23). Graphsar: A sparsity-aware processing-in-memory architecture for large-scale graph processing on rerams. Proceedings of the 24th Asia and South Pacific Design Automation Conference, Tokyo, Japan.","DOI":"10.1145\/3287624.3287637"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Lin, J., Zhu, Z., Wang, Y., and Xie, Y. (2019, January 23). Learning the sparsity for ReRAM: Mapping and pruning sparse neural network for ReRAM based accelerator. Proceedings of the 24th Asia and South Pacific Design Automation Conference, Tokyo, Japan.","DOI":"10.1145\/3287624.3287715"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Ji, Y., Zhang, Y., Xie, X., Li, S., Wang, P., Hu, X., Zhang, Y., and Xie, Y. (2019, January 13\u201317). Fpsa: A full system stack solution for reconfigurable reram-based nn accelerator architecture. Proceedings of the Twenty-Fourth International Conference on Architectural Support for Programming Languages and Operating Systems, Providence, RI, USA.","DOI":"10.1145\/3297858.3304048"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Song, L., Qian, X., Li, H., and Chen, Y. (2017, January 4\u20138). Pipelayer: A Pipelined ReRAM-Based Accelerator for Deep Learning. Proceedings of the IEEE International Symposium on High Performance Computer Architecture, Austin, TX, USA.","DOI":"10.1109\/HPCA.2017.55"},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3392080","article-title":"3D-ReG: A 3D ReRAM-based heterogeneous architecture for training deep neural networks","volume":"16","author":"Li","year":"2020","journal-title":"ACM J. Emerg. Technol. Comput. Syst."},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"4736","DOI":"10.1109\/TCAD.2020.2981055","article-title":"RED: A ReRAM based Efficient Accelerator for Deconvolutional Computation","volume":"14","author":"Li","year":"2020","journal-title":"IEEE Trans. Comput. Aided Des. Integr. Circuits Syst."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Song, L., Zhou, Y., Qian, X., Li, H., and Chen, Y. (2018, January 24\u201328). GraphR: Accelerating Graph Processing Using ReRAM. Proceedings of the IEEE International Symposium on High Performance Computer Architecture, Vienna, Austria.","DOI":"10.1109\/HPCA.2018.00052"},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"2237","DOI":"10.1109\/JPROC.2010.2070830","article-title":"Resistive Random Access Memory (ReRAM) based on metal oxides","volume":"98","author":"Akinaga","year":"2010","journal-title":"Proc. IEEE"},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"13330","DOI":"10.1038\/srep13330","article-title":"A learnable parallel processing architecture towards unity of memory and computing","volume":"5","author":"Li","year":"2015","journal-title":"Sci. Rep."},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Li, J., Yan, G., Lu, W., Jiang, S., Gong, S., Wu, J., and Li, X. (2018, January 19\u201323). SmartShuttle: Optimization off-chip memory accesses for deep learning accelerators. Proceedings of the 2018 Design, Automation & Test in Europe Conference & Exhibition (DATE), Dresden, Germany.","DOI":"10.23919\/DATE.2018.8342033"},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"702","DOI":"10.1109\/TVLSI.2021.3060509","article-title":"ROMANet: Fine-Grained Reuse-Driven Off-Chip Memory Access Management and Data Organization for Deep Neural Network Accelerators","volume":"29","author":"Putra","year":"2021","journal-title":"IEEE Trans. Very Large-Scale Integr. Syst."},{"key":"ref_34","doi-asserted-by":"crossref","first-page":"1768","DOI":"10.1109\/TCAD.2020.3030610","article-title":"DESCNet: Developing Efficient Scratchpad Memories for Capsule Network Hardware","volume":"40","author":"Marchisio","year":"2021","journal-title":"IEEE Trans. Comput. Aided Des. Integr. Circuits Syst."}],"container-title":["Future Internet"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1999-5903\/14\/5\/146\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T23:08:38Z","timestamp":1760137718000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1999-5903\/14\/5\/146"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,5,10]]},"references-count":34,"journal-issue":{"issue":"5","published-online":{"date-parts":[[2022,5]]}},"alternative-id":["fi14050146"],"URL":"https:\/\/doi.org\/10.3390\/fi14050146","relation":{},"ISSN":["1999-5903"],"issn-type":[{"value":"1999-5903","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,5,10]]}}}