{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,7,30]],"date-time":"2025-07-30T16:31:17Z","timestamp":1753893077196,"version":"3.41.2"},"reference-count":44,"publisher":"Frontiers Media SA","license":[{"start":{"date-parts":[[2024,10,4]],"date-time":"2024-10-04T00:00:00Z","timestamp":1728000000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["frontiersin.org"],"crossmark-restriction":true},"short-container-title":["Front. Artif. Intell."],"abstract":"<jats:p>In-memory computing (IMC) with non-volatile memories (NVMs) has emerged as a promising approach to address the rapidly growing computational demands of Deep Neural Networks (DNNs). Mapping DNN layers spatially onto NVM-based IMC accelerators achieves high degrees of parallelism. However, two challenges that arise in this approach are the highly non-uniform distribution of layer processing times and high area requirements. We propose LRMP, a method to jointly apply layer replication and mixed precision quantization to improve the performance of DNNs when mapped to area-constrained IMC accelerators. LRMP uses a combination of reinforcement learning and mixed integer linear programming to search the replication-quantization design space using a model that is closely informed by the target hardware architecture. Across five DNN benchmarks, LRMP achieves 2.6\u20139.3\u00d7 latency and 8\u201318\u00d7 throughput improvement at minimal (&amp;lt;1%) degradation in accuracy.<\/jats:p>","DOI":"10.3389\/frai.2024.1268317","type":"journal-article","created":{"date-parts":[[2024,10,4]],"date-time":"2024-10-04T04:40:26Z","timestamp":1728016826000},"update-policy":"https:\/\/doi.org\/10.3389\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["LRMP: Layer Replication with Mixed Precision for spatial in-memory DNN accelerators"],"prefix":"10.3389","volume":"7","author":[{"given":"Abinand","family":"Nallathambi","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Christin David","family":"Bose","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Wilfried","family":"Haensch","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Anand","family":"Raghunathan","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1965","published-online":{"date-parts":[[2024,10,4]]},"reference":[{"key":"B1","doi-asserted-by":"crossref","first-page":"715","DOI":"10.1145\/3297858.3304049","article-title":"\u201cPuma: A programmable ultra-efficient memristor-based accelerator for machine learning inference,\u201d","author":"Ankit","year":"2019","journal-title":"Proceedings of the Twenty-Fourth International Conference on Architectural Support for Programming Languages and Operating Systems"},{"key":"B2","doi-asserted-by":"publisher","first-page":"283","DOI":"10.3390\/math10020283","article-title":"Transformation and linearization techniques in optimization: a state-of-the-art survey","volume":"10","author":"Asghari","year":"2022","journal-title":"Mathematics"},{"key":"B3","doi-asserted-by":"publisher","first-page":"3498","DOI":"10.1109\/TED.2015.2439635","article-title":"Experimental demonstration and tolerancing of a large-scale neural network (165 000 synapses) using phase-change memory as the synaptic weight element","volume":"62","author":"Burr","year":"2015","journal-title":"IEEE Trans. Electron Devices"},{"key":"B4","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1109\/ISSCC42614.2022.9731679","article-title":"\u201cA 40nm 60.64tops\/w ecc-capable compute-in-memory\/digital 2.25mb\/768kb rram\/sram system with embedded cortex m3 microprocessor for edge recommendation systems,\u201d","volume-title":"2022 IEEE International Solid- State Circuits Conference (ISSCC)","author":"Chang","year":"2022"},{"key":"B5","doi-asserted-by":"publisher","first-page":"27","DOI":"10.1109\/JXCDC.2020.2987605","article-title":"Accurate inference with inaccurate rram devices: A joint algorithm-design solution","volume":"6","author":"Charan","year":"2020","journal-title":"IEEE J. Explorat. Solid-State Comp. Dev. Circ"},{"key":"B6","doi-asserted-by":"publisher","first-page":"105","DOI":"10.1145\/2996864","article-title":"Diannao family: energy-efficient hardware accelerators for machine learning","volume":"59","author":"Chen","year":"2016","journal-title":"Commun. ACM"},{"key":"B7","doi-asserted-by":"publisher","first-page":"27","DOI":"10.1145\/3007787.3001140","article-title":"Prime: A novel processing-in-memory architecture for neural network computation in reram-based main memory","volume":"44","author":"Chi","year":"2016","journal-title":"ACM SIGARCH Comp. Arch. News"},{"key":"B8","first-page":"30318","article-title":"Gpt3. int8 (): 8-bit matrix multiplication for transformers at scale","volume":"35","author":"Dettmers","year":"2022","journal-title":"Adv. Neural Inf. Process. Syst"},{"key":"B9","doi-asserted-by":"crossref","first-page":"100","DOI":"10.1145\/3352460.3358260","article-title":"\u201cComputedram: In-memory compute using off-the-shelf drams,\u201d","author":"Gao","year":"2019","journal-title":"Proceedings of the 52nd Annual IEEE\/ACM International Symposium on Microarchitecture"},{"key":"B10","doi-asserted-by":"publisher","first-page":"907","DOI":"10.3389\/fnins.2020.00907","article-title":"Hfnet: A CNN architecture co-designed for neuromorphic hardware with a crossbar array of synapses","volume":"14","author":"Gopalakrishnan","year":"2020","journal-title":"Front. Neurosci"},{"key":"B11","doi-asserted-by":"crossref","DOI":"10.1145\/3508352.3549453","article-title":"\u201cDesign space and memory technology co-exploration for in-memory computing based machine learning accelerators,\u201d","volume-title":"Proceedings of the 41st IEEE\/ACM International Conference on Computer-Aided Design, ICCAD '22","author":"He","year":"2022"},{"key":"B12","article-title":"\u201cMixed precision quantization for reram-based dnn inference accelerators,\u201d","author":"Huang","year":"2021","journal-title":"2021 26th Asia and South Pacific Design Automation Conference (ASP-DAC"},{"key":"B13","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3362035","article-title":"Cxdnn: Hardware-software compensation methods for deep neural networks on resistive crossbar systems","volume":"18","author":"Jain","year":"2019","journal-title":"ACM Trans. Embed. Comput. Syst"},{"key":"B14","doi-asserted-by":"publisher","first-page":"470","DOI":"10.1109\/TVLSI.2017.2776954","article-title":"Computing in memory with spin-transfer torque magnetic ram","volume":"26","author":"Jain","year":"2018","journal-title":"IEEE Trans. Very Large Scale Integrat"},{"key":"B15","doi-asserted-by":"publisher","first-page":"326","DOI":"10.1109\/TCAD.2020.3000185","article-title":"Rxnn: A framework for evaluating deep neural networks on resistive crossbars","volume":"40","author":"Jain","year":"2020","journal-title":"IEEE Trans. Comp.-Aided Desig. Integrat. Circ. Syst"},{"key":"B16","doi-asserted-by":"publisher","first-page":"114","DOI":"10.1109\/TVLSI.2022.3221390","article-title":"A heterogeneous and programmable compute-in-memory accelerator architecture for analog-ai using dense 2-d mesh","volume":"31","author":"Jain","year":"2023","journal-title":"IEEE Trans. Very Large Scale Integrat"},{"key":"B17","first-page":"1","article-title":"Analog-to-digital converter design exploration for compute-in-memory accelerators","volume":"38","author":"Jiang","year":"2021","journal-title":"IEEE Design Test of Comp"},{"key":"B18","doi-asserted-by":"publisher","first-page":"10","DOI":"10.1109\/MM.2018.032271057","article-title":"Motivation for and evaluation of the first tensor processing unit","volume":"38","author":"Jouppi","year":"2018","journal-title":"IEEE Micro"},{"key":"B19","doi-asserted-by":"publisher","first-page":"649","DOI":"10.1109\/JETCAS.2021.3127129","article-title":"Genetic algorithm based energy-aware CNN quantization for processing-in-memory architecture","volume":"11","author":"Kang","year":"2021","journal-title":"IEEE J. Emerg. Select. Topics Circ. Syst"},{"key":"B20","doi-asserted-by":"publisher","first-page":"2251","DOI":"10.1109\/JPROC.2020.3034117","article-title":"Deep in-memory architectures in sram: an analog approach to approximate computing","volume":"108","author":"Kang","year":"2020","journal-title":"Proc. IEEE"},{"key":"B21","doi-asserted-by":"crossref","first-page":"1","DOI":"10.23919\/VLSICircuits52068.2021.9492362","article-title":"\u201cHermes core-a 14nm cmos and pcm-based in-memory compute core using an array of 300ps\/lsb linearized cco-based adcs and local digital processing,\u201d","volume-title":"2021 Symposium on VLSI Circuits","author":"Khaddam-Aljameh","year":"2021"},{"key":"B22","article-title":"\u201cHITM,\u201d","author":"Li","year":"2020","journal-title":"Proceedings of the 39th International Conference on Computer-Aided Design"},{"key":"B23","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3631523","article-title":"Mathematical framework for optimizing crossbar allocation for reram-based CNN accelerators","volume":"29","author":"Li","year":"2023","journal-title":"ACM Trans. Des. Autom. Electron. Syst"},{"key":"B24","first-page":"17535","article-title":"\u201cQ-diffusion: Quantizing diffusion models,\u201d","author":"Li","year":"2023","journal-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision"},{"key":"B25","doi-asserted-by":"publisher","first-page":"370","DOI":"10.1016\/j.neucom.2021.07.045","article-title":"Pruning and quantization for deep neural network acceleration: a survey","volume":"461","author":"Liang","year":"2021","journal-title":"Neurocomputing"},{"key":"B26","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1109\/HCS55958.2022.9895479","article-title":"\u201cCerebras architecture deep dive: first look inside the hw\/sw co-design for deep learning: Cerebras systems,\u201d","volume-title":"2022 IEEE Hot Chips 34 Symposium (HCS)","author":"Lie","year":"2022"},{"key":"B27","doi-asserted-by":"publisher","first-page":"659060","DOI":"10.3389\/frai.2021.659060","article-title":"Neurosim simulator for compute-in-memory hardware accelerator: validation and benchmark","volume":"4","author":"Lu","year":"2021","journal-title":"Front. Artif. Intellig"},{"key":"B28","doi-asserted-by":"publisher","first-page":"89","DOI":"10.1109\/MM.2021.3131114","article-title":"Temperature-resilient rram-based in-memory computing for DNN inference","volume":"42","author":"Meng","year":"","journal-title":"IEEE Micro"},{"key":"B29","doi-asserted-by":"publisher","first-page":"1576","DOI":"10.1109\/TCSII.2021.3069011","article-title":"Structured pruning of rram crossbars for efficient in-memory computing acceleration of deep neural networks","volume":"68","author":"Meng","year":"","journal-title":"IEEE Trans. Circ. Syst. II-Express Briefs"},{"key":"B30","doi-asserted-by":"publisher","first-page":"6629","DOI":"10.1109\/TED.2021.3115993","article-title":"Fully on-chip mac at 14 nm enabled by accurate row-wise programming of PCM-based weights and parallel vector-transport in duration-format","volume":"68","author":"Narayanan","year":"2021","journal-title":"IEEE Trans. Electron Devices"},{"key":"B31","doi-asserted-by":"publisher","first-page":"4124","DOI":"10.1109\/TCAD.2022.3197495","article-title":"CMQ: Crossbar-aware neural network mixed-precision quantization via differentiable architecture search","volume":"41","author":"Peng","year":"2022","journal-title":"IEEE Trans. Comput.-Aided Des. Integr"},{"key":"B32","first-page":"1","article-title":"\u201cOptimizing weight mapping and data flow for convolutional neural networks on rram based processing-in-memory architecture,\u201d","volume-title":"2019 IEEE International Symposium on Circuits and Systems (ISCAS)","author":"Peng","year":"2019"},{"key":"B33","doi-asserted-by":"publisher","first-page":"2693","DOI":"10.1109\/TED.2021.3072868","article-title":"Variability and energy consumption tradeoffs in multilevel programming of rram arrays","volume":"68","author":"Perez","year":"2021","journal-title":"IEEE Trans. Electron Devices"},{"key":"B34","doi-asserted-by":"publisher","first-page":"753","DOI":"10.3389\/fnins.2019.00753","article-title":"Rapa-convnets: Modified convolutional networks for accelerated training on architectures with analog arrays","volume":"13","author":"Rasch","year":"2019","journal-title":"Front. Neurosci"},{"key":"B35","doi-asserted-by":"publisher","first-page":"730","DOI":"10.1109\/TVLSI.2021.3063543","article-title":"Txsim: modeling training of deep neural networks on resistive crossbar systems","volume":"29","author":"Roy","year":"2021","journal-title":"IEEE Trans. Very Large Scale Integrat"},{"key":"B36","article-title":"\u201cTowards ADC-less compute-in-memory accelerators for energy efficient deep learning,\u201d","volume-title":"Design, Automation and Test in Europe","author":"Saxena","year":"2022"},{"key":"B37","doi-asserted-by":"publisher","first-page":"14","DOI":"10.1145\/3007787.3001139","article-title":"Isaac: a convolutional neural network accelerator with in-situ analog arithmetic in crossbars","volume":"44","author":"Shafiee","year":"2016","journal-title":"ACM SIGARCH Comp. Arch. News"},{"key":"B38","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1109\/IRPS46558.2021.9405210","article-title":"\u201cImpact of multilevel retention characteristics on rram based DNN inference engine,\u201d","volume-title":"2021 IEEE International Reliability Physics Symposium (IRPS)","author":"Shim","year":"2021"},{"key":"B39","doi-asserted-by":"crossref","first-page":"541","DOI":"10.1109\/HPCA.2017.55","article-title":"\u201cPipelayer: A pipelined reram-based accelerator for deep learning,\u201d","volume-title":"2017 IEEE International Symposium on High Performance Computer Architecture (HPCA)","author":"Song","year":"2017"},{"key":"B40","first-page":"8612","article-title":"\u201cHAQ: Hardware-aware automated quantization with mixed precision,\u201d","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Wang","year":"2019"},{"key":"B41","doi-asserted-by":"crossref","first-page":"149","DOI":"10.1145\/3205289.3205297","article-title":"\u201cCelia: a device and architecture co-design framework for stt-mram-based deep learning acceleration,\u201d","volume-title":"Proceedings of the 2018 International Conference on Supercomputing","author":"Yan","year":"2018"},{"key":"B42","doi-asserted-by":"publisher","first-page":"1733","DOI":"10.1109\/JSSC.2019.2963616","article-title":"Xnor-sram: In-memory computing sram macro for binary\/ternary deep neural networks","volume":"55","author":"Yin","year":"2020","journal-title":"IEEE J. Solid-State Circuits"},{"key":"B43","doi-asserted-by":"publisher","first-page":"915","DOI":"10.1109\/JSSC.2016.2642198","article-title":"In-memory computation of a machine-learning classifier in a standard 6t sram array","volume":"52","author":"Zhang","year":"2017","journal-title":"IEEE J. Solid-State Circuits"},{"key":"B44","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3316781.3317739","article-title":"\u201cA configurable multi-precision CNN computing framework based on single bit RRAM,\u201d","volume-title":"Proceedings of the 56th Annual Design Automation Conference 2019","author":"Zhu","year":"2019"}],"container-title":["Frontiers in Artificial Intelligence"],"original-title":[],"link":[{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/frai.2024.1268317\/full","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,10,4]],"date-time":"2024-10-04T04:40:34Z","timestamp":1728016834000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/frai.2024.1268317\/full"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,10,4]]},"references-count":44,"alternative-id":["10.3389\/frai.2024.1268317"],"URL":"https:\/\/doi.org\/10.3389\/frai.2024.1268317","relation":{},"ISSN":["2624-8212"],"issn-type":[{"type":"electronic","value":"2624-8212"}],"subject":[],"published":{"date-parts":[[2024,10,4]]},"article-number":"1268317"}}