{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,11,11]],"date-time":"2025-11-11T15:52:36Z","timestamp":1762876356904,"version":"3.41.0"},"reference-count":31,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2022,3,8]],"date-time":"2022-03-08T00:00:00Z","timestamp":1646697600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"NSF-CCF","award":["1903951"],"award-info":[{"award-number":["1903951"]}]},{"name":"ASCENT"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Des. Autom. Electron. Syst."],"published-print":{"date-parts":[[2022,7,31]]},"abstract":"<jats:p>On-device embedded artificial intelligence prefers the adaptive learning capability when deployed in the field, and thus in situ training is required. The compute-in-memory approach, which exploits the analog computation within the memory array, is a promising solution for deep neural network (DNN) on-chip acceleration. Emerging non-volatile memories are of great interest, serving as analog synapses due to their multilevel programmability. However, the asymmetry and nonlinearity in the conductance tuning remain grand challenges for achieving high in situ training accuracy. In addition, analog-to-digital converters at the edge of the memory array introduce quantization errors. In this work, we present an algorithm-hardware co-optimization to overcome these challenges. We incorporate the device\/circuit non-ideal effects into the DNN propagation and weight update steps. By introducing the adaptive \u201cmomentum\u201d in the weight update rule, in situ training accuracy on CIFAR-10 could approach its software baseline even under severe asymmetry\/nonlinearity and analog-to-digital converter quantization error. The hardware performance of the on-chip training architecture and the overhead for adding \u201cmomentum\u201d are also evaluated. By optimizing the backpropagation dataflow, 23.59 TOPS\/W training energy efficiency (12\u00d7 improvement compared to na\u00efve dataflow) is achieved. The circuits that handle \u201cmomentum\u201d introduce only 4.2% energy overhead. Our results show great potential and more relaxed requirements that enable emerging non-volatile memories for DNN acceleration on the embedded artificial intelligence platforms.<\/jats:p>","DOI":"10.1145\/3500929","type":"journal-article","created":{"date-parts":[[2022,3,8]],"date-time":"2022-03-08T16:32:11Z","timestamp":1646757131000},"page":"1-19","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":3,"title":["Achieving High In Situ Training Accuracy and Energy Efficiency with Analog Non-Volatile Synaptic Devices"],"prefix":"10.1145","volume":"27","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-1760-7656","authenticated-orcid":false,"given":"Shanshi","family":"Huang","sequence":"first","affiliation":[{"name":"Georgia Institute of Technology, Atlanta, GA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xiaoyu","family":"Sun","sequence":"additional","affiliation":[{"name":"Georgia Institute of Technology, Atlanta, GA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xiaochen","family":"Peng","sequence":"additional","affiliation":[{"name":"Georgia Institute of Technology, Atlanta, GA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hongwu","family":"Jiang","sequence":"additional","affiliation":[{"name":"Georgia Institute of Technology, Atlanta, GA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Shimeng","family":"Yu","sequence":"additional","affiliation":[{"name":"Georgia Institute of Technology, Atlanta, GA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2022,3,8]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.1145\/3079856.3080246"},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1109\/JPROC.2018.2790840"},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSI.2019.2958568"},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISSCC.2018.8310401"},{"key":"e_1_3_1_6_2","doi-asserted-by":"publisher","DOI":"10.1109\/VLSIT.2018.8510690"},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.23919\/VLSIT.2019.8776551"},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.1109\/IEDM.2017.8268338"},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1109\/IEDM.2018.8614551"},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/IEDM.2015.7409625"},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","DOI":"10.1109\/IJCNN.2016.7727298"},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.1038\/s41586-018-0180-5"},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1109\/IEDM.2018.8614611"},{"key":"e_1_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.3389\/fnins.2020.00103"},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.1109\/IEDM.2017.8268337"},{"key":"e_1_3_1_16_2","volume-title":"Proceedings of the International Conference on Learning Representations (ICLR\u201918)","author":"Wu S.","year":"2018","unstructured":"S. Wu, G. Li, F. Chen, and L. Shi. 2018. Training and inference with integers in deep neural networks. In Proceedings of the International Conference on Learning Representations (ICLR\u201918)."},{"key":"e_1_3_1_17_2","first-page":"5151","volume-title":"Advances in Neural Information Processing Systems (NIPS\u201918)","author":"Banner R.","year":"2018","unstructured":"R. Banner, I. Hubara, E. Hoffer, and D. Soudry. 2018. Scalable methods for 8-bit training of neural networks. In Advances in Neural Information Processing Systems (NIPS\u201918). 5151\u20135159."},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.23919\/DATE48585.2020.9116215"},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","DOI":"10.1109\/MNANO.2018.2844902"},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1109\/JETCAS.2019.2933148"},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.1038\/s41563-017-0001-5"},{"issue":"6","key":"e_1_3_1_22_2","first-page":"Article 8","article-title":"AI hardware acceleration with analog memory: Micro-architectures for low energy at high speed","volume":"64","author":"Chang H.-Y.","year":"2019","unstructured":"H.-Y. Chang, P. Narayanan, S. C. Lewis, N. C. P. Farinha, K. Hosokawa, C. Mackin, H. Tsai, S. Ambrogio, A. Chen, and G. W. Burr. 2019. AI hardware acceleration with analog memory: Micro-architectures for low energy at high speed. IBM Journal of Research and Development 64, 6 (2019), Article 8, 14 pages.","journal-title":"IBM Journal of Research and Development"},{"key":"e_1_3_1_23_2","first-page":"315","volume-title":"Proceedings of the 14th International Conference on Artificial Intelligence and Statistics","author":"Glorot X.","year":"2011","unstructured":"X. Glorot, A. Bordes, and Y. Bengio. 2011. Deep sparse rectifier neural networks. In Proceedings of the 14th International Conference on Artificial Intelligence and Statistics. 315\u2013323."},{"key":"e_1_3_1_24_2","first-page":"448","volume-title":"Proceedings of the International Conference on Machine Learning (ICML\u201915)","author":"Sergey I.","year":"2015","unstructured":"I. Sergey and C. Szegedy. 2015. Batch normalization: Accelerating deep network training by reducing internal covariate shift. In Proceedings of the International Conference on Machine Learning (ICML\u201915). 448\u2013456."},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISCAS45731.2020.9181020"},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2017.55"},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.3389\/fnins.2016.00333"},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCAD.2020.3043731"},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.23919\/DATE.2018.8342235"},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISSCC19947.2020.9062979"},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.1109\/JSSC.2021.3101209"},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISSCC42613.2021.9365769"}],"container-title":["ACM Transactions on Design Automation of Electronic Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3500929","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3500929","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T17:49:40Z","timestamp":1750182580000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3500929"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,3,8]]},"references-count":31,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2022,7,31]]}},"alternative-id":["10.1145\/3500929"],"URL":"https:\/\/doi.org\/10.1145\/3500929","relation":{},"ISSN":["1084-4309","1557-7309"],"issn-type":[{"type":"print","value":"1084-4309"},{"type":"electronic","value":"1557-7309"}],"subject":[],"published":{"date-parts":[[2022,3,8]]},"assertion":[{"value":"2021-07-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2021-11-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2022-03-08","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}