{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,14]],"date-time":"2026-04-14T16:01:02Z","timestamp":1776182462021,"version":"3.50.1"},"reference-count":42,"publisher":"Association for Computing Machinery (ACM)","issue":"5","license":[{"start":{"date-parts":[[2022,6,6]],"date-time":"2022-06-06T00:00:00Z","timestamp":1654473600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Des. Autom. Electron. Syst."],"published-print":{"date-parts":[[2022,9,30]]},"abstract":"<jats:p>Conventionally, DNN models are trained once in the cloud and deployed in edge devices such as cars, robots, or unmanned aerial vehicles (UAVs) for real-time inference. However, there are many cases that require the models to adapt to new environments, domains, or users. In order to realize such domain adaption or personalization, the models on devices need to be continuously trained on the device. In this work, we design EF-Train, an efficient DNN training accelerator with a unified channel-level parallelism-based convolution kernel that can achieve end-to-end training on resource-limited low-power edge-level FPGAs. It is challenging to implement on-device training on resource-limited FPGAs due to the low efficiency caused by different memory access patterns among forward and backward propagation and weight update. Therefore, we developed a data reshaping approach with intra-tile continuous memory allocation and weight reuse. An analytical model is established to automatically schedule computation and memory resources to achieve high energy efficiency on edge FPGAs. The experimental results show that our design achieves 46.99 GFLOPS and 6.09 GFLOPS\/W in terms of throughput and energy efficiency, respectively.<\/jats:p>","DOI":"10.1145\/3505633","type":"journal-article","created":{"date-parts":[[2022,2,24]],"date-time":"2022-02-24T17:13:41Z","timestamp":1645722821000},"page":"1-36","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":24,"title":["EF-Train: Enable Efficient On-device CNN Training on FPGA through Data Reshaping for Online Adaptation or Personalization"],"prefix":"10.1145","volume":"27","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-3253-8003","authenticated-orcid":false,"given":"Yue","family":"Tang","sequence":"first","affiliation":[{"name":"University of Pittsburgh, Pittsburgh, PA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9307-1654","authenticated-orcid":false,"given":"Xinyi","family":"Zhang","sequence":"additional","affiliation":[{"name":"University of Pittsburgh, Pittsburgh, PA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0493-1844","authenticated-orcid":false,"given":"Peipei","family":"Zhou","sequence":"additional","affiliation":[{"name":"University of Pittsburgh, Pittsburgh, PA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4029-4034","authenticated-orcid":false,"given":"Jingtong","family":"Hu","sequence":"additional","affiliation":[{"name":"University of Pittsburgh, Pittsburgh, PA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2022,6,6]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.5858\/arpa.2016-0471-ED"},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICC40277.2020.9148776"},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.1145\/3316781.3317829"},{"key":"e_1_3_1_5_2","unstructured":"Xilinx. ([n. d.]). Corazon AI. http:\/\/www.xilinx.com\/products\/boards-and-kits\/1-1bua5s3.html."},{"key":"e_1_3_1_6_2","unstructured":"Xilinx. ([n.d.]). Pony.ai Sensor Fusion using multiple Xilinx devices. https:\/\/www.xilinx.com\/applications\/automotive\/automated-driving.html."},{"key":"e_1_3_1_7_2","unstructured":"Xilinx. ([n.d.]). ZF ProAI Gen 3 using Xilinx Zynq UltraScale+ MPSoC. https:\/\/www.xilinx.com\/applications\/automotive\/automated-driving.html."},{"issue":"18","key":"e_1_3_1_8_2","first-page":"19","article-title":"Real-time data analysis for medical diagnosis using FPGA-accelerated neural networks","volume":"19","author":"Sanaullah Ahmed","year":"2018","unstructured":"Ahmed Sanaullah, Chen Yang, Yuri Alexeev, Kazutomo Yoshii, and Martin C. Herbordt. 2018. Real-time data analysis for medical diagnosis using FPGA-accelerated neural networks. BMC Bioinformatics 19, 18 (2018), 19\u201331.","journal-title":"BMC Bioinformatics"},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1109\/PerComWorkshops48775.2020.9156260"},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2019.2911709"},{"key":"e_1_3_1_11_2","article-title":"Skynet: A champion model for DAC-SDC on low power object detection","author":"Zhang Xiaofan","year":"2019","unstructured":"Xiaofan Zhang, Cong Hao, Haoming Lu, Jiachen Li, Yuhong Li, Yuchen Fan, Kyle Rupnow, Jinjun Xiong, Thomas Huang, Honghui Shi, et\u00a0al. 2019. Skynet: A champion model for DAC-SDC on low power object detection. arXiv preprint arXiv:1906.10327 (2019).","journal-title":"arXiv preprint arXiv:1906.10327"},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.1109\/JIOT.2020.2976762"},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1109\/PERCOM.2018.8444585"},{"key":"e_1_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.artmed.2021.102019"},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.5369\/JSST.2019.29.1.19"},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.1145\/3287075"},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1145\/2684746.2689060"},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1145\/3240765.3240801"},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","DOI":"10.1145\/3316781.3317875"},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1145\/3218603.3218625"},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW50498.2020.00208"},{"key":"e_1_3_1_22_2","first-page":"107","volume-title":"2016 IEEE 27th International Conference on Application-specific Systems, Architectures and Processors (ASAP\u201916)","author":"Zhao Wenlai","year":"2016","unstructured":"Wenlai Zhao, Haohuan Fu, Wayne Luk, Teng Yu, Shaojun Wang, Bo Feng, Yuchun Ma, and Guangwen Yang. 2016. F-CNN: An FPGA-based framework for training convolutional neural networks. In 2016 IEEE 27th International Conference on Application-specific Systems, Architectures and Processors (ASAP\u201916). IEEE, 107\u2013114."},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.1109\/FPL.2019.00034"},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.1088\/1674-4926\/41\/2\/022403"},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1145\/3358192"},{"key":"e_1_3_1_26_2","first-page":"622","volume-title":"2020 53rd Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO\u201920)","author":"Kao Sheng-Chun","year":"2020","unstructured":"Sheng-Chun Kao, Geonhwa Jeong, and Tushar Krishna. 2020. Confuciux: Autonomous hardware resource assignment for DNN accelerators using reinforcement learning. In 2020 53rd Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO\u201920). IEEE, 622\u2013636."},{"key":"e_1_3_1_27_2","doi-asserted-by":"crossref","first-page":"152","DOI":"10.1109\/FCCM.2017.25","volume-title":"2017 IEEE 25th Annual International Symposium on Field-programmable Custom Computing Machines (FCCM\u201917)","author":"Guan Yijin","year":"2017","unstructured":"Yijin Guan, Hao Liang, Ningyi Xu, Wenqiang Wang, Shaoshuai Shi, Xi Chen, Guangyu Sun, Wei Zhang, and Jason Cong. 2017. FP-DNN: An automated framework for mapping deep neural networks onto FPGAs with RTL-HLS hybrid templates. In 2017 IEEE 25th Annual International Symposium on Field-programmable Custom Computing Machines (FCCM\u201917). IEEE, 152\u2013159."},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.1109\/TVLSI.2021.3060509"},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.1109\/DAC18072.2020.9218672"},{"key":"e_1_3_1_30_2"},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCAD.2017.2785257"},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.1145\/3373087.3375321"},{"key":"e_1_3_1_33_2","doi-asserted-by":"publisher","DOI":"10.1109\/TC.2008.29"},{"key":"e_1_3_1_34_2","doi-asserted-by":"publisher","DOI":"10.1109\/FPL.2019.00036"},{"key":"e_1_3_1_35_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICFPT47387.2019.00009"},{"key":"e_1_3_1_36_2","first-page":"1","volume-title":"2020 IEEE Workshop on Signal Processing Systems (SiPS\u201920)","author":"Lu Jinming","year":"2020","unstructured":"Jinming Lu, Jun Lin, and Zhongfeng Wang. 2020. A reconfigurable DNN training accelerator on FPGA. In 2020 IEEE Workshop on Signal Processing Systems (SiPS\u201920). IEEE, 1\u20136."},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.1109\/FPT.2017.8280142"},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCAD.2017.2705069"},{"key":"e_1_3_1_39_2","article-title":"CUDNN: Efficient primitives for deep learning","author":"Chetlur Sharan","year":"2014","unstructured":"Sharan Chetlur, Cliff Woolley, Philippe Vandermersch, Jonathan Cohen, John Tran, Bryan Catanzaro, and Evan Shelhamer. 2014. CUDNN: Efficient primitives for deep learning. arXiv preprint arXiv:1410.0759 (2014).","journal-title":"arXiv preprint arXiv:1410.0759"},{"key":"e_1_3_1_40_2","unstructured":"OpenVINO. ([n. d.]). Optimization Guide. https:\/\/docs.openvino.ai\/2020.2\/_docs_optimization_guide_dldt_optimization_guide.html."},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","DOI":"10.1145\/3400302.3415643"},{"key":"e_1_3_1_42_2","doi-asserted-by":"publisher","DOI":"10.1145\/3373087.3375389"},{"key":"e_1_3_1_43_2","first-page":"1","volume-title":"2020 57th ACM\/IEEE Design Automation Conference (DAC\u201920)","author":"Li Yuhong","year":"2020","unstructured":"Yuhong Li, Cong Hao, Xiaofan Zhang, Xinheng Liu, Yao Chen, Jinjun Xiong, Wen-mei Hwu, and Deming Chen. 2020. AEdd: Efficient differentiable DNN architecture and implementation co-search for embedded ai solutions. In 2020 57th ACM\/IEEE Design Automation Conference (DAC\u201920). IEEE, 1\u20136."}],"container-title":["ACM Transactions on Design Automation of Electronic Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3505633","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3505633","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T17:49:24Z","timestamp":1750182564000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3505633"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,6,6]]},"references-count":42,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2022,9,30]]}},"alternative-id":["10.1145\/3505633"],"URL":"https:\/\/doi.org\/10.1145\/3505633","relation":{},"ISSN":["1084-4309","1557-7309"],"issn-type":[{"value":"1084-4309","type":"print"},{"value":"1557-7309","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,6,6]]},"assertion":[{"value":"2021-06-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2021-12-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2022-06-06","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}