{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T04:30:59Z","timestamp":1750221059142,"version":"3.41.0"},"reference-count":54,"publisher":"Association for Computing Machinery (ACM)","issue":"1","license":[{"start":{"date-parts":[[2018,12,21]],"date-time":"2018-12-21T00:00:00Z","timestamp":1545350400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100011002","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["61572470, 61872336, 61532017,61432017, 61521092, 61376043"],"award-info":[{"award-number":["61572470, 61872336, 61532017,61432017, 61521092, 61376043"]}],"id":[{"id":"10.13039\/501100011002","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100004739","name":"Youth Innovation Promotion Association of the Chinese Academy of Sciences","doi-asserted-by":"publisher","award":["Y404441000"],"award-info":[{"award-number":["Y404441000"]}],"id":[{"id":"10.13039\/501100004739","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Des. Autom. Electron. Syst."],"published-print":{"date-parts":[[2019,1,31]]},"abstract":"<jats:p>Neural networks (NNs) have achieved great success in a broad range of applications. As NN-based methods are often both computation and memory intensive, accelerator solutions have been proved to be highly promising in terms of both performance and energy efficiency. Although prior solutions can deliver high computational throughput for convolutional layers, they could incur severe performance degradation when accommodating the entire network model, because there exist very diverse computing and memory bandwidth requirements between convolutional layers and fully connected layers and, furthermore, among different NN models. To overcome this problem, we proposed an elastic accelerator architecture, called SynergyFlow, which intrinsically supports layer-level and model-level parallelism for large-scale deep neural networks. SynergyFlow boosts the resource utilization by exploiting the complementary effect of resource demanding in different layers and different NN models. SynergyFlow can dynamically reconfigure itself according to the workload characteristics, maintaining a high performance and high resource utilization among various models. As a case study, we implement SynergyFlow on a P395-AB FPGA board. Under 100MHz working frequency, our implementation improves the performance by 33.8% on average (up to 67.2% on AlexNet) compared to comparable provisioned previous architectures.<\/jats:p>","DOI":"10.1145\/3275243","type":"journal-article","created":{"date-parts":[[2018,12,21]],"date-time":"2018-12-21T13:39:21Z","timestamp":1545399561000},"page":"1-27","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":5,"title":["SynergyFlow"],"prefix":"10.1145","volume":"24","author":[{"given":"Jiajun","family":"Li","sequence":"first","affiliation":[{"name":"State Key Laboratory of Computer Architecture, Institute of Computing Technology, Chinese Academy of Sciences, People\u2019s Republic of China, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Guihai","family":"Yan","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Computer Architecture, Institute of Computing Technology, Chinese Academy of Sciences, People\u2019s Republic of China, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Wenyan","family":"Lu","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Computer Architecture, Institute of Computing Technology, Chinese Academy of Sciences, People\u2019s Republic of China, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Shijun","family":"Gong","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Computer Architecture, Institute of Computing Technology, Chinese Academy of Sciences, People\u2019s Republic of China, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Shuhao","family":"Jiang","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Computer Architecture, Institute of Computing Technology, Chinese Academy of Sciences, People\u2019s Republic of China, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jingya","family":"Wu","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Computer Architecture, Institute of Computing Technology, Chinese Academy of Sciences, People\u2019s Republic of China, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xiaowei","family":"Li","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Computer Architecture, Institute of Computing Technology, Chinese Academy of Sciences, People\u2019s Republic of China, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2018,12,21]]},"reference":[{"doi-asserted-by":"publisher","key":"e_1_2_1_1_1","DOI":"10.1109\/ISCA.2016.11"},{"doi-asserted-by":"publisher","key":"e_1_2_1_2_1","DOI":"10.5555\/3195638.3195664"},{"doi-asserted-by":"publisher","key":"e_1_2_1_3_1","DOI":"10.1109\/TPAMI.2013.50"},{"doi-asserted-by":"publisher","key":"e_1_2_1_4_1","DOI":"10.1145\/1854273.1854309"},{"doi-asserted-by":"publisher","key":"e_1_2_1_5_1","DOI":"10.1145\/1816038.1815993"},{"doi-asserted-by":"publisher","key":"e_1_2_1_6_1","DOI":"10.1145\/2644865.2541967"},{"doi-asserted-by":"publisher","key":"e_1_2_1_7_1","DOI":"10.1109\/MICRO.2014.58"},{"doi-asserted-by":"publisher","key":"e_1_2_1_8_1","DOI":"10.1109\/JSSC.2016.2616357"},{"key":"e_1_2_1_9_1","first-page":"9","article-title":"Balance principles for algorithm-architecture co-design","volume":"11","author":"Czechowski Kent","year":"2011","journal-title":"HotPar"},{"volume-title":"2013 IEEE International Conference on Acoustics, Speech and Signal Processing. IEEE, 8609--8613","author":"Dahl George E.","key":"e_1_2_1_10_1"},{"unstructured":"Jeffrey Dean Greg Corrado Rajat Monga Kai Chen Matthieu Devin Mark Mao Andrew Senior Paul Tucker Ke Yang and Quoc V. Le. 2012. Large scale distributed deep networks. In Advances in Neural Information Processing Systems. 1223--1231.   Jeffrey Dean Greg Corrado Rajat Monga Kai Chen Matthieu Devin Mark Mao Andrew Senior Paul Tucker Ke Yang and Quoc V. Le. 2012. Large scale distributed deep networks. In Advances in Neural Information Processing Systems. 1223--1231.","key":"e_1_2_1_11_1"},{"doi-asserted-by":"publisher","key":"e_1_2_1_12_1","DOI":"10.1145\/2749469.2750389"},{"doi-asserted-by":"publisher","key":"e_1_2_1_13_1","DOI":"10.1145\/144965.145834"},{"doi-asserted-by":"publisher","key":"e_1_2_1_14_1","DOI":"10.1109\/MICRO.2012.48"},{"doi-asserted-by":"publisher","key":"e_1_2_1_15_1","DOI":"10.1109\/HPCA.2009.4798266"},{"doi-asserted-by":"publisher","key":"e_1_2_1_16_1","DOI":"10.1109\/FPL.2009.5272559"},{"doi-asserted-by":"publisher","key":"e_1_2_1_17_1","DOI":"10.1145\/2020408.2020426"},{"doi-asserted-by":"publisher","key":"e_1_2_1_18_1","DOI":"10.1145\/2637166.2637229"},{"doi-asserted-by":"publisher","key":"e_1_2_1_19_1","DOI":"10.1145\/1816038.1815968"},{"doi-asserted-by":"publisher","key":"e_1_2_1_20_1","DOI":"10.1109\/ISCA.2016.30"},{"doi-asserted-by":"publisher","key":"e_1_2_1_21_1","DOI":"10.1145\/3218603.3218643"},{"unstructured":"X. He W. Lu G. Yan and X. Zhang. 2018b. Joint design of training and hardware towards efficient and accuracy-scalable neural network inference. IEEE Journal on Emerging and Selected Topics in Circuits and Systems (2018) 1--1.  X. He W. Lu G. Yan and X. Zhang. 2018b. Joint design of training and hardware towards efficient and accuracy-scalable neural network inference. IEEE Journal on Emerging and Selected Topics in Circuits and Systems (2018) 1--1.","key":"e_1_2_1_22_1"},{"doi-asserted-by":"publisher","key":"e_1_2_1_23_1","DOI":"10.1162\/neco.2006.18.7.1527"},{"doi-asserted-by":"publisher","key":"e_1_2_1_24_1","DOI":"10.1145\/2505515.2505665"},{"volume-title":"2014 International Joint Conference on Neural Networks (IJCNN\u201914)","author":"Huang W.","key":"e_1_2_1_25_1"},{"doi-asserted-by":"publisher","key":"e_1_2_1_26_1","DOI":"10.1109\/TPAMI.2012.59"},{"doi-asserted-by":"publisher","key":"e_1_2_1_27_1","DOI":"10.1145\/2647868.2654889"},{"doi-asserted-by":"publisher","key":"e_1_2_1_28_1","DOI":"10.1145\/3218603.3218647"},{"unstructured":"Alex Krizhevsky Ilya Sutskever and Geoffrey E. Hinton. 2012. Imagenet classification with deep convolutional neural networks. In Advances in Neural Information Processing Systems. 1097--1105.   Alex Krizhevsky Ilya Sutskever and Geoffrey E. Hinton. 2012. Imagenet classification with deep convolutional neural networks. In Advances in Neural Information Processing Systems. 1097--1105.","key":"e_1_2_1_29_1"},{"doi-asserted-by":"publisher","key":"e_1_2_1_30_1","DOI":"10.1145\/3173162.3173176"},{"doi-asserted-by":"publisher","key":"e_1_2_1_31_1","DOI":"10.1145\/1273496.1273556"},{"doi-asserted-by":"publisher","key":"e_1_2_1_32_1","DOI":"10.1109\/5.726791"},{"doi-asserted-by":"crossref","unstructured":"B. Li Y. Wang Y. Wang Y. Chen and H. Yang. 2014. Training itself: Mixed-signal training acceleration for Memristor-based neural network. In 2014 19th Asia and South Pacific (ASP-DAC\u201914). IEEE 361--366.  B. Li Y. Wang Y. Wang Y. Chen and H. Yang. 2014. Training itself: Mixed-signal training acceleration for Memristor-based neural network. In 2014 19th Asia and South Pacific (ASP-DAC\u201914). IEEE 361--366.","key":"e_1_2_1_33_1","DOI":"10.1109\/ASPDAC.2014.6742916"},{"volume-title":"Automation Test in Europe Conference Exhibition (DATE\u201918)","author":"Li J.","key":"e_1_2_1_34_1"},{"volume-title":"Automation Test in Europe Conference Exhibition (DATE\u201918)","author":"Li J.","key":"e_1_2_1_35_1"},{"doi-asserted-by":"publisher","key":"e_1_2_1_36_1","DOI":"10.1007\/s11227-015-1463-3"},{"doi-asserted-by":"publisher","key":"e_1_2_1_37_1","DOI":"10.1145\/2786763.2694358"},{"doi-asserted-by":"publisher","key":"e_1_2_1_38_1","DOI":"10.1109\/HPCA.2017.29"},{"doi-asserted-by":"publisher","key":"e_1_2_1_39_1","DOI":"10.1109\/TC.2015.2419655"},{"volume-title":"Proceedings of the 29th ICML (ICML\u201912)","year":"2012","author":"Mnih Volodymyr","key":"e_1_2_1_40_1"},{"doi-asserted-by":"publisher","key":"e_1_2_1_41_1","DOI":"10.1145\/3079856.3080254"},{"doi-asserted-by":"publisher","key":"e_1_2_1_42_1","DOI":"10.1109\/ICCD.2013.6657019"},{"doi-asserted-by":"publisher","key":"e_1_2_1_43_1","DOI":"10.1145\/2847263.2847265"},{"doi-asserted-by":"publisher","key":"e_1_2_1_44_1","DOI":"10.1109\/ASAP.2009.25"},{"volume-title":"2017 IEEE International Solid-State Circuits Conference (ISSCC\u201917)","author":"Shin D.","key":"e_1_2_1_45_1"},{"doi-asserted-by":"crossref","unstructured":"David Silver Aja Huang Chris J. Maddison Arthur Guez Laurent Sifre George Van Den Driessche Julian Schrittwieser Ioannis Antonoglou Veda Panneershelvam and Marc Lanctot. 2016. Mastering the game of Go with deep neural networks and tree search. Nature 529 7587 (2016) 484--489.  David Silver Aja Huang Chris J. Maddison Arthur Guez Laurent Sifre George Van Den Driessche Julian Schrittwieser Ioannis Antonoglou Veda Panneershelvam and Marc Lanctot. 2016. Mastering the game of Go with deep neural networks and tree search. Nature 529 7587 (2016) 484--489.","key":"e_1_2_1_46_1","DOI":"10.1038\/nature16961"},{"unstructured":"Karen Simonyan and Andrew Zisserman. 2014. Very deep convolutional networks for large-scale image recognition. arXiv Preprint arXiv:1409.1556 (2014).  Karen Simonyan and Andrew Zisserman. 2014. Very deep convolutional networks for large-scale image recognition. arXiv Preprint arXiv:1409.1556 (2014).","key":"e_1_2_1_47_1"},{"doi-asserted-by":"publisher","key":"e_1_2_1_48_1","DOI":"10.1145\/2847263.2847276"},{"doi-asserted-by":"crossref","unstructured":"F. Tu S. Yin P. Ouyang S. Tang L. Liu and S. Wei. 2017. Deep convolutional neural network architecture with reconfigurable computation patterns. IEEE Transactions on Very Large Scale Integration Systems (VLSI\u201917) 25 8 (2017) 2220--2233.  F. Tu S. Yin P. Ouyang S. Tang L. Liu and S. Wei. 2017. Deep convolutional neural network architecture with reconfigurable computation patterns. IEEE Transactions on Very Large Scale Integration Systems (VLSI\u201917) 25 8 (2017) 2220--2233.","key":"e_1_2_1_49_1","DOI":"10.1109\/TVLSI.2017.2688340"},{"volume":"1","volume-title":"Proceeding of the Deep Learning and Unsupervised Feature Learning NIPS Workshop","author":"Vanhoucke Vincent","key":"e_1_2_1_50_1"},{"doi-asserted-by":"publisher","key":"e_1_2_1_51_1","DOI":"10.1145\/1498765.1498785"},{"unstructured":"Samuel Webb Williams. 2008. Auto-Tuning Performance on Multicore Computers. University of California Berkeley.  Samuel Webb Williams. 2008. Auto-Tuning Performance on Multicore Computers. University of California Berkeley.","key":"e_1_2_1_52_1"},{"doi-asserted-by":"publisher","key":"e_1_2_1_53_1","DOI":"10.1145\/2684746.2689060"},{"doi-asserted-by":"publisher","key":"e_1_2_1_54_1","DOI":"10.5555\/3195638.3195662"}],"container-title":["ACM Transactions on Design Automation of Electronic Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3275243","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3275243","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T00:44:38Z","timestamp":1750207478000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3275243"}},"subtitle":["An Elastic Accelerator Architecture Supporting Batch Processing of Large-Scale Deep Neural Networks"],"short-title":[],"issued":{"date-parts":[[2018,12,21]]},"references-count":54,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2019,1,31]]}},"alternative-id":["10.1145\/3275243"],"URL":"https:\/\/doi.org\/10.1145\/3275243","relation":{},"ISSN":["1084-4309","1557-7309"],"issn-type":[{"type":"print","value":"1084-4309"},{"type":"electronic","value":"1557-7309"}],"subject":[],"published":{"date-parts":[[2018,12,21]]},"assertion":[{"value":"2018-01-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2018-09-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2018-12-21","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}