{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,12,8]],"date-time":"2025-12-08T22:16:57Z","timestamp":1765232217405,"version":"3.41.0"},"reference-count":45,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2015,5,22]],"date-time":"2015-05-22T00:00:00Z","timestamp":1432252800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"10,000 and 1,000 talent programs"},{"name":"NSF of China","award":["61100163, 61133004, 61222204, 61221062, 61303158, 61432016, 61472396, 61473275"],"award-info":[{"award-number":["61100163, 61133004, 61222204, 61221062, 61303158, 61432016, 61472396, 61473275"]}]},{"name":"Strategic Priority Research Program of the CAS","award":["XDA06010403"],"award-info":[{"award-number":["XDA06010403"]}]},{"name":"973 Program of China","award":["2015CB358800"],"award-info":[{"award-number":["2015CB358800"]}]},{"name":"Google Faculty Research Award"},{"name":"French ANR MHANN and NEMESIS"},{"name":"International Collaboration Key Program of the CAS","award":["171111KYSB20130002"],"award-info":[{"award-number":["171111KYSB20130002"]}]},{"name":"Intel Collaborative Research Institute for Computational Intelligence (ICRI-CI)"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Comput. Syst."],"published-print":{"date-parts":[[2015,6,8]]},"abstract":"<jats:p>Machine-learning tasks are becoming pervasive in a broad range of domains, and in a broad range of systems (from embedded systems to data centers). At the same time, a small set of machine-learning algorithms (especially Convolutional and Deep Neural Networks, i.e., CNNs and DNNs) are proving to be state-of-the-art across many applications. As architectures evolve toward heterogeneous multicores composed of a mix of cores and accelerators, a machine-learning accelerator can achieve the rare combination of efficiency (due to the small number of target algorithms) and broad application scope.<\/jats:p>\n          <jats:p>Until now, most machine-learning accelerator designs have been focusing on efficiently implementing the computational part of the algorithms. However, recent state-of-the-art CNNs and DNNs are characterized by their large size. In this study, we design an accelerator for large-scale CNNs and DNNs, with a special emphasis on the impact of memory on accelerator design, performance, and energy.<\/jats:p>\n          <jats:p>We show that it is possible to design an accelerator with a high throughput, capable of performing 452 GOP\/s (key NN operations such as synaptic weight multiplications and neurons outputs additions) in a small footprint of 3.02mm&lt;sup&gt;2&lt;\/sup&gt; and 485mW; compared to a 128-bit 2GHz SIMD processor, the accelerator is 117.87 \u00d7 faster, and it can reduce the total energy by 21.08 \u00d7. The accelerator characteristics are obtained after layout at 65nm. Such a high throughput in a small footprint can open up the usage of state-of-the-art machine-learning algorithms in a broad set of systems and for a broad set of applications.<\/jats:p>","DOI":"10.1145\/2701417","type":"journal-article","created":{"date-parts":[[2015,5,26]],"date-time":"2015-05-26T14:36:05Z","timestamp":1432650965000},"page":"1-27","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":8,"title":["A Small-Footprint Accelerator for Large-Scale Neural Networks"],"prefix":"10.1145","volume":"33","author":[{"given":"Tianshi","family":"Chen","sequence":"first","affiliation":[{"name":"SKLCA, ICT, CAS, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Shijin","family":"Zhang","sequence":"additional","affiliation":[{"name":"SKLCA, ICT, CAS, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Shaoli","family":"Liu","sequence":"additional","affiliation":[{"name":"SKLCA, ICT, CAS, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zidong","family":"Du","sequence":"additional","affiliation":[{"name":"SKLCA, ICT, CAS, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Tao","family":"Luo","sequence":"additional","affiliation":[{"name":"SKLCA, ICT, CAS, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yuan","family":"Gao","sequence":"additional","affiliation":[{"name":"TNLIST, Tsinghua University, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Junjie","family":"Liu","sequence":"additional","affiliation":[{"name":"TNLIST, Tsinghua University, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Dongsheng","family":"Wang","sequence":"additional","affiliation":[{"name":"TNLIST, Tsinghua University, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Chengyong","family":"Wu","sequence":"additional","affiliation":[{"name":"SKLCA, ICT, CAS, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ninghui","family":"Sun","sequence":"additional","affiliation":[{"name":"SKLCA, ICT, CAS, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yunji","family":"Chen","sequence":"additional","affiliation":[{"name":"SKLCA, ICT, CAS, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Olivier","family":"Temam","sequence":"additional","affiliation":[{"name":"Inria, Saclay, France"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2015,5,22]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2008.4771812"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1145\/1454115.1454128"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/1815961.1815993"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/IISWC.2012.6402898"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2014.58"},{"volume-title":"International Conference on Machine Learning. http:\/\/jmlr.org\/proceedings\/papers\/v28\/coates13","author":"Coates Adam","key":"e_1_2_1_6_1","unstructured":"Adam Coates , Brody Huval , Tao Wang , David J. Wu , and Andrew Y. Ng . 2013. Deep learning with cots hpc systems . In International Conference on Machine Learning. http:\/\/jmlr.org\/proceedings\/papers\/v28\/coates13 .html. Adam Coates, Brody Huval, Tao Wang, David J. Wu, and Andrew Y. Ng. 2013. Deep learning with cots hpc systems. In International Conference on Machine Learning. http:\/\/jmlr.org\/proceedings\/papers\/v28\/coates13.html."},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1023\/A:1022627411411"},{"volume-title":"International Conference on Acoustics, Speech and Signal Processing. http:\/\/www.cs.toronto.edu\/&sim;gdahl\/papers\/reluDropoutBN&lowbar;icassp2013","author":"Dahl George E.","key":"e_1_2_1_8_1","unstructured":"George E. Dahl , Tara N. Sainath , and Geoffrey E. Hinton . 2013. Improving deep neural networks for LVCSR using rectified linear units and dropout . In International Conference on Acoustics, Speech and Signal Processing. http:\/\/www.cs.toronto.edu\/&sim;gdahl\/papers\/reluDropoutBN&lowbar;icassp2013 .pdf. George E. Dahl, Tara N. Sainath, and Geoffrey E. Hinton. 2013. Improving deep neural networks for LVCSR using rectified linear units and dropout. In International Conference on Acoustics, Speech and Signal Processing. http:\/\/www.cs.toronto.edu\/&sim;gdahl\/papers\/reluDropoutBN&lowbar;icassp2013.pdf."},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1016\/S0893-6080(02)00032-1"},{"key":"e_1_2_1_10_1","volume-title":"Asia and South Pacific Design Automation Conference.","author":"Du Zidong","year":"2014","unstructured":"Zidong Du , Avinash Lingamneni , Yunji Chen , Krishna V. Palem , Olivier Temam , and Chengyong Wu . 2014 . Leveraging the error resilience of machine-learning applications for designing highly energy efficient accelerators . In Asia and South Pacific Design Automation Conference. Zidong Du, Avinash Lingamneni, Yunji Chen, Krishna V. Palem, Olivier Temam, and Chengyong Wu. 2014. Leveraging the error resilience of machine-learning applications for designing highly energy efficient accelerators. In Asia and South Pacific Design Automation Conference."},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/2000064.2000108"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2012.48"},{"key":"e_1_2_1_13_1","volume-title":"Mahlke","author":"Fan Kevin","year":"2009","unstructured":"Kevin Fan , Manjunath Kudlur , Ganesh S. Dasika , and Scott A . Mahlke . 2009 . Bridging the computation gap between programmable processors and hardwired accelerators. In HPCA. IEEE Computer Society , 313--322. Kevin Fan, Manjunath Kudlur, Ganesh S. Dasika, and Scott A. Mahlke. 2009. Bridging the computation gap between programmable processors and hardwired accelerators. In HPCA. IEEE Computer Society, 313--322."},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW.2011.5981829"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/1815961.1815968"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/1950365.1950385"},{"key":"e_1_2_1_17_1","unstructured":"Geoffrey E. Hinton and N. Srivastava. 2012. Improving neural networks by preventing co-adaptation of feature detectors. arXiv preprint arXiv (2012) 1--18. http:\/\/arxiv.org\/abs\/1207.0580  Geoffrey E. Hinton and N. Srivastava. 2012. Improving neural networks by preventing co-adaptation of feature detectors. arXiv preprint arXiv (2012) 1--18. http:\/\/arxiv.org\/abs\/1207.0580"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/12.210171"},{"volume-title":"\u201cfloating gate","author":"Holler Mark","key":"e_1_2_1_19_1","unstructured":"Mark Holler , Simon Tam , Hernan Castro , and Ronald Benson . 1990. An electrically trainable artificial neural network (ETANN) with 10240 \u201cfloating gate \u201d synapses. In Artificial Neural Networks. IEEE Press , Piscataway, NJ, 50--55. DOI:http:\/\/dx.doi.org\/10.1109\/IJCNN.1989.118698 10.1109\/IJCNN.1989.118698 Mark Holler, Simon Tam, Hernan Castro, and Ronald Benson. 1990. An electrically trainable artificial neural network (ETANN) with 10240 \u201cfloating gate\u201d synapses. In Artificial Neural Networks. IEEE Press, Piscataway, NJ, 50--55. DOI:http:\/\/dx.doi.org\/10.1109\/IJCNN.1989.118698"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/2505515.2505665"},{"key":"e_1_2_1_21_1","volume-title":"IEEE International Joint Conference on Neural Networks (IJCNN). IEEE, 2849--2856","author":"Khan Muhammad Mukaram","year":"2008","unstructured":"Muhammad Mukaram Khan , David R. Lester , Luis A. Plana , Alexander D. Rast , Xin Jin , Eustace Painkras , and Stephen B. Furber . 2008. SpiNNaker: Mapping neural networks onto a massively-parallel chip multiprocessor . In IEEE International Joint Conference on Neural Networks (IJCNN). IEEE, 2849--2856 . DOI:http:\/\/dx.doi.org\/10.1109\/IJCNN. 2008 .4634199 10.1109\/IJCNN.2008.4634199 Muhammad Mukaram Khan, David R. Lester, Luis A. Plana, Alexander D. Rast, Xin Jin, Eustace Painkras, and Stephen B. Furber. 2008. SpiNNaker: Mapping neural networks onto a massively-parallel chip multiprocessor. In IEEE International Joint Conference on Neural Networks (IJCNN). IEEE, 2849--2856. DOI:http:\/\/dx.doi.org\/10.1109\/IJCNN.2008.4634199"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1109\/JSSC.2009.2031768"},{"key":"e_1_2_1_23_1","volume-title":"Conference Record of the 31st Asilomar Conference on Signals, Systems &amp; Computers","volume":"2","author":"Eric","unstructured":"Eric J. King and Earl E. Swartzlander Jr. 1997. Data-dependent truncation scheme for parallel multipliers . In Conference Record of the 31st Asilomar Conference on Signals, Systems &amp; Computers , Vol. 2 . IEEE, 1178--1182. Eric J. King and Earl E. Swartzlander Jr. 1997. Data-dependent truncation scheme for parallel multipliers. In Conference Record of the 31st Asilomar Conference on Signals, Systems &amp; Computers, Vol. 2. IEEE, 1178--1182."},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1007\/11760191_192"},{"key":"e_1_2_1_25_1","volume-title":"O\u2019Connor","author":"Larkin Daniel","year":"2006","unstructured":"Daniel Larkin , Andrew Kinane , and Noel E . O\u2019Connor . 2006 a. Towards hardware acceleration of neuroevolution for multimedia processing applications on mobile devices. In ICONIP ( 3). 1178--1188. Daniel Larkin, Andrew Kinane, and Noel E. O\u2019Connor. 2006a. Towards hardware acceleration of neuroevolution for multimedia processing applications on mobile devices. In ICONIP (3). 1178--1188."},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1145\/1273496.1273556"},{"volume-title":"International Conference on Machine Learning.","author":"Le Quoc V.","key":"e_1_2_1_27_1","unstructured":"Quoc V. Le , MarcAurelio Aurelio Ranzato , Rajat Monga , Matthieu Devin , Kai Chen , Greg S. Corrado , Jeffrey Dean , and Andrew Y. Ng . 2012. Building high-level features using large scale unsupervised learning . In International Conference on Machine Learning. Quoc V. Le, MarcAurelio Aurelio Ranzato, Rajat Monga, Matthieu Devin, Kai Chen, Greg S. Corrado, Jeffrey Dean, and Andrew Y. Ng. 2012. Building high-level features using large scale unsupervised learning. In International Conference on Machine Learning."},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1109\/5.726791"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/1669112.1669172"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1145\/2228360.2228465"},{"volume-title":"IEEE Custom Integrated Circuits Conference. IEEE, 1--4.","author":"Merolla Paul","key":"e_1_2_1_31_1","unstructured":"Paul Merolla , John Arthur , Filipp Akopyan , Nabil Imam , Rajit Manohar , and D. S. Modha . 2011. A digital neurosynaptic core using embedded crossbar memory with 45pJ per spike in 45nm . In IEEE Custom Integrated Circuits Conference. IEEE, 1--4. Paul Merolla, John Arthur, Filipp Akopyan, Nabil Imam, Rajit Manohar, and D. S. Modha. 2011. A digital neurosynaptic core using embedded crossbar memory with 45pJ per spike in 45nm. In IEEE Custom Integrated Circuits Conference. IEEE, 1--4."},{"key":"e_1_2_1_32_1","volume-title":"Proceedings of the 29th International Conference on Machine Learning (ICML\u201912)","author":"Mnih Volodymyr","year":"2012","unstructured":"Volodymyr Mnih and Geoffrey Hinton . 2012 . Learning to label aerial images from noisy data . In Proceedings of the 29th International Conference on Machine Learning (ICML\u201912) . 567--574. Volodymyr Mnih and Geoffrey Hinton. 2012. Learning to label aerial images from noisy data. In Proceedings of the 29th International Conference on Machine Learning (ICML\u201912). 567--574."},{"key":"e_1_2_1_33_1","volume-title":"Virtual Conference.","author":"Muller Mike","year":"2010","unstructured":"Mike Muller . 2010 . Dark silicon and the internet. In EE Times \u201cDesigning with ARM \u201d Virtual Conference. Mike Muller. 2010. Dark silicon and the internet. In EE Times \u201cDesigning with ARM\u201d Virtual Conference."},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1145\/2485922.2485925"},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1109\/IJCNN.2008.4633828"},{"volume-title":"International Conference on Pattern Recognition. http:\/\/ieeexplore.ieee.org\/xpls\/abs_all.jsp&quest;arnumber&equals;6460867","author":"Sermanet Pierre","key":"e_1_2_1_36_1","unstructured":"Pierre Sermanet , Soumith Chintala , and Y. LeCun . 2012. Convolutional neural networks applied to house numbers digit classification . In International Conference on Pattern Recognition. http:\/\/ieeexplore.ieee.org\/xpls\/abs_all.jsp&quest;arnumber&equals;6460867 . Pierre Sermanet, Soumith Chintala, and Y. LeCun. 2012. Convolutional neural networks applied to house numbers digit classification. In International Conference on Pattern Recognition. http:\/\/ieeexplore.ieee.org\/xpls\/abs_all.jsp&quest;arnumber&equals;6460867."},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1109\/IJCNN.2011.6033589"},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2007.56"},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.5555\/2337159.2337200"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1016\/0167-739X(95)00022-K"},{"key":"e_1_2_1_41_1","unstructured":"Shyamkumar Thoziyoor Naveen Muralimanohar and JH Ahn. 2008. CACTI 5.1. HP Labs Palo Alto Tech (2008). http:\/\/www.hpl.hp.com\/techreports\/2008\/HPL-2008-20.pdf&quest;q&equals;cacti.  Shyamkumar Thoziyoor Naveen Muralimanohar and JH Ahn. 2008. CACTI 5.1. HP Labs Palo Alto Tech (2008). http:\/\/www.hpl.hp.com\/techreports\/2008\/HPL-2008-20.pdf&quest;q&equals;cacti."},{"key":"e_1_2_1_42_1","volume-title":"Deep Learning and Unsupervised Feature Learning Workshop, NIPS","author":"Vanhoucke Vincent","year":"2011","unstructured":"Vincent Vanhoucke , Andrew Senior , and Mark Z. Mao . 2011. Improving the speed of neural networks on CPUs . In Deep Learning and Unsupervised Feature Learning Workshop, NIPS 2011 . Vincent Vanhoucke, Andrew Senior, and Mark Z. Mao. 2011. Improving the speed of neural networks on CPUs. In Deep Learning and Unsupervised Feature Learning Workshop, NIPS 2011."},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1145\/2155620.2155640"},{"key":"e_1_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1109\/TNN.2006.883007"},{"key":"e_1_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2009.4798263"}],"container-title":["ACM Transactions on Computer Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2701417","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/2701417","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T06:13:10Z","timestamp":1750227190000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2701417"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2015,5,22]]},"references-count":45,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2015,6,8]]}},"alternative-id":["10.1145\/2701417"],"URL":"https:\/\/doi.org\/10.1145\/2701417","relation":{},"ISSN":["0734-2071","1557-7333"],"issn-type":[{"type":"print","value":"0734-2071"},{"type":"electronic","value":"1557-7333"}],"subject":[],"published":{"date-parts":[[2015,5,22]]},"assertion":[{"value":"2014-09-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2014-12-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2015-05-22","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}