{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2024,5,1]],"date-time":"2024-05-01T12:19:29Z","timestamp":1714565969651},"reference-count":49,"publisher":"Institute of Electronics, Information and Communications Engineers (IEICE)","issue":"6","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["IEICE Trans. Electron."],"published-print":{"date-parts":[[2022,6,1]]},"DOI":"10.1587\/transele.2021lhp0003","type":"journal-article","created":{"date-parts":[[2021,12,2]],"date-time":"2021-12-02T22:09:29Z","timestamp":1638482969000},"page":"209-221","source":"Crossref","is-referenced-by-count":2,"title":["In Search of the Performance- and Energy-Efficient CNN Accelerators"],"prefix":"10.1587","volume":"E105.C","author":[{"given":"Stanislav","family":"SEDUKHIN","sequence":"first","affiliation":[{"name":"University of Aizu"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yoichi","family":"TOMIOKA","sequence":"additional","affiliation":[{"name":"University of Aizu"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Kohei","family":"YAMAMOTO","sequence":"additional","affiliation":[{"name":"Oki Electric Industry Co., Ltd."}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"532","reference":[{"key":"1","unstructured":"[1] K. Simonyan and A. Zisserman, \u201cVery deep convolutional networks for large-scale image recognition,\u201d 3rd Int. Conf. Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceedings, ed. Y. Bengio and Y. LeCun, 2015. 10.48550\/arXiv.1409.1556"},{"key":"2","doi-asserted-by":"crossref","unstructured":"[2] W. Liu, D. Anguelov, D. Erhan, C. Szegedy, S. Reed, C.Y. Fu, and A.C. Berg, \u201cSSD: Single Shot MultiBox Detector,\u201d Lect. Notes Comput. Sci., p.21-37, 2016. 10.1007\/978-3-319-46448-0_2","DOI":"10.1007\/978-3-319-46448-0_2"},{"key":"3","unstructured":"[3] WikiChip Fuse, \u201cInside Tesla&apos;s neural processor in the FSD chip,\u201d https:\/\/fuse.wikichip.org\/news\/2707\/inside-teslas-neural-processor-in-the-fsd-chip\/, 2019."},{"key":"4","doi-asserted-by":"publisher","unstructured":"[4] P.S. Labini, M. Cianfriglia, D. Perri, O. Gervasi, G. Fursin, A. Lokhmotov, C. Nugteren, B. Carpentieri, F. Zollo, and F. Vella, \u201cOn the anatomy of predictive models for accelerating GPU convolution kernels and beyond,\u201d ACM Trans. Archit. Code Optim., vol.18, no.1, Jan. 2021. 10.1145\/3434402","DOI":"10.1145\/3434402"},{"key":"5","doi-asserted-by":"publisher","unstructured":"[5] J.J. Dongarra, J. Du Croz, S. Hammarling, and I.S. Duff, \u201cA set of level 3 basic linear algebra subprograms,\u201d ACM Trans. Math. Softw., vol.16, no.1, pp.1-17, March 1990. 10.1145\/77626.79170","DOI":"10.1145\/77626.79170"},{"key":"6","unstructured":"[6] P. Warden, \u201cWhy GEMM is at the heart of deep learning,\u201d https:\/\/shorturl.at\/htQW8, 2015."},{"key":"7","doi-asserted-by":"crossref","unstructured":"[7] A. Vasudevan, A. Anderson, and D. Gregg, \u201cParallel multi channel convolution using general matrix multiplication,\u201d 2017 IEEE 28th Int. Conf. Application-specific Systems, Architectures and Processors (ASAP), pp.19-24, 2017. 10.1109\/ASAP.2017.7995254","DOI":"10.1109\/ASAP.2017.7995254"},{"key":"8","unstructured":"[8] H. Kung and C. Leiserson, \u201cAlgorithms for VLSI processor arrays,\u201d in Introduction to VLSI Systems, ed. C. Mead and L. Conway, ch. 8, pp.271-292, Addison-Wesley, Reading, MA, 1980."},{"key":"9","doi-asserted-by":"publisher","unstructured":"[9] H.T. Kung, \u201cWhy Systolic Architectures?,\u201d Computer, vol.15, no.1, pp.37-46, Jan. 1982. 10.1109\/MC.1982.1653825","DOI":"10.1109\/MC.1982.1653825"},{"key":"10","doi-asserted-by":"crossref","unstructured":"[10] N.P. Jouppi, C. Young, N. Patil, D. Patterson, G. Agrawal, R. Bajwa, S. Bates, S. Bhatia, N. Boden, A. Borchers, R. Boyle, P.-l. Cantin, C. Chao, C. Clark, J. Coriell, M. Daley, M. Dau, J. Dean, B. Gelb, T.V. Ghaemmaghami, R. Gottipati, W. Gulland, R. Hagmann, C. Richard Ho, D. Hogberg, J. Hu, R. Hundt, D. Hurt, J. Ibarz, A. Jaffey, A. Jaworski, A. Kaplan, H. Khaitan, D. Killebrew, A. Koch, N. Kumar, S. Lacy, J. Laudon, J. Law, D. Le, C. Leary, Z. Liu, K. Lucke, A. Lundin, G. MacKean, A. Maggiore, M. Mahony, K. Miller, R. Nagarajan, R. Narayanaswami, R. Ni, K. Nix, T. Norrie, M. Omernick, N. Penukonda, A. Phelps, J. Ross, M. Ross, A. Salek, E. Samadiani, C. Severn, G. Sizikov, M. Snelham, J. Souter, D. Steinberg, A. Swing, M. Tan, G. Thorson, B. Tian, H. Toma, E. Tuttle, V. Vasudevan, R. Walter, W. Wang, E. Wilcox, and D.H. Yoon, \u201cIn-datacenter performance analysis of a tensor processing unit,\u201d Proc. 44th Annual Int. Symp. Computer Architecture, ISCA &apos;17, pp.1-12, ACM, New York, NY, USA, June 2017. 10.1145\/3079856.3080246","DOI":"10.1145\/3079856.3080246"},{"key":"11","unstructured":"[11] Google&apos;s Cloud TPUs, \u201cSystem architecture,\u201d https:\/\/shorturl.at\/huINS, 2019."},{"key":"12","doi-asserted-by":"publisher","unstructured":"[12] N.P. Jouppi, D.H. Yoon, G. Kurian, S. Li, N. Patil, J. Laudon, C. Young, and D. Patterson, \u201cA domain-specific supercomputer for training deep neural networks,\u201d Commun. ACM, vol.63, no.7, pp.67-78, June 2020. 10.1145\/3360307","DOI":"10.1145\/3360307"},{"key":"13","doi-asserted-by":"crossref","unstructured":"[13] N.P. Jouppi, D.H. Yoon, M. Ashcraft, M. Gottscho, T.B. Jablin, G. Kurian, J. Laudon, S. Li, P. Ma, X. Ma, T. Norrie, N. Patil, S. Prasad, C. Young, Z. Zhou, and D. Patterson, \u201cTen lessons from three generations shaped Google&apos;s TPUv4i: Industrial product,\u201d 2021 ACM\/IEEE 48th Annual Int. Symp. Computer Architecture (ISCA), pp.1-14, IEEE Computer Society, Los Alamitos, CA, USA, June 2021. 10.1109\/ISCA52012.2021.00010","DOI":"10.1109\/ISCA52012.2021.00010"},{"key":"14","unstructured":"[14] H. Genc, A. Haj-Ali, V. Iyer, A. Amid, H. Mao, J. Wright, C. Schmidt, J. Zhao, A.J. Ou, M. Banister, Y.S. Shao, B. Nikolic, I. Stoica, and K. Asanovic, \u201cGemmini: An agile systolic array generator enabling systematic evaluations of deep-learning architectures,\u201d CoRR, vol.abs\/1911.09925, 2019."},{"key":"15","unstructured":"[15] C. Mead and L. Conway, \u201cIntroduction to VLSI systems,\u201d Reading, MA, Addison-Wesley Publishing Co., 1980. 426 p., vol.-1, Jan. 1980."},{"key":"16","unstructured":"[16] S. Sedukhin, \u201cDesign and analysis of systolic algorithms and structures,\u201d Programming and Computer Software, vol.17, no.2, pp.73-88, 1992."},{"key":"17","doi-asserted-by":"publisher","unstructured":"[17] R.C. Agarwal, F.G. Gustavson, and M. Zubair, \u201cA high-performance matrix-multiplication algorithm on a distributed-memory parallel computer, using overlapped communication,\u201d IBM J. Res. Dev., vol.38, no.6, pp.673-681, Nov. 1994. 10.1147\/rd.386.0673","DOI":"10.1147\/rd.386.0673"},{"key":"18","unstructured":"[18] R.A. van de Geijn and J. Watts, \u201cSUMMA: Scalable universal matrix multiplication algorithm,\u201d Tech. Rep., University of Texas at Austin, USA, TR-95-13, 1995. 10.1002\/(SICI)1096-9128(199704)9:4%3C255::AID-CPE250%3E3.0.CO;2-2"},{"key":"19","doi-asserted-by":"publisher","unstructured":"[19] E. Talpes, D.D. Sarma, G. Venkataramanan, P. Bannon, B. McGee, B. Floering, A. Jalote, C. Hsiong, S. Arora, A. Gorti, and G.S. Sachdev, \u201cCompute solution for Tesla&apos;s full self-driving computer,\u201d IEEE Micro, vol.40, no.2, pp.25-35, March-April 2020. 10.1109\/MM.2020.2975764","DOI":"10.1109\/MM.2020.2975764"},{"key":"20","doi-asserted-by":"crossref","unstructured":"[20] H. Kwon, P. Chatarasi, M. Pellauer, A. Parashar, V. Sarkar, and T. Krishna, \u201cUnderstanding reuse, performance, and hardware cost of DNN dataflow: A data-centric approach,\u201d Proc. 52nd Annual IEEE\/ACM Int. Symp. Microarchitecture, MICRO &apos;52, pp.754-768, Association for Computing Machinery, New York, NY, USA, 2019. 10.1145\/3352460.3358252","DOI":"10.1145\/3352460.3358252"},{"key":"21","doi-asserted-by":"crossref","unstructured":"[21] K. He, X. Zhang, S. Ren, and J. Sun, \u201cDeep residual learning for image recognition,\u201d 2016 IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR), pp.770-778, 2016. 10.1109\/CVPR.2016.90","DOI":"10.1109\/CVPR.2016.90"},{"key":"22","unstructured":"[22] T. Liu, M. Chen, M. Zhou, S.S. Du, E. Zhou, and T. Zhao, \u201cTowards understanding the importance of shortcut connections in residual networks,\u201d 2019. 10.48550\/arXiv.1909.04653"},{"key":"23","doi-asserted-by":"crossref","unstructured":"[23] M. Horowitz, \u201cComputing&apos;s energy problem (and what we can do about it),\u201d 2014 IEEE International Solid-State Circuits Conference Digest of Technical Papers (ISSCC), pp.10-14, 2014. 10.1109\/ISSCC.2014.6757323","DOI":"10.1109\/ISSCC.2014.6757323"},{"key":"24","doi-asserted-by":"publisher","unstructured":"[24] B. Murmann, \u201cMixed-signal computing for deep neural network inference,\u201d IEEE Trans. Very Large Scale Integr. (VLSI) Syst., vol.29, no.1, pp.3-13, Jan. 2021. 10.1109\/TVLSI.2020.3020286","DOI":"10.1109\/TVLSI.2020.3020286"},{"key":"25","doi-asserted-by":"publisher","unstructured":"[25] M.H. Sunwoo and J.K. Aggarwal, \u201cA sliding memory plane array processor,\u201d IEEE Trans. Parallel Distrib. Syst., vol.4, no.6, pp.601-612, June 1993. 10.1109\/71.242162","DOI":"10.1109\/71.242162"},{"key":"26","unstructured":"[26] N. Li, Y. Tomioka, and H. Kitazawa, \u201cAn FPGA implementation of deep convolutional neural network using synchronous shift data transfer,\u201d Tech. Rep. 426, IEICE Technical Report (VLD2014-140), Jan. 2015."},{"key":"27","doi-asserted-by":"crossref","unstructured":"[27] S. Sedukhin, K. Matsumoto, and Y. Tomioka, \u201cBrain-inspired co-design of algorithm\/architecture for CNN accelerators,\u201d 8th International Congress on Advanced Applied Informatics, IIAI-AAI 2019, Toyama, Japan, July 7-11, 2019, pp.556-560, IEEE, 2019. 10.1109\/IIAI-AAI.2019.00119","DOI":"10.1109\/IIAI-AAI.2019.00119"},{"key":"28","doi-asserted-by":"crossref","unstructured":"[28] G. Li, F. Li, T. Zhao, and J. Cheng, \u201cBlock convolution: Towards memory-efficient inference of large-scale cnns on FPGA,\u201d 2018 Design, Automation &amp; Test in Europe Conference &amp; Exhibition, DATE 2018, Dresden, Germany, March 19-23, 2018, ed. J. Madsen and A.K. Coskun, pp.1163-1166, IEEE, 2018. 10.23919\/DATE.2018.8342188","DOI":"10.23919\/DATE.2018.8342188"},{"key":"29","unstructured":"[29] Blog, \u201cIteratively load image block-by-block where blocks are partially overlapped,\u201d https:\/\/shorturl.at\/hvDW3, 2019. [Online; accessed 12-June-2021]."},{"key":"30","doi-asserted-by":"crossref","unstructured":"[30] S. Sedukhin, Y. Tomioka, and K. Yamamoto, \u201cIn search of the performance- and energy-efficient CNN accelerators,\u201d 2021 IEEE Symposium in Low-Power and High-Speed Chips (COOL CHIPS), pp.1-6, IEEE Computer Society, Los Alamitos, CA, USA, April 2021. 10.1109\/COOLCHIPS52128.2021.9410350","DOI":"10.1109\/COOLCHIPS52128.2021.9410350"},{"key":"31","doi-asserted-by":"publisher","unstructured":"[31] V. Sze, Y. Chen, T. Yang, and J.S. Emer, Efficient Processing of Deep Neural Networks, Synthesis Lectures on Computer Architecture, Morgan &amp; Claypool Publishers, 2020. 10.2200\/S01004ED1V01Y202004CAC050","DOI":"10.2200\/S01004ED1V01Y202004CAC050"},{"key":"32","doi-asserted-by":"crossref","unstructured":"[32] X. Ding, X. Zhang, N. Ma, J. Han, G. Ding, and J. Sun, \u201cRepVGG: Making VGG-style ConvNets great again,\u201d CoRR, vol.abs\/2101.03697, 2021. 10.1109\/CVPR46437.2021.01352","DOI":"10.1109\/CVPR46437.2021.01352"},{"key":"33","unstructured":"[33] S. Alyamkin, M. Ardi, A. Brighton, A.C. Berg, Y. Chen, H.-P. Cheng, B. Chen, Z. Fan, C. Feng, B. Fu, K. Gauen, J. Go, A. Goncharenko, X. Guo, H.H. Nguyen, A. Howard, Y. Huang, D. Kang, J. Kim, A. Kondratyev, S. Lee, S. Lee, J. Lee, Z. Liang, X. Liu, J. Liu, Z. Li, Y. Lu, Y.-H. Lu, D. Malik, E. Park, D. Repin, T. Sheng, L. Shen, F. Sun, D. Svitov, G.K. Thiruvathukal, B. Zhang, J. Zhang, X. Zhang, and S. Zhuo, \u201c2018 low-power image recognition challenge,\u201d CoRR, vol.abs\/1810.01732, 2018. 10.48550\/arXiv.1810.01732"},{"key":"34","doi-asserted-by":"crossref","unstructured":"[34] N. Shazeer, K. Fatahalian, W.R. Mark, and R.T. Mullapudi, \u201cHydranets: Specialized dynamic architectures for efficient inference,\u201d 2018 IEEE\/CVF Conf. Comput. Vis. Pattern Recognit., pp.8080-8089, 2018. 10.1109\/CVPR.2018.00843","DOI":"10.1109\/CVPR.2018.00843"},{"key":"35","unstructured":"[35] T. Garipov, D. Podoprikhin, A. Novikov, and D.P. Vetrov, \u201cUltimate tensorization: compressing convolutional and FC layers alike,\u201d http:\/\/arxiv.org\/abs\/1611.03214, 2016. 10.48550\/arXiv.1611.03214"},{"key":"36","doi-asserted-by":"publisher","unstructured":"[36] D.I. Moldovan, \u201cOn the design of algorithms for vlsi systolic arrays,\u201d Proc. IEEE, vol.71, no.1, pp.113-120, Jan. 1983. 10.1109\/PROC.1983.12532","DOI":"10.1109\/PROC.1983.12532"},{"key":"37","doi-asserted-by":"publisher","unstructured":"[37] P. Quinton, \u201cAutomatic synthesis of systolic arrays from uniform recurrent equations,\u201d SIGARCH Comput. Archit. News, vol.12, no.3, pp.208-214, Jan. 1984. 10.1145\/773453.808184","DOI":"10.1145\/773453.808184"},{"key":"38","doi-asserted-by":"crossref","unstructured":"[38] P.R. Cappello and K. Steiglitz, \u201cSelecting systolic designs using linear transformations of space-time,\u201d Real-Time Signal Processing VII, ed. K. Bromley, pp.75-85, International Society for Optics and Photonics, SPIE, 1984. 10.1117\/12.944011","DOI":"10.1117\/12.944011"},{"key":"39","doi-asserted-by":"publisher","unstructured":"[39] S. Kung, \u201cVLSI array processors,\u201d IEEE ASSP Magazine, vol.2, no.3, pp.4-22, July 1985. 10.1109\/MASSP.1985.1163741","DOI":"10.1109\/MASSP.1985.1163741"},{"key":"40","doi-asserted-by":"crossref","unstructured":"[40] S.G. Sedukhin, A.S. Zekri, and T. Myiazaki, \u201cOrbital algorithms and unified array processor for computing 2d separable transforms,\u201d 2010 39th Int. Conf. Parallel Processing Workshops, pp.127-134, 2010. 10.1109\/ICPPW.2010.29","DOI":"10.1109\/ICPPW.2010.29"},{"key":"41","doi-asserted-by":"publisher","unstructured":"[41] X. Zhang, J. Zou, K. He, and J. Sun, \u201cAccelerating very deep convolutional networks for classification and detection,\u201d IEEE Trans. Pattern Anal. Mach. Intell., vol.38, no.10, pp.1943-1955, Oct. 2016. 10.1109\/TPAMI.2015.2502579","DOI":"10.1109\/TPAMI.2015.2502579"},{"key":"42","unstructured":"[42] J. Redmon and A. Farhadi, \u201cYOLOv3: An incremental improvement,\u201d http:\/\/arxiv.org\/abs\/1804.02767, 2018. 10.48550\/arXiv.1804.02767"},{"key":"43","unstructured":"[43] K. Yamamoto and K. Maeno, \u201cPCAS: pruning channels with attention statistics for deep network compression,\u201d 30th British Machine Vision Conference 2019, BMVC 2019, Cardiff, UK, Sept. 9-12, 2019, p.138, BMVA Press, 2019."},{"key":"44","doi-asserted-by":"publisher","unstructured":"[44] Y. Chen, T. Krishna, J.S. Emer, and V. Sze, \u201cEyeriss: An energy-efficient reconfigurable accelerator for deep convolutional neural networks,\u201d IEEE J. Solid State Circuits, vol.52, no.1, pp.127-138, Jan. 2017. 10.1109\/JSSC.2016.2616357","DOI":"10.1109\/JSSC.2016.2616357"},{"key":"45","doi-asserted-by":"crossref","unstructured":"[45] A. Bytyn, R. Leupers, and G. Ascheid, \u201cAn application-specific VLIW processor with vector instruction set for CNN acceleration,\u201d IEEE Int. Symp. Circuits and Systems, ISCAS 2019, Sapporo, Japan, May 26-29, 2019, pp.1-5, IEEE, 2019. 10.1109\/ISCAS.2019.8702357","DOI":"10.1109\/ISCAS.2019.8702357"},{"key":"46","doi-asserted-by":"publisher","unstructured":"[46] T. Norrie, N. Patil, D.H. Yoon, G. Kurian, S. Li, J. Laudon, C. Young, N. Jouppi, and D. Patterson, \u201cThe design process for Google&apos;s training chips: TPUv2 and TPUv3,\u201d IEEE Micro, vol.41, no.2, pp.56-63, March-April 2021. 10.1109\/MM.2021.3058217","DOI":"10.1109\/MM.2021.3058217"},{"key":"47","doi-asserted-by":"crossref","unstructured":"[47] M. Alwani, H. Chen, M. Ferdman, and P. Milder, \u201cFused-layer CNN accelerators,\u201d 2016 49th Annual IEEE\/ACM Int. Symp. Microarchitecture (MICRO), pp.1-12, 2016. 10.1109\/MICRO.2016.7783725","DOI":"10.1109\/MICRO.2016.7783725"},{"key":"48","doi-asserted-by":"crossref","unstructured":"[48] P. Seitz, \u201cSmart image sensors: an emerging key technology for advanced optical measurement and microsystems,\u201d Micro-Optical Technologies for Measurement, Sensors, and Microsystems, ed. O.M. Parriaux, pp.244-255, International Society for Optics and Photonics, SPIE, 1996. 10.1117\/12.248493","DOI":"10.1117\/12.248493"},{"key":"49","doi-asserted-by":"publisher","unstructured":"[49] E.H.M. Heijne, \u201cGigasensors for an attoscope: Catching quanta in CMOS,\u201d IEEE Solid-State Circuits Society Newsletter, vol.13, no.4, pp.28-34, 2008. 10.1109\/N-SSC.2008.4785820","DOI":"10.1109\/N-SSC.2008.4785820"}],"container-title":["IEICE Transactions on Electronics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.jstage.jst.go.jp\/article\/transele\/E105.C\/6\/E105.C_2021LHP0003\/_pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,6,4]],"date-time":"2022-06-04T04:14:24Z","timestamp":1654316064000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.jstage.jst.go.jp\/article\/transele\/E105.C\/6\/E105.C_2021LHP0003\/_article"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,6,1]]},"references-count":49,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2022]]}},"URL":"https:\/\/doi.org\/10.1587\/transele.2021lhp0003","relation":{},"ISSN":["0916-8524","1745-1353"],"issn-type":[{"value":"0916-8524","type":"print"},{"value":"1745-1353","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,6,1]]},"article-number":"2021LHP0003"}}