{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,11,21]],"date-time":"2025-11-21T11:29:53Z","timestamp":1763724593202,"version":"3.41.0"},"reference-count":73,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2022,10,13]],"date-time":"2022-10-13T00:00:00Z","timestamp":1665619200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/100000181","name":"Air Force Office of Scientific Research","doi-asserted-by":"crossref","award":["FA9550-19-1-0277"],"award-info":[{"award-number":["FA9550-19-1-0277"]}],"id":[{"id":"10.13039\/100000181","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["J. Emerg. Technol. Comput. Syst."],"published-print":{"date-parts":[[2022,10,31]]},"abstract":"<jats:p>\n            The acceleration of\n            <jats:bold>Deep Neural Networks (DNNs)<\/jats:bold>\n            has attracted much attention in research. Many critical real-time applications benefit from DNN accelerators but are limited by their compute-intensive nature. This work introduces an accelerator for\n            <jats:bold>Convolutional Neural Network (CNN)<\/jats:bold>\n            , based on a hybrid optoelectronic computing architecture and\n            <jats:bold>residue number system (RNS)<\/jats:bold>\n            . The RNS reduces the optical critical path and lowers the power requirements. In addition, the\n            <jats:bold>wavelength division multiplexing (WDM)<\/jats:bold>\n            allows high-speed operation at the system level by enabling high-level parallelism. The proposed RNS compute modules use one-hot encoding, and thus enable fast switching between the electrical and optical domains. We propose a new architecture that combines residue electrical adders and optical multipliers as the matrix-vector multiplication unit. Moreover, we enhance the implementation of different CNN computational kernels using WDM-enabled RNS based integrated photonics. The area and power efficiency of the proposed accelerator are 0.39 TOPS\/s\/mm\n            <jats:sup>2<\/jats:sup>\n            and 3.22 TOPS\/s\/W, respectively. In terms of computation capability, the proposed chip is 12.7\u00d7 and 4.02\u00d7 better than other optical implementation and memristor implementation, respectively. Our experimental evaluation using DNN benchmarks illustrates that our architecture can perform on average more than 72 times faster than GPU under the same power budget.\n          <\/jats:p>","DOI":"10.1145\/3550273","type":"journal-article","created":{"date-parts":[[2022,7,21]],"date-time":"2022-07-21T12:17:11Z","timestamp":1658405831000},"page":"1-26","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":11,"title":["A Deep Neural Network Accelerator using Residue Arithmetic in a Hybrid Optoelectronic System"],"prefix":"10.1145","volume":"18","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-2943-6305","authenticated-orcid":false,"given":"Jiaxin","family":"Peng","sequence":"first","affiliation":[{"name":"The George Washington University, Washington, D.C., USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8806-8146","authenticated-orcid":false,"given":"Yousra","family":"Alkabani","sequence":"additional","affiliation":[{"name":"Halmstad University, Halmstad, Sweden"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1372-5361","authenticated-orcid":false,"given":"Krunal","family":"Puri","sequence":"additional","affiliation":[{"name":"The George Washington University, Washington, D.C., USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7051-6761","authenticated-orcid":false,"given":"Xiaoxuan","family":"Ma","sequence":"additional","affiliation":[{"name":"The George Washington University, Washington, D.C., USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5152-4766","authenticated-orcid":false,"given":"Volker","family":"Sorger","sequence":"additional","affiliation":[{"name":"The George Washington University, Washington, D.C., USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9687-7939","authenticated-orcid":false,"given":"Tarek","family":"El-Ghazawi","sequence":"additional","affiliation":[{"name":"The George Washington University, Washington, D.C., USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2022,10,13]]},"reference":[{"key":"e_1_3_1_2_2","first-page":"655","volume-title":"Proceedings of SAI Intelligent Systems Conference","author":"Ahmed Hossam O.","year":"2018","unstructured":"Hossam O. Ahmed, Maged Ghoneima, and Mohamed Dessouky. 2018. High-speed 2D parallel MAC unit hardware accelerator for convolutional neural network. In Proceedings of SAI Intelligent Systems Conference. Springer, 655\u2013663."},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neunet.2004.11.006"},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.1145\/3297858.3304049"},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11082-016-0412-6"},{"key":"e_1_3_1_6_2","doi-asserted-by":"publisher","DOI":"10.1109\/JSTQE.2019.2945540"},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.1002\/j.1538-7305.1964.tb04102.x"},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.1038\/s41377-020-0272-5"},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1142\/S0129626497000292"},{"key":"e_1_3_1_10_2","first-page":"269","volume-title":"ACM Sigplan Notices","author":"Chen Tianshi","year":"2014","unstructured":"Tianshi Chen, Zidong Du, Ninghui Sun, Jia Wang, Chengyong Wu, Yunji Chen, and Olivier Temam. 2014. DianNao: A small-footprint high-throughput accelerator for ubiquitous machine-learning. In ACM Sigplan Notices, Vol. 49. ACM, 269\u2013284."},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2014.58"},{"key":"e_1_3_1_12_2","volume-title":"Twenty-Second International Joint Conference on Artificial Intelligence","author":"Ciresan Dan Claudiu","year":"2011","unstructured":"Dan Claudiu Ciresan, Ueli Meier, Jonathan Masci, Luca Maria Gambardella, and J\u00fcrgen Schmidhuber. 2011. Flexible, high performance convolutional neural networks for image classification. In Twenty-Second International Joint Conference on Artificial Intelligence."},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-11179-7_36"},{"key":"e_1_3_1_14_2","unstructured":"NVIDIA Corporation. 2019. NVIDIA TESLA V100 TENSOR CORE GPU. (2019). https:\/\/www.nvidia.com\/en-us\/data-center\/tesla-v100\/."},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.1049\/el.2013.3098"},{"key":"e_1_3_1_16_2","volume-title":"239th ECS Meeting with the 18th International Meeting on Chemical Sensors (IMCS)","author":"Dong Boqun","year":"2021","unstructured":"Boqun Dong, Mengfei Liu, and Yangyang Zhao. 2021. Simulation of a nano plasmonic pillar-based optical sensor with AI-assisted signal processing. In 239th ECS Meeting with the 18th International Meeting on Chemical Sensors (IMCS) (May 30-June 3, 2021). ECS."},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1038\/s41586-020-03070-1"},{"key":"e_1_3_1_18_2","first-page":"JTh2B\u20138","volume-title":"CLEO: Applications and Technology","author":"Feng Chenghao","year":"2020","unstructured":"Chenghao Feng, Zheng Zhao, Zhoufeng Ying, Jiaqi Gu, David Z. Pan, and Ray T. Chen. 2020. Compact design of on-chip Elman optical recurrent neural network. In CLEO: Applications and Technology. Optical Society of America, JTh2B\u20138."},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMAG.2011.2150238"},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.asej.2020.05.005"},{"issue":"1","key":"e_1_3_1_21_2","first-page":"501","article-title":"Comparative analysis of energy-efficient low power 1-bit full adders at 120nm technology","volume":"4","author":"Goyal Candy","year":"2012","unstructured":"Candy Goyal and Ashish Kumar. 2012. Comparative analysis of energy-efficient low power 1-bit full adders at 120nm technology. International Journal of Advances in Engineering & Technology 4, 1 (2012), 501.","journal-title":"International Journal of Advances in Engineering & Technology"},{"key":"e_1_3_1_22_2","article-title":"A survey of FPGA-based neural network accelerator","author":"Guo Kaiyuan","year":"2017","unstructured":"Kaiyuan Guo, Shulin Zeng, Jincheng Yu, Yu Wang, and Huazhong Yang. 2017. A survey of FPGA-based neural network accelerator. arXiv preprint arXiv:1712.08934 (2017).","journal-title":"arXiv preprint arXiv:1712.08934"},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-04274-4_39"},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1364\/AO.18.000149"},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.1109\/OIC.2015.7115680"},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/PDGC.2012.6449931"},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.1145\/3360307"},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.procs.2014.11.068"},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.1109\/JPHOT.2020.3037834"},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.1109\/5.726791"},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.1126\/science.aat8084"},{"key":"e_1_3_1_33_2","first-page":"369","volume-title":"ACM SIGARCH Computer Architecture News","author":"Liu Daofu","year":"2015","unstructured":"Daofu Liu, Tianshi Chen, Shaoli Liu, Jinhong Zhou, Shengyuan Zhou, Olivier Teman, Xiaobing Feng, Xuehai Zhou, and Yunji Chen. 2015. PuDianNao: A polyvalent machine learning accelerator. In ACM SIGARCH Computer Architecture News, Vol. 43. ACM, 369\u2013381."},{"key":"e_1_3_1_34_2","doi-asserted-by":"publisher","DOI":"10.23919\/DATE.2019.8715195"},{"key":"e_1_3_1_35_2","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2007.30"},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","DOI":"10.1109\/FPL.2015.7293933"},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.23919\/DATE.2017.7927113"},{"key":"e_1_3_1_38_2","first-page":"1","volume-title":"2018 IEEE International Symposium on Circuits and Systems (ISCAS)","author":"Olsen Eric B.","year":"2018","unstructured":"Eric B. Olsen. 2018. RNS hardware matrix multiplier for high precision neural network acceleration: \u201cRNS TPU\u201d. In 2018 IEEE International Symposium on Circuits and Systems (ISCAS). IEEE, 1\u20135."},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","DOI":"10.5555\/1543620"},{"key":"e_1_3_1_40_2","doi-asserted-by":"publisher","DOI":"10.1038\/nature08364"},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","DOI":"10.1109\/MDAT.2020.2982628"},{"key":"e_1_3_1_42_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICRC.2019.8914700"},{"key":"e_1_3_1_43_2","doi-asserted-by":"publisher","DOI":"10.1145\/3404397.3404467"},{"key":"e_1_3_1_44_2","doi-asserted-by":"publisher","DOI":"10.1364\/OL.43.002026"},{"key":"e_1_3_1_45_2","doi-asserted-by":"publisher","DOI":"10.1109\/AE.2016.7577275"},{"key":"e_1_3_1_46_2","doi-asserted-by":"publisher","DOI":"10.4018\/978-1-7998-0261-7.ch005"},{"key":"e_1_3_1_47_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSII.2020.2999458"},{"issue":"3","key":"e_1_3_1_48_2","doi-asserted-by":"crossref","first-page":"56","DOI":"10.26634\/jfet.14.3.15149","article-title":"Artificial intelligence in autonomous vehicles\u2014a literature review","volume":"14","author":"Sagar Vinyas D.","year":"2019","unstructured":"Vinyas D. Sagar and T. S. Nanjundeswaraswamy. 2019. Artificial intelligence in autonomous vehicles\u2014a literature review. i-Manager\u2019s Journal on Future Engineering and Technology 14, 3 (2019), 56.","journal-title":"i-Manager\u2019s Journal on Future Engineering and Technology"},{"key":"e_1_3_1_49_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICRC.2018.8638592"},{"key":"e_1_3_1_50_2","doi-asserted-by":"publisher","DOI":"10.1021\/acsphotonics.8b00525"},{"key":"e_1_3_1_51_2","doi-asserted-by":"publisher","DOI":"10.1109\/PROC.1984.12937"},{"key":"e_1_3_1_52_2","doi-asserted-by":"publisher","DOI":"10.1145\/3007787.3001139"},{"key":"e_1_3_1_53_2","doi-asserted-by":"publisher","DOI":"10.1038\/nphoton.2017.93"},{"key":"e_1_3_1_54_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA52012.2021.00072"},{"key":"e_1_3_1_55_2","doi-asserted-by":"publisher","DOI":"10.1109\/DCAS.2018.8620116"},{"key":"e_1_3_1_56_2","article-title":"Very deep convolutional networks for large-scale image recognition","author":"Simonyan Karen","year":"2014","unstructured":"Karen Simonyan and Andrew Zisserman. 2014. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556 (2014).","journal-title":"arXiv preprint arXiv:1409.1556"},{"key":"e_1_3_1_57_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2017.55"},{"issue":"2","key":"e_1_3_1_58_2","first-page":"1","article-title":"Hybrid photonic-plasmonic nonblocking broadband 5 \u00d7 5 router for optical networks","volume":"10","author":"Sun Shuai","year":"2017","unstructured":"Shuai Sun, Vikram K. Narayana, Ibrahim Sarpkaya, Joseph Crandall, Richard A. Soref, Hamed Dalir, Tarek El-Ghazawi, and Volker J. Sorger. 2017. Hybrid photonic-plasmonic nonblocking broadband 5 \u00d7 5 router for optical networks. IEEE Photonics Journal 10, 2 (2017), 1\u201312.","journal-title":"IEEE Photonics Journal"},{"key":"e_1_3_1_59_2","doi-asserted-by":"publisher","DOI":"10.1109\/DAC18074.2021.9586161"},{"key":"e_1_3_1_60_2","volume-title":"Residue Arithmetic and its Applications to Computer Technology","author":"Szabo Nicholas S.","year":"1967","unstructured":"Nicholas S. Szabo and Richard I. Tanaka. 1967. Residue Arithmetic and its Applications to Computer Technology. McGraw-Hill."},{"key":"e_1_3_1_61_2","doi-asserted-by":"publisher","DOI":"10.1109\/I2CT.2018.8529482"},{"key":"e_1_3_1_62_2","doi-asserted-by":"publisher","DOI":"10.1364\/AO.18.002812"},{"key":"e_1_3_1_63_2","doi-asserted-by":"publisher","DOI":"10.1038\/s41598-017-07754-z"},{"key":"e_1_3_1_64_2","doi-asserted-by":"publisher","DOI":"10.1088\/1361-6463\/aac8a5"},{"key":"e_1_3_1_65_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICECS.2017.8292029"},{"key":"e_1_3_1_66_2","unstructured":"International Communication Union. 2012. G.694.1 : Spectral grids for WDM applications: DWDM frequency grid. (2012). https:\/\/www.itu.int\/rec\/T-REC-G.694.1-201202-I\/en."},{"key":"e_1_3_1_67_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA52012.2021.00021"},{"key":"e_1_3_1_68_2","doi-asserted-by":"publisher","DOI":"10.1145\/3316781.3317753"},{"key":"e_1_3_1_69_2","doi-asserted-by":"publisher","DOI":"10.1038\/s41586-020-03063-0"},{"key":"e_1_3_1_70_2","doi-asserted-by":"publisher","DOI":"10.1109\/JSTQE.2018.2836955"},{"key":"e_1_3_1_71_2","doi-asserted-by":"publisher","DOI":"10.1145\/2684746.2689060"},{"key":"e_1_3_1_72_2","doi-asserted-by":"crossref","first-page":"10","DOI":"10.1109\/EMC2-NIPS53020.2019.00010","volume-title":"2019 Fifth Workshop on Energy Efficient Machine Learning and Cognitive Computing-NeurIPS Edition (EMC2-NIPS)","author":"Zhang Tianyi","year":"2019","unstructured":"Tianyi Zhang, Zhiqiu Lin, Guandao Yang, and Christopher De Sa. 2019. QPyTorch: A low-precision arithmetic simulation framework. In 2019 Fifth Workshop on Energy Efficient Machine Learning and Cognitive Computing-NeurIPS Edition (EMC2-NIPS). IEEE, 10\u201313."},{"key":"e_1_3_1_73_2","doi-asserted-by":"publisher","DOI":"10.1038\/s41467-019-10282-1"},{"key":"e_1_3_1_74_2","doi-asserted-by":"publisher","DOI":"10.23919\/DATE48585.2020.9116494"}],"container-title":["ACM Journal on Emerging Technologies in Computing Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3550273","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3550273","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T18:43:23Z","timestamp":1750272203000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3550273"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,10,13]]},"references-count":73,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2022,10,31]]}},"alternative-id":["10.1145\/3550273"],"URL":"https:\/\/doi.org\/10.1145\/3550273","relation":{},"ISSN":["1550-4832","1550-4840"],"issn-type":[{"type":"print","value":"1550-4832"},{"type":"electronic","value":"1550-4840"}],"subject":[],"published":{"date-parts":[[2022,10,13]]},"assertion":[{"value":"2021-12-28","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2022-07-05","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2022-10-13","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}