{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,13]],"date-time":"2026-03-13T14:56:40Z","timestamp":1773413800706,"version":"3.50.1"},"reference-count":54,"publisher":"Association for Computing Machinery (ACM)","issue":"5s","license":[{"start":{"date-parts":[[2023,9,9]],"date-time":"2023-09-09T00:00:00Z","timestamp":1694217600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Embed. Comput. Syst."],"published-print":{"date-parts":[[2023,10,31]]},"abstract":"<jats:p>Deep Neural Networks (DNNs) have demonstrated great success in many fields such as image recognition and text analysis. However, the ever-increasing sizes of both DNN models and training datasets make deep leaning extremely computation- and memory-intensive. Recently, photonic computing has emerged as a promising technology for accelerating DNNs. While the design of photonic accelerators for DNN inference and forward propagation of DNN training has been widely investigated, the architectural acceleration for equally important backpropagation of DNN training has not been well studied. In this paper, we propose a novel silicon photonic-based backpropagation accelerator for high performance DNN training. Specifically, a general-purpose photonic gradient descent unit named STADIA is designed to implement the multiplication, accumulation, and subtraction operations required for computing gradients using mature optical devices including Mach-Zehnder Interferometer (MZI) and Mircoring Resonator (MRR), which can significantly reduce the training latency and improve the energy efficiency of backpropagation. To demonstrate efficient parallel computing, we propose a STADIA-based backpropagation acceleration architecture and design a dataflow by using wavelength-division multiplexing (WDM). We analyze the precision of STADIA by quantifying the precision limitations imposed by losses and noises. Furthermore, we evaluate STADIA with different element sizes by analyzing the power, area and time delay for photonic accelerators based on DNN models such as AlexNet, VGG19 and ResNet. Simulation results show that the proposed architecture STADIA can achieve significant improvement by 9.7\u00d7 in time efficiency and 147.2\u00d7 in energy efficiency, compared with the most advanced optical-memristor based backpropagation accelerator.<\/jats:p>","DOI":"10.1145\/3607920","type":"journal-article","created":{"date-parts":[[2023,9,9]],"date-time":"2023-09-09T13:33:18Z","timestamp":1694266398000},"page":"1-23","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["STADIA: Photonic Stochastic Gradient Descent for Neural Network Accelerators"],"prefix":"10.1145","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-9520-0229","authenticated-orcid":false,"given":"Chengpeng","family":"Xia","sequence":"first","affiliation":[{"name":"University of Otago, New Zealand"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7006-2459","authenticated-orcid":false,"given":"Yawen","family":"Chen","sequence":"additional","affiliation":[{"name":"University of Otago, New Zealand"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3752-0806","authenticated-orcid":false,"given":"Haibo","family":"Zhang","sequence":"additional","affiliation":[{"name":"University of Otago, New Zealand"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6470-9794","authenticated-orcid":false,"given":"Jigang","family":"Wu","sequence":"additional","affiliation":[{"name":"Guangdong University of Technology, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2023,9,9]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.1364\/OE.20.002911"},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1063\/5.0070992"},{"issue":"1","key":"e_1_3_1_4_2","first-page":"1","article-title":"Optical RAM and integrated optical memories: A survey","volume":"9","author":"Alexoudi Theoni","year":"2020","unstructured":"Theoni Alexoudi, George Theodore Kanellos, and Nikos Pleros. 2020. Optical RAM and integrated optical memories: A survey. Light: Science & Applications 9, 1 (2020), 1\u201316.","journal-title":"Light: Science & Applications"},{"issue":"2","key":"e_1_3_1_5_2","doi-asserted-by":"crossref","first-page":"973","DOI":"10.1109\/TMTT.2017.2752170","article-title":"A linear differential transimpedance amplifier for 100-Gb\/s integrated coherent optical fiber receivers","volume":"66","author":"Awny Ahmed","year":"2017","unstructured":"Ahmed Awny, Rajasekhar Nagulapalli, Marcel Kroh, Jan Hoffmann, Patrick Runge, Daniel Micusik, Gunter Fischer, Ahmet Cagri Ulusoy, Minsu Ko, and Dietmar Kissinger. 2017. A linear differential transimpedance amplifier for 100-Gb\/s integrated coherent optical fiber receivers. IEEE Transactions on Microwave Theory and Techniques 66, 2 (2017), 973\u2013986.","journal-title":"IEEE Transactions on Microwave Theory and Techniques"},{"issue":"12","key":"e_1_3_1_6_2","doi-asserted-by":"crossref","first-page":"2101","DOI":"10.1109\/JPROC.2018.2854372","article-title":"The emergence of silicon photonics as a flexible technology platform","volume":"106","author":"Chen Xia","year":"2018","unstructured":"Xia Chen, Milan M Milosevic, Stevan Stankovi\u0107, Scott Reynolds, Thalia Dominguez Bucio, Ke Li, David J Thomson, Frederic Gardes, and Graham T Reed. 2018. The emergence of silicon photonics as a flexible technology platform. Proc. IEEE 106, 12 (2018), 2101\u20132116.","journal-title":"Proc. IEEE"},{"key":"e_1_3_1_7_2","article-title":"cudnn: Efficient primitives for deep learning","author":"Chetlur Sharan","year":"2014","unstructured":"Sharan Chetlur, Cliff Woolley, Philippe Vandermersch, Jonathan Cohen, John Tran, Bryan Catanzaro, and Evan Shelhamer. 2014. cudnn: Efficient primitives for deep learning. arXiv preprint arXiv:1410.0759 (2014).","journal-title":"arXiv preprint arXiv:1410.0759"},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.1145\/3446212"},{"key":"e_1_3_1_9_2","first-page":"27","volume-title":"Proceedings of the Great Lakes Symposium on VLSI","author":"Dang Dharanidhar","year":"2020","unstructured":"Dharanidhar Dang, Aurosmita Khansama, Rabi Mahapatra, and Debashis Sahoo. 2020. BPhoton-CNN: An ultrafast photonic backpropagation accelerator for deep learning. In Proceedings of the Great Lakes Symposium on VLSI. 27\u201332."},{"key":"e_1_3_1_10_2","first-page":"1","volume-title":"2020 57th ACM\/IEEE Design Automation Conference (DAC)","author":"Dang Dharanidhar","year":"2020","unstructured":"Dharanidhar Dang, Sahar Taheri, Bill Lin, and Debashis Sahoo. 2020. MEMTONIC: A neuromorphic accelerator for energy efficient deep learning. In 2020 57th ACM\/IEEE Design Automation Conference (DAC). IEEE, 1\u20132."},{"key":"e_1_3_1_11_2","first-page":"561","volume-title":"Proceedings of the 44th Annual International Symposium on Computer Architecture","author":"Sa Christopher De","year":"2017","unstructured":"Christopher De Sa, Matthew Feldman, Christopher R\u00e9, and Kunle Olukotun. 2017. Understanding and optimizing asynchronous low-precision stochastic gradient descent. In Proceedings of the 44th Annual International Symposium on Computer Architecture. 561\u2013574."},{"key":"e_1_3_1_12_2","first-page":"1223","article-title":"Large scale distributed deep networks","volume":"25","author":"Dean Jeffrey","year":"2012","unstructured":"Jeffrey Dean, Greg Corrado, Rajat Monga, Kai Chen, Matthieu Devin, Mark Mao, Marc\u2019aurelio Ranzato, Andrew Senior, Paul Tucker, Ke Yang, et\u00a0al. 2012. Large scale distributed deep networks. Advances in Neural Information Processing Systems 25 (2012), 1223\u20131231.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"e_1_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.1109\/MSP.2012.2211477"},{"issue":"6","key":"e_1_3_1_15_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1109\/JSTQE.2018.2836985","article-title":"All-optical reservoir computing on a photonic chip using silicon-based ring resonators","volume":"24","author":"Coarer Florian Denis-Le","year":"2018","unstructured":"Florian Denis-Le Coarer, Marc Sciamanna, Andrew Katumba, Matthias Freiberger, Joni Dambre, Peter Bienstman, and Damien Rontani. 2018. All-optical reservoir computing on a photonic chip using silicon-based ring resonators. IEEE Journal of Selected Topics in Quantum Electronics 24, 6 (2018), 1\u20138.","journal-title":"IEEE Journal of Selected Topics in Quantum Electronics"},{"key":"e_1_3_1_16_2","first-page":"1","volume-title":"39th European Conference and Exhibition on Optical Communication (ECOC 2013)","author":"Descos A","year":"2013","unstructured":"A Descos, C Jany, D Bordel, H Duprez, G Beninca de Farias, P Brianceau, S Menezo, and B Ben Bakir. 2013. Heterogeneously integrated III-V\/Si distributed Bragg reflector laser with adiabatic coupling. In 39th European Conference and Exhibition on Optical Communication (ECOC 2013). IET, 1\u20133."},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1364\/OE.27.014009"},{"key":"e_1_3_1_18_2","unstructured":"William Fedus Barret Zoph and Noam Shazeer. 2021. Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity."},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","DOI":"10.1038\/s41586-020-03070-1"},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11023-020-09548-1"},{"issue":"3","key":"e_1_3_1_21_2","first-page":"B71\u2013B80","article-title":"Backpropagation through nonlinear units for the all-optical training of neural networks","volume":"9","author":"Guo Xianxin","year":"2021","unstructured":"Xianxin Guo, Thomas D Barrett, Zhiming M Wang, and AI Lvovsky. 2021. Backpropagation through nonlinear units for the all-optical training of neural networks. Photonics Research 9, 3 (2021), B71\u2013B80.","journal-title":"Photonics Research"},{"key":"e_1_3_1_22_2","article-title":"Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding","author":"Han Song","year":"2015","unstructured":"Song Han, Huizi Mao, and William J Dally. 2015. Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding. arXiv preprint arXiv:1510.00149 (2015).","journal-title":"arXiv preprint arXiv:1510.00149"},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.13140\/RG.2.2.19501.49124"},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_1_25_2","volume-title":"Introduction to Fiber-optic Communications","author":"Hui Rongqing","year":"2019","unstructured":"Rongqing Hui. 2019. Introduction to Fiber-optic Communications. Academic Press."},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00286"},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1145\/2716282.2716289"},{"key":"e_1_3_1_28_2","article-title":"Photonic Supercomputer For AI: 10X faster, 90% less energy, plus runway for 100X speed boost","author":"Koetsier John","year":"2021","unstructured":"John Koetsier. 2021. Photonic Supercomputer For AI: 10X faster, 90% less energy, plus runway for 100X speed boost. Forbes (2021). https:\/\/www.forbes.com\/sites\/johnkoetsier\/2021\/04\/07\/photonic-supercomputer-for-ai-10x-faster-90-less-energy-plus-runway-for-100x-speed-boost\/?sh=4589d9b67260","journal-title":"Forbes"},{"key":"e_1_3_1_29_2","first-page":"84","article-title":"Imagenet classification with deep convolutional neural networks","volume":"25","author":"Krizhevsky Alex","year":"2012","unstructured":"Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. 2012. Imagenet classification with deep convolutional neural networks. Advances in Neural Information Processing Systems 25 (2012), 84\u201390.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.1109\/5.726791"},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.1515\/nanoph-2022-0109"},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.1126\/science.aat8084"},{"key":"e_1_3_1_33_2","article-title":"A threshold-based bioluminescence detector with a CMOS-integrated photodiode array in 65 nm for a multi-diagnostic ingestible capsule","author":"Liu Qijun","year":"2022","unstructured":"Qijun Liu, Miguel Jimenez, Maria Eugenia Inda, Arslan Riaz, Timur Zirtiloglu, Anantha P Chandrakasan, Timothy K Lu, Giovanni Traverso, Phillip Nadeau, and Rabia Tugce Yazicigil. 2022. A threshold-based bioluminescence detector with a CMOS-integrated photodiode array in 65 nm for a multi-diagnostic ingestible capsule. IEEE Journal of Solid-State Circuits (2022).","journal-title":"IEEE Journal of Solid-State Circuits"},{"key":"e_1_3_1_34_2","first-page":"169","volume-title":"2018 31st International System-on-Chip Conference","author":"Mehrabian Armin","year":"2018","unstructured":"Armin Mehrabian, Yousra Al-Kabani, Volker J Sorger, and Tarek El-Ghazawi. 2018. PCNNA: A photonic convolutional neural network accelerator. In 2018 31st International System-on-Chip Conference. IEEE, 169\u2013173."},{"key":"e_1_3_1_35_2","doi-asserted-by":"publisher","DOI":"10.5555\/2414756"},{"issue":"6643","key":"e_1_3_1_36_2","doi-asserted-by":"crossref","first-page":"398","DOI":"10.1126\/science.ade8450","article-title":"Experimentally realized in situ backpropagation for deep learning in photonic neural networks","volume":"380","author":"Pai Sunil","year":"2023","unstructured":"Sunil Pai, Zhanghao Sun, Tyler W Hughes, Taewon Park, Ben Bartlett, Ian AD Williamson, Momchil Minkov, Maziyar Milanizadeh, Nathnael Abebe, Francesco Morichetti, et\u00a0al. 2023. Experimentally realized in situ backpropagation for deep learning in photonic neural networks. Science 380, 6643 (2023), 398\u2013404.","journal-title":"Science"},{"key":"e_1_3_1_37_2","doi-asserted-by":"crossref","first-page":"290","DOI":"10.1109\/ISVLSI.2014.94","volume-title":"2014 IEEE Computer Society Annual Symposium on VLSI","author":"Shafaei Alireza","year":"2014","unstructured":"Alireza Shafaei, Yanzhi Wang, and Xue Lin. 2014. FinCACTI: Architectural analysis and modeling of caches with deeply-scaled FinFET devices. In 2014 IEEE Computer Society Annual Symposium on VLSI. 290\u2013295."},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","DOI":"10.1145\/3007787.3001139"},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","DOI":"10.1038\/nphoton.2017.93"},{"key":"e_1_3_1_40_2","first-page":"860","volume-title":"ACM\/IEEE 48th Annual International Symposium on Computer Architecture (ISCA)","author":"Shiflett Kyle","year":"2021","unstructured":"Kyle Shiflett, Avinash Karanth, Razvan Bunescu, and Ahmed Louri. 2021. Albireo: Energy-efficient acceleration of convolutional neural networks via silicon photonics. In ACM\/IEEE 48th Annual International Symposium on Computer Architecture (ISCA). 860\u2013873."},{"key":"e_1_3_1_41_2","doi-asserted-by":"crossref","first-page":"474","DOI":"10.1109\/HPCA47549.2020.00046","volume-title":"2020 IEEE International Symposium on High Performance Computer Architecture (HPCA)","author":"Shiflett Kyle","year":"2020","unstructured":"Kyle Shiflett, Dylan Wright, Avinash Karanth, and Ahmed Louri. 2020. PIXEL: Photonic neural network accelerator. In 2020 IEEE International Symposium on High Performance Computer Architecture (HPCA). 474\u2013487."},{"key":"e_1_3_1_42_2","doi-asserted-by":"publisher","DOI":"10.1109\/JPHOT.2019.2952562"},{"key":"e_1_3_1_43_2","first-page":"1","article-title":"Very deep convolutional networks for large-scale image recognition","author":"Simonyan Karen","year":"2014","unstructured":"Karen Simonyan and Andrew Zisserman. 2014. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556 (2014), 1\u201314.","journal-title":"arXiv preprint arXiv:1409.1556"},{"key":"e_1_3_1_44_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCOMM.2009.12.080559"},{"key":"e_1_3_1_45_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2017.55"},{"key":"e_1_3_1_46_2","article-title":"Hybrid 8-bit floating point (HFP8) training and inference for deep neural networks","volume":"32","author":"Sun Xiao","year":"2019","unstructured":"Xiao Sun, Jungwook Choi, Chia-Yu Chen, Naigang Wang, Swagath Venkataramani, Vijayalakshmi Viji Srinivasan, Xiaodong Cui, Wei Zhang, and Kailash Gopalakrishnan. 2019. Hybrid 8-bit floating point (HFP8) training and inference for deep neural networks. Advances in Neural Information Processing Systems 32 (2019).","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_47_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298594"},{"key":"e_1_3_1_48_2","doi-asserted-by":"publisher","DOI":"10.1109\/JLT.2014.2345652"},{"key":"e_1_3_1_49_2","unstructured":"Texas. 2023. Texas Instruments ADS1285 32-Bit Low-Power ADC. https:\/\/nz.mouser.com\/new\/texas-instruments\/ti-ads1285-low-power-adc\/"},{"issue":"10","key":"e_1_3_1_50_2","doi-asserted-by":"crossref","first-page":"14618","DOI":"10.1364\/OE.394783","article-title":"Compact ultrabroad-bandwidth cascaded arrayed waveguide gratings","volume":"28","author":"Wijk Arthur van","year":"2020","unstructured":"Arthur van Wijk, Christopher R Doerr, Zain Ali, Mustafa Karabiyik, and B Imran Akca. 2020. Compact ultrabroad-bandwidth cascaded arrayed waveguide gratings. Optics Express 28, 10 (2020), 14618\u201314626.","journal-title":"Optics Express"},{"issue":"1","key":"e_1_3_1_51_2","doi-asserted-by":"crossref","first-page":"6650","DOI":"10.1038\/s41467-021-26804-9","article-title":"High-performance lasers for fully integrated silicon nitride photonics","volume":"12","author":"Xiang Chao","year":"2021","unstructured":"Chao Xiang, Joel Guo, Warren Jin, Lue Wu, Jonathan Peters, Weiqiang Xie, Lin Chang, Boqiang Shen, Heming Wang, Qi-Fan Yang, et\u00a0al. 2021. High-performance lasers for fully integrated silicon nitride photonics. Nature Communications 12, 1 (2021), 6650.","journal-title":"Nature Communications"},{"issue":"2","key":"e_1_3_1_52_2","doi-asserted-by":"crossref","first-page":"023105","DOI":"10.1088\/1674-4926\/42\/2\/023105","article-title":"A review: Photonics devices, architectures, and algorithms for optical neural computing","volume":"42","author":"Xiang Shuiying","year":"2021","unstructured":"Shuiying Xiang, Yanan Han, Ziwei Song, Xingxing Guo, Yahui Zhang, Zhenxing Ren, Suhong Wang, Yuanting Ma, Weiwen Zou, Bowen Ma, et\u00a0al. 2021. A review: Photonics devices, architectures, and algorithms for optical neural computing. Journal of Semiconductors 42, 2 (2021), 023105.","journal-title":"Journal of Semiconductors"},{"key":"e_1_3_1_53_2","first-page":"79","volume-title":"Proceedings of the 26th International Symposium on High-Performance Parallel and Distributed Computing","author":"Xie Xiaolong","year":"2017","unstructured":"Xiaolong Xie, Wei Tan, Liana L Fong, and Yun Liang. 2017. CuMF_SGD: Parallelized stochastic gradient descent for matrix factorization on GPUs. In Proceedings of the 26th International Symposium on High-Performance Parallel and Distributed Computing. 79\u201392."},{"key":"e_1_3_1_54_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2020.2974744"},{"key":"e_1_3_1_55_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00204"}],"container-title":["ACM Transactions on Embedded Computing Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3607920","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3607920","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T16:38:06Z","timestamp":1750178286000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3607920"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,9,9]]},"references-count":54,"journal-issue":{"issue":"5s","published-print":{"date-parts":[[2023,10,31]]}},"alternative-id":["10.1145\/3607920"],"URL":"https:\/\/doi.org\/10.1145\/3607920","relation":{},"ISSN":["1539-9087","1558-3465"],"issn-type":[{"value":"1539-9087","type":"print"},{"value":"1558-3465","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,9,9]]},"assertion":[{"value":"2023-03-23","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-07-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-09-09","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}