{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,17]],"date-time":"2026-07-17T10:27:20Z","timestamp":1784284040142,"version":"3.55.0"},"reference-count":51,"publisher":"Association for Computing Machinery (ACM)","issue":"6","license":[{"start":{"date-parts":[[2023,11,9]],"date-time":"2023-11-09T00:00:00Z","timestamp":1699488000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"United Kingdom EPSRC","award":["EP\/P010040\/1, EP\/S030069\/1"],"award-info":[{"award-number":["EP\/P010040\/1, EP\/S030069\/1"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Embed. Comput. Syst."],"published-print":{"date-parts":[[2023,11,30]]},"abstract":"<jats:p>\n            The ever-growing computational demands of increasingly complex machine learning models frequently necessitate the use of powerful cloud-based infrastructure for their training. Binary neural networks are known to be promising candidates for on-device inference due to their extreme compute and memory savings over higher-precision alternatives. However, their existing training methods require the concurrent storage of high-precision activations for all layers, generally making learning on memory-constrained devices infeasible. In this article, we demonstrate that the backward propagation operations needed for binary neural network training are strongly robust to quantization, thereby making on-the-edge learning with modern models a practical proposition. We introduce a low-cost binary neural network training strategy exhibiting sizable memory footprint reductions while inducing little to no accuracy loss\n            <jats:italic>vs<\/jats:italic>\n            Courbariaux &amp; Bengio\u2019s standard approach. These decreases are primarily enabled through the retention of activations exclusively in binary format. Against the latter algorithm, our drop-in replacement sees memory requirement reductions of 3\u20135\u00d7, while reaching similar test accuracy (\u00b1 2\u00a0pp) in comparable time, across a range of small-scale models trained to classify popular datasets. We also demonstrate from-scratch ImageNet training of binarized ResNet-18, achieving a 3.78\u00d7 memory reduction. Our work is open-source, and includes the Raspberry Pi-targeted prototype we used to verify our modeled memory decreases and capture the associated energy drops. Such savings will allow for unnecessary cloud offloading to be avoided, reducing latency, increasing energy efficiency, and safeguarding end-user privacy.\n          <\/jats:p>","DOI":"10.1145\/3626100","type":"journal-article","created":{"date-parts":[[2023,10,4]],"date-time":"2023-10-04T15:53:38Z","timestamp":1696434818000},"page":"1-19","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":9,"title":["Enabling Binary Neural Network Training on the Edge"],"prefix":"10.1145","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-3603-6852","authenticated-orcid":false,"given":"Erwei","family":"Wang","sequence":"first","affiliation":[{"name":"Imperial College London, United Kingdom"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4910-3188","authenticated-orcid":false,"given":"James J.","family":"Davis","sequence":"additional","affiliation":[{"name":"Imperial College London, United Kingdom"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5544-830X","authenticated-orcid":false,"given":"Daniele","family":"Moro","sequence":"additional","affiliation":[{"name":"Google, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0000-1887-496X","authenticated-orcid":false,"given":"Piotr","family":"Zielinski","sequence":"additional","affiliation":[{"name":"Google, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-9383-5324","authenticated-orcid":false,"given":"Jia Jie","family":"Lim","sequence":"additional","affiliation":[{"name":"iSize, United Kingdom"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9637-1890","authenticated-orcid":false,"given":"Claudionor","family":"Coelho","sequence":"additional","affiliation":[{"name":"Advantest, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8135-8378","authenticated-orcid":false,"given":"Satrajit","family":"Chatterjee","sequence":"additional","affiliation":[{"name":"United States"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8236-1816","authenticated-orcid":false,"given":"Peter Y. K.","family":"Cheung","sequence":"additional","affiliation":[{"name":"Imperial College London, United Kingdom"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0201-310X","authenticated-orcid":false,"given":"George A.","family":"Constantinides","sequence":"additional","affiliation":[{"name":"Imperial College London, United Kingdom"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2023,11,9]]},"reference":[{"key":"e_1_3_2_2_2","volume-title":"International Conference on Neural Information Processing Systems","author":"Agarwal Naman","year":"2018","unstructured":"Naman Agarwal, Ananda Theertha Suresh, Felix Yu, Sanjiv Kumar, and H. Brendan McMahan. 2018. CpSGD: Communication-efficient and differentially-private distributed SGD. In International Conference on Neural Information Processing Systems."},{"key":"e_1_3_2_3_2","volume-title":"International Conference on Learning Representations","author":"Alizadeh Milad","year":"2018","unstructured":"Milad Alizadeh, Javier Fern\u00e1ndez-Marqu\u00e9s, Nicholas D. Lane, and Yarin Gal. 2018. An empirical study of binary neural networks\u2019 optimisation. In International Conference on Learning Representations."},{"key":"e_1_3_2_4_2","volume-title":"International Conference on Machine Learning","author":"Bernstein Jeremy","year":"2018","unstructured":"Jeremy Bernstein, Yu-Xiang Wang, Kamyar Azizzadenesheli, and Animashree Anandkumar. 2018. SignSGD: Compressed optimisation for non-convex problems. In International Conference on Machine Learning."},{"key":"e_1_3_2_5_2","article-title":"Back to simplicity: How to train accurate BNNs from scratch?","author":"Bethge Joseph","year":"2019","unstructured":"Joseph Bethge, Haojin Yang, Marvin Bornstein, and Christoph Meinel. 2019. Back to simplicity: How to train accurate BNNs from scratch? arXiv preprint arXiv:1906.08637 (2019).","journal-title":"arXiv preprint arXiv:1906.08637"},{"key":"e_1_3_2_6_2","doi-asserted-by":"crossref","unstructured":"L. Susan Blackford Antoine Petitet Roldan Pozo Karin Remington R. Clint Whaley James Demmel Jack Dongarra Iain Duff Sven Hammarling Greg Henry Michael Heroux Linda Kaufman and Andrew Lumsdaine. 2002. An updated set of basic linear algebra subprograms (BLAS). ACM Trans. Math. Software 28 2 (2002) 135\u2013151.","DOI":"10.1145\/567806.567807"},{"key":"e_1_3_2_7_2","volume-title":"Conference on Machine Learning and Systems","author":"Bonawitz Keith","year":"2019","unstructured":"Keith Bonawitz, Hubert Eichner, Wolfgang Grieskamp, Dzmitry Huba, Alex Ingerman, Vladimir Ivanov, Chloe Kiddon, Jakub Kone\u010dn\u1ef3, Stefano Mazzocchi, H. Brendan McMahan, Timon van Overveldt, David Petrou, Daniel Ramage, and Jason Roselander. 2019. Towards federated learning at scale: System design. In Conference on Machine Learning and Systems."},{"key":"e_1_3_2_8_2","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition","author":"Cai Han","year":"2020","unstructured":"Han Cai, Chuang Gan, Ligeng Zhu, and Song Han. 2020. Tiny transfer learning: Towards memory-efficient on-device learning. In IEEE Conference on Computer Vision and Pattern Recognition."},{"key":"e_1_3_2_9_2","volume-title":"Advances in Neural Information Processing Systems","author":"Chakrabarti Ayan","year":"2019","unstructured":"Ayan Chakrabarti and Benjamin Moseley. 2019. Backprop with approximate activations for memory-efficient network training. In Advances in Neural Information Processing Systems."},{"key":"e_1_3_2_10_2","article-title":"On the generalization mystery in deep learning","author":"Chatterjee Satrajit","year":"2022","unstructured":"Satrajit Chatterjee and Piotr Zielinski. 2022. On the generalization mystery in deep learning. arXiv preprint arXiv:2203.10036 (2022).","journal-title":"arXiv preprint arXiv:2203.10036"},{"key":"e_1_3_2_11_2","article-title":"Training deep nets with sublinear memory cost","author":"Chen Tianqi","year":"2016","unstructured":"Tianqi Chen, Bing Xu, Chiyuan Zhang, and Carlos Guestrin. 2016. Training deep nets with sublinear memory cost. arXiv preprint arXiv:1604.06174 (2016).","journal-title":"arXiv preprint arXiv:1604.06174"},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW53098.2021.00520"},{"issue":"2166","key":"e_1_3_2_13_2","article-title":"Rethinking arithmetic for deep neural networks","volume":"378","author":"Constantinides George A.","year":"2019","unstructured":"George A. Constantinides. 2019. Rethinking arithmetic for deep neural networks. Philosophical Transactions of the Royal Society A 378, 2166 (2019).","journal-title":"Philosophical Transactions of the Royal Society A"},{"key":"e_1_3_2_14_2","article-title":"BinaryNet: Training deep neural networks with weights and activations Constrained to +1 or -1","author":"Courbariaux Matthieu","year":"2016","unstructured":"Matthieu Courbariaux and Yoshua Bengio. 2016. BinaryNet: Training deep neural networks with weights and activations Constrained to +1 or -1. arXiv preprint arXiv:1602.02830 (2016).","journal-title":"arXiv preprint arXiv:1602.02830"},{"key":"e_1_3_2_15_2","volume-title":"Conference on Neural Information Processing Systems","author":"Courbariaux Matthieu","year":"2015","unstructured":"Matthieu Courbariaux, Yoshua Bengio, and Jean-Pierre David. 2015. BinaryConnect: Training deep neural networks with binary weights during propagations. In Conference on Neural Information Processing Systems."},{"key":"e_1_3_2_16_2","unstructured":"Sajad Darabi Mouloud Belbahri Matthieu Courbariaux and Vahid Partovi Nia. 2018. BNN+: Improved Binary Network Training. (2018). https:\/\/openreview.net\/pdf?id=SJfHg2A5tQ"},{"key":"e_1_3_2_17_2","volume-title":"IEEE International Symposium on Field-Programmable Custom Computing Machines","author":"Ghasemzadeh Mohammad","year":"2018","unstructured":"Mohammad Ghasemzadeh, Mohammad Samragh, and Farinaz Koushanfar. 2018. ReBNet: Residual binarized neural network. In IEEE International Symposium on Field-Programmable Custom Computing Machines."},{"key":"e_1_3_2_18_2","volume-title":"Nvidia GPU Technology Conference","author":"Ginsburg Boris","year":"2017","unstructured":"Boris Ginsburg, Sergei Nikolaev, and Paulius Micikevicius. 2017. Training of deep networks with half-precision float. In Nvidia GPU Technology Conference."},{"key":"e_1_3_2_19_2","volume-title":"International Conference on Artificial Intelligence and Statistics","author":"Glorot Xavier","year":"2010","unstructured":"Xavier Glorot and Yoshua Bengio. 2010. Understanding the difficulty of training deep feedforward neural networks. In International Conference on Artificial Intelligence and Statistics."},{"key":"e_1_3_2_20_2","article-title":"Low-precision batch-normalized activations","author":"Graham Benjamin","year":"2017","unstructured":"Benjamin Graham. 2017. Low-precision batch-normalized activations. arXiv preprint arXiv:1702.08231 (2017).","journal-title":"arXiv preprint arXiv:1702.08231"},{"key":"e_1_3_2_21_2","volume-title":"Advances in Neural Information Processing Systems","author":"Gruslys Audrunas","year":"2016","unstructured":"Audrunas Gruslys, R\u00e9mi Munos, Ivo Danihelka, Marc Lanctot, and Alex Graves. 2016. Memory-efficient backpropagation through time. In Advances in Neural Information Processing Systems."},{"key":"e_1_3_2_22_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58580-8_14"},{"key":"e_1_3_2_23_2","volume-title":"Advances in Neural Information Processing Systems","author":"Helwegen Koen","year":"2019","unstructured":"Koen Helwegen, James Widdicombe, Lukas Geiger, Zechun Liu, Kwang-Ting Cheng, and Roeland Nusselder. 2019. Latent weights do not exist: Rethinking binarized neural network optimization. In Advances in Neural Information Processing Systems."},{"key":"e_1_3_2_24_2","volume-title":"Advances in Neural Information Processing Systems","author":"Hoffer Elad","year":"2018","unstructured":"Elad Hoffer, Ron Banner, Itay Golan, and Daniel Soudry. 2018. Norm matters: Efficient and accurate normalization schemes in deep networks. In Advances in Neural Information Processing Systems."},{"key":"e_1_3_2_25_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v35i2.16263"},{"key":"e_1_3_2_26_2","unstructured":"Keras. memory leak in tf.keras.Model.predict. (n.d.). https:\/\/github.com\/tensorflow\/tensorflow\/issues\/44711"},{"key":"e_1_3_2_27_2","volume-title":"International Conference on Learning Representations","author":"Kingma Diederik P.","year":"2015","unstructured":"Diederik P. Kingma and Jimmy Ba. 2015. Adam: A method for stochastic optimization. In International Conference on Learning Representations."},{"key":"e_1_3_2_28_2","volume-title":"Conference on Neural Information Processing Systems","author":"Lin Xiaofan","year":"2017","unstructured":"Xiaofan Lin, Cong Zhao, and Wei Pan. 2017. Towards accurate binary convolutional neural network. In Conference on Neural Information Processing Systems."},{"key":"e_1_3_2_29_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58568-6_9"},{"key":"e_1_3_2_30_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01267-0_44"},{"key":"e_1_3_2_31_2","volume-title":"International Conference on Learning Representations","author":"Martinez Brais","year":"2020","unstructured":"Brais Martinez, Jing Yang, Adrian Bulat, and Georgios Tzimiropoulos. 2020. Training binary neural networks with real-to-binary convolutions. In International Conference on Learning Representations."},{"key":"e_1_3_2_32_2","volume-title":"International Conference on Artificial Intelligence and Statistics","author":"McMahan Brendan","year":"2017","unstructured":"Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, and Blaise Ag\u00fcera y Arcas. 2017. Communication-efficient learning of deep networks from decentralized data. In International Conference on Artificial Intelligence and Statistics."},{"key":"e_1_3_2_33_2","volume-title":"International Conference on Learning Representations","author":"Micikevicius Paulius","year":"2018","unstructured":"Paulius Micikevicius, Sharan Narang, Jonah Alben, Gregory Diamos, Erich Elsen, David Garcia, Boris Ginsburg, Michael Houston, Oleksii Kuchaiev, Ganesh Venkatesh, and Hao Wu. 2018. Mixed precision training. In International Conference on Learning Representations."},{"key":"e_1_3_2_34_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2020.107281"},{"key":"e_1_3_2_35_2","unstructured":"reichelt. RPI USB METER2. (n.d.). https:\/\/www.reichelt.com\/de\/en\/raspberry-pi-amp-voltmeter-2-way-usb-rpi-usb-meter2-p223623.html?r=1"},{"key":"e_1_3_2_36_2","article-title":"How does batch normalization help binary training","author":"Sari Eyy\u00fcb","year":"2019","unstructured":"Eyy\u00fcb Sari, Mouloud Belbahri, and Vahid P. Nia. 2019. How does batch normalization help binary training. arXiv preprint arXiv:1909.09139 (2019).","journal-title":"arXiv preprint arXiv:1909.09139"},{"key":"e_1_3_2_37_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCVW.2019.00256"},{"key":"e_1_3_2_38_2","article-title":"Low-memory neural network training: A technical report","author":"Sohoni Nimit S.","year":"2019","unstructured":"Nimit S. Sohoni, Christopher R. Aberger, Megan Leszczynski, Jian Zhang, and Christopher R\u00e9. 2019. Low-memory neural network training: A technical report. arXiv preprint arXiv:1904.10631 (2019).","journal-title":"arXiv preprint arXiv:1904.10631"},{"key":"e_1_3_2_39_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICFPT56656.2022.9974324"},{"key":"e_1_3_2_40_2","doi-asserted-by":"publisher","DOI":"10.1109\/FPL50879.2020.00055"},{"key":"e_1_3_2_41_2","volume-title":"ACM\/SIGDA International Symposium on Field-Programmable Gate Arrays","author":"Umuroglu Yaman","year":"2017","unstructured":"Yaman Umuroglu, Nicholas J. Fraser, Giulio Gambardella, Michaela Blott, Philip H. W. Leong, Magnus Jahre, and Kees Vissers. 2017. FINN: A framework for fast, scalable binarized neural network inference. In ACM\/SIGDA International Symposium on Field-Programmable Gate Arrays."},{"key":"e_1_3_2_42_2","volume-title":"IEEE International Symposium on Field-Programmable Custom Computing Machines","author":"Wang Erwei","year":"2019","unstructured":"Erwei Wang, James J. Davis, Peter Y. K. Cheung, and George A. Constantinides. 2019. LUTNet: Rethinking inference in FPGA soft logic. In IEEE International Symposium on Field-Programmable Custom Computing Machines."},{"key":"e_1_3_2_43_2","doi-asserted-by":"publisher","DOI":"10.1109\/TC.2020.2978817"},{"key":"e_1_3_2_44_2","doi-asserted-by":"publisher","DOI":"10.1145\/3309551"},{"key":"e_1_3_2_45_2","volume-title":"Advances in Neural Information Processing Systems","author":"Wang Naigang","year":"2018","unstructured":"Naigang Wang, Jungwook Choi, Daniel Brand, Chia-Yu Chen, and Kailash Gopalakrishnan. 2018. Training deep neural networks with 8-bit floating point numbers. In Advances in Neural Information Processing Systems."},{"key":"e_1_3_2_46_2","volume-title":"Advances in Neural Information Processing Systems","author":"Wen Wei","year":"2017","unstructured":"Wei Wen, Cong Xu, Feng Yan, Chunpeng Wu, Yandan Wang, Yiran Chen, and Hai Li. 2017. TernGrad: Ternary gradients to reduce communication in distributed deep learning. In Advances in Neural Information Processing Systems."},{"key":"e_1_3_2_47_2","volume-title":"Advances in Neural Information Processing Systems","author":"Wilson Ashia C.","year":"2017","unstructured":"Ashia C. Wilson, Rebecca Roelofs, Mitchell Stern, Nati Srebro, and Benjamin Recht. 2017. The marginal value of adaptive gradient methods in machine learning. In Advances in Neural Information Processing Systems."},{"key":"e_1_3_2_48_2","volume-title":"International Conference on Learning Representations","author":"Wu Shuang","year":"2018","unstructured":"Shuang Wu, Guoqi Li, Feng Chen, and Luping Shi. 2018. Training and inference with integers in deep neural networks. In International Conference on Learning Representations."},{"key":"e_1_3_2_49_2","doi-asserted-by":"publisher","DOI":"10.1109\/TNNLS.2018.2876179"},{"key":"e_1_3_2_50_2","doi-asserted-by":"publisher","DOI":"10.1145\/3489517.3530496"},{"key":"e_1_3_2_51_2","doi-asserted-by":"publisher","DOI":"10.1145\/3446776"},{"key":"e_1_3_2_52_2","article-title":"DoReFa-Net: Training low bitwidth convolutional neural networks with low bitwidth gradients","author":"Zhou Shuchang","year":"2016","unstructured":"Shuchang Zhou, Zekun Ni, Xinyu Zhou, He Wen, Yuxin Wu, and Yuheng Zou. 2016. DoReFa-Net: Training low bitwidth convolutional neural networks with low bitwidth gradients. arXiv preprint arXiv:1606.06160 (2016).","journal-title":"arXiv preprint arXiv:1606.06160"}],"container-title":["ACM Transactions on Embedded Computing Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3626100","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3626100","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T22:53:59Z","timestamp":1750287239000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3626100"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,11,9]]},"references-count":51,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2023,11,30]]}},"alternative-id":["10.1145\/3626100"],"URL":"https:\/\/doi.org\/10.1145\/3626100","relation":{},"ISSN":["1539-9087","1558-3465"],"issn-type":[{"value":"1539-9087","type":"print"},{"value":"1558-3465","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,11,9]]},"assertion":[{"value":"2023-03-31","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-09-11","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-11-09","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}