{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,9]],"date-time":"2026-05-09T04:53:13Z","timestamp":1778302393048,"version":"3.51.4"},"reference-count":66,"publisher":"Association for Computing Machinery (ACM)","issue":"6","license":[{"start":{"date-parts":[[2022,10,18]],"date-time":"2022-10-18T00:00:00Z","timestamp":1666051200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Embed. Comput. Syst."],"published-print":{"date-parts":[[2022,11,30]]},"abstract":"<jats:p>Recently, there has been an explosive growth of mobile and embedded applications using convolutional neural networks (CNNs). To alleviate their excessive computational demands, developers have traditionally resorted to cloud offloading, inducing high infrastructure costs and a strong dependence on networking conditions. On the other end, the emergence of powerful SoCs is gradually enabling on-device execution. Nonetheless, low- and mid-tier platforms still struggle to run state-of-the-art CNNs sufficiently. In this article, we present DynO, a distributed inference framework that combines the best of both worlds to address several challenges, such as device heterogeneity, varying bandwidth, and multi-objective requirements. Key components that enable this are its novel CNN-specific data packing method, which exploits the variability of precision needs in different parts of the CNN when onloading computation, and its novel scheduler, which jointly tunes the partition point and transferred data precision at runtime to adapt inference to its execution environment. Quantitative evaluation shows that DynO outperforms the current state of the art, improving throughput by over an order of magnitude over device-only execution and up to 7.9\u00d7 over competing CNN offloading systems, with up to 60\u00d7 less data transferred.<\/jats:p>","DOI":"10.1145\/3510831","type":"journal-article","created":{"date-parts":[[2022,1,26]],"date-time":"2022-01-26T18:01:53Z","timestamp":1643220113000},"page":"1-24","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":29,"title":["DynO: Dynamic Onloading of Deep Neural Networks from Cloud to Device"],"prefix":"10.1145","volume":"21","author":[{"given":"Mario","family":"Almeida","sequence":"first","affiliation":[{"name":"Samsung AI Center, Cambridge, UK"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Stefanos","family":"Laskaridis","sequence":"additional","affiliation":[{"name":"Samsung AI Center, Cambridge, UK"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5181-6251","authenticated-orcid":false,"given":"Stylianos I.","family":"Venieris","sequence":"additional","affiliation":[{"name":"Samsung AI Center, Cambridge, UK"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ilias","family":"Leontiadis","sequence":"additional","affiliation":[{"name":"Samsung AI Center, Cambridge, UK"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Nicholas D.","family":"Lane","sequence":"additional","affiliation":[{"name":"Samsung AI Center, Cambridge, &amp; University of Cambridge, Cambridge, UK"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2022,10,18]]},"reference":[{"key":"e_1_3_2_2_2","first-page":"1","volume-title":"3rd International Workshop on Deep Learning for Mobile Systems and Applications (EMDL\u201919)","author":"Almeida Mario","year":"2019","unstructured":"Mario Almeida, Stefanos Laskaridis, Ilias Leontiadis, Stylianos I. Venieris, and Nicholas D. Lane. 2019. EmBench: Quantifying performance variations of deep neural networks across modern commodity devices. In 3rd International Workshop on Deep Learning for Mobile Systems and Applications (EMDL\u201919). ACM, 1\u20136."},{"key":"e_1_3_2_3_2","doi-asserted-by":"publisher","DOI":"10.1145\/3487552.3487863"},{"key":"e_1_3_2_4_2","doi-asserted-by":"publisher","DOI":"10.1145\/3363347.3363363"},{"key":"e_1_3_2_5_2","first-page":"129","volume-title":"Proceedings of Machine Learning and Systems (MLSys\u201920)","volume":"2","author":"Blalock Davis","year":"2020","unstructured":"Davis Blalock, Jose Javier Gonzalez Ortiz, Jonathan Frankle, and John Guttag. 2020. What is the state of neural network pruning? In Proceedings of Machine Learning and Systems (MLSys\u201920). Vol. 2. 129\u2013146."},{"key":"e_1_3_2_6_2","doi-asserted-by":"publisher","DOI":"10.1109\/HOTCHIPS.2019.8875651"},{"key":"e_1_3_2_7_2","doi-asserted-by":"publisher","DOI":"10.1145\/3301418.3313946"},{"key":"e_1_3_2_8_2","doi-asserted-by":"publisher","DOI":"10.1145\/1966445.1966473"},{"key":"e_1_3_2_9_2","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2018.022071131"},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","DOI":"10.1145\/1814433.1814441"},{"key":"e_1_3_2_11_2","first-page":"1597","volume-title":"2017 IEEE 60th International Midwest Symposium on Circuits and Systems (MWSCAS\u201917)","author":"Dey Rahul","year":"2017","unstructured":"Rahul Dey and Fathi M. Salem. 2017. Gate-Variants of gated recurrent unit (GRU) neural networks. In 2017 IEEE 60th International Midwest Symposium on Circuits and Systems (MWSCAS\u201917). IEEE, 1597\u20131600."},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.1145\/3288599.3288634"},{"key":"e_1_3_2_13_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00725"},{"key":"e_1_3_2_14_2","volume-title":"USENIX Conference on Operating Systems Design and Implementation (OSDI\u201912)","author":"Gordon Mark S.","year":"2012","unstructured":"Mark S. Gordon, D. Anoushe Jamshidi, Scott Mahlke, Z. Morley Mao, and Xu Chen. 2012. COMET: Code Offload by Migrating Execution Transparently. In USENIX Conference on Operating Systems Design and Implementation (OSDI\u201912)."},{"key":"e_1_3_2_15_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCAD.2017.2705069"},{"key":"e_1_3_2_16_2","doi-asserted-by":"publisher","DOI":"10.1109\/TNNLS.2018.2808319"},{"key":"e_1_3_2_17_2","volume-title":"International Conference on Learning Representations (ICLR\u201916)","author":"Han Song","year":"2016","unstructured":"Song Han, Huizi Mao, and William J. Dally. 2016. Deep compression: Compressing deep neural networks with pruning, trained quantization and Huffman coding. In International Conference on Learning Representations (ICLR\u201916)."},{"key":"e_1_3_2_18_2","doi-asserted-by":"publisher","DOI":"10.1145\/2906388.2906396"},{"key":"e_1_3_2_19_2","doi-asserted-by":"crossref","first-page":"620","DOI":"10.1109\/HPCA.2018.00059","volume-title":"2018 IEEE International Symposium on High Performance Computer Architecture (HPCA\u201918)","author":"Hazelwood K.","year":"2018","unstructured":"K. Hazelwood, S. Bird, D. Brooks, S. Chintala, U. Diril, D. Dzhulgakov, M. Fawzy, B. Jia, Y. Jia, A. Kalro, J. Law, K. Lee, J. Lu, P. Noordhuis, M. Smelyanskiy, L. Xiong, and X. Wang. 2018. Applied machine learning at Facebook: A datacenter infrastructure perspective. In 2018 IEEE International Symposium on High Performance Computer Architecture (HPCA\u201918). 620\u2013629."},{"key":"e_1_3_2_20_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.123"},{"key":"e_1_3_2_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_2_22_2","doi-asserted-by":"publisher","DOI":"10.1162\/neco.1997.9.8.1735"},{"key":"e_1_3_2_23_2","doi-asserted-by":"crossref","unstructured":"Chuang Hu Wei Bao Dan Wang and Fengming Liu. 2019. Dynamic adaptive DNN surgery for inference acceleration on the edge. In Proceedings of the IEEE Conference on Computer Communications (INFOCOM\u201919) . 1423\u20131431.","DOI":"10.1109\/INFOCOM.2019.8737614"},{"key":"e_1_3_2_24_2","doi-asserted-by":"crossref","unstructured":"Jin Huang Colin Samplawski Deepak Ganesan Benjamin Marlin and Heesung Kwon. 2020. CLIO: Enabling automatic compilation of deep learning pipelines across IoT and Cloud. In Proceedings of the 26th Annual International Conference on Mobile Computing and Networking (MobiCom\u201920) .","DOI":"10.1145\/3372224.3419215"},{"key":"e_1_3_2_25_2","doi-asserted-by":"crossref","unstructured":"Andrey Ignatov Radu Timofte Andrei Kulik Seungsoo Yang Ke Wang Felix Baum Max Wu Lirong Xu and Luc Van Gool. 2019. AI benchmark: All about deep learning on smartphones in 2019. In 2019 IEEE\/CVF International Conference on Computer Vision Workshop (ICCVW) .","DOI":"10.1109\/ICCVW.2019.00447"},{"key":"e_1_3_2_26_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00286"},{"key":"e_1_3_2_27_2","doi-asserted-by":"publisher","DOI":"10.1145\/3366636"},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCAD.2018.2857338"},{"key":"e_1_3_2_29_2","doi-asserted-by":"publisher","DOI":"10.1145\/3400302.3415675"},{"key":"e_1_3_2_30_2","doi-asserted-by":"publisher","DOI":"10.1145\/3267809.3267828"},{"key":"e_1_3_2_31_2","volume-title":"International Symposium on Computer Architecture (ISCA\u201917)","author":"Jouppi Norman P.","year":"2017","unstructured":"Norman P. Jouppi et\u00a0al. 2017. In-datacenter performance analysis of a tensor processing unit. In International Symposium on Computer Architecture (ISCA\u201917)."},{"key":"e_1_3_2_32_2","first-page":"615","article-title":"Neurosurgeon: Collaborative intelligence between the cloud and mobile edge","author":"Kang Yiping","year":"2017","unstructured":"Yiping Kang, Johann Hauswald, Cao Gao, Austin Rovinski, Trevor Mudge, Jason Mars, and Lingjia Tang. 2017. Neurosurgeon: Collaborative intelligence between the cloud and mobile edge. International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS\u201917), 615\u2013629.","journal-title":"International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS\u201917)"},{"key":"e_1_3_2_33_2","doi-asserted-by":"publisher","DOI":"10.1109\/IROS.2018.8594204"},{"key":"e_1_3_2_34_2","first-page":"155","volume-title":"2018 28th International Conference on Field Programmable Logic and Applications (FPL\u201918)","author":"Kouris A.","year":"2018","unstructured":"A. Kouris, S. I. Venieris, and C. Bouganis. 2018. CascadeCNN: Pushing the performance limits of quantisation in convolutional neural networks. In 2018 28th International Conference on Field Programmable Logic and Applications (FPL\u201918). 155\u20131557."},{"key":"e_1_3_2_35_2","volume-title":"arXiv","author":"Kouris Alexandros","year":"2021","unstructured":"Alexandros Kouris, Stylianos I. Venieris, Stefanos Laskaridis, and Nicholas D. Lane. 2021. Multi-exit semantic segmentation networks. In arXiv."},{"key":"e_1_3_2_36_2","doi-asserted-by":"publisher","DOI":"10.1109\/MCE.2018.2828440"},{"key":"e_1_3_2_37_2","doi-asserted-by":"publisher","DOI":"10.1145\/3469116.3470012"},{"key":"e_1_3_2_38_2","doi-asserted-by":"publisher","DOI":"10.1145\/3372224.3419194"},{"key":"e_1_3_2_39_2","doi-asserted-by":"publisher","DOI":"10.1145\/3400302.3415698"},{"key":"e_1_3_2_40_2","doi-asserted-by":"publisher","DOI":"10.1145\/3300061.3345447"},{"key":"e_1_3_2_41_2","doi-asserted-by":"publisher","DOI":"10.1109\/TWC.2019.2946140"},{"key":"e_1_3_2_42_2","first-page":"671","volume-title":"International Conference on Parallel and Distributed Systems (ICPADS\u201919)","author":"Li Hongshan","year":"2019","unstructured":"Hongshan Li, Chenghao Hu, Jingyan Jiang, Zhi Wang, Yonggang Wen, and Wenwu Zhu. 2019. JALAD: Joint accuracy-and latency-aware deep structure decoupling for edge-cloud execution. In International Conference on Parallel and Distributed Systems (ICPADS\u201919). 671\u2013678."},{"key":"e_1_3_2_43_2","volume-title":"Design, Automation and Test in Europe (DATE\u201917)","author":"Mao Jiachen","year":"2017","unstructured":"Jiachen Mao, Xiang Chen, Kent W. Nixon, Christopher Krieger, and Yiran Chen. 2017. MoDNN: Local distributed mobile computing system for Deep Neural Network. In Design, Automation and Test in Europe (DATE\u201917)."},{"key":"e_1_3_2_44_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCAD.2017.8203852"},{"key":"e_1_3_2_45_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.ascom.2015.07.002"},{"key":"e_1_3_2_46_2","doi-asserted-by":"publisher","DOI":"10.1145\/3386901.3388946"},{"key":"e_1_3_2_47_2","first-page":"807","volume-title":"International Conference on Machine Learning (ICML\u201910)","author":"Nair Vinod","year":"2010","unstructured":"Vinod Nair and Geoffrey E. Hinton. 2010. Rectified linear units improve restricted Boltzmann machines. In International Conference on Machine Learning (ICML\u201910). 807\u2013814."},{"key":"e_1_3_2_48_2","doi-asserted-by":"crossref","first-page":"86","DOI":"10.1109\/IISWC.2018.8573509","volume-title":"2018 IEEE International Symposium on Workload Characterization (IISWC\u201918)","author":"Nikoli\u0107 M.","year":"2018","unstructured":"M. Nikoli\u0107, M. Mahmoud, and A. Moshovos. 2018. Characterizing sources of ineffectual computations in deep learning networks. In 2018 IEEE International Symposium on Workload Characterization (IISWC\u201918). 86\u201387."},{"key":"e_1_3_2_49_2","unstructured":"Ofcom. 2014. 3G and 4G Network Speeds. https:\/\/www.ofcom.org.uk\/about-ofcom\/latest\/media\/media-releases\/2014\/3g-4g-bb-speeds."},{"key":"e_1_3_2_50_2","doi-asserted-by":"publisher","DOI":"10.5555\/2971808.2971918"},{"key":"e_1_3_2_51_2","unstructured":"PyTorch. 2021. Dynamic Quantization. Retrieved 2022\/10\/14 12:38:53 from https:\/\/pytorch.org\/tutorials\/recipes\/recipes\/dynamic_quantization.html."},{"key":"e_1_3_2_52_2","volume-title":"International Conference on Learning Representations (ICLR) Workshops","author":"Ramachandran Prajit","year":"2018","unstructured":"Prajit Ramachandran, Barret Zoph, and Quoc V. Le. 2018. Searching for activation functions. In International Conference on Learning Representations (ICLR) Workshops."},{"key":"e_1_3_2_53_2","first-page":"78","volume-title":"International Symposium on High Performance Computer Architecture (HPCA\u201918)","author":"Rhu M.","year":"2018","unstructured":"M. Rhu, M. O\u2019Connor, N. Chatterjee, J. Pool, Y. Kwon, and S. W. Keckler. 2018. Compressing DMA engine: Leveraging activation sparsity for training deep neural networks. In International Symposium on High Performance Computer Architecture (HPCA\u201918). 78\u201391."},{"key":"e_1_3_2_54_2","doi-asserted-by":"publisher","DOI":"10.1145\/1866739.1866751"},{"key":"e_1_3_2_55_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00474"},{"key":"e_1_3_2_56_2","volume-title":"International Conference on Learning Representations (ICLR\u201915)","author":"Simonyan K.","year":"2015","unstructured":"K. Simonyan and A. Zisserman. 2015. Very deep convolutional networks for large-scale image recognition. In International Conference on Learning Representations (ICLR\u201915)."},{"key":"e_1_3_2_57_2","doi-asserted-by":"publisher","DOI":"10.1109\/JPROC.2017.2761740"},{"key":"e_1_3_2_58_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v31i1.11231"},{"key":"e_1_3_2_59_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-1-84996-241-4_20"},{"key":"e_1_3_2_60_2","volume-title":"IEEE International Symposium on High Performance Computer Architecture (HPCA\u201919)","author":"Wu C.","year":"2019","unstructured":"C. Wu, D. Brooks, K. Chen, D. Chen, S. Choudhury, M. Dukhan, K. Hazelwood, E. Isaac, Y. Jia, B. Jia, T. Leyvand, H. Lu, Y. Lu, L. Qiao, B. Reagen, J. Spisak, F. Sun, A. Tulloch, P. Vajda, X. Wang, Y. Wang, B. Wasti, Y. Wu, R. Xian, S. Yoo, and P. Zhang. 2019. Machine learning at Facebook: Understanding inference at the edge. In IEEE International Symposium on High Performance Computer Architecture (HPCA\u201919)."},{"key":"e_1_3_2_61_2","doi-asserted-by":"crossref","unstructured":"Xiaowei Xu Yukun Ding Sharon Xiaobo Hu Michael Niemier Jason Cong Yu Hu and Yiyu Shi. 2018. Scaling for edge inference of deep neural networks. Nature Electronics 1 (2018) 216\u2013222.","DOI":"10.1038\/s41928-018-0059-3"},{"key":"e_1_3_2_62_2","doi-asserted-by":"publisher","DOI":"10.1145\/3038912.3052577"},{"key":"e_1_3_2_63_2","doi-asserted-by":"publisher","DOI":"10.1109\/CloudNet.2012.6483660"},{"key":"e_1_3_2_64_2","first-page":"951","volume-title":"2018 USENIX Annual Technical Conference","author":"Zhang Minjia","year":"2018","unstructured":"Minjia Zhang, Samyam Rajbhandari, Wenhan Wang, and Yuxiong He. 2018. DeepCPU: Serving RNN-based deep learning models 10x faster. In 2018 USENIX Annual Technical Conference. 951\u2013965."},{"key":"e_1_3_2_65_2","doi-asserted-by":"publisher","DOI":"10.1145\/3447993.3448628"},{"key":"e_1_3_2_66_2","doi-asserted-by":"crossref","unstructured":"Zhuoran Zhao Kamyar Mirzazad Barijough and Andreas Gerstlauer. 2018. DeepThings: Distributed adaptive deep learning inference on resource-constrained IoT edge clusters. IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems (TCAD) 37 11 (2018) 2348\u20132359.","DOI":"10.1109\/TCAD.2018.2858384"},{"key":"e_1_3_2_67_2","volume-title":"International Conference on Learning Representations (ICLR\u201917)","author":"Zhou Aojun","year":"2017","unstructured":"Aojun Zhou et\u00a0al. 2017. Incremental network quantization: Towards lossless CNNs with low-precision weights. In International Conference on Learning Representations (ICLR\u201917)."}],"container-title":["ACM Transactions on Embedded Computing Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3510831","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3510831","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T19:02:11Z","timestamp":1750186931000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3510831"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,10,18]]},"references-count":66,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2022,11,30]]}},"alternative-id":["10.1145\/3510831"],"URL":"https:\/\/doi.org\/10.1145\/3510831","relation":{},"ISSN":["1539-9087","1558-3465"],"issn-type":[{"value":"1539-9087","type":"print"},{"value":"1558-3465","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,10,18]]},"assertion":[{"value":"2021-04-16","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2022-01-07","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2022-10-18","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}