{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,4]],"date-time":"2026-06-04T12:57:55Z","timestamp":1780577875362,"version":"3.54.1"},"reference-count":54,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2023,4,3]],"date-time":"2023-04-03T00:00:00Z","timestamp":1680480000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"National Science Foundation","award":["CNS-2038986\/2038566, and CNS-2146449"],"award-info":[{"award-number":["CNS-2038986\/2038566, and CNS-2146449"]}]},{"name":"Amazon Research Award"},{"DOI":"10.13039\/100006754","name":"Army Research Lab","doi-asserted-by":"crossref","award":["W911NF-2020-221"],"award-info":[{"award-number":["W911NF-2020-221"]}],"id":[{"id":"10.13039\/100006754","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Des. Autom. Electron. Syst."],"published-print":{"date-parts":[[2023,5,31]]},"abstract":"<jats:p>\n            Efficient and adaptive computer vision systems have been proposed to make computer vision tasks, such as image classification and object detection, optimized for embedded or mobile devices. These solutions, quite recent in their origin, focus on optimizing the model (a deep neural network) or the system by designing an adaptive system with approximation knobs. Despite several recent efforts, we show that existing solutions suffer from two major drawbacks.\n            <jats:italic>First<\/jats:italic>\n            , while mobile devices or systems-on-chips usually come with limited resources including battery power, most systems do not consider the energy consumption of the models during inference.\n            <jats:italic>Second<\/jats:italic>\n            , they do not consider the interplay between the three metrics of interest in their configurations, namely, latency, accuracy, and energy. In this work, we propose an efficient and adaptive video object detection system\u2014\n            <jats:sc>Virtuoso<\/jats:sc>\n            , which is jointly optimized for accuracy, energy efficiency, and latency. Underlying\n            <jats:sc>Virtuoso<\/jats:sc>\n            is a multi-branch execution kernel that is capable of running at different operating points in the accuracy-energy-latency axes, and a lightweight runtime scheduler to select the best fit execution branch to satisfy the user requirement. We position this work as a first step in understanding the suitability of various object detection kernels on embedded boards in the accuracy-latency-energy axes, opening the door for further development in solutions customized to embedded systems and for benchmarking such solutions.\n            <jats:sc>Virtuoso<\/jats:sc>\n            is able to achieve up to 286 FPS on the NVIDIA Jetson AGX Xavier board, which is up to 45\u00d7 faster than the baseline EfficientDet D3 and 15\u00d7 faster than the baseline EfficientDet D0. In addition, we also observe up to 97.2% energy reduction using\n            <jats:sc>Virtuoso<\/jats:sc>\n            compared to the baseline YOLO (v3)\u2014a widely used object detector designed for mobiles. To fairly compare with\n            <jats:sc>Virtuoso<\/jats:sc>\n            , we benchmark 15 state-of-the-art or widely used protocols, including Faster R-CNN (FRCNN) [NeurIPS\u201915], YOLO v3 [CVPR\u201916], SSD [ECCV\u201916], EfficientDet [CVPR\u201920], SELSA [ICCV\u201919], MEGA [CVPR\u201920], REPP [IROS\u201920], FastAdapt [EMDL\u201921], and our in-house adaptive variants of FRCNN+, YOLO+, SSD+, and EfficientDet+ (our variants have enhanced efficiency for mobiles). With this comprehensive benchmark,\n            <jats:sc>Virtuoso<\/jats:sc>\n            has shown superiority to all the above protocols, leading the accuracy frontier at every efficiency level on NVIDIA Jetson mobile GPUs. Specifically,\n            <jats:sc>Virtuoso<\/jats:sc>\n            has achieved an accuracy of 63.9%, which is more than 10% higher than some of the popular object detection models, FRCNN at 51.1% and YOLO at 49.5%.\n          <\/jats:p>","DOI":"10.1145\/3564289","type":"journal-article","created":{"date-parts":[[2022,10,3]],"date-time":"2022-10-03T12:26:14Z","timestamp":1664799974000},"page":"1-32","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":6,"title":["<scp>Virtuoso<\/scp>\n            : Energy- and Latency-aware Streamlining of Streaming Videos on Systems-on-Chips"],"prefix":"10.1145","volume":"28","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-1011-6002","authenticated-orcid":false,"given":"Jayoung","family":"Lee","sequence":"first","affiliation":[{"name":"Purdue University, West Lafayette, Indiana"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2797-6973","authenticated-orcid":false,"given":"Pengcheng","family":"Wang","sequence":"additional","affiliation":[{"name":"Purdue University, West Lafayette, Indiana"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2913-9420","authenticated-orcid":false,"given":"Ran","family":"Xu","sequence":"additional","affiliation":[{"name":"Purdue University, West Lafayette, Indiana"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8748-0603","authenticated-orcid":false,"given":"Sarthak","family":"Jain","sequence":"additional","affiliation":[{"name":"Los Gatos High School, CA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9795-6567","authenticated-orcid":false,"given":"Venkat","family":"Dasari","sequence":"additional","affiliation":[{"name":"Army Research Lab, Adelphi, MD"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2152-8854","authenticated-orcid":false,"given":"Noah","family":"Weston","sequence":"additional","affiliation":[{"name":"Army Research Lab, Adelphi, MD"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4173-9453","authenticated-orcid":false,"given":"Yin","family":"Li","sequence":"additional","affiliation":[{"name":"University of Wisconsin at Madison, WI"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4239-5632","authenticated-orcid":false,"given":"Saurabh","family":"Bagchi","sequence":"additional","affiliation":[{"name":"Purdue University, West Lafayette, Indiana"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3651-6362","authenticated-orcid":false,"given":"Somali","family":"Chaterji","sequence":"additional","affiliation":[{"name":"Purdue University, West Lafayette, Indiana"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2023,4,3]]},"reference":[{"key":"e_1_3_2_2_2","doi-asserted-by":"crossref","first-page":"96","DOI":"10.1145\/3356250.3360044","volume-title":"Proceedings of the 17th Conference on Embedded Networked Sensor Systems","author":"Apicharttrisorn Kittipat","year":"2019","unstructured":"Kittipat Apicharttrisorn, Xukan Ran, Jiasi Chen, Srikanth V. Krishnamurthy, and Amit K. Roy-Chowdhury. 2019. Frugal following: Power thrifty object detection and tracking for mobile augmented reality. In Proceedings of the 17th Conference on Embedded Networked Sensor Systems. 96\u2013109."},{"issue":"10","key":"e_1_3_2_3_2","doi-asserted-by":"crossref","first-page":"3782","DOI":"10.1109\/TITS.2019.2892405","article-title":"A survey on 3D object detection methods for autonomous driving applications","volume":"20","author":"Arnold Eduardo","year":"2019","unstructured":"Eduardo Arnold, Omar Y. Al-Jarrah, Mehrdad Dianati, Saber Fallah, David Oxtoby, and Alex Mouzakitis. 2019. A survey on 3D object detection methods for autonomous driving applications. IEEE Trans. Intell. Transport. Syst. 20, 10 (2019), 3782\u20133795.","journal-title":"IEEE Trans. Intell. Transport. Syst."},{"key":"e_1_3_2_4_2","first-page":"1218","volume-title":"Proceedings of the IEEE conference on Computer Vision and Pattern Recognition (CVPR\u201914)","author":"Bae Seung-Hwan","year":"2014","unstructured":"Seung-Hwan Bae and Kuk-Jin Yoon. 2014. Robust online multi-object tracking based on tracklet confidence and online discriminative appearance learning. In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition (CVPR\u201914). 1218\u20131225."},{"key":"e_1_3_2_5_2","first-page":"975","volume-title":"Proceedings of the IEEE International Conference on Computer Vision","author":"Buckler Mark","year":"2017","unstructured":"Mark Buckler, Suren Jayasuriya, and Adrian Sampson. 2017. Reconfiguring the imaging pipeline for computer vision. In Proceedings of the IEEE International Conference on Computer Vision. 975\u2013984."},{"key":"e_1_3_2_6_2","first-page":"13607","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR\u201920)","author":"Chen Bo","year":"2020","unstructured":"Bo Chen, Golnaz Ghiasi, Hanxiao Liu, Tsung-Yi Lin, Dmitry Kalenichenko, Hartwig Adam, and Quoc V. Le. 2020. MnasFPN: Learning latency-aware pyramid architecture for object detection on mobile devices. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR\u201920). 13607\u201313616."},{"key":"e_1_3_2_7_2","first-page":"7814","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Chen Kai","year":"2018","unstructured":"Kai Chen, Jiaqi Wang, Shuo Yang, Xingcheng Zhang, Yuanjun Xiong, Chen Change Loy, and Dahua Lin. 2018. Optimizing video object detection via a scale-time lattice. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 7814\u20137823."},{"key":"e_1_3_2_8_2","first-page":"10337","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR\u201920)","author":"Chen Yihong","year":"2020","unstructured":"Yihong Chen, Yue Cao, Han Hu, and Liwei Wang. 2020. Memory enhanced global-local aggregation for video object detection. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR\u201920). 10337\u201310346."},{"key":"e_1_3_2_9_2","first-page":"431","article-title":"Adascale: Towards real-time video object detection using adaptive scaling","volume":"1","author":"Chin Ting-Wu","year":"2019","unstructured":"Ting-Wu Chin, Ruizhou Ding, and Diana Marculescu. 2019. Adascale: Towards real-time video object detection using adaptive scaling. Proc. Mach. Learn. Syst. 1 (2019), 431\u2013441.","journal-title":"Proc. Mach. Learn. Syst."},{"key":"e_1_3_2_10_2","first-page":"91","volume-title":"Proceedings of the IEEE International Symposium on Workload Characterization (IISWC\u201911)","author":"Clemons Jason","year":"2011","unstructured":"Jason Clemons, Haishan Zhu, Silvio Savarese, and Todd Austin. 2011. MEVBench: A mobile computer vision benchmarking suite. In Proceedings of the IEEE International Symposium on Workload Characterization (IISWC\u201911). IEEE, 91\u2013102."},{"key":"e_1_3_2_11_2","first-page":"379","volume-title":"Proceedings of the Advances in Neural Information Processing Systems (NeurIPS\u201916)","author":"Dai Jifeng","year":"2016","unstructured":"Jifeng Dai, Yi Li, Kaiming He, and Jian Sun. 2016. R-FCN: Object detection via region-based fully convolutional networks. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS\u201916). 379\u2013387."},{"key":"e_1_3_2_12_2","first-page":"7023","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Deng Jiajun","year":"2019","unstructured":"Jiajun Deng, Yingwei Pan, Ting Yao, Wengang Zhou, Houqiang Li, and Tao Mei. 2019. Relation distillation networks for video object detection. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 7023\u20137032."},{"key":"e_1_3_2_13_2","doi-asserted-by":"crossref","first-page":"115","DOI":"10.1145\/3241539.3241559","volume-title":"Proceedings of the 24th Annual International Conference on Mobile Computing and Networking","author":"Fang Biyi","year":"2018","unstructured":"Biyi Fang, Xiao Zeng, and Mi Zhang. 2018. Nestdnn: Resource-aware multi-tenant on-device deep learning for continuous mobile vision. In Proceedings of the 24th Annual International Conference on Mobile Computing and Networking. 115\u2013127."},{"key":"e_1_3_2_14_2","first-page":"3038","volume-title":"Proceedings of the IEEE International Conference on Computer Vision","author":"Feichtenhofer Christoph","year":"2017","unstructured":"Christoph Feichtenhofer, Axel Pinz, and Andrew Zisserman. 2017. Detect to track and track to detect. In Proceedings of the IEEE International Conference on Computer Vision. 3038\u20133046."},{"issue":"3","key":"e_1_3_2_15_2","doi-asserted-by":"crossref","first-page":"1341","DOI":"10.1109\/TITS.2020.2972974","article-title":"Deep multi-modal object detection and semantic segmentation for autonomous driving: Datasets, methods, and challenges","volume":"22","author":"Feng Di","year":"2020","unstructured":"Di Feng, Christian Haase-Sch\u00fctz, Lars Rosenbaum, Heinz Hertlein, Claudius Glaeser, Fabian Timm, Werner Wiesbeck, and Klaus Dietmayer. 2020. Deep multi-modal object detection and semantic segmentation for autonomous driving: Datasets, methods, and challenges. IEEE Trans. Intell. Transport. Syst. 22, 3 (2020), 1341\u20131360.","journal-title":"IEEE Trans. Intell. Transport. Syst."},{"key":"e_1_3_2_16_2","doi-asserted-by":"crossref","first-page":"403","DOI":"10.1145\/2808719.2808761","volume-title":"Proceedings of the 6th ACM Conference on Bioinformatics, Computational Biology and Health Informatics","author":"Ghoshal Asish","year":"2015","unstructured":"Asish Ghoshal, Ananth Grama, Saurabh Bagchi, and Somali Chaterji. 2015. An ensemble svm model for the accurate prediction of non-canonical microrna targets. In Proceedings of the 6th ACM Conference on Bioinformatics, Computational Biology and Health Informatics. 403\u2013412."},{"key":"e_1_3_2_17_2","first-page":"1580","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Han Kai","year":"2020","unstructured":"Kai Han, Yunhe Wang, Qi Tian, Jianyuan Guo, Chunjing Xu, and Chang Xu. 2020. Ghostnet: More features from cheap operations. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 1580\u20131589."},{"issue":"3","key":"e_1_3_2_18_2","doi-asserted-by":"crossref","first-page":"583","DOI":"10.1109\/TPAMI.2014.2345390","article-title":"High-speed tracking with kernelized correlation filters","volume":"37","author":"Henriques Jo\u00e3o F.","year":"2014","unstructured":"Jo\u00e3o F. Henriques, Rui Caseiro, Pedro Martins, and Jorge Batista. 2014. High-speed tracking with kernelized correlation filters. IEEE Trans. Pattern Anal. Mach. Intell. 37, 3 (2014), 583\u2013596.","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"e_1_3_2_19_2","first-page":"18","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Jiang Zhengkai","year":"2020","unstructured":"Zhengkai Jiang, Yu Liu, Ceyuan Yang, Jihao Liu, Peng Gao, Qian Zhang, Shiming Xiang, and Chunhong Pan. 2020. Learning where to focus for efficient video object detection. In Proceedings of the European Conference on Computer Vision. Springer, 18\u201334."},{"key":"e_1_3_2_20_2","first-page":"2756","volume-title":"Proceedings of the 20th International Conference on Pattern Recognition","author":"Kalal Zdenek","year":"2010","unstructured":"Zdenek Kalal, Krystian Mikolajczyk, and Jiri Matas. 2010. Forward-backward error: Automatic detection of tracking failures. In Proceedings of the 20th International Conference on Pattern Recognition. IEEE, 2756\u20132759."},{"key":"e_1_3_2_21_2","first-page":"1","volume-title":"Proceedings of the 4th International Conference on Reliability, Infocom Technologies, and Optimization (ICRITO\u201915)","author":"Kale Kiran","year":"2015","unstructured":"Kiran Kale, Sushant Pawar, and Pravin Dhulekar. 2015. Moving object tracking using optical flow and motion vector estimation. In Proceedings of the 4th International Conference on Reliability, Infocom Technologies, and Optimization (ICRITO\u201915). IEEE, 1\u20136."},{"key":"e_1_3_2_22_2","first-page":"19","volume-title":"Proceedings of the 5th International Workshop on Embedded and Mobile Deep Learning","author":"Lee Jayoung","year":"2021","unstructured":"Jayoung Lee, Pengcheng Wang, Ran Xu, Venkat Dasari, Noah Weston, Yin Li, Saurabh Bagchi, and Somali Chaterji. 2021. Benchmarking video object detection systems on embedded devices under resource contention. In Proceedings of the 5th International Workshop on Embedded and Mobile Deep Learning. 19\u201324."},{"key":"e_1_3_2_23_2","first-page":"1019","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR\u201919)","author":"Li Buyu","year":"2019","unstructured":"Buyu Li, Wanli Ouyang, Lu Sheng, Xingyu Zeng, and Xiaogang Wang. 2019. GS3D: An efficient 3D object detection framework for autonomous driving. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR\u201919). 1019\u20131028."},{"key":"e_1_3_2_24_2","first-page":"740","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV\u201914)","author":"Lin Tsung-Yi","year":"2014","unstructured":"Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll\u00e1r, and C. Lawrence Zitnick. 2014. Microsoft COCO: Common objects in context. In Proceedings of the European Conference on Computer Vision (ECCV\u201914). Springer, 740\u2013755."},{"key":"e_1_3_2_25_2","first-page":"21","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV\u201916)","volume":"9907","author":"Liu Wei","year":"2016","unstructured":"Wei Liu, Dragomir Anguelov, Dumitru Erhan, Christian Szegedy, Scott Reed, Cheng-Yang Fu, and Alexander C. Berg. 2016. SSD: Single shot multibox detector. In Proceedings of the European Conference on Computer Vision (ECCV\u201916), Vol. 9907. 21\u201337."},{"key":"e_1_3_2_26_2","first-page":"6309","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201917)","author":"Lukezic Alan","year":"2017","unstructured":"Alan Lukezic, Tomas Vojir, Luka \u010cehovin Zajc, Jiri Matas, and Matej Kristan. 2017. Discriminative correlation filter with channel and spatial reliability. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201917). 6309\u20136318."},{"key":"e_1_3_2_27_2","first-page":"31","volume-title":"Proceedings of the International Symposium on Benchmarking, Measuring and Optimization","author":"Luo Chunjie","year":"2018","unstructured":"Chunjie Luo, Fan Zhang, Cheng Huang, Xingwang Xiong, Jianan Chen, Lei Wang, Wanling Gao, Hainan Ye, Tong Wu, Runsong Zhou, et\u00a0al. 2018. AIoT bench: Towards comprehensive benchmarking mobile and embedded device intelligence. In Proceedings of the International Symposium on Benchmarking, Measuring and Optimization. Springer, 31\u201335."},{"key":"e_1_3_2_28_2","first-page":"189","volume-title":"Proceedings of the USENIX Annual Technical Conference (USENIXATC\u201920)","author":"Mahgoub Ashraf","year":"2020","unstructured":"Ashraf Mahgoub, Alexander Michaelson Medoff, Rakesh Kumar, Subrata Mitra, Ana Klimovic, Somali Chaterji, and Saurabh Bagchi. 2020. OPTIMUSCLOUD: Heterogeneous configuration optimization for distributed databases in the cloud. In Proceedings of the USENIX Annual Technical Conference (USENIXATC\u201920). 189\u2013203."},{"key":"e_1_3_2_29_2","first-page":"300","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV\u201918)","author":"Muller Matthias","year":"2018","unstructured":"Matthias Muller, Adel Bibi, Silvio Giancola, Salman Alsubaihi, and Bernard Ghanem. 2018. Trackingnet: A large-scale dataset and benchmark for object tracking in the wild. In Proceedings of the European Conference on Computer Vision (ECCV\u201918). 300\u2013317."},{"key":"e_1_3_2_30_2","unstructured":"NVIDIA Corporation. 2020. NVIDIA Jetson AGX Xavier Board. Retrieved from https:\/\/developer.nvidia.com\/embedded\/jetson-agx-xavier-developer-kit."},{"key":"e_1_3_2_31_2","unstructured":"NVIDIA Corporation. 2020. NVIDIA Jetson Linux Developer Guide. Retrieved from https:\/\/docs.nvidia.com\/jetson\/l4t\/index.html#page\/Tegra%20Linux%20Driver%20Package%20Development%20Guide\/power_management_jetson_xavier.html#wwpID0E0VO0HA."},{"key":"e_1_3_2_32_2","unstructured":"NVIDIA Corporation. 2020. NVIDIA Jetson TX2 Board. Retrieved from https:\/\/developer.nvidia.com\/embedded\/jetson-tx2."},{"key":"e_1_3_2_33_2","unstructured":"NVIDIA Corporation. 2020. NVIDIA Jetson Xavier NX Board. Retrieved from https:\/\/developer.nvidia.com\/embedded\/jetson-xavier-nx-devkit."},{"key":"e_1_3_2_34_2","unstructured":"NVIDIA Corporation. 2020. Tegrastats Utility. Retrieved from https:\/\/docs.nvidia.com\/jetson\/archives\/l4t-archived\/l4t-3231\/index.html#page\/Tegra%20Linux%20Driver%20Package%20Development%20Guide\/AppendixTegraStats.html."},{"key":"e_1_3_2_35_2","doi-asserted-by":"crossref","first-page":"101896","DOI":"10.1016\/j.sysarc.2020.101896","article-title":"Benchmarking vision kernels and neural network inference accelerators on embedded platforms","volume":"113","author":"Qasaimeh Murad","year":"2021","unstructured":"Murad Qasaimeh, Kristof Denolf, Alireza Khodamoradi, Michaela Blott, Jack Lo, Lisa Halder, Kees Vissers, Joseph Zambreno, and Phillip H. Jones. 2021. Benchmarking vision kernels and neural network inference accelerators on embedded platforms. J. Syst. Architect. 113 (2021), 101896.","journal-title":"J. Syst. Architect."},{"issue":"9","key":"e_1_3_2_36_2","doi-asserted-by":"crossref","first-page":"1951","DOI":"10.3390\/s17091951","article-title":"A mobile outdoor augmented reality method combining deep learning object detection and spatial relationships for geovisualization","volume":"17","author":"Rao Jinmeng","year":"2017","unstructured":"Jinmeng Rao, Yanjun Qiao, Fu Ren, Junxing Wang, and Qingyun Du. 2017. A mobile outdoor augmented reality method combining deep learning object detection and spatial relationships for geovisualization. Sensors 17, 9 (2017), 1951.","journal-title":"Sensors"},{"key":"e_1_3_2_37_2","first-page":"779","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201916)","author":"Redmon Joseph","year":"2016","unstructured":"Joseph Redmon, Santosh Divvala, Ross Girshick, and Ali Farhadi. 2016. You only look once: Unified, real-time object detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201916). 779\u2013788."},{"key":"e_1_3_2_38_2","unstructured":"Joseph Redmon and Ali Farhadi. 2018. YOLOv3: An Incremental Improvement. Retrieved from https:\/\/arxiv.org\/abs\/1804.02767. 10.48550\/ARXIV.1804.02767"},{"key":"e_1_3_2_39_2","first-page":"91","volume-title":"Proceedings of the Advances in Neural Information Processing Systems (NeurIPS\u201915)","author":"Ren Shaoqing","year":"2015","unstructured":"Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. 2015. Faster R-CNN: Towards real-time object detection with region proposal networks. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS\u201915). 91\u201399."},{"key":"e_1_3_2_40_2","doi-asserted-by":"crossref","first-page":"69575","DOI":"10.1109\/ACCESS.2019.2919332","article-title":"Convolutional neural network-based real-time object detection and tracking for parrot AR drone 2","volume":"7","author":"Rohan Ali","year":"2019","unstructured":"Ali Rohan, Mohammed Rabah, and Sung-Ho Kim. 2019. Convolutional neural network-based real-time object detection and tracking for parrot AR drone 2. IEEE Access 7 (2019), 69575\u201369584.","journal-title":"IEEE Access"},{"issue":"3","key":"e_1_3_2_41_2","doi-asserted-by":"crossref","first-page":"211","DOI":"10.1007\/s11263-015-0816-y","article-title":"ImageNet large-scale visual recognition challenge","volume":"115","author":"Russakovsky Olga","year":"2015","unstructured":"Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei. 2015. ImageNet large-scale visual recognition challenge. Int. J. Comput. Vision 115, 3 (2015), 211\u2013252.","journal-title":"Int. J. Comput. Vision"},{"key":"e_1_3_2_42_2","first-page":"10536","volume-title":"Proceedings of the IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS\u201920)","author":"Sabater Alberto","year":"2020","unstructured":"Alberto Sabater, Luis Montesano, and Ana C. Murillo. 2020. Robust and efficient post-processing for video object detection. In Proceedings of the IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS\u201920). IEEE, 10536\u201310542."},{"key":"e_1_3_2_43_2","first-page":"4510","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201918)","author":"Sandler Mark","year":"2018","unstructured":"Mark Sandler, Andrew Howard, Menglong Zhu, Andrey Zhmoginov, and Liang-Chieh Chen. 2018. Mobilenetv2: Inverted residuals and linear bottlenecks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201918). 4510\u20134520."},{"key":"e_1_3_2_44_2","first-page":"2820","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Tan Mingxing","year":"2019","unstructured":"Mingxing Tan, Bo Chen, Ruoming Pang, Vijay Vasudevan, Mark Sandler, Andrew Howard, and Quoc V. Le. 2019. Mnasnet: Platform-aware neural architecture search for mobile. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 2820\u20132828."},{"key":"e_1_3_2_45_2","first-page":"6105","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Tan Mingxing","year":"2019","unstructured":"Mingxing Tan and Quoc Le. 2019. Efficientnet: Rethinking model scaling for convolutional neural networks. In Proceedings of the International Conference on Machine Learning. PMLR, 6105\u20136114."},{"key":"e_1_3_2_46_2","first-page":"10781","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR\u201920)","author":"Tan Mingxing","year":"2020","unstructured":"Mingxing Tan, Ruoming Pang, and Quoc V. Le. 2020. EfficientDet: Scalable and efficient object detection. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR\u201920). 10781\u201310790."},{"key":"e_1_3_2_47_2","first-page":"10734","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Wu Bichen","year":"2019","unstructured":"Bichen Wu, Xiaoliang Dai, Peizhao Zhang, Yanghan Wang, Fei Sun, Yiming Wu, Yuandong Tian, Peter Vajda, Yangqing Jia, and Kurt Keutzer. 2019. Fbnet: Hardware-aware efficient convnet design via differentiable neural architecture search. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 10734\u201310742."},{"key":"e_1_3_2_48_2","first-page":"9217","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV\u201919)","author":"Wu Haiping","year":"2019","unstructured":"Haiping Wu, Yuntao Chen, Naiyan Wang, and Zhaoxiang Zhang. 2019. Sequence level semantics aggregation for video object detection. In Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV\u201919). 9217\u20139225."},{"key":"e_1_3_2_49_2","doi-asserted-by":"publisher","DOI":"10.1145\/3463530"},{"key":"e_1_3_2_50_2","first-page":"449","volume-title":"Proceedings of the 18th Conference on Embedded Networked Sensor Systems (SenSys\u201920)","author":"Xu Ran","year":"2020","unstructured":"Ran Xu, Chen-lin Zhang, Pengcheng Wang, Jayoung Lee, Subrata Mitra, Somali Chaterji, Yin Li, and Saurabh Bagchi. 2020. ApproxDet: Content and contention-aware approximate object detection for mobiles. In Proceedings of the 18th Conference on Embedded Networked Sensor Systems (SenSys\u201920). 449\u2013462."},{"key":"e_1_3_2_51_2","first-page":"160","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Yao Chun-Han","year":"2020","unstructured":"Chun-Han Yao, Chen Fang, Xiaohui Shen, Yangyue Wan, and Ming-Hsuan Yang. 2020. Video object detection via object-level temporal aggregation. In Proceedings of the European Conference on Computer Vision. Springer, 160\u2013177."},{"key":"e_1_3_2_52_2","first-page":"4203","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201918)","author":"Zhang Shifeng","year":"2018","unstructured":"Shifeng Zhang, Longyin Wen, Xiao Bian, Zhen Lei, and Stan Z. Li. 2018. Single-shot refinement neural network for object detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201918). 4203\u20134212."},{"key":"e_1_3_2_53_2","first-page":"7210","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Zhu Xizhou","year":"2018","unstructured":"Xizhou Zhu, Jifeng Dai, Lu Yuan, and Yichen Wei. 2018. Towards high-performance video object detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 7210\u20137218."},{"key":"e_1_3_2_54_2","first-page":"408","volume-title":"Proceedings of the IEEE International Conference on Computer Vision","author":"Zhu Xizhou","year":"2017","unstructured":"Xizhou Zhu, Yujie Wang, Jifeng Dai, Lu Yuan, and Yichen Wei. 2017. Flow-guided feature aggregation for video object detection. In Proceedings of the IEEE International Conference on Computer Vision. 408\u2013417."},{"key":"e_1_3_2_55_2","first-page":"547","volume-title":"Proceedings of the 45th Annual International Symposium on Computer Architecture (ISCA\u201918)","author":"Zhu Yuhao","year":"2018","unstructured":"Yuhao Zhu, Anand Samajdar, Matthew Mattina, and Paul Whatmough. 2018. Euphrates: Algorithm-SoC co-design for low-power mobile continuous vision. In Proceedings of the 45th Annual International Symposium on Computer Architecture (ISCA\u201918). IEEE Press, 547\u2013560. 10.1109\/ISCA.2018.00052"}],"container-title":["ACM Transactions on Design Automation of Electronic Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3564289","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3564289","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T18:09:10Z","timestamp":1750183750000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3564289"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,4,3]]},"references-count":54,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2023,5,31]]}},"alternative-id":["10.1145\/3564289"],"URL":"https:\/\/doi.org\/10.1145\/3564289","relation":{},"ISSN":["1084-4309","1557-7309"],"issn-type":[{"value":"1084-4309","type":"print"},{"value":"1557-7309","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,4,3]]},"assertion":[{"value":"2022-01-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2022-09-08","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-04-03","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}