{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,11,21]],"date-time":"2025-11-21T18:14:38Z","timestamp":1763748878529,"version":"3.41.0"},"reference-count":54,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2023,4,19]],"date-time":"2023-04-19T00:00:00Z","timestamp":1681862400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/100019923","name":"DEVCOM Army Research Laboratory","doi-asserted-by":"crossref","award":["W911NF-09-2-0053 and W911NF-17-2-0196"],"award-info":[{"award-number":["W911NF-09-2-0053 and W911NF-17-2-0196"]}],"id":[{"id":"10.13039\/100019923","id-type":"DOI","asserted-by":"crossref"}]},{"name":"NSF CPS","award":["1544969"],"award-info":[{"award-number":["1544969"]}]},{"name":"NSF CNS","award":["2038817"],"award-info":[{"award-number":["2038817"]}]},{"name":"ONR","award":["N00014-19-1-2264"],"award-info":[{"award-number":["N00014-19-1-2264"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Cyber-Phys. Syst."],"published-print":{"date-parts":[[2023,4,30]]},"abstract":"<jats:p>\n            Steerable cameras that can be controlled via a network, to retrieve telemetries of interest have become popular. In this paper, we develop a framework called\n            <jats:monospace>AcTrak<\/jats:monospace>\n            , to automate a camera\u2019s motion to appropriately switch between (a) zoom ins on existing targets in a scene to track their activities, and (b) zoom out to search for new targets arriving to the area of interest. Specifically, we seek to achieve a good trade-off between the two tasks, i.e., we want to ensure that new targets are observed by the camera before they leave the scene, while also zooming in on existing targets frequently enough to monitor their activities. There exist prior control algorithms for steering cameras to optimize certain objectives; however, to the best of our knowledge, none have considered this problem, and do not perform well when target activity tracking is required.\n            <jats:monospace>AcTrak<\/jats:monospace>\n            \u00a0automatically controls the camera\u2019s PTZ configurations using\n            <jats:bold>reinforcement learning (RL<\/jats:bold>\n            ), to select the best camera position given the current state. Via simulations using real datasets, we show that\n            <jats:monospace>AcTrak<\/jats:monospace>\n            detects newly arriving targets\n            <jats:italic>30%<\/jats:italic>\n            faster than a non-adaptive baseline and rarely misses targets, unlike the baseline which can miss up to 5% of the targets. We also implement\n            <jats:monospace>AcTrak<\/jats:monospace>\n            to control a real camera and demonstrate that in comparison with the baseline, it acquires about\n            <jats:italic>2\u00d7<\/jats:italic>\n            more high resolution images of targets.\n          <\/jats:p>","DOI":"10.1145\/3585316","type":"journal-article","created":{"date-parts":[[2023,3,3]],"date-time":"2023-03-03T11:08:46Z","timestamp":1677841726000},"page":"1-27","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":5,"title":["AcTrak: Controlling a Steerable Surveillance Camera using Reinforcement Learning"],"prefix":"10.1145","volume":"7","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-0801-8737","authenticated-orcid":false,"given":"Abdulrahman","family":"Fahim","sequence":"first","affiliation":[{"name":"University of California, Riverside, California, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3411-8483","authenticated-orcid":false,"given":"Evangelos","family":"Papalexakis","sequence":"additional","affiliation":[{"name":"University of California, Riverside, California, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6533-4381","authenticated-orcid":false,"given":"Srikanth V.","family":"Krishnamurthy","sequence":"additional","affiliation":[{"name":"University of California, Riverside, California, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6690-9725","authenticated-orcid":false,"given":"Amit","family":"K. Roy Chowdhury","sequence":"additional","affiliation":[{"name":"University of California, Riverside, California, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3627-4471","authenticated-orcid":false,"given":"Lance","family":"Kaplan","sequence":"additional","affiliation":[{"name":"DEVCOM Army Research Laboratory, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3883-7220","authenticated-orcid":false,"given":"Tarek","family":"Abdelzaher","sequence":"additional","affiliation":[{"name":"University of Illinois at Urbana Champaign, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2023,4,19]]},"reference":[{"key":"e_1_3_3_2_2","unstructured":"The Verge. Dec 2019. The US like China has about one surveillance camera for every four people says report. https:\/\/www.theverge.com\/2019\/12\/9\/21002515\/surveillance-cameras-globally-us-china-amount-citizens. (Dec 2019). [Online; accessed March-30-2020]."},{"key":"e_1_3_3_3_2","unstructured":"Andy Penfold. [n. d.]. How IoT is reshaping the future of Video Surveillance. https:\/\/www.securityandsafetythings.com\/insights\/iot-reshaping-future-surveillance."},{"key":"e_1_3_3_4_2","doi-asserted-by":"crossref","first-page":"771","DOI":"10.1007\/978-981-10-3874-7_73","volume-title":"Computational Intelligence in Data Mining","author":"Gulve Sonali P.","year":"2017","unstructured":"Sonali P. Gulve, Suchitra A. Khoje, and Prajakta Pardeshi. 2017. Implementation of IoT-based smart video surveillance system. In Computational Intelligence in Data Mining. Springer, 771\u2013780."},{"key":"e_1_3_3_5_2","unstructured":"Avipas. [n. d.]. AViPAS model AV-1081 manual. https:\/\/e7aba150-670b-4b8b-9a25-311a84251d5f.filesusr.com\/ugd\/6b6a18_34ff6afc20914be9943327dc3f5a6211.pdf."},{"key":"e_1_3_3_6_2","unstructured":"Sony. [n. d.]. Remotely controlled PTZ color video camera with IP streaming. https:\/\/pro.sony\/ue_US\/products\/ptz-network-cameras\/srg-300se."},{"key":"e_1_3_3_7_2","unstructured":"alsecuritycamera. [n. d.]. Axis outdoor multi sensor IP security camera. https:\/\/www.a1securitycameras.com."},{"key":"e_1_3_3_8_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2015.2426575"},{"key":"e_1_3_3_9_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2016.2593620"},{"key":"e_1_3_3_10_2","first-page":"1","volume-title":"ICDSC","author":"Qureshi Faisal Z.","year":"2009","unstructured":"Faisal Z. Qureshi and Demetri Terzopoulos. 2009. Planning ahead for PTZ camera assignment and handoff. In ICDSC. IEEE, 1\u20138."},{"key":"e_1_3_3_11_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.cviu.2010.01.003"},{"key":"e_1_3_3_12_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.cviu.2009.04.004"},{"key":"e_1_3_3_13_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICDSC.2011.6042928"},{"key":"e_1_3_3_14_2","doi-asserted-by":"publisher","DOI":"10.1145\/2529992"},{"key":"e_1_3_3_15_2","first-page":"1","volume-title":"AVSS","author":"Neves Joao C.","year":"2015","unstructured":"Joao C. Neves and Hugo Proen\u00e7a. 2015. Dynamic camera scheduling for visual surveillance in crowded scenes using Markov random fields. In AVSS. IEEE, 1\u20136."},{"key":"e_1_3_3_16_2","doi-asserted-by":"crossref","first-page":"119","DOI":"10.1145\/3055031.3055085","volume-title":"IPSN","author":"Jain Shubham","year":"2017","unstructured":"Shubham Jain, Viet Nguyen, Marco Gruteser, and Paramvir Bahl. 2017. Panoptes: Servicing multiple applications simultaneously using steerable cameras. In IPSN. 119\u2013130."},{"key":"e_1_3_3_17_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v32i1.11796"},{"key":"e_1_3_3_18_2","volume-title":"DAC","author":"Wei Tianshu","year":"2017","unstructured":"Tianshu Wei, Yanzhi Wang, and Qi Zhu. 2017. Deep reinforcement learning for building HVAC control. In DAC."},{"key":"e_1_3_3_19_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00708"},{"key":"e_1_3_3_20_2","doi-asserted-by":"publisher","DOI":"10.1109\/INFOCOM42981.2021.9488770"},{"key":"e_1_3_3_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICDCS51616.2021.00058"},{"key":"e_1_3_3_22_2","doi-asserted-by":"publisher","DOI":"10.1016\/S0004-3702(99)00052-1"},{"key":"e_1_3_3_23_2","doi-asserted-by":"publisher","DOI":"10.1007\/BF00992698"},{"key":"e_1_3_3_24_2","volume-title":"Reinforcement Learning: An Introduction","author":"Sutton Richard S.","year":"2018","unstructured":"Richard S. Sutton and Andrew G. Barto. 2018. Reinforcement Learning: An Introduction. A Bradford Book, Cambridge, MA, USA."},{"key":"e_1_3_3_25_2","doi-asserted-by":"publisher","DOI":"10.5555\/3016100.3016191"},{"key":"e_1_3_3_26_2","article-title":"Dueling network architectures for deep reinforcement learning","author":"Wang Ziyu","year":"2015","unstructured":"Ziyu Wang, Tom Schaul, Matteo Hessel, Hado Van Hasselt, Marc Lanctot, and Nando De Freitas. 2015. Dueling network architectures for deep reinforcement learning. arXiv preprint arXiv:1511.06581 (2015).","journal-title":"arXiv preprint arXiv:1511.06581"},{"key":"e_1_3_3_27_2","unstructured":"darknet. [n. d.]. YOLO: Real-Time Object Detection. https:\/\/pjreddie.com\/darknet\/yolo\/. [Online; accessed August-10-2022]."},{"key":"e_1_3_3_28_2","article-title":"YOLOv3: An incremental improvement","author":"Redmon Joseph","year":"2018","unstructured":"Joseph Redmon and Ali Farhadi. 2018. YOLOv3: An incremental improvement. arXiv (2018).","journal-title":"arXiv"},{"key":"e_1_3_3_29_2","first-page":"791","volume-title":"ECCV","author":"Varior Rahul Rama","year":"2016","unstructured":"Rahul Rama Varior, Mrinal Haloi, and Gang Wang. 2016. Gated siamese convolutional neural network architecture for human re-identification. In ECCV. Springer, 791\u2013808."},{"key":"e_1_3_3_30_2","volume-title":"European Conference on Computer Vision","author":"Assari Shayan Modiri","year":"2016","unstructured":"Shayan Modiri Assari, Haroon Idrees, and Mubarak Shah. 2016. Human re-identification in crowd videos using personal, social and environmental constraints. In European Conference on Computer Vision. Springer."},{"key":"e_1_3_3_31_2","article-title":"MobileNets: Efficient convolutional neural networks for mobile vision applications","author":"Howard Andrew G.","year":"2017","unstructured":"Andrew G. Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam. 2017. MobileNets: Efficient convolutional neural networks for mobile vision applications. arXiv preprint arXiv:1704.04861 (2017).","journal-title":"arXiv preprint arXiv:1704.04861"},{"key":"e_1_3_3_32_2","article-title":"Monocular depth estimation: A survey","author":"Bhoi Amlaan","year":"2019","unstructured":"Amlaan Bhoi. 2019. Monocular depth estimation: A survey. arXiv preprint arXiv:1901.09402 (2019).","journal-title":"arXiv preprint arXiv:1901.09402"},{"key":"e_1_3_3_33_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00043"},{"key":"e_1_3_3_34_2","first-page":"789","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV)","author":"Garanderie Greire Payen de La","year":"2018","unstructured":"Greire Payen de La Garanderie, Amir Atapour Abarghouei, and Toby P. Breckon. 2018. Eliminating the blind spot: Adapting 3D object detection and monocular depth estimation to 360 panoramic imagery. In Proceedings of the European Conference on Computer Vision (ECCV). 789\u2013807."},{"key":"e_1_3_3_35_2","article-title":"Calibrating self-supervised monocular depth estimation","author":"McCraith Robert","year":"2020","unstructured":"Robert McCraith, Lukas Neumann, and Andrea Vedaldi. 2020. Calibrating self-supervised monocular depth estimation. arXiv preprint arXiv:2009.07714 (2020).","journal-title":"arXiv preprint arXiv:2009.07714"},{"key":"e_1_3_3_36_2","doi-asserted-by":"publisher","DOI":"10.1109\/34.888718"},{"key":"e_1_3_3_37_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2012.250"},{"volume-title":"ICML","author":"Drucker Harris","key":"e_1_3_3_38_2","unstructured":"Harris Drucker. Improving regressors using boosting techniques. In ICML."},{"volume-title":"CVPR 2011","author":"Oh Sangmin","key":"e_1_3_3_39_2","unstructured":"Sangmin Oh, Anthony Hoogs, Amitha Perera, Naresh Cuntoor, Chia-Chih Chen, Jong Taek Lee, Saurajit Mukherjee, J. K. Aggarwal, Hyungtae Lee, Larry Davis, et\u00a0al. 2011. A large-scale benchmark dataset for event recognition in surveillance video. In CVPR 2011."},{"key":"e_1_3_3_40_2","volume-title":"Computer Graphics Forum","author":"Lerner Alon","year":"2007","unstructured":"Alon Lerner, Yiorgos Chrysanthou, and Dani Lischinski. 2007. Crowds by example. In Computer Graphics Forum. Wiley Online Library."},{"key":"e_1_3_3_41_2","volume-title":"International Conference on Learning Representations","author":"Yoon Jinsung","year":"2018","unstructured":"Jinsung Yoon, William R. Zame, and Mihaela Van Der Schaar. 2018. Deep sensing: Active sensing using multi-directional recurrent neural networks. In International Conference on Learning Representations."},{"key":"e_1_3_3_42_2","first-page":"12345","volume-title":"CVPR","author":"Uzkent Burak","year":"2020","unstructured":"Burak Uzkent and Stefano Ermon. 2020. Learning when and where to zoom with deep reinforcement learning. In CVPR. 12345\u201312354."},{"key":"e_1_3_3_43_2","doi-asserted-by":"publisher","DOI":"10.1109\/TVT.2012.2231973"},{"key":"e_1_3_3_44_2","unstructured":"Amazon.com. [n. d.]. Amazon.com. https:\/\/www.amazon.com\/."},{"key":"e_1_3_3_45_2","unstructured":"bhphotovideo.com. [n. d.]. bhphotovideo.com. https:\/\/www.bhphotovideo.com\/. [Online; accessed June-29-2021]."},{"key":"e_1_3_3_46_2","doi-asserted-by":"publisher","DOI":"10.1109\/IHMSC.2013.240"},{"key":"e_1_3_3_47_2","article-title":"Hybrid reward architecture for reinforcement learning","volume":"30","author":"Seijen Harm Van","year":"2017","unstructured":"Harm Van Seijen, Mehdi Fatemi, Joshua Romoff, Romain Laroche, Tavian Barnes, and Jeffrey Tsang. 2017. Hybrid reward architecture for reinforcement learning. Advances in Neural Information Processing Systems 30 (2017).","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_3_48_2","article-title":"Population based training of neural networks","author":"Jaderberg Max","year":"2017","unstructured":"Max Jaderberg, Valentin Dalibard, Simon Osindero, Wojciech M. Czarnecki, Jeff Donahue, Ali Razavi, Oriol Vinyals, Tim Green, Iain Dunning, Karen Simonyan, et\u00a0al. 2017. Population based training of neural networks. arXiv preprint arXiv:1711.09846 (2017).","journal-title":"arXiv preprint arXiv:1711.09846"},{"key":"e_1_3_3_49_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11760-013-0484-8"},{"key":"e_1_3_3_50_2","doi-asserted-by":"publisher","DOI":"10.3390\/s20174691"},{"key":"e_1_3_3_51_2","doi-asserted-by":"publisher","DOI":"10.1145\/2659021.2659052"},{"key":"e_1_3_3_52_2","article-title":"Visual sensor network reconfiguration with deep reinforcement learning","author":"Jasek Paul","year":"2018","unstructured":"Paul Jasek and Bernard Abayowa. 2018. Visual sensor network reconfiguration with deep reinforcement learning. arXiv preprint arXiv:1808.04287 (2018).","journal-title":"arXiv preprint arXiv:1808.04287"},{"key":"e_1_3_3_53_2","volume-title":"ACM International Workshop on Video Surveillance and Sensor Networks","author":"Bagdanov Andrew D.","year":"2006","unstructured":"Andrew D. Bagdanov, Alberto Del Bimbo, Walter Nunziati, and Federico Pernici. 2006. A reinforcement learning approach to active camera foveation. In ACM International Workshop on Video Surveillance and Sensor Networks."},{"key":"e_1_3_3_54_2","volume-title":"ICEIC","author":"Kim Dongchil","year":"2019","unstructured":"Dongchil Kim, Kyoungman Kim, and Sungjoo Park. 2019. Automatic PTZ camera control based on deep-Q network in video surveillance system. In ICEIC. IEEE."},{"key":"e_1_3_3_55_2","doi-asserted-by":"publisher","DOI":"10.1145\/1943552.1943556"}],"container-title":["ACM Transactions on Cyber-Physical Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3585316","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3585316","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T16:37:56Z","timestamp":1750178276000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3585316"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,4,19]]},"references-count":54,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2023,4,30]]}},"alternative-id":["10.1145\/3585316"],"URL":"https:\/\/doi.org\/10.1145\/3585316","relation":{},"ISSN":["2378-962X","2378-9638"],"issn-type":[{"type":"print","value":"2378-962X"},{"type":"electronic","value":"2378-9638"}],"subject":[],"published":{"date-parts":[[2023,4,19]]},"assertion":[{"value":"2022-04-28","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-02-20","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-04-19","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}