{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,12,6]],"date-time":"2025-12-06T17:07:09Z","timestamp":1765040829192,"version":"build-2065373602"},"reference-count":49,"publisher":"MDPI AG","issue":"1","license":[{"start":{"date-parts":[[2018,12,20]],"date-time":"2018-12-20T00:00:00Z","timestamp":1545264000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100010198","name":"Ministerio de Econom\u00eda, Industria y Competitividad, Gobierno de Espa\u00f1a","doi-asserted-by":"publisher","award":["TEC2014-53176-R (HAVideo)"],"award-info":[{"award-number":["TEC2014-53176-R (HAVideo)"]}],"id":[{"id":"10.13039\/501100010198","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Applying people detectors to unseen data is challenging since patterns distributions, such as viewpoints, motion, poses, backgrounds, occlusions and people sizes, may significantly differ from the ones of the training dataset. In this paper, we propose a coarse-to-fine framework to adapt frame by frame people detectors during runtime classification, without requiring any additional manually labeled ground truth apart from the offline training of the detection model. Such adaptation make use of multiple detectors mutual information, i.e., similarities and dissimilarities of detectors estimated and agreed by pair-wise correlating their outputs. Globally, the proposed adaptation discriminates between relevant instants in a video sequence, i.e., identifies the representative frames for an adaptation of the system. Locally, the proposed adaptation identifies the best configuration (i.e., detection threshold) of each detector under analysis, maximizing the mutual information to obtain the detection threshold of each detector. The proposed coarse-to-fine approach does not require training the detectors for each new scenario and uses standard people detector outputs, i.e., bounding boxes. The experimental results demonstrate that the proposed approach outperforms state-of-the-art detectors whose optimal threshold configurations are previously determined and fixed from offline training data.<\/jats:p>","DOI":"10.3390\/s19010004","type":"journal-article","created":{"date-parts":[[2018,12,20]],"date-time":"2018-12-20T12:54:36Z","timestamp":1545310476000},"page":"4","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":2,"title":["Coarse-to-Fine Adaptive People Detection for Video Sequences by Maximizing Mutual Information \u2020"],"prefix":"10.3390","volume":"19","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-1705-3972","authenticated-orcid":false,"given":"\u00c1lvaro","family":"Garc\u00eda-Mart\u00edn","sequence":"first","affiliation":[{"name":"Video Processing and Understanding Lab (VPULab), Universidad Aut\u00f3noma de Madrid, 28049 Madrid, Spain"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4999-2851","authenticated-orcid":false,"given":"Juan C.","family":"SanMiguel","sequence":"additional","affiliation":[{"name":"Video Processing and Understanding Lab (VPULab), Universidad Aut\u00f3noma de Madrid, 28049 Madrid, Spain"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2236-1769","authenticated-orcid":false,"given":"Jos\u00e9 M.","family":"Mart\u00ednez","sequence":"additional","affiliation":[{"name":"Video Processing and Understanding Lab (VPULab), Universidad Aut\u00f3noma de Madrid, 28049 Madrid, Spain"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2018,12,20]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"1137","DOI":"10.1109\/TPAMI.2016.2577031","article-title":"Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks","volume":"39","author":"Ren","year":"2017","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Redmon, J., and Farhadi, A. (2017, January 21\u201326). YOLO9000: Better, Faster, Stronger. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.690"},{"key":"ref_3","unstructured":"Xingyu, Z., Wanli, O., Meng, W., and Xiaogang, W. (2014, January 6\u201312). Deep Learning of Scene-Specific Classifier for Pedestrian Detection. Proceedings of the European Conference on Computer Vision (ECCV), Zurich, Switzerland."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"361","DOI":"10.1109\/TPAMI.2013.124","article-title":"Scene-Specific Pedestrian Detection for Static Video Surveillance","volume":"36","author":"Wang","year":"2014","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Royer, A., and Lampert, C.H. (2015, January 7\u201312). Classifier adaptation at prediction time. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298746"},{"key":"ref_6","unstructured":"Kalinke, T., Tzomakas, C., and Seelen, W.V. (1998, January 28\u201330). A Texture-based Object Detection and an adaptive Model-based Classification. Proceedings of the IEEE Intelligent Vehicles Symposium, Stuttgart, Germany."},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Zhang, S., Zhu, Q., and Roy-Chowdhury, A. (2016, January 25\u201328). Adaptive algorithm selection, with applications in pedestrian detection. Proceedings of the IEEE International Conference on Image Processing (ICIP), Phoenix, AZ, USA.","DOI":"10.1109\/ICIP.2016.7533064"},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"233","DOI":"10.1109\/TIP.2015.2499702","article-title":"Detect2Rank: Combining Object Detectors Using Learning to Rank","volume":"25","author":"Karaoglu","year":"2016","journal-title":"IEEE Trans. Image Process."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"142","DOI":"10.1016\/j.engappai.2016.01.029","article-title":"Adapting pedestrian detectors to new domains: A comprehensive review","volume":"50","author":"Htike","year":"2016","journal-title":"Eng. Appl. Artif. Intell."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Dimou, A., and Alvarez, F. (2016, January 25\u201328). Multi-target detection in CCTV footage for tracking applications using deep learning techniques. Proceedings of the IEEE International Conference on Image Processing (ICIP), Phoenix, AZ, USA.","DOI":"10.1109\/ICIP.2016.7532493"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Mees, O., Eitel, A., and Burgard, W. (2016, January 9\u201314). Choosing Smartly: Adaptive Multimodal Fusion for Object Detection in Changing Environments. Proceedings of the IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS), Daejeon, Korea.","DOI":"10.1109\/IROS.2016.7759048"},{"key":"ref_12","unstructured":"Verma, A., Hebbalaguppe, R., Vig, L., Kumar, S., and Hassan, E. (July, January 26). Pedestrian Detection via Mixture of CNN Experts and Thresholded Aggregated Channel Features. Proceedings of the IEEE International Conference on Computer Vision (ICCV), Las Vegas, NV."},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"1409","DOI":"10.1109\/TPAMI.2011.239","article-title":"Tracking-Learning-Detection","volume":"34","author":"Kalal","year":"2012","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_14","unstructured":"Gaidon, A., Zen, G., and Rodriguez, J. (2014, January 6\u201312). Self-Learning Camera: Autonomous Adaption of Object Detectors to Unlabeled Video Streams. Proceedings of the European Conference on Computer Vision (ECCV), Zurich, Switzerland."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Garcia-Martin, A., and SanMiguel, J.C. (2017, January 17\u201320). Adaptive people detection based on cross-correlation maximization. Proceedings of the IEEE International Conference on Image Processing (ICIP), Beijing, China.","DOI":"10.1109\/ICIP.2017.8296910"},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"2367","DOI":"10.1109\/TPAMI.2014.2327973","article-title":"Domain Adaptation of Deformable Part-Based Models","volume":"36","author":"Xu","year":"2014","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Roth, P.M., Sternig, S., Grabner, H., and Bischof, H. (2009, January 20\u201326). Classifier grids for robust adaptive object detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Miami, FL, USA.","DOI":"10.1109\/CVPRW.2009.5206616"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Liu, S., and Kovashka, A. (2016, January 7\u201310). Adapting attributes by selecting features similar across domains. Proceedings of the IEEE Winter Conference on Applications of Computer Vision (WACV), Lake Placid, NY, USA.","DOI":"10.1109\/WACV.2016.7477731"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Shu, G., Dehghan, A., and Shah, M. (2013, January 23\u201328). Improving an object detector and extracting regions using superpixels. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Portland, OR, USA.","DOI":"10.1109\/CVPR.2013.477"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Ye, Q., Zhang, T., Ke, W., Qiu, Q., Chen, J., Sapiro, G., and Zhang, B. (2017, January 21\u201326). Self-Learning Scene-Specific Pedestrian Detectors Using a Progressive Latent Model. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.222"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Chen, Y., Li, W., Sakaridis, C., Dai, D., and Van Gool, L. (2018, January 18\u201322). Domain Adaptive Faster R-CNN for Object Detection in the Wild. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00352"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Hattori, H., Boddeti, V.N., Kitani, K., and Kanade, T. (2015, January 7\u201312). Learning scene-specific pedestrian detectors without real data. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7299006"},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"797","DOI":"10.1109\/TPAMI.2013.163","article-title":"Virtual and real world adaptationfor pedestrian detection","volume":"36","author":"Vazquez","year":"2014","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"345","DOI":"10.1016\/j.imavis.2012.03.005","article-title":"On collaborative people detection and tracking in complex scenarios","volume":"30","author":"Martinez","year":"2012","journal-title":"Image Vis. Comput."},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"932","DOI":"10.1016\/j.robot.2013.05.002","article-title":"Indoor scene recognition by a mobile robot through adaptive object detection","volume":"61","author":"Espinace","year":"2013","journal-title":"Robot. Auton. Syst."},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"1865","DOI":"10.1049\/el.2015.3099","article-title":"Context-aware part-based people detection for video monitoring","volume":"51","author":"SanMiguel","year":"2015","journal-title":"Electron. Lett."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Singh, K.K., Divvala, S., Farhadi, A., and Lee, Y.J. (2018, January 8\u201314). DOCK: Detecting Objects by transferring Common-sense Knowledge. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01261-8_30"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Kang, J.K., Hong, H.G., and Park, K.R. (2017). Pedestrian Detection Based on Adaptive Selection of Visible Light or Far-Infrared Light Camera Image by Fuzzy Inference System and Convolutional Neural Network-Based Verification. Sensors, 17.","DOI":"10.3390\/s17071598"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Fleet, D., Pajdla, T., Schiele, B., and Tuytelaars, T. (2014). Statistical and Spatial Consensus Collection for Detector Adaptation. Computer Vision\u2014ECCV 2014: 13th European Conference, Zurich, Switzerland, 6\u201312 September 2014, Proceedings, Part III, Springer International Publishing.","DOI":"10.1007\/978-3-319-10578-9"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Conaire, C.O., O\u2019Connor, N.E., and Smeaton, A.F. (2007, January 18\u201323). Detector adaptation by maximising agreement between independent data sources. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Minneapolis, MN, USA.","DOI":"10.1109\/CVPR.2007.383448"},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"2102","DOI":"10.1016\/j.patrec.2013.07.016","article-title":"Skin detection by dual maximization of detectors agreement for video monitoring","volume":"34","author":"SanMiguel","year":"2013","journal-title":"Pattern Recognit. Lett."},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"1627","DOI":"10.1109\/TPAMI.2009.167","article-title":"Object Detection with Discriminatively Trained Part-Based Models","volume":"32","author":"Felzenszwalb","year":"2010","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"76","DOI":"10.1016\/j.cviu.2014.09.010","article-title":"Post-processing approaches for improving people detection performance","volume":"133","author":"Martinez","year":"2015","journal-title":"Comput. Vis. Image Underst."},{"key":"ref_34","unstructured":"Leibe, B., Seemann, E., and Schiele, B. (2005, January 20\u201326). Pedestrian Detection in Crowded Scenes. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), San Diego, CA, USA."},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"779","DOI":"10.1049\/iet-cvi.2014.0148","article-title":"People detection in surveillance: Classification and evaluation","volume":"9","author":"Martinez","year":"2015","journal-title":"IET Comput. Vis."},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Ionescu, B., Benois-Pineau, J., Piatrik, T., and Quenot, G. (2014). Fusion in Computer Vision: Understanding Complex Visual Content, Springer.","DOI":"10.1007\/978-3-319-05696-8"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Baruque, B., and Corchado, E. (2011). Fusion Methods for Unsupervised Learning Ensembles, Springer.","DOI":"10.1007\/978-3-642-16205-3"},{"key":"ref_38","doi-asserted-by":"crossref","first-page":"671","DOI":"10.1126\/science.220.4598.671","article-title":"Optimization by Simulated Annealing","volume":"220","author":"Kirkpatrick","year":"1983","journal-title":"Science"},{"key":"ref_39","doi-asserted-by":"crossref","first-page":"146","DOI":"10.1007\/BF01386306","article-title":"Cauchy\u2019s method of minimization","volume":"4","author":"Goldstein","year":"1962","journal-title":"Numer. Math."},{"key":"ref_40","doi-asserted-by":"crossref","first-page":"559","DOI":"10.1049\/el.2014.3795","article-title":"PDbm: People detection benchmark repository","volume":"51","author":"Alcedo","year":"2015","journal-title":"Electron. Lett."},{"key":"ref_41","doi-asserted-by":"crossref","first-page":"1532","DOI":"10.1109\/TPAMI.2014.2300479","article-title":"Fast Feature Pyramids for Object Detection","volume":"36","author":"Dollar","year":"2014","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_42","unstructured":"Everingham, M., Van Gool, L., Williams, C.K.I., Winn, J., and Zisserman, A. (2018, December 19). The PASCAL Visual Object Classes Challenge 2012 (VOC2012) Results. Available online: http:\/\/host.robots.ox.ac.uk\/pascal\/VOC\/voc2012\/."},{"key":"ref_43","doi-asserted-by":"crossref","first-page":"58","DOI":"10.1109\/TPAMI.2013.103","article-title":"Continuous Energy Minimization for Multitarget Tracking","volume":"36","author":"Milan","year":"2014","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_44","unstructured":"PETS (2018, December 19). International Workshop on Performance Evaluation of Tracking and Surveillance. Available online: http:\/\/www.cvg.reading.ac.uk\/PETS2009\/a.html."},{"key":"ref_45","doi-asserted-by":"crossref","first-page":"438","DOI":"10.1109\/76.313138","article-title":"A new three-step search algorithm for block motion estimation","volume":"4","author":"Li","year":"1994","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_46","doi-asserted-by":"crossref","first-page":"313","DOI":"10.1109\/76.499840","article-title":"A novel four-step search algorithm for fast block motion estimation","volume":"6","year":"1996","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_47","doi-asserted-by":"crossref","first-page":"287","DOI":"10.1109\/TIP.2000.826791","article-title":"A new diamond search algorithm for fast block-matching motion estimation","volume":"9","author":"Zhu","year":"2000","journal-title":"IEEE Trans. Image Process."},{"key":"ref_48","doi-asserted-by":"crossref","first-page":"212","DOI":"10.1145\/321062.321069","article-title":"\u201c Direct Search\u201d Solution of Numerical and Statistical Problems","volume":"8","author":"Hooke","year":"1961","journal-title":"J. ACM"},{"key":"ref_49","unstructured":"Kennedy, J., and Eberhart, R. (December, January 27). Particle swarm optimization. Proceedings of the IEEE International Conference on Neural Networks, Perth, Australia."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/19\/1\/4\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T15:35:09Z","timestamp":1760196909000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/19\/1\/4"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2018,12,20]]},"references-count":49,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2019,1]]}},"alternative-id":["s19010004"],"URL":"https:\/\/doi.org\/10.3390\/s19010004","relation":{},"ISSN":["1424-8220"],"issn-type":[{"type":"electronic","value":"1424-8220"}],"subject":[],"published":{"date-parts":[[2018,12,20]]}}}