{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,11]],"date-time":"2026-07-11T01:48:22Z","timestamp":1783734502190,"version":"3.55.0"},"reference-count":43,"publisher":"Association for Computing Machinery (ACM)","issue":"1","license":[{"start":{"date-parts":[[2021,2,28]],"date-time":"2021-02-28T00:00:00Z","timestamp":1614470400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2021,2,28]]},"abstract":"<jats:p>Pedestrian detection is a canonical problem for safety and security applications, and it remains a challenging problem due to the highly variable lighting conditions in which pedestrians must be detected. This article investigates several domain adaptation approaches to adapt RGB-trained detectors to the thermal domain. Building on our earlier work on domain adaptation for privacy-preserving pedestrian detection, we conducted an extensive experimental evaluation comparing top-down and bottom-up domain adaptation and also propose two new bottom-up domain adaptation strategies. For top-down domain adaptation, we leverage a detector pre-trained on RGB imagery and efficiently adapt it to perform pedestrian detection in the thermal domain. Our bottom-up domain adaptation approaches include two steps: first, training an adapter segment corresponding to initial layers of the RGB-trained detector adapts to the new input distribution; then, we reconnect the adapter segment to the original RGB-trained detector for final adaptation with a top-down loss. To the best of our knowledge, our bottom-up domain adaptation approaches outperform the best-performing single-modality pedestrian detection results on KAIST and outperform the state of the art on FLIR.<\/jats:p>","DOI":"10.1145\/3418213","type":"journal-article","created":{"date-parts":[[2021,4,16]],"date-time":"2021-04-16T12:42:08Z","timestamp":1618576928000},"page":"1-19","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":52,"title":["Bottom-up and Layerwise Domain Adaptation for Pedestrian Detection in Thermal Images"],"prefix":"10.1145","volume":"17","author":[{"given":"My","family":"Kieu","sequence":"first","affiliation":[{"name":"Media Integration and Communication Center (MICC), University of Florence, Italy"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Andrew D.","family":"Bagdanov","sequence":"additional","affiliation":[{"name":"Media Integration and Communication Center (MICC), University of Florence, Italy"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Marco","family":"Bertini","sequence":"additional","affiliation":[{"name":"Media Integration and Communication Center (MICC), University of Florence, Italy"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2021,4,16]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2019.8683026"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.5244\/C.29.32"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.3390\/s17081850"},{"key":"e_1_2_1_4_1","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV\u201914)","volume":"8926","author":"Benenson Rodrigo","year":"2014","unstructured":"Rodrigo Benenson , Mohamed Omran , Jan Hosang , and Bernt Schiele . 2014 . Ten years of pedestrian detection, what have we learned? . In Proceedings of the European Conference on Computer Vision (ECCV\u201914) , Vol. 8926 . Springer, 613\u2013627. Rodrigo Benenson, Mohamed Omran, Jan Hosang, and Bernt Schiele. 2014. Ten years of pedestrian detection, what have we learned?. In Proceedings of the European Conference on Computer Vision (ECCV\u201914), Vol. 8926. Springer, 613\u2013627."},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.530"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.isprsjprs.2019.02.005"},{"key":"e_1_2_1_7_1","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW\u201919)","author":"Devaguptapu Chaitanya","year":"2019","unstructured":"Chaitanya Devaguptapu , Ninad Akolekar , Manuj M. Sharma , and Vineeth N. Balasubramanian . 2019. Borrow from anywhere: Pseudo multi-modal object detection in thermal imagery . In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW\u201919) . 1029\u20131038. DOI:https:\/\/doi.org\/10.1109\/CVPRW. 2019 .00135 10.1109\/CVPRW.2019.00135 Chaitanya Devaguptapu, Ninad Akolekar, Manuj M. Sharma, and Vineeth N. Balasubramanian. 2019. Borrow from anywhere: Pseudo multi-modal object detection in thermal imagery. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW\u201919). 1029\u20131038. DOI:https:\/\/doi.org\/10.1109\/CVPRW.2019.00135"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2011.155"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/WACV.2017.111"},{"key":"e_1_2_1_10_1","unstructured":"FLIR. 2018. FLIR starter thermal dataset. Retrieved from https:\/\/www.flir.com\/oem\/adas\/adas-dataset-form\/.  FLIR. 2018. FLIR starter thermal dataset. Retrieved from https:\/\/www.flir.com\/oem\/adas\/adas-dataset-form\/."},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1117\/12.2520705"},{"key":"e_1_2_1_12_1","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW\u201919)","author":"Ghose Debasmita","year":"2019","unstructured":"Debasmita Ghose , Shasvat M. Desai , Sneha Bhattacharya , Deep Chakraborty , Madalina Fiterau , and Tauhidur Rahman . 2019 . Pedestrian detection in thermal images using saliency maps . In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW\u201919) . Debasmita Ghose, Shasvat M. Desai, Sneha Bhattacharya, Deep Chakraborty, Madalina Fiterau, and Tauhidur Rahman. 2019. Pedestrian detection in thermal images using saliency maps. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW\u201919)."},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.inffus.2018.11.017"},{"key":"e_1_2_1_14_1","volume-title":"AdapterNet\u2014Learning input transformation for domain adaptation. CoRRabs\/1805.11601","author":"Hazan Alon","year":"2018","unstructured":"Alon Hazan , Yoel Shoshan , Daniel Khapun , Roy Aladjem , and Vadim Ratner . 2018. AdapterNet\u2014Learning input transformation for domain adaptation. CoRRabs\/1805.11601 ( 2018 ). arXiv:1805.11601. http:\/\/arxiv.org\/abs\/1805.11601. Alon Hazan, Yoel Shoshan, Daniel Khapun, Roy Aladjem, and Vadim Ratner. 2018. AdapterNet\u2014Learning input transformation for domain adaptation. CoRRabs\/1805.11601 (2018). arXiv:1805.11601. http:\/\/arxiv.org\/abs\/1805.11601."},{"key":"e_1_2_1_15_1","volume-title":"Proceedings of the Autonomous Systems: Sensors, Vehicles, Security, and the Internet of Everything","volume":"10643","author":"Herrmann Christian","year":"2018","unstructured":"Christian Herrmann , Miriam Ruf , and J\u00fcrgen Beyerer . 2018 . CNN-based thermal infrared person detection by domain adaptation . In Proceedings of the Autonomous Systems: Sensors, Vehicles, Security, and the Internet of Everything , Vol. 10643 . International Society for Optics and Photonics, 1064308. Christian Herrmann, Miriam Ruf, and J\u00fcrgen Beyerer. 2018. CNN-based thermal infrared person detection by domain adaptation. In Proceedings of the Autonomous Systems: Sensors, Vehicles, Security, and the Internet of Everything, Vol. 10643. International Society for Optics and Photonics, 1064308."},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298706"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1109\/MVA.2015.7153177"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-30645-8_19"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW.2017.36"},{"key":"e_1_2_1_20_1","unstructured":"Wouter M. Kouw. 2018. An introduction to domain adaptation and transfer learning. arxiv:1812.11806. Retrieved fromhttp:\/\/arxiv.org\/abs\/1812.11806.  Wouter M. Kouw. 2018. An introduction to domain adaptation and transfer learning. arxiv:1812.11806. Retrieved fromhttp:\/\/arxiv.org\/abs\/1812.11806."},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.23919\/APSIPA.2018.8659688"},{"key":"e_1_2_1_22_1","volume-title":"Proceedings of the British Machine Vision Conference (BMVC\u201918)","author":"Li Chengyang","year":"2018","unstructured":"Chengyang Li , Dan Song , Ruofeng Tong , and Min Tang . 2018 . Multispectral pedestrian detection via simultaneous detection and segmentation . In Proceedings of the British Machine Vision Conference (BMVC\u201918) . BMVA Press, 225. Chengyang Li, Dan Song, Ruofeng Tong, and Min Tang. 2018. Multispectral pedestrian detection via simultaneous detection and segmentation. In Proceedings of the British Machine Vision Conference (BMVC\u201918). BMVA Press, 225."},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2018.08.005"},{"key":"e_1_2_1_24_1","first-page":"985","article-title":"Scale-aware fast R-CNN for pedestrian detection","volume":"20","author":"Li Jianan","year":"2017","unstructured":"Jianan Li , Xiaodan Liang , ShengMei Shen , Tingfa Xu , Jiashi Feng , and Shuicheng Yan . 2017 . Scale-aware fast R-CNN for pedestrian detection . IEEE Trans. Multimedia 20 , 4 (2017), 985 \u2013 996 . Jianan Li, Xiaodan Liang, ShengMei Shen, Tingfa Xu, Jiashi Feng, and Shuicheng Yan. 2017. Scale-aware fast R-CNN for pedestrian detection. IEEE Trans. Multimedia 20, 4 (2017), 985\u2013996.","journal-title":"IEEE Trans. Multimedia"},{"key":"e_1_2_1_25_1","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV\u201914)","author":"Lin Tsung-Yi","unstructured":"Tsung-Yi Lin , Michael Maire , Serge Belongie , James Hays , Pietro Perona , Deva Ramanan , Piotr Doll\u00e1r , and C. Lawrence Zitnick . 2014. Microsoft coco: Common objects in context . In Proceedings of the European Conference on Computer Vision (ECCV\u201914) . Springer, 740\u2013755. Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll\u00e1r, and C. Lawrence Zitnick. 2014. Microsoft coco: Common objects in context. In Proceedings of the European Conference on Computer Vision (ECCV\u201914). Springer, 740\u2013755."},{"key":"e_1_2_1_26_1","volume-title":"Proceedings of the British Machine Vision Conference (BMVC\u201916)","author":"Liu Jingjing","unstructured":"Jingjing Liu , Shaoting Zhang , Shu Wang , and Dimitris N. Metaxas . 2016. Multispectral deep neural networks for pedestrian detection . In Proceedings of the British Machine Vision Conference (BMVC\u201916) . BMVA Press. Jingjing Liu, Shaoting Zhang, Shu Wang, and Dimitris N. Metaxas. 2016. Multispectral deep neural networks for pedestrian detection. In Proceedings of the British Machine Vision Conference (BMVC\u201916). BMVA Press."},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00533"},{"key":"e_1_2_1_28_1","volume-title":"Proceedings of the International Conference on Machine Learning (ICML\u201915)","author":"Long Mingsheng","unstructured":"Mingsheng Long , Yue Cao , Jianmin Wang , and Michael I. Jordan . 2015. Learning transferable features with deep adaptation networks . In Proceedings of the International Conference on Machine Learning (ICML\u201915) . 97\u2013105. Mingsheng Long, Yue Cao, Jianmin Wang, and Michael I. Jordan. 2015. Learning transferable features with deep adaptation networks. In Proceedings of the International Conference on Machine Learning (ICML\u201915). 97\u2013105."},{"key":"e_1_2_1_29_1","volume-title":"245 million video surveillance cameras installed globally","author":"Markit IHS","year":"2014","unstructured":"IHS Markit . 2019. 245 million video surveillance cameras installed globally in 2014 . Retrieved May 5, 2019 from https:\/\/technology.ihs.com\/532501\/245-million-video-surveillance-cameras-installed-globally-in-2014. IHS Markit. 2019. 245 million video surveillance cameras installed globally in 2014. Retrieved May 5, 2019 from https:\/\/technology.ihs.com\/532501\/245-million-video-surveillance-cameras-installed-globally-in-2014."},{"key":"e_1_2_1_30_1","volume-title":"Proceedings of the IEEE International Conference on Computer Vision (ICCV\u201917)","author":"Masana Marc","unstructured":"Marc Masana , Joost van de Weijer, Luis Herranz, Andrew D. Bagdanov, and Jose M. Alvarez. 2017. Domain-adaptive deep network compression . In Proceedings of the IEEE International Conference on Computer Vision (ICCV\u201917) . 4289\u20134297. Marc Masana, Joost van de Weijer, Luis Herranz, Andrew D. Bagdanov, and Jose M. Alvarez. 2017. Domain-adaptive deep network compression. In Proceedings of the IEEE International Conference on Computer Vision (ICCV\u201917). 4289\u20134297."},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.sbspro.2010.01.038"},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-016-0890-9"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.690"},{"key":"e_1_2_1_34_1","unstructured":"Joseph Redmon and Ali Farhadi. 2018. YOLOv3: An incremental improvement. arxiv:1804.02767. Retrieved from http:\/\/arxiv.org\/abs\/1804.02767.  Joseph Redmon and Ali Farhadi. 2018. YOLOv3: An incremental improvement. arxiv:1804.02767. Retrieved from http:\/\/arxiv.org\/abs\/1804.02767."},{"key":"e_1_2_1_35_1","volume-title":"Proceedings of the IEEE International Conference on Computer Vision (ICCV\u201917)","author":"Selvaraju R. R.","unstructured":"R. R. Selvaraju , M. Cogswell , A. Das , R. Vedantam , D. Parikh , and D. Batra . 2017. Grad-CAM: Visual explanations from deep networks via gradient-based localization . In Proceedings of the IEEE International Conference on Computer Vision (ICCV\u201917) . 618\u2013626. R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, and D. Batra. 2017. Grad-CAM: Visual explanations from deep networks via gradient-based localization. In Proceedings of the IEEE International Conference on Computer Vision (ICCV\u201917). 618\u2013626."},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.3390\/computation7020020"},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7299143"},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-93000-8_47"},{"key":"e_1_2_1_39_1","volume-title":"Proceedings of the European Symposium on Artificial Neural Networks, Computational Intelligence and Machine Learning (ESANN\u201916)","author":"Wagner J\u00f6rg","year":"2016","unstructured":"J\u00f6rg Wagner , Volker Fischer , Michael Herman , and Sven Behnke . 2016 . Multispectral pedestrian detection using deep fusion convolutional neural networks . In Proceedings of the European Symposium on Artificial Neural Networks, Computational Intelligence and Machine Learning (ESANN\u201916) . 509\u2013514. J\u00f6rg Wagner, Volker Fischer, Michael Herman, and Sven Behnke. 2016. Multispectral pedestrian detection using deep fusion convolutional neural networks. In Proceedings of the European Symposium on Artificial Neural Networks, Computational Intelligence and Machine Learning (ESANN\u201916). 509\u2013514."},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.451"},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46475-6_28"},{"key":"e_1_2_1_42_1","unstructured":"Lu Zhang Zhiyong Liu Xiangyu Chen and Xu Yang. 2019. The cross-modality disparity problem in multispectral pedestrian detection. arxiv:1901.02645. Retrieved fromhttp:\/\/arxiv.org\/abs\/1901.02645.  Lu Zhang Zhiyong Liu Xiangyu Chen and Xu Yang. 2019. The cross-modality disparity problem in multispectral pedestrian detection. arxiv:1901.02645. Retrieved fromhttp:\/\/arxiv.org\/abs\/1901.02645."},{"key":"e_1_2_1_43_1","unstructured":"Yang Zheng Izzat H. Izzat and Shahrzad Ziaee. 2019. GFD-SSD: Gated fusion double SSD for multispectral pedestrian detection. arxiv:1903.06999. Retrieved fromhttp:\/\/arxiv.org\/abs\/1903.06999.  Yang Zheng Izzat H. Izzat and Shahrzad Ziaee. 2019. GFD-SSD: Gated fusion double SSD for multispectral pedestrian detection. arxiv:1903.06999. Retrieved fromhttp:\/\/arxiv.org\/abs\/1903.06999."}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3418213","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3418213","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T21:31:36Z","timestamp":1750195896000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3418213"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,2,28]]},"references-count":43,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2021,2,28]]}},"alternative-id":["10.1145\/3418213"],"URL":"https:\/\/doi.org\/10.1145\/3418213","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,2,28]]},"assertion":[{"value":"2019-12-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2020-07-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2021-04-16","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}