{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,18]],"date-time":"2026-07-18T16:19:45Z","timestamp":1784391585138,"version":"3.55.0"},"reference-count":59,"publisher":"MDPI AG","issue":"21","license":[{"start":{"date-parts":[[2022,11,7]],"date-time":"2022-11-07T00:00:00Z","timestamp":1667779200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"European project INFINITY","award":["883293"],"award-info":[{"award-number":["883293"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Object detection is a computer vision task that involves localisation and classification of objects in an image. Video data implicitly introduces several challenges, such as blur, occlusion and defocus, making video object detection more challenging in comparison to still image object detection, which is performed on individual and independent images. This paper tackles these challenges by proposing an attention-heavy framework for video object detection that aggregates the disentangled features extracted from individual frames. The proposed framework is a two-stage object detector based on the Faster R-CNN architecture. The disentanglement head integrates scale, spatial and task-aware attention and applies it to the features extracted by the backbone network across all the frames. Subsequently, the aggregation head incorporates temporal attention and improves detection in the target frame by aggregating the features of the support frames. These include the features extracted from the disentanglement network along with the temporal features. We evaluate the proposed framework using the ImageNet VID dataset and achieve a mean Average Precision (mAP) of 49.8 and 52.5 using the backbones of ResNet-50 and ResNet-101, respectively. The improvement in performance over the individual baseline methods validates the efficacy of the proposed approach.<\/jats:p>","DOI":"10.3390\/s22218583","type":"journal-article","created":{"date-parts":[[2022,11,8]],"date-time":"2022-11-08T08:17:12Z","timestamp":1667895432000},"page":"8583","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":7,"title":["Attention-Guided Disentangled Feature Aggregation for Video Object Detection"],"prefix":"10.3390","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-7942-4698","authenticated-orcid":false,"given":"Shishir","family":"Muralidhara","sequence":"first","affiliation":[{"name":"Department of Computer Science, Technical University of Kaiserslautern, 67663 Kaiserslautern, Germany"},{"name":"Mindgarage, Technical University of Kaiserslautern, 67663 Kaiserslautern, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0456-6493","authenticated-orcid":false,"given":"Khurram Azeem","family":"Hashmi","sequence":"additional","affiliation":[{"name":"Department of Computer Science, Technical University of Kaiserslautern, 67663 Kaiserslautern, Germany"},{"name":"Mindgarage, Technical University of Kaiserslautern, 67663 Kaiserslautern, Germany"},{"name":"German Research Institute for Artificial Intelligence (DFKI), 67663 Kaiserslautern, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Alain","family":"Pagani","sequence":"additional","affiliation":[{"name":"German Research Institute for Artificial Intelligence (DFKI), 67663 Kaiserslautern, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4029-6574","authenticated-orcid":false,"given":"Marcus","family":"Liwicki","sequence":"additional","affiliation":[{"name":"Department of Computer Science, Lule\u00e5 University of Technology, 971 87 Lule\u00e5, Sweden"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Didier","family":"Stricker","sequence":"additional","affiliation":[{"name":"Department of Computer Science, Technical University of Kaiserslautern, 67663 Kaiserslautern, Germany"},{"name":"German Research Institute for Artificial Intelligence (DFKI), 67663 Kaiserslautern, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0536-6867","authenticated-orcid":false,"given":"Muhammad Zeshan","family":"Afzal","sequence":"additional","affiliation":[{"name":"Department of Computer Science, Technical University of Kaiserslautern, 67663 Kaiserslautern, Germany"},{"name":"Mindgarage, Technical University of Kaiserslautern, 67663 Kaiserslautern, Germany"},{"name":"German Research Institute for Artificial Intelligence (DFKI), 67663 Kaiserslautern, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2022,11,7]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1561\/0600000079","article-title":"Computer vision for autonomous vehicles: Problems, datasets and state of the art","volume":"12","author":"Janai","year":"2020","journal-title":"Found. Trends Comput. Graph. Vis."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"3027","DOI":"10.1109\/TVT.2021.3065250","article-title":"Deep learning-based computer vision for surveillance in its: Evaluation of state-of-the-art methods","volume":"70","author":"Xie","year":"2021","journal-title":"IEEE Trans. Veh. Technol."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"69","DOI":"10.1016\/j.compag.2018.08.001","article-title":"Computer vision and artificial intelligence in precision agriculture for grain crops: A systematic review","volume":"153","author":"Rieder","year":"2018","journal-title":"Comput. Electron. Agric."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Gao, J., Yang, Y., Lin, P., and Park, D.S. (2018). Computer vision in healthcare applications. J. Healthc. Eng., 2018.","DOI":"10.1155\/2018\/5157020"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Liu, W., Anguelov, D., Erhan, D., Szegedy, C., Reed, S., Fu, C.Y., and Berg, A.C. (2016, January 11\u201314). Ssd: Single shot multibox detector. Proceedings of the European Conference on Computer Vision, Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46448-0_2"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Redmon, J., Divvala, S., Girshick, R., and Farhadi, A. (2016, January 27\u201330). You only look once: Unified, real-time object detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.91"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Lin, T.Y., Goyal, P., Girshick, R., He, K., and Doll\u00e1r, P. (2017, January 22\u201329). Focal loss for dense object detection. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.324"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Girshick, R., Donahue, J., Darrell, T., and Malik, J. (2014, January 23\u201328). Rich feature hierarchies for accurate object detection and semantic segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA.","DOI":"10.1109\/CVPR.2014.81"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Girshick, R. (2015, January 7\u201313). Fast r-cnn. Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.169"},{"key":"ref_10","unstructured":"Cortes, C., Lawrence, N., Lee, D., Sugiyama, M., and Garnett, R. (2015). Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks. Proceedings of the Advances in Neural Information Processing Systems, Montreal, QC, Canada, 7\u201312 December 2015, Curran Associates, Inc."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Sultana, F., Sufian, A., and Dutta, P. (2020). A review of object detection models based on convolutional neural network. Intelligent Computing: Image Processing Based Applications, Springer.","DOI":"10.1007\/978-981-15-4288-6_1"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Ahmed, M., Hashmi, K.A., Pagani, A., Liwicki, M., Stricker, D., and Afzal, M.Z. (2021). Survey and performance analysis of deep learning based object detection in challenging environments. Sensors, 21.","DOI":"10.20944\/preprints202106.0590.v1"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Hashmi, K.A., Pagani, A., Liwicki, M., Stricker, D., and Afzal, M.Z. (2022). Exploiting Concepts of Instance Segmentation to Boost Detection in Challenging Environments. Sensors, 22.","DOI":"10.20944\/preprints202204.0279.v1"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Jiao, L., Zhang, R., Liu, F., Yang, S., Hou, B., Li, L., and Tang, X. (2021). New Generation Deep Learning for Video Object Detection: A Survey. IEEE Transactions on Neural Networks and Learning Systems, IEEE.","DOI":"10.1109\/TNNLS.2021.3053249"},{"key":"ref_15","unstructured":"Wu, H., Chen, Y., Wang, N., and Zhang, Z. (November, January 27). Sequence Level Semantics Aggregation for Video Object Detection. Proceedings of the IEEE International Conference on Computer Vision, Seoul, Korea."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Gong, T., Chen, K., Wang, X., Chu, Q., Zhu, F., Lin, D., Yu, N., and Feng, H. (2021, January 2\u20139). Temporal ROI Align for Video Object Recognition. Proceedings of the AAAI Conference on Artificial Intelligence, Virtual.","DOI":"10.1609\/aaai.v35i2.16234"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Hashmi, K.A., Pagani, A., Stricker, D., and Afzal, M.Z. (2022). BoxMask: Revisiting Bounding Box Supervision for Video Object Detection. arXiv.","DOI":"10.1109\/WACV56688.2023.00207"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Dai, X., Chen, Y., Xiao, B., Chen, D., Liu, M., Yuan, L., and Zhang, L. (2021, January 20\u201325). Dynamic Head: Unifying Object Detection Heads With Attentions. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.00729"},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"886","DOI":"10.1109\/CVPR.2005.177","article-title":"Histograms of oriented gradients for human detection","volume":"Volume 1","author":"Dalal","year":"2005","journal-title":"Proceedings of the 2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR\u201905)"},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"1150","DOI":"10.1109\/ICCV.1999.790410","article-title":"Object recognition from local scale-invariant features","volume":"Volume 2","author":"Lowe","year":"1999","journal-title":"Proceedings of the Seventh IEEE International Conference on Computer Vision"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Bay, H., Tuytelaars, T., and Gool, L.V. (2006, January 7\u201313). Surf: Speeded up robust features. Proceedings of the European Conference on Computer Vision, Graz, Austria.","DOI":"10.1007\/11744023_32"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Calonder, M., Lepetit, V., Strecha, C., and Fua, P. (2010, January 5\u201311). Brief: Binary robust independent elementary features. Proceedings of the European Conference on Computer Vision, Heraklion, Greece.","DOI":"10.1007\/978-3-642-15561-1_56"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"O\u2019Mahony, N., Campbell, S., Carvalho, A., Harapanahalli, S., Hernandez, G.V., Krpalkova, L., Riordan, D., and Walsh, J. (2020). Deep Learning vs. Traditional Computer Vision. Advances in Computer Vision, Springer.","DOI":"10.1007\/978-3-030-17795-9_10"},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"3212","DOI":"10.1109\/TNNLS.2018.2876865","article-title":"Object detection with deep learning: A review","volume":"30","author":"Zhao","year":"2019","journal-title":"IEEE Trans. Neural Netw. Learn. Syst."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Zhai, M., Xiang, X., Lv, N., and Kong, X. (2021). Optical flow and scene flow estimation: A survey. Pattern Recognit., 114.","DOI":"10.1016\/j.patcog.2021.107861"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Dosovitskiy, A., Fischer, P., Ilg, E., Hausser, P., Hazirbas, C., Golkov, V., Smagt, P., Cremers, D., and Brox, T. (2015, January 7\u201313). FlowNet: Learning Optical Flow with Convolutional Networks. Proceedings of the 2015 IEEE International Conference on Computer Vision (ICCV), Santiago, Chile.","DOI":"10.1109\/ICCV.2015.316"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Zhu, X., Xiong, Y., Dai, J., Yuan, L., and Wei, Y. (2017, January 21\u201326). Deep feature flow for video recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.441"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Zhu, X., Wang, Y., Dai, J., Yuan, L., and Wei, Y. (2017, January 22\u201329). Flow-guided feature aggregation for video object detection. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.52"},{"key":"ref_29","unstructured":"Shi, X., Chen, Z., Wang, H., Yeung, D.Y., Wong, W.K., and Woo, W.c. (2015, January 7\u201312). Convolutional LSTM network: A machine learning approach for precipitation nowcasting. Proceedings of the 28th International Conference on Neural Information Processing Systems\u2014Volume 1, Montreal, QC, Canada."},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Lu, Y., Lu, C., and Tang, C.K. (2017, January 22\u201329). Online Video Object Detection Using Association LSTM. Proceedings of the 2017 IEEE International Conference on Computer Vision (ICCV), Venice, Italy.","DOI":"10.1109\/ICCV.2017.257"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Xiao, F., and Lee, Y.J. (2018, January 8\u201314). Video Object Detection with an Aligned Spatial-Temporal Memory. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01237-3_30"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Guo, M.H., Xu, T.X., Liu, J.J., Liu, Z.N., Jiang, P.T., Mu, T.J., Zhang, S.H., Martin, R.R., Cheng, M.M., and Hu, S.M. (2022). Attention mechanisms in computer vision: A survey. Computational Visual Media, Springer.","DOI":"10.1007\/s41095-022-0271-y"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Chen, Y., Cao, Y., Hu, H., and Wang, L. (2020, January 13\u201319). Memory Enhanced Global-Local Aggregation for Video Object Detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.01035"},{"key":"ref_34","unstructured":"Guo, C., Fan, B., Gu, J., Zhang, Q., Xiang, S., Prinet, V., and Pan, C. (November, January 27). Progressive sparse local attention for video object detection. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Seoul, Korea."},{"key":"ref_35","unstructured":"Mao, H., Kong, T., and Dally, B. (April, January 31). CaTDet: Cascaded Tracked Detector for Efficient Object Detection from Video. Proceedings of the Machine Learning and Systems, Stanford, CA, USA."},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Feichtenhofer, C., Pinz, A., and Zisserman, A. (2017, January 22\u201329). Detect to track and track to detect. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.330"},{"key":"ref_37","unstructured":"Dai, J., Li, Y., He, K., and Sun, J. (2016, January 5\u201310). R-fcn: Object detection via region-based fully convolutional networks. Proceedings of the 30th International Conference on Neural Information Processing Systems, Barcelona, Spain."},{"key":"ref_38","first-page":"1","article-title":"Object detection based on an adaptive attention mechanism","volume":"10","author":"Li","year":"2020","journal-title":"Sci. Rep."},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Hu, J., Shen, L., and Sun, G. (2018, January 18\u201323). Squeeze-and-excitation networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00745"},{"key":"ref_40","doi-asserted-by":"crossref","first-page":"94508","DOI":"10.1109\/ACCESS.2019.2928522","article-title":"Multi-Attention Object Detection Model in Remote Sensing Images Based on Multi-Scale","volume":"7","author":"Ying","year":"2019","journal-title":"IEEE Access"},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Carion, N., Massa, F., Synnaeve, G., Usunier, N., Kirillov, A., and Zagoruyko, S. (2020, January 23\u201328). End-to-end object detection with transformers. Proceedings of the European Conference on Computer Vision, Glasgow, UK.","DOI":"10.1007\/978-3-030-58452-8_13"},{"key":"ref_42","unstructured":"Zhu, X., Su, W., Lu, L., Li, B., Wang, X., and Dai, J. (2020). Deformable detr: Deformable transformers for end-to-end object detection. arXiv."},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Dai, X., Chen, Y., Yang, J., Zhang, P., Yuan, L., and Zhang, L. (2021, January 11\u201317). Dynamic DETR: End-to-End Object Detection With Dynamic Attention. Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.00298"},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"Wang, H., Wang, Z., Jia, M., Li, A., Feng, T., Zhang, W., and Jiao, L. (2019, January 17\u201328). Spatial attention for multi-scale feature refinement for object detection. Proceedings of the IEEE\/CVF International Conference on Computer Vision Workshops, Seoul, Korea.","DOI":"10.1109\/ICCVW.2019.00014"},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Huang, Z., Ke, W., and Huang, D. (2020, January 1\u20135). Improving object detection with inverted attention. Proceedings of the 2020 IEEE Winter Conference on Applications of Computer Vision (WACV), Snowmass Village, CO, USA.","DOI":"10.1109\/WACV45572.2020.9093507"},{"key":"ref_46","unstructured":"Chen, K., Wang, J., Pang, J., Cao, Y., Xiong, Y., Li, X., Sun, S., Feng, W., Liu, Z., and Xu, J. (2019). MMDetection: Open MMLab Detection Toolbox and Benchmark. arXiv."},{"key":"ref_47","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep Residual Learning for Image Recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_48","doi-asserted-by":"crossref","unstructured":"He, K., Gkioxari, G., Doll\u00e1r, P., and Girshick, R. (2017, January 22\u201329). Mask r-cnn. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.322"},{"key":"ref_49","doi-asserted-by":"crossref","first-page":"211","DOI":"10.1007\/s11263-015-0816-y","article-title":"ImageNet Large Scale Visual Recognition Challenge","volume":"115","author":"Russakovsky","year":"2015","journal-title":"Int. J. Comput. Vis."},{"key":"ref_50","unstructured":"Deng, J., Pan, Y., Yao, T., Zhou, W., Li, H., and Mei, T. (November, January 27). Relation distillation networks for video object detection. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Seoul, Korea."},{"key":"ref_51","unstructured":"Hetang, C., Qin, H., Liu, S., and Yan, J. (2017). Impression Network for Video Object Detection. arXiv."},{"key":"ref_52","doi-asserted-by":"crossref","unstructured":"Jiang, Z., Liu, Y., Yang, C., Liu, J., Gao, P., Zhang, Q., Xiang, S., and Pan, C. (2020, January 23\u201328). Learning where to focus for efficient video object detection. Proceedings of the European Conference on Computer Vision, Glasgow, UK.","DOI":"10.1007\/978-3-030-58517-4_2"},{"key":"ref_53","doi-asserted-by":"crossref","unstructured":"Wang, S., Zhou, Y., Yan, J., and Deng, Z. (2018, January 8\u201314). Fully Motion-Aware Network for Video Object Detection. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01261-8_33"},{"key":"ref_54","doi-asserted-by":"crossref","unstructured":"Zhu, X., Dai, J., Yuan, L., and Wei, Y. (2018, January 18\u201323). Towards high performance video object detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00753"},{"key":"ref_55","doi-asserted-by":"crossref","unstructured":"Bertasius, G., Torresani, L., and Shi, J. (2018, January 8\u201314). Object detection in video with spatiotemporal sampling networks. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01258-8_21"},{"key":"ref_56","unstructured":"Deng, H., Hua, Y., Song, T., Zhang, Z., Xue, Z., Ma, R., Robertson, N., and Guan, H. (November, January 27). Object guided external memory network for video object detection. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Seoul, Korea."},{"key":"ref_57","doi-asserted-by":"crossref","unstructured":"Shvets, M., Liu, W., and Berg, A. (November, January 27). Leveraging Long-Range Temporal Relationships Between Proposals for Video Object Detection. Proceedings of the 2019 IEEE\/CVF International Conference on Computer Vision (ICCV), Seoul, Korea.","DOI":"10.1109\/ICCV.2019.00985"},{"key":"ref_58","doi-asserted-by":"crossref","unstructured":"Vedaldi, A., Bischof, H., Brox, T., and Frahm, J.M. (2020). Mining Inter-Video Proposal Relations for Video Object Detection. Proceedings of the Computer Vision\u2014ECCV 2020, Virtual, 23\u201328 August 2020, Springer.","DOI":"10.1007\/978-3-030-58592-1"},{"key":"ref_59","doi-asserted-by":"crossref","unstructured":"Zhou, P., Zhou, C., Peng, P., Du, J., Sun, X., Guo, X., and Huang, F. (2020, January 22\u201326). NOH-NMS: Improving Pedestrian Detection by Nearby Objects Hallucination. Proceedings of the 28th ACM International Conference on Multimedia, Seoul, Korea.","DOI":"10.1145\/3394171.3413617"}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/22\/21\/8583\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T01:12:15Z","timestamp":1760145135000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/22\/21\/8583"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,11,7]]},"references-count":59,"journal-issue":{"issue":"21","published-online":{"date-parts":[[2022,11]]}},"alternative-id":["s22218583"],"URL":"https:\/\/doi.org\/10.3390\/s22218583","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,11,7]]}}}