{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,18]],"date-time":"2026-02-18T23:51:30Z","timestamp":1771458690454,"version":"3.50.1"},"reference-count":41,"publisher":"MDPI AG","issue":"5","license":[{"start":{"date-parts":[[2021,3,5]],"date-time":"2021-03-05T00:00:00Z","timestamp":1614902400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["61702032,61573057 and 61771042"],"award-info":[{"award-number":["61702032,61573057 and 61771042"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100012226","name":"Fundamental Research Funds for the Central Universities","doi-asserted-by":"publisher","award":["2017JBZ002"],"award-info":[{"award-number":["2017JBZ002"]}],"id":[{"id":"10.13039\/501100012226","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Fund","award":["61404130316, 61400010302"],"award-info":[{"award-number":["61404130316, 61400010302"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>The existing pedestrian detection algorithms cannot effectively extract features of heavily occluded targets which results in lower detection accuracy. To solve the heavy occlusion in crowds, we propose a multi-scale feature pyramid network based on ResNet (MFPN) to enhance the features of occluded targets and improve the detection accuracy. MFPN includes two modules, namely double feature pyramid network (FPN) integrated with ResNet (DFR) and repulsion loss of minimum (RLM). We propose the double FPN which improves the architecture to further enhance the semantic information and contours of occluded pedestrians, and provide a new way for feature extraction of occluded targets. The features extracted by our network can be more separated and clearer, especially those heavily occluded pedestrians. Repulsion loss is introduced to improve the loss function which can keep predicted boxes away from the ground truths of the unrelated targets. Experiments carried out on the public CrowdHuman dataset, we obtain 90.96% AP which yields the best performance, 5.16% AP gains compared to the FPN-ResNet50 baseline. Compared with the state-of-the-art works, the performance of the pedestrian detection system has been boosted with our method.<\/jats:p>","DOI":"10.3390\/s21051820","type":"journal-article","created":{"date-parts":[[2021,3,5]],"date-time":"2021-03-05T11:46:09Z","timestamp":1614944769000},"page":"1820","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":19,"title":["Multi-Scale Feature Pyramid Network: A Heavily Occluded Pedestrian Detection Network Based on ResNet"],"prefix":"10.3390","volume":"21","author":[{"given":"Xiaotao","family":"Shao","sequence":"first","affiliation":[{"name":"School of Electronic and Information Engineering, Beijing Jiaotong University, Beijing 100044, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Qing","family":"Wang","sequence":"additional","affiliation":[{"name":"School of Electronic and Information Engineering, Beijing Jiaotong University, Beijing 100044, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Wei","family":"Yang","sequence":"additional","affiliation":[{"name":"School of Electronic and Information Engineering, Beijing Jiaotong University, Beijing 100044, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yun","family":"Chen","sequence":"additional","affiliation":[{"name":"Shanghai Aerospace Control Technology Institute, Shanghai 201109, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yi","family":"Xie","sequence":"additional","affiliation":[{"name":"Beijing Xinghang Mechanical-Electrical Equipment Co., Ltd., Beijing 100074, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9287-1206","authenticated-orcid":false,"given":"Yan","family":"Shen","sequence":"additional","affiliation":[{"name":"School of Electronic and Information Engineering, Beijing Jiaotong University, Beijing 100044, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zhongli","family":"Wang","sequence":"additional","affiliation":[{"name":"School of Electronic and Information Engineering, Beijing Jiaotong University, Beijing 100044, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2021,3,5]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Shen, Y., Zhang, L., Wang, Z.L., Hao, X.L., and Hou, Y.L. (2019, January 22\u201325). Multi-Level Residual Up-Projection Activation Network for Image SuperResolution. Proceedings of the 2019 IEEE International Conference on Image Processing (ICIP), Taipei, Taiwan.","DOI":"10.1109\/ICIP.2019.8803331"},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"900","DOI":"10.1109\/TITS.2019.2901817","article-title":"Autonomous vehicles that interact with pedestrians: A survey of theory and practice","volume":"21","author":"Rasouli","year":"2020","journal-title":"IEEE Trans. Intell. Transp. Syst."},{"key":"ref_3","unstructured":"Shao, S., Zhao, Z.J., Li, B.X., Xiao, T.T., Yu, G., Zhang, X.Y., and Sun, J. (2018). Crowdhuman: A benchmark for detecting humans in a crowd. arXiv, preprint."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Vimal, S.P., Ajay, B., and Thiruvikiraman, P.K. (2013, January 22\u201323). Context pruned histogram of oriented gradients for pedestrian detection. Proceedings of the 2013 International Mutli-Conference on Automation, Computing, Communication, Control and Compressed Sensing (iMac4s),  Kottayam, India.","DOI":"10.1109\/iMac4s.2013.6526501"},{"key":"ref_5","unstructured":"Zhuang, J. (2016, January 14\u201317). Compressive tracking based on HOG and extended Haar-like feature. Proceedings of the 2016 2nd IEEE International Conference on Computer and Communications (ICCC), Chengdu, China."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Cosma, C., Brehar, R., and Nedevschi, S. (2013, January 5\u20137). Pedestrians detection using a cascade of LBP and HOG classifiers. Proceedings of the 2013 IEEE 9th International Conference on Intelligent Computer Communication and Processing (ICCP), Cluj-Napoca, Romania.","DOI":"10.1109\/ICCP.2013.6646084"},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"3540","DOI":"10.1109\/TITS.2017.2726140","article-title":"Pedestrian Movement Direction Recognition Using Convolutional Neural Networks","volume":"18","author":"Cazorla","year":"2017","journal-title":"IEEE Trans. Intell. Transp. Syst."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Zhao, J.X., Li, J., and Ma, Y.D. (2018, January 3\u20135). RPN+ fast boosted tree: Combining deep neural network with traditional classifier for pedestrian detection. Proceedings of the 2018 4th International Conference on Computer and Technology Applications (ICCTA), Istanbul, Turkey.","DOI":"10.1109\/CATA.2018.8398672"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Zhang, Z.S., Gao, J.Y., Mao, J.H., and Liu, Y.K. (2020, January 13\u201319). STINet: Spatio-Temporal-Interactive Network for Pedestrian Detection and Trajectory Prediction. Proceedings of the 2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.01136"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Ferrari, V., Hebert, M., Sminchisescu, C., and Weiss, Y. (2018, January 8\u201314). Occlusion-Aware R-CNN: Detecting Pedestrians in a Crowd. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01234-2"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Wang, X.L., Xiao, T.T., Jiang, Y.N., Shao, S., Sun, J., and Shen, C.H. (2018, January 18\u201323). Repulsion Loss: Detecting Pedestrians in a Crowd. Proceedings of the 2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00811"},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"849","DOI":"10.1109\/ACCESS.2020.3046498","article-title":"Feature Enhancement Based on CycleGAN for Nighttime Vehicle Detection","volume":"9","author":"Shao","year":"2021","journal-title":"IEEE Access."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Ke, W., Zhang, T.L., Huang, Z.Y., Ye, Q.X., Liu, J.Z., and Huang, D. (2020, January 13\u201319). Multiple Anchor Learning for Visual Object Detection. Proceedings of the 2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.01022"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Chen, Y.H., Cao, Y., Hu, H., and Wang, L.W. (2020, January 13\u201319). Memory Enhanced Global-Local Aggregation for Video Object Detection. Proceedings of the 2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.01035"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Li, Y., Wang, T., Kang, B.Y., Tang, S., Wang, C.F., Li, J.T., and Feng, J.S. (2020, January 13\u201319). Overcoming Classifier Imbalance for Long-Tail Object Detection with Balanced Group Softmax. Proceedings of the 2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.01100"},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"12955","DOI":"10.1109\/ACCESS.2021.3052241","article-title":"Cross-View Image Translation Based on Local and Global Information Guidance","volume":"9","author":"Shen","year":"2021","journal-title":"IEEE Access"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Tan, M.X., Pang, R.M., and Le, Q.V. (2020, January 13\u201319). EfficientDet: Scalable and Efficient Object Detection. Proceedings of the 2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.01079"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Redmon, J., Divvala, S., Girshick, R., and Farhadi, A. (2016, January 27\u201330). You Only Look Once: Unified, Real-Time Object Detection. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.91"},{"key":"ref_19","unstructured":"Leibe, B., Matas, J., Sebe, N., and Welling, M. (2016, January 11\u201314). SSD: Single Shot MultiBox Detector. Proceedings of the 14th European Conference on Computer Vision, Amsterdam, The Netherlands."},{"key":"ref_20","unstructured":"Bochkovskiy, A., Wan, C.Y., and Liao, H.-Y.M. (2020). YOLOv4: Optimal Speed and Accuracy of Object Detection. arXiv, preprint."},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Zhu, C.C., He, Y.H., and Savvides, M. (2019, January 15\u201320). Feature Selective Anchor-Free Module for Single-Shot Object Detection. Proceedings of the 2019 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00093"},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"318","DOI":"10.1109\/TPAMI.2018.2858826","article-title":"Focal Loss for Dense Object Detection","volume":"42","author":"Lin","year":"2017","journal-title":"IEEE Trans Pattern Anal. Mach. Intell."},{"key":"ref_23","unstructured":"Tan, M.X., and Le, Q.V. (2019, January 9\u201315). EfficientNet: Rethinking model scaling for convolutional neural networks. Proceedings of the 36th International Conference on Machine Learning (ICML), Long Beach, CA, USA."},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Girshick, R. (2015, January 7\u201313). Fast R-CNN. Proceedings of the 2015 IEEE International Conference on Computer Vision (ICCV), Santiago, Chile.","DOI":"10.1109\/ICCV.2015.169"},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"1137","DOI":"10.1109\/TPAMI.2016.2577031","article-title":"Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks","volume":"39","author":"Ren","year":"2017","journal-title":"IEEE Trans Pattern Anal. Mach. Intell."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"He, K.M., Gkioxari, G., Doll\u00e1r, P., and Girshick, R. (2017, January 22\u201329). Mask R-CNN. Proceedings of the 2017 IEEE International Conference on Computer Vision (ICCV), Venice, Italy.","DOI":"10.1109\/ICCV.2017.322"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Cai, Z.W., and Vasconcelos, N. (2019). Cascade R-CNN: High Quality Object Detection and Instance Segmentation. IEEE Trans Pattern Anal. Mach. Intell.","DOI":"10.1109\/CVPR.2018.00644"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Bodla, N., Singh, B., Chellappa, R., and Davis, L.S. (2017, January 22\u201329). Soft-NMS-Improving Object Detection with One Line of Code. Proceedings of the 2017 IEEE International Conference on Computer Vision (ICCV), Venice, Italy.","DOI":"10.1109\/ICCV.2017.593"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"He, Y.H., Zhu, C.C., Wang, J.R., Savvides, M., and Zhang, X.Y. (2019, January 16\u201320). Bounding Box Regression with Uncertainty for Accurate Object Detection. Proceedings of the 32nd IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00300"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Liu, S.T., Huang, D., and Wang, Y.H. (2019, January 15\u201320). Adaptive NMS: Refining Pedestrian Detection in a Crowd. Proceedings of the 2019 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00662"},{"key":"ref_31","unstructured":"Ferrari, V., Hebert, M., Sminchisescu, C., and Weiss, Y. (2018, January 8\u201314). Bi-box Regression for Pedestrian Detection and Occlusion Estimation. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany."},{"key":"ref_32","unstructured":"Chi, C., Zhang, S.F., Xing, J.L., Lei, Z., Li, S.Z., and Zou, X.D. (2020, January 7\u201312). PedHunter: Occlusion Robust Pedestrian Detector in Crowded Scenes. Proceedings of the 2020 AAAI Conference on Artificial Intelligence, New York, NY, USA."},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Pang, Y.W., Xie, J., Khan, M.H., Anwer, R.M., Khan, F.S., and Shao, L. (November, January 27). Mask-Guided Attention Network for Occluded Pedestrian Detection. Proceedings of the 2019 IEEE\/CVF International Conference on Computer Vision (ICCV), Seoul, Korea.","DOI":"10.1109\/ICCV.2019.00507"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Wang, A.T., Sun, Y.H., Kortylewski, A., and Yuille, A. (2020, January 13\u201319). Robust Object Detection under Occlusion with Context-Aware CompositionalNets. Proceedings of the 2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.01266"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Wu, J.L., Zhou, C.L., Yang, M., Zhang, Q., Li, Y., and Yuan, J.S. (2020, January 13\u201319). Temporal-Context Enhanced Detection of Heavily Occluded Pedestrians. Proceedings of the 2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.01344"},{"key":"ref_36","doi-asserted-by":"crossref","first-page":"3143","DOI":"10.1109\/TIP.2019.2957927","article-title":"Taking a Look at Small-Scale Pedestrians and Occluded Pedestrians","volume":"29","author":"Cao","year":"2020","journal-title":"IEEE Trans. Image Process."},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Lin, T.Y., Doll\u00e1r, P., Girshick, R., He, K.M., Hariharan, B., and Belongie, S. (2017, January 21\u201326). Feature Pyramid Networks for Object Detection. Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.106"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Liu, S., Qi, L., Qin, H.F., Shi, J.P., and Jia, J.Y. (2018, January 18\u201323). Path Aggregation Network for Instance Segmentation. Proceedings of the 2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00913"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"He, K.M., Zhang, X.Y., Ren, S.Q., and Sun, J. (2016, January 27\u201330). Deep Residual Learning for Image Recognition. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_40","unstructured":"Sifre, L. (2014). Rigid-Motion Scattering for Image Classification. [Ph.D. Thesis, Ecole Polytechnique]."},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Chu, X.G., Zheng, A.L., Zhang, X.Y., and Sun, J. (2020, January 13\u201319). Detection in Crowded Scenes: One Proposal, Multiple Predictions. Proceedings of the 2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.01223"}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/21\/5\/1820\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T05:33:33Z","timestamp":1760160813000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/21\/5\/1820"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,3,5]]},"references-count":41,"journal-issue":{"issue":"5","published-online":{"date-parts":[[2021,3]]}},"alternative-id":["s21051820"],"URL":"https:\/\/doi.org\/10.3390\/s21051820","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,3,5]]}}}