{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,27]],"date-time":"2026-05-27T14:19:51Z","timestamp":1779891591864,"version":"3.53.1"},"reference-count":53,"publisher":"MDPI AG","issue":"6","license":[{"start":{"date-parts":[[2024,3,12]],"date-time":"2024-03-12T00:00:00Z","timestamp":1710201600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation","doi-asserted-by":"publisher","award":["61976177"],"award-info":[{"award-number":["61976177"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation","doi-asserted-by":"publisher","award":["U21A20524"],"award-info":[{"award-number":["U21A20524"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation","doi-asserted-by":"publisher","award":["2023-YBGY-222"],"award-info":[{"award-number":["2023-YBGY-222"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Key R&amp;D Project in Shaanxi Province of China","award":["61976177"],"award-info":[{"award-number":["61976177"]}]},{"name":"Key R&amp;D Project in Shaanxi Province of China","award":["U21A20524"],"award-info":[{"award-number":["U21A20524"]}]},{"name":"Key R&amp;D Project in Shaanxi Province of China","award":["2023-YBGY-222"],"award-info":[{"award-number":["2023-YBGY-222"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Crowd counting is an important task that serves as a preprocessing step in many applications. Despite obvious improvement reported by various convolutional-neural-network-based approaches, they only focus on the role of deep feature maps while neglecting the importance of shallow features for crowd counting. In order to surmount this issue, a dilated convolutional-neural-network-based cross-level contextual information extraction network is proposed in this work, which is abbreviated as CL-DCNN. Specifically, a dilated contextual module (DCM) is constructed by importing cross-level connection between different feature maps. It can effectively integrate contextual information while conserving the local details of crowd scenes. Extensive experiments show that the proposed approach outperforms state-of-the-art approaches using five public datasets, i.e., ShanghaiTech part A, ShanghaiTech part B, Mall, UCF_CC_50 and UCF-QNRF, achieving MAE 52.6, 8.1, 1.55, 181.8, and 96.4, respectively.<\/jats:p>","DOI":"10.3390\/s24061816","type":"journal-article","created":{"date-parts":[[2024,3,12]],"date-time":"2024-03-12T12:14:40Z","timestamp":1710245680000},"page":"1816","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":11,"title":["A Dilated Convolutional Neural Network for Cross-Layers of Contextual Information for Congested Crowd Counting"],"prefix":"10.3390","volume":"24","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-2475-7177","authenticated-orcid":false,"given":"Zhiqiang","family":"Zhao","sequence":"first","affiliation":[{"name":"The School of Computer Science and Engineering, Xi\u2019an University of Technology, Xi\u2019an 710048, China"},{"name":"The Shaanxi Key Laboratory of Network Computing and Security Technology, Xi\u2019an 710048, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Peihong","family":"Ma","sequence":"additional","affiliation":[{"name":"The School of Computer Science and Engineering, Xi\u2019an University of Technology, Xi\u2019an 710048, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Meng","family":"Jia","sequence":"additional","affiliation":[{"name":"The School of Computer Science and Engineering, Xi\u2019an University of Technology, Xi\u2019an 710048, China"},{"name":"The Shaanxi Key Laboratory of Network Computing and Security Technology, Xi\u2019an 710048, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xiaofan","family":"Wang","sequence":"additional","affiliation":[{"name":"The School of Computer Science and Engineering, Xi\u2019an University of Technology, Xi\u2019an 710048, China"},{"name":"The Shaanxi Key Laboratory of Network Computing and Security Technology, Xi\u2019an 710048, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xinhong","family":"Hei","sequence":"additional","affiliation":[{"name":"The School of Computer Science and Engineering, Xi\u2019an University of Technology, Xi\u2019an 710048, China"},{"name":"The Shaanxi Key Laboratory of Network Computing and Security Technology, Xi\u2019an 710048, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2024,3,12]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"645","DOI":"10.1109\/TMM.2017.2751966","article-title":"PROVID: Progressive and Multimodal Vehicle Reidentification for Large-Scale Urban Surveillance","volume":"20","author":"Liu","year":"2018","journal-title":"IEEE Trans. Multimed."},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Zhang, Y., Zhou, D., Chen, S., Gao, S., and Ma, Y. (2016, January 27\u201330). Single-image crowd counting via multi-column convolutional neural network. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.70"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Liu, W., Salzmann, M., and Fua, P. (2019, January 15\u201320). Context-aware crowd counting. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00524"},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"2818","DOI":"10.1007\/s10489-020-01688-2","article-title":"A hybrid model of convolutional neural networks and deep regression forests for crowd counting","volume":"50","author":"Ji","year":"2020","journal-title":"Appl. Intell."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"10445","DOI":"10.1109\/ACCESS.2022.3144607","article-title":"FSC-set: Counting, localization of football supporters crowd in the stadiums","volume":"10","author":"Elharrouss","year":"2022","journal-title":"IEEE Access"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Lin, H., Ma, Z., Ji, R., Wang, Y., and Hong, X. (2022, January 18\u201324). Boosting crowd counting via multifaceted attention. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.01901"},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"6821","DOI":"10.1109\/TCSVT.2022.3171235","article-title":"Lw-count: An effective lightweight encoding-decoding crowd counting network","volume":"32","author":"Liu","year":"2022","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Song, Q., Wang, C., Wang, Y., Tai, Y., Wang, C., Li, J., Wu, J., and Ma, J. (2021, January 2\u20139). To choose or to fuse? Scale selection for crowd counting. Proceedings of the AAAI Conference on Artificial Intelligence, Virtual Event.","DOI":"10.1609\/aaai.v35i3.16360"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Li, Y., Zhang, X., and Chen, D. (2018, January 18\u201323). Csrnet: Dilated convolutional neural networks for understanding the highly congested scenes. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00120"},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"91","DOI":"10.1016\/j.neucom.2019.03.065","article-title":"Atrous convolutions spatial pyramid network for crowd counting and density estimation","volume":"350","author":"Ma","year":"2019","journal-title":"Neurocomputing"},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"2633","DOI":"10.1109\/TMM.2021.3086709","article-title":"Crowd Counting via Perspective-Guided Fractional-Dilation Convolution","volume":"24","author":"Yan","year":"2022","journal-title":"IEEE Trans. Multimed."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"127444","DOI":"10.1109\/ACCESS.2021.3112174","article-title":"U-ASD Net: Supervised Crowd Counting Based on Semantic Segmentation and Adaptive Scenario Discovery","volume":"9","author":"Hafeezallah","year":"2021","journal-title":"IEEE Access"},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"136032","DOI":"10.1109\/ACCESS.2021.3115963","article-title":"SRNet: Scale-aware representation learning network for dense crowd counting","volume":"9","author":"Huang","year":"2021","journal-title":"IEEE Access"},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"1557","DOI":"10.1007\/s10115-021-01563-7","article-title":"Metro passengers counting and density estimation via dilated-transposed fully convolutional neural network","volume":"63","author":"Zhu","year":"2021","journal-title":"Knowl. Inf. Syst."},{"key":"ref_15","unstructured":"Zhang, C., Li, H., Wang, X., and Yang, X. (2015, January 7\u201312). Cross-scene crowd counting via deep convolutional neural networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA."},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"228","DOI":"10.1109\/TCSVT.2022.3187194","article-title":"Spatial-Temporal Graph Network for Video Crowd Counting","volume":"33","author":"Wu","year":"2023","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Idrees, H., Saleemi, I., Seibert, C., and Shah, M. (2013, January 23\u201328). Multi-source multi-scale counting in extremely dense crowd images. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Portland, OR, USA.","DOI":"10.1109\/CVPR.2013.329"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Idrees, H., Tayyab, M., Athrey, K., Zhang, D., Al-Maadeed, S., Rajpoot, N., and Shah, M. (2018, January 8\u201314). Composition loss for counting, density map estimation and localization in dense crowds. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01216-8_33"},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"323","DOI":"10.1109\/TIP.2019.2928634","article-title":"Ha-ccn: Hierarchical attention-based crowd counting network","volume":"29","author":"Sindagi","year":"2019","journal-title":"IEEE Trans. Image Process."},{"key":"ref_20","first-page":"2594","article-title":"JHU-CROWD++: Large-Scale Crowd Counting Dataset and A Benchmark Method","volume":"44","author":"Sindagi","year":"2020","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_21","unstructured":"Sindagi, V.A., and Patel, V.M. (September, January 29). CNN-based Cascaded Multi-task Learning of High-level Prior and Density Estimation for Crowd Counting. Proceedings of the IEEE International Conference on Advanced Video and Signal Based Surveillance, Lecce, Italy. Number 17287241."},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Sam, D.B., Surya, S., and Babu, R.V. (2017, January 21\u201326). Switching Convolutional Neural Network for Crowd Counting. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.429"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Sindagi, V.A., and Patel, V.M. (2017, January 22\u201329). Generating High-Quality Crowd Density Maps Using Contextual Pyramid CNNs. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.206"},{"key":"ref_24","unstructured":"Xiong, H., Lu, H., Liu, C., Liu, L., Cao, Z., and Shen, C. (November, January 27). From Open Set to Closed Set: Counting Objects by Spatial Divide-and-Conquer. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Seoul, Republic of Korea."},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"3486","DOI":"10.1109\/TCSVT.2019.2919139","article-title":"PCC Net: Perspective Crowd Counting via Spatial Convolutional Network","volume":"30","author":"Gao","year":"2020","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_26","first-page":"2739","article-title":"Locate, size, and count: Accurately resolving people in dense crowds via detection","volume":"43","author":"Sam","year":"2020","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Bai, S., He, Z., Qiao, Y., Hu, H., Wu, W., and Yan, J. (2020, January 13\u201319). Adaptive Dilated Network with Self-Correction Supervision for Counting. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00465"},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"3988","DOI":"10.1109\/TAES.2021.3087821","article-title":"Drone-SCNet: Scaled cascade network for crowd counting on drone images","volume":"57","author":"Elharrouss","year":"2021","journal-title":"IEEE Trans. Aerosp. Electron. Syst."},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"443","DOI":"10.1109\/TMM.2020.2980945","article-title":"Density-Aware Multi-Task Learning for Crowd Counting","volume":"23","author":"Jiang","year":"2021","journal-title":"IEEE Trans. Multimed."},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"1395","DOI":"10.1109\/TIP.2020.3043122","article-title":"Embedding Perspective Analysis Into Multi-Column Convolutional Neural Network for Crowd Counting","volume":"30","author":"Yang","year":"2021","journal-title":"IEEE Trans. Image Process."},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Lin, H., Hong, X., Ma, Z., Wei, X., Qiu, Y., Wang, Y., and Gong, Y. (2021, January 19\u201326). Direct Measure Matching for Crowd Counting. Proceedings of the International Joint Conferences on Artificial Intelligence Organization, Virtual Event.","DOI":"10.24963\/ijcai.2021\/116"},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"168","DOI":"10.1007\/s44196-021-00016-x","article-title":"A Deep-Fusion Network for Crowd Counting in High-Density Crowded Scenes","volume":"14","author":"Khan","year":"2021","journal-title":"Int. J. Comput. Intell. Syst."},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Shu, W., Wan, J., Tan, K.C., Kwong, S., and Chan, A.B. (2022, January 18\u201324). Crowd counting in the frequency domain. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.01900"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Chen, K., Loy, C.C., Gong, S., and Xiang, T. (2012, January 3\u20137). Feature mining for localised crowd counting. Proceedings of the British Machine Vision Conference, Guildford, UK.","DOI":"10.5244\/C.26.21"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Wang, Y., and Zou, Y. (2016, January 25\u201328). Fast visual object counting via example-based density estimation. Proceedings of the IEEE International Conference on Image Processing (ICIP), Phoenix, AZ, USA.","DOI":"10.1109\/ICIP.2016.7533041"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Xiong, F., Shi, X., and Yeung, D.Y. (2017, January 22\u201329). Spatiotemporal Modeling for Crowd Counting in Videos. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.551"},{"key":"ref_37","doi-asserted-by":"crossref","first-page":"1788","DOI":"10.1109\/TCSVT.2016.2637379","article-title":"Crowd Counting via Weighted VLAD on a Dense Attribute Feature Map","volume":"28","author":"Sheng","year":"2018","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Liu, L., Wang, H., Li, G., Ouyang, W., and Lin, L. (2018, January 9\u201319). Crowd Counting Using Deep Recurrent Spatial-Aware Network. Proceedings of the International Joint Conference on Artificial Intelligence, Stockholm, Sweden.","DOI":"10.24963\/ijcai.2018\/118"},{"key":"ref_39","doi-asserted-by":"crossref","first-page":"66215","DOI":"10.1109\/ACCESS.2019.2918936","article-title":"An Automatic Scale-Adaptive Approach with Attention Mechanism-Based Crowd Spatial Information for Crowd Counting","volume":"7","author":"Kong","year":"2019","journal-title":"IEEE Access"},{"key":"ref_40","doi-asserted-by":"crossref","first-page":"35317","DOI":"10.1109\/ACCESS.2019.2904712","article-title":"Crowd counting in low-resolution crowded scenes using region-based deep convolutional neural networks","volume":"7","author":"Saqib","year":"2019","journal-title":"IEEE Access"},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Fang, Y., Zhan, B., Cai, W., Gao, S., and Hu, B. (2019, January 8\u201312). Locality-Constrained Spatial Transformer Network for Video Crowd Counting. Proceedings of the IEEE International Conference on Multimedia and Expo, Shanghai, China.","DOI":"10.1109\/ICME.2019.00145"},{"key":"ref_42","doi-asserted-by":"crossref","first-page":"113","DOI":"10.1016\/j.patrec.2019.04.012","article-title":"ST-CNN: Spatial-Temporal Convolutional Neural Network for crowd counting in videos","volume":"125","author":"Miao","year":"2019","journal-title":"Pattern Recognit. Lett."},{"key":"ref_43","doi-asserted-by":"crossref","first-page":"98","DOI":"10.1016\/j.neucom.2020.01.087","article-title":"Multi-level feature fusion based Locality-Constrained Spatial Transformer network for video crowd counting","volume":"392","author":"Fang","year":"2020","journal-title":"Neurocomputing"},{"key":"ref_44","doi-asserted-by":"crossref","first-page":"13","DOI":"10.1016\/j.neucom.2020.04.071","article-title":"Fast video crowd counting with a Temporal Aware Network","volume":"403","author":"Wu","year":"2020","journal-title":"Neurocomputing"},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Han, T., Gao, J., Yuan, Y., and Wang, Q. (2020, January 4\u20138). Focus on Semantic Consistency for Cross-Domain Crowd Understanding. Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing, Barcelona, Spain.","DOI":"10.1109\/ICASSP40776.2020.9054768"},{"key":"ref_46","doi-asserted-by":"crossref","first-page":"5222","DOI":"10.1109\/TMM.2022.3189246","article-title":"Global Representation Guided Adaptive Fusion Network for Stable Video Crowd Counting","volume":"25","author":"Cai","year":"2022","journal-title":"IEEE Trans. Multimed."},{"key":"ref_47","doi-asserted-by":"crossref","unstructured":"Wang, Q., Gao, J., Lin, W., and Yuan, Y. (2019, January 15\u201320). Learning from Synthetic Data for Crowd Counting in the Wild. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00839"},{"key":"ref_48","doi-asserted-by":"crossref","first-page":"71576","DOI":"10.1109\/ACCESS.2019.2918650","article-title":"Scale Driven Convolutional Neural Network Model for People Counting and Localization in Crowd Scenes","volume":"7","author":"Basalamah","year":"2019","journal-title":"IEEE Access"},{"key":"ref_49","doi-asserted-by":"crossref","first-page":"3051","DOI":"10.1007\/s13369-020-04990-w","article-title":"Sparse to Dense Scale Prediction for Crowd Couting in High Density Crowds","volume":"46","author":"Khan","year":"2021","journal-title":"Arab. J. Sci. Eng."},{"key":"ref_50","doi-asserted-by":"crossref","first-page":"2127","DOI":"10.1007\/s00371-020-01974-7","article-title":"Scale and density invariant head detection deep model for crowd counting in pedestrian crowds","volume":"37","author":"Khan","year":"2021","journal-title":"Vis. Comput."},{"key":"ref_51","doi-asserted-by":"crossref","first-page":"1357","DOI":"10.1109\/TPAMI.2020.3022878","article-title":"Kernel-Based Density Map Generation for Dense Object Counting","volume":"44","author":"Wan","year":"2022","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_52","doi-asserted-by":"crossref","unstructured":"Wang, C.Y., Liao, H.Y.M., Wu, Y.H., Chen, P.Y., Hsieh, J.W., and Yeh, I.H. (2020, January 14\u201319). CSPNet: A new backbone that can enhance learning capability of CNN. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops, Seattle, WA, USA.","DOI":"10.1109\/CVPRW50498.2020.00203"},{"key":"ref_53","doi-asserted-by":"crossref","first-page":"160104","DOI":"10.1007\/s11432-021-3445-y","article-title":"Transcrowd: Weakly-supervised crowd counting with transformers","volume":"65","author":"Liang","year":"2022","journal-title":"Sci. China Inf. Sci."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/24\/6\/1816\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T14:12:19Z","timestamp":1760105539000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/24\/6\/1816"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,3,12]]},"references-count":53,"journal-issue":{"issue":"6","published-online":{"date-parts":[[2024,3]]}},"alternative-id":["s24061816"],"URL":"https:\/\/doi.org\/10.3390\/s24061816","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,3,12]]}}}