{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,22]],"date-time":"2026-03-22T06:38:02Z","timestamp":1774161482945,"version":"3.50.1"},"reference-count":34,"publisher":"MDPI AG","issue":"6","license":[{"start":{"date-parts":[[2019,3,18]],"date-time":"2019-03-18T00:00:00Z","timestamp":1552867200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/100007219","name":"Natural Science Foundation of Shanghai","doi-asserted-by":"publisher","award":["16ZR1413300"],"award-info":[{"award-number":["16ZR1413300"]}],"id":[{"id":"10.13039\/100007219","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["61802250"],"award-info":[{"award-number":["61802250"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Estimating the number of people in highly clustered crowd scenes is an extremely challenging task on account of serious occlusion and non-uniformity distribution in one crowd image. Traditional works on crowd counting take advantage of different CNN like networks to regress crowd density map, and further predict the count. In contrast, we investigate a simple but valid deep learning model that concentrates on accurately predicting the density map and simultaneously training a density level classifier to relax parameters of the network to prevent dangerous stampede with a smart camera. First, a combination of atrous and fractional stride convolutional neural network (CAFN) is proposed to deliver larger receptive fields and reduce the loss of details during down-sampling by using dilated kernels. Second, the expanded architecture is offered to not only precisely regress the density map, but also classify the density level of the crowd in the meantime (MTCAFN, multiple tasks CAFN for both regression and classification). Third, experimental results demonstrated on four datasets (Shanghai Tech A (MAE = 88.1) and B (MAE = 18.8), WorldExpo\u201910(average MAE = 8.2), NS UCF_CC_50(MAE = 303.2) prove our proposed method can deliver effective performance.<\/jats:p>","DOI":"10.3390\/s19061346","type":"journal-article","created":{"date-parts":[[2019,3,18]],"date-time":"2019-03-18T12:18:53Z","timestamp":1552911533000},"page":"1346","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":9,"title":["Smart Camera Aware Crowd Counting via Multiple Task Fractional Stride Deep Learning"],"prefix":"10.3390","volume":"19","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-2589-5755","authenticated-orcid":false,"given":"Minglei","family":"Tong","sequence":"first","affiliation":[{"name":"School of Electronics and Information Engineering, Shanghai University of Electric Power, Shanghai 200090, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Lyuyuan","family":"Fan","sequence":"additional","affiliation":[{"name":"School of Electronics and Information Engineering, Shanghai University of Electric Power, Shanghai 200090, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hao","family":"Nan","sequence":"additional","affiliation":[{"name":"School of Electronics and Information Engineering, Shanghai University of Electric Power, Shanghai 200090, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yan","family":"Zhao","sequence":"additional","affiliation":[{"name":"School of Electronics and Information Engineering, Shanghai University of Electric Power, Shanghai 200090, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2019,3,18]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Ryan, D., Denman, S., Fookes, C., and Sridharan, S. (2009). Crowd counting using multiple local features. Proc. Digit. Image Comput. Techn. Appl., 81\u201388.","DOI":"10.1109\/DICTA.2009.22"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Zeng, L., Xu, X., Cai, B., Qiu, S., and Zhang, T. (2017, January 17\u201320). Multi-scale convolutional neural networks for crowd counting. Proceedings of the IEEE International Conference on Image Processing, Beijing, China.","DOI":"10.1109\/ICIP.2017.8296324"},{"key":"ref_3","unstructured":"Zhang, C., Li, H., Wang, X., and Yang, X. (2015, January 8\u201310). Cross-scene crowd counting via deep convolutional neural networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA."},{"key":"ref_4","unstructured":"Marana, A.N., Costa, L.F., Lotufo, R.A., and Velastin, S.A. (1998, January 20\u201323). On the efficacy of texture analysis for crowd monitoring. Proceedings of the SIBGRAPI\u201998. International Symposium on Computer Graphics, Image Processing, and Vision (Cat. No.98EX237), Rio de Janeiro, Brazil."},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Topkaya, I.S., Erdogan, H., and Porikli, F. (2014, January 26\u201329). Counting people by clustering person detector outputs. Proceedings of the 11th IEEE International Conference on Advanced Video and Signal-Based Surveillance, Seoul, Korea.","DOI":"10.1109\/AVSS.2014.6918687"},{"key":"ref_6","unstructured":"Jingwei, G. (2013). Research and Implementation of People Counting in High Crowd Scenes of Tongji Bridge, Sun Yat-sen University."},{"key":"ref_7","first-page":"129","article-title":"Crowd Density Estimation Based on Pixel and Texture","volume":"28","author":"Qiang","year":"2015","journal-title":"Electron. Sci. Technol."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Chen, K., Loy, C.C., Gong, S., and Xiang, T. (2012, January 03\u201307). Feature mining for localised crowd counting. Proceedings of the 23rd British Machine Vision Conference, Surrey, Guildford, UK.","DOI":"10.5244\/C.26.21"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Rahmalan, H., Nixon, M.S., and Carter, J.N. (2006, January 13\u201314). On Crowd Density Estimation for Surveillance. Proceedings of the The Institution of Engineering and Technology Conference on Crime and Security, IET, London, UK.","DOI":"10.1049\/ic:20060360"},{"key":"ref_10","unstructured":"Lempitsky, V., and Zisserman, A. (2010, January 06\u201311). Learning to count objects in images. Proceedings of the 24th Neural Information Processing Systems, Vancouver, BC, Canada."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Pham, V., Kozakaya, T., Yamaguchi, O., and Okada, R. (2015, January 07\u201313). COUNT Forest: CO-Voting Uncertain Number of Targets Using Random Forest for Crowd Density Estimation. Proceedings of the 17th IEEE International Conference on Computer Vision (ICCV), Santiago, Chile.","DOI":"10.1109\/ICCV.2015.372"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Xu, B., and Qiu, G. (2016, January 7\u201310). Crowd density estimation based on rich features and random projection forest. Proceedings of the 21th IEEE Winter Conference on Applications of Computer Vision (WACV), Lake Placid, NY, USA.","DOI":"10.1109\/WACV.2016.7477682"},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"276","DOI":"10.1109\/TIP.2016.2624140","article-title":"Weakly Supervised Deep Matrix Factorization for Social Image Understanding","volume":"26","author":"Li","year":"2017","journal-title":"IEEE Trans. Image Process."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Li, Z., Tang, J., and Mei, T. (2018). Deep Collaborative Embedding for Social Image Understanding. IEEE Trans. Patt. Anal. Mach. Intel.","DOI":"10.1109\/TPAMI.2018.2852750"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Boominathan, L., Kruthiventi, S.S., and Babu, R.V. (2016, January 15\u201319). Crowdnet: A deep convolutional network for dense crowd counting. Proceedings of the 24th Proceedings of the ACM on Multimedia Conference, Amsterdam, The Netherlands.","DOI":"10.1145\/2964284.2967300"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Zhang, Y., Zhou, D., Chen, S., Gao, S., and Ma, Y. (2016, January 27\u201330). Single-image crowd counting via multi-column convolutional neural network. Proceedings of the 34th IEEE International Conference on Computer Vision and Pattern Recognition, Las Vegas, NE, USA.","DOI":"10.1109\/CVPR.2016.70"},{"key":"ref_17","unstructured":"Sindagi, V.A., and Patel, V.M. (September, January 29). CNN-Based cascaded multi-task learning of high-level prior and density estimation for crowd counting. Proceedings of the 14th IEEE International Conference on Advanced Video and Signal Based Surveillance, Lecce, Italy."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Idrees, H., Saleemi, I., Seibert, C., and Shah, M. (2013, January 23\u201328). Multi-source multi-scale counting in extremely dense crowd images. Proceedings of the 31st IEEE International Conference on Computer Vision and Pattern Recognition, Portland, OR, USA.","DOI":"10.1109\/CVPR.2013.329"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Wang, C., Zhang, H., Yang, L., Liu, S., and Cao, X. (2015, January 26\u201330). Deep People Counting in Extremely Dense Crowds. Proceedings of the 23rd International Conference ACM on Multimedia, Brisbane, Australia.","DOI":"10.1145\/2733373.2806337"},{"key":"ref_20","first-page":"615","article-title":"Towards Perspective-Free Object Counting with Deep Learning","volume":"Volume 911","author":"Leibe","year":"2016","journal-title":"Computer Vision\u2014ECCV 2016. Lecture Notes in Computer Science"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Shang, C., Ai, H., and Bai, B. (2016, January 25\u201328). End-to-end crowd counting via joint learning local and global count. Proceedings of the 23rd IEEE International Conference on Image Processing (ICIP), Phoenix, AZ, USA.","DOI":"10.1109\/ICIP.2016.7532551"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Sam, D.B., Surya, S., and Babu, R.V. (2017, January 21\u201326). Switching convolutional neural network for crowd counting. Proceedings of the 35th IEEE International Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.429"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Sindagi, V.A., and Patel, V.M. (2017, January 22\u201329). Generating high-quality crowd density maps using contextual pyramid cnns. Proceedings of the 19th IEEE International Conference on Computer Vision (ICCV), Venice, Italy.","DOI":"10.1109\/ICCV.2017.206"},{"key":"ref_24","unstructured":"Simonyan, K., and Zisserman, A. (2015, January 7\u20139). Very deep convolutional networks for large-scale image recognition. Proceedings of the 6th International Conference on Learning Representations, San Diego, CA, USA."},{"key":"ref_25","unstructured":"Yuhong, L., Xiaofan, Z., and Deming, C. (2018, January 18\u201323). CSRNet: Dilated Convolutional Neural Networks for Understanding the Highly Congested Scenes. Proceedings of the 36th IEEE International Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA."},{"key":"ref_26","unstructured":"Fisher, Y., and Vladlen, K. (arXiv, 2015). Multi-Scale Context Aggregation by Dilated Convolutions, arXiv."},{"key":"ref_27","unstructured":"Chen, L.C., Papandreou, G., Schroff, F., and Adam, H. (arXiv, 2017). Rethinking Atrous Convolution for Semantic Image Segmentation, arXiv."},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Eigen, D., and Fergus, R. (2015, January 7\u201313). Predicting depth, surface normals and semantic labels with a common multi-scale convolutional architecture. Proceedings of the 2015 IEEE International Conference on Computer Vision (ICCV), Santiago, Chile.","DOI":"10.1109\/ICCV.2015.304"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Kokkinos, I. (2017, January 21\u201326). UberNet: Training aniversal\u2019convolutional neural network for low-, mid-, and high-level vision using diverse datasets and limited memory. Proceedings of the 35th IEEE International Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.579"},{"key":"ref_30","unstructured":"Gkioxari, G., Hariharan, B., Girshick, R., and Malik, J. (arXiv, 2014). R-CNNS for pose estimation and action detection, arXiv."},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"121","DOI":"10.1109\/TPAMI.2017.2781233","article-title":"Hyperface: A deep multi-task learning framework for face detection, landmark localization, pose estimation, and gender recognition","volume":"41","author":"Ranjan","year":"2019","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_32","unstructured":"Dumoulin, V., and Visin, F. (arXiv, 2016). A guide to convolution arithmetic for deep learning, arXiv."},{"key":"ref_33","unstructured":"Collobert, R., Kavukcuoglu, K., and Farabet, C. Torch7: A Matlab-like Environment for Machine Learning. Proceedings of the 26th Annual Conference on Neural Information Processing Systems (NIPS)."},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Marsden, M., Mcguinness, K., Little, S., and O\u2019Connor, N.E. (arXiv, 2016). Fully Convolutional Crowd Counting on Highly Congested Scenes, arXiv.","DOI":"10.5220\/0006097300270033"}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/19\/6\/1346\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T12:38:44Z","timestamp":1760186324000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/19\/6\/1346"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019,3,18]]},"references-count":34,"journal-issue":{"issue":"6","published-online":{"date-parts":[[2019,3]]}},"alternative-id":["s19061346"],"URL":"https:\/\/doi.org\/10.3390\/s19061346","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2019,3,18]]}}}