{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,17]],"date-time":"2026-02-17T15:00:56Z","timestamp":1771340456016,"version":"3.50.1"},"reference-count":35,"publisher":"Walter de Gruyter GmbH","issue":"1","license":[{"start":{"date-parts":[[2020,7,8]],"date-time":"2020-07-08T00:00:00Z","timestamp":1594166400000},"content-version":"unspecified","delay-in-days":0,"URL":"http:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2020,7,8]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>The purpose of crowd counting is to estimate the number of pedestrians in crowd images. Crowd counting or density estimation is an extremely challenging task in computer vision, due to large scale variations and dense scene. Current methods solve these issues by compounding multi-scale Convolutional Neural Network with different receptive fields. In this paper, a novel end-to-end architecture based on Multi-Scale Adversarial Convolutional Neural Network (MSA-CNN) is proposed to generate crowd density and estimate the amount of crowd. Firstly, a multi-scale network is used to extract the globally relevant features in the crowd image, and then fractionally-strided convolutional layers are designed for up-sampling the output to recover the loss of crucial details caused by the earlier max pooling layers. An adversarial loss is directly employed to shrink the estimated value into the realistic subspace to reduce the blurring effect of density estimation. Joint training is performed in an end-to-end fashion using a combination of Adversarial loss and Euclidean loss. The two losses are integrated via a joint training scheme to improve density estimation performance.We conduct some extensive experiments on available datasets to show the significant improvements and supremacy of the proposed approach over the available state-of-the-art approaches.<\/jats:p>","DOI":"10.1515\/jisys-2019-0157","type":"journal-article","created":{"date-parts":[[2020,7,10]],"date-time":"2020-07-10T08:47:24Z","timestamp":1594370844000},"page":"180-191","source":"Crossref","is-referenced-by-count":7,"title":["Crowd counting via Multi-Scale Adversarial Convolutional Neural Networks"],"prefix":"10.1515","volume":"30","author":[{"given":"Liping","family":"Zhu","sequence":"first","affiliation":[{"name":"Beijing Key Lab of Petroleum Data Mining, China University of Petroleum , Beijing , 10224 , China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hong","family":"Zhang","sequence":"additional","affiliation":[{"name":"Beijing Key Lab of Petroleum Data Mining, China University of Petroleum , Beijing , 10224 , China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Sikandar","family":"Ali","sequence":"additional","affiliation":[{"name":"Beijing Key Lab of Petroleum Data Mining, China University of Petroleum , Beijing , 10224 , China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Baoli","family":"Yang","sequence":"additional","affiliation":[{"name":"Beijing Key Lab of Petroleum Data Mining, China University of Petroleum , Beijing , 10224 , China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Chengyang","family":"Li","sequence":"additional","affiliation":[{"name":"Beijing Key Lab of Petroleum Data Mining, China University of Petroleum , Beijing , 10224 , China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"374","published-online":{"date-parts":[[2020,7,8]]},"reference":[{"key":"2025120523494748263_j_jisys-2019-0157_ref_001","doi-asserted-by":"crossref","unstructured":"B.B. Zhan, D.N. Monekosso, P. Remagnino, S.A. Velastin, and L.Q. Xu. Crowd analysis: a survey. Machine Vision and Applications 19(5):345\u2013357, 2008.","DOI":"10.1007\/s00138-008-0132-4"},{"key":"2025120523494748263_j_jisys-2019-0157_ref_002","doi-asserted-by":"crossref","unstructured":"J. Shao, K. Kang, C.C. Loy, and X.G. Wang. Deeply learned attributes for crowded scene understanding. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR) pages 4657\u20134666, 2015.","DOI":"10.1109\/CVPR.2015.7299097"},{"key":"2025120523494748263_j_jisys-2019-0157_ref_003","doi-asserted-by":"crossref","unstructured":"K. Chen, C.L. Chen, S.G. Gong, and T. Xiang. Feature mining for localised crowd counting. In British Machine Vision Conference 2012.","DOI":"10.5244\/C.26.21"},{"key":"2025120523494748263_j_jisys-2019-0157_ref_004","unstructured":"V. Lempitsky and A. Zisserman. Learning to count objects in images. In Neural Information Processing Systems pages 1324\u20131332, 2010."},{"key":"2025120523494748263_j_jisys-2019-0157_ref_005","doi-asserted-by":"crossref","unstructured":"E. Walach and L. Wolf. Learning to count with cnn boosting. In European Conference on Computer Vision Computer Vision '1C ECCV 2016, pages 660\u2013676. Springer International Publishing, 2016.","DOI":"10.1007\/978-3-319-46475-6_41"},{"key":"2025120523494748263_j_jisys-2019-0157_ref_006","doi-asserted-by":"crossref","unstructured":"Y. Wang and Y. Zou. Fast visual object counting via example-based density estimation. In IEEE International Conference on Image Processing (ICIP) pages 3653\u20133657, 2016.","DOI":"10.1109\/ICIP.2016.7533041"},{"key":"2025120523494748263_j_jisys-2019-0157_ref_007","doi-asserted-by":"crossref","unstructured":"M.R. Hsieh, Y.L. Lin, and W.H. Hsu. Drone-based object counting by spatially regularized regional proposal network. In IEEE International Conference on Computer Vision pages 4165\u20134173, 2017.","DOI":"10.1109\/ICCV.2017.446"},{"key":"2025120523494748263_j_jisys-2019-0157_ref_008","doi-asserted-by":"crossref","unstructured":"R.D. O\u00f1oro and S.R.J. L\u00f3pez. Towards perspective-free object counting with deep learning. In Bastian Leibe, Jiri Matas, Nicu  Sebe, and Max Welling, editors, Computer Vision '1C ECCV 2016 pages 615\u2013629. Springer International Publishing, 2016.","DOI":"10.1007\/978-3-319-46478-7_38"},{"key":"2025120523494748263_j_jisys-2019-0157_ref_009","doi-asserted-by":"crossref","unstructured":"E. Toropov, L.Y. Gui, S.H. Zhang, and S. Kottur. Traflc flow from a low frame rate city camera. In IEEE International Conference on Image Processing pages 3802\u20133806, 2015.","DOI":"10.1109\/ICIP.2015.7351516"},{"key":"2025120523494748263_j_jisys-2019-0157_ref_010","doi-asserted-by":"crossref","unstructured":"S.H. Zhang, G.H. Wu, J.P. Costeira, and J.M.F. Moura. Fcn-rlstm: Deep spatio-temporal neural networks for vehicle counting in city cameras. In IEEE International Conference on Computer Vision pages 3687\u20133696, 2017.","DOI":"10.1109\/ICCV.2017.396"},{"key":"2025120523494748263_j_jisys-2019-0157_ref_011","doi-asserted-by":"crossref","unstructured":"S.H. Zhang, G.H. Wu, J.P. Costeira, and J.M.F. Moura. Understanding traflc density from large-scale web camera data. In IEEE Computer Vision and Pattern Recognition 2017.","DOI":"10.1109\/CVPR.2017.454"},{"key":"2025120523494748263_j_jisys-2019-0157_ref_012","doi-asserted-by":"crossref","unstructured":"H. Idrees, I. Saleemi, C. Seibert, and M. Shah. Multi-source multi-scale counting in extremely dense crowd images. In Computer Vision and Pattern Recognition pages 2547\u20132554, 2013.","DOI":"10.1109\/CVPR.2013.329"},{"key":"2025120523494748263_j_jisys-2019-0157_ref_013","unstructured":"A. Bansal and K.S. Venkatesh. People counting in high density crowds from still images. Computer Science 2015."},{"key":"2025120523494748263_j_jisys-2019-0157_ref_014","doi-asserted-by":"crossref","unstructured":"L. Boominathan, S.S.S. Kruthiventi, and V.R. Babu. Crowdnet: A deep convolutional network for dense crowd counting. In ACM on Multimedia Conference pages 640\u2013644, 2016.","DOI":"10.1145\/2964284.2967300"},{"key":"2025120523494748263_j_jisys-2019-0157_ref_015","unstructured":"C. Zhang, H.S. Li, X.G. Wang, and X.K. Yang. Cross-scene crowd counting via deep convolutional neural networks. In IEEE Conference on Computer Vision and Pattern Recognition pages 833\u2013841, 2015."},{"key":"2025120523494748263_j_jisys-2019-0157_ref_016","doi-asserted-by":"crossref","unstructured":"N. Paragios and V. Ramesh. A mrf-based approach for real-time subway monitoring. In Proceedings of the 2001 IEEE Computer Society Conference on Computer Vision and Pattern Recognition volume 1, pages I\u2013I, 2001.","DOI":"10.1109\/CVPR.2001.990644"},{"key":"2025120523494748263_j_jisys-2019-0157_ref_017","unstructured":"K. Tota and H. Idrees. Counting in dense crowds using deep features. Center for Research in Computer Vision 2015."},{"key":"2025120523494748263_j_jisys-2019-0157_ref_018","doi-asserted-by":"crossref","unstructured":"D.B. Sam, S. Surya, and R.V. Babu. Switching convolutional neural network for crowd counting. In IEEE Conference on Computer Vision and Pattern Recognition 2017.","DOI":"10.1109\/CVPR.2017.429"},{"key":"2025120523494748263_j_jisys-2019-0157_ref_019","doi-asserted-by":"crossref","unstructured":"Y.Y. Zhang, D. Zhou, S. Chen, S.H. Gao, and Y. Ma. Single-image crowd counting via multi-column convolutional neural network. In IEEE Conference on Computer Vision and Pattern Recognition 2016.","DOI":"10.1109\/CVPR.2016.70"},{"key":"2025120523494748263_j_jisys-2019-0157_ref_020","doi-asserted-by":"crossref","unstructured":"P. Isola, J.Y. Zhu, T.H. Zhou, and A.A. Efros. Image-to-image translation with conditional adversarial networks. In Computer Vision and Pattern Recognition pages 5967\u20135976, 2016.","DOI":"10.1109\/CVPR.2017.632"},{"key":"2025120523494748263_j_jisys-2019-0157_ref_021","doi-asserted-by":"crossref","unstructured":"D. Cires, U. Meier, and J. Schmidhuber. Multi-column deep neural networks for image classification. In Computer Vision and Pattern Recognition volume 157, pages 3642\u20133649, 2012.","DOI":"10.1109\/CVPR.2012.6248110"},{"key":"2025120523494748263_j_jisys-2019-0157_ref_022","doi-asserted-by":"crossref","unstructured":"M. Li, Z.X. Zhang, K.Q. Huang, and T.N. Tan. Estimating the number of people in crowded scenes by mid based foreground segmentation and head-shoulder detection. In International Conference on Pattern Recognition pages 1\u20134, 2009.","DOI":"10.1109\/ICPR.2008.4761705"},{"key":"2025120523494748263_j_jisys-2019-0157_ref_023","doi-asserted-by":"crossref","unstructured":"C. Szegedy, W. Liu, Y.Q. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich. Going deeper with convolutions. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR) pages 1\u20139, 2015.","DOI":"10.1109\/CVPR.2015.7298594"},{"key":"2025120523494748263_j_jisys-2019-0157_ref_024","unstructured":"K. Dan, G. Douglas, and T. Hai. Counting pedestrians in crowds using viewpoint invariant training. In British Machine Vision Conference 2005."},{"key":"2025120523494748263_j_jisys-2019-0157_ref_025","doi-asserted-by":"crossref","unstructured":"C.S. Regazzoni and A. Tesei. Distributed data fusion for real-time crowding estimation. Signal Processing 53(1):47\u201363, 1996.","DOI":"10.1016\/0165-1684(96)00075-8"},{"key":"2025120523494748263_j_jisys-2019-0157_ref_026","doi-asserted-by":"crossref","unstructured":"C. Wang, H. Zhang, L. Yang, S. Liu, and X.C. Cao. Deep people counting in extremely dense crowds. In ACM International Conference on Multimedia pages 1299\u20131302, 2015.","DOI":"10.1145\/2733373.2806337"},{"key":"2025120523494748263_j_jisys-2019-0157_ref_027","unstructured":"H. Zhang, V. Sindagi, and V.M. Patel. Image de-raining using a conditional generative adversarial network. In IEEE Conference on Computer Vision and Pattern Recognition 2017."},{"key":"2025120523494748263_j_jisys-2019-0157_ref_028","doi-asserted-by":"crossref","unstructured":"D.G. Lowe. Distinctive image features from scale-invariant keypoints. International Journal of Computer Vision 60(2):91\u2013 110, 2004.","DOI":"10.1023\/B:VISI.0000029664.99615.94"},{"key":"2025120523494748263_j_jisys-2019-0157_ref_029","doi-asserted-by":"crossref","unstructured":"L. Zhang and M.J. Shi. Crowd counting via scale-adaptive convolutional neural network. In IEEE Conference on Computer Vision and Pattern Recognition 2017.","DOI":"10.1109\/WACV.2018.00127"},{"key":"2025120523494748263_j_jisys-2019-0157_ref_030","doi-asserted-by":"crossref","unstructured":"V.A. Sindagi and V.M. Patel. Generating high-quality crowd density maps using contextual pyramid cnns. In International Conference on Computer Vision 2017.","DOI":"10.1109\/ICCV.2017.206"},{"key":"2025120523494748263_j_jisys-2019-0157_ref_031","unstructured":"A. Boesen-Lindbo-Larsen, S. Kaae-S\u2205nderby, H. Larochelle, and O. Winther. Autoencoding beyond pixels using a learned similarity metric. In International Conference on Machine Learning 2015."},{"key":"2025120523494748263_j_jisys-2019-0157_ref_032","unstructured":"I.J. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio. Generative adversarial networks. In Neural Information Processing Systems 2014."},{"key":"2025120523494748263_j_jisys-2019-0157_ref_033","doi-asserted-by":"crossref","unstructured":"D. Pathak, P. Krahenbuhl, J. Donahue, T. Darrell, and A.A. Efros. Context encoders: Feature learning by inpainting. In IEEE Computer Vision and Pattern Recognition 2016.","DOI":"10.1109\/CVPR.2016.278"},{"key":"2025120523494748263_j_jisys-2019-0157_ref_034","doi-asserted-by":"crossref","unstructured":"J. Johnson, A. Alahi, and F.F. Li. Perceptual losses for real-time style transfer and super-resolution. In Bastian Leibe, Jiri Matas, Nicu Sebe, and Max Welling, editors, European Conference on Computer Vision pages 694\u2013711. Springer International Publishing, 2016.","DOI":"10.1007\/978-3-319-46475-6_43"},{"key":"2025120523494748263_j_jisys-2019-0157_ref_035","doi-asserted-by":"crossref","unstructured":"V.A. Sindagi and V.M. Patel. Cnn-based cascaded multi-task learning of high-level prior and density estimation for crowd counting. In Advanced Video and Signal Based Surveillance 2017.","DOI":"10.1109\/AVSS.2017.8078491"}],"container-title":["Journal of Intelligent Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.degruyter.com\/view\/journals\/jisys\/30\/1\/article-p180.xml","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/www.degruyterbrill.com\/document\/doi\/10.1515\/jisys-2019-0157\/xml","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/www.degruyterbrill.com\/document\/doi\/10.1515\/jisys-2019-0157\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,12,5]],"date-time":"2025-12-05T23:50:10Z","timestamp":1764978610000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.degruyterbrill.com\/document\/doi\/10.1515\/jisys-2019-0157\/html"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,7,8]]},"references-count":35,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2020,8,15]]},"published-print":{"date-parts":[[2020,8,15]]}},"alternative-id":["10.1515\/jisys-2019-0157"],"URL":"https:\/\/doi.org\/10.1515\/jisys-2019-0157","relation":{},"ISSN":["2191-026X"],"issn-type":[{"value":"2191-026X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,7,8]]}}}