{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,12,30]],"date-time":"2025-12-30T08:50:49Z","timestamp":1767084649938,"version":"build-2065373602"},"reference-count":47,"publisher":"MDPI AG","issue":"2","license":[{"start":{"date-parts":[[2021,1,31]],"date-time":"2021-01-31T00:00:00Z","timestamp":1612051200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100004955","name":"\u00d6sterreichische Forschungsf\u00f6rderungsgesellschaft","doi-asserted-by":"publisher","award":["854747 and 879709"],"award-info":[{"award-number":["854747 and 879709"]}],"id":[{"id":"10.13039\/501100004955","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["J. Imaging"],"abstract":"<jats:p>Many scientific studies deal with person counting and density estimation from single images. Recently, convolutional neural networks (CNNs) have been applied for these tasks. Even though often better results are reported, it is often not clear where the improvements are resulting from, and if the proposed approaches would generalize. Thus, the main goal of this paper was to identify the critical aspects of these tasks and to show how these limit state-of-the-art approaches. Based on these findings, we show how to mitigate these limitations. To this end, we implemented a CNN-based baseline approach, which we extended to deal with identified problems. These include the discovery of bias in the reference data sets, ambiguity in ground truth generation, and mismatching of evaluation metrics w.r.t. the training loss function. The experimental results show that our modifications allow for significantly outperforming the baseline in terms of the accuracy of person counts and density estimation. In this way, we get a deeper understanding of CNN-based person density estimation beyond the network architecture. Furthermore, our insights would allow to advance the field of person density estimation in general by highlighting current limitations in the evaluation protocols.<\/jats:p>","DOI":"10.3390\/jimaging7020021","type":"journal-article","created":{"date-parts":[[2021,1,31]],"date-time":"2021-01-31T21:31:56Z","timestamp":1612128716000},"page":"21","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":8,"title":["Critical Aspects of Person Counting and Density Estimation"],"prefix":"10.3390","volume":"7","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-3374-4201","authenticated-orcid":false,"given":"Roland","family":"Perko","sequence":"first","affiliation":[{"name":"Joanneum Research Forschungsgesellschaft mbH, DIGITAL, Remote Sensing and Geoinformation, 8010 Graz, Austria"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Manfred","family":"Klopschitz","sequence":"additional","affiliation":[{"name":"Joanneum Research Forschungsgesellschaft mbH, DIGITAL, Remote Sensing and Geoinformation, 8010 Graz, Austria"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Alexander","family":"Almer","sequence":"additional","affiliation":[{"name":"Joanneum Research Forschungsgesellschaft mbH, DIGITAL, Remote Sensing and Geoinformation, 8010 Graz, Austria"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9566-1298","authenticated-orcid":false,"given":"Peter M.","family":"Roth","sequence":"additional","affiliation":[{"name":"Data Science in Earth Observation, Technical University of Munich, 82024 Taufkirchen, Ottobrunn, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2021,1,31]]},"reference":[{"key":"ref_1","unstructured":"Hopkins, I.H.G., Pountney, S.J., Hayes, P., and Sheppard, M.A. (1993). Crowd pressure monitoring. Eng. Crowd Saf., 389\u2013398."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"712","DOI":"10.1016\/j.aap.2006.01.001","article-title":"Prediction of human crowd pressures","volume":"38","author":"Lee","year":"2006","journal-title":"Accid. Anal. Prev."},{"key":"ref_3","unstructured":"Perko, R., Schnabel, T., Fritz, G., Almer, A., and Paletta, L. (2013). Counting people from above: Airborne video based crowd analysis. Workshop of the Austrian Association for Pattern Recognition. arXiv."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Helbing, D., and Johansson, A. (2009). Pedestrian, crowd and evacuation dynamics. Encyclopedia of Complexity and Systems Science, Springer.","DOI":"10.1007\/978-0-387-30440-3_382"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"046109","DOI":"10.1103\/PhysRevE.75.046109","article-title":"Dynamics of crowd disasters: An empirical study","volume":"75","author":"Helbing","year":"2007","journal-title":"Phys. Rev. E"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Almer, A., Perko, R., Schrom-Feiertag, H., Schnabel, T., and Paletta, L. (2016). Critical Situation Monitoring at Large Scale Events from Airborne Video based Crowd Dynamics Analysis. Geospatial Data in a Changing World\u2014Selected papers of the AGILE Conference on Geographic Information Science, Springer.","DOI":"10.1007\/978-3-319-33783-8_20"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Perko, R., Schnabel, T., Fritz, G., Almer, A., and Paletta, L. (2013). Airborne based High Performance Crowd Monitoring for Security Applications. Scandinavian Conference on Image Analysis, Springer.","DOI":"10.1007\/978-3-642-38886-6_62"},{"key":"ref_8","unstructured":"Perko, R., Schnabel, T., Almer, A., and Paletta, L. (2014, January 1\u20132). Towards View Invariant Person Counting and Crowd Density Estimation for Remote Vision-Based Services. Proceedings of the IEEE Electrotechnical and Computer Science Conference, Bhopal, India."},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Li, Y., Zhang, X., and Chen, D. (2018, January 18\u201322). CSRNet: Dilated convolutional neural networks for understanding the highly congested scenes. Proceedings of the Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00120"},{"key":"ref_10","unstructured":"Lempitsky, V., and Zisserman, A. (2021, January 31). Learning to Count Objects in Images. Advances in Neural Information Processing Systems. Available online: https:\/\/proceedings.neurips.cc\/paper\/2010."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"98","DOI":"10.1007\/s11263-014-0733-5","article-title":"The PASCAL visual object classes challenge: A retrospective","volume":"111","author":"Everingham","year":"2015","journal-title":"Int. J. Comput. Vis."},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Redmon, J., Divvala, S., Girshick, R., and Farhadi, A. (2016, January 27\u201330). You only look once: Unified, real-time object detection. Proceedings of the Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.91"},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"604","DOI":"10.1109\/TPAMI.2009.204","article-title":"Shape-based human detection and segmentation via hierarchical part-template matching","volume":"32","author":"Lin","year":"2010","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_14","unstructured":"Wu, B., and Nevatia, R. (2005, January 17\u201321). Detection of multiple, partially occluded humans in a single image by bayesian combination of edgelet part detectors. Proceedings of the International Conference on Computer Vision, Beijing, China."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Kong, D., Gray, D., and Tao, H. (2006, January 20\u201324). A viewpoint invariant approach for crowd counting. Proceedings of the International Conference on Pattern Recognition, Hong Kong, China.","DOI":"10.1109\/ICPR.2006.197"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Marana, A.N., Velastin, S., Costa, L., and Lotufo, R. (1997, January 10). Estimation of crowd density using image processing. Proceedings of the IEE Colloquium on Image Processing for Security Applications, London, UK.","DOI":"10.1049\/ic:19970387"},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"535","DOI":"10.1109\/3477.775269","article-title":"A neural-based crowd estimation by hybrid global learning algorithm","volume":"29","author":"Cho","year":"1999","journal-title":"IEEE Trans. Syst. Man Cybern. Part B"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Arteta, C., Lempitsky, V., Noble, J.A., and Zisserman, A. (2014, January 6\u201312). Interactive object counting. Proceedings of the European Conference on Computer Vision, Zurich, Switzerland.","DOI":"10.1007\/978-3-319-10578-9_33"},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"91","DOI":"10.1023\/B:VISI.0000029664.99615.94","article-title":"Distinctive image features from scale-invariant keypoints","volume":"60","author":"Lowe","year":"2004","journal-title":"Int. J. Comput. Vis."},{"key":"ref_20","unstructured":"Dalal, N., and Triggs, B. (2005, January 20\u201325). Histograms of oriented gradients for human detection. Proceedings of the Conference on Computer Vision and Pattern Recognition, San Diego, CA, USA."},{"key":"ref_21","unstructured":"Fiaschi, L., K\u00f6the, U., Nair, R., and Hamprecht, F.A. (2012, January 11\u201315). Learning to count with regression forest and structured labels. Proceedings of the International Conference on Pattern Recognition, Tsukuba, Japan."},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Sam, D.B., Surya, S., and Babu, R.V. (2017, January 21\u201326). Switching convolutional neural network for crowd counting. Proceedings of the Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.429"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Cao, X., Wang, Z., Zhao, Y., and Su, F. (2018, January 8\u201314). Scale aggregation network for accurate and efficient crowd counting. Proceedings of the European Conference on Computer Vision, Munich, Germany.","DOI":"10.1007\/978-3-030-01228-1_45"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Zhang, Y., Zhou, D., Chen, S., Gao, S., and Ma, Y. (2016, January 27\u201330). Single-image crowd counting via multi-column convolutional neural network. Proceedings of the Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.70"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Idrees, H., Tayyab, M., Athrey, K., Zhang, D., Al-Maadeed, S., Rajpoot, N., and Shah, M. (2018, January 8\u201314). Composition loss for counting, density map estimation and localization in dense crowds. Proceedings of the European Conference on Computer Vision, Munich, Germany.","DOI":"10.1007\/978-3-030-01216-8_33"},{"key":"ref_26","unstructured":"Bahmanyar, R., Vig, E., and Reinartz, P. (2019, January 9\u201312). MRCNet: Crowd counting and density map estimation in aerial and ground imagery. Proceedings of the British Machine Vision Conference Workshop on Object Detection and Recognition for Security Screening, Cardiff, UK."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Sindagi, V.A., and Patel, V.M. (2017, January 22\u201329). Generating high-quality crowd density maps using contextual pyramid CNNs. Proceedings of the International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.206"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Liu, W., Salzmann, M., and Fua, P. (2019, January 16\u201320). Context-aware Crowd Counting. Proceedings of the Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00524"},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"225","DOI":"10.1016\/j.neucom.2019.10.096","article-title":"A GPSO-optimized convolutional neural networks for EEG-based emotion recognition","volume":"380","author":"Gao","year":"2020","journal-title":"Neurocomputing"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Bacanin, N., Bezdan, T., Tuba, E., Strumberger, I., and Tuba, M. (2020). Monarch Butterfly Optimization Based Convolutional Neural Network Design. Mathematics, 8.","DOI":"10.3390\/math8060936"},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"156139","DOI":"10.1109\/ACCESS.2020.3019245","article-title":"Automating configuration of convolutional neural network hyperparameters using genetic algorithm","volume":"8","author":"Johnson","year":"2020","journal-title":"IEEE Access"},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"520","DOI":"10.1016\/j.tics.2007.09.009","article-title":"The role of context in object recognition","volume":"11","author":"Oliva","year":"2007","journal-title":"Trends Cogn. Sci."},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"700","DOI":"10.1016\/j.cviu.2010.03.005","article-title":"A framework for visual-context-aware object detection in still images","volume":"114","author":"Perko","year":"2010","journal-title":"Comput. Vis. Image Underst."},{"key":"ref_34","unstructured":"Simonyan, K., and Zisserman, A. (2014). Very deep convolutional networks for large-scale image recognition. arXiv."},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"1904","DOI":"10.1109\/TPAMI.2015.2389824","article-title":"Spatial pyramid pooling in deep convolutional networks for visual recognition","volume":"37","author":"He","year":"2015","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Liu, W., Anguelov, D., Erhan, D., Szegedy, C., Reed, S., Fu, C.Y., and Berg, A.C. (2016, January 8\u201316). SSD: Single shot multibox detector. Proceedings of the European Conference on Computer Vision, Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46448-0_2"},{"key":"ref_37","unstructured":"Zhang, C., Li, H., Wang, X., and Yang, X. (2015, January 7\u201312). Cross-scene crowd counting via deep convolutional neural networks. Proceedings of the Conference on Computer Vision and Pattern Recognition, Boston, MA, USA."},{"key":"ref_38","unstructured":"Wang, Q., Gao, J., Lin, W., and Li, X. (2020). NWPU-Crowd: A large-scale benchmark for crowd counting. IEEE Trans. Pattern Anal. Mach. Intell., 1\u201310."},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Torralba, A., and Efros, A.A. (2011, January 20\u201325). Unbiased look at dataset bias. Proceedings of the Conference on Computer Vision and Pattern Recognition, Colorado Springs, CO, USA.","DOI":"10.1109\/CVPR.2011.5995347"},{"key":"ref_40","unstructured":"Dijk, T.v., and Croon, G.d. (November, January 27). How do neural networks see depth in single images?. Proceedings of the International Conference on Computer Vision, Seoul, Korea."},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Idrees, H., Saleemi, I., Seibert, C., and Shah, M. (2013, January 23\u201328). Multi-source multi-scale counting in extremely dense crowd images. Proceedings of the Conference on Computer Vision and Pattern Recognition, Portland, OR, USA.","DOI":"10.1109\/CVPR.2013.329"},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Chan, A.B., Liang, Z.S.J., and Vasconcelos, N. (2008, January 23\u201328). Privacy preserving crowd monitoring: Counting people without people models or tracking. Proceedings of the Conference on Computer Vision and Pattern Recognition, Anchorage, AK, USA.","DOI":"10.1109\/CVPR.2008.4587569"},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Liu, J., Gao, C., Meng, D., and Hauptmann, A.G. (2018, January 18\u201322). DecideNet: Counting varying density crowds through attention guided detection and density estimation. Proceedings of the Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00545"},{"key":"ref_44","doi-asserted-by":"crossref","first-page":"7","DOI":"10.1023\/A:1014573219977","article-title":"A taxonomy and evaluation of dense two-frame stereo correspondence algorithms","volume":"47","author":"Scharstein","year":"2002","journal-title":"Int. J. Comput. Vis."},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Cordts, M., Omran, M., Ramos, S., Rehfeld, T., Enzweiler, M., Benenson, R., Franke, U., Roth, S., and Schiele, B. (2016, January 27\u201330). The Cityscapes Dataset for Semantic Urban Scene Understanding. Proceedings of the Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.350"},{"key":"ref_46","unstructured":"Kingma, D.P., and Ba, J. (2014). Adam: A method for stochastic optimization. arXiv."},{"key":"ref_47","doi-asserted-by":"crossref","first-page":"211","DOI":"10.1007\/s11263-015-0816-y","article-title":"Imagenet large scale visual recognition challenge","volume":"115","author":"Russakovsky","year":"2015","journal-title":"Int. J. Comput. Vis."}],"container-title":["Journal of Imaging"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2313-433X\/7\/2\/21\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T05:17:52Z","timestamp":1760159872000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2313-433X\/7\/2\/21"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,1,31]]},"references-count":47,"journal-issue":{"issue":"2","published-online":{"date-parts":[[2021,2]]}},"alternative-id":["jimaging7020021"],"URL":"https:\/\/doi.org\/10.3390\/jimaging7020021","relation":{},"ISSN":["2313-433X"],"issn-type":[{"type":"electronic","value":"2313-433X"}],"subject":[],"published":{"date-parts":[[2021,1,31]]}}}