{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,15]],"date-time":"2026-05-15T15:43:53Z","timestamp":1778859833896,"version":"3.51.4"},"reference-count":53,"publisher":"Association for Computing Machinery (ACM)","issue":"6","license":[{"start":{"date-parts":[[2023,7,12]],"date-time":"2023-07-12T00:00:00Z","timestamp":1689120000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"Science and Technology Innovation (STI) 2030-Major Projects","award":["2022ZD0208700"],"award-info":[{"award-number":["2022ZD0208700"]}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["61976127"],"award-info":[{"award-number":["61976127"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2023,11,30]]},"abstract":"<jats:p>As an emerging computer vision task, crowd localization has received increasing attention due to its ability to produce more accurate spatially predictions. However, continuous scale variations in complex crowd scenes lead to tiny individuals at the edges, so that existing methods cannot achieve precise crowd localization. Aiming at alleviating the above problems, we propose a novel Dilated Convolution-based Feature Refinement Network (DFRNet) to enhance the representation learning capability. Specifically, the DFRNet is built with three branches that can capture the information of each individual in crowd scenes more precisely. More specifically, we introduce a Feature Perception Module to model long-range contextual information at different scales by adopting multiple dilated convolutions, thus providing sufficient feature information to perceive tiny individuals at the edge of images. Afterwards, a Feature Refinement Module is deployed at multiple stages of the three branches to facilitate the mutual refinement of feature information at different scales, thus further improving the expression capability of multi-scale contextual information. By incorporating the above modules, DFRNet can locate individuals in complex scenes more precisely. Extensive experiments on multiple datasets demonstrate that the proposed method has more advanced performance compared to existing methods and can be more accurately adapted to complex crowd scenes.<\/jats:p>","DOI":"10.1145\/3571134","type":"journal-article","created":{"date-parts":[[2022,12,20]],"date-time":"2022-12-20T12:50:33Z","timestamp":1671540633000},"page":"1-16","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":18,"title":["Dilated Convolution-based Feature Refinement Network for Crowd Localization"],"prefix":"10.1145","volume":"19","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-4660-8092","authenticated-orcid":false,"given":"Xingyu","family":"Gao","sequence":"first","affiliation":[{"name":"Institute of Microelectronics, Chinese Academy of Sciences, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6668-0710","authenticated-orcid":false,"given":"Jinyang","family":"Xie","sequence":"additional","affiliation":[{"name":"School of Information Science and Engineering, Shandong Normal University, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4989-7109","authenticated-orcid":false,"given":"Zhenyu","family":"Chen","sequence":"additional","affiliation":[{"name":"Big Data Center, State Grid Corporation of China, and China Electric Power Research Institute, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5755-9145","authenticated-orcid":false,"given":"An-An","family":"Liu","sequence":"additional","affiliation":[{"name":"School of Electrical and Information Engineering, Tianjin University, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4029-9935","authenticated-orcid":false,"given":"Zhenan","family":"Sun","sequence":"additional","affiliation":[{"name":"Institute of Automation, Chinese Academy of Sciences, and University of Chinese Academy of Sciences, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9521-6039","authenticated-orcid":false,"given":"Lei","family":"Lyu","sequence":"additional","affiliation":[{"name":"School of Information Science and Engineering, Shandong Normal University, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2023,7,12]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v35i2.16170"},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2009.5206754"},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCVW54120.2021.00028"},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2016.2644615"},{"key":"e_1_3_1_6_2","article-title":"Enhanced information fusion network for crowd counting","author":"Chen Geng","year":"2021","unstructured":"Geng Chen and Peirong Guo. 2021. Enhanced information fusion network for crowd counting. arXiv:2101.04279. Retrieved from https:\/\/arxiv.org\/abs\/2101.04279.","journal-title":"arXiv:2101.04279"},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00968"},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.1109\/WACV.2019.00211"},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-88004-0_17"},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2021.3055631"},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","DOI":"10.1214\/009053604000000201"},{"key":"e_1_3_1_12_2","article-title":"Counting and locating high-density objects using convolutional neural network","author":"Arruda Mauro dos Santos de","year":"2021","unstructured":"Mauro dos Santos de Arruda, Lucas Prado Osco, Plabiany Rodrigo Acosta, Diogo Nunes Gon\u00e7alves, Jos\u00e9 Marcato Junior, Ana Paula Marques Ramos, Edson Takashi Matsubara, Zhipeng Luo, Jonathan Li, Jonathan de Andrade Silva, et\u00a0al. 2021. Counting and locating high-density objects using convolutional neural network. arXiv:2102.04366. Retrieved from https:\/\/arxiv.org\/abs\/2102.04366.","journal-title":"arXiv:2102.04366"},{"key":"e_1_3_1_13_2","article-title":"Retinaface: Single-stage dense face localisation in the wild","author":"Deng Jiankang","year":"2019","unstructured":"Jiankang Deng, Jia Guo, Yuxiang Zhou, Jinke Yu, Irene Kotsia, and Stefanos Zafeiriou. 2019. Retinaface: Single-stage dense face localisation in the wild. arXiv:1905.00641. Retrieved from https:\/\/arxiv.org\/abs\/1905.00641.","journal-title":"arXiv:1905.00641"},{"key":"e_1_3_1_14_2","article-title":"Congested crowd instance localization with dilated convolutional swin transformer","author":"Gao Junyu","year":"2021","unstructured":"Junyu Gao, Maoguo Gong, and Xuelong Li. 2021. Congested crowd instance localization with dilated convolutional swin transformer. arXiv:2108.00584. Retrieved from https:\/\/arxiv.org\/abs\/2108.00584.","journal-title":"arXiv:2108.00584"},{"key":"e_1_3_1_15_2","article-title":"Domain-adaptive crowd counting via inter-domain features segregation and gaussian-prior reconstruction","author":"Gao Junyu","year":"2019","unstructured":"Junyu Gao, Tao Han, Qi Wang, and Yuan Yuan. 2019. Domain-adaptive crowd counting via inter-domain features segregation and gaussian-prior reconstruction. arXiv:1912.03677. Retrieved from https:\/arxiv.org\/abs\/1912.03677.","journal-title":"arXiv:1912.03677"},{"key":"e_1_3_1_16_2","article-title":"Learning independent instance maps for crowd localization","author":"Gao Junyu","year":"2020","unstructured":"Junyu Gao, Tao Han, Yuan Yuan, and Qi Wang. 2020. Learning independent instance maps for crowd localization. arXiv:2012.04164. Retrieved from https:\/\/arxiv.org\/abs\/2012.04164.","journal-title":"arXiv:2012.04164"},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.01502"},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1145\/3343031.3350881"},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","DOI":"10.1109\/WACV.2019.00141"},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00745"},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.166"},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01216-8_33"},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2021.06.055"},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00629"},{"key":"e_1_3_1_25_2","first-page":"1097","article-title":"Imagenet classification with deep convolutional neural networks","volume":"25","author":"Krizhevsky Alex","year":"2012","unstructured":"Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton. 2012. Imagenet classification with deep convolutional neural networks. Adv. Neural Inf. Process. Syst. 25 (2012), 1097\u20131105.","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v31i1.11208"},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2020.2985284"},{"key":"e_1_3_1_28_2","article-title":"Pyramidbox++: High performance detector for finding tiny face","author":"Li Zhihang","year":"2019","unstructured":"Zhihang Li, Xu Tang, Junyu Han, Jingtuo Liu, and Ran He. 2019. Pyramidbox++: High performance detector for finding tiny face. arXiv:1904.00386. Retrieved from https:\/\/arxiv.org\/abs\/1904.00386.","journal-title":"arXiv:1904.00386"},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00192"},{"key":"e_1_3_1_30_2","article-title":"Focal inverse distance transform maps for crowd localization and counting in dense crowd","author":"Liang Dingkang","year":"2021","unstructured":"Dingkang Liang, Wei Xu, Yingying Zhu, and Yu Zhou. 2021. Focal inverse distance transform maps for crowd localization and counting in dense crowd. arXiv:2102.07925. Retrieved from https:\/\/arxiv.org\/abs\/2102.07925.","journal-title":"arXiv:2102.07925"},{"key":"e_1_3_1_31_2","doi-asserted-by":"crossref","unstructured":"Dingkang Liang Wei Xu Yingying Zhu and Yu Zhou. 2021. Reciprocal distance transform maps for crowd counting and people localization in dense crowd (unpublished).","DOI":"10.1109\/TMM.2022.3203870"},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00131"},{"key":"e_1_3_1_33_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00545"},{"key":"e_1_3_1_34_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00479"},{"key":"e_1_3_1_35_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00186"},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00943"},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.91"},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","DOI":"10.1109\/WACV45572.2020.9093386"},{"key":"e_1_3_1_39_2","article-title":"Locate, size and count: Accurately resolving people in dense crowds via detection","author":"Sam Deepak Babu","year":"2020","unstructured":"Deepak Babu Sam, Skand Vishwanath Peri, Mukuntha Narayanan Sundararaman, Amogh Kamath, and Venkatesh Babu Radhakrishnan. 2020. Locate, size and count: Accurately resolving people in dense crowds via detection. IEEE Trans. Pattern Anal. Mach. Intell. (2020).","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"e_1_3_1_40_2","article-title":"Very deep convolutional networks for large-scale image recognition","author":"Simonyan Karen","year":"2014","unstructured":"Karen Simonyan and Andrew Zisserman. 2014. Very deep convolutional networks for large-scale image recognition. arXiv:1409.1556. Retrieved from https:\/\/arxiv.org\/abs\/1409.1556.","journal-title":"arXiv:1409.1556"},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00335"},{"key":"e_1_3_1_42_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.255"},{"key":"e_1_3_1_43_2","first-page":"620","volume-title":"Proceedings of the International Conference on Computer Vision Theory and Applications (VISAPP\u201911)","author":"Oosterhout Tim Van","year":"2011","unstructured":"Tim Van Oosterhout, Sander Bakkes, Ben J. A. Kr\u00f6se, et\u00a0al. 2011. Head detection in stereo data for people counting and segmentation. In Proceedings of the International Conference on Computer Vision Theory and Applications (VISAPP\u201911). 620\u2013625."},{"key":"e_1_3_1_44_2","article-title":"Modeling noisy annotations for crowd counting","volume":"33","author":"Wan Jia","year":"2020","unstructured":"Jia Wan and Antoni Chan. 2020. Modeling noisy annotations for crowd counting. Adv. Neural Inf. Process. Syst. 33 (2020).","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"e_1_3_1_45_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00201"},{"key":"e_1_3_1_46_2","article-title":"Dynamic fusion module evolves drivable area and road anomaly detection: A benchmark and algorithms","author":"Wang Hengli","year":"2021","unstructured":"Hengli Wang, Rui Fan, Yuxiang Sun, and Ming Liu. 2021. Dynamic fusion module evolves drivable area and road anomaly detection: A benchmark and algorithms. IEEE Trans. Cybernet. (2021).","journal-title":"IEEE Trans. Cybernet."},{"key":"e_1_3_1_47_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2020.3013269"},{"key":"e_1_3_1_48_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2021.3055632"},{"key":"e_1_3_1_49_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICMEW53276.2021.9455954"},{"key":"e_1_3_1_50_2","doi-asserted-by":"publisher","DOI":"10.1109\/MSP.2008.930649"},{"key":"e_1_3_1_51_2","doi-asserted-by":"publisher","DOI":"10.3354\/cr030079"},{"key":"e_1_3_1_52_2","unstructured":"Chenfeng Xu Dingkang Liang Yongchao Xu Song Bai Wei Zhan Xiang Bai and Masayoshi Tomizuka. 2019. Autoscale: Learning to scale for crowd counting (unpublished)."},{"key":"e_1_3_1_53_2","doi-asserted-by":"publisher","DOI":"10.1109\/WACV45572.2020.9093394"},{"key":"e_1_3_1_54_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.70"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3571134","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3571134","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T18:08:20Z","timestamp":1750183700000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3571134"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,7,12]]},"references-count":53,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2023,11,30]]}},"alternative-id":["10.1145\/3571134"],"URL":"https:\/\/doi.org\/10.1145\/3571134","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,7,12]]},"assertion":[{"value":"2022-04-13","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2022-09-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-07-12","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}