{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,6]],"date-time":"2026-06-06T19:33:22Z","timestamp":1780774402780,"version":"3.54.1"},"reference-count":27,"publisher":"World Scientific Pub Co Pte Ltd","issue":"08","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Int. J. Patt. Recogn. Artif. Intell."],"published-print":{"date-parts":[[2025,6,30]]},"abstract":"<jats:p> Object counting is a fundamental task in computer vision, with critical applications in areas such as crowd monitoring and ecological conservation. Traditional methods typically rely on large-scale annotated datasets, which are costly and time-consuming to obtain. Few-shot object counting has emerged as a promising solution, enabling accurate counting with minimal annotated samples. However, in real-world scenarios, objects often exhibit significant scale variations due to factors such as view distortion, varying shooting distances, and inherent size differences. Existing few-shot methods usually struggle to address this challenge effectively. To address these issues, we propose a Scale-Aware Vision Transformer (SAViT) framework. Specifically, we design a multi-scale dilated convolution module in SAViT, which can adaptively adjust convolution kernel sampling rates to handle objects of varying sizes. Additionally, we incorporate a global channel attention mechanism to strengthen the model\u2019s ability to capture robust feature representations, thereby improving detection accuracy. For practical usability, we integrate the Segment Anything Model (SAM) to create an exemplar box selection module, simplifying the process by allowing users to generate precise exemplar boxes with a single line drawn on the target object. Extensive experiments on the FSC-147 dataset demonstrate the effectiveness of our approach, achieving a Mean Absolute Error (MAE) of 8.92 and a Root Mean Squared Error (RMSE) of 31.26. Compared to the state-of-the-art method, CACViT, our model reduces MAE by 0.21 (2.30% improvement) and RMSE by 17.7 (36.15% improvement). Our approach not only provides an effective solution for few-shot object counting but also provides a new practical paradigm for extending few-shot learning to complex vision tasks requiring multi-scale reasoning. The code of our paper is available at https:\/\/github.com\/BlouseDong\/SAViT . <\/jats:p>","DOI":"10.1142\/s0218001425560014","type":"journal-article","created":{"date-parts":[[2025,4,26]],"date-time":"2025-04-26T00:20:12Z","timestamp":1745626812000},"source":"Crossref","is-referenced-by-count":1,"title":["Few-Shot Counting with Multi-Scale Vision Transformers and Attention Mechanisms"],"prefix":"10.1142","volume":"39","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-2766-9367","authenticated-orcid":false,"given":"Xiaopan","family":"Chen","sequence":"first","affiliation":[{"name":"School of Computer and Information Engineering, Henan University, Kaifeng, P. R. China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0004-7349-7596","authenticated-orcid":false,"given":"Zhiwei","family":"Dong","sequence":"additional","affiliation":[{"name":"Henan Key Laboratory of Big Data Analysis and Processing, Henan University, Kaifeng, P. R. China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0664-1832","authenticated-orcid":false,"given":"Xiaoke","family":"Zhu","sequence":"additional","affiliation":[{"name":"School of Computer and Information Engineering, Henan University, Kaifeng, P. R. China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2176-3835","authenticated-orcid":false,"given":"Fan","family":"Zhang","sequence":"additional","affiliation":[{"name":"Henan Engineering Research Center of Intelligent Technology and Application, Henan University, Kaifeng, P. R. China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0004-9622-4179","authenticated-orcid":false,"given":"Caihong","family":"Yuan","sequence":"additional","affiliation":[{"name":"School of Computer and Information Engineering, Henan University, Kaifeng, P. R. China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"219","published-online":{"date-parts":[[2025,5,19]]},"reference":[{"key":"S0218001425560014BIB001","doi-asserted-by":"publisher","DOI":"10.1016\/j.media.2015.03.002"},{"key":"S0218001425560014BIB002","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2024.3515835"},{"key":"S0218001425560014BIB003","volume-title":"Int. Conf. Learning Representations","author":"Dosovitskiy A.","year":"2021"},{"key":"S0218001425560014BIB004","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2020.01.087"},{"key":"S0218001425560014BIB005","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00537"},{"key":"S0218001425560014BIB006","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-19827-4_23"},{"key":"S0218001425560014BIB007","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.446"},{"key":"S0218001425560014BIB009","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.00371"},{"key":"S0218001425560014BIB010","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00192"},{"key":"S0218001425560014BIB012","first-page":"313","volume-title":"British Machine Vision Conf.","author":"Lin W.","year":"2022"},{"key":"S0218001425560014BIB013","first-page":"370","volume-title":"British Machine Vision Conf.","author":"Liu C.","year":"2022"},{"key":"S0218001425560014BIB014","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2024.3349978"},{"key":"S0218001425560014BIB015","first-page":"669","volume-title":"Asian Conf. Computer Vision","author":"Lu E.","year":"2018"},{"key":"S0218001425560014BIB016","doi-asserted-by":"publisher","DOI":"10.1016\/j.agrformet.2018.10.013"},{"key":"S0218001425560014BIB017","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46487-9_48"},{"key":"S0218001425560014BIB018","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46478-7_38"},{"key":"S0218001425560014BIB019","doi-asserted-by":"publisher","DOI":"10.1109\/ICPR.2018.8545068"},{"key":"S0218001425560014BIB020","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2024.3351805"},{"key":"S0218001425560014BIB021","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW56347.2022.00467"},{"key":"S0218001425560014BIB022","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00340"},{"key":"S0218001425560014BIB023","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.00931"},{"key":"S0218001425560014BIB024","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01900"},{"key":"S0218001425560014BIB025","volume-title":"Annual Conf. Neural Information Processing Systems","author":"Wang B.","year":"2020"},{"key":"S0218001425560014BIB026","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2020.3013269"},{"key":"S0218001425560014BIB027","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v38i6.28396"},{"key":"S0218001425560014BIB028","doi-asserted-by":"publisher","DOI":"10.1080\/21681163.2016.1149104"},{"key":"S0218001425560014BIB029","doi-asserted-by":"publisher","DOI":"10.1109\/WACV56688.2023.00625"}],"container-title":["International Journal of Pattern Recognition and Artificial Intelligence"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.worldscientific.com\/doi\/pdf\/10.1142\/S0218001425560014","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,5,29]],"date-time":"2025-05-29T05:34:51Z","timestamp":1748496891000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.worldscientific.com\/doi\/10.1142\/S0218001425560014"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,5,19]]},"references-count":27,"journal-issue":{"issue":"08","published-print":{"date-parts":[[2025,6,30]]}},"alternative-id":["10.1142\/S0218001425560014"],"URL":"https:\/\/doi.org\/10.1142\/s0218001425560014","relation":{},"ISSN":["0218-0014","1793-6381"],"issn-type":[{"value":"0218-0014","type":"print"},{"value":"1793-6381","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,5,19]]},"article-number":"2556001"}}