{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,2]],"date-time":"2026-08-02T20:14:19Z","timestamp":1785701659040,"version":"3.56.0"},"reference-count":59,"publisher":"MDPI AG","issue":"3","license":[{"start":{"date-parts":[[2022,1,20]],"date-time":"2022-01-20T00:00:00Z","timestamp":1642636800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Remote Sensing"],"abstract":"<jats:p>The number of trees and their spatial distribution are key information for forest management. In recent years, deep learning-based approaches have been proposed and shown promising results in lowering the expensive labor cost of a forest inventory. In this paper, we propose a new efficient deep learning model called density transformer or DENT for automatic tree counting from aerial images. The architecture of DENT contains a multi-receptive field convolutional neural network to extract visual feature representation from local patches and their wide context, a transformer encoder to transfer contextual information across correlated positions, a density map generator to generate spatial distribution map of trees, and a fast tree counter to estimate the number of trees in each input image. We compare DENT with a variety of state-of-art methods, including one-stage and two-stage, anchor-based and anchor-free deep neural detectors, and different types of fully convolutional regressors for density estimation. The methods are evaluated on a new large dataset we built and an existing cross-site dataset. DENT achieves top accuracy on both datasets, significantly outperforming most of the other methods. We have released our new dataset, called Yosemite Tree Dataset, containing a 10 km2 rectangular study area with around 100k trees annotated, as a benchmark for public access.<\/jats:p>","DOI":"10.3390\/rs14030476","type":"journal-article","created":{"date-parts":[[2022,1,20]],"date-time":"2022-01-20T22:51:06Z","timestamp":1642719066000},"page":"476","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":38,"title":["Transformer for Tree Counting in Aerial Images"],"prefix":"10.3390","volume":"14","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-1487-8044","authenticated-orcid":false,"given":"Guang","family":"Chen","sequence":"first","affiliation":[{"name":"Department of Electrical Engineering and Computer Science (EECS), University of Missouri, Columbia, MO 65211, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7771-4034","authenticated-orcid":false,"given":"Yi","family":"Shang","sequence":"additional","affiliation":[{"name":"Department of Electrical Engineering and Computer Science (EECS), University of Missouri, Columbia, MO 65211, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2022,1,20]]},"reference":[{"key":"ref_1","first-page":"1097","article-title":"ImageNet classification with deep convolutional neural networks","volume":"Volume 1","author":"Krizhevsky","year":"2012","journal-title":"Proceedings of the 25th International Conference on Neural Information Processing Systems, NIPS\u201912"},{"key":"ref_2","unstructured":"Simonyan, K., and Zisserman, A. (2015, January 7\u20139). Very deep convolutional networks for large-scale image recognition. Proceedings of the International Conference on Learning Representations, San Diego, CA, USA."},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep residual learning for image recognition. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S., Anguelov, D., Erhan, D., Vanhoucke, V., and Rabinovich, A. (2015, January 7\u201312). Going deeper with convolutions. Proceedings of the 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298594"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Szegedy, C., Vanhoucke, V., Ioffe, S., Shlens, J., and Wojna, Z. (2016, January 27\u201330). Rethinking the inception architecture for computer vision. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.308"},{"key":"ref_6","first-page":"91","article-title":"Faster R-CNN: Towards real-time object detection with region proposal networks","volume":"Volume 1","author":"Ren","year":"2015","journal-title":"Proceedings of the 28th International Conference on Neural Information Processing Systems, NIPS\u201915"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Leibe, B., Matas, J., Sebe, N., and Welling, M. (2016). SSD: Single shot multiBox detector. European Conference on Computer Vision (ECCV), Proceedings of the 14th European Conference, Amsterdam, The Netherlands, 11\u201314 October 2016, Springer International Publishing.","DOI":"10.1007\/978-3-319-46475-6"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Lin, T.Y., Goyal, P., Girshick, R., He, K., and Doll\u00e1r, P. (2017, January 22\u201329). Focal loss for dense object detection. Proceedings of the IEEE International Conference on Computer Vision (ICCV), Venice, Italy.","DOI":"10.1109\/ICCV.2017.324"},{"key":"ref_9","unstructured":"Redmon, J., and Farhadi, A. (2018). Yolov3: An incremental improvement. arXiv."},{"key":"ref_10","unstructured":"Zhou, X., Wang, D., and Kr\u00e4henb\u00fchl, P. (2019). Objects as points. arXiv."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Vedaldi, A., Bischof, H., Brox, T., and Frahm, J.M. (2020). End-to-end object detection with transformers. European Conference on Computer Vision (ECCV), Proceedings of the 16th European Conference, Glasgow, UK, 23\u201328 August 2020, Springer International Publishing.","DOI":"10.1007\/978-3-030-58548-8"},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"7500","DOI":"10.1080\/01431161.2019.1569282","article-title":"Young and mature oil palm tree detection and counting using convolutional neural network deep learning method","volume":"40","author":"Mubin","year":"2019","journal-title":"Int. J. Remote Sens."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Li, W., Fu, H., Yu, L., and Cracknell, A. (2017). Deep learning based oil palm tree detection and counting for high-resolution remote sensing images. Remote Sens., 9.","DOI":"10.3390\/rs9010022"},{"key":"ref_14","first-page":"65","article-title":"Fast and robust detection of oil palm trees using high-resolution remote sensing images","volume":"Volume 10988","author":"Hammoud","year":"2019","journal-title":"Automatic Target Recognition XXIX"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Machefer, M., Lemarchand, F., Bonnefond, V., Hitchins, A., and Sidiropoulos, P. (2020). Mask R-CNN Refitting Strategy for Plant Counting and Sizing in UAV Imagery. Remote Sens., 12.","DOI":"10.3390\/rs12183015"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Weinstein, B.G., Marconi, S., Bohlman, S., Zare, A., and White, E. (2019). Individual tree-crown detection in RGB imagery using semi-supervised deep learning neural networks. Remote Sens., 11.","DOI":"10.1101\/532952"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Roslan, Z., Awang, Z., Husen, M.N., Ismail, R., and Hamzah, R. (2020, January 3\u20135). Deep learning for tree crown detection in tropical forest. Proceedings of the 2020 14th International Conference on Ubiquitous Information Management and Communication (IMCOM), Taichung, Taiwan.","DOI":"10.1109\/IMCOM48794.2020.9001817"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Zheng, J., Li, W., Xia, M., Dong, R., Fu, H., and Yuan, S. (August, January 28). Large-scale oil palm tree detection from high-resolution remote sensing images using faster-rcnn. Proceedings of the IGARSS 2019-2019 IEEE International Geoscience and Remote Sensing Symposium, Yokohama, Japan.","DOI":"10.1109\/IGARSS.2019.8898360"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Zhang, Y., Zhou, D., Chen, S., Gao, S., and Ma, Y. (2016, January 27\u201330). Single-image crowd counting via multi-column convolutional neural network. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.70"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Sam, D.B., Surya, S., and Babu, R.V. (2017, January 21\u201326). Switching convolutional neural network for crowd counting. Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.429"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Li, Y., Zhang, X., and Chen, D. (2018, January 18\u201323). CSRNet: Dilated convolutional neural networks for understanding the highly congested scenes. Proceedings of the 2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00120"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Ferrari, V., Hebert, M., Sminchisescu, C., and Weiss, Y. (2018, January 8\u201314). Scale aggregation network for accurate and efficient crowd counting. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01249-6"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Liu, W., Salzmann, M., and Fua, P. (2019, January 15\u201320). Context-aware crowd counting. Proceedings of the 2019 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00524"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Djerriri, K., Ghabi, M., Karoui, M.S., and Adjoudj, R. (2018, January 22\u201327). Palm trees counting in remote sensing imagery using regression convolutional neural network. Proceedings of the IGARSS 2018-2018 IEEE International Geoscience and Remote Sensing Symposium, Valencia, Spain.","DOI":"10.1109\/IGARSS.2018.8519188"},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"107591","DOI":"10.1016\/j.ecolind.2021.107591","article-title":"Tree counting with high spatial-resolution satellite imagery based on deep neural networks","volume":"125","author":"Yao","year":"2021","journal-title":"Ecol. Indic."},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"101061","DOI":"10.1016\/j.ecoinf.2020.101061","article-title":"Cross-site learning in deep learning RGB tree crown detection","volume":"56","author":"Weinstein","year":"2020","journal-title":"Ecol. Inform."},{"key":"ref_27","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, L., and Polosukhin, I. (2017, January 4\u20139). Attention is all you need. Proceedings of the 31st International Conference on Neural Information Processing Systems, NIPS\u201917, Long Beach, CA, USA."},{"key":"ref_28","unstructured":"Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., and Gelly, S. (2021, January 3\u20137). An Image is Worth 16\u00d716 Words: Transformers for Image Recognition at Scale. Proceedings of the International Conference on Learning Representations, Virtual."},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1109\/LGRS.2021.3085139","article-title":"Contrasting YOLOv5, Transformer, and EfficientDet Detectors for Crop Circle Detection in Desert","volume":"19","author":"Mekhalfi","year":"2022","journal-title":"IEEE Geosci. Remote Sens. Lett."},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Bazi, Y., Bashmal, L., Rahhal, M.M.A., Dayil, R.A., and Ajlan, N.A. (2021). Vision Transformers for Remote Sensing Image Classification. Remote Sens., 13.","DOI":"10.3390\/rs13030516"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Ronneberger, O., Fischer, P., and Brox, T. (2015). U-net: Convolutional networks for biomedical image segmentation. International Conference on Medical Image Computing and Computer-Assisted Intervention (MICCAI), Proceedings of the 18th International Conference, Munich, Germany, 5\u20139 October 2015, Springer International Publishing.","DOI":"10.1007\/978-3-319-24574-4_28"},{"key":"ref_32","unstructured":"Touretzky, D., Mozer, M.C., and Hasselmo, M. (1996). Human face detection in visual scenes. Advances in Neural Information Processing Systems, MIT Press."},{"key":"ref_33","unstructured":"Viola, P., and Jones, M. (2001, January 8\u201314). Rapid object detection using a boosted cascade of simple features. Proceedings of the 2001 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR), Kauai, HI, USA."},{"key":"ref_34","unstructured":"Dalal, N., and Triggs, B. (2005, January 20\u201326). Histograms of oriented gradients for human detection. Proceedings of the 2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR), San Diego, CA, USA."},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"1627","DOI":"10.1109\/TPAMI.2009.167","article-title":"Object detection with discriminatively trained part-based models","volume":"32","author":"Felzenszwalb","year":"2009","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Harzallah, H., Jurie, F., and Schmid, C. (October, January 27). Combining efficient object localization and image classification. Proceedings of the 2009 IEEE 12th International Conference on Computer Vision (ICCV), Kyoto, Japan.","DOI":"10.1109\/ICCV.2009.5459257"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Lowe, D. (1999, January 20\u201325). Object recognition from local scale-invariant features. Proceedings of the Seventh IEEE International Conference on Computer Vision (ICCV), Corfu, Greece.","DOI":"10.1109\/ICCV.1999.790410"},{"key":"ref_38","unstructured":"Pollock, R. (1996). The Automatic Recognition of Individual Trees in Aerial Images of Forests Based on a Synthetic Tree Crown Image Model. [Ph.D. Thesis, University of British Columbia]."},{"key":"ref_39","unstructured":"Larsen, M., and Rudemo, M. (1997, January 9\u201311). Using ray-traced templates to find individual trees in aerial photographs. Proceedings of the Scandinavian Conference on Image Analysis, Lappenranta, Finland."},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Vibha, L., Shenoy, P.D., Venugopal, K., and Patnaik, L. (2009, January 6\u20137). Robust technique for segmentation and counting of trees from remotely sensed data. Proceedings of the 2009 IEEE International Advance Computing Conference, Patiala, India.","DOI":"10.1109\/IADCC.2009.4809228"},{"key":"ref_41","unstructured":"Hung, C., Bryson, M., and Sukkarieh, S. (2011, January 10\u201315). Vision-based shadow-aided tree crown detection and classification algorithm using imagery from an unmanned airborne vehicle. Proceedings of the 34th International Symposium for Remote Sensing of the Environment (ISRSE), Sydney, Australia."},{"key":"ref_42","doi-asserted-by":"crossref","first-page":"465","DOI":"10.5194\/isprs-annals-III-3-465-2016","article-title":"Palm tree detection using circular autocorrelation of polar shape matrix","volume":"3","author":"Manandhar","year":"2016","journal-title":"ISPRS Ann. Photogramm. Remote Sens. Spat. Inf. Sci."},{"key":"ref_43","doi-asserted-by":"crossref","first-page":"7356","DOI":"10.1080\/01431161.2018.1513669","article-title":"Automatic detection of individual oil palm trees from UAV images using HOG features and an SVM classifier","volume":"40","author":"Wang","year":"2019","journal-title":"Int. J. Remote Sens."},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"Li, W., Fu, H., and Yu, L. (2017, January 23\u201328). Deep convolutional neural network based large-scale oil palm tree detection for high-resolution remote sensing images. Proceedings of the 2017 IEEE International Geoscience and Remote Sensing Symposium (IGARSS), Fort Worth, TX, USA.","DOI":"10.1109\/IGARSS.2017.8127085"},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Li, W., Dong, R., Fu, H., and Yu, L. (2019). Large-scale oil palm tree detection from high-resolution satellite images using two-stage convolutional neural networks. Remote Sens., 11.","DOI":"10.3390\/rs11010011"},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Huang, G., Liu, Z., Maaten, L.V.D., and Weinberger, K.Q. (2017, January 21\u201326). Densely connected convolutional networks. Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.243"},{"key":"ref_47","doi-asserted-by":"crossref","unstructured":"Freudenberg, M., N\u00f6lke, N., Agostini, A., Urban, K., W\u00f6rg\u00f6tter, F., and Kleinn, C. (2019). Large scale palm tree detection in high resolution satellite images using U-Net. Remote Sens., 11.","DOI":"10.3390\/rs11030312"},{"key":"ref_48","doi-asserted-by":"crossref","unstructured":"Miyoshi, G.T., Arruda, M.d.S., Osco, L.P., Marcato Junior, J., Gon\u00e7alves, D.N., Imai, N.N., Tommaselli, A.M.G., Honkavaara, E., and Gon\u00e7alves, W.N. (2020). A novel deep learning method to identify single tree species in UAV-based hyperspectral images. Remote Sens., 12.","DOI":"10.3390\/rs12081294"},{"key":"ref_49","doi-asserted-by":"crossref","first-page":"e21","DOI":"10.23915\/distill.00021","article-title":"Computing receptive fields of convolutional neural networks","volume":"4","author":"Araujo","year":"2019","journal-title":"Distill"},{"key":"ref_50","unstructured":"Ba, J., Kiros, J.R., and Hinton, G.E. (2016). Layer normalization. arXiv."},{"key":"ref_51","unstructured":"Parmar, N.J., Vaswani, A., Uszkoreit, J., Kaiser, L., Shazeer, N., Ku, A., and Tran, D. (2018, January 10\u201315). Image transformer. Proceedings of the International Conference on Machine Learning (ICML), Stockholm, Sweden."},{"key":"ref_52","unstructured":"Devlin, J., Chang, M.W., Lee, K., and Toutanova, K. (2018). Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv."},{"key":"ref_53","doi-asserted-by":"crossref","unstructured":"Lei, J., Wang, L., Shen, Y., Yu, D., Berg, T.L., and Bansal, M. (2020). Mart: Memory-augmented recurrent transformer for coherent video paragraph captioning. arXiv.","DOI":"10.18653\/v1\/2020.acl-main.233"},{"key":"ref_54","doi-asserted-by":"crossref","first-page":"2481","DOI":"10.1109\/TPAMI.2016.2644615","article-title":"Segnet: A deep convolutional encoder-decoder architecture for image segmentation","volume":"39","author":"Badrinarayanan","year":"2017","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_55","unstructured":"Paszke, A., Gross, S., Chintala, S., Chanan, G., Yang, E., DeVito, Z., Lin, Z., Desmaison, A., Antiga, L., and Lerer, A. (2017, January 4\u20139). Automatic differentiation in pytorch. Proceedings of the Neural Information Processing Systems Workshop, Long Beach, CA, USA."},{"key":"ref_56","doi-asserted-by":"crossref","unstructured":"Deng, J., Dong, W., Socher, R., Li, L.J., Li, K., and Li, F.-F. (2009, January 20\u201325). Imagenet: A large-scale hierarchical image database. Proceedings of the 2009 IEEE Conference on Computer Vision and Pattern Recognition, Miami, FL, USA.","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"ref_57","doi-asserted-by":"crossref","first-page":"211","DOI":"10.1007\/s11263-015-0816-y","article-title":"ImageNet large scale visual recognition challenge","volume":"115","author":"Russakovsky","year":"2015","journal-title":"Int. J. Comput. Vis. (IJCV)"},{"key":"ref_58","unstructured":"Glorot, X., and Bengio, Y. (2010). Understanding the difficulty of training deep feedforward neural networks. JMLR Workshop and Conference Proceedings, Proceedings of the Thirteenth International Conference on Artificial Intelligence and Statistics, Sardinia, Italy, 13\u201315 May 2010, PMLR."},{"key":"ref_59","unstructured":"Kingma, D.P., and Ba, J. (2014). Adam: A method for stochastic optimization. arXiv."}],"container-title":["Remote Sensing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2072-4292\/14\/3\/476\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T22:04:28Z","timestamp":1760133868000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2072-4292\/14\/3\/476"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,1,20]]},"references-count":59,"journal-issue":{"issue":"3","published-online":{"date-parts":[[2022,2]]}},"alternative-id":["rs14030476"],"URL":"https:\/\/doi.org\/10.3390\/rs14030476","relation":{},"ISSN":["2072-4292"],"issn-type":[{"value":"2072-4292","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,1,20]]}}}