{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,24]],"date-time":"2026-07-24T16:41:09Z","timestamp":1784911269538,"version":"3.55.0"},"reference-count":38,"publisher":"MDPI AG","issue":"13","license":[{"start":{"date-parts":[[2023,6,25]],"date-time":"2023-06-25T00:00:00Z","timestamp":1687651200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>With the increasing popularity of online fruit sales, accurately predicting fruit yields has become crucial for optimizing logistics and storage strategies. However, existing manual vision-based systems and sensor methods have proven inadequate for solving the complex problem of fruit yield counting, as they struggle with issues such as crop overlap and variable lighting conditions. Recently CNN-based object detection models have emerged as a promising solution in the field of computer vision, but their effectiveness is limited in agricultural scenarios due to challenges such as occlusion and dissimilarity among the same fruits. To address this issue, we propose a novel variant model that combines the self-attentive mechanism of Vision Transform, a non-CNN network architecture, with Yolov7, a state-of-the-art object detection model. Our model utilizes two attention mechanisms, CBAM and CA, and is trained and tested on a dataset of apple images. In order to enable fruit counting across video frames in complex environments, we incorporate two multi-objective tracking methods based on Kalman filtering and motion trajectory prediction, namely SORT, and Cascade-SORT. Our results show that the Yolov7-CA model achieved a 91.3% mAP and 0.85 F1 score, representing a 4% improvement in mAP and 0.02 improvement in F1 score compared to using Yolov7 alone. Furthermore, three multi-object tracking methods demonstrated a significant improvement in MAE for inter-frame counting across all three test videos, with an 0.642 improvement over using yolov7 alone achieved using our multi-object tracking method. These findings suggest that our proposed model has the potential to improve fruit yield assessment methods and could have implications for decision-making in the fruit industry.<\/jats:p>","DOI":"10.3390\/s23135903","type":"journal-article","created":{"date-parts":[[2023,6,26]],"date-time":"2023-06-26T05:28:02Z","timestamp":1687757282000},"page":"5903","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":38,"title":["Fruit Detection and Counting in Apple Orchards Based on Improved Yolov7 and Multi-Object Tracking Methods"],"prefix":"10.3390","volume":"23","author":[{"given":"Jing","family":"Hu","sequence":"first","affiliation":[{"name":"School of Mathematics and Computer Science, Wuhan Polytechnic University, Wuhan 430024, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0005-6623-8622","authenticated-orcid":false,"given":"Chuang","family":"Fan","sequence":"additional","affiliation":[{"name":"School of Mathematics and Computer Science, Wuhan Polytechnic University, Wuhan 430024, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Zhoupu","family":"Wang","sequence":"additional","affiliation":[{"name":"School of Mathematics and Computer Science, Wuhan Polytechnic University, Wuhan 430024, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jinglin","family":"Ruan","sequence":"additional","affiliation":[{"name":"School of Mathematics and Computer Science, Wuhan Polytechnic University, Wuhan 430024, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Suyin","family":"Wu","sequence":"additional","affiliation":[{"name":"School of Mathematics and Computer Science, Wuhan Polytechnic University, Wuhan 430024, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2023,6,25]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"103639","DOI":"10.1109\/ACCESS.2019.2925812","article-title":"Window Zooming\u2013Based Localization Algorithm of Fruit and Vegetable for Harvesting Robot","volume":"7","author":"Wang","year":"2019","journal-title":"IEEE Access"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Pawara, P., Boshchenko, A., Schomaker, L., and Wiering, M.A. (2020, January 12\u201314). Deep Learning with Data Augmentation for Fruit Counting. Artificial Intelligence and Soft Computing. Proceedings of the Artificial Intelligence and Soft Computing: 19th International Conference, ICAISC 2020, Zakopane, Poland.","DOI":"10.1007\/978-3-030-61401-0_20"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"105714","DOI":"10.1016\/j.compag.2020.105714","article-title":"Invariant leaf image recognition with histogram of Gaussian convolution vectors","volume":"178","author":"Chen","year":"2020","journal-title":"Comput. Electron. Agric."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"626","DOI":"10.1016\/j.ijleo.2016.11.177","article-title":"A robust fruit image segmentation algorithm against varying illumination for vision system of fruit harvesting robot","volume":"131","author":"Wang","year":"2017","journal-title":"Optik"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"831","DOI":"10.1109\/70.897793","article-title":"Robotic melon harvesting","volume":"16","author":"Edan","year":"2000","journal-title":"IEEE Trans. Robot. Autom."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"296","DOI":"10.21273\/HORTTECH03965-18","article-title":"Hand and mechanical fruit-zone leaf removal at prebloom and fruit-set was more effective in reducing crop yield than reducing bunch rot in \u2018riesling\u2019 grapevines","volume":"28","author":"Hed","year":"2018","journal-title":"Horttechnology"},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"8","DOI":"10.1016\/j.compag.2015.05.021","article-title":"Sensors and systems for fruit detection and localization: A review","volume":"116","author":"Gongal","year":"2015","journal-title":"Comput. Electron. Agric."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"51","DOI":"10.1016\/j.biosystemseng.2013.07.007","article-title":"Identification and determination of the number of immature green citrus fruit in a canopy under different ambient light conditions","volume":"117","author":"Sengupta","year":"2014","journal-title":"Biosyst. Eng."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"7390","DOI":"10.1016\/j.eswa.2014.06.013","article-title":"Detecting corn tassels using computer vision and support vector machines","volume":"41","author":"Kavdir","year":"2014","journal-title":"Expert Syst. Appl."},{"key":"ref_10","unstructured":"Simonyan, K., and Zisserman, A. (2014). Very Deep Convolutional Networks for Large-Scale Image Recognition. arXiv."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep residual learning for image recognition. Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Girshick, R. (2015, January 7\u201313). Fast R-CNN. Proceedings of the 2015 IEEE International Conference on Computer Vision (ICCV), Santiago, Chile.","DOI":"10.1109\/ICCV.2015.169"},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"1137","DOI":"10.1109\/TPAMI.2016.2577031","article-title":"Faster R-CNN: Towards real-time object detection with region proposal networks","volume":"39","author":"Ren","year":"2017","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Sun, P., Zhang, R., Jiang, Y., Kong, T., Xu, C., Zhan, W., Tomizuka, M., Li, L., Yuan, Z., and Wang, C. (2021, January 19\u201325). Sparse r-cnn: End-to-end object detection with learnable proposals. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Virtual.","DOI":"10.1109\/CVPR46437.2021.01422"},{"key":"ref_15","unstructured":"Redmon, J., Divvala, S., Girshick, R., and Farhadi, A. (July, January 26). You only look once: Unified, real-time object detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Redmon, J., and Farhadi, A. (2017, January 21\u201326). YOLO9000: Better, faster, stronger. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.690"},{"key":"ref_17","unstructured":"Redmon, J., and Farhadi, A. (2018). Yolov3: An incremental improvement. arXiv."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Leibe, B., Matas, J., Sebe, N., and Welling, M. (2016). Computer Vision\u2014ECCV 2016, Springer. Lecture Notes in Computer Science.","DOI":"10.1007\/978-3-319-46487-9"},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"1395","DOI":"10.1007\/s00217-022-03971-7","article-title":"Fast olive quality assessment through RGB images and advanced convolutional neural network modeling","volume":"248","author":"Salvucci","year":"2022","journal-title":"Eur. Food Res. Technol."},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"105348","DOI":"10.1016\/j.compag.2020.105348","article-title":"Comparison of convolutional neural networks in fruit detection and counting: A comprehensive evaluation","volume":"173","author":"Vasconez","year":"2020","journal-title":"Comput. Electron. Agric."},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Bewley, A., Ge, Z., Ott, L., Ramos, F., and Upcroft, B. (2016, January 25\u201328). Simple online and realtime tracking. Proceedings of the 2016 IEEE International Conference on Image Processing (ICIP), Phoenix, AZ, USA.","DOI":"10.1109\/ICIP.2016.7533003"},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"83","DOI":"10.1002\/nav.3800020109","article-title":"The Hungarian method for the assignment problem","volume":"2","author":"Kuhn","year":"1955","journal-title":"Nav. Res. Logist. Q."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Wojke, N., Bewley, A., and Paulus, D. (2017, January 17\u201320). Simple online and realtime tracking with a deep association metric. Proceedings of the 2017 IEEE International Conference on Image Processing (ICIP), Beijing, China.","DOI":"10.1109\/ICIP.2017.8296962"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Wang, Z., Zheng, L., Liu, Y., and Wang, S. (2020, January 23\u201328). Towards Real-Time Multi-Object Tracking. Proceedings of the Computer Vision\u2013ECCV 2020: 16th European Conference, Glasgow, UK.","DOI":"10.1007\/978-3-030-58621-8_7"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Wang, C.Y., Bochkovskiy, A., and Liao HY, M. (2023, January 18\u201322). YOLOv7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada.","DOI":"10.1109\/CVPR52729.2023.00721"},{"key":"ref_26","first-page":"24","article-title":"Real-time visual inspection system for grading fruits using computer vision and deep learning techniques","volume":"9","author":"Ismail","year":"2021","journal-title":"Inf. Process. Agric."},{"key":"ref_27","unstructured":"(2017, May 03). Tzutalin: LabelImg Homepage. Available online: https:\/\/github.com\/tzutalin\/labelImg."},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Zhang, H., Cisse, M., Dauphin, Y.N., and Lopez-Paz, D. (2017). Mixup: Beyond empirical risk minimization. arXiv.","DOI":"10.1007\/978-1-4899-7687-1_79"},{"key":"ref_29","unstructured":"Bochkovskiy, A., Wang, C.Y., and Liao, H.Y.M. (2020). Yolov4: Optimal speed and accuracy of object detection. arXiv."},{"key":"ref_30","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, \u0141., and Polosukhin, I. (2023, May 21). Attention Is All You Need. Available online: https:\/\/proceedings.neurips.cc\/paper\/2017\/file\/3f5ee243547dee91fbd053c1c4a845aa-Paper.pdf."},{"key":"ref_31","unstructured":"Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., and Houlsby, N. (2020). An image is worth 16x16 words: Transformers for image recognition at scale. arXiv."},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Hu, J., Shen, L., and Sun, G. (2018, January 18\u201323). Squeeze-and-Excitation Networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00745"},{"key":"ref_33","unstructured":"Zhu, X., Cheng, D., Zhang, Z., Lin, S., and Dai, J. (November, January 27). An empirical study of spatial attention mechanisms in deep networks. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Seoul, Republic of Korea."},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Woo, S., Park, J., Lee, J.Y., and Kweon, I.S. (2018, January 8\u201314). Cbam: Convolutional block attention module. Proceedings of the European Conference on Computer Vision, Munich, Germany.","DOI":"10.1007\/978-3-030-01234-2_1"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Hou, Q., Zhou, D., and Feng, J. (2021, January 19\u201325). Coordinate attention for efficient mobile network design. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Virtual.","DOI":"10.1109\/CVPR46437.2021.01350"},{"key":"ref_36","doi-asserted-by":"crossref","first-page":"107223","DOI":"10.1016\/j.compag.2022.107223","article-title":"Cascade-SORT: A robust fruit counting approach using multiple features cascade matching","volume":"200","author":"He","year":"2022","journal-title":"Comput. Electron. Agric."},{"key":"ref_37","doi-asserted-by":"crossref","first-page":"35","DOI":"10.1115\/1.3662552","article-title":"A New Approach to Linear Filtering and Prediction Problems","volume":"82","author":"Kalman","year":"1960","journal-title":"J. Basic Eng."},{"key":"ref_38","unstructured":"Gennari, M., Fawcett, R., and Prisacariu, V.A. (November, January 27). DSConv: Efficient Convolution Operator. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Seoul, Republic of Korea."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/13\/5903\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T20:00:40Z","timestamp":1760126440000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/13\/5903"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,6,25]]},"references-count":38,"journal-issue":{"issue":"13","published-online":{"date-parts":[[2023,7]]}},"alternative-id":["s23135903"],"URL":"https:\/\/doi.org\/10.3390\/s23135903","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,6,25]]}}}