{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,8,30]],"date-time":"2025-08-30T16:56:28Z","timestamp":1756572988621,"version":"3.41.2"},"reference-count":41,"publisher":"Frontiers Media SA","license":[{"start":{"date-parts":[[2025,2,7]],"date-time":"2025-02-07T00:00:00Z","timestamp":1738886400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["frontiersin.org"],"crossmark-restriction":true},"short-container-title":["Front. Neurorobot."],"abstract":"<jats:sec><jats:title>Introduction<\/jats:title><jats:p>To enhance the detection of litchi fruits in natural scenes, address challenges such as dense occlusion and small target identification, this paper proposes a novel multimodal target detection method, denoted as YOLOv5-Litchi.<\/jats:p><\/jats:sec><jats:sec><jats:title>Methods<\/jats:title><jats:p>Initially, the Neck layer network of YOLOv5s is simplified by changing its FPN+PAN structure to an FPN structure and increasing the number of detection heads from 3 to 5. Additionally, the detection heads with resolutions of 80 \u00d7 80 pixels and 160 \u00d7 160 pixels are replaced by TSCD detection heads to enhance the model's ability to detect small targets. Subsequently, the positioning loss function is replaced with the EIoU loss function, and the confidence loss is substituted by VFLoss to further improve the accuracy of the detection bounding box and reduce the missed detection rate in occluded targets. A sliding slice method is then employed to predict image targets, thereby reducing the miss rate of small targets.<\/jats:p><\/jats:sec><jats:sec><jats:title>Results<\/jats:title><jats:p>Experimental results demonstrate that the proposed model improves accuracy, recall, and mean average precision (mAP) by 9.5, 0.9, and 12.3 percentage points, respectively, compared to the original YOLOv5s model. When benchmarked against other models such as YOLOx, YOLOv6, and YOLOv8, the proposed model's AP value increases by 4.0, 6.3, and 3.7 percentage points, respectively.<\/jats:p><\/jats:sec><jats:sec><jats:title>Discussion<\/jats:title><jats:p>The improved network exhibits distinct improvements, primarily focusing on enhancing the recall rate and AP value, thereby reducing the missed detection rate which exhibiting a reduced number of missed targets and a more accurate prediction frame, indicating its suitability for litchi fruit detection. Therefore, this method significantly enhances the detection accuracy of mature litchi fruits and effectively addresses the challenges of dense occlusion and small target detection, providing crucial technical support for subsequent litchi yield estimation.<\/jats:p><\/jats:sec>","DOI":"10.3389\/fnbot.2024.1518878","type":"journal-article","created":{"date-parts":[[2025,2,7]],"date-time":"2025-02-07T06:50:14Z","timestamp":1738911014000},"update-policy":"https:\/\/doi.org\/10.3389\/crossmark-policy","source":"Crossref","is-referenced-by-count":2,"title":["A scalable multi-modal learning fruit detection algorithm for dynamic environments"],"prefix":"10.3389","volume":"18","author":[{"given":"Liang","family":"Mao","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zihao","family":"Guo","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Mingzhe","family":"Liu","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yue","family":"Li","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Linlin","family":"Wang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jie","family":"Li","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1965","published-online":{"date-parts":[[2025,2,7]]},"reference":[{"key":"B1","first-page":"1","article-title":"\u201cMultimodal machine learning for pedestrian detection,\u201d","volume-title":"2021 IEEE 93rd Vehicular Technology Conference (VTC2021-Spring)","author":"Aledhari","year":"2021"},{"key":"B2","doi-asserted-by":"publisher","first-page":"102217","DOI":"10.1016\/j.ecoinf.2023.102217","article-title":"A review of deep learning techniques used in agriculture","volume":"2023","author":"Attri","year":"2023","journal-title":"Ecol. Informat"},{"key":"B3","doi-asserted-by":"publisher","first-page":"14804","DOI":"10.1109\/ACCESS.2023.3243854","article-title":"A systematic literature review on multimodal machine learning: applications, challenges, gaps and future directions","volume":"11","author":"Barua","year":"2023","journal-title":"IEEE Access"},{"key":"B4","doi-asserted-by":"publisher","first-page":"692","DOI":"10.1109\/TEVC.2017.2744328","article-title":"Evolutionary multiobjective optimization-based multimodal optimization: fitness landscape approximation and peak detection","volume":"22","author":"Cheng","year":"2017","journal-title":"IEEE Trans. Evol. Comput"},{"key":"B5","doi-asserted-by":"publisher","first-page":"1365","DOI":"10.3233\/JIFS-213251","article-title":"Improved YOLO object detection algorithm to detect ripe pineapple phase","volume":"43","author":"Cuong","year":"2022","journal-title":"J. Intell. Fuzzy Syst"},{"key":"B6","doi-asserted-by":"publisher","first-page":"9675628","DOI":"10.1155\/2022\/9675628","article-title":"An improved image classification method for cervical precancerous lesions based on shufflenet","volume":"2022","author":"Fang","year":"2022","journal-title":"Comput. Intell. Neurosci"},{"key":"B7","first-page":"21056","article-title":"\u201cDeep residual learning in spiking neural networks,\u201d","author":"Fang","year":"2021","journal-title":"Advances in Neural Information Processing Systems 34: NeurIPS 2021, December 6-14, 2021, Virtual"},{"key":"B8","doi-asserted-by":"publisher","first-page":"63373","DOI":"10.1109\/ACCESS.2019.2916887","article-title":"Deep multimodal representation learning: a survey","volume":"7","author":"Guo","year":"2019","journal-title":"IEEE Access"},{"key":"B9","doi-asserted-by":"publisher","first-page":"1588","DOI":"10.1109\/TNNLS.2021.3105602","article-title":"Graph fusion network-based multimodal learning for freezing of gait detection","volume":"34","author":"Hu","year":"2021","journal-title":"IEEE Trans. Neural Netw. Learn. Syst"},{"key":"B10","doi-asserted-by":"publisher","first-page":"446","DOI":"10.3390\/rs11040446","article-title":"Fusing multimodal video data for detecting moving objects\/targets in challenging indoor and outdoor scenes","volume":"11","author":"Kandylakis","year":"2019","journal-title":"Rem. Sens"},{"key":"B11","doi-asserted-by":"publisher","first-page":"56","DOI":"10.1016\/j.neunet.2017.12.005","article-title":"Stdp-based spiking deep convolutional neural networks for object recognition","volume":"99","author":"Kheradpisheh","year":"2018","journal-title":"Neural Netw"},{"key":"B12","doi-asserted-by":"publisher","first-page":"219","DOI":"10.1016\/j.compag.2019.04.017","article-title":"Deep learning\u2013method overview and review of use for fruit detection and yield estimation","volume":"162","author":"Koirala","year":"2019","journal-title":"Comput. Electron. Agricult"},{"key":"B13","doi-asserted-by":"publisher","first-page":"104628","DOI":"10.1016\/j.imavis.2023.104628","article-title":"Intelligent multimodal pedestrian detection using hybrid metaheuristic optimization with deep learning model","volume":"131","author":"Kolluri","year":"2023","journal-title":"Image Vis. Comput"},{"key":"B14","first-page":"1","article-title":"\u201cFruits and vegetables recognition using YOLO,\u201d","volume-title":"2022 International Conference on Computer Communication and Informatics (ICCCI)","author":"Latha","year":"2022"},{"key":"B15","doi-asserted-by":"publisher","DOI":"10.1002\/cpe.4523","article-title":"Verifiable chebyshev maps-based chaotic encryption schemes with outsourcing computations in the cloud\/fog scenarios","author":"Li","year":"2019","journal-title":"Concurr. Comput. Pract. Exp"},{"key":"B16","first-page":"3367","article-title":"\u201cRecurrent convolutional neural network for object recognition,\u201d","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Liang","year":"2015"},{"key":"B17","first-page":"2117","article-title":"\u201cFeature pyramid networks for object detection,\u201d","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Lin","year":"2017"},{"key":"B18","first-page":"8759","article-title":"\u201cPath aggregation network for instance segmentation,\u201d","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Liu","year":"2018"},{"key":"B19","doi-asserted-by":"publisher","first-page":"36516","DOI":"10.1109\/ACCESS.2019.2903826","article-title":"Secure remote sensing image registration based on compressed sensing in cloud setting","volume":"7","author":"Liu","year":"2019","journal-title":"IEEE Access"},{"key":"B20","first-page":"689","article-title":"\u201cMultimodal deep learning,\u201d","volume-title":"Proceedings of the 28th International Conference on Machine Learning (ICML-11)","author":"Ngiam","year":"2011"},{"key":"B21","doi-asserted-by":"publisher","first-page":"211","DOI":"10.25165\/j.ijabe.20221502.6541","article-title":"Litchi detection in the field using an improved YOLOv3 model","volume":"15","author":"Peng","year":"2022","journal-title":"Int. J. Agricult. Biol. Eng"},{"key":"B22","doi-asserted-by":"publisher","first-page":"203","DOI":"10.1016\/j.inffus.2021.12.003","article-title":"Multimodal co-learning: challenges, applications with datasets, recent advances and future directions","volume":"81","author":"Rahate","year":"2022","journal-title":"Inform. Fus"},{"key":"B23","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2202.06218","article-title":"Emotion based hate speech detection using multimodal learning","author":"Rana","year":"2022","journal-title":"arXiv preprint arXiv:2202.06218"},{"key":"B24","doi-asserted-by":"publisher","first-page":"1222","DOI":"10.3390\/s16081222","article-title":"Deepfruits: a fruit detection system using deep neural networks","volume":"16","author":"Sa","year":"2016","journal-title":"Sensors"},{"key":"B25","doi-asserted-by":"publisher","first-page":"2053","DOI":"10.1007\/s11119-021-09806-x","article-title":"Automation in agriculture by machine and deep learning techniques: a review of recent developments","volume":"22","author":"Saleem","year":"2021","journal-title":"Precis. Agricult"},{"key":"B26","doi-asserted-by":"publisher","first-page":"1497","DOI":"10.1109\/JSTARS.2020.3041316","article-title":"YOLOrs: object detection in multimodal remote sensing imagery","volume":"14","author":"Sharma","year":"2020","journal-title":"IEEE J. Select. Top. Appl. Earth Observ. Rem. Sens"},{"key":"B27","doi-asserted-by":"publisher","first-page":"569","DOI":"10.1016\/j.neuroimage.2014.06.077","article-title":"Hierarchical feature representation and multimodal fusion with deep learning for ad\/mci diagnosis","volume":"101","author":"Suk","year":"2014","journal-title":"NeuroImage"},{"key":"B28","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1007\/978-981-15-4288-6_1","article-title":"A review of object detection models based on convolutional neural network","volume":"1","author":"Sultana","year":"2020","journal-title":"Intell. Comput. Image Process. Bas. Appl"},{"key":"B29","doi-asserted-by":"publisher","first-page":"196","DOI":"10.3390\/agriculture8120196","article-title":"Detection of key organs in tomato based on deep migration learning in a complex background","volume":"8","author":"Sun","year":"2018","journal-title":"Agriculture"},{"key":"B30","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1016\/j.inpa.2019.09.006","article-title":"Computer vision technology in agricultural automation\u2014a review","volume":"7","author":"Tian","year":"2020","journal-title":"Inform. Process. Agricult"},{"key":"B31","doi-asserted-by":"publisher","first-page":"417","DOI":"10.1016\/j.compag.2019.01.012","article-title":"Apple detection during different growth stages in orchards using the improved YOLO-v3 model","volume":"157","author":"Tian","year":"2019","journal-title":"Comput. Electron. Agricult"},{"key":"B32","doi-asserted-by":"publisher","first-page":"9210947","DOI":"10.1155\/2022\/9210947","article-title":"Recent advancements in fruit detection and classification using deep learning techniques","volume":"2022","author":"Ukwuoma","year":"2022","journal-title":"Math. Probl. Eng"},{"key":"B33","doi-asserted-by":"publisher","first-page":"965425","DOI":"10.3389\/fpls.2022.965425","article-title":"Fast and precise detection of litchi fruits for yield estimation based on the improved YOLOv5 model","volume":"13","author":"Wang","year":"2022","journal-title":"Front. Plant Sci"},{"key":"B34","doi-asserted-by":"publisher","first-page":"294","DOI":"10.1016\/j.ins.2019.07.023","article-title":"Multilevel similarity model for high-resolution remote sensing image registration","volume":"505","author":"Wang","year":"2019","journal-title":"Inform. Sci"},{"key":"B35","doi-asserted-by":"publisher","first-page":"111808","DOI":"10.1016\/j.postharvbio.2021.111808","article-title":"Apple stem\/calyx real-time recognition using YOLO-v5 algorithm for fruit automatic loading system","volume":"185","author":"Wang","year":"2022","journal-title":"Postharv. Biol. Technol"},{"key":"B36","doi-asserted-by":"publisher","first-page":"3783","DOI":"10.3390\/s24123783","article-title":"EMA-YOLO: A novel target-detection algorithm for immature yellow peach based on YOLOv8","volume":"24","author":"Xu","year":"2024","journal-title":"Sensors"},{"key":"B37","doi-asserted-by":"crossref","first-page":"303","DOI":"10.1109\/ECTI-CON51831.2021.9454904","article-title":"\u201cFig fruit recognition method based on YOLO v4 deep learning,\u201d","volume-title":"2021 18th International Conference on Electrical Engineering\/Electronics, Computer, Telecommunications and Information Technology (ECTI-CON)","author":"Yijing","year":"2021"},{"key":"B38","doi-asserted-by":"publisher","first-page":"478","DOI":"10.1109\/JSTSP.2020.2987728","article-title":"Multimodal intelligence: representation learning, information fusion, and applications","volume":"14","author":"Zhang","year":"2020","journal-title":"IEEE J. Select. Top. Sign. Process"},{"key":"B39","first-page":"6848","article-title":"\u201cShuffleNet: an extremely efficient convolutional neural network for mobile devices,\u201d","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Zhang","year":"2018"},{"key":"B40","doi-asserted-by":"crossref","first-page":"6655","DOI":"10.1145\/3637528.3671462","article-title":"\u201cA survey on safe multi-modal learning systems,\u201d","volume-title":"Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining","author":"Zhao","year":"2024"},{"key":"B41","doi-asserted-by":"publisher","first-page":"1514","DOI":"10.3390\/math12101514","article-title":"Enhancing emergency vehicle detection: a deep learning approach with multimodal fusion","volume":"12","author":"Zohaib","year":"2024","journal-title":"Mathematics"}],"container-title":["Frontiers in Neurorobotics"],"original-title":[],"link":[{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/fnbot.2024.1518878\/full","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,2,7]],"date-time":"2025-02-07T06:50:22Z","timestamp":1738911022000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/fnbot.2024.1518878\/full"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,2,7]]},"references-count":41,"alternative-id":["10.3389\/fnbot.2024.1518878"],"URL":"https:\/\/doi.org\/10.3389\/fnbot.2024.1518878","relation":{},"ISSN":["1662-5218"],"issn-type":[{"type":"electronic","value":"1662-5218"}],"subject":[],"published":{"date-parts":[[2025,2,7]]},"article-number":"1518878"}}