{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,23]],"date-time":"2026-07-23T19:57:50Z","timestamp":1784836670681,"version":"3.55.0"},"reference-count":38,"publisher":"MDPI AG","issue":"8","license":[{"start":{"date-parts":[[2016,8,3]],"date-time":"2016-08-03T00:00:00Z","timestamp":1470182400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001793","name":"Queensland University of Technology","doi-asserted-by":"publisher","id":[{"id":"10.13039\/501100001793","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>This paper presents a novel approach to fruit detection using deep convolutional neural networks. The aim is to build an accurate, fast and reliable fruit detection system, which is a vital element of an autonomous agricultural robotic platform; it is a key element for fruit yield estimation and automated harvesting. Recent work in deep neural networks has led to the development of a state-of-the-art object detector termed Faster Region-based CNN (Faster R-CNN). We adapt this model, through transfer learning, for the task of fruit detection using imagery obtained from two modalities: colour (RGB) and Near-Infrared (NIR). Early and late fusion methods are explored for combining the multi-modal (RGB and NIR) information. This leads to a novel multi-modal Faster R-CNN model, which achieves state-of-the-art results compared to prior work with the F1 score, which takes into account both precision and recall performances improving from     0 . 807     to     0 . 838     for the detection of sweet pepper. In addition to improved accuracy, this approach is also much quicker to deploy for new fruits, as it requires bounding box annotation rather than pixel-level annotation (annotating bounding boxes is approximately an order of magnitude quicker to perform). The model is retrained to perform the detection of seven fruits, with the entire process taking four hours to annotate and train the new model per fruit.<\/jats:p>","DOI":"10.3390\/s16081222","type":"journal-article","created":{"date-parts":[[2016,8,3]],"date-time":"2016-08-03T10:10:31Z","timestamp":1470219031000},"page":"1222","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":954,"title":["DeepFruits: A Fruit Detection System Using Deep Neural Networks"],"prefix":"10.3390","volume":"16","author":[{"given":"Inkyu","family":"Sa","sequence":"first","affiliation":[{"name":"Science and Engineering Faculty, Queensland University of Technology, Brisbane 4000, Australia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Zongyuan","family":"Ge","sequence":"additional","affiliation":[{"name":"Science and Engineering Faculty, Queensland University of Technology, Brisbane 4000, Australia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Feras","family":"Dayoub","sequence":"additional","affiliation":[{"name":"Science and Engineering Faculty, Queensland University of Technology, Brisbane 4000, Australia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ben","family":"Upcroft","sequence":"additional","affiliation":[{"name":"Science and Engineering Faculty, Queensland University of Technology, Brisbane 4000, Australia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Tristan","family":"Perez","sequence":"additional","affiliation":[{"name":"Science and Engineering Faculty, Queensland University of Technology, Brisbane 4000, Australia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Chris","family":"McCool","sequence":"additional","affiliation":[{"name":"Science and Engineering Faculty, Queensland University of Technology, Brisbane 4000, Australia"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2016,8,3]]},"reference":[{"key":"ref_1","unstructured":"ABARE (2015). Australian Vegetable Growing Farms: An Economic Survey, 2013\u201314 and 2014\u201315, Research report."},{"key":"ref_2","unstructured":"Kondo, N., Monta, M., and Noguchi, N. (2011). Agricultural Robots: Mechanisms and Practice, Trans Pacific Press."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"888","DOI":"10.1002\/rob.21525","article-title":"Harvesting Robots for High-Value Crops: State-of-the-Art Review and Challenges Ahead","volume":"31","author":"Bac","year":"2014","journal-title":"J. Field Robot."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"McCool, C., Sa, I., Dayoub, F., Lehnert, C., Perez, T., and Upcroft, B. (2016, January 16\u201321). Visual Detection of Occluded Crop: For automated harvesting. Proceedings of the International Conference on Robotics and Automation, Stockholm, Sweden.","DOI":"10.1109\/ICRA.2016.7487405"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"211","DOI":"10.1007\/s11263-015-0816-y","article-title":"Imagenet large scale visual recognition challenge","volume":"115","author":"Russakovsky","year":"2015","journal-title":"Int. J. Comput. Vis."},{"key":"ref_6","unstructured":"Ge, Z.Y., and Sa, I. Open datasets and tutorial documentation. Available online: http:\/\/goo.gl\/9LmmOU."},{"key":"ref_7","unstructured":"Wikipedia F1 Score. Available online: https:\/\/en.wikipedia.org\/wiki\/F1_score."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Nuske, S.T., Achar, S., Bates, T., Narasimhan, S.G., and Singh, S. (2011, January 25\u201330). Yield Estimation in Vineyards by Visual Grape Detection. Proceedings of the 2011 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS \u201911), San Francisco, CA, USA.","DOI":"10.1109\/IROS.2011.6048830"},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"837","DOI":"10.1002\/rob.21541","article-title":"Automated visual yield estimation in vineyards","volume":"31","author":"Nuske","year":"2014","journal-title":"J. Field Robot."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"12191","DOI":"10.3390\/s140712191","article-title":"On plant detection of intact tomato fruits using image analysis and machine learning methods","volume":"14","author":"Yamamoto","year":"2014","journal-title":"Sensors"},{"key":"ref_11","unstructured":"Wang, Q., Nuske, S.T., Bergerman, M., and Singh, S. (2012, January 17\u201322). Automated Crop Yield Estimation for Apple Orchards. Proceedings of the 13th Internation Symposium on Experimental Robotics (ISER 2012), Qu\u00e9bec City, QC, Canada."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"148","DOI":"10.1016\/j.compag.2013.05.004","article-title":"Robust pixel-based classification of obstacles for robotic harvesting of sweet-pepper","volume":"96","author":"Bac","year":"2013","journal-title":"Comput. Electron. Agric."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Hung, C., Nieto, J., Taylor, Z., Underwood, J., and Sukkarieh, S. (2013, January 3\u20137). Orchard fruit segmentation using multi-spectral feature learning. Proceedings of the 2013 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS), Tokyo, Japan.","DOI":"10.1109\/IROS.2013.6697125"},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"4","DOI":"10.1504\/IJCVR.2012.046419","article-title":"Computer vision for fruit harvesting robots-state of the art and challenges ahead","volume":"3","author":"Kapach","year":"2012","journal-title":"Int. J. Comput. Vis. Robot."},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"203","DOI":"10.1016\/j.biosystemseng.2013.12.008","article-title":"Automatic fruit recognition and counting from multiple images","volume":"118","author":"Song","year":"2014","journal-title":"Biosyst. Eng."},{"key":"ref_16","unstructured":"Simonyan, K., and Zisserman, A. (2014, January 8\u201313). Two-stream convolutional networks for action recognition in videos. Proceedings of the Advances in Neural Information Processing Systems, Montr\u00e9al, QC, Canada."},{"key":"ref_17","unstructured":"Krizhevsky, A., Sutskever, I., and Hinton, G.E. (2012, January 3\u20138). Imagenet classification with deep convolutional neural networks. Proceedings of the Advances in Neural Information Processing Systems, Tahoe City, CA, USA."},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"98","DOI":"10.1007\/s11263-014-0733-5","article-title":"The pascal visual object classes challenge: A retrospective","volume":"111","author":"Everingham","year":"2015","journal-title":"Int. J. Comput. Vis."},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"154","DOI":"10.1007\/s11263-013-0620-5","article-title":"Selective search for object recognition","volume":"104","author":"Uijlings","year":"2013","journal-title":"Int. J. Comput. Vis."},{"key":"ref_20","unstructured":"Zitnick, C.L., and Doll\u00e1r, P. (2014). Computer Vision\u2013ECCV 2014, Springer."},{"key":"ref_21","unstructured":"Ren, S., He, K., Girshick, R., and Sun, J. (2015, January 7\u201312). Faster R-CNN: Towards real-time object detection with region proposal networks. Proceedings of the Advances in Neural Information Processing Systems, Montr\u00e9al, QC, Canada."},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"1904","DOI":"10.1109\/TPAMI.2015.2389824","article-title":"Spatial pyramid pooling in deep convolutional networks for visual recognition","volume":"37","author":"He","year":"2015","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Girshick, R. (2015, January 13\u201316). Fast r-cnn. Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.169"},{"key":"ref_24","unstructured":"Ngiam, J., Khosla, A., Kim, M., Nam, J., Lee, H., and Ng, A.Y. (July, January 28). Multimodal deep learning. Proceedings of the 28th international conference on machine learning (ICML-11), Bellevue, WA, USA."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Eitel, A., Springenberg, J.T., Spinello, L., Riedmiller, M., and Burgard, W. (October, January 28). Multimodal deep learning for robust RGB-D object recognition. Proceedings of the 2015 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS), Hamburg, Germany.","DOI":"10.1109\/IROS.2015.7353446"},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"705","DOI":"10.1177\/0278364914549607","article-title":"Deep learning for detecting robotic grasps","volume":"34","author":"Lenz","year":"2015","journal-title":"Int. J. Robot. Res."},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"2454","DOI":"10.1109\/TPAMI.2013.31","article-title":"Learning graphical model parameters with approximate marginal inference","volume":"35","author":"Domke","year":"2013","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"971","DOI":"10.1109\/TPAMI.2002.1017623","article-title":"Multiresolution gray-scale and rotation invariant texture classification with local binary patterns","volume":"24","author":"Ojala","year":"2002","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_29","unstructured":"Dalal, N., and Triggs, B. (2005, January 25). Histograms of oriented gradients for human detection. Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR 2005), San Diego, CA, USA."},{"key":"ref_30","unstructured":"Simonyan, K., and Zisserman, A. Very Deep Convolutional Networks for Large-Scale Image Recognition. Available online: https:\/\/arxiv.org\/abs\/1409.1556."},{"key":"ref_31","unstructured":"Donahue, J., Jia, Y., Vinyals, O., Hoffman, J., Zhang, N., Tzeng, E., and Darrell, T. DeCAF: A Deep Convolutional Activation Feature for Generic Visual Recognition. Available online: https:\/\/arxiv.org\/abs\/1310.1531."},{"key":"ref_32","first-page":"2579","article-title":"Visualizing data using t-SNE","volume":"9","author":"Hinton","year":"2008","journal-title":"J. Mach. Learn. Res."},{"key":"ref_33","unstructured":"Zeiler, M.D., and Fergus, R. (2014). Computer Vision\u2013ECCV 2014, Springer."},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S., Anguelov, D., Erhan, D., Vanhoucke, V., and Rabinovich, A. (2015, January 7\u201312). Going deeper with convolutions. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298594"},{"key":"ref_35","unstructured":"Stanford University CS231n: Convolutional Neural Networks for Visual Recognition (2016). Available online: http:\/\/cs231n.github.io\/transfer-learning\/."},{"key":"ref_36","unstructured":"University of California, Berkeley Fine-Tuning CaffeNet for Style Recognition on Flickr Style Data (2016). Available online: http:\/\/caffe.berkeleyvision.org\/gathered\/examples\/finetune_flickr_style.html."},{"key":"ref_37","doi-asserted-by":"crossref","first-page":"283","DOI":"10.1007\/BF01469346","article-title":"Detecting salient blob-like image structures and their scales with a scale-space primal sketch: A method for focus-of-attention","volume":"11","author":"Lindeberg","year":"1993","journal-title":"Int. J. Comput. Vis."},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Razavian, A., Azizpour, H., Sullivan, J., and Carlsson, S. (2014, January 23\u201328). CNN features off-the-shelf: an astounding baseline for recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, Columbus, OH, USA.","DOI":"10.1109\/CVPRW.2014.131"}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/16\/8\/1222\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T19:27:39Z","timestamp":1760210859000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/16\/8\/1222"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2016,8,3]]},"references-count":38,"journal-issue":{"issue":"8","published-online":{"date-parts":[[2016,8]]}},"alternative-id":["s16081222"],"URL":"https:\/\/doi.org\/10.3390\/s16081222","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2016,8,3]]}}}