{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,16]],"date-time":"2026-07-16T14:10:57Z","timestamp":1784211057963,"version":"3.55.0"},"reference-count":33,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2024,4,1]],"date-time":"2024-04-01T00:00:00Z","timestamp":1711929600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2024,4,12]],"date-time":"2024-04-12T00:00:00Z","timestamp":1712880000000},"content-version":"vor","delay-in-days":11,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J Comput Soc Sc"],"published-print":{"date-parts":[[2024,4]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>The retail industry encounters huge obstacles with computer vision (CV) technology due to frequent model retraining with changing products and time-consuming, costly data annotation. Previous research in this field has been primarily focused on optimizing model performance rather than minimizing annotation effort. Therefore, the main idea of this paper is to evaluate active learning as a method to minimize annotation effort in the retail industry. The MVTEC Densely Segmented Supermarket dataset is used to evaluate various active learning methods such as the Least Confident, Entropy and Cost-Effective Active Learning (CEAL) along with Mask R-CNN model. The results demonstrate that annotating only 20.83<jats:inline-formula><jats:alternatives><jats:tex-math>$$-$$<\/jats:tex-math><mml:math xmlns:mml=\"http:\/\/www.w3.org\/1998\/Math\/MathML\">\n                  <mml:mo>-<\/mml:mo>\n                <\/mml:math><\/jats:alternatives><\/jats:inline-formula>24.34% of the data achieves 95% of the full dataset\u2019s performance. When training, out-of-sample data share similar characteristics, the Least Confident and CEAL methods reduce annotation requirements by 7.7<jats:inline-formula><jats:alternatives><jats:tex-math>$$-$$<\/jats:tex-math><mml:math xmlns:mml=\"http:\/\/www.w3.org\/1998\/Math\/MathML\">\n                  <mml:mo>-<\/mml:mo>\n                <\/mml:math><\/jats:alternatives><\/jats:inline-formula>15.7% while maintaining 95% and 97% of the full dataset\u2019s performance. However, the Entropy method under-performs compared to the random selection baseline. Ultimately, none of the methods show a clear advantage when the data characteristics differ between training and out-of-sample data. Finally, the proposed active learning methods on an industry-specific retail dataset remarkably propels the development of highly efficient and cost-effective CV solutions meticulously tailored for the retail industry.<\/jats:p>","DOI":"10.1007\/s42001-024-00266-7","type":"journal-article","created":{"date-parts":[[2024,4,12]],"date-time":"2024-04-12T12:02:02Z","timestamp":1712923322000},"page":"1039-1070","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":4,"title":["Efficient annotation reduction with active learning for computer vision-based Retail Product Recognition"],"prefix":"10.1007","volume":"7","author":[{"given":"Niels","family":"Griffioen","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9910-5886","authenticated-orcid":false,"given":"Nevena","family":"Rankovic","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Federico","family":"Zamberlan","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Monisha","family":"Punith","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2024,4,12]]},"reference":[{"key":"266_CR1","doi-asserted-by":"publisher","unstructured":"Wei, Y., Tran, S., Xu, S., Kang, B., & Springer, M. (2020). Deep learning for retail product recognition: Challenges and techniques. Computational Intelligence and Neuroscience. 2020: 1\u201323. https:\/\/doi.org\/10.1155\/2020\/8875910","DOI":"10.1155\/2020\/8875910"},{"key":"266_CR2","doi-asserted-by":"publisher","DOI":"10.1007\/s11432-022-3513-y","author":"X-S Wei","year":"2022","unstructured":"Wei, X.-S., Cui, Q., Yang, L., Wang, P., Liu, L., & Yang, J. (2022). RPC: a large-scale and fine-grained retail product checkout dataset. Science China Information Sciences. https:\/\/doi.org\/10.1007\/s11432-022-3513-y","journal-title":"Science China Information Sciences"},{"issue":"3","key":"266_CR3","doi-asserted-by":"publisher","first-page":"177","DOI":"10.1561\/0600000071","volume":"10","author":"A Kovashka","year":"2016","unstructured":"Kovashka, A., Russakovsky, O., Fei-Fei, L., & Grauman, K. (2016). Crowdsourcing in computer vision. Foundations and Trends. Computer Graphics and Vision, 10(3), 177\u2013243. https:\/\/doi.org\/10.1561\/0600000071","journal-title":"Computer Graphics and Vision"},{"key":"266_CR4","doi-asserted-by":"publisher","first-page":"45","DOI":"10.1016\/j.imavis.2019.03.005","volume":"86","author":"B Santra","year":"2019","unstructured":"Santra, B., & Mukherjee, D. P. (2019). A comprehensive survey on computer vision based approaches for automatic identification of products in retail store. Image and Vision Computing., 86, 45\u201363. https:\/\/doi.org\/10.1016\/j.imavis.2019.03.005","journal-title":"Image and Vision Computing."},{"issue":"6","key":"266_CR5","doi-asserted-by":"publisher","first-page":"2902","DOI":"10.3390\/app12062902","volume":"12","author":"C-H Hsia","year":"2022","unstructured":"Hsia, C.-H., Chang, T.-H.W., Chiang, C.-Y., & Chan, H.-T. (2022). Mask r-CNN with new data augmentation features for smart detection of retail products. Applied Sciences, 12(6), 2902. https:\/\/doi.org\/10.3390\/app12062902","journal-title":"Applied Sciences"},{"key":"266_CR6","doi-asserted-by":"publisher","unstructured":"Bartl, V., Spanhel, J., & Herout, A. (2022). PersonGONE: Image inpainting for automated checkout solution. In: 2022 IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW). pp. 3115\u20133123. https:\/\/doi.org\/10.1109\/cvprw56347.2022.00351","DOI":"10.1109\/cvprw56347.2022.00351"},{"key":"266_CR7","doi-asserted-by":"publisher","unstructured":"Fuchs, K., Grundmann, T., & Fleisch, E. (2019). Towards identification of packaged products via computer vision. In: Proceedings of the 9th International Conference on the Internet of Things. https:\/\/doi.org\/10.1145\/3365871.3365899","DOI":"10.1145\/3365871.3365899"},{"issue":"11","key":"266_CR8","doi-asserted-by":"publisher","first-page":"2278","DOI":"10.1109\/5.726791","volume":"86","author":"Y Lecun","year":"1998","unstructured":"Lecun, Y., Bottou, L., Bengio, Y., & Haffner, P. (1998). Gradient-based learning applied to document recognition. Proceedings of the IEEE, 86(11), 2278\u20132324. https:\/\/doi.org\/10.1109\/5.726791","journal-title":"Proceedings of the IEEE"},{"issue":"11","key":"266_CR9","doi-asserted-by":"publisher","first-page":"3212","DOI":"10.1109\/tnnls.2018.2876865","volume":"30","author":"Z-Q Zhao","year":"2019","unstructured":"Zhao, Z.-Q., Zheng, P., Xu, S.-T., & Wu, X. (2019). Object detection with deep learning: A review. IEEE Transactions on Neural Networks and Learning Systems, 30(11), 3212\u20133232. https:\/\/doi.org\/10.1109\/tnnls.2018.2876865","journal-title":"IEEE Transactions on Neural Networks and Learning Systems"},{"issue":"2","key":"266_CR10","doi-asserted-by":"publisher","first-page":"154","DOI":"10.1007\/s11263-013-0620-5","volume":"104","author":"JRR Uijlings","year":"2013","unstructured":"Uijlings, J. R. R., Sande, K. E. A., Gevers, T., & Smeulders, A. W. M. (2013). Selective search for object recognition. International Journal of Computer Vision, 104(2), 154\u2013171. https:\/\/doi.org\/10.1007\/s11263-013-0620-5","journal-title":"International Journal of Computer Vision"},{"issue":"6","key":"266_CR11","doi-asserted-by":"publisher","first-page":"1137","DOI":"10.1109\/tpami.2016.2577031","volume":"39","author":"S Ren","year":"2017","unstructured":"Ren, S., He, K., Girshick, R., & Sun, J. (2017). Faster r-CNN: Towards real-time object detection with region proposal networks. IEEE Transactions on Pattern Analysis and Machine Intelligence, 39(6), 1137\u20131149. https:\/\/doi.org\/10.1109\/tpami.2016.2577031","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"266_CR12","doi-asserted-by":"publisher","unstructured":"He, K., Gkioxari, G., Dollar, P., & Girshick, R. (2017). Mask r-CNN. In: 2017 IEEE International Conference on Computer Vision (ICCV). pp. 2961\u20132969. https:\/\/doi.org\/10.1109\/iccv.2017.322","DOI":"10.1109\/iccv.2017.322"},{"key":"266_CR13","doi-asserted-by":"publisher","unstructured":"Follmann, P., B\u00f6ttger, T., H\u00e4rtinger, P., K\u00f6nig, R., & Ulrich, M. (2018). MVTec d2s: Densely segmented supermarket dataset. In: Computer Vision \u2013 ECCV 2018. pp. 581\u2013597. https:\/\/doi.org\/10.1007\/978-3-030-01249-6_35","DOI":"10.1007\/978-3-030-01249-6_35"},{"issue":"4","key":"266_CR14","doi-asserted-by":"publisher","first-page":"640","DOI":"10.1109\/tpami.2016.2572683","volume":"39","author":"E Shelhamer","year":"2017","unstructured":"Shelhamer, E., Long, J., & Darrell, T. (2017). Fully convolutional networks for semantic segmentation. IEEE Transactions on Pattern Analysis and Machine Intelligence, 39(4), 640\u2013651. https:\/\/doi.org\/10.1109\/tpami.2016.2572683","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"266_CR15","doi-asserted-by":"publisher","unstructured":"Bai, M., & Urtasun, R. (2017). Deep watershed transform for instance segmentation. In: 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). pp. 5221\u20135229. https:\/\/doi.org\/10.1109\/cvpr.2017.305","DOI":"10.1109\/cvpr.2017.305"},{"key":"266_CR16","doi-asserted-by":"publisher","DOI":"10.1016\/j.imavis.2022.104571","volume":"128","author":"S Chen","year":"2022","unstructured":"Chen, S., Liu, D., Pu, Y., & Zhong, Y. (2022). Advances in deep learning-based image recognition of product packaging. Image and Vision Computing., 128, 104571.","journal-title":"Image and Vision Computing."},{"key":"266_CR17","doi-asserted-by":"publisher","first-page":"292","DOI":"10.1016\/j.neucom.2021.07.040","volume":"461","author":"K Hameed","year":"2021","unstructured":"Hameed, K., Chai, D., & Rassau, A. (2021). Class distribution-aware adaptive margins and cluster embedding for classification of fruit and vegetables at supermarket self-checkouts. Neurocomputing, 461, 292\u2013309.","journal-title":"Neurocomputing"},{"key":"266_CR18","doi-asserted-by":"publisher","unstructured":"Karlinsky, L., Shtok, J., Tzur, Y., & Tzadok, A. (2017). Fine-grained recognition of thousands of object categories with single-example training. In: 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). https:\/\/doi.org\/10.1109\/cvpr.2017.109","DOI":"10.1109\/cvpr.2017.109"},{"key":"266_CR19","doi-asserted-by":"crossref","unstructured":"Budd, S., Robinson, E. C., & Kainz, B. (2021). A survey on active learning and human-in-the-loop deep learning for medical image analysis. Medical Image Analysis, 71: 102062.","DOI":"10.1016\/j.media.2021.102062"},{"key":"266_CR20","doi-asserted-by":"publisher","unstructured":"Settles, B. (2012). Active Learning. https:\/\/doi.org\/10.1007\/978-3-031-01560-1","DOI":"10.1007\/978-3-031-01560-1"},{"issue":"12","key":"266_CR21","doi-asserted-by":"publisher","first-page":"2591","DOI":"10.1109\/tcsvt.2016.2589879","volume":"27","author":"K Wang","year":"2017","unstructured":"Wang, K., Zhang, D., Li, Y., Zhang, R., & Lin, L. (2017). Cost-effective active learning for deep image classification. IEEE Transactions on Circuits and Systems for Video Technology, 27(12), 2591\u20132600. https:\/\/doi.org\/10.1109\/tcsvt.2016.2589879","journal-title":"IEEE Transactions on Circuits and Systems for Video Technology"},{"key":"266_CR22","unstructured":"Kirsch, A., Van\u00a0Amersfoort, J., Gal, Y. (2019). Batchbald: Efficient and diverse batch acquisition for deep bayesian active learning. Advances in Neural Information Processing Systems, 32"},{"issue":"9","key":"266_CR23","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3472291","volume":"54","author":"P Ren","year":"2021","unstructured":"Ren, P., Xiao, Y., Chang, X., Huang, P.-Y., Li, Z., Gupta, B. B., Chen, X., & Wang, X. (2021). A survey of deep active learning. ACM computing surveys (CSUR), 54(9), 1\u201340.","journal-title":"ACM computing surveys (CSUR)"},{"key":"266_CR24","doi-asserted-by":"crossref","unstructured":"Kim, K., Park, D., Kim, K.I., & Chun, S.Y. (2021). Task-aware variational adversarial active learning. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. pp. 8166\u20138175","DOI":"10.1109\/CVPR46437.2021.00807"},{"key":"266_CR25","first-page":"11933","volume":"34","author":"G Citovsky","year":"2021","unstructured":"Citovsky, G., DeSalvo, G., Gentile, C., Karydas, L., Rajagopalan, A., Rostamizadeh, A., & Kumar, S. (2021). Batch active learning at scale. Advances in Neural Information Processing Systems, 34, 11933\u201311944.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"266_CR26","doi-asserted-by":"publisher","unstructured":"Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Doll\u00e1r, P., & Zitnick, C.L. (2014). Microsoft COCO: Common objects in context. In: Computer Vision \u2013 ECCV 2014. pp. 740\u2013755. https:\/\/doi.org\/10.1007\/978-3-319-10602-1_48","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"266_CR27","doi-asserted-by":"publisher","unstructured":"Cordts, M., Omran, M., Ramos, S., Rehfeld, T., Enzweiler, M., Benenson, R., Franke, U., Roth, S., & Schiele, B. (2016). The cityscapes dataset for semantic urban scene understanding. In: 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). https:\/\/doi.org\/10.1109\/cvpr.2016.350","DOI":"10.1109\/cvpr.2016.350"},{"key":"266_CR28","doi-asserted-by":"publisher","unstructured":"Culotta, A., & McCallum, A. (jan 2005). Reducing labeling effort for structured prediction tasks. Technical Report. https:\/\/doi.org\/10.21236\/ada440382","DOI":"10.21236\/ada440382"},{"key":"266_CR29","doi-asserted-by":"publisher","unstructured":"Settles, B., & Craven, M. (2008). An analysis of active learning strategies for sequence labeling tasks. In: Proceedings of the Conference on Empirical Methods in Natural Language Processing\u2013EMNLP \u201908. https:\/\/doi.org\/10.3115\/1613715.1613855","DOI":"10.3115\/1613715.1613855"},{"key":"266_CR30","unstructured":"Lafferty, J.D., McCallum, A., & Pereira, F.C.N. (2001). Conditional random fields: Probabilistic models for segmenting and labeling sequence data. In: Proceedings of the Eighteenth International Conference on Machine Learning. ICML \u201901, pp. 282\u2013289. Morgan Kaufmann Publishers Inc., San Francisco, CA, USA."},{"issue":"1","key":"266_CR31","doi-asserted-by":"publisher","first-page":"98","DOI":"10.1007\/s11263-014-0733-5","volume":"111","author":"M Everingham","year":"2014","unstructured":"Everingham, M., Eslami, S. M. A., Gool, L. V., Williams, C. K. I., Winn, J., & Zisserman, A. (2014). The pascal visual object classes challenge: A retrospective. International Journal of Computer Vision, 111(1), 98\u2013136. https:\/\/doi.org\/10.1007\/s11263-014-0733-5","journal-title":"International Journal of Computer Vision"},{"key":"266_CR32","doi-asserted-by":"publisher","unstructured":"Hu, T., Deng, Y., Deng, Y., & Ge, A. (2021). Fully convolutional network variations and method on small dataset. In: 2021 IEEE International Conference on Consumer Electronics and Computer Engineering (ICCECE). pp. 40\u201346. https:\/\/doi.org\/10.1109\/iccece51280.2021.9342059","DOI":"10.1109\/iccece51280.2021.9342059"},{"key":"266_CR33","doi-asserted-by":"publisher","unstructured":"He, K., Zhang, X., Ren, S., & Sun, J. (2016). Deep residual learning for image recognition. In: 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). pp. 770\u2013778. https:\/\/doi.org\/10.1109\/cvpr.2016.90","DOI":"10.1109\/cvpr.2016.90"}],"container-title":["Journal of Computational Social Science"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s42001-024-00266-7.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s42001-024-00266-7\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s42001-024-00266-7.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,7,16]],"date-time":"2024-07-16T04:07:46Z","timestamp":1721102866000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s42001-024-00266-7"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,4]]},"references-count":33,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2024,4]]}},"alternative-id":["266"],"URL":"https:\/\/doi.org\/10.1007\/s42001-024-00266-7","relation":{},"ISSN":["2432-2717","2432-2725"],"issn-type":[{"value":"2432-2717","type":"print"},{"value":"2432-2725","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,4]]},"assertion":[{"value":"27 September 2023","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"27 February 2024","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"12 April 2024","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare no conflict of interest.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}},{"value":"Not applicable.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethical approval"}},{"value":"Not applicable.","order":4,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent to participate"}},{"value":"All authors have read and agreed to the published version of the manuscript.","order":5,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent for publication"}}]}}