{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,9,30]],"date-time":"2025-09-30T00:11:12Z","timestamp":1759191072882,"version":"3.44.0"},"reference-count":67,"publisher":"MDPI AG","issue":"10","license":[{"start":{"date-parts":[[2025,9,27]],"date-time":"2025-09-27T00:00:00Z","timestamp":1758931200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Young Doctoral Research Initiation Fund Project of Harbin University","award":["HUDF2022110"],"award-info":[{"award-number":["HUDF2022110"]}]},{"DOI":"10.13039\/501100005046","name":"Natural Science Foundation of Heilongjiang Province","doi-asserted-by":"publisher","award":["LH2024F047"],"award-info":[{"award-number":["LH2024F047"]}],"id":[{"id":"10.13039\/501100005046","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Harbin Science and Technology Plan","award":["ZC2022ZJ010027"],"award-info":[{"award-number":["ZC2022ZJ010027"]}]}],"content-domain":{"domain":["www.mdpi.com"],"crossmark-restriction":true},"short-container-title":["J. Imaging"],"abstract":"<jats:p>The Automatic Checkout (ACO) task aims to accurately generate complete shopping lists from checkout images. Severe product occlusions, numerous categories, and cluttered layouts impose high demands on detection models\u2019 robustness and generalization. To address these challenges, we propose the Edge-Embedded Multi-Feature Fusion Network (E2MF2Net), which jointly optimizes synthetic image generation and feature modeling. We introduce the Hierarchical Mask-Guided Composition (HMGC) strategy to select natural product poses based on mask compactness, incorporating geometric priors and occlusion tolerance to produce photorealistic, structurally coherent synthetic images. Mask-structure supervision further enhances boundary and spatial awareness. Architecturally, the Edge-Embedded Enhancement Module (E3) embeds salient structural cues to explicitly capture boundary details and facilitate cross-layer edge propagation, while the Multi-Feature Fusion Module (MFF) integrates multi-scale semantic cues, improving feature discriminability. Experiments on the RPC dataset demonstrate that E2MF2Net outperforms state-of-the-art methods, achieving checkout accuracy (cAcc) of 98.52%, 97.95%, 96.52%, and 97.62% on Easy, Medium, Hard, and Average mode, respectively. Notably, it improves by 3.63 percentage points in the heavily occluded Hard mode and exhibits strong robustness and adaptability in incremental learning and domain generalization scenarios.<\/jats:p>","DOI":"10.3390\/jimaging11100337","type":"journal-article","created":{"date-parts":[[2025,9,29]],"date-time":"2025-09-29T12:59:46Z","timestamp":1759150786000},"page":"337","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Edge-Embedded Multi-Feature Fusion Network for Automatic Checkout"],"prefix":"10.3390","volume":"11","author":[{"ORCID":"https:\/\/orcid.org\/0009-0008-8068-2772","authenticated-orcid":false,"given":"Jicai","family":"Li","sequence":"first","affiliation":[{"name":"College of Computer and Control Engineering, Northeast Forestry University, Harbin 150040, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0005-4326-6399","authenticated-orcid":false,"given":"Meng","family":"Zhu","sequence":"additional","affiliation":[{"name":"College of Information Engineering, Harbin University, Harbin 150076, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5334-7636","authenticated-orcid":false,"given":"Honge","family":"Ren","sequence":"additional","affiliation":[{"name":"College of Computer and Control Engineering, Northeast Forestry University, Harbin 150040, China"},{"name":"Heilongjiang Forestry Intelligent Equipment Engineering Research Center, Harbin 150040, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2025,9,27]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Redmon, J., and Farhadi, A. (2017, January 21\u201326). YOLO9000: Better, Faster, Stronger. Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.690"},{"key":"ref_2","unstructured":"Redmon, J., and Farhadi, A. (2018). YOLOv3: An Incremental Improvement. arXiv."},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Girshick, R. (2015, January 7\u201313). Fast R-CNN. Proceedings of the IEEE International Conference on Computer Vision (ICCV), Santiago, Chile.","DOI":"10.1109\/ICCV.2015.169"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"He, K., Gkioxari, G., Doll\u00c1R, P., and Girshick, R. (2017, January 22\u201329). Mask R-CNN. Proceedings of the 2017 IEEE International Conference on Computer Vision (ICCV), Venice, Italy.","DOI":"10.1109\/ICCV.2017.322"},{"key":"ref_5","unstructured":"Simonyan, K., and Zisserman, A. (2015, January 7\u20139). Very Deep Convolutional Networks for Large-Scale Image Recognition. Proceedings of the International Conference on Learning Representations, San Diego, CA, USA."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep Residual Learning for Image Recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Fan, H., Xiong, B., Mangalam, K., Li, Y., Yan, Z., Malik, J., and Feichtenhofer, C. (2021, January 10\u201317). Multiscale Vision Transformers. Proceedings of the 2021 IEEE\/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.00675"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Li, Y., Wu, C.Y., Fan, H., Mangalam, K., Xiong, B., Malik, J., and Feichtenhofer, C. (2022, January 18\u201324). MViTv2: Improved Multiscale Vision Transformers for Classification and Detection. Proceedings of the 2022 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.00476"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Lin, T.Y., Doll\u00c1R, P., Girshick, R., He, K., Hariharan, B., and Belongie, S. (2017, January 21\u201326). Feature Pyramid Networks for Object Detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.106"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Cai, Z., and Vasconcelos, N. (2018, January 18\u201322). Cascade R-CNN: Delving into High Quality Object Detection. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00644"},{"key":"ref_11","first-page":"1","article-title":"Channel-Layer-Oriented Lightweight Spectral\u2013Spatial Network for Hyperspectral Image Classification","volume":"62","author":"Li","year":"2024","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"2557","DOI":"10.1109\/JSTARS.2023.3344635","article-title":"A Lightweight Change Detection Network Based on Feature Interleaved Fusion and Bistage Decoding","volume":"17","author":"Wang","year":"2024","journal-title":"IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Lv, Y., Tian, B., Guo, Q., and Zhang, D. (2025). A Lightweight Small Target Detection Algorithm for UAV Platforms. Appl. Sci., 15.","DOI":"10.3390\/app15010012"},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"652","DOI":"10.1109\/TMM.2022.3230330","article-title":"LRDNet: Lightweight LiDAR Aided Cascaded Feature Pools for Free Road Space Detection","volume":"27","author":"Khan","year":"2025","journal-title":"IEEE Trans. Multimed."},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"107617","DOI":"10.1016\/j.aap.2024.107617","article-title":"Enhancing rail safety through real-time defect detection: A novel lightweight network approach","volume":"203","author":"Cao","year":"2024","journal-title":"Accid. Anal. Prev."},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"2332","DOI":"10.1109\/TCSVT.2023.3307693","article-title":"Boosting Salient Object Detection with Transformer-Based Asymmetric Bilateral U-Net","volume":"34","author":"Qiu","year":"2024","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"e1755","DOI":"10.7717\/peerj-cs.1755","article-title":"Lightweight transformer image feature extraction network","volume":"10","author":"Zheng","year":"2024","journal-title":"PeerJ Comput. Sci."},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1109\/TGRS.2023.3334492","article-title":"Rethinking Transformers for Semantic Segmentation of Remote Sensing Images","volume":"61","author":"Liu","year":"2023","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"3123","DOI":"10.1109\/TPAMI.2023.3341806","article-title":"CrossFormer++: A Versatile Vision Transformer Hinging on Cross-Scale Attention","volume":"46","author":"Wang","year":"2024","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Liu, S., Qi, L., Qin, H., Shi, J., and Jia, J. (2018, January 18\u201322). Path Aggregation Network for Instance Segmentation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00913"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Tan, M., Pang, R., and Le, Q.V. (2020, January 13\u201319). EfficientDet: Scalable and Efficient Object Detection. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.01079"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Wang, X., Zhang, S., Yu, Z., Feng, L., and Zhang, W. (2020, January 13\u201319). Scale-equalizing pyramid convolution for object detection. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.01337"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Park, H.J., Kang, J.W., and Kim, B.G. (2023). ssFPN: Scale sequence (S2) feature-based feature pyramid network for object detection. Sensors, 23.","DOI":"10.3390\/s23094432"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Tran, D., Bourdev, L., Fergus, R., Torresani, L., and Paluri, M. (2015, January 7\u201313). Learning spatiotemporal features with 3d convolutional networks. Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.510"},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"105057","DOI":"10.1016\/j.imavis.2024.105057","article-title":"ASF-YOLO: A novel YOLO model with attentional scale sequence fusion for cell instance segmentation","volume":"147","author":"Kang","year":"2024","journal-title":"Image Vis. Comput."},{"key":"ref_26","unstructured":"Zhao, J.X., Liu, J.J., Fan, D.P., Cao, Y., Yang, J., and Cheng, M.M. (November, January 27). EGNet: Edge guidance network for salient object detection. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Seoul, Republic of Korea."},{"key":"ref_27","first-page":"026512","article-title":"SE2Net: Semantic segmentation of remote sensing images based on self-attention and edge enhancement modules","volume":"15","author":"Liu","year":"2021","journal-title":"J. Appl. Remote Sens."},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Yang, L., Gu, Y., and Feng, H. (2025). Multi-scale Feature Fusion and Feature Calibration with Edge Information Enhancement for Remote Sensing Object Detection. Sci. Rep., 15.","DOI":"10.1038\/s41598-025-99835-7"},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"5465","DOI":"10.1109\/TIP.2023.3318967","article-title":"Context and Spatial Feature Calibration for Real-Time Semantic Segmentation","volume":"32","author":"Li","year":"2023","journal-title":"IEEE Trans. Image Process."},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Gao, S., Zhang, P., Yan, T., and Lu, H. (2024\u20131, January 28). Multi-scale and Detail-Enhanced Segment Anything Model for Salient Object Detection. Proceedings of the ACM International Conference on Multimedia, Melbourne, Australia.","DOI":"10.1145\/3664647.3680650"},{"key":"ref_31","unstructured":"Koubaroulis, D., Matas, J., and Kittler, J. (2002, January 23\u201325). Evaluating Colour-Based Object Recognition Algorithms Using the SOIL-47 Database. Proceedings of the Asian Conference on Computer Vision, Melbourne, Australia."},{"key":"ref_32","unstructured":"Jund, P., Abdo, N., Eitel, A., and Burgard, W. (2016). The Freiburg Groceries Dataset. arXiv."},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Rivera-Rubio, J., Idrees, S., Alexiou, I., Hadjilucas, L., and Bharath, A.A. (2014, January 27\u201330). A dataset for Hand-Held Object Recognition. Proceedings of the IEEE International Conference on Image Processing, Paris, France.","DOI":"10.1109\/WACV.2014.6836057"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Follmann, P., B\u00f6ttger, T., H\u00e4rtinger, P., K\u00f6nig, R., and Ulrich, M. (2018, January 8\u201314). MVTec D2S: Densely Segmented Supermarket Dataset. Proceedings of the Computer Vision\u2013ECCV 2018, Munich, Germany.","DOI":"10.1007\/978-3-030-01249-6_35"},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"96","DOI":"10.1016\/j.compag.2009.09.002","article-title":"Automatic fruit and vegetable classification from images","volume":"70","author":"Rocha","year":"2010","journal-title":"Comput. Electron. Agric."},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Merler, M., Galleguillos, C., and Belongie, S. (2007, January 17\u201322). Recognizing Groceries in situ Using in vitro Training Data. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Minneapolis, MN, USA.","DOI":"10.1109\/CVPR.2007.383486"},{"key":"ref_37","doi-asserted-by":"crossref","first-page":"440","DOI":"10.1007\/978-3-319-10605-2_29","article-title":"Recognizing Products: A Per-exemplar Multi-label Image Classification Approach","volume":"Volume 8690","author":"Fleet","year":"2014","journal-title":"Computer Vision\u2013ECCV 2014"},{"key":"ref_38","unstructured":"Peng, J., Xiao, C., Wei, X., and Li, Y. (2020). RP2K: A Large-Scale Retail Product Dataset for Fine-Grained Image Classification. arXiv."},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Goldman, E., Herzig, R., Eisenschtat, A., Goldberger, J., and Hassner, T. (2019, January 16\u201320). Precise Detection in Densely Packed Scenes. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00537"},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Cai, Y., Wen, L., Zhang, L., Du, D., and Wang, W. (2021, January 2\u20139). Rethinking Object Detection in Retail Stores. Proceedings of the AAAI Conference on Artificial Intelligence, Virtual.","DOI":"10.1609\/aaai.v35i2.16178"},{"key":"ref_41","doi-asserted-by":"crossref","first-page":"7722","DOI":"10.1109\/TII.2019.2954956","article-title":"Toward New Retail: A Benchmark Dataset for Smart Unmanned Vending Machines","volume":"16","author":"Zhang","year":"2020","journal-title":"IEEE Trans. Ind. Inform."},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Hao, Y., Fu, Y., and Jiang, Y.G. (2019, January 10\u201313). Take Goods from Shelves: A Dataset for Class-Incremental Object Detection. Proceedings of the 2019 on International Conference on Multimedia Retrieval, Ottawa, ON, Canada.","DOI":"10.1145\/3323873.3325033"},{"key":"ref_43","unstructured":"Georgiadis, K., Kordopatis-Zilos, G., Kalaganis, F., Migkotzidis, P., Chatzilari, E., Panakidou, V., Pantouvakis, K., Tortopidis, S., Papadopoulos, S., and Nikolopoulos, S. (July, January 29). Products-6K: A Large-Scale Groceries Product Recognition Dataset. Proceedings of the PETRA Conference, Virtual."},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"Cheng, L., Zhou, X., Zhao, L., Li, D., Shang, H., Zheng, Y., Pan, P., and Xu, Y. (2020, January 23\u201328). Weakly Supervised Learning with Side Information for Noisy Labeled Images. Proceedings of the Computer Vision\u2013ECCV 2020, Glasgow, UK.","DOI":"10.1007\/978-3-030-58577-8_19"},{"key":"ref_45","unstructured":"Bai, Y., Chen, Y., Yu, W., Wang, L., and Zhang, W. (2020). Products-10K: A Large-scale Product Recognition Dataset. arXiv."},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Zhan, X., Wu, Y., Dong, X., Wei, Y., Lu, M., Zhang, Y., Xu, H., and Liang, X. (2021, January 10\u201317). Product1M: Towards Weakly Supervised Instance-Level Product Retrieval via Cross-modal Pretraining. Proceedings of the 2021 IEEE International Conference on Computer Vision (ICCV), Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.01157"},{"key":"ref_47","unstructured":"Wei, X.S., Cui, Q., Yang, L., Wang, P., and Liu, L. (2019). RPC: A Large-Scale Retail Product Checkout Dataset. arXiv."},{"key":"ref_48","unstructured":"Chu, C., Zhmoginov, A., and Sandler, M. (2017). CycleGAN, a Master of Steganography. arXiv."},{"key":"ref_49","unstructured":"Amsaleg, L., Huet, B., Larson, M.A., Gravier, G., and Hung, H. (2019, January 21\u201325). Data Priming Network for Automatic Check-Out. Proceedings of the 27th ACM International Conference on Multimedia, Nice, France."},{"key":"ref_50","doi-asserted-by":"crossref","unstructured":"Yang, Y., Sheng, L., Jiang, X., Wang, H., Xu, D., and Cao, X. (2021, January 3\u20138). IncreACO: Incrementally Learned Automatic Check-out with Photorealistic Exemplar Augmentation. Proceedings of the IEEE Winter Conference on Applications of Computer Vision, Waikoloa, HI, USA.","DOI":"10.1109\/WACV48630.2021.00067"},{"key":"ref_51","doi-asserted-by":"crossref","first-page":"4158","DOI":"10.1109\/TMM.2020.3037502","article-title":"Iterative Knowledge Distillation for Automatic Check-Out","volume":"23","author":"Zhang","year":"2021","journal-title":"IEEE Trans. Multim."},{"key":"ref_52","doi-asserted-by":"crossref","first-page":"593","DOI":"10.1007\/s00521-021-06394-9","article-title":"Context-guided feature enhancement network for automatic check-out","volume":"34","author":"Sun","year":"2022","journal-title":"Neural Comput. Appl."},{"key":"ref_53","first-page":"277","article-title":"Automatic Check-Out via Prototype-Based Classifier Learning from Single-Product Exemplars","volume":"Volume 13685","author":"Avidan","year":"2022","journal-title":"Proceedings of the Computer Vision\u2013ECCV 2022"},{"key":"ref_54","doi-asserted-by":"crossref","first-page":"9147","DOI":"10.1109\/TMM.2023.3247219","article-title":"Prototype Learning for Automatic Check-Out","volume":"25","author":"Chen","year":"2023","journal-title":"IEEE Trans. Multim."},{"key":"ref_55","doi-asserted-by":"crossref","first-page":"3004","DOI":"10.1109\/TIP.2022.3163527","article-title":"Self-Supervised Multi-Category Counting Networks for Automatic Check-Out","volume":"31","author":"Chen","year":"2022","journal-title":"IEEE Trans. Image Process."},{"key":"ref_56","first-page":"74","article-title":"Incremental learning method for intelligent retail automatic check-out","volume":"48","author":"Chen","year":"2024","journal-title":"J. Nanjing Univ. Sci. Technol."},{"key":"ref_57","doi-asserted-by":"crossref","first-page":"2350049:1","DOI":"10.1142\/S0129065723500491","article-title":"Decoupled Edge Guidance Network for Automatic Checkout","volume":"33","author":"You","year":"2023","journal-title":"Int. J. Neural Syst."},{"key":"ref_58","doi-asserted-by":"crossref","unstructured":"Tan, L., Liu, S., Gao, J., Liu, X., Chu, L., and Jiang, H. (2024). Enhanced Self-Checkout System for Retail Based on Improved YOLOv10. J. Imaging, 10.","DOI":"10.3390\/jimaging10100248"},{"key":"ref_59","doi-asserted-by":"crossref","first-page":"1649","DOI":"10.1016\/j.procs.2024.09.644","article-title":"Solving Automatic Check-out with Fine-Tuned YOLO Models","volume":"246","author":"Skinderowicz","year":"2024","journal-title":"Procedia Comput. Sci."},{"key":"ref_60","first-page":"324","article-title":"Product Detection and Counting Algorithm for Automatic Check-Out","volume":"60","author":"Yang","year":"2024","journal-title":"Comput. Eng. Appl."},{"key":"ref_61","doi-asserted-by":"crossref","first-page":"4940","DOI":"10.1109\/TPAMI.2025.3546356","article-title":"Reliable Representation Learning for Incomplete Multi-View Missing Multi-Label Classification","volume":"47","author":"Liu","year":"2025","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_62","doi-asserted-by":"crossref","first-page":"5539","DOI":"10.1109\/TMM.2022.3194332","article-title":"Localized Sparse Incomplete Multi-View Clustering","volume":"25","author":"Liu","year":"2023","journal-title":"IEEE Trans. Multimed."},{"key":"ref_63","doi-asserted-by":"crossref","first-page":"62","DOI":"10.1109\/TSMC.1979.4310076","article-title":"A threshold selection method from gray-level histograms","volume":"9","author":"Otsu","year":"1979","journal-title":"IEEE Trans. Syst. Man Cybern."},{"key":"ref_64","doi-asserted-by":"crossref","unstructured":"Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Doll\u00c1R, P., and Zitnick, C.L. (2014, January 6\u201312). Microsoft COCO: Common Objects in Context. Proceedings of the Computer Vision\u2013ECCV 2014, Zurich, Switzerland.","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"ref_65","unstructured":"Wu, Y., Kirillov, A., Massa, F., Lo, W.Y., and Girshick, R. (2025, March 28). Detectron2. Available online: https:\/\/github.com\/facebookresearch\/detectron2."},{"key":"ref_66","first-page":"8026","article-title":"PyTorch: An Imperative Style, High-Performance Deep Learning Library","volume":"32","author":"Paszke","year":"2019","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_67","doi-asserted-by":"crossref","unstructured":"Dai, L., and Liu, H. (2023, January 10\u201312). DeCo-DETR: Densely Packed Commodity Detection with Transformer. Proceedings of the 2023 7th Asian Conference on Artificial Intelligence Technology (ACAIT), Jiaxing, China.","DOI":"10.1109\/ACAIT60137.2023.10528472"}],"container-title":["Journal of Imaging"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2313-433X\/11\/10\/337\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,9,29]],"date-time":"2025-09-29T13:44:59Z","timestamp":1759153499000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2313-433X\/11\/10\/337"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,9,27]]},"references-count":67,"journal-issue":{"issue":"10","published-online":{"date-parts":[[2025,10]]}},"alternative-id":["jimaging11100337"],"URL":"https:\/\/doi.org\/10.3390\/jimaging11100337","relation":{},"ISSN":["2313-433X"],"issn-type":[{"value":"2313-433X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,9,27]]}}}