{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,1]],"date-time":"2026-06-01T23:21:42Z","timestamp":1780356102439,"version":"3.54.1"},"reference-count":57,"publisher":"MDPI AG","issue":"6","license":[{"start":{"date-parts":[[2022,6,8]],"date-time":"2022-06-08T00:00:00Z","timestamp":1654646400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"European project INFINITY","award":["883293"],"award-info":[{"award-number":["883293"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Future Internet"],"abstract":"<jats:p>The graphical page object detection classifies and localizes objects such as Tables and Figures in a document. As deep learning techniques for object detection become increasingly successful, many supervised deep neural network-based methods have been introduced to recognize graphical objects in documents. However, these models necessitate a substantial amount of labeled data for the training process. This paper presents an end-to-end semi-supervised framework for graphical object detection in scanned document images to address this limitation. Our method is based on a recently proposed Soft Teacher mechanism that examines the effects of small percentage-labeled data on the classification and localization of graphical objects. On both the PubLayNet and the IIIT-AR-13K datasets, the proposed approach outperforms the supervised models by a significant margin in all labeling ratios (1%,\u00a05%, and 10%). Furthermore, the 10% PubLayNet Soft Teacher model improves the average precision of Table, Figure, and List by +5.4,+1.2, and +3.2 points, respectively, with a similar total mAP as the Faster-RCNN baseline. Moreover, our model trained on 10% of IIIT-AR-13K labeled data beats the previous fully supervised method +4.5 points.<\/jats:p>","DOI":"10.3390\/fi14060176","type":"journal-article","created":{"date-parts":[[2022,6,10]],"date-time":"2022-06-10T00:22:39Z","timestamp":1654820559000},"page":"176","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":10,"title":["Toward Semi-Supervised Graphical Object Detection in Document Images"],"prefix":"10.3390","volume":"14","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-1121-0885","authenticated-orcid":false,"given":"Goutham","family":"Kallempudi","sequence":"first","affiliation":[{"name":"Department of Computer Science, Technical University of Kaiserslautern, 67663 Kaiserslautern, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0456-6493","authenticated-orcid":false,"given":"Khurram Azeem","family":"Hashmi","sequence":"additional","affiliation":[{"name":"Department of Computer Science, Technical University of Kaiserslautern, 67663 Kaiserslautern, Germany"},{"name":"Mindgarage, Technical University of Kaiserslautern, 67663 Kaiserslautern, Germany"},{"name":"German Research Institute for Artificial Intelligence (DFKI), 67663 Kaiserslautern, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Alain","family":"Pagani","sequence":"additional","affiliation":[{"name":"German Research Institute for Artificial Intelligence (DFKI), 67663 Kaiserslautern, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4029-6574","authenticated-orcid":false,"given":"Marcus","family":"Liwicki","sequence":"additional","affiliation":[{"name":"Department of Computer Science, Lule\u00e5 University of Technology, 97187 Lulea, Sweden"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Didier","family":"Stricker","sequence":"additional","affiliation":[{"name":"Department of Computer Science, Technical University of Kaiserslautern, 67663 Kaiserslautern, Germany"},{"name":"German Research Institute for Artificial Intelligence (DFKI), 67663 Kaiserslautern, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0536-6867","authenticated-orcid":false,"given":"Muhammad Zeshan","family":"Afzal","sequence":"additional","affiliation":[{"name":"Department of Computer Science, Technical University of Kaiserslautern, 67663 Kaiserslautern, Germany"},{"name":"Mindgarage, Technical University of Kaiserslautern, 67663 Kaiserslautern, Germany"},{"name":"German Research Institute for Artificial Intelligence (DFKI), 67663 Kaiserslautern, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2022,6,8]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Orosz, T., V\u00e1gi, R., Cs\u00e1nyi, G.M., Nagy, D., \u00dcveges, I., Vad\u00e1sz, J.P., and Megyeri, A. (2021). Evaluating Human versus Machine Learning Performance in a LegalTech Problem. Appl. Sci., 12.","DOI":"10.3390\/app12010297"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Fang, J., Gao, L., Bai, K., Qiu, R., Tao, X., and Tang, Z. (2011, January 18\u201321). A Table Detection Method for Multipage PDF Documents via Visual Seperators and Tabular Structures. Proceedings of the 2011 International Conference on Document Analysis and Recognition, ICDAR 2011, Beijing, China.","DOI":"10.1109\/ICDAR.2011.304"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Chen, J., and Lopresti, D.P. (2011, January 18\u201321). Table Detection in Noisy Off-line Handwritten Documents. Proceedings of the 2011 International Conference on Document Analysis and Recognition, ICDAR 2011, Beijing, China.","DOI":"10.1109\/ICDAR.2011.88"},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"116","DOI":"10.1109\/ICDARW.2019.40091","article-title":"Feedback learning: Automating the process of correcting and completing the extracted information","volume":"Volume 5","author":"Hashmi","year":"2019","journal-title":"Proceedings of the 2019 International Conference on Document Analysis and Recognition Workshops (ICDARW)"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Saha, R., Mondal, A., and Jawahar, C.V. (2019, January 20\u201325). Graphical Object Detection in Document Images. Proceedings of the 2019 International Conference on Document Analysis and Recognition, ICDAR 2019, Sydney, Australia.","DOI":"10.1109\/ICDAR.2019.00018"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Girshick, R.B. (2015, January 7\u201313). Fast R-CNN. Proceedings of the 2015 IEEE International Conference on Computer Vision, ICCV 2015, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.169"},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"1137","DOI":"10.1109\/TPAMI.2016.2577031","article-title":"Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks","volume":"39","author":"Ren","year":"2017","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Redmon, J., and Farhadi, A. (2016). YOLO9000: Better, Faster, Stronger. arXiv.","DOI":"10.1109\/CVPR.2017.690"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"He, K., Gkioxari, G., Doll\u00e1r, P., and Girshick, R.B. (2017, January 21\u201326). Mask R-CNN. Proceedings of the IEEE International Conference on Computer Vision, Honolulu, HI, USA.","DOI":"10.1109\/ICCV.2017.322"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Xu, M., Zhang, Z., Hu, H., Wang, J., Wang, L., Wei, F., Bai, X., and Liu, Z. (2021, January 11\u201317). End-to-End Semi-Supervised Object Detection with Soft Teacher. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.00305"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Wang, K., Yan, X., Zhang, D., Zhang, L., and Lin, L. (2018, January 18\u201323). Towards Human-Machine Cooperation: Self-supervised Sample Mining for Object Detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00173"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Tang, P., Ramaiah, C., Xu, R., and Xiong, C. (2021, January 11\u201317). Proposal Learning for Semi-Supervised Object Detection. Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision, Montreal, QC, Canada.","DOI":"10.1109\/WACV48630.2021.00234"},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"109","DOI":"10.1016\/j.cogsys.2017.05.006","article-title":"Active and semi-supervised learning for object detection with imperfect data","volume":"45","author":"Rhee","year":"2017","journal-title":"Cogn. Syst. Res."},{"key":"ref_14","unstructured":"Xie, Q., Dai, Z., Hovy, E.H., Luong, T., and Le, Q. (2020, January 6\u201312). Unsupervised Data Augmentation for Consistency Training. Proceedings of the Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, Virtual. Available online: https:\/\/proceedings.neurips.cc\/paper\/2020\/file\/44feb0096faa8326192570788b38c1d1-Paper.pdf."},{"key":"ref_15","unstructured":"Doermann, D.S., Govindaraju, V., Lopresti, D.P., and Natarajan, P. (2010, January 9\u201311). Table detection in heterogeneous documents. Proceedings of the The Ninth IAPR International Workshop on Document Analysis Systems, DAS 2010, Boston, MA, USA."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Kasar, T., Barlas, P., Adam, S., Chatelain, C., and Paquet, T. (2013, January 25\u201328). Learning to Detect Tables in Scanned Document Images Using Line Information. Proceedings of the 12th International Conference on Document Analysis and Recognition, ICDAR 2013, Washington, DC, USA.","DOI":"10.1109\/ICDAR.2013.240"},{"key":"ref_17","unstructured":"Cesarini, F., Marinai, S., Sarti, L., and Soda, G. (2002, January 11\u201315). Trainable Table Location in Document Images. Proceedings of the 16th International Conference on Pattern Recognition, ICPR 2002, Quebec, QC, Canada."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"e Silva, A.C. (2009, January 26\u201329). Learning Rich Hidden Markov Models in Document Analysis: Table Location. Proceedings of the 10th International Conference on Document Analysis and Recognition, ICDAR 2009, Barcelona, Spain.","DOI":"10.1109\/ICDAR.2009.185"},{"key":"ref_19","first-page":"255","article-title":"The T-Recs Table Recognition and Analysis System","volume":"Volume 1655","author":"Lee","year":"1998","journal-title":"Proceedings of the Document Analysis Systems: Theory and Practice, Third IAPR Workshop, DAS\u201998"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Hao, L., Gao, L., Yi, X., and Tang, Z. (2016, January 11\u201314). A Table Detection Method for PDF Documents Based on Convolutional Neural Networks. Proceedings of the 12th IAPR Workshop on Document Analysis Systems, DAS 2016, Santorini, Greece.","DOI":"10.1109\/DAS.2016.23"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Schreiber, S., Agne, S., Wolf, I., Dengel, A., and Ahmed, S. (2017, January 9\u201315). DeepDeSRT: Deep Learning for Detection and Structure Recognition of Tables in Document Images. Proceedings of the 14th IAPR International Conference on Document Analysis and Recognition, ICDAR 2017, Kyoto, Japan.","DOI":"10.1109\/ICDAR.2017.192"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Hashmi, K.A., Pagani, A., Liwicki, M., Stricker, D., and Afzal, M.Z. (2021). CasTabDetectoRS: Cascade Network for Table Detection in Document Images with Recursive Feature Pyramid and Switchable Atrous Convolution. J. Imaging, 7.","DOI":"10.20944\/preprints202109.0059.v1"},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"68150L","DOI":"10.1117\/12.767295","article-title":"Segmentation-based retrieval of document images from diverse collections","volume":"Volume 6815","author":"Yanikoglu","year":"2008","journal-title":"Proceedings of the Document Recognition and Retrieval XV, part of the IS&T-SPIE Electronic Imaging Symposium"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Nayef, N., and Ogier, J. (2015, January 23\u201326). Text zone classification using unsupervised feature learning. Proceedings of the 13th International Conference on Document Analysis and Recognition, ICDAR 2015, Nancy, France.","DOI":"10.1109\/ICDAR.2015.7333867"},{"key":"ref_25","first-page":"200","article-title":"Text\/Graphics Separation Revisited","volume":"Volume 2423","author":"Lopresti","year":"2002","journal-title":"Proceedings of the Document Analysis Systems V, 5th International Workshop, DAS 2002"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Zhong, X., Tang, J., and Jimeno-Yepes, A. (2019, January 20\u201325). PubLayNet: Largest dataset ever for document layout analysis. Proceedings of the 2019 International Conference on Document Analysis and Recognition (ICDAR), Sydney, NSW, Australia.","DOI":"10.1109\/ICDAR.2019.00166"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Zach, C., S\u00e1nchez, A.P., and Pham, M. (2015, January 7\u201312). A dynamic programming approach for fast and robust object pose recognition from range images. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2015, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298615"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Bhatt, J., Hashmi, K.A.A., Afzal, M.Z., and Stricker, D. (2021). A Survey of Graphical Page Object Detection with Deep Neural Networks. Appl. Sci., 11.","DOI":"10.20944\/preprints202104.0739.v1"},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"87663","DOI":"10.1109\/ACCESS.2021.3087865","article-title":"Current Status and Performance Analysis of Table Recognition in Document Images with Deep Neural Networks","volume":"9","author":"Hashmi","year":"2021","journal-title":"IEEE Access"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Nazir, D., Hashmi, K.A., Pagani, A., Liwicki, M., Stricker, D., and Afzal, M.Z. (2021). HybridTabNet: Towards better table detection in scanned document images. Appl. Sci., 11.","DOI":"10.3390\/app11188396"},{"key":"ref_31","first-page":"311","article-title":"Semi-supervised Deep Learning for Fully Convolutional Networks","volume":"Volume 10435","author":"Descoteaux","year":"2017","journal-title":"Proceedings of the Medical Image Computing and Computer Assisted Intervention - MICCAI 2017-20th International Conference"},{"key":"ref_32","first-page":"370","article-title":"ASDNet: Attention Based Semi-supervised Deep Networks for Medical Image Segmentation","volume":"Volume 11073","author":"Frangi","year":"2018","journal-title":"Proceedings of the Medical Image Computing and Computer Assisted Intervention - MICCAI 2018-21st International Conference"},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"925","DOI":"10.1007\/s11548-018-1772-0","article-title":"Exploiting the potential of unlabeled endoscopic video data with self-supervised learning","volume":"13","author":"Zimmerer","year":"2018","journal-title":"Int. J. Comput. Assist. Radiol. Surg."},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Iscen, A., Tolias, G., Avrithis, Y., and Chum, O. (2019, January 16\u201320). Label Propagation for Deep Semi-supervised Learning. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00521"},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"108602","DOI":"10.1016\/j.knosys.2022.108602","article-title":"Deep semi-supervised learning with contrastive learning and partial label propagation for image data","volume":"245","author":"Gan","year":"2022","journal-title":"Knowl. Based Syst."},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Kiran, B.R., Thomas, D.M., and Parakkal, R. (2018). An overview of deep learning based methods for unsupervised and semi-supervised anomaly detection in videos. J. Imaging, 4.","DOI":"10.3390\/jimaging4020036"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Papandreou, G., Chen, L., Murphy, K.P., and Yuille, A.L. (2015, January 7\u201313). Weakly-and Semi-Supervised Learning of a Deep Convolutional Network for Semantic Image Segmentation. Proceedings of the 2015 IEEE International Conference on Computer Vision, ICCV 2015, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.203"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Olsson, V., Tranheden, W., Pinto, J., and Svensson, L. (2021, January 11\u201317). ClassMix: Segmentation-Based Data Augmentation for Semi-Supervised Learning. Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision, Montreal, QC, Canada.","DOI":"10.1109\/WACV48630.2021.00141"},{"key":"ref_39","unstructured":"Wallach, H.M., Larochelle, H., Beygelzimer, A., d\u2019Alch\u00e9-Buc, F., Fox, E.B., and Garnett, R. (2019, January 8\u201314). Consistency-based Semi-supervised Learning for Object detection. Proceedings of the Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, Vancouver, BC, Canada. Available online: https:\/\/papers.nips.cc\/paper\/2019\/hash\/d0f4dae80c3d0277922f8371d5827292-Abstract.html."},{"key":"ref_40","doi-asserted-by":"crossref","first-page":"1979","DOI":"10.1109\/TPAMI.2018.2858821","article-title":"Virtual Adversarial Training: A Regularization Method for Supervised and Semi-supervised Learning","volume":"41","author":"Miyato","year":"2018","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_41","unstructured":"Sajjadi, M., Javanmardi, M., and Tasdizen, T. (2022, April 28). Regularization With Stochastic Transformations and Perturbations for Deep Semi-Supervised Learning. Available online: https:\/\/proceedings.neurips.cc\/paper\/2016\/file\/30ef30b64204a3088a26bc2e6ecf7602-Paper.pdf."},{"key":"ref_42","unstructured":"Grandvalet, Y., and Bengio, Y. (2004, January 13\u201318). Semi-supervised Learning by Entropy Minimization. Proceedings of the Neural Information Processing Systems 17 Neural Information Processing Systems, NIPS 2004, Vancouver, BC, Canada."},{"key":"ref_43","unstructured":"Berthelot, D., Carlini, N., Goodfellow, I.J., Papernot, N., Oliver, A., and Raffel, C. (2022, April 28). MixMatch: A Holistic Approach to Semi-Supervised Learning. Available online: https:\/\/proceedings.neurips.cc\/paper\/2019\/file\/1cd138d0499a68f4bb72bee04bbec2d7-Paper.pdf."},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"Jeong, J., Verma, V., Hyun, M., Kannala, J., and Kwak, N. (2020, January 13\u201319). Interpolation-based semi-supervised learning for object detection. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR46437.2021.01143"},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Radosavovic, I., Doll\u00e1r, P., Girshick, R.B., Gkioxari, G., and He, K. (2017, January 21\u201326). Data Distillation: Towards Omni-Supervised Learning. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, Hawaii, USA.","DOI":"10.1109\/CVPR.2018.00433"},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Yang, Q., Wei, X., Wang, B., Hua, X., and Zhang, L. (2021, January 19\u201325). Interactive Self-Training With Mean Teachers for Semi-Supervised Object Detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2021, Virtual.","DOI":"10.1109\/CVPR46437.2021.00588"},{"key":"ref_47","doi-asserted-by":"crossref","unstructured":"Tang, Y., Chen, W., Luo, Y., and Zhang, Y. (2021, January 20\u201325). Humble Teachers Teach Better Students for Semi-Supervised Object Detection. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.00315"},{"key":"ref_48","doi-asserted-by":"crossref","unstructured":"Lin, T., Maire, M., Belongie, S.J., Bourdev, L.D., Girshick, R.B., Hays, J., Perona, P., Ramanan, D., Doll\u00e1r, P., and Zitnick, C.L. (2014, January 6\u201412). Microsoft COCO: Common Objects in Context. Proceedings of the European Conference on Computer Vision, Zurich, Switzerland.","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"ref_49","first-page":"596","article-title":"FixMatch: Simplifying Semi-Supervised Learning with Consistency and Confidence","volume":"33","author":"Sohn","year":"2020","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_50","doi-asserted-by":"crossref","unstructured":"Mondal, A., Lipps, P., and Jawahar, C.V. (2020, January 26\u201329). IIIT-AR-13K: A New Dataset for Graphical Object Detection in Documents. Proceedings of the International Workshop on Document Analysis Systems, Wuhan, China.","DOI":"10.1007\/978-3-030-57058-3_16"},{"key":"ref_51","doi-asserted-by":"crossref","unstructured":"Li, M., Xu, Y., Cui, L., Huang, S., Wei, F., Li, Z., and Zhou, M. (2020). DocBank: A Benchmark Dataset for Document Layout Analysis. arXiv.","DOI":"10.18653\/v1\/2020.coling-main.82"},{"key":"ref_52","unstructured":"Sohn, K., Zhang, Z., Li, C., Zhang, H., Lee, C., and Pfister, T. (2020). A Simple Semi-Supervised Learning Framework for Object Detection. arXiv."},{"key":"ref_53","unstructured":"Powers, D.M.W. (2020). Evaluation: From precision, recall and F-measure to ROC, informedness, markedness and correlation. arXiv."},{"key":"ref_54","doi-asserted-by":"crossref","unstructured":"Lin, T., Doll\u00e1r, P., Girshick, R.B., He, K., Hariharan, B., and Belongie, S.J. (2017, January 21\u201326). Feature Pyramid Networks for Object Detection. Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2017, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.106"},{"key":"ref_55","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2015, January 7\u201312). Deep Residual Learning for Image Recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_56","doi-asserted-by":"crossref","unstructured":"Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., and Guo, B. (2021, January 11\u201317). Swin Transformer: Hierarchical Vision Transformer using Shifted Windows. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.00986"},{"key":"ref_57","doi-asserted-by":"crossref","unstructured":"Nguyen, P., Ngo, L., Truong, T., Nguyen, T.T., Vo, N.D., and Nguyen, K. (2021, January 21\u201322). Page Object Detection with YOLOF. Proceedings of the 2021 8th NAFOSTED Conference on Information and Computer Science (NICS), Hanoi, Vietnam.","DOI":"10.1109\/NICS54270.2021.9701449"}],"container-title":["Future Internet"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1999-5903\/14\/6\/176\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T23:26:06Z","timestamp":1760138766000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1999-5903\/14\/6\/176"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,6,8]]},"references-count":57,"journal-issue":{"issue":"6","published-online":{"date-parts":[[2022,6]]}},"alternative-id":["fi14060176"],"URL":"https:\/\/doi.org\/10.3390\/fi14060176","relation":{},"ISSN":["1999-5903"],"issn-type":[{"value":"1999-5903","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,6,8]]}}}