{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,27]],"date-time":"2026-03-27T23:05:59Z","timestamp":1774652759104,"version":"3.50.1"},"reference-count":74,"publisher":"MDPI AG","issue":"10","license":[{"start":{"date-parts":[[2021,10,16]],"date-time":"2021-10-16T00:00:00Z","timestamp":1634342400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["J. Imaging"],"abstract":"<jats:p>Table detection is a preliminary step in extracting reliable information from tables in scanned document images. We present CasTabDetectoRS, a novel end-to-end trainable table detection framework that operates on Cascade Mask R-CNN, including Recursive Feature Pyramid network and Switchable Atrous Convolution in the existing backbone architecture. By utilizing a comparativelyightweight backbone of ResNet-50, this paper demonstrates that superior results are attainable without relying on pre- and post-processing methods, heavier backbone networks (ResNet-101, ResNeXt-152), and memory-intensive deformable convolutions. We evaluate the proposed approach on five different publicly available table detection datasets. Our CasTabDetectoRS outperforms the previous state-of-the-art results on four datasets (ICDAR-19, TableBank, UNLV, and Marmot) and accomplishes comparable results on ICDAR-17 POD. Upon comparing with previous state-of-the-art results, we obtain a significant relative error reduction of 56.36%, 20%, 4.5%, and 3.5% on the datasets of ICDAR-19, TableBank, UNLV, and Marmot, respectively. Furthermore, this paper sets a new benchmark by performing exhaustive cross-datasets evaluations to exhibit the generalization capabilities of the proposed method.<\/jats:p>","DOI":"10.3390\/jimaging7100214","type":"journal-article","created":{"date-parts":[[2021,10,17]],"date-time":"2021-10-17T23:05:47Z","timestamp":1634511947000},"page":"214","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":16,"title":["CasTabDetectoRS: Cascade Network for Table Detection in Document Images with Recursive Feature Pyramid and Switchable Atrous Convolution"],"prefix":"10.3390","volume":"7","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-0456-6493","authenticated-orcid":false,"given":"Khurram Azeem","family":"Hashmi","sequence":"first","affiliation":[{"name":"Department of Computer Science, Technical University of Kaiserslautern, 67663 Kaiserslautern, Germany"},{"name":"Mindgarage, Technical University of Kaiserslautern, 67663 Kaiserslautern, Germany"},{"name":"German Research Institute for Artificial Intelligence (DFKI), 67663 Kaiserslautern, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Alain","family":"Pagani","sequence":"additional","affiliation":[{"name":"German Research Institute for Artificial Intelligence (DFKI), 67663 Kaiserslautern, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4029-6574","authenticated-orcid":false,"given":"Marcus","family":"Liwicki","sequence":"additional","affiliation":[{"name":"Department of Computer Science, Lule\u00e5 University of Technology, 971 87 Lule\u00e5, Sweden"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Didier","family":"Stricker","sequence":"additional","affiliation":[{"name":"Department of Computer Science, Technical University of Kaiserslautern, 67663 Kaiserslautern, Germany"},{"name":"German Research Institute for Artificial Intelligence (DFKI), 67663 Kaiserslautern, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0536-6867","authenticated-orcid":false,"given":"Muhammad Zeshan","family":"Afzal","sequence":"additional","affiliation":[{"name":"Department of Computer Science, Technical University of Kaiserslautern, 67663 Kaiserslautern, Germany"},{"name":"Mindgarage, Technical University of Kaiserslautern, 67663 Kaiserslautern, Germany"},{"name":"German Research Institute for Artificial Intelligence (DFKI), 67663 Kaiserslautern, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2021,10,16]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Gao, L., Yi, X., Jiang, Z., Hao, L., and Tang, Z. (2017, January 9\u201315). ICDAR2017 competition on page object detection. Proceedings of the 14th IAPR International Conference on Document Analysis and Recognition. (ICDAR), Kyoto, Japan.","DOI":"10.1109\/ICDAR.2017.231"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Bhatt, J., Hashmi, K.A., Afzal, M.Z., and Stricker, D. (2021). A Survey of Graphical Page Object Detection with Deep Neural Networks. Applied Sci., 11.","DOI":"10.20944\/preprints202104.0739.v1"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Zhao, Z., Jiang, M., Guo, S., Wang, Z., Chao, F., and Tan, K.C. (2020, January 19\u201324). Improving deepearning based optical character recognition via neural architecture search. Proceedings of the IEEE Congress on Evolutionary Computation (CEC), Glasgow, UK.","DOI":"10.1109\/CEC48606.2020.9185798"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Hashmi, K.A., Ponnappa, R.B., Bukhari, S.S., Jenckel, M., and Dengel, A. (2019, January 20\u201325). Feedback Learning: Automating the Process of Correcting and Completing the Extracted Information. Proceedings of the International Conference on Document Analysis and Recognition Workshops (ICDARW), Sydney, Australia.","DOI":"10.1109\/ICDARW.2019.40091"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"van Strien, D., Beelen, K., Ardanuy, M.C., Hosseini, K., McGillivray, B., and Colavizza, G. (2020, January 22\u201324). Assessing the Impact of OCR Quality on Downstream NLP Tasks. Proceedings of the ICAART (1), Valletta, Malta.","DOI":"10.5220\/0009169004840496"},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"22","DOI":"10.1117\/12.304642","article-title":"Table structure recognition based on robust block segmentation","volume":"Volume 3305","author":"Kieninger","year":"1998","journal-title":"Document Recognition V"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Schreiber, S., Agne, S., Wolf, I., Dengel, A., and Ahmed, S. (2017, January 9\u201315). Deepdesrt: Deepearning for detection and structure recognition of tables in document images. Proceedings of the 14th IAPR International Conference on Document Analysis and Recognition (ICDAR), Kyoto, Japan.","DOI":"10.1109\/ICDAR.2017.192"},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"87663","DOI":"10.1109\/ACCESS.2021.3087865","article-title":"Current Status and Performance Analysis of Table Recognition in Document Images with Deep Neural Networks","volume":"9","author":"Hashmi","year":"2021","journal-title":"IEEE Access"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Hashmi, K.A., Pagani, A., Liwicki, M., Stricker, D., and Afzal, M.Z. (2021). Cascade Network with Deformable Composite Backbone for Formula Detection in Scanned Document Images. Appl. Sci., 11.","DOI":"10.20944\/preprints202107.0165.v1"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Smith, R. (2007, January 23\u201326). An overview of the Tesseract OCR engine. Proceedings of the Ninth International Conference on Document Analysis and Recognition (ICDAR 2007), Curitiba, Brazil.","DOI":"10.1109\/ICDAR.2007.4376991"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Prasad, D., Gadpal, A., Kapadni, K., Visave, M., and Sultanpure, K. (2020, January 14\u201319). CascadeTabNet: An approach for end to end table detection and structure recognition from image-based documents. Proceedings of the IEEE\/CVF Conference Computer Vision Pattern Recognition Workshops, Seattle, WA, USA.","DOI":"10.1109\/CVPRW50498.2020.00294"},{"key":"ref_12","unstructured":"Agarwal, M., Mondal, A., and Jawahar, C. (2020). CDeC-Net: Composite Deformable Cascade Network for Table Detection in Document Images. arXiv."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Zheng, X., Burdick, D., Popa, L., Zhong, X., and Wang, N.X.R. (2021, January 5\u20139). Global table extractor (gte): A framework for joint table identification and cell structure recognition using visual context. Proceedings of the IEEE\/CVF Winter Conference Applied Computer Vision, Virtual (Online).","DOI":"10.1109\/WACV48630.2021.00074"},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"8396","DOI":"10.3390\/app11188396","article-title":"HybridTabNet: Towards Better Table Detection in Scanned Document Images","volume":"11","author":"Afzal","year":"2021","journal-title":"Appl. Sci."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Doermann, D., and Tombre, K. (2014). Handbook of Document Image Processing and Recognition, Chapter Recognition of Tables and Forms, Springer.","DOI":"10.1007\/978-0-85729-859-1"},{"key":"ref_16","first-page":"1","article-title":"A survey of table recognition","volume":"7","author":"Zanibbi","year":"2004","journal-title":"Doc. Anal. Recognit."},{"key":"ref_17","unstructured":"Kieninger, T., and Dengel, A. (2001, January 10\u201313). Applying the T-RECS table recognition system to the businessetter domain. Proceedings of the 6th International Conference on Document Analysis and Recognition, Seattle, WA, USA."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Shigarov, A., Mikhailov, A., and Altaev, A. (2016, January 13\u201316). Configurable table structure recognition in untagged PDF documents. Proceedings of the 2016 ACM Symposium Document Engineering, Vienna, Austria.","DOI":"10.1145\/2960811.2967152"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Gilani, A., Qasim, S.R., Malik, I., and Shafait, F. (2017, January 9\u201315). Table detection using deepearning. Proceedings of the 14th IAPR International Conference on Document Analysis and Recognition (ICDAR), Kyoto, Japan.","DOI":"10.1109\/ICDAR.2017.131"},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"74151","DOI":"10.1109\/ACCESS.2018.2880211","article-title":"Decnt: Deep deformable cnn for table detection","volume":"6","author":"Siddiqui","year":"2018","journal-title":"IEEE Access"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Hashmi, K.A., Stricker, D., Liwicki, M., Afzal, M.N., and Afzal, M.Z. (2021). Guided Table Structure Recognition through Anchor Optimization. arXiv.","DOI":"10.1109\/ACCESS.2021.3103413"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Xie, S., Girshick, R., Doll\u00e1r, P., Tu, Z., and He, K. (2016). Aggregated Residual Transformations for Deep Neural Networks. arXiv.","DOI":"10.1109\/CVPR.2017.634"},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"3349","DOI":"10.1109\/TPAMI.2020.2983686","article-title":"Deep high-resolution representationearning for visual recognition","volume":"43","author":"Wang","year":"2020","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Qiao, S., Chen, L.C., and Yuille, A. (2020). Detectors: Detecting objects with recursive feature pyramid and switchable atrous convolution. arXiv.","DOI":"10.1109\/CVPR46437.2021.01008"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Cai, Z., and Vasconcelos, N. (2018, January 30\u201331). Cascade r-cnn: Delving into high quality object detection. Proceedings of the IEEE Conference Computer vision pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00644"},{"key":"ref_26","unstructured":"Itonori, K. (1993, January 20\u201322). Table structure recognition based on textblock arrangement and ruled ine position. Proceedings of the 2nd International Conference on Document Analysis and Recognition (ICDAR\u201993), Tsukuba City, Japan."},{"key":"ref_27","unstructured":"Chandran, S., and Kasturi, R. (1993, January 20\u201322). Structural recognition of tabulated data. Proceedings of the 2nd International Conference on Document Analysis and Recognition (ICDAR\u201993), Sukuba, Japan."},{"key":"ref_28","unstructured":"Hirayama, Y. (1995, January 14\u201315). A method for table structure analysis using DP matching. Proceedings of the 3rd International Conference on Document Analysis and Recognition, Montreal, QC, Canada."},{"key":"ref_29","unstructured":"Green, E., and Krishnamoorthy, M. (1995, January 24\u201326). Recognition of tables using table grammars. Proceedings of the 4th Annual Symposium Document Analysis Information Retrieval, Desert Inn Hotel, Las Vegas, NV, USA."},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Huang, Y., Yan, Q., Li, Y., Chen, Y., Wang, X., Gao, L., and Tang, Z. (2019, January 20\u201325). A YOLO-based table detection method. Proceedings of the International Conference Document Analysis Recognition (ICDAR), Sydney, Australia.","DOI":"10.1109\/ICDAR.2019.00135"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Casado-Garc\u00eda, \u00c1., Dom\u00ednguez, C., Heras, J., Mata, E., and Pascual, V. (2020). The benefits of close-domain fine-tuning for table detection in document images. International Workshop Document Analysis System, Springer.","DOI":"10.1007\/978-3-030-57058-3_15"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Arif, S., and Shafait, F. (2018, January 10\u201313). Table detection in document images using foreground and background features. Proceedings of the Digital Image Computing: Techniques Applied (DICTA), Canberra, Australia.","DOI":"10.1109\/DICTA.2018.8615795"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Sun, N., Zhu, Y., and Hu, X. (2019, January 20\u201325). Faster R-CNN based table detection combining cornerocating. Proceedings of the International Conference on Document Analysis and Recognition (ICDAR), Sydney, Australia.","DOI":"10.1109\/ICDAR.2019.00212"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Qasim, S.R., Mahmood, H., and Shafait, F. (2019, January 20\u201325). Rethinking table recognition using graph neural networks. Proceedings of the International Conference on Document Analysis and Recognition (ICDAR), Sydney, Australia.","DOI":"10.1109\/ICDAR.2019.00031"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Pyreddy, P., and Croft, W.B. (1997, January 14\u201316). Tintin: A system for retrieval in text tables. Proceedings of the 2nd ACM International Conference Digit Libraries, Ottawa, ON, Canada.","DOI":"10.1145\/263690.263816"},{"key":"ref_36","doi-asserted-by":"crossref","first-page":"567","DOI":"10.1016\/j.datak.2006.04.002","article-title":"Transforming arbitrary tables intoogical form with TARTAR","volume":"3","author":"Pivk","year":"2007","journal-title":"Data Knowl. Eng."},{"key":"ref_37","doi-asserted-by":"crossref","first-page":"291","DOI":"10.1117\/12.373506","article-title":"Medium-independent table detection","volume":"Volume 3967","author":"Hu","year":"1999","journal-title":"Document Recognition Retrieval VII"},{"key":"ref_38","doi-asserted-by":"crossref","first-page":"144","DOI":"10.1007\/s10032-005-0001-x","article-title":"Design of an end-to-end method to extract information from tables","volume":"8","author":"Jorge","year":"2006","journal-title":"Int. Doc. Anal. Recognit. (IJDAR)"},{"key":"ref_39","doi-asserted-by":"crossref","first-page":"41","DOI":"10.1177\/0165551514551903","article-title":"On methods and tools of table detection, extraction and annotation in PDF documents","volume":"41","author":"Khusro","year":"2015","journal-title":"J. Information Sci."},{"key":"ref_40","doi-asserted-by":"crossref","first-page":"66","DOI":"10.1007\/s10032-006-0017-x","article-title":"Table-processing paradigms: A research survey","volume":"8","author":"Embley","year":"2006","journal-title":"Int. Doc. Anal. Recognit. (IJDAR)"},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Kieninger, T., and Dengel, A. (1998). The t-recs table recognition and analysis system. International Workshop on Document Analysis System, Springer.","DOI":"10.1007\/3-540-48172-9_21"},{"key":"ref_42","unstructured":"Cesarini, F., Marinai, S., Sarti, L., and Soda, G. (2002, January 11\u201315). Trainable tableocation in document images. Proceedings of the Object Recognition Supported User Interaction Service Robots, International Conference on Pattern Recognition, Quebec City, QC, Canada."},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Kasar, T., Barlas, P., Adam, S., Chatelain, C., and Paquet, T. (2013, January 25\u201328). Learning to detect tables in scanned document images usingine information. Proceedings of the 12th International Conference on Document Analysis and Recognition, Washington, DC, USA.","DOI":"10.1109\/ICDAR.2013.240"},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"e Silva, A.C. (2009, January 26\u201329). Learning rich hidden markov models in document analysis: Table ocation. Proceedings of the 10th International Conference on Document Analysis and Recognition, Barcelona, Spain.","DOI":"10.1109\/ICDAR.2009.185"},{"key":"ref_45","unstructured":"Silva, A. (2010). Parts That Add Up to a Whole: A Framework for the Analysis of Tables, Edinburgh University."},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Hao, L., Gao, L., Yi, X., and Tang, Z. (2016, January 11\u201314). A table detection method for pdf documents based on convolutional neural networks. Proceedings of the 12th IAPR Workshop Document Analysis System (DAS), Santorini, Greece.","DOI":"10.1109\/DAS.2016.23"},{"key":"ref_47","doi-asserted-by":"crossref","unstructured":"Kavasidis, I., Palazzo, S., Spampinato, C., Pino, C., Giordano, D., Giuffrida, D., and Messina, P. (2018). A saliency-based convolutional neural network for table and chart detection in digitized documents. arXiv.","DOI":"10.1007\/978-3-030-30645-8_27"},{"key":"ref_48","doi-asserted-by":"crossref","unstructured":"Paliwal, S.S., Vishwanath, D., Rahul, R., Sharma, M., and Vig, L. (2019, January 20\u201325). Tablenet: Deepearning model for end-to-end table detection and tabular data extraction from scanned document images. Proceedings of the International Conference on Document Analysis and Recognition (ICDAR), Sydney, Australia.","DOI":"10.1109\/ICDAR.2019.00029"},{"key":"ref_49","doi-asserted-by":"crossref","unstructured":"Hole\u010dek, M., Hoskovec, A., Baudi\u0161, P., and Klinger, P. (2019, January 20\u201325). Table understanding in structured documents. Proceedings of the International Conference on Document Analysis and Recognition Workshops (ICDARW), Sydney, Australia.","DOI":"10.1109\/ICDARW.2019.40098"},{"key":"ref_50","unstructured":"Ren, S., He, K., Girshick, R., and Sun, J. (2015). Faster r-cnn: Towards real-time object detection with region proposal networks. arXiv."},{"key":"ref_51","doi-asserted-by":"crossref","unstructured":"Zeiler, M.D., and Fergus, R. (2014, January 6\u201312). Visualizing and understanding convolutional networks. Proceedings of the European Conference Computer Vision, Zurich, Switzerland.","DOI":"10.1007\/978-3-319-10590-1_53"},{"key":"ref_52","unstructured":"Simonyan, K., and Zisserman, A. (2014). Very deep convolutional networks forarge-scale image recognition. arXiv."},{"key":"ref_53","doi-asserted-by":"crossref","unstructured":"Dai, J., Qi, H., Xiong, Y., Li, Y., Zhang, G., Hu, H., and Wei, Y. (2017, January 22\u201329). Deformable convolutional networks. Proceedings of the IEEE International Conference Computer vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.89"},{"key":"ref_54","doi-asserted-by":"crossref","unstructured":"Saha, R., Mondal, A., and Jawahar, C. (2019, January 20\u201325). Graphical object detection in document images. Proceedings of the International Conference on Document Analysis and Recognition (ICDAR), Sydney, Australia.","DOI":"10.1109\/ICDAR.2019.00018"},{"key":"ref_55","doi-asserted-by":"crossref","unstructured":"He, K., Gkioxari, G., Doll\u00e1r, P., and Girshick, R. (2017, January 22\u201329). Mask r-cnn. Proceedings of the IEEE International Conference Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.322"},{"key":"ref_56","doi-asserted-by":"crossref","unstructured":"Zhong, X., ShafieiBavani, E., and Yepes, A.J. (2019). Image-based table recognition: Data, model, and evaluation. arXiv.","DOI":"10.1007\/978-3-030-58589-1_34"},{"key":"ref_57","unstructured":"Redmon, J., and Farhadi, A. (2018). Yolov3: An incremental improvement. arXiv."},{"key":"ref_58","doi-asserted-by":"crossref","unstructured":"Liu, W., Anguelov, D., Erhan, D., Szegedy, C., Reed, S., Fu, C.Y., and Berg, A.C. (2016, January 11\u201314). Ssd: Single shot multibox detector. Proceedings of the European Conference Computer Vision, Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46448-0_2"},{"key":"ref_59","doi-asserted-by":"crossref","unstructured":"Lin, T.Y., Goyal, P., Girshick, R., He, K., and Doll\u00e1r, P. (2017, January 22\u201329). Focaloss for dense object detection. Proceedings of the IEEE International Conference Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.324"},{"key":"ref_60","doi-asserted-by":"crossref","unstructured":"Chen, K., Pang, J., Wang, J., Xiong, Y., Li, X., Sun, S., Feng, W., Liu, Z., Shi, J., and Ouyang, W. (2019, January 16\u201320). Hybrid task cascade for instance segmentation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00511"},{"key":"ref_61","doi-asserted-by":"crossref","unstructured":"Lin, T.Y., Doll\u00e1r, P., Girshick, R., He, K., Hariharan, B., and Belongie, S. (2017, January 21\u201326). Feature pyramid networks for object detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.106"},{"key":"ref_62","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (July, January 26). Deep residualearning for image recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA."},{"key":"ref_63","doi-asserted-by":"crossref","first-page":"834","DOI":"10.1109\/TPAMI.2017.2699184","article-title":"Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs","volume":"40","author":"Chen","year":"2017","journal-title":"IEEE Trans. Pattern Analysis Mach. Intell."},{"key":"ref_64","doi-asserted-by":"crossref","unstructured":"Lin, T.Y., Maire, M., Belongie, S., Bourdev, L., Girshick, R., Hays, J., Perona, P., Ramanan, D., Zitnick, C.L., and Dollar, P. (2019). Microsoft COCO: Common objects in context (2014). arXiv.","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"ref_65","doi-asserted-by":"crossref","unstructured":"Gao, L., Huang, Y., D\u00e9jean, H., Meunier, J.L., Yan, Q., Fang, Y., Kleber, F., and Lang, E. (2019, January 20\u201325). ICDAR 2019 competition on table detection and recognition (cTDaR). Proceedings of the International Conference on Document Analysis and Recognition (ICDAR), Sydney, Australia.","DOI":"10.1109\/ICDAR.2019.00243"},{"key":"ref_66","unstructured":"Li, M., Cui, L., Huang, S., Wei, F., Zhou, M., and Li, Z. (2020, January 11\u201316). Tablebank: Table benchmark for image-based table detection and recognition. Proceedings of the 12th Language Resource Evaluation Conference, Marseille, France."},{"key":"ref_67","doi-asserted-by":"crossref","unstructured":"Shahab, A., Shafait, F., Kieninger, T., and Dengel, A. (2010, January 9\u201310). An open approach towards the benchmarking of table structure recognition systems. Proceedings of the 9th IAPR International Workshop Document Analysis System, Boston, MA, USA.","DOI":"10.1145\/1815330.1815345"},{"key":"ref_68","doi-asserted-by":"crossref","unstructured":"Fang, J., Tao, X., Tang, Z., Qiu, R., and Liu, Y. (2012, January 27\u201329). Dataset, ground-truth and performance metrics for table detection evaluation. Proceedings of the 10th IAPR International Workshop Document Analysis System, Gold Coast, Australia.","DOI":"10.1109\/DAS.2012.29"},{"key":"ref_69","unstructured":"Chen, K., Wang, J., Pang, J., Cao, Y., Xiong, Y., Li, X., Sun, S., Feng, W., Liu, Z., and Xu, J. (2019). MMDetection: Open MMLab Detection Toolbox and Benchmark. arXiv."},{"key":"ref_70","first-page":"1097","article-title":"Imagenet classification with deep convolutional neural networks","volume":"25","author":"Krizhevsky","year":"2012","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_71","unstructured":"Powers, D.M. (2020). Evaluation: From precision, recall and F-measure to ROC, informedness, markedness and correlation. arXiv."},{"key":"ref_72","doi-asserted-by":"crossref","unstructured":"Blaschko, M.B., and Lampert, C.H. (2008, January 12\u201318). Learning toocalize objects with structured output regression. Proceedings of the European Conference Computer Vision, Marseille, France.","DOI":"10.1007\/978-3-540-88682-2_2"},{"key":"ref_73","unstructured":"Zhu, X., Su, W., Lu, L., Li, B., Wang, X., and Dai, J. (2020). Deformable detr: Deformable transformers for end-to-end object detection. arXiv."},{"key":"ref_74","doi-asserted-by":"crossref","unstructured":"Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., and Guo, B. (2021). Swin transformer: Hierarchical vision transformer using shifted windows. arXiv.","DOI":"10.1109\/ICCV48922.2021.00986"}],"container-title":["Journal of Imaging"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2313-433X\/7\/10\/214\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T07:16:19Z","timestamp":1760166979000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2313-433X\/7\/10\/214"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,10,16]]},"references-count":74,"journal-issue":{"issue":"10","published-online":{"date-parts":[[2021,10]]}},"alternative-id":["jimaging7100214"],"URL":"https:\/\/doi.org\/10.3390\/jimaging7100214","relation":{"has-preprint":[{"id-type":"doi","id":"10.20944\/preprints202109.0059.v1","asserted-by":"object"}]},"ISSN":["2313-433X"],"issn-type":[{"value":"2313-433X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,10,16]]}}}