{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,22]],"date-time":"2026-07-22T23:07:10Z","timestamp":1784761630155,"version":"3.55.0"},"reference-count":33,"publisher":"Springer Science and Business Media LLC","issue":"2","license":[{"start":{"date-parts":[[2025,12,1]],"date-time":"2025-12-01T00:00:00Z","timestamp":1764547200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2025,12,1]],"date-time":"2025-12-01T00:00:00Z","timestamp":1764547200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100000270","name":"Natural Environment Research Council","doi-asserted-by":"publisher","award":["NE\/S015604\/1"],"award-info":[{"award-number":["NE\/S015604\/1"]}],"id":[{"id":"10.13039\/501100000270","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100000270","name":"Natural Environment Research Council","doi-asserted-by":"publisher","award":["NE\/S015604\/1"],"award-info":[{"award-number":["NE\/S015604\/1"]}],"id":[{"id":"10.13039\/501100000270","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["IJDAR"],"published-print":{"date-parts":[[2026,6]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>\n                    This study uses a novel semi-supervised learning framework to explore Tabular Structure Recognition (TSR) for digitizing historical documents, specifically employing the CascadeTabNet model. TSR is crucial for transforming archival tabular data into digital formats, enhancing accessibility and analysis across various research fields. Challenges like physical degradation, inconsistent lighting, and non-standard handwriting hinder the generation of high-quality annotations of historical documents needed for effective model training. To address these issues, this research explores two research questions: (i)\n                    <jats:italic>Can a semi-supervised training approach reduce the need for expensive data annotations?<\/jats:italic>\n                    and (ii)\n                    <jats:italic>Does semi-supervised training improve model robustness?<\/jats:italic>\n                    We applied our methodology across three datasets: the GloSAT and ICDAR-2019 datasets based on historical documents, and the predominantly modern documents PubTabNet dataset. Our results indicate that semi-supervised learning substantially increases TSR accuracy and decreases dependency on extensive labelled datasets, providing a robust solution for large-scale digitization initiatives and contributing to the preservation and improved accessibility of historical data. All code from this paper is freely available on GitHub (\n                    <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"https:\/\/github.com\/stuartemiddleton\/glosat_table_dataset\" ext-link-type=\"uri\">https:\/\/github.com\/stuartemiddleton\/glosat_table_dataset<\/jats:ext-link>\n                    ).\n                  <\/jats:p>","DOI":"10.1007\/s10032-025-00562-6","type":"journal-article","created":{"date-parts":[[2025,12,1]],"date-time":"2025-12-01T17:51:40Z","timestamp":1764611500000},"page":"429-445","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":1,"title":["Data rescue of historical tables through semi-supervised table structure recognition"],"prefix":"10.1007","volume":"29","author":[{"given":"Loitongbam Gyanendro","family":"Singh","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Stuart E.","family":"Middleton","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2025,12,1]]},"reference":[{"issue":"1","key":"562_CR1","doi-asserted-by":"publisher","first-page":"28","DOI":"10.1002\/gdj3.56","volume":"5","author":"S Br\u00f6nnimann","year":"2018","unstructured":"Br\u00f6nnimann, S., Brugnara, Y., Allan, R.J., Brunet, M., Compo, G.P., Crouthamel, R.I., Jones, P.D., Jourdain, S., Luterbacher, J., Siegmund, P., et al.: A roadmap to climate data rescue services. Geosci. Data J. 5(1), 28\u201339 (2018)","journal-title":"Geosci. Data J."},{"issue":"1\u20132","key":"562_CR2","doi-asserted-by":"publisher","first-page":"29","DOI":"10.3354\/cr00960","volume":"47","author":"M Brunet","year":"2011","unstructured":"Brunet, M., Jones, P.: Data rescue initiatives: bringing historical climate data into the 21st century. Climate Res. 47(1\u20132), 29\u201340 (2011)","journal-title":"Climate Res."},{"issue":"11","key":"562_CR3","doi-asserted-by":"publisher","first-page":"40","DOI":"10.1175\/2011BAMS3124.1","volume":"92","author":"PW Thorne","year":"2011","unstructured":"Thorne, P.W., Willett, K.M., Allan, R.J., Bojinski, S., Christy, J.R., Fox, N., Gilbert, S., Jolliffe, I., Kennedy, J.J., Kent, E., et al.: Guiding the creation of a comprehensive surface temperature resource for twenty-first-century climate science. Bull. Am. Meteor. Soc. 92(11), 40\u201347 (2011)","journal-title":"Bull. Am. Meteor. Soc."},{"issue":"10","key":"562_CR4","doi-asserted-by":"publisher","first-page":"1483","DOI":"10.1175\/BAMS-85-10-1483","volume":"85","author":"CM Page","year":"2004","unstructured":"Page, C.M., Nicholls, N., Plummer, N., Trewin, B., Manton, M., Alexander, L., Chambers, L.E., Choi, Y., Collins, D.A., Gosai, A., et al.: Data rescue in the southeast asia and south pacific region: challenges and opportunities. Bull. Am. Meteor. Soc. 85(10), 1483\u20131490 (2004)","journal-title":"Bull. Am. Meteor. Soc."},{"issue":"3","key":"562_CR5","doi-asserted-by":"publisher","first-page":"39","DOI":"10.3390\/cli12030039","volume":"12","author":"J Luterbacher","year":"2024","unstructured":"Luterbacher, J., Allan, R., Wilkinson, C., Hawkins, E., Teleti, P., Lorrey, A., Br\u00f6nnimann, S., Hechler, P., Velikou, K., Xoplaki, E.: The importance and scientific value of long weather and climate records; examples of historical marine data efforts across the globe. Climate 12(3), 39 (2024)","journal-title":"Climate"},{"issue":"4","key":"562_CR6","doi-asserted-by":"publisher","first-page":"1465","DOI":"10.5194\/nhess-23-1465-2023","volume":"23","author":"E Hawkins","year":"2023","unstructured":"Hawkins, E., Brohan, P., Burgess, S.N., Burt, S., Compo, G.P., Gray, S.L., Haigh, I.D., Hersbach, H., Kuijjer, K., Mart\u00ednez-Alvarado, O., et al.: Rescuing historical weather observations improves quantification of severe windstorm risks. Nat. Hazard. 23(4), 1465\u20131482 (2023)","journal-title":"Nat. Hazard."},{"issue":"7","key":"562_CR7","doi-asserted-by":"publisher","first-page":"1159","DOI":"10.1002\/asl.1159","volume":"24","author":"EL Yule","year":"2023","unstructured":"Yule, E.L., Hegerl, G., Schurer, A., Hawkins, E.: Using early extremes to place the 2022 uk heat waves into historical context. Atmos. Sci. Lett. 24(7), 1159 (2023)","journal-title":"Atmos. Sci. Lett."},{"key":"562_CR8","unstructured":"Peng, S., Chakravarthy, A., Lee, S., Wang, X., Balasubramaniyan, R., Chau, D.H.: Unitable: Towards a unified framework for table recognition via self-supervised pretraining. In: NeurIPS 2024 Third Table Representation Learning Workshop (2024)"},{"key":"562_CR9","doi-asserted-by":"crossref","unstructured":"Kim, G., Hong, T., Yim, M., Park, J., Yim, J., Hwang, W., Yun, S., Han, D., Park, S.: Donut: document understanding transformer without ocr. ArXiv abs\/2111.15664 (2021)","DOI":"10.1007\/978-3-031-19815-1_29"},{"key":"562_CR10","doi-asserted-by":"crossref","unstructured":"Fischer, P., Smajic, A., Abrami, G., Mehler, A.: Multi-type-td-tsr\u2013extracting tables from document images using a multi-stage pipeline for table detection and table structure recognition: From ocr to structured table representations. In: KI 2021: Advances in Artificial Intelligence: 44th German Conference on AI, Virtual Event, September 27\u2013October 1, 2021, Proceedings 44, pp. 95\u2013108 (2021). Springer","DOI":"10.1007\/978-3-030-87626-5_8"},{"key":"562_CR11","doi-asserted-by":"crossref","unstructured":"Raja, S., Mondal, A., Jawahar, C.: Table structure recognition using top-down and bottom-up cues. In: Computer Vision \u2013 ECCV 2020, pp. 70\u201386. Springer, Cham (2020)","DOI":"10.1007\/978-3-030-58604-1_5"},{"key":"562_CR12","doi-asserted-by":"crossref","unstructured":"Raja, S., Mondal, A., Jawahar, C.: Visual understanding of complex table structures from document images. In: Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision, pp. 2299\u20132308 (2022)","DOI":"10.1109\/WACV51458.2022.00260"},{"key":"562_CR13","first-page":"1","volume":"7","author":"R Zanibbi","year":"2004","unstructured":"Zanibbi, R., Blostein, D., Cordy, J.R.: A survey of table recognition: Models, observations, transformations, and inferences. Document Analysis and Recognition 7, 1\u201316 (2004)","journal-title":"Document Analysis and Recognition"},{"key":"562_CR14","doi-asserted-by":"crossref","unstructured":"V\u00f6gtlin, L., Scius-Bertrand, A., Maergner, P., Fischer, A., Ingold, R.: Diva-daf: a deep learning framework for historical document image analysis. In: Proceedings of the 7th International Workshop on Historical Document Imaging and Processing, pp. 61\u201366 (2023)","DOI":"10.1145\/3604951.3605511"},{"key":"562_CR15","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2022.109006","volume":"133","author":"C Ma","year":"2023","unstructured":"Ma, C., Lin, W., Sun, L., Huo, Q.: Robust table detection and structure recognition from heterogeneous document images. Pattern Recogn. 133, 109006 (2023)","journal-title":"Pattern Recogn."},{"key":"562_CR16","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2022.108946","volume":"132","author":"X-H Li","year":"2022","unstructured":"Li, X.-H., Yin, F., Dai, H.-S., Liu, C.-L.: Table structure recognition and form parsing by end-to-end object detection and relation parsing. Pattern Recogn. 132, 108946 (2022)","journal-title":"Pattern Recogn."},{"key":"562_CR17","doi-asserted-by":"crossref","unstructured":"Zhong, X., ShafieiBavani, E., Jimeno\u00a0Yepes, A.: Image-based table recognition: data, model, and evaluation. In: European Conference on Computer Vision, pp. 564\u2013580 (2020). Springer","DOI":"10.1007\/978-3-030-58589-1_34"},{"key":"562_CR18","doi-asserted-by":"crossref","unstructured":"Prasad, D., Gadpal, A., Kapadni, K., Visave, M., Sultanpure, K.: Cascadetabnet: An approach for end to end table detection and structure recognition from image-based documents. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops, pp. 572\u2013573 (2020)","DOI":"10.1109\/CVPRW50498.2020.00294"},{"key":"562_CR19","doi-asserted-by":"crossref","unstructured":"Gao, L., Huang, Y., D\u00e9jean, H., Meunier, J.-L., Yan, Q., Fang, Y., Kleber, F., Lang, E.: Icdar 2019 competition on table detection and recognition (ctdar). In: 2019 International Conference on Document Analysis and Recognition (ICDAR), pp. 1510\u20131515 (2019). IEEE","DOI":"10.1109\/ICDAR.2019.00243"},{"key":"562_CR20","doi-asserted-by":"crossref","unstructured":"Zheng, X., Burdick, D., Popa, L., Zhong, X., Wang, N.X.R.: Global table extractor (gte): a framework for joint table identification and cell structure recognition using visual context. In: Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision, pp. 697\u2013706 (2021)","DOI":"10.1109\/WACV48630.2021.00074"},{"key":"562_CR21","unstructured":"Nurminen, A.: Algorithmic extraction of data in tables in pdf documents. Master\u2019s thesis (2013)"},{"key":"562_CR22","doi-asserted-by":"crossref","unstructured":"Oro, E., Ruffolo, M.: Trex: An approach for recognizing and extracting tables from pdf documents. In: 2009 10th International Conference on Document Analysis and Recognition, pp. 906\u2013910 (2009). IEEE","DOI":"10.1109\/ICDAR.2009.12"},{"key":"562_CR23","doi-asserted-by":"crossref","unstructured":"Girshick, R.: Fast r-cnn. In: Proceedings of the IEEE International Conference on Computer Vision, pp. 1440\u20131448 (2015)","DOI":"10.1109\/ICCV.2015.169"},{"key":"562_CR24","doi-asserted-by":"crossref","unstructured":"Lin, T.-Y., Goyal, P., Girshick, R., He, K., Doll\u00e1r, P.: Focal loss for dense object detection. In: Proceedings of the IEEE International Conference on Computer Vision, pp. 2980\u20132988 (2017)","DOI":"10.1109\/ICCV.2017.324"},{"key":"562_CR25","doi-asserted-by":"crossref","unstructured":"Redmon, J., Divvala, S., Girshick, R., Farhadi, A.: You only look once: Unified, real-time object detection. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 779\u2013788 (2016)","DOI":"10.1109\/CVPR.2016.91"},{"key":"562_CR26","doi-asserted-by":"crossref","unstructured":"Paliwal, S.S., Vishwanath, D., Rahul, R., Sharma, M., Vig, L.: Tablenet: Deep learning model for end-to-end table detection and tabular data extraction from scanned document images. In: 2019 International Conference on Document Analysis and Recognition (ICDAR), pp. 128\u2013133 (2019). IEEE","DOI":"10.1109\/ICDAR.2019.00029"},{"key":"562_CR27","doi-asserted-by":"crossref","unstructured":"Li, J., Xu, Y., Lv, T., Cui, L., Zhang, C., Wei, F.: Dit: Self-supervised pre-training for document image transformer. In: Proceedings of the 30th ACM International Conference on Multimedia, pp. 3530\u20133539 (2022)","DOI":"10.1145\/3503161.3547911"},{"key":"562_CR28","doi-asserted-by":"crossref","unstructured":"Nassar, A., Livathinos, N., Lysak, M., Staar, P.: Tableformer: Table structure understanding with transformers. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pp. 4614\u20134623 (2022)","DOI":"10.1109\/CVPR52688.2022.00457"},{"key":"562_CR29","doi-asserted-by":"crossref","unstructured":"Carion, N., Massa, F., Synnaeve, G., Usunier, N., Kirillov, A., Zagoruyko, S.: End-to-end object detection with transformers. In: European Conference on Computer Vision, pp. 213\u2013229 (2020). Springer","DOI":"10.1007\/978-3-030-58452-8_13"},{"key":"562_CR30","doi-asserted-by":"crossref","unstructured":"Cai, Z., Vasconcelos, N.: Cascade r-cnn: Delving into high quality object detection. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 6154\u20136162 (2018)","DOI":"10.1109\/CVPR.2018.00644"},{"key":"562_CR31","unstructured":"Li, M., Cui, L., Huang, S., Wei, F., Zhou, M., Li, Z.: Tablebank: Table benchmark for image-based table detection and recognition. In: Proceedings of the Twelfth Language Resources and Evaluation Conference, pp. 1918\u20131925 (2020)"},{"key":"562_CR32","doi-asserted-by":"crossref","unstructured":"Ziomek, J., Middleton, S.E.: Glosat historical measurement table dataset: enhanced table structure recognition annotation for downstream historical data rescue. In: Proceedings of the 6th International Workshop on Historical Document Imaging and Processing, pp. 49\u201354 (2021)","DOI":"10.1145\/3476887.3476890"},{"issue":"10","key":"562_CR33","doi-asserted-by":"publisher","first-page":"3349","DOI":"10.1109\/TPAMI.2020.2983686","volume":"43","author":"J Wang","year":"2020","unstructured":"Wang, J., Sun, K., Cheng, T., Jiang, B., Deng, C., Zhao, Y., Liu, D., Mu, Y., Tan, M., Wang, X., et al.: Deep high-resolution representation learning for visual recognition. IEEE Trans. Pattern Anal. Mach. Intell. 43(10), 3349\u20133364 (2020)","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."}],"container-title":["International Journal on Document Analysis and Recognition (IJDAR)"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10032-025-00562-6.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10032-025-00562-6","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10032-025-00562-6.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,20]],"date-time":"2026-06-20T07:07:13Z","timestamp":1781939233000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10032-025-00562-6"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,12,1]]},"references-count":33,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2026,6]]}},"alternative-id":["562"],"URL":"https:\/\/doi.org\/10.1007\/s10032-025-00562-6","relation":{"has-preprint":[{"id-type":"doi","id":"10.21203\/rs.3.rs-5842111\/v1","asserted-by":"object"}]},"ISSN":["1433-2833","1433-2825"],"issn-type":[{"value":"1433-2833","type":"print"},{"value":"1433-2825","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,12,1]]},"assertion":[{"value":"16 January 2025","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"3 August 2025","order":2,"name":"revised","label":"Revised","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"17 November 2025","order":3,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"1 December 2025","order":4,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare no competing interests.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}]}}