{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,28]],"date-time":"2026-08-28T16:47:49Z","timestamp":1787935669549,"version":"build-2784847793"},"reference-count":47,"publisher":"Springer Science and Business Media LLC","issue":"23","license":[{"start":{"date-parts":[[2020,5,9]],"date-time":"2020-05-09T00:00:00Z","timestamp":1588982400000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2020,5,9]],"date-time":"2020-05-09T00:00:00Z","timestamp":1588982400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100001823","name":"Ministerstvo \u0160kolstv\u00ec, Ml\u00e1de\u017ee a T\u011blov\u00fdchovy","doi-asserted-by":"publisher","award":["CZ.02.1.01\/0.0\/0.0\/17_048\/0007267"],"award-info":[{"award-number":["CZ.02.1.01\/0.0\/0.0\/17_048\/0007267"]}],"id":[{"id":"10.13039\/501100001823","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Neural Comput &amp; Applic"],"published-print":{"date-parts":[[2020,12]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>As the number of digitized historical documents has increased rapidly during the last a few decades, it is necessary to provide efficient methods of information retrieval and knowledge extraction to make the data accessible. Such methods are dependent on optical character recognition (OCR) which converts the document images into textual representations. Nowadays, OCR methods are often not adapted to the historical domain; moreover, they usually need a significant amount of annotated documents. Therefore, this paper introduces a set of methods that allows performing an OCR on historical document images using only a small amount of real, manually annotated training data. The presented complete OCR system includes two main tasks: page layout analysis including text block and line segmentation and OCR. Our segmentation methods are based on fully convolutional networks, and the OCR approach utilizes recurrent neural networks. Both approaches are state of the art in the relevant fields. We have created a novel real dataset for OCR from Porta fontium portal. This corpus is freely available for research, and all proposed methods are evaluated on these data. We show that both the segmentation and OCR tasks are feasible with only a few annotated real data samples. The experiments aim at determining the best way how to achieve good performance with the given small set of data. We also demonstrate that obtained scores are comparable or even better than the scores of several state-of-the-art systems. To sum up, this paper shows a way how to create an efficient OCR system for historical documents with a need for only a little annotated training data.<\/jats:p>","DOI":"10.1007\/s00521-020-04910-x","type":"journal-article","created":{"date-parts":[[2020,5,9]],"date-time":"2020-05-09T04:02:48Z","timestamp":1588996968000},"page":"17209-17227","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":54,"title":["Building an efficient OCR system for historical documents with little training data"],"prefix":"10.1007","volume":"32","author":[{"given":"Ji\u0159\u00ed","family":"Mart\u00ednek","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ladislav","family":"Lenc","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Pavel","family":"Kr\u00e1l","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2020,5,9]]},"reference":[{"key":"4910_CR1","unstructured":"Pascanu R, Mikolov T, Bengio Y (2013) On the difficulty of training recurrent neural networks. In: International conference on machine learning, pp 1310\u20131318"},{"key":"4910_CR2","doi-asserted-by":"crossref","unstructured":"Graves A, Fern\u00e1ndez S, Gomez F, Schmidhuber J (2006) Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks. In: Proceedings of the 23rd international conference on machine learning (ACM), pp 369\u2013376","DOI":"10.1145\/1143844.1143891"},{"issue":"3","key":"4910_CR3","doi-asserted-by":"publisher","first-page":"285","DOI":"10.1007\/s10032-019-00332-1","volume":"22","author":"T Gr\u00fcning","year":"2019","unstructured":"Gr\u00fcning T, Leifert G, Strau\u00df T, Michael J, Labahn R (2019) A two-stage method for text line detection in historical documents. Int J Doc Anal Recognit (IJDAR) 22(3):285","journal-title":"Int J Doc Anal Recognit (IJDAR)"},{"key":"4910_CR4","doi-asserted-by":"crossref","unstructured":"Breuel TM, Ul-Hasan A, Azawi MIAA, Shafait F (2013) High-performance OCR for printed English and Fraktur using LSTM networks. In: 2013 12th international conference on document analysis and recognition, pp 683\u2013687","DOI":"10.1109\/ICDAR.2013.140"},{"issue":"11","key":"4910_CR5","doi-asserted-by":"publisher","first-page":"2298","DOI":"10.1109\/TPAMI.2016.2646371","volume":"39","author":"B Shi","year":"2017","unstructured":"Shi B, Bai X, Yao C (2017) An end-to-end trainable neural network for image-based sequence recognition and its application to scene text recognition. IEEE Trans Pattern Anal Mach Intell 39(11):2298","journal-title":"IEEE Trans Pattern Anal Mach Intell"},{"key":"4910_CR6","doi-asserted-by":"crossref","unstructured":"Sabir E, Rawls S, Natarajan P (2017) Implicit Language Model in LSTM for OCR. In: 2017 14th IAPR international conference on document analysis and recognition (ICDAR), vol\u00a07. IEEE, pp 27\u201331","DOI":"10.1109\/ICDAR.2017.361"},{"key":"4910_CR7","unstructured":"Karpathy A, Johnson J, Fei-Fei L (2015) Visualizing and understanding recurrent networks. arXiv preprint arXiv:1506.02078"},{"key":"4910_CR8","doi-asserted-by":"crossref","unstructured":"Ul-Hasan A, Breuel TM (2013) Can we build language-independent OCR using LSTM networks?. In: Proceedings of the 4th international workshop on multilingual OCR , pp 1\u20135","DOI":"10.1145\/2505377.2505394"},{"key":"4910_CR9","doi-asserted-by":"crossref","unstructured":"Isola P, Zhu JY, Zhou T, Efros AA (2017) Image-to-image translation with conditional adversarial networks. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 1125\u20131134","DOI":"10.1109\/CVPR.2017.632"},{"key":"4910_CR10","doi-asserted-by":"crossref","unstructured":"Afzal MZ, Pastor-Pellicer J, Shafait F, Breuel TM, Dengel A, Liwicki M (2015) Document image binarization using lstm: A sequence learning approach. In: Proceedings of the 3rd international workshop on historical document imaging and processing (ACM), pp 79\u201384","DOI":"10.1145\/2809544.2809561"},{"key":"4910_CR11","doi-asserted-by":"crossref","unstructured":"Ronneberger O, Fischer P, Brox T (2015) U-net: Convolutional networks for biomedical image segmentation. In: International conference on medical image computing and computer-assisted intervention. Springer, pp 234\u2013241","DOI":"10.1007\/978-3-319-24574-4_28"},{"key":"4910_CR12","doi-asserted-by":"crossref","unstructured":"Xie S, Tu Z (2015) Holistically-nested edge detection. In: Proceedings of the IEEE international conference on computer vision, pp 1395\u20131403","DOI":"10.1109\/ICCV.2015.164"},{"key":"4910_CR13","doi-asserted-by":"crossref","unstructured":"Huang X, Liu MY, Belongie S, Kautz J (2018) Multimodal Unsupervised Image-to-image Translation. In: The European conference on computer vision (ECCV)","DOI":"10.1007\/978-3-030-01219-9_11"},{"key":"4910_CR14","doi-asserted-by":"crossref","unstructured":"Bukhari SS, Shafait F, Breuel TM (2011) Improved document image segmentation algorithm using multiresolution morphology. Document recognition and retrieval XVIII, vol 7874. International Society for Optics and Photonics, p 78740D","DOI":"10.1117\/12.873461"},{"issue":"4","key":"4910_CR15","doi-asserted-by":"publisher","first-page":"640","DOI":"10.1109\/TPAMI.2016.2572683","volume":"39","author":"E Shelhamer","year":"2017","unstructured":"Shelhamer E, Long J, Darrell T (2017) Fully convolutional networks for semantic segmentation. IEEE Trans Pattern Anal Mach Intell 39(4):640. https:\/\/doi.org\/10.1109\/TPAMI.2016.2572683","journal-title":"IEEE Trans Pattern Anal Mach Intell"},{"key":"4910_CR16","doi-asserted-by":"crossref","unstructured":"He K, Gkioxari G, Doll\u00e1r P, Girshick R (2017) Mask r-cnn. In: Proceedings of the IEEE international conference on computer vision, pp 2961\u20132969","DOI":"10.1109\/ICCV.2017.322"},{"key":"4910_CR17","doi-asserted-by":"crossref","unstructured":"Breuel TM (2017) Robust, simple page segmentation using hybrid convolutional mdlstm networks. In: 2017 14th IAPR international conference on document analysis and recognition (ICDAR), vol\u00a01. IEEE, pp 733\u2013740","DOI":"10.1109\/ICDAR.2017.125"},{"issue":"8","key":"4910_CR18","doi-asserted-by":"publisher","first-page":"1735","DOI":"10.1162\/neco.1997.9.8.1735","volume":"9","author":"S Hochreiter","year":"1997","unstructured":"Hochreiter S, Schmidhuber J (1997) Long short-term memory. Neural Comput 9(8):1735","journal-title":"Neural Comput"},{"key":"4910_CR19","doi-asserted-by":"crossref","unstructured":"Breuel TM (2017) High performance text recognition using a hybrid convolutional-lstm implementation. In: 2017 14th IAPR international conference on document analysis and recognition (ICDAR), vol\u00a01. IEEE, pp 11\u201316","DOI":"10.1109\/ICDAR.2017.12"},{"issue":"10","key":"4910_CR20","first-page":"1995","volume":"3361","author":"Y LeCun","year":"1995","unstructured":"LeCun Y, Bengio Y et al (1995) Convolutional networks for images, speech, and time series. Handb Brain Theory Neural Netw 3361(10):1995","journal-title":"Handb Brain Theory Neural Netw"},{"key":"4910_CR21","doi-asserted-by":"crossref","unstructured":"Elagouni K, Garcia C, Mamalet F, S\u00e9billot P (2012) Text recognition in videos using a recurrent connectionist approach. In: International conference on artificial neural networks. Springer, pp 172\u2013179","DOI":"10.1007\/978-3-642-33266-1_22"},{"key":"4910_CR22","doi-asserted-by":"crossref","unstructured":"He P, Huang W, Qiao Y, Loy CC, Tang X (2016) Reading scene text in deep convolutional sequences. In: Thirtieth AAAI conference on artificial intelligence","DOI":"10.1609\/aaai.v30i1.10465"},{"key":"4910_CR23","doi-asserted-by":"crossref","unstructured":"Graves A (2012) Sequence transduction with recurrent neural networks. arXiv preprint arXiv:1211.3711","DOI":"10.1007\/978-3-642-24797-2_3"},{"key":"4910_CR24","unstructured":"Bahdanau D, Cho K, Bengio Y (2014) Neural machine translation by jointly learning to align and translate. arXiv preprint arXiv:1409.0473"},{"key":"4910_CR25","doi-asserted-by":"crossref","unstructured":"Bluche T, Louradour J, Messina R (2017) Scan, attend and read: End-to-end handwritten paragraph recognition with mdlstm attention. In: 2017 14th IAPR international conference on document analysis and recognition (ICDAR), vol\u00a01. IEEE, pp 1050\u20131055","DOI":"10.1109\/ICDAR.2017.174"},{"key":"4910_CR26","unstructured":"Jaderberg M, Simonyan K, Vedaldi A, Zisserman A (2014) Synthetic data and artificial neural networks for natural scene text recognition. arXiv preprint arXiv:1406.2227"},{"key":"4910_CR27","doi-asserted-by":"crossref","unstructured":"Margner V, Pechwitz M (2001) Synthetic data for Arabic OCR system development. In: Sixth international conference on Document analysis and recognition, 2001. Proceedings. IEEE, pp 1159\u20131163","DOI":"10.1109\/ICDAR.2001.953967"},{"key":"4910_CR28","doi-asserted-by":"crossref","unstructured":"Gaur S, Sonkar S, Roy PP (2015) Generation of synthetic training data for handwritten Indic script recognition. In: 2015 13th international conference on document analysis and recognition (ICDAR). IEEE, pp 491\u2013495","DOI":"10.1109\/ICDAR.2015.7333810"},{"key":"4910_CR29","unstructured":"Perez L, Wang J (2017) The effectiveness of data augmentation in image classification using deep learning. arXiv preprint arXiv:1712.04621"},{"key":"4910_CR30","unstructured":"Clausner C, Pletschacher S, Antonacopoulos A (2014) Efficient OCR training data generation with aletheia. In: Proceedings of the international association for pattern recognition (IAPR), Tours, France pp 7\u201310"},{"key":"4910_CR31","doi-asserted-by":"crossref","unstructured":"Pletschacher S, Antonacopoulos A (2010) The page (page analysis and ground-truth elements) format framework. In: 2010 20th international conference on pattern recognition (IEEE), pp 257\u2013260","DOI":"10.1109\/ICPR.2010.72"},{"key":"4910_CR32","doi-asserted-by":"crossref","unstructured":"Breuel TM (2008) The OCRopus open source OCR system. Document recognition and retrieval XV, vol 6815. International Society for Optics and Photonics, p 68150F","DOI":"10.1117\/12.783598"},{"key":"4910_CR33","unstructured":"Vincent L, Lead UT (2006) Announcing tesseract OCR, Google Code. http:\/\/googlecode.blogspot.com.au\/2006\/08\/announcing-tesseract-ocr.html. Accessed 1 Nov 2015"},{"key":"4910_CR34","unstructured":"Leifert G, Strauss T, Gr\u00fcning T, Labahn R (2016) Citlab argus for historical handwritten documents"},{"key":"4910_CR35","unstructured":"Strauss T, Weidemann M, Michael J, Leifert G, Gr\u00fcning T, Labahn R (2018) System description of citlab\u2019s recognition & retrieval engine for ICDAR 2017 competition on information extraction in historical handwritten records"},{"key":"4910_CR36","doi-asserted-by":"crossref","unstructured":"Wick C, Puppe F (2018) Fully convolutional neural networks for page segmentation of historical document images. In: 2018 13th IAPR international workshop on document analysis systems (DAS). IEEE, pp 287\u2013292","DOI":"10.1109\/DAS.2018.39"},{"issue":"1","key":"4910_CR37","doi-asserted-by":"publisher","first-page":"62","DOI":"10.1109\/TSMC.1979.4310076","volume":"9","author":"N Otsu","year":"1979","unstructured":"Otsu N (1979) A threshold selection method from gray-level histograms. IEEE Trans Syst Man Cybern 9(1):62","journal-title":"IEEE Trans Syst Man Cybern"},{"key":"4910_CR38","doi-asserted-by":"publisher","unstructured":"Mart\u00ednek J, Lenc L, Kr\u00e1l P, Nicolaou A, Christlein V (2019) Hybrid Training Data for Historical Text OCR. In: 15th international conference on document analysis and recognition (ICDAR 2019), Sydney, Australia, pp 565\u2013570. https:\/\/doi.org\/10.1109\/ICDAR.2019.00096","DOI":"10.1109\/ICDAR.2019.00096"},{"key":"4910_CR39","unstructured":"Glorot X, Bengio Y (2010) Understanding the difficulty of training deep feedforward neural networks. In: Proceedings of the thirteenth international conference on artificial intelligence and statistics, pp 249\u2013256"},{"key":"4910_CR40","doi-asserted-by":"crossref","unstructured":"Clausner C, Papadopoulos C, Pletschacher S, Antonacopoulos A (2015) The ENP image and ground truth dataset of historical newspapers. In: 2015 13th international conference on document analysis and recognition (ICDAR). IEEE, pp 931\u2013935","DOI":"10.1109\/ICDAR.2015.7333898"},{"key":"4910_CR41","unstructured":"Tong X, Evans DA (1996) A statistical approach to automatic OCR error correction in context. In: Fourth workshop on very large corpora"},{"key":"4910_CR42","first-page":"361","volume":"5","author":"DD Lewis","year":"2004","unstructured":"Lewis DD, Yang Y, Rose TG, Li F (2004) RCV1: a new benchmark collection for text categorization research. J Mach Learn Res 5:361","journal-title":"J Mach Learn Res"},{"key":"4910_CR43","doi-asserted-by":"crossref","unstructured":"Oquab M, Bottou L, Laptev I, Sivic J (2014) Learning and transferring mid-level image representations using convolutional neural networks. In: The IEEE conference on computer vision and pattern recognition (CVPR)","DOI":"10.1109\/CVPR.2014.222"},{"key":"4910_CR44","unstructured":"Shang W, Sohn K, Almeida D, Lee H (2016) Understanding and improving convolutional neural networks via concatenated rectified linear units. In: International conference on machine learning, pp 2217\u20132225"},{"key":"4910_CR45","unstructured":"Kingma DP, Ba J (2014) Adam: a method for stochastic optimization. arXiv preprint arXiv:1412.6980"},{"key":"4910_CR46","doi-asserted-by":"publisher","unstructured":"Alberti M, Bouillon M, Ingold R, Liwicki M (2017) Open evaluation tool for layout analysis of document images. In: 2017 14th IAPR international conference on document analysis and recognition (ICDAR), Kyoto, Japan, pp 43\u201347. https:\/\/doi.org\/10.1109\/ICDAR.2017.311","DOI":"10.1109\/ICDAR.2017.311"},{"issue":"1","key":"4910_CR47","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1007\/s11263-015-0823-z","volume":"116","author":"M Jaderberg","year":"2016","unstructured":"Jaderberg M, Simonyan K, Vedaldi A, Zisserman A (2016) Reading text in the wild with convolutional neural networks. Int J Comput Vis 116(1):1","journal-title":"Int J Comput Vis"}],"updated-by":[{"DOI":"10.1007\/s00521-020-05563-6","type":"correction","label":"Correction","source":"publisher","updated":{"date-parts":[[2021,3,10]],"date-time":"2021-03-10T00:00:00Z","timestamp":1615334400000}}],"container-title":["Neural Computing and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s00521-020-04910-x.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s00521-020-04910-x\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s00521-020-04910-x.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,10,23]],"date-time":"2022-10-23T09:58:39Z","timestamp":1666519119000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s00521-020-04910-x"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,5,9]]},"references-count":47,"journal-issue":{"issue":"23","published-print":{"date-parts":[[2020,12]]}},"alternative-id":["4910"],"URL":"https:\/\/doi.org\/10.1007\/s00521-020-04910-x","relation":{"correction":[{"id-type":"doi","id":"10.1007\/s00521-020-05563-6","asserted-by":"object"}]},"ISSN":["0941-0643","1433-3058"],"issn-type":[{"value":"0941-0643","type":"print"},{"value":"1433-3058","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,5,9]]},"assertion":[{"value":"25 December 2019","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"6 April 2020","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"9 May 2020","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"10 March 2021","order":4,"name":"change_date","label":"Change Date","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"Correction","order":5,"name":"change_type","label":"Change Type","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"A Correction to this paper has been published:","order":6,"name":"change_details","label":"Change Details","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"https:\/\/doi.org\/10.1007\/s00521-020-05563-6","URL":"https:\/\/doi.org\/10.1007\/s00521-020-05563-6","order":7,"name":"change_details","label":"Change Details","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Compliance with ethical standards"}},{"value":"The authors whose names are listed in this paper certify that they have NO affiliations with or involvement in any organization or entity with any financial interest (such as honoraria; educational grants; participation in speakers\u2019 bureaus; membership, employment, consultancies, stock ownership or other equity interest; and expert testimony or patent-licencing arrangements) or non-financial interest (such as personal or professional relationships, affiliations, knowledge or beliefs) in the subject matter or materials discussed in this manuscript.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}}]}}