{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,10]],"date-time":"2026-06-10T13:02:08Z","timestamp":1781096528444,"version":"3.54.1"},"reference-count":45,"publisher":"Springer Science and Business Media LLC","issue":"4","license":[{"start":{"date-parts":[[2013,5,29]],"date-time":"2013-05-29T00:00:00Z","timestamp":1369785600000},"content-version":"tdm","delay-in-days":0,"URL":"http:\/\/www.springer.com\/tdm"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Lang Resources &amp; Evaluation"],"published-print":{"date-parts":[[2013,12]]},"DOI":"10.1007\/s10579-013-9236-1","type":"journal-article","created":{"date-parts":[[2013,5,28]],"date-time":"2013-05-28T07:12:55Z","timestamp":1369725175000},"page":"1233-1259","source":"Crossref","is-referenced-by-count":5,"title":["Dealing with orthographic variation in a tagger-lemmatizer for fourteenth century Dutch charters"],"prefix":"10.1007","volume":"47","author":[{"given":"Hans","family":"van Halteren","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Margit","family":"Rem","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2013,5,29]]},"reference":[{"key":"9236_CR1","unstructured":"Archer, D., McEnery, T., Rayson, P., & Hardie, A. (2003). Developing an automated semantic analysis system for Early Modern English. In Proceedings of the Corpus Linguistics 2003 Conference. UCREL Technical Paper Number 16. UCREL, Lancaster University."},{"key":"9236_CR2","unstructured":"Baron, A., & Rayson, P. (2008). VARD 2: A tool for dealing with spelling variation in historical corpora. In Proceedings of the Postgraduate Conference in Corpus Linguistics. Birmingham, UK: Aston University, May 22, 2008."},{"issue":"1","key":"9236_CR3","first-page":"41","volume":"20","author":"A Baron","year":"2009","unstructured":"Baron, A., Rayson, P., & Archer, D. (2009). Word frequency and key word statistics in corpus linguistics. Anglistik, 20(1), 41\u201367.","journal-title":"Anglistik"},{"key":"9236_CR4","unstructured":"Baron, A., Rayson, P., & Archer, D. (2011). Quantifying early modern English spelling variation: Change over time and genre. Presented at Conference on New Methods in Historical Corpora, Manchester, United Kingdom, 29\u201330 April 2011."},{"key":"9236_CR5","unstructured":"Berteloot, A. (1984). Bijdrage tot een klankatlas van het dertiende-eeuwse Middelnederlands. Koninklijke Academie voor Nederlandse Taal- en Letterkunde, Gent."},{"key":"9236_CR6","unstructured":"Bollmann, M. (2012). (Semi-) Automatic normalization of historical texts using distance measures and the norma tool. In Proceedings of the Second Workshop on Annotation of Corpora for Research in the Humanities (ACRH-2). Lisbon, Portugal, November 29, 2012."},{"key":"9236_CR7","unstructured":"Bollmann, M., Petran, F., & Dipper, S. (2011). Rule-based normalization of historical texts. In Proceedings of the RANLP Workshop on Language Technologies for Digital Humanities and Cultural Heritage. Hissar, Bulgaria, pp. 34\u201342."},{"key":"9236_CR8","doi-asserted-by":"crossref","unstructured":"Brants, T. (2000). TnT\u2014a statistical part-of-speech tagger. In Proceedings of the Sixth Applied Natural Language Processing Conference. ANLP-2000, pp. 224\u2013231.","DOI":"10.3115\/974147.974178"},{"key":"9236_CR9","unstructured":"Burgers, J. W. J. (1995). De paleografie van de documentaire bronnen in Holland en Zeeland in de dertiende eeuw, 3 dln. Leuven: Schrift en Schriftdragers in de Nederlanden in de Middeleeuwen."},{"key":"9236_CR10","first-page":"709","volume":"1","author":"WW Cohen","year":"1996","unstructured":"Cohen, W. W. (1996). Learning trees and rules with set-valued features. AAAI\/IAAI, 1, 709\u2013716.","journal-title":"AAAI\/IAAI"},{"key":"9236_CR11","unstructured":"de Vries, M., te Winkel, L. A., et al. (1956). Woordenboek der Nederlandsche Taal. Delen I-XXIX.\u2019s-Gravenhage\/Leiden etc.: M. Nijhoff\/A.W. Sijthoff etc., 1882\u20131998. Supplement I.\u2019s-Gravenhage\/Leiden etc.: M. Nijhoff\/A.W. Sijthoff etc. Supplements parts I-III.\u2019s-Gravenhage: Sdu Uitgevers."},{"key":"9236_CR12","unstructured":"Depuydt, K., & de Does, J. (2009). Computational tools and lexica to improve access to text. In Beijk, E., Colman, L., et al. (Eds.), Fons Verborum, F. Feestbundel voor Prof. Dr. A.M.F.J. (Fons) Moerdijk, aangeboden door vrienden en collega\u2019s bij zijn afscheid van het INL. Leiden\/Amsterdam, pp. 187\u2013199."},{"key":"9236_CR13","unstructured":"Dijkhof, E. C. (1997). Het oorkondewezen van enige kloosters en steden in Holland en Zeeland. PhD Thesis. University of Amsterdam."},{"key":"9236_CR14","doi-asserted-by":"crossref","unstructured":"Dipper, S. (2011). Morphological and part-of-speech tagging of historical language data: A comparison. In Journal for Language Technology and Computational Linguistics, Special Issue, 26(2), 25\u201337. (=\u00a0Proceedings of the TLT-Workshop on Annotation of Corpora for Research in the Humanities, 2012).","DOI":"10.21248\/jlcl.26.2011.144"},{"key":"9236_CR15","unstructured":"Gerritsen, M. (1990). The relationship between punctuation and syntax in Middle Dutch. In Fisiak, J. (Ed.), Historical Linguistics and Philology. Trends in Linguistics. Studies and Monographs 46. Berlin: Mouton de Gruyter."},{"key":"9236_CR16","unstructured":"Gim\u00e9nez, J., & M\u00e1rquez, L. (2004). SVMTool: A general POS tagger generator based on Support Vector Machines. In Proceedings of the 4th International Conference on Language Resources and Evaluation (LREC\u201904). Lisbon, Portugal."},{"key":"9236_CR17","volume-title":"Beginselen en ontwikkeling van de interpunctie, in\u2019t biezonder in de Nederlanden","author":"J Greidanus","year":"1926","unstructured":"Greidanus, J. (1926). Beginselen en ontwikkeling van de interpunctie, in\u2019t biezonder in de Nederlanden. Zeist: Vonk & Co."},{"key":"9236_CR18","unstructured":"Gysseling, M. (1977\u20131981). Corpus van Middelnederlandse teksten (tot en met het jaar 1300). M.m.v. Willy Pijnenburg. Nijhoff,\u2019s-Gravenhage."},{"issue":"2","key":"9236_CR19","doi-asserted-by":"crossref","first-page":"65","DOI":"10.21248\/jlcl.26.2011.147","volume":"26","author":"I Hendrickx","year":"2011","unstructured":"Hendrickx, I., & Marquilhas, R. (2011). From old texts to modern spellings: An experiment in automatic normalisation. Journal for Language Technology and Computational Linguistics (JLCL), 26(2), 65\u201376.","journal-title":"Journal for Language Technology and Computational Linguistics (JLCL)"},{"key":"9236_CR20","doi-asserted-by":"crossref","unstructured":"Hinrichs, E., & Zastrow, T. (2012). Linguistic annotations for a diachronic corpus of German. Linguistic Issues in Language Technology, 7(7).","DOI":"10.33011\/lilt.v7i.1271"},{"issue":"1","key":"9236_CR21","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/1731035.1731036","volume":"9","author":"CC Hsu","year":"2010","unstructured":"Hsu, C. C., & Chen, C. H. (2010). Mining synonymous transliterations from the World Wide Web. ACM Transactions on Asian Language Information Processing (TALIP), 9(1), 1\u201328.","journal-title":"ACM Transactions on Asian Language Information Processing (TALIP)"},{"key":"9236_CR22","unstructured":"Huber, O. (1988). Het kodeerprogramma L.I.M.A. In van Reenen-Stein, K. H., van Reenen, P. Th., & Dees, A. (eds.), Corpus Gebaseerde Woordanalyse (VWF VULET 88\/9), Jaarboek 1987-1988, pp. 61\u201370."},{"key":"9236_CR23","volume-title":"The construction of lexically analysed text corpora on the computer","author":"O Huber","year":"1989","unstructured":"Huber, O. (1989). The construction of lexically analysed text corpora on the computer. Amsterdam: Vrije Universiteit."},{"key":"9236_CR25","unstructured":"Jurish, B. (2010). Comparing canonicalizations of historical German text. In Proceedings of the 11th Meeting of SIGMORPHON, Uppsala, Sweden, July 15, 2010."},{"issue":"3","key":"9236_CR26","doi-asserted-by":"crossref","first-page":"287","DOI":"10.1093\/llc\/fqq011","volume":"25","author":"M Kestemont","year":"2010","unstructured":"Kestemont, M., Daelemans, W., & De Pauw, G. (2010). Weight your words\u2014memory based lemmatization for Middle Dutch. Literary and Linguistic Computing, 25(3), 287\u2013301.","journal-title":"Literary and Linguistic Computing"},{"key":"9236_CR46","unstructured":"\u041be\u0432e\u043d\u0448\u0442e\u0439\u043d, B. \u0418. (1965). \u0414\u0432o\u0438\u0447\u043d\u044be \u043ao\u0434\u044b c \u0438c\u043fpa\u0432\u043be\u043d\u0438e\u043c \u0432\u044b\u043fa\u0434e\u043d\u0438\u0439, \u0432c\u0442a\u0432o\u043a \u0438 \u0437a\u043ce\u0449e\u043d\u0438\u0439 c\u0438\u043c\u0432o\u043bo\u0432. \u0414o\u043a\u043ba\u0434\u044b A\u043aa\u0434e\u043c\u0438\u0439 Hay\u043a CCCP, 163(4), 845\u2013848. Appeared in English as: Levenshtein, V.I. 1966. Binary codes capable of correcting deletions, insertions, and reversals. Soviet Physics Doklady, 10, 707\u2013710."},{"key":"9236_CR27","unstructured":"Mooijaart, M. A. (1992). Atlas van Vroegmiddelnederlandse taalvarianten, Led, Utrecht 1992. PhD Thesis. University of Leiden."},{"issue":"1","key":"9236_CR28","doi-asserted-by":"crossref","first-page":"65","DOI":"10.1093\/llc\/fqm044","volume":"23","author":"T Pilz","year":"2008","unstructured":"Pilz, T., Ernst-Gerlach, A., Kempken, S., Rayson, P., & Archer, D. (2008). The identification of spelling variants in English and German historical texts: manual or automatic? Literary and Linguistic Computing, 23(1), 65\u201372.","journal-title":"Literary and Linguistic Computing"},{"key":"9236_CR29","doi-asserted-by":"crossref","first-page":"179","DOI":"10.1093\/llc\/fql020","volume":"21","author":"T Pilz","year":"2006","unstructured":"Pilz, T., Luther, W., Ammon, U., & Fuhr, N. (2006). Rule-based search in text databases with nonstandard orthography. Literary and Linguistic Computing, 21, 179\u2013186.","journal-title":"Literary and Linguistic Computing"},{"key":"9236_CR30","unstructured":"Rem, M. (2003). De taal van de klerken uit de Hollandse grafelijke kanselarij (1300-1340). Naar een lokaliseringsprocedure voor het veertiende-eeuws Middelnederlands. PhD Thesis, Vrije Universiteit, Amsterdam."},{"key":"9236_CR31","unstructured":"Reynaert, M. (2005). Text-Induced Spelling Correction. PhD Thesis. University of Tilburg."},{"key":"9236_CR32","unstructured":"Reynaert, M., Hendrickx, I., & Marquilhas, R. (2012). Historical spelling normalization. A comparison of two statistical methods: TICCL and VARD2. In Proceedings of the Second Workshop on Annotation of Corpora for Research in the Humanities (ACRH-2). Lisbon, Portugal, November 29, 2012."},{"key":"9236_CR33","unstructured":"Santorini, B. (2010). Annotation manual for the Penn Historical Corpora and the PCEEC. http:\/\/www.ling.upenn.edu\/hist-corpora\/annotation\/index.html ."},{"key":"9236_CR34","unstructured":"Schnurrenberger, M. (2010). Methods for graphemic normalization of unstandardized written language from Middle High German Corpora. Master thesis, Ruhr University Bochum."},{"issue":"3","key":"9236_CR35","doi-asserted-by":"crossref","first-page":"379","DOI":"10.1002\/j.1538-7305.1948.tb01338.x","volume":"27","author":"CE Shannon","year":"1948","unstructured":"Shannon, C. E. (1948). A mathematical theory of communication. Bell System Technical Journal, 27(3), 379\u2013423.","journal-title":"Bell System Technical Journal"},{"key":"9236_CR36","unstructured":"Tagg, C. (2009). A corpus linguistic study of SMS text messaging. PhD thesis. University of Birmingham."},{"key":"9236_CR37","unstructured":"van Camp, V. (2011). De oorkonden en de kanselarij van de graven van Henegouwen, Holland en Zeeland. Schriftelijke communicatie tijdens een personele unie: Henegouwen, 1280-1345, 2 dln. Hilversum: Verloren."},{"key":"9236_CR38","volume-title":"Woordenboek van voornamen","author":"J Schaar van der","year":"1992","unstructured":"van der Schaar, J. (1992). Woordenboek van voornamen. Utrecht: Het Spectrum."},{"key":"9236_CR39","unstructured":"Van Eynde, F. (2005) Part of speech tagging en lemmatisering van het D-Coi corpus. Centrum voor Computerlingu\u00efstiek, Leuven. http:\/\/www.ccl.kuleuven.be\/Papers\/DCOIpos.pdf ."},{"key":"9236_CR40","doi-asserted-by":"crossref","unstructured":"van Halteren, H. (2000a). A default first order family weight determination procedure for WPDV models. In Proceedings of the 2nd Workshop on Learning Language in Logic and the 4th Conference on Computational Natural Language Learning, pp. 119\u2013122.","DOI":"10.3115\/1117601.1117628"},{"key":"9236_CR41","unstructured":"van Halteren, H. (2000b). The detection of inconsistency in manually tagged text. In Proceedings of the LINC2000."},{"issue":"2","key":"9236_CR42","doi-asserted-by":"crossref","first-page":"199","DOI":"10.1162\/089120101750300508","volume":"27","author":"H Halteren van","year":"2001","unstructured":"van Halteren, H., Daelemans, W., & Zavrel, J. (2001). Improving accuracy in word class tagging through the combination of machine learning systems. Computational Linguistics, 27(2), 199\u2013229.","journal-title":"Computational Linguistics"},{"key":"9236_CR43","first-page":"2","volume":"2","author":"H Halteren van","year":"2012","unstructured":"van Halteren, H., & Oostdijk, N. (2012). Towards identifying normal forms for various word form spellings on Twitter. CLIN Journal, 2, 2\u201322.","journal-title":"CLIN Journal"},{"key":"9236_CR44","unstructured":"Van Uytvank, D. (2007). Orthography-based dating and localisation of Middle Dutch charters. MA Thesis, Radboud University Nijmegen."},{"key":"9236_CR45","doi-asserted-by":"crossref","unstructured":"Wieling, M., Proki\u0107, J., & Nerbonne, J. (2009). Evaluating the pairwise string alignment of pronunciations. In Borin, L., & Lendvai, P. (Eds.), Proceedings Language Technology and Resources for Cultural Heritage, Social Sciences, Humanities, and Education (LaTeCH\u2014SHELT&R 2009) Workshop at the 12th Meeting of the European Chapter of the Association for Computational Linguistics. Athens, March 30, 2009, pp. 26\u201334","DOI":"10.3115\/1642049.1642053"}],"container-title":["Language Resources and Evaluation"],"original-title":[],"language":"en","link":[{"URL":"http:\/\/link.springer.com\/content\/pdf\/10.1007\/s10579-013-9236-1.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/link.springer.com\/article\/10.1007\/s10579-013-9236-1\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/link.springer.com\/content\/pdf\/10.1007\/s10579-013-9236-1","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,5,11]],"date-time":"2024-05-11T06:58:15Z","timestamp":1715410695000},"score":1,"resource":{"primary":{"URL":"http:\/\/link.springer.com\/10.1007\/s10579-013-9236-1"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2013,5,29]]},"references-count":45,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2013,12]]}},"alternative-id":["9236"],"URL":"https:\/\/doi.org\/10.1007\/s10579-013-9236-1","relation":{},"ISSN":["1574-020X","1574-0218"],"issn-type":[{"value":"1574-020X","type":"print"},{"value":"1574-0218","type":"electronic"}],"subject":[],"published":{"date-parts":[[2013,5,29]]}}}