{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2024,3,14]],"date-time":"2024-03-14T02:49:57Z","timestamp":1710384597501},"reference-count":35,"publisher":"Cambridge University Press (CUP)","issue":"2","license":[{"start":{"date-parts":[[2019,4,1]],"date-time":"2019-04-01T00:00:00Z","timestamp":1554076800000},"content-version":"unspecified","delay-in-days":31,"URL":"https:\/\/www.cambridge.org\/core\/terms"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Nat. Lang. Eng."],"published-print":{"date-parts":[[2019,3]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>This paper presents a study about methods for normalization of historical texts. The aim of these methods is learning relations between historical and contemporary word forms. We have compiled training and test corpora for different languages and scenarios, and we have tried to read the results related to the features of the corpora and languages. Our proposed method, based on weighted finite-state transducers, is compared to previously published ones. Our method learns to map phonological changes using a noisy channel model; it is a simple solution that can use a limited amount of supervision in order to achieve adequate performance. The compiled corpora are ready to be used for other researchers in order to compare results. Concerning the amount of supervision for the task, we investigate how the size of training corpus affects the results and identify some interesting factors to anticipate the difficulty of the task.<\/jats:p>","DOI":"10.1017\/s1351324918000505","type":"journal-article","created":{"date-parts":[[2019,4,1]],"date-time":"2019-04-01T10:22:07Z","timestamp":1554114127000},"page":"307-321","source":"Crossref","is-referenced-by-count":3,"title":["Weighted finite-state transducers for normalization of historical texts"],"prefix":"10.1017","volume":"25","author":[{"given":"Izaskun","family":"Etxeberria","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"I\u00f1aki","family":"Alegria","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Larraitz","family":"Uria","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"56","published-online":{"date-parts":[[2019,4,1]]},"reference":[{"key":"S1351324918000505_ref17","first-page":"72","volume-title":"Proceedings of the 11th Meeting of the ACL Special Interest Group on Computational Morphology and Phonology","author":"Jurish","year":"2010"},{"key":"S1351324918000505_ref30","first-page":"124","volume-title":"Proceedings of the 5th Linguistic Annotation Workshop","author":"Scheible","year":"2011"},{"key":"S1351324918000505_ref3","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-540-76336-9_3"},{"key":"S1351324918000505_ref32","doi-asserted-by":"publisher","DOI":"10.1017\/S1351324915000236"},{"key":"S1351324918000505_ref16","first-page":"372","volume-title":"Proceedings of HLT-NAACL\u201907","author":"Jiampojamarn","year":"2007"},{"key":"S1351324918000505_ref12","first-page":"1064","volume-title":"Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC2016)","author":"Etxeberria","year":"2016"},{"key":"S1351324918000505_ref29","first-page":"1977","volume-title":"Proceedings of the Eighth International Conference on Language Resources and Evaluation (LREC2012)","author":"Rognvaldsson","year":"2012"},{"key":"S1351324918000505_ref13","doi-asserted-by":"publisher","DOI":"10.1080\/03585522.2011.617576"},{"key":"S1351324918000505_ref22","doi-asserted-by":"publisher","DOI":"10.1016\/0743-1066(94)90035-3"},{"key":"S1351324918000505_ref27","doi-asserted-by":"publisher","DOI":"10.2200\/S00436ED1V01Y201207HLT017"},{"key":"S1351324918000505_ref33","first-page":"224","volume-title":"The Evolution of Functional Left Peripheries in Hungarian Syntax","author":"Simon","year":"2014"},{"key":"S1351324918000505_ref31","doi-asserted-by":"publisher","DOI":"10.3115\/1557835.1557847"},{"key":"S1351324918000505_ref7","first-page":"131","volume-title":"Proceedings of the 26th International Conference on Computational Linguistics (COLING 2016)","author":"Bollmann","year":"2016"},{"key":"S1351324918000505_ref18","doi-asserted-by":"publisher","DOI":"10.1093\/llc\/fqq011"},{"key":"S1351324918000505_ref1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-14684-8_11"},{"key":"S1351324918000505_ref20","first-page":"146","volume-title":"Proceedings of the 13th Conference on Natural Language Processing (KONVENS 2016)","author":"Ljube\u0161ic","year":"2016"},{"key":"S1351324918000505_ref2","first-page":"15","volume-title":"Proceedings of the Tweet Normalization Workshop at the conference of the Spanish Society for Natural Language Processing (SEPLN)","author":"Alegria","year":"2013"},{"key":"S1351324918000505_ref4","first-page":"227","volume-title":"Proceedings of the Seventh International Conference on Language Resources and Evaluation (LREC2010)","author":"Almeida","year":"2010"},{"key":"S1351324918000505_ref5","volume-title":"Finite-state Morphology: Xerox Tools and Techniques","author":"Beesley","year":"2003"},{"key":"S1351324918000505_ref28","first-page":"70","volume-title":"Proceedings of the Workshop on Computational Historical Linguistics at NODALIDA 2013","volume":"18","author":"Porta","year":"2013"},{"key":"S1351324918000505_ref6","first-page":"342","volume-title":"Proceedings of the First International Workshop on Language Technology for Historical Text(s)","author":"Bollmann","year":"2012"},{"key":"S1351324918000505_ref8","first-page":"239","volume-title":"Proceedings of the Fourth International Conference on Language Resources and Evaluation (LREC2004)","author":"Carreras","year":"2004"},{"key":"S1351324918000505_ref9","doi-asserted-by":"publisher","DOI":"10.1007\/s10579-015-9294-7"},{"key":"S1351324918000505_ref10","first-page":"13","article-title":"Learning to map variation-standard forms using a limited parallel corpus and the standard morphology","volume":"52","author":"Etxeberria","year":"2014","journal-title":"Procesamiento del Lenguaje Natural"},{"key":"S1351324918000505_ref14","first-page":"29","volume-title":"Proceedings of the 12th Conference of the European Chapter of the Association for Computational Linguistics: Demonstrations Session","author":"Hulden","year":"2009"},{"key":"S1351324918000505_ref25","volume-title":"Spelling Normalisation and Linguistic Analysis of Historical Text for Information Extraction","author":"Pettersson","year":"2016"},{"key":"S1351324918000505_ref15","first-page":"39","volume-title":"Proceedings of the First Workshop on Algorithms and Resources for Modelling of Dialects and Language Varieties","author":"Hulden","year":"2011"},{"key":"S1351324918000505_ref19","first-page":"12","volume-title":"Proceedings of the NoDaLiDa 2017 Workshop on Processing Historical Language","author":"Korchagina","year":"2017"},{"key":"S1351324918000505_ref21","first-page":"151","volume-title":"Proceedings of the Second Meeting of the North American Chapter of the Association for Computational Linguistics on Language Technologies","author":"Mann","year":"2001"},{"key":"S1351324918000505_ref23","first-page":"45","volume-title":"Proceedings of the 10th International Workshop on Finite State Methods and Natural Language Processing (FSMNLP2012)","author":"Novak","year":"2012"},{"key":"S1351324918000505_ref24","doi-asserted-by":"publisher","DOI":"10.1017\/S1351324915000315"},{"key":"S1351324918000505_ref26","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/W14-0605"},{"key":"S1351324918000505_ref34","first-page":"102","volume-title":"Proceedings of the Demonstrations at the 13th Conference of the European Chapter of the Association for Computational Linguistics","author":"Stenetorp","year":"2012"},{"key":"S1351324918000505_ref35","first-page":"53","article-title":"The CLIN27 shared task: Translating historical text to contemporary language for improving automatic linguistic annotation","volume":"7","author":"Tjong","year":"2017","journal-title":"Computational Linguistics in the Netherlands Journal"},{"key":"S1351324918000505_ref11","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/W16-2112"}],"container-title":["Natural Language Engineering"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.cambridge.org\/core\/services\/aop-cambridge-core\/content\/view\/S1351324918000505","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2019,4,12]],"date-time":"2019-04-12T00:36:36Z","timestamp":1555029396000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.cambridge.org\/core\/product\/identifier\/S1351324918000505\/type\/journal_article"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019,3]]},"references-count":35,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2019,3]]}},"alternative-id":["S1351324918000505"],"URL":"https:\/\/doi.org\/10.1017\/s1351324918000505","relation":{},"ISSN":["1351-3249","1469-8110"],"issn-type":[{"value":"1351-3249","type":"print"},{"value":"1469-8110","type":"electronic"}],"subject":[],"published":{"date-parts":[[2019,3]]}}}