{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T04:39:06Z","timestamp":1750307946683,"version":"3.41.0"},"reference-count":31,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2007,10,1]],"date-time":"2007-10-01T00:00:00Z","timestamp":1191196800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Speech Lang. Process."],"published-print":{"date-parts":[[2007,10]]},"abstract":"<jats:p>For the success of lexical text correction, high coverage of the underlying background dictionary is crucial. Still, most correction tools are built on top of static dictionaries that represent fixed collections of expressions of a given language. When treating texts from specific domains and areas, often a significant part of the vocabulary is missed. In this situation, both automated and interactive correction systems produce suboptimal results. In this article, we describe strategies for crawling Web pages that fit the thematic domain of the given input text. Special filtering techniques are introduced to avoid pages with many orthographic errors. Collecting the vocabulary of filtered pages that meet the vocabulary of the input text, dynamic dictionaries of modest size are obtained that reach excellent coverage values. A tool has been developed that automatically crawls dictionaries in the indicated way. Our correction experiments with crawled dictionaries, which address English and German document collections from a variety of thematic fields, show that with these dictionaries even the error rate of highly accurate texts can be reduced, using completely automated correction methods. For interactive text correction, more sensible candidate sets for correcting erroneous words are obtained and the manual effort is reduced in a significant way. To complete this picture, we study the effect when using word trigram models for correction. Again, trigram models from crawled corpora outperform those obtained from static corpora.<\/jats:p>","DOI":"10.1145\/1289600.1289602","type":"journal-article","created":{"date-parts":[[2012,10,26]],"date-time":"2012-10-26T18:34:02Z","timestamp":1351276442000},"page":"9","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":9,"title":["Adaptive text correction with Web-crawled domain-dependent dictionaries"],"prefix":"10.1145","volume":"4","author":[{"given":"Christoph","family":"Ringlstetter","sequence":"first","affiliation":[{"name":"University of Alberta, Edmonton, Canada"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Klaus U.","family":"Schulz","sequence":"additional","affiliation":[{"name":"University of Munich, Munich, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Stoyan","family":"Mihov","sequence":"additional","affiliation":[{"name":"Bulgarian Academy of Sciences, Sofia, Bulgaria"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2007,10]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"crossref","first-page":"255","DOI":"10.1016\/0306-4573(83)90022-5","article-title":"Automatic spelling correction using a trigram similarity measure","volume":"19","author":"Angell R. C.","year":"1983","unstructured":"Angell , R. C. , Freund , G. E. , and Willett , P. 1983 . Automatic spelling correction using a trigram similarity measure . Inform. Proces. Manage. 19 , 255 -- 261 . Angell, R. C., Freund, G. E., and Willett, P. 1983. Automatic spelling correction using a trigram similarity measure. Inform. Proces. Manage. 19, 255--261.","journal-title":"Inform. Proces. Manage."},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1145\/3532.315102"},{"volume-title":"Users reference guide for the British National Corpus. Tech. rep","author":"Burnard L.","key":"e_1_2_1_3_1","unstructured":"Burnard , L. 1995. Users reference guide for the British National Corpus. Tech. rep ., British National Corpus Consortium . Burnard, L. 1995. Users reference guide for the British National Corpus. Tech. rep., British National Corpus Consortium."},{"volume-title":"Proceedings of the 6th European Conference on Speech Communication and Technology (EUROSPEECH'99)","author":"Chelba C.","key":"e_1_2_1_4_1","unstructured":"Chelba , C. and Jelinek , F . 2002. Recognition performance of a structured language model . In Proceedings of the 6th European Conference on Speech Communication and Technology (EUROSPEECH'99) . Budapest, Hungary, 1567--1570. Chelba, C. and Jelinek, F. 2002. Recognition performance of a structured language model. In Proceedings of the 6th European Conference on Speech Communication and Technology (EUROSPEECH'99). Budapest, Hungary, 1567--1570."},{"volume-title":"Proceedings of the Eurospeech '97","author":"Clarkson P.","key":"e_1_2_1_5_1","unstructured":"Clarkson , P. and Rosenfeld , R . 1997. Statistical language modeling using the CMU--cambridge toolkit . In Proceedings of the Eurospeech '97 . Rhodes, Greece, 2707--2710. Clarkson, P. and Rosenfeld, R. 1997. Statistical language modeling using the CMU--cambridge toolkit. In Proceedings of the Eurospeech '97. Rhodes, Greece, 2707--2710."},{"key":"e_1_2_1_6_1","volume-title":"Proceedings of EMNLP","author":"Cucerzan S.","year":"2004","unstructured":"Cucerzan , S. and Brill , E . 2004. Spelling correction as an iterative process that exploits the collective knowledge of Web users . In Proceedings of EMNLP 2004 . Cucerzan, S. and Brill, E. 2004. Spelling correction as an iterative process that exploits the collective knowledge of Web users. In Proceedings of EMNLP 2004."},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/363958.363994"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1016\/0306-4573(89)90099-X"},{"key":"e_1_2_1_9_1","doi-asserted-by":"crossref","unstructured":"Dengel A. Hoch R. H\u00f6nes F. J\u00e4ger T. Malburg M. and Weigel A. 1997. Techniques for improving OCR results. In Handbook of Character Recognition and Document Image Analysis H. Bunke and P. S. Wang Eds. World Scientific 227--258. Dengel A. Hoch R. H\u00f6nes F. J\u00e4ger T. Malburg M. and Weigel A. 1997. Techniques for improving OCR results. In Handbook of Character Recognition and Document Image Analysis H. Bunke and P. S. Wang Eds. World Scientific 227--258.","DOI":"10.1142\/9789812830968_0008"},{"volume-title":"Proceedings of the Workshop on Computational Terminology for Medical and Biological Applications, 2nd International Conference on Natural Language Processing (NLP-2000)","author":"Gaizauskas R.","key":"e_1_2_1_10_1","unstructured":"Gaizauskas , R. , Demetriou , G. , and Humphreys , K . 2000. Term recognition in biological science journal articles . In Proceedings of the Workshop on Computational Terminology for Medical and Biological Applications, 2nd International Conference on Natural Language Processing (NLP-2000) , Patras, Greece. 37--44. Gaizauskas, R., Demetriou, G., and Humphreys, K. 2000. Term recognition in biological science journal articles. In Proceedings of the Workshop on Computational Terminology for Medical and Biological Applications, 2nd International Conference on Natural Language Processing (NLP-2000), Patras, Greece. 37--44."},{"volume-title":"Computational Linguistics in the Netherlands 2000: Selected Papers from the 11th CLIN Meeting. Language and Computers. Rodopi.","author":"Grefenstette G.","key":"e_1_2_1_11_1","unstructured":"Grefenstette , G. 2001. Very large lexicons . In Computational Linguistics in the Netherlands 2000: Selected Papers from the 11th CLIN Meeting. Language and Computers. Rodopi. Grefenstette, G. 2001. Very large lexicons. In Computational Linguistics in the Netherlands 2000: Selected Papers from the 11th CLIN Meeting. Language and Computers. Rodopi."},{"key":"e_1_2_1_12_1","unstructured":"Ho T. K. Hull J. J. and Srihari S. N. 1992. A word shape analysis approach to lexicon-based word recognition. Patte. Recogn. Lett. 10.1016\/0167-8655(92)90133-K Ho T. K. Hull J. J. and Srihari S. N. 1992. A word shape analysis approach to lexicon-based word recognition. Patte. Recogn. Lett. 10.1016\/0167-8655(92)90133-K"},{"key":"e_1_2_1_13_1","doi-asserted-by":"crossref","first-page":"273","DOI":"10.1142\/S0218001496000207","article-title":"On virtual partitioning of large dictionaries for contextual post-processing to improve character recognition","volume":"10","author":"Hoch R.","year":"1996","unstructured":"Hoch , R. and Kieninger , T. 1996 . On virtual partitioning of large dictionaries for contextual post-processing to improve character recognition . Int. J. Patt. Recog. AI. 10 , 4, 273 -- 289 . Hoch, R. and Kieninger, T. 1996. On virtual partitioning of large dictionaries for contextual post-processing to improve character recognition. Int. J. Patt. Recog. AI. 10, 4, 273--289.","journal-title":"Int. J. Patt. Recog. AI."},{"volume-title":"Statistical Methods for Speech Recognition","author":"Jelinek F.","key":"e_1_2_1_14_1","unstructured":"Jelinek , F. 1997. Statistical Methods for Speech Recognition . MIT Press , Cambridge, MA . Jelinek, F. 1997. Statistical Methods for Speech Recognition. MIT Press, Cambridge, MA."},{"key":"e_1_2_1_15_1","unstructured":"Ku&ccirc;era H. and Francis W. N. 1967. Computational Analysis of Present-Day American English. Brown University Press Providence RI. Ku&ccirc;era H. and Francis W. N. 1967. Computational Analysis of Present-Day American English. Brown University Press Providence RI."},{"key":"e_1_2_1_16_1","doi-asserted-by":"crossref","unstructured":"Kukich K. 1992. Techniques for automatically correcting words in texts. ACM Comput. Surv. 377--439. 10.1145\/146370.146380 Kukich K. 1992. Techniques for automatically correcting words in texts. ACM Comput. Surv. 377--439. 10.1145\/146370.146380","DOI":"10.1145\/146370.146380"},{"key":"e_1_2_1_17_1","unstructured":"Levenshtein V. I. 1966. Binary codes capable of correcting deletions insertions and reversals. Sov. Phys. Dokl. Levenshtein V. I. 1966. Binary codes capable of correcting deletions insertions and reversals. Sov. Phys. Dokl."},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1162\/0891201042544938"},{"volume-title":"ANLP\/NAACL 2000 Workshop on Conversational Systems. 27--32","author":"Oh A. H.","key":"e_1_2_1_19_1","unstructured":"Oh , A. H. and Rudickny , A. I . 2000. Stochastic language generation for spoken dialogue systems . In ANLP\/NAACL 2000 Workshop on Conversational Systems. 27--32 . 10.3115\/1117562.1117568 Oh, A. H. and Rudickny, A. I. 2000. Stochastic language generation for spoken dialogue systems. In ANLP\/NAACL 2000 Workshop on Conversational Systems. 27--32. 10.3115\/1117562.1117568"},{"key":"e_1_2_1_20_1","doi-asserted-by":"crossref","first-page":"360","DOI":"10.1109\/89.536930","article-title":"From HMMs to segment models: A unified view of stochastic modeling for speech recognition","volume":"4","author":"Ostendorf M.","year":"1996","unstructured":"Ostendorf , M. , Digalakis , V. V. , and Kimball , O. A. 1996 . From HMMs to segment models: A unified view of stochastic modeling for speech recognition . IEEE Trans. Speech Audio Proces. 4 , 5, 360 -- 378 . Ostendorf, M., Digalakis, V. V., and Kimball, O. A. 1996. From HMMs to segment models: A unified view of stochastic modeling for speech recognition. IEEE Trans. Speech Audio Proces. 4, 5, 360--378.","journal-title":"IEEE Trans. Speech Audio Proces."},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/359038.359041"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/6138.6146"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1162\/coli.2006.32.3.295"},{"key":"e_1_2_1_24_1","volume-title":"Proceedings of the International Conference on New Methods in Language Processing. Manchester, England, 44--49","author":"Schmid H.","year":"1994","unstructured":"Schmid , H. 1994 . Probabilistic part-of-speech tagging using decision trees . In Proceedings of the International Conference on New Methods in Language Processing. Manchester, England, 44--49 . Schmid, H. 1994. Probabilistic part-of-speech tagging using decision trees. In Proceedings of the International Conference on New Methods in Language Processing. Manchester, England, 44--49."},{"key":"e_1_2_1_25_1","doi-asserted-by":"crossref","unstructured":"Strohmaier C. Ringlstetter C. Schulz K. U. and Mihov S. 2003a. Lexical postcorrection of OCR-results: The Web as a dynamic secondary dictionary&quest; In Proceedings of the 7th International Conference on Document Analysis and Recognition (ICDAR'03). 1133--1137. Strohmaier C. Ringlstetter C. Schulz K. U. and Mihov S. 2003a. Lexical postcorrection of OCR-results: The Web as a dynamic secondary dictionary&quest; In Proceedings of the 7th International Conference on Document Analysis and Recognition (ICDAR'03). 1133--1137.","DOI":"10.1109\/ICDAR.2003.1227833"},{"volume-title":"Proceedings of the IEEE Workshop on Document Image Analysis and Recognition (DIAR'03)","author":"Strohmaier C.","key":"e_1_2_1_26_1","unstructured":"Strohmaier , C. , Ringlstetter , C. , Schulz , K. U. , and Mihov , S . 2003b. A visual and interactive tool for optimizing lexical postcorrection of OCR results . In Proceedings of the IEEE Workshop on Document Image Analysis and Recognition (DIAR'03) . Strohmaier, C., Ringlstetter, C., Schulz, K. U., and Mihov, S. 2003b. A visual and interactive tool for optimizing lexical postcorrection of OCR results. In Proceedings of the IEEE Workshop on Document Image Analysis and Recognition (DIAR'03)."},{"volume-title":"Proceedings of the 1st ACM Workshop on Hardcopy Document Processing (HDP'04)","author":"Taghva K.","key":"e_1_2_1_27_1","unstructured":"Taghva , K. , Nartker , T. , and Borsack , J . 2004. Information access in the presence of OCR errors . In Proceedings of the 1st ACM Workshop on Hardcopy Document Processing (HDP'04) . ACM Press, New York, NY, 1--8. 10.1145\/1031442.1031443 Taghva, K., Nartker, T., and Borsack, J. 2004. Information access in the presence of OCR errors. In Proceedings of the 1st ACM Workshop on Hardcopy Document Processing (HDP'04). ACM Press, New York, NY, 1--8. 10.1145\/1031442.1031443"},{"key":"e_1_2_1_28_1","doi-asserted-by":"crossref","first-page":"125","DOI":"10.1007\/PL00013558","article-title":"OCRSpell: an interactive spelling correction system for OCR errors in text","volume":"3","author":"Taghva K.","year":"2001","unstructured":"Taghva , K. and Stofsky , E. 2001 . OCRSpell: an interactive spelling correction system for OCR errors in text . Int. J. Docum. Anal. Recogn. 3 , 125 -- 137 . Taghva, K. and Stofsky, E. 2001. OCRSpell: an interactive spelling correction system for OCR errors in text. Int. J. Docum. Anal. Recogn. 3, 125--137.","journal-title":"Int. J. Docum. Anal. Recogn."},{"volume-title":"Proceedings of the 3rd International Conference on Document Analysis and Recognition (ICDAR'95)","author":"Weigel A.","key":"e_1_2_1_29_1","unstructured":"Weigel , A. , Baumann , S. , and Rohrschneider , J . 1995. Lexical postprocessing by heuristic search and automatic determination of the edit costs . In Proceedings of the 3rd International Conference on Document Analysis and Recognition (ICDAR'95) . 857--860. Weigel, A., Baumann, S., and Rohrschneider, J. 1995. Lexical postprocessing by heuristic search and automatic determination of the edit costs. In Proceedings of the 3rd International Conference on Document Analysis and Recognition (ICDAR'95). 857--860."},{"key":"e_1_2_1_30_1","doi-asserted-by":"crossref","first-page":"99","DOI":"10.1007\/s00799-003-0050-z","article-title":"Searchable words on the","volume":"5","author":"Williams H.","year":"2005","unstructured":"Williams , H. and Zobel , J. 2005 . Searchable words on the Web. Int. J. Digit. Libra. 5 , 2, 99 -- 105 . Williams, H. and Zobel, J. 2005. Searchable words on the Web. Int. J. Digit. Libra. 5, 2, 99--105.","journal-title":"Web. Int. J. Digit. Libra."},{"key":"e_1_2_1_31_1","doi-asserted-by":"crossref","first-page":"1085","DOI":"10.1109\/18.87000","article-title":"The zero-frequency problem: Estimating the probabilities of novel events in adaptive text compression","volume":"37","author":"Witten I. H.","year":"1991","unstructured":"Witten , I. H. and Bell , T. C. 1991 . The zero-frequency problem: Estimating the probabilities of novel events in adaptive text compression . IEEE Trans. Inform. Theory 37 , 4, 1085 -- 1094 . Witten, I. H. and Bell, T. C. 1991. The zero-frequency problem: Estimating the probabilities of novel events in adaptive text compression. IEEE Trans. Inform. Theory 37, 4, 1085--1094.","journal-title":"IEEE Trans. Inform. Theory"}],"container-title":["ACM Transactions on Speech and Language Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/1289600.1289602","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/1289600.1289602","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T14:52:31Z","timestamp":1750258351000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/1289600.1289602"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2007,10]]},"references-count":31,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2007,10]]}},"alternative-id":["10.1145\/1289600.1289602"],"URL":"https:\/\/doi.org\/10.1145\/1289600.1289602","relation":{},"ISSN":["1550-4875","1550-4883"],"issn-type":[{"type":"print","value":"1550-4875"},{"type":"electronic","value":"1550-4883"}],"subject":[],"published":{"date-parts":[[2007,10]]},"assertion":[{"value":"2007-10-01","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}