{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,14]],"date-time":"2026-05-14T23:01:39Z","timestamp":1778799699245,"version":"3.51.4"},"reference-count":0,"publisher":"Cambridge University Press (CUP)","issue":"2","license":[{"start":{"date-parts":[[2003,8,4]],"date-time":"2003-08-04T00:00:00Z","timestamp":1059955200000},"content-version":"unspecified","delay-in-days":64,"URL":"https:\/\/www.cambridge.org\/core\/terms"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Nat. Lang. Eng."],"published-print":{"date-parts":[[2003,6]]},"abstract":"<jats:p>This paper presents the results of a study on information extraction from unrestricted Turkish \ntext using statistical language processing methods. In languages like English, there is a very \nsmall number of possible word forms with a given root word. However, languages like Turkish \nhave very productive agglutinative morphology. Thus, it is an issue to build statistical models \nfor specific tasks using the surface forms of the words, mainly because of the data sparseness \nproblem. In order to alleviate this problem, we used additional syntactic information, i.e. the \nmorphological structure of the words. We have successfully applied statistical methods using \nboth the lexical and morphological information to sentence segmentation, topic segmentation, \nand name tagging tasks. For sentence segmentation, we have modeled the final inflectional \ngroups of the words and combined it with the lexical model, and decreased the error rate \nto 4.34%, which is 21% better than the result obtained using only the surface forms of the \nwords. For topic segmentation, stems of the words (especially nouns) have been found to \nbe more effective than using the surface forms of the words and we have achieved 10.90% \nsegmentation error rate on our test set according to the weighted TDT-2 segmentation cost \nmetric. This is 32% better than the word-based baseline model. For name tagging, we used \nfour different information sources to model names. Our first information source is based on the \nsurface forms of the words. Then we combined the contextual cues with the lexical model, and \nobtained some improvement. After this, we modeled the morphological analyses of the words, and finally we modeled the tag sequence, and reached an F-Measure of 91.56%, according\nto the MUC evaluation criteria. Our results are important in the sense that, using linguistic \ninformation, i.e. morphological analyses of the words, and a corpus large enough to train a statistical model significantly improves these basic information extraction tasks for Turkish.<\/jats:p>","DOI":"10.1017\/s135132490200284x","type":"journal-article","created":{"date-parts":[[2003,10,16]],"date-time":"2003-10-16T10:50:27Z","timestamp":1066301427000},"page":"181-210","source":"Crossref","is-referenced-by-count":60,"title":["A statistical information extraction system for Turkish"],"prefix":"10.1017","volume":"9","author":[{"given":"G\u00d6KHAN","family":"T\u00dcR","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"DILEK","family":"HAKKANI-T\u00dcR","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"KEMAL","family":"OFLAZER","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"56","published-online":{"date-parts":[[2003,8,4]]},"container-title":["Natural Language Engineering"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.cambridge.org\/core\/services\/aop-cambridge-core\/content\/view\/S135132490200284X","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2019,3,30]],"date-time":"2019-03-30T18:55:00Z","timestamp":1553972100000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.cambridge.org\/core\/product\/identifier\/S135132490200284X\/type\/journal_article"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2003,6]]},"references-count":0,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2003,6]]}},"alternative-id":["S135132490200284X"],"URL":"https:\/\/doi.org\/10.1017\/s135132490200284x","relation":{},"ISSN":["1351-3249","1469-8110"],"issn-type":[{"value":"1351-3249","type":"print"},{"value":"1469-8110","type":"electronic"}],"subject":[],"published":{"date-parts":[[2003,6]]}}}