{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,27]],"date-time":"2026-01-27T22:32:45Z","timestamp":1769553165964,"version":"3.49.0"},"reference-count":20,"publisher":"Oxford University Press (OUP)","issue":"3","license":[{"start":{"date-parts":[[2023,4,12]],"date-time":"2023-04-12T00:00:00Z","timestamp":1681257600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/academic.oup.com\/pages\/standard-publication-reuse-rights"}],"funder":[{"DOI":"10.13039\/501100012456","name":"National Social Science Foundation of China","doi-asserted-by":"publisher","id":[{"id":"10.13039\/501100012456","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2023,8,31]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>There exists no sentence boundary in most classical Chinese literature texts. Since it is difficult to read literature of this kind, experts in literature or linguistics would segment the sentence manually. This article explores the effectiveness of classical Chinese sentence segmentation method so as to provide a reference for classical Chinese punctuation. On the basis of the machine learning methods, we chose three components of machine learning, namely models, tagging schemes, and features, to compare the learning results. The models include conditional random field (CRF) models, long short term memory (LSTM) models, BiLSTM\u2013CRF models, and three Bidirectional Encoder Representation from Transformers (BERT) models. There are five tagging schemes in this article and three features including the statistical feature, Guangyun, and Fanqie. Finally, the performance of the combined feature template is evaluated by ten-fold cross-validation on four classical Chinese texts in different genres. The SikuBERT model is proved to be the most effective model for sentence segmentation at present. Different tagging schemes and various features are introduced. The results show that 5-tag-J tagging schemes can improve performance. Statistical feature, as an important clue for classical Chinese sentence segmentation, is useful in related tasks, but Guangyun and Fanqie have little impact. Other important factors of sentence segmentation are genres and writing styles.<\/jats:p>","DOI":"10.1093\/llc\/fqad016","type":"journal-article","created":{"date-parts":[[2023,4,12]],"date-time":"2023-04-12T20:47:15Z","timestamp":1681332435000},"page":"1067-1077","source":"Crossref","is-referenced-by-count":1,"title":["Automatic sentence segmentation for classical Chinese: <i>The Spring and Autumn Annals<\/i> as an example"],"prefix":"10.1093","volume":"38","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-7667-2629","authenticated-orcid":false,"given":"Wenjie","family":"Fan","sequence":"first","affiliation":[{"name":"College of Information Management, Nanjing Agricultural University , Nanjing, China"},{"name":"Research Center for Humanities and Social Computing, Nanjing Agricultural University , Nanjing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9894-9550","authenticated-orcid":false,"given":"Dongbo","family":"Wang","sequence":"additional","affiliation":[{"name":"College of Information Management, Nanjing Agricultural University , Nanjing, China"},{"name":"Research Center for Humanities and Social Computing, Nanjing Agricultural University , Nanjing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1646-9300","authenticated-orcid":false,"given":"Shuiqing","family":"Huang","sequence":"additional","affiliation":[{"name":"College of Information Management, Nanjing Agricultural University , Nanjing, China"},{"name":"Research Center for Humanities and Social Computing, Nanjing Agricultural University , Nanjing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"286","published-online":{"date-parts":[[2023,4,12]]},"reference":[{"issue":"3","key":"2023083111394515800_fqad016-B1","first-page":"192","article-title":"Archaic Chinese punctuating sentences based on context N-gram model","volume":"33","author":"Chen","year":"2007","journal-title":"Computer Engineering"},{"key":"2023083111394515800_fqad016-B2","doi-asserted-by":"crossref","first-page":"3504","DOI":"10.1109\/TASLP.2021.3124365","article-title":"Pre-training with whole word masking for Chinese BERT","volume":"29","author":"Cui","year":"2021","journal-title":"IEEE\/ACM Transactions on Audio, Speech, and Language Processing"},{"key":"2023083111394515800_fqad016-B4","author":"Flanders","year":"2012"},{"key":"2023083111394515800_fqad016-B3","first-page":"1512","author":"Hammerton","year":"2003"},{"key":"2023083111394515800_fqad016-B5","doi-asserted-by":"crossref","first-page":"225","DOI":"10.1016\/j.aiopen.2021.08.002","article-title":"Pre-trained models: Past, present and future","volume":"2","author":"Han","year":"2021","journal-title":"AI Open"},{"key":"2023083111394515800_fqad016-B6","first-page":"473","author":"Hochreiter","year":"1996"},{"key":"2023083111394515800_fqad016-B7","author":"Huang","year":"2010"},{"issue":"4","key":"2023083111394515800_fqad016-B8","first-page":"31","article-title":"On sentence segmentation and punctuation model for ancient books on agriculture","volume":"22","author":"Huang","year":"2008","journal-title":"Journal of Chinese Information Processing"},{"issue":"11","key":"2023083111394515800_fqad016-B9","first-page":"27","article-title":"Exploring of word segmentation for fore-Qin literature based on the domain glossary of Sinological Index Series","volume":"59","author":"Huang","year":"2015","journal-title":"Library and Information Service"},{"key":"2023083111394515800_fqad016-B10","author":"Jacob","year":"2018"},{"issue":"2","key":"2023083111394515800_fqad016-B11","first-page":"67","article-title":"Cultural interpretation of non-punctuation","volume":"31","author":"Kan","year":"2005","journal-title":"Journal of Jiangsu Normal University (Philosophy and Social Sciences Edition)"},{"key":"2023083111394515800_fqad016-B12","author":"Laffrtty","year":"2001"},{"issue":"6","key":"2023083111394515800_fqad016-B13","first-page":"6","article-title":"The construction of a segmented and part-of-speech tagged archaic Chinese corpus: A case study on Huainanzi","volume":"27","author":"Lau","year":"2013","journal-title":"Journal of Chinese Information Processing"},{"issue":"231","key":"2023083111394515800_fqad016-B14","first-page":"32","article-title":"Exploring technical system and theoretical structure of digital humanities","volume":"43","author":"Liu","year":"2017","journal-title":"Journal of Library Science in China"},{"issue":"2","key":"2023083111394515800_fqad016-B15","first-page":"66","article-title":"Visual analysis and exploration of ancient texts for digital humanities research","volume":"42","author":"Ouyang","year":"2016","journal-title":"Journal of Library Science in China"},{"key":"2023083111394515800_fqad016-B16","first-page":"39","article-title":"The integrated research about ancient Chinese word segmentation and POS tagging based on CRF","volume":"2","author":"Shi","year":"2010","journal-title":"Journal of Chinese Information Processing"},{"issue":"2","key":"2023083111394515800_fqad016-B17","first-page":"255","article-title":"A sentence segmentation method for ancient Chinese texts based on recurrent neural network","volume":"53","author":"Wang","year":"2017","journal-title":"Workshop on Chinese Lexical Semantics"},{"key":"2023083111394515800_fqad016-B18","volume-title":"Word Segmentation Model of the Ancient Chinese based on Root Word Algorithm","author":"Yang","year":"2007"},{"issue":"09","key":"2023083111394515800_fqad016-B19","first-page":"3326","article-title":"CRF-based approach to sentence segmentation and punctuation for ancient Chinese prose","volume":"26","author":"Zhang","year":"2009","journal-title":"Journal of Tsinghua University (Science and Technology)"},{"issue":"9","key":"2023083111394515800_fqad016-B20","first-page":"3326","article-title":"Method of sentence segmentation and punctuating for ancient Chinese literature based on cascaded CRF","volume":"26","author":"Zhang","year":"2009","journal-title":"Application Research of Computers"}],"container-title":["Digital Scholarship in the Humanities"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/dsh\/article-pdf\/38\/3\/1067\/51309589\/fqad016.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/dsh\/article-pdf\/38\/3\/1067\/51309589\/fqad016.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,8,31]],"date-time":"2023-08-31T11:44:20Z","timestamp":1693482260000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/dsh\/article\/38\/3\/1067\/7116309"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,4,12]]},"references-count":20,"journal-issue":{"issue":"3","published-online":{"date-parts":[[2023,4,12]]},"published-print":{"date-parts":[[2023,8,31]]}},"URL":"https:\/\/doi.org\/10.1093\/llc\/fqad016","relation":{},"ISSN":["2055-7671","2055-768X"],"issn-type":[{"value":"2055-7671","type":"print"},{"value":"2055-768X","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2023,9,1]]},"published":{"date-parts":[[2023,4,12]]}}}