{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,28]],"date-time":"2026-08-28T10:48:04Z","timestamp":1787914084572,"version":"build-2784847793"},"reference-count":36,"publisher":"Oxford University Press (OUP)","issue":"2","license":[{"start":{"date-parts":[[2025,3,4]],"date-time":"2025-03-04T00:00:00Z","timestamp":1741046400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/academic.oup.com\/pages\/standard-publication-reuse-rights"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"name":"The Construction of the Knowledge Graph for the History of Chinese Confucianism","award":["72010107003"],"award-info":[{"award-number":["72010107003"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2025,6,1]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>This work contributes to the digital humanities approach for studying premodern Chinese history and culture by creating a large-scale dataset annotated with named entities and relations. Through careful annotation guidelines and labeling of over 200,000 characters, we developed a dataset containing 30,000 named entities across six types and 7,000 relations spanning twenty categories. Experiments on named entity recognition (NER) using pre-trained language models and large language models on this dataset achieved an initial performance of NER (91.32 percent F1). In addition, relationship extraction (RE) on the pretrained language model achieves an 85.32 percent F1 score. While there is still room for improvement, our annotated dataset and models provide a useful starting point for extracting semantic information from premodern Chinese texts. It represents an effort to connect history and technology, increasing accessibility and preservation of premodern Chinese cultural treasures. Furthermore, our dataset can facilitate downstream tasks like culture analysis, knowledge graph construction, and computational understanding of premodern Chinese. Overall, this research represents a significant step toward digitally exploring premodern Chinese documents, providing a pathway for future work on knowledge organization and computational analysis of this valuable cultural legacy. Our code and data are available at: https:\/\/github.com\/tangxuemei1995\/AnChineseNERE<\/jats:p>","DOI":"10.1093\/llc\/fqaf001","type":"journal-article","created":{"date-parts":[[2025,3,5]],"date-time":"2025-03-05T02:02:34Z","timestamp":1741140154000},"page":"617-638","source":"Crossref","is-referenced-by-count":2,"title":["ChisNERE: a premodern Chinese corpus with named entity and relation annotation"],"prefix":"10.1093","volume":"40","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-4230-3838","authenticated-orcid":false,"given":"Xuemei","family":"Tang","sequence":"first","affiliation":[{"name":"Department of Information Management, Peking University , 100871 Beijing,","place":["China"]},{"name":"Research Center for Digital Humanities, Peking University , 100871 Beijing,","place":["China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Zekun","family":"Deng","sequence":"additional","affiliation":[{"name":"Department of Information Management, Peking University , 100871 Beijing,","place":["China"]},{"name":"Research Center for Digital Humanities, Peking University , 100871 Beijing,","place":["China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jun","family":"Wang","sequence":"additional","affiliation":[{"name":"Department of Information Management, Peking University , 100871 Beijing,","place":["China"]},{"name":"Research Center for Digital Humanities, Peking University , 100871 Beijing,","place":["China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4769-2812","authenticated-orcid":false,"given":"Qi","family":"Su","sequence":"additional","affiliation":[{"name":"Research Center for Digital Humanities, Peking University , 100871 Beijing,","place":["China"]},{"name":"School of Foreign Languages, Peking University , 100871 Beijing,","place":["China"]}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"286","published-online":{"date-parts":[[2025,3,4]]},"reference":[{"key":"2025053003523402500_fqaf001-B1","author":"Brown","year":"2020"},{"key":"2025053003523402500_fqaf001-B2","doi-asserted-by":"crossref","first-page":"571","DOI":"10.1007\/s10579-017-9382-y","article-title":"A French Clinical Corpus with Comprehensive Semantic Annotations: Development of the Medical Entity and Relation LIMSI Annotated Text Corpus (merlot)","volume":"52","author":"Campillos","year":"2018","journal-title":"Language Resources and Evaluation"},{"key":"2025053003523402500_fqaf001-B3","doi-asserted-by":"crossref","first-page":"357","DOI":"10.1162\/tacl_a_00104","article-title":"Named Entity Recognition with Bidirectional LSTM-CNNS","volume":"4","author":"Chiu","year":"2016","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"2025053003523402500_fqaf001-B4","author":"Devlin","year":"2019"},{"key":"2025053003523402500_fqaf001-B5","doi-asserted-by":"publisher","author":"Geng","year":"2022","DOI":"10.48550\/arXiv.2202.09022"},{"key":"2025053003523402500_fqaf001-B6","doi-asserted-by":"publisher","author":"Gui","year":"2023","DOI":"10.48550\/arXiv.2305.11527"},{"key":"2025053003523402500_fqaf001-B7","first-page":"94","author":"Hendrickx","year":"2009"},{"key":"2025053003523402500_fqaf001-B8","doi-asserted-by":"publisher","author":"Hu","year":"2021","DOI":"10.48550\/arXiv.2106.09685"},{"key":"2025053003523402500_fqaf001-B9","first-page":"1","article-title":"Research on Information Extraction Methods for Historical Classics under the Perspective of Digital Humanities\u2019","volume":"8","author":"Ji","year":"2021","journal-title":"Big Data Research"},{"key":"2025053003523402500_fqaf001-B10","author":"Lafferty","year":"2001"},{"key":"2025053003523402500_fqaf001-B11","doi-asserted-by":"publisher","author":"Lample","year":"2016","DOI":"10.48550\/arXiv.1603.01360"},{"key":"2025053003523402500_fqaf001-B12","first-page":"159","author":"Landis","year":"1977"},{"key":"2025053003523402500_fqaf001-B13","first-page":"108","author":"Levow","year":"2006"},{"key":"2025053003523402500_fqaf001-B14","first-page":"145","author":"Li","year":"2013"},{"key":"2025053003523402500_fqaf001-B15","doi-asserted-by":"publisher","author":"Li","year":"2020","DOI":"10.48550\/arXiv.1812.09449"},{"key":"2025053003523402500_fqaf001-B16","doi-asserted-by":"crossref","first-page":"12060","DOI":"10.3390\/app112412060","article-title":"Few-Shot Relation Extraction on Ancient Chinese Documents","volume":"11","author":"Li","year":"2021","journal-title":"Applied Sciences"},{"key":"2025053003523402500_fqaf001-B17","first-page":"1","article-title":"Ancient-Modern Chinese Translation With A Large Training Dataset","volume":"19","author":"Liu","year":"2020","journal-title":"ACM Transactions on Asian and Low-Resource Language Information Processing"},{"key":"2025053003523402500_fqaf001-B18","doi-asserted-by":"crossref","first-page":"8528","DOI":"10.1609\/aaai.v34i05.6374","article-title":"Effective Modeling of Encoder-Decoder Architecture for Joint Entity and Relation Extraction\u2019","volume":"34","author":"Nayak","year":"2020","journal-title":"Proceedings of the AAAI Conference on Artificial Intelligence"},{"key":"2025053003523402500_fqaf001-B19","first-page":"548","author":"Peng","year":"2015"},{"key":"2025053003523402500_fqaf001-B20","doi-asserted-by":"crossref","first-page":"3471","DOI":"10.1007\/s00521-021-05815-z","article-title":"A Joint Model For Entity And Relation Extraction Based On BERT\u2019","volume":"34","author":"Qiao","year":"2022","journal-title":"Neural Computing and Applications"},{"key":"2025053003523402500_fqaf001-B21","author":"Sun","year":"2019"},{"key":"2025053003523402500_fqaf001-B22","author":"Sundheim","year":"1995"},{"key":"2025053003523402500_fqaf001-B23","first-page":"678","author":"Tang","year":"2021"},{"key":"2025053003523402500_fqaf001-B24","first-page":"7830","author":"Tang","year":"2022"},{"key":"2025053003523402500_fqaf001-B25","first-page":"1","author":"Tian","year":"2021"},{"key":"2025053003523402500_fqaf001-B26","doi-asserted-by":"publisher","author":"Touvron","year":"2023","DOI":"10.48550\/arXiv.2307.09288"},{"key":"2025053003523402500_fqaf001-B27","first-page":"220","author":"Wang","year":"2021"},{"key":"2025053003523402500_fqaf001-B28","doi-asserted-by":"publisher","author":"Wang","year":"2023","DOI":"10.48550\/arXiv.2304.08085"},{"key":"2025053003523402500_fqaf001-B29","author":"Weischedel","year":"2011"},{"key":"2025053003523402500_fqaf001-B30","first-page":"178","author":"Xinyuan","year":"2022"},{"key":"2025053003523402500_fqaf001-B31","doi-asserted-by":"publisher","author":"Xu","year":"2019","DOI":"10.48550\/arXiv.1711.07010"},{"key":"2025053003523402500_fqaf001-B32","doi-asserted-by":"crossref","first-page":"181629","DOI":"10.1109\/ACCESS.2020.3026535","article-title":"MoGCN: Mixture of Gated Convolutional Neural Network for Named Entity Recognition of Chinese Historical Texts\u2019","volume":"8","author":"Yan","year":"2020","journal-title":"IEEE Access"},{"key":"2025053003523402500_fqaf001-B33","author":"Yao","year":"2019"},{"key":"2025053003523402500_fqaf001-B34","doi-asserted-by":"publisher","author":"Zhang","year":"2022","DOI":"10.48550\/arXiv.2106.08087"},{"key":"2025053003523402500_fqaf001-B35","doi-asserted-by":"publisher","author":"Zhang","year":"2019","DOI":"10.48550\/arXiv.1803.01557"},{"key":"2025053003523402500_fqaf001-B36","doi-asserted-by":"publisher","author":"Zhang","year":"2018","DOI":"10.48550\/arXiv.1805.02023"}],"container-title":["Digital Scholarship in the Humanities"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/dsh\/article-pdf\/40\/2\/617\/62266276\/fqaf001.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/dsh\/article-pdf\/40\/2\/617\/62266276\/fqaf001.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,5,30]],"date-time":"2025-05-30T07:52:45Z","timestamp":1748591565000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/dsh\/article\/40\/2\/617\/8051847"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,3,4]]},"references-count":36,"journal-issue":{"issue":"2","published-online":{"date-parts":[[2025,3,4]]},"published-print":{"date-parts":[[2025,6,1]]}},"URL":"https:\/\/doi.org\/10.1093\/llc\/fqaf001","relation":{},"ISSN":["2055-7671","2055-768X"],"issn-type":[{"value":"2055-7671","type":"print"},{"value":"2055-768X","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2025,6]]},"published":{"date-parts":[[2025,3,4]]}}}