{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,24]],"date-time":"2026-05-24T14:03:41Z","timestamp":1779631421860,"version":"3.53.1"},"reference-count":29,"publisher":"SAGE Publications","issue":"3","license":[{"start":{"date-parts":[[2026,5,8]],"date-time":"2026-05-08T00:00:00Z","timestamp":1778198400000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/journals.sagepub.com\/page\/policies\/text-and-data-mining-license"}],"content-domain":{"domain":["journals.sagepub.com"],"crossmark-restriction":true},"short-container-title":["Big Data"],"published-print":{"date-parts":[[2026,6,1]]},"abstract":"<jats:p>\n                    Patent text segmentation is a fundamental task in patent data mining, enabling applications such as patent analysis and search. The objective is to decompose structurally complex, lengthy sentences into grammatically complete, semantically equivalent short sentences to facilitate downstream processing. Traditional approaches rely on manually defined rules or feature-based machine learning methods, which are labor-intensive, domain-specific, and exhibit limited generalizability. To overcome these limitations, this study proposes a Deep Segmentation Model for Patent Text (DS\n                    <jats:sup>2<\/jats:sup>\n                    PT), a two-stage fine-grained segmentation framework based on ALBERT. The first stage employs a conditional random field model to perform coarse segmentation of patent paragraphs into shorter clauses based on structural cues. The second stage utilizes the ALBERT model to perform deep, context-aware segmentation of complex clauses into syntactically independent and semantically complete sentences. Compared to conventional methods, DS\n                    <jats:sup>2<\/jats:sup>\n                    PT effectively captures hierarchical contextual information across two stages, significantly improving segmentation accuracy without semantic loss. Furthermore, this research draws inspiration from advancements in cross-lingual speech-to-text systems with low-latency neural networks for real-time applications. While the domains differ, the core technical challenges are analogous: both require models to process sequential, information-dense input (audio streams or long sentences) into structured, meaningful units (transcribed text or segmented clauses) with high accuracy and efficiency. The principles of low-latency neural networks\u2014such as efficient context modeling, parallelizable architectures, and real-time incremental processing\u2014inform the design of our segmentation pipeline to enhance its scalability and potential for integration into real-time patent analysis systems. Similarly, the cross-lingual capability highlights the importance of model generalization, which aligns with our goal of developing a domain-adaptive segmentation tool for diverse patent corpora.\n                  <\/jats:p>","DOI":"10.1177\/2167647x261447812","type":"journal-article","created":{"date-parts":[[2026,5,8]],"date-time":"2026-05-08T11:26:14Z","timestamp":1778239574000},"page":"244-255","update-policy":"https:\/\/doi.org\/10.1177\/sage-journals-update-policy","source":"Crossref","is-referenced-by-count":0,"title":["DS\n                    <sup>2<\/sup>\n                    PT: A Deep Two-Stage Patent Text Segmentation Framework Informed by Low-Latency Neural Network Characteristics"],"prefix":"10.1177","volume":"14","author":[{"ORCID":"https:\/\/orcid.org\/0009-0004-3674-2492","authenticated-orcid":false,"given":"Boting","family":"Geng","sequence":"first","affiliation":[{"name":"School of Computer Science and Technology, Zhejiang University of Water Resources and Electric Power, Hangzhou, China."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Hongxia","family":"Wang","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Zhejiang University of Water Resources and Electric Power, Hangzhou, China."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Pengliang","family":"Zhang","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Zhejiang University of Water Resources and Electric Power, Hangzhou, China."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jin","family":"Xue","sequence":"additional","affiliation":[{"name":"School of Humanities and Foreign Languages, Zhejiang Shuren University, Hangzhou, China."}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"179","published-online":{"date-parts":[[2026,5,8]]},"reference":[{"issue":"2","key":"e_1_3_2_2_1","first-page":"ivae143","article-title":"Effect of patent complete revascularization on the akinetic myocardial segments[J]","volume":"39","author":"Kim MS","year":"2024","unstructured":"1.Kim MS, , Kim MJ, , Jeong HJ, et al. Effect of patent complete revascularization on the akinetic myocardial segments[J]. Interdiscip Cardiovasc Thorac Surg 2024;39(2):ivae143.","journal-title":"Interdiscip Cardiovasc Thorac Surg"},{"key":"e_1_3_2_3_1","unstructured":"2.Nishimura M Buma K Utsuro T et al. Improving Japanese-English patent claim translation with clause segmentation models based on word alignment[C]. Proceedings of Machine Translation Summit XX: Volume 1. 2025:333\u2013343."},{"issue":"11","key":"e_1_3_2_4_1","first-page":"9257","article-title":"Enhancing patent text classification with Bi-LSTM technique and alpine skiing optimization for improved diagnostic accuracy[J]","volume":"84","author":"Wang J","year":"2024","unstructured":"3.Wang J, , Wang L, , Ji N, et al. Enhancing patent text classification with Bi-LSTM technique and alpine skiing optimization for improved diagnostic accuracy[J]. Multimed Tools Appl 2024;84(11):9257\u20139286.","journal-title":"Multimed Tools Appl"},{"key":"e_1_3_2_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2022.3176877"},{"key":"e_1_3_2_6_1","unstructured":"5.Fujii A Ishikawa T. Document structure analysis for the NTCIR-5 patent retrieval task[C]. In Proceedings of the 5th NTCIR Workshop Meeting on Evalution of Information Access Technologies: Information Retrieval Question Answering and Cross-lingual Information Access. National Institute of Informatics Tokyo Japan 2005:1\u20135."},{"key":"e_1_3_2_7_1","doi-asserted-by":"crossref","unstructured":"6.Takaki T Fujii A Ishikawa T. Associative document retrieval by query subtopic analysis and its application to invalidity patent search[C]. In Proceedings of the 13th ACM International Conference on Information and Knowledge Management. Association for Computing Machinery Washington D.C. USA 2004:399\u2013405.","DOI":"10.1145\/1031171.1031251"},{"key":"e_1_3_2_8_1","doi-asserted-by":"crossref","unstructured":"7.Shinmori A Okumura M Marukawa Y et al. Patent claim processing for readability: Structure analysis and term explanation[C]. In Proceedings of the ACL-2003 Workshop on Patent Corpus Processing. Association for Computational Linguistics Sapporo Japan 2003:56\u201365.","DOI":"10.3115\/1119303.1119310"},{"key":"e_1_3_2_9_1","doi-asserted-by":"crossref","unstructured":"8.Sheremetyeva S. Automatic text simplification for handling intellectual property (the case of multiple patent claims)[C]. In Proceedings of the Workshop on Automatic Text Simplification-Methods and Applications in the Multilingual Society. Association for Computational Linguistics Dublin Ireland 2014:41\u201352.","DOI":"10.3115\/v1\/W14-5605"},{"key":"e_1_3_2_10_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.ipm.2007.02.002"},{"key":"e_1_3_2_11_1","doi-asserted-by":"crossref","unstructured":"10.Agatonovic M Aswani N Bontcheva K et al. Large-scale parallel automatic patent annotation[C]. In Proceedings of the 1st ACM Workshop on Patent Information Retrieval. Association for Computing Machinery Napa Valley California USA 2008:1\u20138.","DOI":"10.1145\/1458572.1458574"},{"key":"e_1_3_2_12_1","doi-asserted-by":"crossref","unstructured":"11.Ferraro G Suominen H Nualart J. Segmentation of patent claims for improving their readability[C]. In Proceedings of the 3rd Worshop on Predicting and Improving Text Readability for Target Reader Populations. Association for Computational Linguistics Gothenburg Sweden 2014:66\u201373.","DOI":"10.3115\/v1\/W14-1208"},{"issue":"5","key":"e_1_3_2_13_1","first-page":"71","article-title":"Patent Abstract Text Segmentation Technology Based on Classification Algorithms [J]","volume":"2012","author":"Changlin D","unstructured":"12.Changlin D, , Dongfeng C, , Peiyan W. Patent Abstract Text Segmentation Technology Based on Classification Algorithms [J]. Journal of Shandong University (Science Edition); 2012(5):71\u201375.","journal-title":"Journal of Shandong University (Science Edition)"},{"key":"e_1_3_2_14_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.bdr.2020.100133"},{"key":"e_1_3_2_15_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-64283-3_25"},{"key":"e_1_3_2_16_1","doi-asserted-by":"publisher","DOI":"10.3991\/ijet.v14i19.10182"},{"key":"e_1_3_2_17_1","unstructured":"16.Carvalho DS Nguyen ML. Efficient neural-based patent document segmentation with term order probabilities[C]. In Proceedings of the 25th European Symposium on Artificial Neural Networks. Elsevier Bruges Belgium 2017:171\u2013178."},{"key":"e_1_3_2_18_1","doi-asserted-by":"crossref","unstructured":"17.Ratinov L Roth D. Design challenges and misconceptions in named entity recognition[C]. In Proceedings of the 13th Conference on Computational Natural Language Learning. Association for Computational Linguistics Boulder Colorado USA 2009:147\u2013155.","DOI":"10.3115\/1596374.1596399"},{"key":"e_1_3_2_19_1","doi-asserted-by":"publisher","DOI":"10.3390\/app10248924"},{"key":"e_1_3_2_20_1","unstructured":"19.Devlin J Chang MW Lee K et al. Bert: Pre-training of deep bidirectional transformers for language understanding[C]. In Proceedings of the 17th Annual Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics Minneapolis Minnesota USA 2018: 4171\u20134186."},{"key":"e_1_3_2_21_1","unstructured":"20.Du J Huang Y Moilanen K. AIG investments AI at the FinSBD task: Sentence boundary detection through sequence labelling and BERT fine-tuning[C]. In Proceedings of the 1st Workshop on Financial Technology and Natural Language Processing. Morgan Kaufmann Macao China 2019:81\u201387."},{"key":"e_1_3_2_22_1","doi-asserted-by":"crossref","unstructured":"21.Kumar CSA Maharana A Murali S et al. BERT-based sequence labelling approach for dependency parsing in Tamil[C]. In Proceedings of the 2nd Workshop on Speech and Language Technologies for Dravidian Languages. Association for Computational Linguistics Dublin Ireland 2022:1\u20138.","DOI":"10.18653\/v1\/2022.dravidianlangtech-1.1"},{"key":"e_1_3_2_23_1","doi-asserted-by":"crossref","unstructured":"22.Liu W Fu X Zhang Y et al. Lexicon enhanced Chinese sequence labeling using Bert adapter[C]. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing. Association for Computational Linguistics Online 2021:5847\u20135858.","DOI":"10.18653\/v1\/2021.acl-long.454"},{"key":"e_1_3_2_24_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-540-30586-6_76"},{"key":"e_1_3_2_25_1","doi-asserted-by":"publisher","DOI":"10.1109\/29.31269"},{"key":"e_1_3_2_26_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-21042-1_12"},{"key":"e_1_3_2_27_1","unstructured":"26.Lafferty J Mccallum A Pereira FCN. Conditional random fields: Probabilistic models for segmenting and labeling sequence data[C]. In Proceedings of the 18th International Conference on Machine Learning 2001. Association for Computing Machinery Williams College Williamstown MA USA 2001:282\u2013289."},{"key":"e_1_3_2_28_1","unstructured":"27.Lan Z Chen M Goodman S et al. Albert: A lite Bert for self-supervised learning of language representations[C]. In Proceedings of the 8th International Conference on Learning Representations. OpenReview.net Addis Ababa Ethiopia 2019:1\u201317."},{"key":"e_1_3_2_29_1","unstructured":"28.Sanh V Debut L Chaumond J et al. DistilBERT a distilled version of BERT: Smaller faster cheaper and lighter[C]. In Proceedings of the 33th International Conference On Neural Information Processing Systems. Morgan Kaufmann Vancouver BC Canada 2019:1\u20135."},{"key":"e_1_3_2_30_1","unstructured":"29.Zhuang L Wayne L Ya S et al. A robustly optimized BERT pre-training approach with post-training[C]. In Proceedings of the 20th Chinese National Conference on Computational Linguistics. Chinese Information Processing Society of China Huhhot China 2021:1218\u20131227."}],"container-title":["Big Data"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/2167647X261447812","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/full-xml\/10.1177\/2167647X261447812","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/2167647X261447812","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,5,24]],"date-time":"2026-05-24T13:46:11Z","timestamp":1779630371000},"score":1,"resource":{"primary":{"URL":"https:\/\/journals.sagepub.com\/doi\/10.1177\/2167647X261447812"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,5,8]]},"references-count":29,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2026,6,1]]}},"alternative-id":["10.1177\/2167647X261447812"],"URL":"https:\/\/doi.org\/10.1177\/2167647x261447812","relation":{},"ISSN":["2167-6461","2167-647X"],"issn-type":[{"value":"2167-6461","type":"print"},{"value":"2167-647X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,5,8]]}}}