{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,9,23]],"date-time":"2025-09-23T14:30:15Z","timestamp":1758637815581,"version":"3.40.5"},"reference-count":46,"publisher":"Oxford University Press (OUP)","issue":"7","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2015,4,1]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Motivation: Information structure (IS) analysis is a text mining technique, which classifies text in biomedical articles into categories that capture different types of information, such as objectives, methods, results and conclusions of research. It is a highly useful technique that can support a range of Biomedical Text Mining tasks and can help readers of biomedical literature find information of interest faster, accelerating the highly time-consuming process of literature review. Several approaches to IS analysis have been presented in the past, with promising results in real-world biomedical tasks. However, all existing approaches, even weakly supervised ones, require several hundreds of hand-annotated training sentences specific to the domain in question. Because biomedicine is subject to considerable domain variation, such annotations are expensive to obtain. This makes the application of IS analysis across biomedical domains difficult. In this article, we investigate an unsupervised approach to IS analysis and evaluate the performance of several unsupervised methods on a large corpus of biomedical abstracts collected from PubMed.<\/jats:p><jats:p>Results: Our best unsupervised algorithm (multilevel-weighted graph clustering algorithm) performs very well on the task, obtaining over 0.70 F scores for most IS categories when applied to well-known IS schemes. This level of performance is close to that of lightly supervised IS methods and has proven sufficient to aid a range of practical tasks. Thus, using an unsupervised approach, IS could be applied to support a wide range of tasks across sub-domains of biomedicine. We also demonstrate that unsupervised learning brings novel insights into IS of biomedical literature and discovers information categories that are not present in any of the existing IS schemes.<\/jats:p><jats:p>Availability and Implementation: The annotated corpus and software are available at http:\/\/www.cl.cam.ac.uk\/\u223cdk427\/bio14info.html.<\/jats:p><jats:p>Contact: \u00a0alk23@cam.ac.uk<\/jats:p>","DOI":"10.1093\/bioinformatics\/btu758","type":"journal-article","created":{"date-parts":[[2014,11,20]],"date-time":"2014-11-20T06:56:29Z","timestamp":1416466589000},"page":"1084-1092","source":"Crossref","is-referenced-by-count":5,"title":["Unsupervised discovery of information structure in biomedical documents"],"prefix":"10.1093","volume":"31","author":[{"given":"Douwe","family":"Kiela","sequence":"first","affiliation":[{"name":"1 Computer Laboratory, University of Cambridge, Cambridge CB3 0FD, UK and 2Institute of Environmental Medicine, Karolinska Institutet, Stockholm SE-171 77, Sweden"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yufan","family":"Guo","sequence":"additional","affiliation":[{"name":"1 Computer Laboratory, University of Cambridge, Cambridge CB3 0FD, UK and 2Institute of Environmental Medicine, Karolinska Institutet, Stockholm SE-171 77, Sweden"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ulla","family":"Stenius","sequence":"additional","affiliation":[{"name":"1 Computer Laboratory, University of Cambridge, Cambridge CB3 0FD, UK and 2Institute of Environmental Medicine, Karolinska Institutet, Stockholm SE-171 77, Sweden"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Anna","family":"Korhonen","sequence":"additional","affiliation":[{"name":"1 Computer Laboratory, University of Cambridge, Cambridge CB3 0FD, UK and 2Institute of Environmental Medicine, Karolinska Institutet, Stockholm SE-171 77, Sweden"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"286","published-online":{"date-parts":[[2014,11,18]]},"reference":[{"key":"2023051309142480200_btu758-B1","doi-asserted-by":"crossref","first-page":"3174","DOI":"10.1093\/bioinformatics\/btp548","article-title":") Automatically classifying sentences in full-text biomedical articles into introduction, methods, results and discussion","volume":"25","author":"Agarwal","year":"2009","journal-title":"Bioinformatics"},{"key":"2023051309142480200_btu758-B2","doi-asserted-by":"crossref","first-page":"461","DOI":"10.1007\/s10791-008-9066-8","article-title":"A comparison of extrinsic clustering evaluation metrics based on formal constraints","volume":"12","author":"Amig\u00f3","year":"2009","journal-title":"Inf. Retr."},{"key":"2023051309142480200_btu758-B3","doi-asserted-by":"crossref","first-page":"173","DOI":"10.1016\/j.jbi.2009.11.001","article-title":"Beyond genes, proteins, and abstracts: Identifying scientific claims from full-text biomedical articles","volume":"43","author":"Blake","year":"2009","journal-title":"J. Biomed. Inform."},{"key":"2023051309142480200_btu758-B4","first-page":"993","article-title":"Latent dirichlet allocation","volume":"3","author":"Blei","year":"2003","journal-title":"J. Machine Learn. Res."},{"key":"2023051309142480200_btu758-B5","doi-asserted-by":"crossref","first-page":"757","DOI":"10.1016\/j.jbi.2009.09.001","article-title":"Current issues in biomedical text mining and natural language processing","volume":"5","author":"Chapman","year":"2009","journal-title":"J. Biomed. Inform."},{"key":"2023051309142480200_btu758-B6","first-page":"663","article-title":"Using argumentative zones for extractive summarization of scientific articles","volume-title":"Proceedings of the International Conference on Computational Linguistics (COLING)","author":"Contractor","year":"2012"},{"key":"2023051309142480200_btu758-B7","first-page":"33","article-title":"Linguistically motivated large-scale nlp with c&c and boxer","volume-title":"ACL '07 Proceedings of the 45th Annual Meeting of the ACL on Interactive Poster and Demonstration Sessions, ACL, Prague, Czech Republic","author":"Curran","year":"2007"},{"key":"2023051309142480200_btu758-B8","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1111\/j.2517-6161.1977.tb01600.x","article-title":"Maximum likelihood from incomplete data via the EM algorithm","volume":"39","author":"Dempster","year":"1977","journal-title":"J. R. Stat. Soc. Ser. B"},{"key":"2023051309142480200_btu758-B9","doi-asserted-by":"crossref","first-page":"143","DOI":"10.1023\/A:1007612920971","article-title":"Concept decompositions for large sparse text data using clustering","volume":"42","author":"Dhillon","year":"2001","journal-title":"Machine Learn."},{"key":"2023051309142480200_btu758-B10","doi-asserted-by":"crossref","first-page":"551","DOI":"10.1145\/1014052.1014118","article-title":"Kernel k-means, spectral clustering and normalized cuts","volume-title":"Proceedings of the 10th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining","author":"Dhillon","year":"2004"},{"key":"2023051309142480200_btu758-B11","first-page":"629","article-title":"A fast kernel-based multilevel algorithm for graph clustering","volume-title":"Proceedings of the 11th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining","author":"Dhillon","year":"2005"},{"key":"2023051309142480200_btu758-B12","doi-asserted-by":"crossref","first-page":"1944","DOI":"10.1109\/TPAMI.2007.1115","article-title":"Weighted graph cuts without eigenvectors: a multilevel approach","volume":"29","author":"Dhillon","year":"2007","journal-title":"IEEE Trans. Pattern Anal. Machine Intell."},{"key":"2023051309142480200_btu758-B13","first-page":"99","article-title":"Identifying the information structure of scientific abstracts: an investigation of three different schemes","volume-title":"Proceedings of BioNLP, ACL 2010 in Uppsala, Sweden","author":"Guo","year":"2010"},{"key":"2023051309142480200_btu758-B14","article-title":"A comparison and user-based evaluation of models of textual information structure in the context of cancer risk assessment","volume":"69","author":"Guo","year":"2011","journal-title":"BMC Bioinformatics."},{"key":"2023051309142480200_btu758-B15","doi-asserted-by":"crossref","first-page":"3179","DOI":"10.1093\/bioinformatics\/btr536","article-title":"Weakly-supervised learning of information structure of scientific abstracts\u2013is it accurate enough to benefit real-world tasks in biomedicine?","volume":"27","author":"Guo","year":"2011","journal-title":"Bioinformatics."},{"key":"2023051309142480200_btu758-B16","doi-asserted-by":"crossref","first-page":"1440","DOI":"10.1093\/bioinformatics\/btt163","article-title":"Active learning-based information structure analysis of full scientific articles and two applications for biomedical literature review","volume":"29","author":"Guo","year":"2013","journal-title":"Bioinformatics."},{"key":"2023051309142480200_btu758-B17","doi-asserted-by":"crossref","first-page":"17","DOI":"10.1186\/1479-7364-5-1-17","article-title":"What the papers say: text mining for genomics and systems biology","volume":"5","author":"Harmston","year":"2010","journal-title":"Hum. Genomics"},{"key":"2023051309142480200_btu758-B18","first-page":"381","article-title":"Identifying sections in scientific abstracts using conditional random fields","volume-title":"Proceedings of 3rd International Joint Conference on Natural Language Processing","author":"Hirohata","year":"2008"},{"key":"2023051309142480200_btu758-B19","doi-asserted-by":"crossref","first-page":"11","DOI":"10.1108\/eb026526","article-title":"A statistical interpretation of term specificity and its application in retrieval","volume":"28","author":"Jones","year":"1972","journal-title":"J. Doc."},{"key":"2023051309142480200_btu758-B20","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/1497577.1497578","article-title":"Clustering high-dimensional data: A survey on subspace clustering, pattern-based clustering, and correlation clustering","volume":"3","author":"Kriegel","year":"2009","journal-title":"ACM Trans. Knowledge Discov. Data"},{"key":"2023051309142480200_btu758-B21","first-page":"2054","article-title":"Corpora for the conceptualisation and zoning of scientific papers","volume-title":"Proceedings of the International Conference on Language Resources and Evaluation (LREC)","author":"Liakata","year":"2010"},{"key":"2023051309142480200_btu758-B22","doi-asserted-by":"crossref","first-page":"65","DOI":"10.3115\/1654415.1654427","article-title":"Generative content models for structural analysis of medical abstracts","volume-title":"HLT-NAACL BioNLP Workshop on Linking Natural Language and Biology","author":"Lin","year":"2006"},{"key":"2023051309142480200_btu758-B23","doi-asserted-by":"crossref","first-page":"212","DOI":"10.1186\/1471-2105-12-212","article-title":"Exploring subdomain variation in biomedical language","volume":"12","author":"Lippincott","year":"2011","journal-title":"BMC Bioinformatics"},{"key":"2023051309142480200_btu758-B24","first-page":"281","article-title":"Some methods for classification and analysis of multivariate observations","volume-title":"Proc. of the fifth Berkeley Symposium on Mathematical Statistics and Probability","author":"MacQueen","year":"1967"},{"key":"2023051309142480200_btu758-B25","article-title":"Value and benefits of text mining","volume":"811","author":"McDonald","year":"2012","journal-title":"Technical report"},{"key":"2023051309142480200_btu758-B26","article-title":"Analysing entity type variation across biomedical subdomains","volume-title":"Proceedings of the Third Workshop on Building and Evaluating Resources for Biomedical Text Mining (BioTxtM 2012)","author":"Mih\u0103il\u0103","year":"2012"},{"key":"2023051309142480200_btu758-B27","doi-asserted-by":"crossref","first-page":"468","DOI":"10.1016\/j.ijmedinf.2005.06.013","article-title":"Zone analysis in biology articles as a basis for information extraction","volume":"75","author":"Mizuta","year":"2006","journal-title":"Int. J. Med. Inform."},{"key":"2023051309142480200_btu758-B28","first-page":"52","article-title":"A baseline feature set for learning rhetorical zones using full articles in the biomedical domain","volume":"7","author":"Mullen","year":"2005","journal-title":"Nat. Lang. Process. Text Mining"},{"key":"2023051309142480200_btu758-B29","first-page":"355","article-title":"A view of the EM algorithm that justifies incremental, sparse, and other variants","volume-title":"Learning in Graphical Models","author":"Neal","year":"1999"},{"key":"2023051309142480200_btu758-B30","first-page":"410","article-title":"V-measure: A conditional entropy-based external cluster evaluation measure","volume-title":"Proceedings of the 2007 Joint Conference on Empirical Methods in Natural Language Processing and Computational Natural Language Learning (EMNLP-CoNLL)","author":"Rosenberg","year":"2007"},{"key":"2023051309142480200_btu758-B31","doi-asserted-by":"crossref","first-page":"195","DOI":"10.1016\/j.ijmedinf.2006.05.002","article-title":"Using argumentation to extract key sentences from biomedical abstracts","volume":"76","author":"Ruch","year":"2007","journal-title":"Int. J. Med. Inform."},{"volume-title":"Part-of-speech tagging guidelines for the penn treebank project (3rd revision)","year":"1990","author":"Santorini","key":"2023051309142480200_btu758-B32"},{"key":"2023051309142480200_btu758-B33","doi-asserted-by":"crossref","first-page":"465","DOI":"10.1007\/978-1-4614-3223-4_14","article-title":"Biomedical text mining: a survey of recent progress","volume-title":"Mining Text Data","author":"Simpson","year":"2012"},{"key":"2023051309142480200_btu758-B34","first-page":"129","article-title":"Parsing natural scenes and natural language with recursive neural networks","volume-title":"The 28th International Conference on Machine Learning (ICML)","author":"Socher","year":"2011"},{"key":"2023051309142480200_btu758-B35","first-page":"1631","article-title":"Recursive deep models for semantic compositionality over a sentiment treebank","volume-title":"Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP)","author":"Socher","year":"2013"},{"key":"2023051309142480200_btu758-B36","first-page":"364","article-title":"The introduction, methods, results, and discussion (IMRAD) structure: a fifty-year survey","volume":"92","author":"Sollaci","year":"2004","journal-title":"J. Med. Libr. Assoc."},{"key":"2023051309142480200_btu758-B37","first-page":"638","article-title":"Improving verb clustering with automatically acquired selectional preference","volume-title":"Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP), ACl, Suntec, Singapore","author":"Sun","year":"2009"},{"key":"2023051309142480200_btu758-B38","doi-asserted-by":"crossref","first-page":"488","DOI":"10.1016\/j.ijmedinf.2005.06.007","article-title":"Using argumentation to retrieve articles with similar citations","volume":"75","author":"Tbahriti","year":"2006","journal-title":"Int. J. Med. Inform."},{"key":"2023051309142480200_btu758-B39","doi-asserted-by":"crossref","first-page":"409","DOI":"10.1162\/089120102762671936","article-title":"Summarizing scientific articles: experiments with relevance and rhetorical status","volume":"28","author":"Teufel","year":"2002","journal-title":"Comput. Linguist."},{"key":"2023051309142480200_btu758-B40","first-page":"110","article-title":"An annotation scheme for discourse-level argumentation in research articles","volume-title":"Proceedings of the Conference of the European Chapter of the Association for Computational Linguistics (EACL)","author":"Teufel","year":"1999"},{"key":"2023051309142480200_btu758-B41","first-page":"1493","article-title":"Towards domain-independent argumentative zoning: Evidence from chemistry and computational linguistics","volume-title":"Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP), ACL, Suntec, Singapore","author":"Teufel","year":"2009"},{"key":"2023051309142480200_btu758-B42","doi-asserted-by":"crossref","first-page":"365","DOI":"10.1108\/eb026584","article-title":"Foundation of evaluation","volume":"30","author":"van Rijsbergen","year":"1974","journal-title":"J. Doc."},{"key":"2023051309142480200_btu758-B43","first-page":"1610","article-title":"Unsupervised document zone identification using probabilistic graphical models","volume-title":"Proceedings of the International Conference on Language Resources and Evaluation (LREC)","author":"Varga","year":"2012"},{"key":"2023051309142480200_btu758-B44","doi-asserted-by":"crossref","first-page":"437","DOI":"10.1017\/S1351324911000337","article-title":"Discourse structure and language technology","volume":"18","author":"Webber","year":"2011","journal-title":"Nat. Lang. Eng."},{"key":"2023051309142480200_btu758-B45","doi-asserted-by":"crossref","first-page":"356","DOI":"10.1186\/1471-2105-7-356","article-title":"New directions in biomedical text annotation: definitions, guidelines and corpus construction","volume":"7","author":"Wilbur","year":"2006","journal-title":"BMC Bioinformatics"},{"key":"2023051309142480200_btu758-B46","first-page":"3180","article-title":"Efficient online spherical k-means clustering","volume-title":"Proceedings of IEEE International Joint Conference on Neural Networks (IJCNN 2005)","author":"Zhong","year":"2005"}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/31\/7\/1084\/50306004\/bioinformatics_31_7_1084.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/31\/7\/1084\/50306004\/bioinformatics_31_7_1084.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,5,13]],"date-time":"2025-05-13T19:46:58Z","timestamp":1747165618000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/31\/7\/1084\/180459"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2014,11,18]]},"references-count":46,"journal-issue":{"issue":"7","published-print":{"date-parts":[[2015,4,1]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/btu758","relation":{},"ISSN":["1367-4811","1367-4803"],"issn-type":[{"type":"electronic","value":"1367-4811"},{"type":"print","value":"1367-4803"}],"subject":[],"published-other":{"date-parts":[[2015,4,1]]},"published":{"date-parts":[[2014,11,18]]}}}