{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,18]],"date-time":"2026-03-18T01:00:32Z","timestamp":1773795632965,"version":"3.50.1"},"reference-count":46,"publisher":"Oxford University Press (OUP)","issue":"7","license":[{"start":{"date-parts":[[2016,10,2]],"date-time":"2016-10-02T00:00:00Z","timestamp":1475366400000},"content-version":"vor","delay-in-days":1698,"URL":"http:\/\/creativecommons.org\/licenses\/by-nc\/3.0"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2012,4,1]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>Motivation: Scholarly biomedical publications report on the findings of a research investigation. Scientists use a well-established discourse structure to relate their work to the state of the art, express their own motivation and hypotheses and report on their methods, results and conclusions. In previous work, we have proposed ways to explicitly annotate the structure of scientific investigations in scholarly publications. Here we present the means to facilitate automatic access to the scientific discourse of articles by automating the recognition of 11 categories at the sentence level, which we call Core Scientific Concepts (CoreSCs). These include: Hypothesis, Motivation, Goal, Object, Background, Method, Experiment, Model, Observation, Result and Conclusion. CoreSCs provide the structure and context to all statements and relations within an article and their automatic recognition can greatly facilitate biomedical information extraction by characterizing the different types of facts, hypotheses and evidence available in a scientific publication.<\/jats:p>\n               <jats:p>Results: We have trained and compared machine learning classifiers (support vector machines and conditional random fields) on a corpus of 265 full articles in biochemistry and chemistry to automatically recognize CoreSCs. We have evaluated our automatic classifications against a manually annotated gold standard, and have achieved promising accuracies with \u2018Experiment\u2019, \u2018Background\u2019 and \u2018Model\u2019 being the categories with the highest F1-scores (76%, 62% and 53%, respectively). We have analysed the task of CoreSC annotation both from a sentence classification as well as sequence labelling perspective and we present a detailed feature evaluation. The most discriminative features are local sentence features such as unigrams, bigrams and grammatical dependencies while features encoding the document structure, such as section headings, also play an important role for some of the categories. We discuss the usefulness of automatically generated CoreSCs in two biomedical applications as well as work in progress.<\/jats:p>\n               <jats:p>Availability: A web-based tool for the automatic annotation of articles with CoreSCs and corresponding documentation is available online at http:\/\/www.sapientaproject.com\/softwarehttp:\/\/www.sapientaproject.com also contains detailed information pertaining to CoreSC annotation and links to annotation guidelines as well as a corpus of manually annotated articles, which served as our training data.<\/jats:p>\n               <jats:p>Contact: \u00a0liakata@ebi.ac.uk<\/jats:p>\n               <jats:p>Supplementary information: \u00a0Supplementary data are available at Bioinformatics online.<\/jats:p>","DOI":"10.1093\/bioinformatics\/bts071","type":"journal-article","created":{"date-parts":[[2012,2,10]],"date-time":"2012-02-10T01:22:15Z","timestamp":1328836935000},"page":"991-1000","source":"Crossref","is-referenced-by-count":88,"title":["Automatic recognition of conceptualization zones in scientific articles and two life science applications"],"prefix":"10.1093","volume":"28","author":[{"given":"Maria","family":"Liakata","sequence":"first","affiliation":[{"name":"1 Department of Computer Science, Aberystwyth University, Aberystwyth, Ceredigion, SY23 3DB, UK, 2Rebholz Group, EMBL-EBI, Wellcome Trust Genome Campus, Hinxton, Cambridge, CB10 1SD, UK, 3Department of Philosophy, Linguistics and Theory of Science, University of Gothenburg, Gothenburg, 405 30, Sweden and 4Royal Society of Chemistry, Cambridge, Thomas Graham House, Science Park, Milton Road, Cambridge CB4 0WF, UK"},{"name":"1 Department of Computer Science, Aberystwyth University, Aberystwyth, Ceredigion, SY23 3DB, UK, 2Rebholz Group, EMBL-EBI, Wellcome Trust Genome Campus, Hinxton, Cambridge, CB10 1SD, UK, 3Department of Philosophy, Linguistics and Theory of Science, University of Gothenburg, Gothenburg, 405 30, Sweden and 4Royal Society of Chemistry, Cambridge, Thomas Graham House, Science Park, Milton Road, Cambridge CB4 0WF, UK"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Shyamasree","family":"Saha","sequence":"additional","affiliation":[{"name":"1 Department of Computer Science, Aberystwyth University, Aberystwyth, Ceredigion, SY23 3DB, UK, 2Rebholz Group, EMBL-EBI, Wellcome Trust Genome Campus, Hinxton, Cambridge, CB10 1SD, UK, 3Department of Philosophy, Linguistics and Theory of Science, University of Gothenburg, Gothenburg, 405 30, Sweden and 4Royal Society of Chemistry, Cambridge, Thomas Graham House, Science Park, Milton Road, Cambridge CB4 0WF, UK"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Simon","family":"Dobnik","sequence":"additional","affiliation":[{"name":"1 Department of Computer Science, Aberystwyth University, Aberystwyth, Ceredigion, SY23 3DB, UK, 2Rebholz Group, EMBL-EBI, Wellcome Trust Genome Campus, Hinxton, Cambridge, CB10 1SD, UK, 3Department of Philosophy, Linguistics and Theory of Science, University of Gothenburg, Gothenburg, 405 30, Sweden and 4Royal Society of Chemistry, Cambridge, Thomas Graham House, Science Park, Milton Road, Cambridge CB4 0WF, UK"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Colin","family":"Batchelor","sequence":"additional","affiliation":[{"name":"1 Department of Computer Science, Aberystwyth University, Aberystwyth, Ceredigion, SY23 3DB, UK, 2Rebholz Group, EMBL-EBI, Wellcome Trust Genome Campus, Hinxton, Cambridge, CB10 1SD, UK, 3Department of Philosophy, Linguistics and Theory of Science, University of Gothenburg, Gothenburg, 405 30, Sweden and 4Royal Society of Chemistry, Cambridge, Thomas Graham House, Science Park, Milton Road, Cambridge CB4 0WF, UK"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Dietrich","family":"Rebholz-Schuhmann","sequence":"additional","affiliation":[{"name":"1 Department of Computer Science, Aberystwyth University, Aberystwyth, Ceredigion, SY23 3DB, UK, 2Rebholz Group, EMBL-EBI, Wellcome Trust Genome Campus, Hinxton, Cambridge, CB10 1SD, UK, 3Department of Philosophy, Linguistics and Theory of Science, University of Gothenburg, Gothenburg, 405 30, Sweden and 4Royal Society of Chemistry, Cambridge, Thomas Graham House, Science Park, Milton Road, Cambridge CB4 0WF, UK"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"286","published-online":{"date-parts":[[2012,2,8]]},"reference":[{"key":"2023012512240524400_B1","doi-asserted-by":"crossref","first-page":"381","DOI":"10.1016\/j.tibtech.2010.04.005","article-title":"Event extraction for systems biology by text mining the literature","volume":"28","author":"Ananiadou","year":"2010","journal-title":"Trends Biotechnol."},{"key":"2023012512240524400_B2","doi-asserted-by":"crossref","first-page":"495","DOI":"10.1162\/coli.2009.35.4.35402","article-title":"From annotator agreement to noise models","volume":"35","author":"Beigman","year":"2009","journal-title":"Comput. Linguist."},{"key":"2023012512240524400_B3","doi-asserted-by":"crossref","first-page":"675","DOI":"10.1016\/0306-4573(95)00052-I","article-title":"Automatic condensation of electronic publications by sentence selection","volume":"31","author":"Brandow","year":"1995","journal-title":"Inform. Process. Manag."},{"key":"2023012512240524400_B4","doi-asserted-by":"crossref","first-page":"27:1","DOI":"10.1145\/1961189.1961199","article-title":"LIBSVM: a library for support vector machines","volume":"2","author":"Chang","year":"2011","journal-title":"ACM Trans. Intell. Syst. Technol."},{"key":"2023012512240524400_B5","doi-asserted-by":"crossref","first-page":"739","DOI":"10.1016\/j.jbi.2008.04.010","article-title":"The swan biomedical discourse ontology","volume":"41","author":"Ciccarese","year":"2008","journal-title":"J. Biomed. Inform."},{"key":"2023012512240524400_B6","doi-asserted-by":"crossref","first-page":"57","DOI":"10.1093\/bib\/6.1.57","article-title":"A survey of current work in biomedical text a survey of current work in biomedical text mining","volume":"6","author":"Cohen","year":"2005","journal-title":"Brief. Bioinform."},{"key":"2023012512240524400_B7","doi-asserted-by":"crossref","first-page":"37","DOI":"10.1177\/001316446002000104","article-title":"A coefficient of agreement for nominal scales","volume":"20","author":"Cohen","year":"1960","journal-title":"Educ. Psychol. Meas."},{"key":"2023012512240524400_B8","first-page":"33","article-title":"Linguistically motivated large-scale nlp with c&c and boxer","volume-title":"Proceedings of the 45th Annual Meeting of the Association for Computational Linguistics Companion Volume Proceedings of the Demo and Poster Sessions.","author":"Curran","year":"2007"},{"key":"2023012512240524400_B9","first-page":"351","article-title":"Identifying the epistemic value of discourse segments in biology texts","volume-title":"Proceedings of the Eighth International Conference on Computational Semantics, IWCS-8 '09","author":"de Waard","year":"2009"},{"key":"2023012512240524400_B10","doi-asserted-by":"crossref","first-page":"83","DOI":"10.1145\/1324237.1324247","article-title":"A pragmatic structure for research articles","volume-title":"Proceedings of the 2nd International Conference on Pragmatic Web, ICPW '07","author":"de Waard","year":"2007"},{"key":"2023012512240524400_B11","first-page":"1871","article-title":"LIBLINEAR: a library for large linear classification","volume":"9","author":"Fan","year":"2008","journal-title":"J. Mach. Learn. Res."},{"key":"2023012512240524400_B12","first-page":"99","article-title":"Identifying the information structure of scientific abstracts: an investigation of three different schemes","volume-title":"Proceedings of the 2010 Workshop on Biomedical Natural Language Processing","author":"Guo","year":"2010"},{"key":"2023012512240524400_B13","article-title":"A comparison and user-based evaluation of models of textual information structure in the context of cancer risk assessment","volume":"12:69","author":"Guo","year":"2011","journal-title":"BMC Bioinform."},{"key":"2023012512240524400_B14","article-title":"Identifying sections in scientific abstracts using conditional random fields","author":"Hirohata","year":"2008","journal-title":"Proceedings of the IJCNLP 2008."},{"issue":"Suppl. 11","key":"2023012512240524400_B15","doi-asserted-by":"crossref","first-page":"S10","DOI":"10.1186\/1471-2105-9-S11-S10","article-title":"Recognizing speculative language in biomedical research articles: a linguistically motivated perspective","volume":"9","author":"Kilicoglu","year":"2008","journal-title":"BMC Bioinform."},{"key":"2023012512240524400_B16","first-page":"1","article-title":"Overview of BioNLP shared task 2011","volume-title":"Proceedings of BioNLP Shared Task 2011 Workshop","author":"Kim","year":"2011"},{"key":"2023012512240524400_B17","first-page":"345","article-title":"Faster parsing by supertagger adaptation","author":"Kummerfeld","year":"2010","journal-title":"Proceedings of the 48th Annual Meeting of the Association for Computational Linguistics, ACL '10"},{"key":"2023012512240524400_B18","first-page":"68","article-title":"A trainable document summarizer","volume-title":"Proceedings of the 18th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR '95","author":"Kupiec","year":"1995"},{"key":"2023012512240524400_B19","volume-title":"Guidelines for the annotation of general scientific concepts.","author":"Liakata","year":"2008"},{"key":"2023012512240524400_B20","volume-title":"The ART Corpus.","author":"Liakata","year":"2009"},{"key":"2023012512240524400_B21","doi-asserted-by":"crossref","first-page":"193","DOI":"10.3115\/1572364.1572391","article-title":"Semantic annotation of papers: interface & enrichment tool (SAPIENT)","volume-title":"Proceedings of BioNLP-09","author":"Liakata","year":"2009"},{"key":"2023012512240524400_B22","first-page":"2054","article-title":"Corpora for the conceptualization and zoning of scientific papers","volume-title":"Proceedings of the Seventh International Conference on Language Resources and Evaluation (LREC'10)","author":"Liakata","year":"2010"},{"key":"2023012512240524400_B23","first-page":"17","article-title":"The language of bioscience: facts, speculations, and statements in between","volume-title":"HLT-NAACL 2004 Workshop: BioLINK 2004, Linking Biological Literature, Ontologies and Databases","author":"Light","year":"2004"},{"key":"2023012512240524400_B24","doi-asserted-by":"crossref","first-page":"65","DOI":"10.3115\/1567619.1567631","article-title":"Generative content models for structural analysis of medical abstracts","volume-title":"Proceedings of the Workshop on Linking Natural Language Processing and Biology: Towards Deeper Biological Literature Analysis, BioNLP '06","author":"Lin","year":"2006"},{"key":"2023012512240524400_B25","first-page":"440","article-title":"Categorization of sentence types in medical abstracts","author":"McKnight","year":"2003","journal-title":"AMIA Annu Symp Proc"},{"key":"2023012512240524400_B26","first-page":"23","article-title":"Weakly supervised learning for hedge classification in scientific literature","volume-title":"45th Annual Meeting of the ACL","author":"Medlock","year":"2007"},{"key":"2023012512240524400_B27","doi-asserted-by":"crossref","first-page":"19","DOI":"10.3115\/1699750.1699754","article-title":"Accurate argumentative zoning with maximum entropy models","volume-title":"Proceedings of the 2009 Workshop on Text and Citation Analysis for Scholarly Digital Libraries, NLPIR4DL '09","author":"Merity","year":"2009"},{"key":"2023012512240524400_B28","doi-asserted-by":"crossref","first-page":"207","DOI":"10.1017\/S1351324901002728","article-title":"Applied morphological processing of english","volume":"7","author":"Minnen","year":"2001","journal-title":"Nat. Lang. Eng."},{"key":"2023012512240524400_B29","doi-asserted-by":"crossref","first-page":"468","DOI":"10.1016\/j.ijmedinf.2005.06.013","article-title":"Zone analysis in biology articles as a basis for information extraction","volume":"75","author":"Mizuta","year":"2006","journal-title":"Int. J. Med. Inform."},{"key":"2023012512240524400_B30","doi-asserted-by":"crossref","first-page":"52","DOI":"10.1145\/1089815.1089823","article-title":"A baseline feature set for learning rhetorical zones using full articles in the biomedical domain","volume":"7","author":"Mullen","year":"2005","journal-title":"SIGKDD Explor."},{"key":"2023012512240524400_B31","first-page":"2498","article-title":"Meta-knowledge annotation of bio-events","author":"Nawaz","year":"2010","journal-title":"LREC"},{"key":"2023012512240524400_B32","author":"Okazaki","year":"2007","journal-title":"CRFsuite: a fast implementation of Conditional Random Fields (CRFs)"},{"key":"2023012512240524400_B33","doi-asserted-by":"crossref","first-page":"852","DOI":"10.1016\/j.jbi.2008.12.004","article-title":"Porting a lexicalized-grammar parser to the biomedical domain","volume":"42","author":"Rimell","year":"2009","journal-title":"J. Biomed. Inform."},{"key":"2023012512240524400_B34","doi-asserted-by":"crossref","first-page":"195","DOI":"10.1016\/j.ijmedinf.2006.05.002","article-title":"Using argumentation to extract key sentences from biomedical abstracts","volume":"76","author":"Ruch","year":"2007","journal-title":"Int. J. Med. Inform."},{"key":"2023012512240524400_B35","doi-asserted-by":"crossref","first-page":"2086","DOI":"10.1093\/bioinformatics\/btn381","article-title":"Multi-dimensional classification of biomedical text: toward automated, practical provision of high-utility text to diverse users","volume":"24:18","author":"Shatkay","year":"2008","journal-title":"J. Bioinform."},{"key":"2023012512240524400_B36","first-page":"113","article-title":"Using localmaxs algorithm for the extraction of contiguous and non-contiguous multiword lexical units","volume-title":"Proceedings of the 9th Portuguese Conference on Artificial Intelligence: Progress in Artificial Intelligence, EPIA '99","author":"Silva","year":"1999"},{"key":"2023012512240524400_B37","doi-asserted-by":"crossref","first-page":"795","DOI":"10.1098\/rsif.2006.0134","article-title":"An ontology of scientific experiments","volume":"3","author":"Soldatova","year":"2006","journal-title":"J. Roy. Soc. Interf."},{"key":"2023012512240524400_B38","volume-title":"An ontology methodology and cisp-the proposed core information about scientific papers.","author":"Soldatova","year":"2007"},{"key":"2023012512240524400_B39","doi-asserted-by":"crossref","first-page":"409","DOI":"10.1162\/089120102762671936","article-title":"Summarizing scientific articles: experiments with relevance and rhetorical status","volume":"28","author":"Teufel","year":"2002","journal-title":"Comput. Linguist."},{"key":"2023012512240524400_B40","first-page":"110","article-title":"An annotation scheme for discourse-level argumentation in research articles","author":"Teufel","year":"1999","journal-title":"Proceedings of EACL"},{"key":"2023012512240524400_B41","doi-asserted-by":"crossref","first-page":"1493","DOI":"10.3115\/1699648.1699696","article-title":"Towards discipline-independent argumentative zoning: evidence from chemistry and computational linguistics","volume-title":"Proceedings of the 2009 Conference on Empirical Methods in Natural Language Processing: Volume 3 - Volume 3, EMNLP '09","author":"Teufel","year":"2009"},{"key":"2023012512240524400_B42","article-title":"Argumentative Zoning: Information Extraction from Scientific Text","volume-title":"PhD Thesis","author":"Teufel","year":"2000"},{"key":"2023012512240524400_B43","volume-title":"The Structure of Scientific Articles: Applications to Citation Indexing and Summarization.","author":"Teufel","year":"2010"},{"key":"2023012512240524400_B44","doi-asserted-by":"crossref","first-page":"393","DOI":"10.1186\/1471-2105-12-393","article-title":"Enriching a biomedical event corpus with meta-knowledge annotation","volume":"12","author":"Thompson","year":"2011","journal-title":"BMC Bioinform."},{"key":"2023012512240524400_B45","first-page":"134","article-title":"Hypothesis and evidence extraction from full-text scientific journal articles","volume-title":"Proceedings of BioNLP 2011 Workshop","author":"White","year":"2011"},{"key":"2023012512240524400_B46","doi-asserted-by":"crossref","first-page":"356","DOI":"10.1186\/1471-2105-7-356","article-title":"New directions in biomedical text annotations: definitions, guidelines and corpus construction","volume":"7","author":"Wilbur","year":"2006","journal-title":"BMC Bioinform."}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/28\/7\/991\/48880902\/bioinformatics_28_7_991.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/28\/7\/991\/48880902\/bioinformatics_28_7_991.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,1,25]],"date-time":"2023-01-25T15:54:56Z","timestamp":1674662096000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/28\/7\/991\/210210"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2012,2,8]]},"references-count":46,"journal-issue":{"issue":"7","published-print":{"date-parts":[[2012,4,1]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/bts071","relation":{},"ISSN":["1367-4811","1367-4803"],"issn-type":[{"value":"1367-4811","type":"electronic"},{"value":"1367-4803","type":"print"}],"subject":[],"published-other":{"date-parts":[[2012,4,1]]},"published":{"date-parts":[[2012,2,8]]}}}