{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,9]],"date-time":"2026-07-09T21:01:14Z","timestamp":1783630874832,"version":"3.55.0"},"reference-count":29,"publisher":"Springer Science and Business Media LLC","issue":"S1","license":[{"start":{"date-parts":[[2019,11,1]],"date-time":"2019-11-01T00:00:00Z","timestamp":1572566400000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2019,11,12]],"date-time":"2019-11-12T00:00:00Z","timestamp":1573516800000},"content-version":"vor","delay-in-days":11,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J Biomed Semant"],"published-print":{"date-parts":[[2019,11]]},"abstract":"<jats:title>Abstract<\/jats:title>\n              <jats:sec>\n                <jats:title>Background<\/jats:title>\n                <jats:p>With the improvements to text mining technology and the availability of large unstructured Electronic Healthcare Records (EHR) datasets, it is now possible to extract structured information from raw text contained within EHR at reasonably high accuracy. We describe a text mining system for classifying radiologists\u2019 reports of CT and MRI brain scans, assigning labels indicating occurrence and type of stroke, as well as other observations. Our system, the Edinburgh Information Extraction for Radiology reports (EdIE-R) system, which we describe here, was developed and tested on a collection of radiology reports.The work reported in this paper is based on 1168 radiology reports from the Edinburgh Stroke Study (ESS), a hospital-based register of stroke and transient ischaemic attack patients. We manually created annotations for this data in parallel with developing the rule-based EdIE-R system to identify phenotype information related to stroke in radiology reports. This process was iterative and domain expert feedback was considered at each iteration to adapt and tune the EdIE-R text mining system which identifies entities, negation and relations between entities in each report and determines report-level labels (phenotypes).<\/jats:p>\n              <\/jats:sec>\n              <jats:sec>\n                <jats:title>Results<\/jats:title>\n                <jats:p>The inter-annotator agreement (IAA) for all types of annotations is high at 96.96 for entities, 96.46 for negation, 95.84 for relations and 94.02 for labels. The equivalent system scores on the blind test set are equally high at 95.49 for entities, 94.41 for negation, 98.27 for relations and 96.39 for labels for the first annotator and 96.86, 96.01, 96.53 and 92.61, respectively for the second annotator.<\/jats:p>\n              <\/jats:sec>\n              <jats:sec>\n                <jats:title>Conclusion<\/jats:title>\n                <jats:p>Automated reading of such EHR data at such high levels of accuracies opens up avenues for population health monitoring and audit, and can provide a resource for epidemiological studies. We are in the process of validating EdIE-R in separate larger cohorts in NHS England and Scotland. The manually annotated ESS corpus will be available for research purposes on application.<\/jats:p>\n              <\/jats:sec>","DOI":"10.1186\/s13326-019-0211-7","type":"journal-article","created":{"date-parts":[[2019,11,12]],"date-time":"2019-11-12T01:02:36Z","timestamp":1573520556000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":35,"title":["Text mining brain imaging reports"],"prefix":"10.1186","volume":"10","author":[{"given":"Beatrice","family":"Alex","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Claire","family":"Grover","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Richard","family":"Tobin","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Cathie","family":"Sudlow","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Grant","family":"Mair","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"William","family":"Whiteley","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2019,11,12]]},"reference":[{"key":"211_CR1","unstructured":"EdIE-R project page. https:\/\/www.ltg.ed.ac.uk\/software\/edie-r. Accessed 10 July 2019."},{"key":"211_CR2","volume-title":"Proceedings of EACL 2012","author":"P Stenetorp","year":"2012","unstructured":"Stenetorp P, Pyysalo S, Topi\u0107 G, Ohta T, Ananiadou S, Tsujii J. BRAT: A Web-based Tool for NLP-assisted Text Annotation. In: Proceedings of EACL 2012. Stroudsburg: Association for Computational Linguistics: 2012. p. 102\u20137."},{"key":"211_CR3","doi-asserted-by":"crossref","unstructured":"Tjong Kim Sang EF, De Meulder F. Introduction to the CoNLL-2003 shared task: language-independent named entity recognition. In: Proceedings of CoNLL-2003: 2003. p. 142\u20137. https:\/\/doi.org\/10.3115\/1119176.1119195.","DOI":"10.3115\/1119176.1119195"},{"key":"211_CR4","doi-asserted-by":"crossref","unstructured":"Finkel JR, Grenager T, Manning C. Incorporating non-local information into information extraction systems by Gibbs sampling. In: Proceedings of the 43rd Annual Meeting on Association for Computational Linguistics: 2005. p. 363\u201370. https:\/\/doi.org\/10.3115\/1219840.1219885.","DOI":"10.3115\/1219840.1219885"},{"key":"211_CR5","volume-title":"Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics","author":"H Cunningham","year":"2002","unstructured":"Cunningham H, Maynard D, Bontcheva K, Tablan V. GATE: A framework and graphical development environment for robust NLP tools and applications. In: Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics. Philadelphia: Association for Computational Linguistics: 2002. p. 168\u201375."},{"issue":"1","key":"211_CR6","doi-asserted-by":"publisher","first-page":"21","DOI":"10.1017\/S1351324911000106","volume":"18","author":"B Hachey","year":"2011","unstructured":"Hachey B, Grover C, Tobin R. Datasets for generic relation extraction. J Nat Lang Eng. 2011; 18(1):21\u201359.","journal-title":"J Nat Lang Eng"},{"key":"211_CR7","unstructured":"BioCreative. http:\/\/www.biocreative.org. Accessed 10 July 2019."},{"key":"211_CR8","unstructured":"BioNLP. http:\/\/2016.bionlp-st.org. Accessed 10 July 2019."},{"key":"211_CR9","doi-asserted-by":"crossref","unstructured":"Alex B, Haddow B, Grover C. Recognising nested named entities in biomedical text. In: Proceedings of BioNLP 2007: 2007. p. 65\u201372. https:\/\/doi.org\/10.3115\/1572392.1572404.","DOI":"10.3115\/1572392.1572404"},{"key":"211_CR10","volume-title":"Proceedings of BioCreative II Workshop 2007","author":"C Grover","year":"2007","unstructured":"Grover C, Haddow B, Klein E, Matthews M, Nielsen LA, Tobin R, Wang X. Adapting a relation extraction pipeline for the BioCreative II task. In: Proceedings of BioCreative II Workshop 2007. Madrid: CNIO Centro Nacional de Investigaciones Oncologicas: 2007."},{"key":"211_CR11","unstructured":"LOUHI\u201917. https:\/\/sites.google.com\/site\/louhi17\/home. Accessed 10 July 2019."},{"key":"211_CR12","unstructured":"LOUHI\u201918. https:\/\/louhi2018.fbk.eu. Accessed 10 July 2019."},{"issue":"Suppl. 1","key":"211_CR13","first-page":"128","volume":"47","author":"SM Meystre","year":"2008","unstructured":"Meystre SM, Savova GK, Kipper-Schuler KC, Hurdle JF. Extracting information from textual documents in the electronic health record: a review of recent research. Yearb Med Inform. 2008; 47(Suppl. 1):128\u201344.","journal-title":"Yearb Med Inform"},{"issue":"5","key":"211_CR14","doi-asserted-by":"publisher","first-page":"301","DOI":"10.1006\/jbin.2001.1029","volume":"34","author":"WW Chapman","year":"2001","unstructured":"Chapman WW, Bridewell W, Hanbury P, Cooper GF, Buchanan BG. A simple algorithm for identifying negated findings and diseases in discharge summaries. J Biomed Inform. 2001; 34(5):301\u201310.","journal-title":"J Biomed Inform"},{"issue":"2","key":"211_CR15","doi-asserted-by":"publisher","first-page":"329","DOI":"10.1148\/radiol.16142770","volume":"279","author":"E Pons","year":"2016","unstructured":"Pons E, Braun LMM, Hunink MGM, Kors JA. Natural Language Processing in Radiology: A Systematic Review. Radiology. 2016; 279(2):329\u201343. https:\/\/doi.org\/10.1148\/radiol.16142770.","journal-title":"Radiology"},{"key":"211_CR16","doi-asserted-by":"publisher","first-page":"29","DOI":"10.1016\/j.artmed.2015.09.007","volume":"66","author":"S Hassanpour","year":"2016","unstructured":"Hassanpour S, Langlotz CP. Information extraction from multi-institutional radiology reports. Artif Intell Med. 2016; 66:29\u201339.","journal-title":"Artif Intell Med"},{"key":"211_CR17","doi-asserted-by":"crossref","unstructured":"Cornegruta S, Bakewell R, Withey S, Montana G. Modelling radiological language with bidirectional long short-term memory networks. In: Proceedings of the 7th International Workshop on Health Text Mining and Information Analysis: 2016. p. 17\u201327. https:\/\/doi.org\/10.18653\/v1\/w16-6103.","DOI":"10.18653\/v1\/W16-6103"},{"issue":"6","key":"211_CR18","doi-asserted-by":"publisher","first-page":"1595","DOI":"10.1148\/rg.266065168","volume":"26","author":"CP Langlotz","year":"2006","unstructured":"Langlotz CP. Radlex: a new method for indexing online educational materials. Radiographics. 2006; 26(6):1595\u20137.","journal-title":"Radiographics"},{"key":"211_CR19","unstructured":"United States National Library of Medicine NLM. Medical Subject Headings 2016. https:\/\/www.nlm.nih.gov\/mesh\/meshhome.html. Accessed 10 July 2019."},{"key":"211_CR20","volume-title":"Proceedings of the Ninth International Workshop on Health Text Mining and Information Analysis","author":"Y Zhang","year":"2018","unstructured":"Zhang Y, Ding DY, Qian T, Manning CD, Langlotz CP. Learning to summarize radiology findings. In: Proceedings of the Ninth International Workshop on Health Text Mining and Information Analysis. Brussels: Association for Computational Linguistics: 2018. p. 204\u201313. http:\/\/aclweb.org\/anthology\/W18-5623."},{"issue":"8","key":"211_CR21","doi-asserted-by":"publisher","first-page":"843","DOI":"10.1002\/pds.1981","volume":"19","author":"R Flynn","year":"2010","unstructured":"Flynn R, Macdonald T, Schembri N, Murray G, Doney A. Automated data capture from free text radiology reports to enhance accuracy of hospital inpatient stroke codes. Pharmacoepidemiol Drug Saf. 2010; 19(8):843\u20137.","journal-title":"Pharmacoepidemiol Drug Saf"},{"issue":"4","key":"211_CR22","first-page":"281","volume":"101","author":"C Jackson","year":"2008","unstructured":"Jackson C, Crossland L, Dennis M, Wardlaw J, Sudlow C. Assessing the impact of the requirement for explicit consent in a hospital-based stroke study. QJM Mon J Assoc Phys. 2008; 101(4):281\u20139.","journal-title":"QJM Mon J Assoc Phys"},{"key":"211_CR23","doi-asserted-by":"crossref","unstructured":"Grover C, Matthews M, Tobin R. Tools to address the interdependence between tokenisation and standoff annotation. In: Proceedings of NLPXML 2006: 2006. p. 19\u201326. https:\/\/doi.org\/10.3115\/1621034.1621038.","DOI":"10.3115\/1621034.1621038"},{"issue":"1","key":"211_CR24","doi-asserted-by":"publisher","first-page":"15","DOI":"10.3366\/ijhac.2015.0136","volume":"9","author":"B Alex","year":"2015","unstructured":"Alex B, Byrne K, Grover C, Tobin R. Adapting the Edinburgh Geoparser for historical georeferencing. Int J Humanit Arts Comput. 2015; 9(1):15\u201335.","journal-title":"Int J Humanit Arts Comput"},{"key":"211_CR25","doi-asserted-by":"crossref","unstructured":"Curran J, Clark S. Language independent NER using a maximum entropy tagger. In: Proceedings of CoNLL 2003: 2003. p. 164\u20137. https:\/\/doi.org\/10.3115\/1119176.1119200.","DOI":"10.3115\/1119176.1119200"},{"issue":"Suppl. 1","key":"211_CR26","doi-asserted-by":"publisher","first-page":"180","DOI":"10.1093\/bioinformatics\/btg1023","volume":"19","author":"J-D Kim","year":"2003","unstructured":"Kim J-D, Ohta T, Teteisi Y, Tsujii J. GENIA corpus - a semantically annotated corpus for bio-textmining. Bioinformatics. 2003; 19(Suppl. 1):180\u20132.","journal-title":"Bioinformatics"},{"key":"211_CR27","doi-asserted-by":"crossref","unstructured":"Minnen G, Carroll J, Pearce D. Robust, applied morphological generation. In: Proceedings of INLG 2000: 2000. p. 201\u20138. https:\/\/doi.org\/10.3115\/1118253.1118281.","DOI":"10.3115\/1118253.1118281"},{"key":"211_CR28","volume-title":"Proceedings of the Fifth International Conference on Language Resources and Evaluation","author":"C Grover","year":"2006","unstructured":"Grover C, Tobin R. Rule-based chunking and reusability. In: Proceedings of the Fifth International Conference on Language Resources and Evaluation. Genoa: European Language Resources Association (ELRA): 2006. p. 873\u20138. http:\/\/www.lrec-conf.org\/proceedings\/lrec2006\/pdf\/457_pdf.pdf."},{"key":"211_CR29","unstructured":"Grover C, Tobin R, Alex B, Sudlow C, Mair G, Whiteley W. Text Mining Brain Imaging Reports. In: HealTAC-2018. Manchester: 2018."}],"container-title":["Journal of Biomedical Semantics"],"original-title":[],"language":"en","link":[{"URL":"http:\/\/link.springer.com\/content\/pdf\/10.1186\/s13326-019-0211-7.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/link.springer.com\/article\/10.1186\/s13326-019-0211-7\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/link.springer.com\/content\/pdf\/10.1186\/s13326-019-0211-7.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2020,11,11]],"date-time":"2020-11-11T00:38:13Z","timestamp":1605055093000},"score":1,"resource":{"primary":{"URL":"https:\/\/jbiomedsem.biomedcentral.com\/articles\/10.1186\/s13326-019-0211-7"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019,11]]},"references-count":29,"journal-issue":{"issue":"S1","published-print":{"date-parts":[[2019,11]]}},"alternative-id":["211"],"URL":"https:\/\/doi.org\/10.1186\/s13326-019-0211-7","relation":{},"ISSN":["2041-1480"],"issn-type":[{"value":"2041-1480","type":"electronic"}],"subject":[],"published":{"date-parts":[[2019,11]]},"assertion":[{"value":"12 November 2019","order":1,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"The Edinburgh Stroke Study received ethical approval from the Lothian Research Ethics Committee (LREC\/2001\/4\/46). This is a patient-consented dataset. We also received permission from the NHS Tayside Caldicott Guardian to use the anonymised brain imaging reports for this work.","order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethics approval and consent to participate"}},{"value":"The authors declare that they have no competing interests","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"23"}}