{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,4]],"date-time":"2026-06-04T15:58:39Z","timestamp":1780588719774,"version":"3.54.1"},"reference-count":33,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2021,7,14]],"date-time":"2021-07-14T00:00:00Z","timestamp":1626220800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2021,7,14]],"date-time":"2021-07-14T00:00:00Z","timestamp":1626220800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100005416","name":"Norges forskningsr\u00e5d","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100005416","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J Biomed Semant"],"published-print":{"date-parts":[[2021,12]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:sec>\n                <jats:title>Background<\/jats:title>\n                <jats:p>The limited availability of clinical texts for Natural Language Processing purposes is hindering the progress of the field. This article investigates the use of synthetic data for the annotation and automated extraction of family history information from Norwegian clinical text. We make use of incrementally developed synthetic clinical text describing patients\u2019 family history relating to cases of cardiac disease and present a general methodology which integrates the synthetically produced clinical statements and annotation guideline development. The resulting synthetic corpus contains 477 sentences and 6030 tokens. In this work we experimentally assess the validity and applicability of the annotated synthetic corpus using machine learning techniques and furthermore evaluate the system trained on synthetic text on a corpus of real clinical text, consisting of de-identified records for patients with genetic heart disease.<\/jats:p>\n              <\/jats:sec><jats:sec>\n                <jats:title>Results<\/jats:title>\n                <jats:p>For entity recognition, an SVM trained on synthetic data had class weighted precision, recall and F<jats:sub>1<\/jats:sub>-scores of 0.83, 0.81 and 0.82, respectively. For relation extraction precision, recall and F<jats:sub>1<\/jats:sub>-scores were 0.74, 0.75 and 0.74.<\/jats:p>\n              <\/jats:sec><jats:sec>\n                <jats:title>Conclusions<\/jats:title>\n                <jats:p>A system for extraction of family history information developed on synthetic data generalizes well to real, clinical notes with a small loss of accuracy. The methodology outlined in this paper may be useful in other situations where limited availability of clinical text hinders NLP tasks. Both the annotation guidelines and the annotated synthetic corpus are made freely available and as such constitutes the first publicly available resource of Norwegian clinical text.<\/jats:p>\n              <\/jats:sec>","DOI":"10.1186\/s13326-021-00244-2","type":"journal-article","created":{"date-parts":[[2021,7,14]],"date-time":"2021-07-14T09:03:01Z","timestamp":1626253381000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":11,"title":["Synthetic data for annotation and extraction of family history information from clinical text"],"prefix":"10.1186","volume":"12","author":[{"given":"P\u00e5l H.","family":"Brekke","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4531-6733","authenticated-orcid":false,"given":"Taraka","family":"Rama","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ildik\u00f3","family":"Pil\u00e1n","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"\u00d8ystein","family":"Nytr\u00f8","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Lilja","family":"\u00d8vrelid","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2021,7,14]]},"reference":[{"issue":"Suppl","key":"244_CR1","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1016\/j.jbi.2015.10.007","volume":"58","author":"O Uzuner","year":"2015","unstructured":"Uzuner O, Stubbs A. Practical applications for natural language processing in clinical research: The 2014 i2b2\/uthealth shared tasks. J Biomed Inform. 2015; 58(Suppl):1.","journal-title":"J Biomed Inform"},{"key":"244_CR2","volume-title":"Proceedings of the LREC 2008 Workshop on Building and Evaluating Resources for Biomedical Text Mining","author":"A Roberts","year":"2008","unstructured":"Roberts A, Gaizauskas R, Hepple M, Demetriou G, Guo Y, Setzer A, Roberts I. Semantic annotation of clinical text: The clef corpus. In: Proceedings of the LREC 2008 Workshop on Building and Evaluating Resources for Biomedical Text Mining. Marrakech: European Language Resources Association (ELRA): 2008. p. 19\u201326."},{"key":"244_CR3","doi-asserted-by":"crossref","unstructured":"Dalianis H, Hassel M, Henriksson A, Skeppstedt M. Stockholm EPR Corpus: A Clinical Database Used to Improve Health Care. In: Proceedings of the Fourth Swedish Language Technology Conference: 2012. p. 17\u20138.","DOI":"10.4018\/978-1-60960-741-8.ch002"},{"issue":"1","key":"244_CR4","first-page":"1","volume":"9","author":"A N\u00e9v\u00e9ol","year":"2018","unstructured":"N\u00e9v\u00e9ol A, Dalianis H, Velupillai S, Savova G, Zweigenbaum P. Clinical natural language processing in languages other than English: opportunities and challenges. J Biotechnol Semant. 2018; 9(1):1\u201313.","journal-title":"J Biotechnol Semant"},{"key":"244_CR5","doi-asserted-by":"publisher","unstructured":"Velupillai S, Suominen H, Liakata M, Roberts A, Shah A, Morley K, Osborn D, Hayes J, Stewart R, Downs J, Chapman W, Dutta R. Using clinical natural language processing for health outcomes research: Overview and actionable suggestions for future advances. J Biomed Inform. 2018. https:\/\/doi.org\/10.1016\/j.jbi.2018.10.005.","DOI":"10.1016\/j.jbi.2018.10.005"},{"key":"244_CR6","volume-title":"Proceedings of the Eleventh International Conference on Language Resources and Evaluation","author":"C Lohr","year":"2018","unstructured":"Lohr C, Buechel S, Hahn U. Sharing copies of synthetic clinical corpora without physical distribution \u2013 a case study to get around IPRs and privacy constraints featuring the German JSYNCC corpus. In: Proceedings of the Eleventh International Conference on Language Resources and Evaluation. Miyazaki: European Language Resources Association (ELRA): 2018. p. 1259\u201366."},{"key":"244_CR7","unstructured":"Boag W, Naumann T, Szolovits P. Towards the creation of a large corpus of synthetically-identified clinical notes. CoRR. 2018; abs\/1803.02728. http:\/\/arxiv.org\/abs\/1803.02728."},{"key":"244_CR8","volume-title":"Proceedings of the NAACL HLT 2010 Second Louhi Workshop on Text and Data Mining of Health Documents","author":"H Allvin","year":"2010","unstructured":"Allvin H, Carlsson E, Dalianis H, Danielsson-Ojala R, Daudaravi\u010dius V, Hassel M, Kokkinakis D, Lundgren-Laine H, Nilsson G, Nytr\u00f8 \u00d8, et al. Characteristics and analysis of Finnish and Swedish clinical intensive care nursing narratives. In: Proceedings of the NAACL HLT 2010 Second Louhi Workshop on Text and Data Mining of Health Documents. Los Angeles: Association for Computational Linguistics: 2010. p. 53\u201360."},{"issue":"2","key":"244_CR9","doi-asserted-by":"publisher","first-page":"162","DOI":"10.5626\/JCSE.2008.2.2.162","volume":"2","author":"T R\u00f8st","year":"2008","unstructured":"R\u00f8st T, Huseth O, Nytr\u00f8 \u00d8, Grimsmo A. Lessons from developing an annotated corpus of patient histories. JCSE. 2008; 2(2):162\u201379.","journal-title":"JCSE"},{"key":"244_CR10","volume-title":"Proceedings of the 9th International Workshop on Health Text Mining and Information Analysis (LOUHI 2018)","author":"T Rama","year":"2018","unstructured":"Rama T, Brekke P, Nytr\u00f8 \u00d8, \u00d8vrelid L. Iterative development of family history annotation guidelines using a synthetic corpus of clinical text. In: Proceedings of the 9th International Workshop on Health Text Mining and Information Analysis (LOUHI 2018). Brussels: Association for Computational Linguistics: 2018."},{"issue":"5","key":"244_CR11","doi-asserted-by":"publisher","first-page":"424","DOI":"10.1007\/s10897-008-9169-9","volume":"17","author":"R Bennett","year":"2008","unstructured":"Bennett R, French K, Resta R, Doyle D. Standardized human pedigree nomenclature: update and assessment of the recommendations of the national society of genetic counselors. J Genet Couns. 2008; 17(5):424\u201333.","journal-title":"J Genet Couns"},{"key":"244_CR12","unstructured":"Elliott P, Anastasakis A, Borger M, Borggrefe M, Cecchi F, Charron P, Hagege A, Lafont A, Limongelli G, Mahrholdt H, McKenna W, Mogensen J, Nihoyannopoulos P, Nistri S, Pieper P, Pieske B, Rapezzi C, Rutten F, Tillmanns C, Watkins H, Contributor A, O\u2019Mahony C, for Practice Guidelines (CPG) EC, Zamorano J, Achenbach S, Baumgartner H, Bax J, Bueno H, Dean V, Deaton C, \u00c7etin Erol, Fagard R, Ferrari R, Hasdai D, Hoes A, Kirchhof P, Knuuti J, Kolh P, Lancellotti P, Linhart A, Nihoyannopoulos P, Piepoli M, Ponikowski P, Sirnes P, Tamargo J, Tendera M, Torbicki A, Wijns W, Windecker S, Reviewers D, Hasdai D, Ponikowski P, Achenbach S, Alfonso F, Basso C, Cardim N, Gimeno J, Heymans S, Holm P, Keren A, Kirchhof P, Kolh P, Lionis C, Muneretto C, Priori S, Salvador M, Wolpert C, Zamorano J, Frick M, Aliyev F, Komissarova S, Mairesse G, Smaji\u0107 E, Velchev V, Antoniades L, Linhart A, Bundgaard H, Heli\u00f6 T, Leenhardt A, Katus H, Efthymiadis G, Sepp R, Gunnarsson G, Carasso S, Kerimkulova A, Kamzola G, Skouri H, Eldirsi G, Kavoliuniene A, Felice T, Michels M, Haugaa K, Lenarczyk R, Brito D, Apetrei E, Bokheria L, Lovic D, Hatala R, Pav\u00eda P, Eriksson M, Noble S, Srbinovska E, \u00d6zdemir M, Nesukay E, Sekhri N. 2014 ESC guidelines on diagnosis and management of hypertrophic cardiomyopathy: the task force for the diagnosis and management of hypertrophic cardiomyopathy of the european society of cardiology (ESC). Eur Heart J. 2014; 35(39)."},{"issue":"2","key":"244_CR13","doi-asserted-by":"publisher","first-page":"381","DOI":"10.1007\/s10897-018-0235-7","volume":"27","author":"B Welch","year":"2018","unstructured":"Welch B, Wiley K, Pflieger L, Achiangia R, Baker K, Hughes-Halbert C, Morrison H, Schiffman J, Doerr M. Review and comparison of electronic patient-facing family health history tools. J Genet Couns. 2018; 27(2):381\u201391. https:\/\/doi.org\/10.1007\/s10897-018-0235-7.","journal-title":"J Genet Couns"},{"key":"244_CR14","unstructured":"Stevens R, Matentzoglu N, Sattler U, Stevens M. Informal Proceedings of the 3rd International Workshop on OWL Reasoner Evaluation (ORE 2014) Co-located with the Vienna Summer of Logic (VSL 2014), Vienna, Austria, July 13, 2014 In: Bail S, Glimm B, Jim\u00e9nez-Ruiz E, Matentzoglu N, Parsia B, Steigmiller A, editors. CEUR Workshop Proceedings. CEUR-WS.org: 2014. p. 71\u20136. http:\/\/ceur-ws.org\/Vol-1207\/paper_11.pdf."},{"issue":"1","key":"244_CR15","doi-asserted-by":"publisher","first-page":"16","DOI":"10.1375\/twin.8.1.16","volume":"8","author":"T Hiekkalinna","year":"2005","unstructured":"Hiekkalinna T, Terwilliger J, Sammalisto S, Peltonen L, Perola M. AUTOGSCAN: Powerful tools for automated genome-wide linkage and linkage disequilibrium analysis. Twin Res Hum Genet. 2005; 8(1):16\u201321. https:\/\/doi.org\/10.1375\/twin.8.1.16.","journal-title":"Twin Res Hum Genet"},{"key":"244_CR16","unstructured":"Bill R, Pakhomov S, Chen E, Winden T, Carter E, Melton G. Automated extraction of family history information from clinical notes. In: AMIA Annual Symposium Proceedings. American Medical Informatics Association: 2014. p. 1709."},{"key":"244_CR17","unstructured":"Polubriaginof F, Tatonetti N, Vawdrey D. An assessment of family history information captured in an electronic health record. In: AMIA Annual Symposium Proceedings. American Medical Informatics Association: 2015. p. 2035."},{"key":"244_CR18","unstructured":"Goryachev S, Kim H, Zeng-Treitler Q. Identification and extraction of family history information from clinical reports. In: AMIA Annual Symposium Proceedings. American Medical Informatics Association: 2008. p. 247."},{"key":"244_CR19","unstructured":"Friedlin J, McDonald C. Using a natural language processing system to extract and code family history data from admission reports. In: AMIA Annual Symposium Proceedings. American Medical Informatics Association: 2006. p. 925."},{"issue":"5","key":"244_CR20","doi-asserted-by":"publisher","first-page":"4","DOI":"10.1186\/2041-1480-2-S5-S4","volume":"2","author":"A Abacha","year":"2011","unstructured":"Abacha A, Zweigenbaum P. Automatic extraction of semantic relations between medical entities: a rule based approach. J Biomed Semant. 2011; 2(5):4.","journal-title":"J Biomed Semant"},{"key":"244_CR21","volume-title":"Proceedings of the Workshop on Current Trends in Biomedical Natural Language Processing","author":"A Roberts","year":"2008","unstructured":"Roberts A, Gaizauskas R, Hepple M. Extracting clinical relationships from patient narratives. In: Proceedings of the Workshop on Current Trends in Biomedical Natural Language Processing. Columbus: Association for Computational Linguistics: 2008. p. 10\u20138."},{"key":"244_CR22","volume-title":"Proceedings of the International Conference Recent Advances in Natural Language Processing 2011","author":"A-L Minard","year":"2011","unstructured":"Minard A-L, Ligozat A-L, Grau B. Multi-class SVM for relation extraction from clinical reports. In: Proceedings of the International Conference Recent Advances in Natural Language Processing 2011. Hissar: Association for Computational Linguistics: 2011. p. 604\u20139."},{"key":"244_CR23","doi-asserted-by":"publisher","unstructured":"Hong G. Relation extraction using Support Vector Machine. In: Second International Joint Conference on Natural Language Processing: Full Papers: 2005. p. 366\u201337. https:\/\/doi.org\/10.1007\/11562214_33.","DOI":"10.1007\/11562214_33"},{"key":"244_CR24","volume-title":"Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP)","author":"M Miwa","year":"2014","unstructured":"Miwa M, Sasaki Y. Modeling joint entity and relation extraction with table representation. In: Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP). Doha: Association for Computational Linguistics: 2014. p. 1858\u201369."},{"key":"244_CR25","volume-title":"BioCreative\/OHNLP 2018 Workshop","author":"S Liu","year":"2018","unstructured":"Liu S, Rastegar-Mojarad M, Wang Y, Wang L, Shen F, Fu S, Liu H. Overview of the BioCreative\/OHNLP 2018 family history extraction task. In: BioCreative\/OHNLP 2018 Workshop. Minneapolis: Association for Computational Linguistics: 2018."},{"key":"244_CR26","volume-title":"Proceedings of the Demonstrations Session at EACL 2012","author":"P Stenetorp","year":"2012","unstructured":"Stenetorp P, Pyysalo S, Topi\u0107 G, Ohta T, Ananiadou S, Tsujii J. brat: a web-based tool for nlp-assisted text annotation. In: Proceedings of the Demonstrations Session at EACL 2012. Avignon: Association for Computational Linguistics: 2012. p. 102\u20137."},{"key":"244_CR27","unstructured":"Morante R, Daelemans W. ConanDoyle-neg: Annotation of negation cues and their scope in Conan Doyle stories. In: Proceedings of the Eighth International Conference on Language Resources and Evaluation (LREC-2012). European Language Resources Association (ELRA): 2012. http:\/\/www.aclweb.org\/anthology\/L12-1077."},{"key":"244_CR28","volume-title":"Instruction manual for the annotation of temporal expressions. Technical report","author":"L Ferro","year":"2002","unstructured":"Ferro L, Gerber L, Mani I, Sundheim B, Wilson G. Instruction manual for the annotation of temporal expressions. Technical report. Washington C3 Center, McLean, Virginia: MITRE; 2002."},{"key":"244_CR29","unstructured":"Saur\u00ed R, Littman J, Knippen B, Gaizauskas R, Setzer A, Pustejovsky J. TimeML annotation guidelines version 1.2. 1. Technical report. LDC. 2006."},{"key":"244_CR30","volume-title":"ICML \u201901","author":"J Lafferty","year":"2001","unstructured":"Lafferty J, McCallum A, Pereira F. Conditional random fields: Probabilistic models for segmenting and labeling sequence data. In: ICML \u201901. San Francisco, CA, USA: Morgan Kaufmann Publishers Inc.: 2001. p. 282\u20139."},{"key":"244_CR31","volume-title":"Proceedings of the CoNLL 2017 Shared Task: Multilingual Parsing from Raw Text to Universal Dependencies","author":"D Zeman","year":"2017","unstructured":"Zeman D, Popel M, Straka M, Hajic J, Nivre J, Ginter F, Luotolahti J, Pyysalo S, Petrov S, Potthast M, Tyers F, Badmaeva E, Gokirmak M, Nedoluzhko A, Cinkova S, Hajic jr. J, Hlavacova J, Kettnerov\u00e1 V, Uresova Z, Kanerva J, Ojala S, Missil\u00e4 A, Manning C, Schuster S, Reddy S, Taji D, Habash N, Leung H, de Marneffe M-C, Sanguinetti M, Simi M, Kanayama H, dePaiva V, Droganova K, Mart\u00ednez Alonso H, \u00c7\u00f6ltekin c, Sulubacak U, Uszkoreit H, Macketanz V, Burchardt A, Harris K, Marheinecke K, Rehm G, Kayadelen T, Attia M, Elkahky A, Yu Z, Pitler E, Lertpradit S, Mandl M, Kirchner J, Alcalde H, Strnadov\u00e1 J, Banerjee E, Manurung R, Stella A, Shimada A, Kwak S, Mendonca G, Lando T, Nitisaroj R, Li J. CoNLL 2017 Shared Task: Multilingual Parsing from Raw Text to Universal Dependencies. In: Proceedings of the CoNLL 2017 Shared Task: Multilingual Parsing from Raw Text to Universal Dependencies. Vancouver: Association for Computational Linguistics: 2017. p. 1\u201319."},{"key":"244_CR32","volume-title":"Proceedings of the International Conference on Language Resources and Evaluation (LREC)","author":"L \u00d8vrelid","year":"2016","unstructured":"\u00d8vrelid L, Hohle P. Universal Dependencies for Norwegian. In: Proceedings of the International Conference on Language Resources and Evaluation (LREC). Portoro\u017e: European Language Resources Association (ELRA): 2016."},{"key":"244_CR33","volume-title":"Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC\u201916)","author":"M Straka","year":"2016","unstructured":"Straka M, Hajic J, Strakov\u00e1 J. UDPipe: trainable pipeline for processing CoNLL-U files performing tokenization, morphological analysis, POS tagging and parsing. In: Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC\u201916). Portoro\u017e: European Language Resources Association (ELRA): 2016. p. 4290\u20137."}],"container-title":["Journal of Biomedical Semantics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s13326-021-00244-2.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1186\/s13326-021-00244-2\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s13326-021-00244-2.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2021,7,14]],"date-time":"2021-07-14T09:10:35Z","timestamp":1626253835000},"score":1,"resource":{"primary":{"URL":"https:\/\/jbiomedsem.biomedcentral.com\/articles\/10.1186\/s13326-021-00244-2"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,7,14]]},"references-count":33,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2021,12]]}},"alternative-id":["244"],"URL":"https:\/\/doi.org\/10.1186\/s13326-021-00244-2","relation":{},"ISSN":["2041-1480"],"issn-type":[{"value":"2041-1480","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,7,14]]},"assertion":[{"value":"11 May 2019","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"26 May 2021","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"14 July 2021","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The project was approved by the regional board for medical research ethics (REK 2017\/1931) and individual consent was not required.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethics approval and consent to participate"}},{"value":"Not applicable.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent for publication"}},{"value":"The authors declare that they have no competing interests.","order":4,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"11"}}