{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,15]],"date-time":"2026-07-15T08:48:11Z","timestamp":1784105291814,"version":"3.55.0"},"reference-count":63,"publisher":"Oxford University Press (OUP)","issue":"3","license":[{"start":{"date-parts":[[2016,10,4]],"date-time":"2016-10-04T00:00:00Z","timestamp":1475539200000},"content-version":"vor","delay-in-days":575,"URL":"http:\/\/creativecommons.org\/licenses\/by-nc\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2015,5,1]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Objective Social media is becoming increasingly popular as a platform for sharing personal health-related information. This information can be utilized for public health monitoring tasks, particularly for pharmacovigilance, via the use of natural language processing (NLP) techniques. However, the language in social media is highly informal, and user-expressed medical concepts are often nontechnical, descriptive, and challenging to extract. There has been limited progress in addressing these challenges, and thus far, advanced machine learning-based NLP techniques have been underutilized. Our objective is to design a machine learning-based approach to extract mentions of adverse drug reactions (ADRs) from highly informal text in social media.<\/jats:p><jats:p>Methods We introduce ADRMine, a machine learning-based concept extraction system that uses conditional random fields (CRFs). ADRMine utilizes a variety of features, including a novel feature for modeling words\u2019 semantic similarities. The similarities are modeled by clustering words based on unsupervised, pretrained word representation vectors (embeddings) generated from unlabeled user posts in social media using a deep learning technique.<\/jats:p><jats:p>Results ADRMine outperforms several strong baseline systems in the ADR extraction task by achieving an F-measure of 0.82. Feature analysis demonstrates that the proposed word cluster features significantly improve extraction performance.<\/jats:p><jats:p>Conclusion It is possible to extract complex medical concepts, with relatively high performance, from informal, user-generated content. Our approach is particularly scalable, suitable for social media mining, as it relies on large volumes of unlabeled data, thus diminishing the need for large, annotated training data sets.<\/jats:p>","DOI":"10.1093\/jamia\/ocu041","type":"journal-article","created":{"date-parts":[[2015,3,10]],"date-time":"2015-03-10T01:52:20Z","timestamp":1425952340000},"page":"671-681","source":"Crossref","is-referenced-by-count":417,"title":["Pharmacovigilance from social media: mining adverse drug reaction mentions using sequence labeling with word embedding cluster features"],"prefix":"10.1093","volume":"22","author":[{"given":"Azadeh","family":"Nikfarjam","sequence":"first","affiliation":[{"name":"Department of Biomedical Informatics, Arizona State University, Scottsdale, AZ, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Abeed","family":"Sarker","sequence":"additional","affiliation":[{"name":"Department of Biomedical Informatics, Arizona State University, Scottsdale, AZ, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Karen","family":"O\u2019Connor","sequence":"additional","affiliation":[{"name":"Department of Biomedical Informatics, Arizona State University, Scottsdale, AZ, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Rachel","family":"Ginn","sequence":"additional","affiliation":[{"name":"Department of Biomedical Informatics, Arizona State University, Scottsdale, AZ, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Graciela","family":"Gonzalez","sequence":"additional","affiliation":[{"name":"Department of Biomedical Informatics, Arizona State University, Scottsdale, AZ, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"286","published-online":{"date-parts":[[2015,3,9]]},"reference":[{"key":"2020110613025931100_ocu041-B1","doi-asserted-by":"crossref","first-page":"15","DOI":"10.1136\/bmj.329.7456.15","article-title":"Adverse drug reactions as cause of admission to hospital: prospective analysis of 18\u2009820 patients","volume":"329","author":"Pirmohamed","year":"2004","journal-title":"BMJ"},{"key":"2020110613025931100_ocu041-B2","doi-asserted-by":"crossref","first-page":"S73","DOI":"10.4103\/0976-500X.120957","article-title":"Clinical and economic burden of adverse drug reactions","volume":"4","author":"Sultana","year":"2013","journal-title":"J Pharmacol Pharmacother."},{"key":"2020110613025931100_ocu041-B3","doi-asserted-by":"crossref","first-page":"1067","DOI":"10.2165\/11316680-000000000-00000","article-title":"Consumer reporting of adverse drug reactions: a retrospective analysis of the Danish adverse drug reaction database from 2004 to 2006","volume":"32","author":"Aagaard","year":"2009","journal-title":"Drug Saf."},{"key":"2020110613025931100_ocu041-B4","article-title":"Evaluation of patient reporting of adverse drug reactions to the UK \u201cYellow Card Scheme\u201d: literature review, descriptive and qualitative analyses, and questionnaire surveys","author":"Avery","year":"2011","journal-title":"Southampton: NIHR HTA"},{"key":"2020110613025931100_ocu041-B5","doi-asserted-by":"crossref","first-page":"1193","DOI":"10.1007\/s00228-007-0375-4","article-title":"Evaluation of patients\u2019 experiences with antidepressants reported by means of a medicine reporting system","volume":"63","author":"Van Geffen","year":"2007","journal-title":"Eur J Clin Pharmacol."},{"key":"2020110613025931100_ocu041-B6","doi-asserted-by":"crossref","first-page":"16","DOI":"10.1186\/1472-6904-11-16","article-title":"What can we learn from consumer reports on psychiatric adverse drug reactions with antidepressant medication? Experiences from reports to a consumer association","volume":"11","author":"Vilhelmsson","year":"2011","journal-title":"BMC Clin Pharmacol."},{"key":"2020110613025931100_ocu041-B7","doi-asserted-by":"crossref","first-page":"385","DOI":"10.2165\/00002018-200629050-00003","article-title":"Under-reporting of adverse drug reactions","volume":"29","author":"Hazell","year":"2006","journal-title":"Drug Saf."},{"key":"2020110613025931100_ocu041-B8","article-title":"Mining Twitter for adverse drug reaction mentions: a corpus and classification benchmark","volume-title":"proceedings of the Fourth Workshop on Building and Evaluating Resources for Health and Biomedical Text Processing (BioTxtM)","author":"Ginn","year":"2014"},{"key":"2020110613025931100_ocu041-B9","article-title":"Pharmacovigilance on Twitter? Mining Tweets for adverse drug reactions","volume-title":"American Medical Informatics Association (AMIA) Annual Symposium","author":"O\u2019Connor","year":", 2014"},{"key":"2020110613025931100_ocu041-B10"},{"key":"2020110613025931100_ocu041-B11","first-page":"117","article-title":"Towards internet-age pharmacovigilance: extracting adverse drug reactions from user posts to health-related social networks","volume-title":"Proceedings of the 2010 Workshop on Biomedical Natural Language Processing","author":"Leaman","year":", 2010"},{"key":"2020110613025931100_ocu041-B12","doi-asserted-by":"crossref","first-page":"816","DOI":"10.1007\/978-3-642-36973-5_92","article-title":"ADRTrace: detecting expected and unexpected adverse drug reactions from user reviews on social media sites","volume":"7814 LNCS","author":"Yates","year":"2013","journal-title":"Adv Inf Retr."},{"key":"2020110613025931100_ocu041-B13","article-title":"Detecting signals of adverse drug reactions from health consumer contributed content in social media","volume-title":"Proceedings of ACM SIGKDD Workshop on Health Informatics","author":"Yang","year":", , 2012"},{"key":"2020110613025931100_ocu041-B14","doi-asserted-by":"crossref","first-page":"989","DOI":"10.1016\/j.jbi.2011.07.005","article-title":"Identifying potential adverse effects using the web: a new approach to medical hypothesis generation","volume":"44","author":"Benton","year":"2011","journal-title":"J Biomed Inform."},{"key":"2020110613025931100_ocu041-B15","first-page":"384","article-title":"Word representations: a simple and general method for semi-supervised learning","volume-title":"Proceedings of the 48th Annual Meeting of the Association for Computational Linguistics","author":"Turian","year":"2010"},{"key":"2020110613025931100_ocu041-B16","first-page":"197","article-title":"Deep Learning: Methods and Applications","volume-title":"Foundations and Trends in Signal Processing","author":"Deng","year":"2014"},{"key":"2020110613025931100_ocu041-B17","first-page":"2493","article-title":"Natural language processing (almost) from scratch","volume":"1","author":"Collobert","year":"2011","journal-title":"J Mach Learn Res."},{"key":"2020110613025931100_ocu041-B18","first-page":"739","article-title":"Extraction of adverse drug effects from clinical records","volume":"160","author":"Aramaki","year":"2010","journal-title":"Stud Heal Technol Inf."},{"key":"2020110613025931100_ocu041-B19","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-642-02976-9_1","article-title":"Discovering novel adverse drug events using natural language processing and mining of the electronic health record","volume-title":"AIME \u201909 Proceedings of the 12th Conference on Artificial Intelligence in Medicine: Artificial Intelligence in Medicine","author":"Friedman","year":"2009"},{"key":"2020110613025931100_ocu041-B20","first-page":"1464","article-title":"A drug-adverse event extraction algorithm to support pharmacovigilance knowledge mining from PubMed citations","volume":"2011","author":"Wang","year":"2011","journal-title":"AMIA Annu Symp Proc."},{"key":"2020110613025931100_ocu041-B21","doi-asserted-by":"crossref","first-page":"15","DOI":"10.1186\/2041-1480-3-15","article-title":"Extraction of adverse drug effects from medical case reports","volume":"3","author":"Gurulingappa","year":"2012","journal-title":"J Biomed Semantics"},{"key":"2020110613025931100_ocu041-B22","first-page":"26","article-title":"Automated identification of adverse events from case reports using machine learning","volume-title":"Proceedings XXIV Conference of the European Federation for Medical Informatics. Workshop on Computational Methods in Pharmacovigilance","author":"Toldo","year":"2012"},{"key":"2020110613025931100_ocu041-B23","doi-asserted-by":"crossref","first-page":"1010","DOI":"10.1038\/clpt.2012.50","article-title":"Novel data-mining methodologies for adverse drug event discovery and analysis","volume":"91","author":"Harpaz","year":"2012","journal-title":"Clin Pharmacol Ther."},{"key":"2020110613025931100_ocu041-B24","doi-asserted-by":"crossref","first-page":"e10","DOI":"10.2196\/medinform.3022","article-title":"Automatically recognizing medication and adverse event information from food and drug administration\u2019s adverse event reporting system narratives","volume":"2","author":"Polepalli Ramesh","year":"2014","journal-title":"JMIR Med Informatics."},{"key":"2020110613025931100_ocu041-B25","first-page":"1019","article-title":"Pattern mining for extraction of mentions of adverse drug reactions from user comments","volume":"2011","author":"Nikfarjam","year":"2011","journal-title":"AMIA Annu Symp Proc."},{"key":"2020110613025931100_ocu041-B26","first-page":"134","article-title":"AZDrugMiner: an information extraction system for mining patient-reported adverse drug events","volume-title":"Proceedings of the 2013 international conference on Smart Health","author":"Liu","year":"2013"},{"key":"2020110613025931100_ocu041-B27","first-page":"217","article-title":"Predicting adverse drug events from personal health messages","volume":"2011","author":"Chee","year":"2011","journal-title":"AMIA Annu Symp Proc."},{"key":"2020110613025931100_ocu041-B28","doi-asserted-by":"crossref","first-page":"411","DOI":"10.1038\/nbt.1837","article-title":"Accelerated clinical discovery using self-reported patient data collected online and a patient-matching algorithm","volume":"29","author":"Wicks","year":"2011","journal-title":"Nat Biotechnol."},{"key":"2020110613025931100_ocu041-B29","doi-asserted-by":"crossref","first-page":"33","DOI":"10.1145\/2389707.2389714","article-title":"Social media mining for drug safety signal detection","volume-title":"Proceedings of the 2012 international workshop on Smart health and wellbeing. New York, USA: ACM Press","author":"Yang","year":", 2012"},{"key":"2020110613025931100_ocu041-B30","article-title":"Portable automatic text classification for adverse drug reaction detection via multi-corpus training","volume-title":"Journal of biomedical informatics","author":"Sarker","year":"2014"},{"key":"2020110613025931100_ocu041-B31","doi-asserted-by":"crossref","first-page":"404","DOI":"10.1136\/amiajnl-2012-001482","article-title":"Web-scale pharmacovigilance: listening to signals from the crowd","volume":"20","author":"White","year":"2013","journal-title":"J Am Med Inform Assoc."},{"key":"2020110613025931100_ocu041-B32","doi-asserted-by":"crossref","first-page":"343","DOI":"10.1038\/msb.2009.98","article-title":"A side effect resource to capture phenotypic effects of drugs","volume":"6","author":"Kuhn","year":"2010","journal-title":"Mol Syst Biol."},{"key":"2020110613025931100_ocu041-B33"},{"key":"2020110613025931100_ocu041-B34","doi-asserted-by":"crossref","first-page":"349","DOI":"10.1197\/jamia.M2592","article-title":"Estimating consumer familiarity with health terminology: a context-based approach","volume":"15","author":"Zeng-Treitler","year":"2008","journal-title":"J Am Med Informatics Assoc."},{"key":"2020110613025931100_ocu041-B35","first-page":"65","article-title":"MedDRA: an overview of the medical dictionary for regulatory activities","volume":"23","author":"Mozzicato","year":"2009","journal-title":"Pharmaceut Med."},{"key":"2020110613025931100_ocu041-B36","doi-asserted-by":"crossref","first-page":"25","DOI":"10.1007\/978-3-319-08416-9_3","article-title":"Identifying adverse drug events from health social media: a case study on heart disease discussion","volume-title":"International Conference on Smart Health","author":"Liu","year":"2014"},{"key":"2020110613025931100_ocu041-B37","doi-asserted-by":"crossref","first-page":"885","DOI":"10.1016\/j.jbi.2012.04.008","article-title":"Development of a benchmark corpus to support the automatic extraction of drug-related adverse effects from medical case reports","volume":"45","author":"Gurulingappa","year":"2012","journal-title":"J Biomed Inform."},{"key":"2020110613025931100_ocu041-B38","doi-asserted-by":"crossref","first-page":"164","DOI":"10.1007\/978-3-642-37210-0_18","article-title":"Discovering consumer health expressions from consumer-contributed content","volume-title":"Proceedings of International Conference on Social Computing, Behavioral-Cultural Modeling, and Prediction. Washington, D.C.","author":"Jiang","year":"2013"},{"key":"2020110613025931100_ocu041-B39","doi-asserted-by":"crossref","first-page":"37","DOI":"10.1177\/001316446002000104","article-title":"A coefficient of agreement for nominal scales","volume":"20","author":"Cohen","year":"1960","journal-title":"Educ Psychol Meas."},{"key":"2020110613025931100_ocu041-B40","first-page":"360","article-title":"Understanding interobserver agreement: the kappa statistic","volume":"37","author":"Viera","year":"2005","journal-title":"Fam Med."},{"key":"2020110613025931100_ocu041-B41","first-page":"1524","article-title":"Named entity recognition in tweets: an experimental study","volume-title":"Proceedings of the Conference on Empirical Methods in Natural Language Processing","author":"Ritter","year":"2011"},{"key":"2020110613025931100_ocu041-B42","first-page":"652","article-title":"BANNER: an executable survey of advances in biomedical named entity recognition","volume":"13","author":"Leaman","year":"2008","journal-title":"Pacific Symp Biocomput."},{"key":"2020110613025931100_ocu041-B43","article-title":"CRFsuite: a fast implementation of Conditional Random Fields (CRFs)","author":"Okazaki","year":"2007"},{"key":"2020110613025931100_ocu041-B44"},{"key":"2020110613025931100_ocu041-B45","article-title":"SCOWL (Spell Checker Oriented Word Lists)","author":"Atkinson"},{"key":"2020110613025931100_ocu041-B46","doi-asserted-by":"crossref","first-page":"197","DOI":"10.1109\/ICTAI.2007.117","article-title":"Dragon Toolkit: incorporating auto-learned semantic knowledge into large-scale text retrieval and mining","volume-title":"proceedings of the 19th IEEE International Conference on Tools with Artificial Intelligence (ICTAI)","author":"Zhou","year":"2007"},{"key":"2020110613025931100_ocu041-B47","doi-asserted-by":"crossref","first-page":"39","DOI":"10.1145\/219717.219748","article-title":"WordNet: a lexical database for English","volume":"38","author":"Miller","year":"1995","journal-title":"Commun ACM."},{"key":"2020110613025931100_ocu041-B48","first-page":"423","article-title":"Accurate unlexicalized parsing","volume-title":"Proceedings of the 41st Meeting of the Association for Computational Linguistics","author":"Manning","year":"2003"},{"key":"2020110613025931100_ocu041-B49","first-page":"119","article-title":"Syntactic dependency based heuristics for biological event extraction","volume-title":"Proceedings of the Workshop on Current Trends in Biomedical Natural Language Processing: Shared Task","author":"Kilicoglu","year":"2009"},{"key":"2020110613025931100_ocu041-B50","first-page":"165","article-title":"A hybrid system for emotion extraction from suicide notes","volume":"5","author":"Nikfarjam","year":"2012","journal-title":"Biomed Inform Insights."},{"key":"2020110613025931100_ocu041-B51"},{"key":"2020110613025931100_ocu041-B52","first-page":"1137","article-title":"A neural probabilistic language model","volume":"3","author":"Bengio","year":"2003","journal-title":"J Mach Learn Res."},{"key":"2020110613025931100_ocu041-B53","article-title":"Efficient estimation of word representations in vector space","author":"Mikolov","year":"2013","journal-title":"Proceedings of International Conference on Learning Representations"},{"key":"2020110613025931100_ocu041-B54","first-page":"17","article-title":"Effective mapping of biomedical text to the UMLS Metathesaurus: the MetaMap program","volume-title":"Proc AMIA Symp","author":"Aronson","year":"2001"},{"key":"2020110613025931100_ocu041-B55","first-page":"414","article-title":"PowerMap: mapping the real semantic web on the fly","author":"Lopez","year":"2006","journal-title":"Proceedings of the 5th International Semantic Web Conference"},{"key":"2020110613025931100_ocu041-B56","doi-asserted-by":"crossref","DOI":"10.1093\/database\/bau084","article-title":"Unsupervised Gene Function Extraction using Semantic Vectors","author":"Emadzadeh","year":"2014","journal-title":"Database"},{"key":"2020110613025931100_ocu041-B57","doi-asserted-by":"crossref","first-page":"1032","DOI":"10.1093\/bioinformatics\/btr042","article-title":"GeneTUKit: a software for document-level gene normalization","volume":"27","author":"Huang","year":"2011","journal-title":"Bioinformatics"},{"key":"2020110613025931100_ocu041-B58","first-page":"137","article-title":"Text categorization with support vector machines: learning with many relevant features","volume":"1398","author":"Joachims","year":"1998","journal-title":"Mach Learn ECML-98"},{"key":"2020110613025931100_ocu041-B59","first-page":"169","article-title":"Making large scale SVM learning practical","volume-title":"Advances in kernel methods - support vector learning","author":"Joachims","year":"1999"},{"key":"2020110613025931100_ocu041-B60","doi-asserted-by":"crossref","first-page":"92","DOI":"10.1186\/1471-2105-7-92","article-title":"Various criteria in the evaluation of biomedical named entity recognition","volume":"7","author":"Tsai","year":"2006","journal-title":"BMC Bioinformatics"},{"key":"2020110613025931100_ocu041-B61","doi-asserted-by":"crossref","first-page":"947","DOI":"10.3115\/992730.992783","article-title":"More accurate tests for the statistical significance of result differences","volume-title":"Proceedings of the 18th Conference on Computational linguistics","author":"Yeh","year":"2000"},{"key":"2020110613025931100_ocu041-B62","article-title":"User\u2019s guide to sigf: significance testing by approximate randomisation","author":"Pado","year":"2006"},{"key":"2020110613025931100_ocu041-B63","first-page":"109","article-title":"Sentence simplification aids protein-protein interaction extraction","author":"Jonnalagadda","year":"2009","journal-title":"Proceedings of the 3rd International Symposium on Languages in Biology and Medicine"}],"container-title":["Journal of the American Medical Informatics Association"],"original-title":[],"language":"en","link":[{"URL":"http:\/\/academic.oup.com\/jamia\/article-pdf\/22\/3\/671\/34146284\/ocu041.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"http:\/\/academic.oup.com\/jamia\/article-pdf\/22\/3\/671\/34146284\/ocu041.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,5,1]],"date-time":"2022-05-01T19:43:16Z","timestamp":1651434196000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/jamia\/article\/22\/3\/671\/776531"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2015,3,9]]},"references-count":63,"journal-issue":{"issue":"3","published-online":{"date-parts":[[2015,3,9]]},"published-print":{"date-parts":[[2015,5,1]]}},"URL":"https:\/\/doi.org\/10.1093\/jamia\/ocu041","relation":{},"ISSN":["1527-974X","1067-5027"],"issn-type":[{"value":"1527-974X","type":"electronic"},{"value":"1067-5027","type":"print"}],"subject":[],"published-other":{"date-parts":[[2015,5]]},"published":{"date-parts":[[2015,3,9]]}}}