{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,8,2]],"date-time":"2025-08-02T17:39:55Z","timestamp":1754156395461,"version":"3.41.2"},"reference-count":51,"publisher":"Emerald","issue":"4","license":[{"start":{"date-parts":[[2021,10,12]],"date-time":"2021-10-12T00:00:00Z","timestamp":1633996800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/www.emerald.com\/insight\/site-policies"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["OIR"],"published-print":{"date-parts":[[2022,7,18]]},"abstract":"<jats:sec><jats:title content-type=\"abstract-subheading\">Purpose<\/jats:title><jats:p>This paper proposes a framework that automatically assesses content coverage and information quality of health websites for end-users.<\/jats:p><\/jats:sec><jats:sec><jats:title content-type=\"abstract-subheading\">Design\/methodology\/approach<\/jats:title><jats:p>The study investigates the impact of textual and content-based features in predicting the quality of health-related texts. Content-based features were acquired using an evidence-based practice guideline in diabetes. A set of textual features inspired by professional health literacy guidelines and the features commonly used for assessing information quality in other domains were also used. In this study, 60 websites about type 2 diabetes were methodically selected for inclusion. Two general practitioners used DISCERN to assess each website in terms of its content coverage and quality.<\/jats:p><\/jats:sec><jats:sec><jats:title content-type=\"abstract-subheading\">Findings<\/jats:title><jats:p>The proposed framework outputs were compared with the experts' evaluation scores. The best accuracy was obtained as 88 and 92% with textual features and content-based features for coverage assessment respectively. When both types of features were used, the proposed framework achieved 90% accuracy. For information quality assessment, the content-based features resulted in a higher accuracy of 92% against 88% obtained using the textual features.<\/jats:p><\/jats:sec><jats:sec><jats:title content-type=\"abstract-subheading\">Research limitations\/implications<\/jats:title><jats:p>The experiments were conducted for websites about type 2 diabetes. As the whole process is costly and requires extensive expert human labelling, the study was carried out in a single domain. However, the methodology is generalizable to other health domains for which evidence-based practice guidelines are available.<\/jats:p><\/jats:sec><jats:sec><jats:title content-type=\"abstract-subheading\">Practical implications<\/jats:title><jats:p>Finding high-quality online health information is becoming increasingly difficult due to the high volume of information generated by non-experts in the area. The search engines fail to rank objective health websites higher within the search results. The proposed framework can aid search engine and information platform developers to implement better retrieval techniques, in turn, facilitating end-users' access to high-quality health information.<\/jats:p><\/jats:sec><jats:sec><jats:title content-type=\"abstract-subheading\">Social implications<\/jats:title><jats:p>Erroneous, biased or partial health information is a serious problem for end-users who need access to objective information on their health problems. Such information may cause patients to stop their treatments provided by professionals. It might also have adverse financial implications by causing unnecessary expenditures on ineffective treatments. The ability to access high-quality health information has a positive effect on the health of both individuals and the whole society.<\/jats:p><\/jats:sec><jats:sec><jats:title content-type=\"abstract-subheading\">Originality\/value<\/jats:title><jats:p>The paper demonstrates that automatic assessment of health websites is a domain-specific problem, which cannot be addressed with the general information quality assessment methodologies in the literature. Content coverage of health websites has also been studied in the health domain for the first time in the literature.<\/jats:p><\/jats:sec>","DOI":"10.1108\/oir-02-2021-0089","type":"journal-article","created":{"date-parts":[[2021,10,13]],"date-time":"2021-10-13T18:54:19Z","timestamp":1634151259000},"page":"715-732","source":"Crossref","is-referenced-by-count":6,"title":["Quality assessment of web-based information on type 2 diabetes"],"prefix":"10.1108","volume":"46","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-7736-1021","authenticated-orcid":false,"given":"Didem","family":"\u00d6l\u00e7er","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7387-8621","authenticated-orcid":false,"given":"Tu\u011fba","family":"Ta\u015fkaya Temizel","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"140","published-online":{"date-parts":[[2021,10,12]]},"reference":[{"first-page":"1","volume-title":"University of Surrey Participation in TREC8: Weirdness Indexing for Logical Document Extrapolation and Retrieval (WILDER)","year":"1999","key":"key2022071415212390800_ref001"},{"first-page":"69","article-title":"Weirdness coefficient as a feature selection method for Arabic special domain text classification","year":"2012","key":"key2022071415212390800_ref002"},{"key":"key2022071415212390800_ref003","article-title":"Standards of medical care in diabetes-2016","volume":"39","author":"American Diabetes Association","year":"2016","journal-title":"Diabetes Care"},{"key":"key2022071415212390800_ref004","unstructured":"American Heart Association (2019), \u201cGuidelines and statements\u201d, available at: https:\/\/professional.heart.org\/professional\/GuidelinesStatements\/UCM_316885_Guidelines-Statements.jsp (accessed September 2019)."},{"issue":"1","key":"key2022071415212390800_ref005","doi-asserted-by":"crossref","first-page":"413","DOI":"10.14687\/jhs.v15i1.5256","article-title":"Utilization of active and passive constructions in English academic writing","volume":"15","year":"2018","journal-title":"Journal of Human Sciences"},{"first-page":"1095","article-title":"Size matters: word count as a measure of quality on Wikipedia","year":"2008","key":"key2022071415212390800_ref006"},{"issue":"6","key":"key2022071415212390800_ref007","article-title":"Optimal classifier for imbalanced data using Matthews Correlation Coefficient metric","volume":"12","year":"2017","journal-title":"PloS One"},{"issue":"2","key":"key2022071415212390800_ref008","doi-asserted-by":"crossref","first-page":"105","DOI":"10.1136\/jech.53.2.105","article-title":"DISCERN: an instrument for judging the quality of written consumer health information on treatment choices","volume":"53","year":"1999","journal-title":"Journal of Epidemiology and Community Health"},{"issue":"1","key":"key2022071415212390800_ref009","first-page":"1","article-title":"Automatic deception detection: methods for finding fake news","volume":"52","year":"2015","journal-title":"Proceedings of the Association for Information Science and Technology"},{"issue":"2","key":"key2022071415212390800_ref010","doi-asserted-by":"publisher","DOI":"10.2196\/18444","article-title":"Misinformation of COVID-19 on the internet: infodemiology study","volume":"6","year":"2020","journal-title":"JMIR Public Health and Surveillance"},{"issue":"3","key":"key2022071415212390800_ref011","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/2063504.2063507","article-title":"Automatic assessment of document quality in web collaborative digital libraries","volume":"2","year":"2011","journal-title":"Journal of Data and Information Quality (JDIQ)"},{"first-page":"295","article-title":"Quality assessment of collaborative content with minimal information","year":"2014","key":"key2022071415212390800_ref012"},{"first-page":"1","article-title":"The Stanford typed dependencies representation","year":"2008","key":"key2022071415212390800_ref013"},{"key":"key2022071415212390800_ref014","unstructured":"Department of Health (2003), \u201cNHS Toolkit for producing patient information\u201d, available at: https:\/\/www.uea.ac.uk\/documents\/246046\/0\/Toolkit+for+producing+patient+information.pdf (accessed March 2018)."},{"issue":"20","key":"key2022071415212390800_ref015","doi-asserted-by":"crossref","first-page":"2691","DOI":"10.1001\/jama.287.20.2691","article-title":"Empirical studies assessing the quality of health information for consumers on the world wide web: a systematic review","volume":"287","year":"2002","journal-title":"Jama"},{"issue":"1","key":"key2022071415212390800_ref016","doi-asserted-by":"crossref","first-page":"24","DOI":"10.4066\/AMJ.2014.1900","article-title":"Quality of patient health information on the internet: reviewing a complex and evolving landscape","volume":"7","year":"2014","journal-title":"Australasian Medical Journal"},{"volume-title":"The Philosophy of Information Quality","year":"2014","key":"key2022071415212390800_ref017"},{"first-page":"997","article-title":"Fact checking and analyzing the web","year":"2013","key":"key2022071415212390800_ref018"},{"issue":"4","key":"key2022071415212390800_ref019","doi-asserted-by":"crossref","first-page":"276","DOI":"10.1111\/jar.12127","article-title":"Evaluation of Autism-related health information on the web","volume":"28","year":"2015","journal-title":"Journal of Applied Research in Intellectual Disabilities"},{"issue":"5","key":"key2022071415212390800_ref020","doi-asserted-by":"crossref","first-page":"e59","DOI":"10.2196\/jmir.7.5.e59","article-title":"Automated assessment of the quality of depression websites","volume":"7","year":"2005","journal-title":"Journal of Medical Internet Research"},{"key":"key2022071415212390800_ref021","unstructured":"Health On the Net (2013), \u201cTrustworthy health and medical information: the Health on the Net\u201d, available at: https:\/\/www.hon.ch\/Global\/pdf\/TrustworthyOct2006.pdf (accessed November 2020)."},{"first-page":"366","article-title":"A semantic and syntactic text simplification tool for health content","year":"2010","key":"key2022071415212390800_ref022"},{"issue":"5","key":"key2022071415212390800_ref023","doi-asserted-by":"crossref","first-page":"461","DOI":"10.1002\/da.20381","article-title":"Quality of web-based information on social phobia: a cross sectional study","volume":"25","year":"2008","journal-title":"Depression and Anxiety"},{"first-page":"1729","article-title":"Mining quality phrases from massive text corpora","year":"2015","key":"key2022071415212390800_ref024"},{"issue":"1","key":"key2022071415212390800_ref025","first-page":"452","article-title":"Elastic net hypergraph learning for image clustering and semi-supervised classification","volume":"26","year":"2016","journal-title":"IEEE Transactions on Image Processing"},{"key":"key2022071415212390800_ref026","doi-asserted-by":"crossref","first-page":"154480","DOI":"10.1109\/ACCESS.2019.2946624","article-title":"Analysis and detection of health-related misinformation on Chinese social media","volume":"7","year":"2019","journal-title":"IEEE Access"},{"key":"key2022071415212390800_ref027","unstructured":"Medicines and Healthcare products Regulatory Agency (2014), \u201cBest practice guidance on patient information leaflets\u201d, available at: https:\/\/www.gov.uk\/government\/publications\/best-practice-guidance-on-patient-information-leaflets (accessed November 2020)."},{"key":"key2022071415212390800_ref028","unstructured":"MedlinePlus (2018), \u201cMedlinePlus trusted health information for you\u201d, available at: https:\/\/medlineplus.gov\/etr.html (accessed June 2018)."},{"first-page":"404","article-title":"Textrank: bringing order into text","year":"2004","key":"key2022071415212390800_ref029"},{"volume-title":"Existential Sentences in English (RLE Linguistics D: English Linguistics)","year":"2014","key":"key2022071415212390800_ref030"},{"key":"key2022071415212390800_ref031","unstructured":"National Comprehensive Cancer Network (2019), \u201cNCCN guidelines\u201d, available at: https:\/\/www.nccn.org\/professionals\/physician_gls\/default.aspx (accessed September 2019)."},{"key":"key2022071415212390800_ref032","unstructured":"National Library of Medicine (2018), \u201cMedical subject headings\u201d, available at: https:\/\/www.nlm.nih.gov\/mesh (accessed June 2018)."},{"key":"key2022071415212390800_ref033","doi-asserted-by":"crossref","first-page":"34","DOI":"10.1016\/j.breast.2015.10.001","article-title":"Evaluating the quality of internet information for breast cancer","volume":"25","year":"2016","journal-title":"The Breast"},{"issue":"5","key":"key2022071415212390800_ref034","doi-asserted-by":"crossref","first-page":"376","DOI":"10.1016\/j.breast.2004.03.003","article-title":"Breast cancer on the Internet: the quality of Swedish breast cancer websites","volume":"13","year":"2004","journal-title":"The Breast"},{"issue":"1","key":"key2022071415212390800_ref035","doi-asserted-by":"crossref","first-page":"20","DOI":"10.1177\/1465312518824100","article-title":"The quality of Internet information on lingual orthodontics in the English language, with DISCERN and JAMA","volume":"46","year":"2019","journal-title":"Journal of Orthodontics"},{"issue":"7","key":"key2022071415212390800_ref036","doi-asserted-by":"crossref","first-page":"1024","DOI":"10.1108\/OIR-01-2017-0028","article-title":"Predicting the quality of health web documents using their characteristics","volume":"42","year":"2018","journal-title":"Online Information Review"},{"first-page":"2931","article-title":"Truth of varying shades: analyzing language in fake news and political fact-checking","year":"2017","key":"key2022071415212390800_ref037"},{"year":"2007","key":"key2022071415212390800_ref038","article-title":"Exploring the feasibility of automatically rating online article quality"},{"issue":"14","key":"key2022071415212390800_ref039","doi-asserted-by":"publisher","first-page":"2012","DOI":"10.1093\/bioinformatics\/btaa535","article-title":"Predictive and interpretable models via the stacked elastic net","volume":"37","year":"2021","journal-title":"Bioinformatics"},{"issue":"10","key":"key2022071415212390800_ref040","doi-asserted-by":"crossref","first-page":"2293","DOI":"10.1002\/lary.26521","article-title":"Quality and readability assessment of websites related to recurrent respiratory papillomatosis","volume":"127","year":"2017","journal-title":"The Laryngoscope"},{"issue":"2","key":"key2022071415212390800_ref041","doi-asserted-by":"crossref","first-page":"104","DOI":"10.1016\/j.cmpb.2014.07.014","article-title":"Automatic information timeliness assessment of diabetes web sites by evidence-based medicine","volume":"117","year":"2014","journal-title":"Computer Methods and Programs in Biomedicine"},{"issue":"4","key":"key2022071415212390800_ref042","article-title":"Design and testing of a tool for evaluating the quality of diabetes consumer-information Web sites","volume":"5","year":"2003","journal-title":"Journal of Medical Internet Research"},{"first-page":"43","article-title":"A hybrid model for quality assessment of Wikipedia articles","year":"2017","key":"key2022071415212390800_ref043"},{"issue":"15","key":"key2022071415212390800_ref044","doi-asserted-by":"crossref","first-page":"1244","DOI":"10.1001\/jama.1997.03540390074039","article-title":"Assessing, controlling, and assuring the quality of medical information on the Internet: caveat lector et viewor\u2014let the reader and viewer beware","volume":"277","year":"1997","journal-title":"Jama"},{"key":"key2022071415212390800_ref045","unstructured":"The iWeb corpus (2018), \u201ciWeb: the 14 billion world web corpus\u201d, available at: https:\/\/www.english-corpora.org\/iweb\/ (accessed November 2020)."},{"first-page":"63","article-title":"Enriching the knowledge sources used in a maximum entropy part-of-speech tagger","year":"2000","key":"key2022071415212390800_ref046"},{"key":"key2022071415212390800_ref047","unstructured":"US Department of Health and Human Services, O. O (2010), \u201cHealth literacy online: a guide to writing and designing easy-to-use health web sites\u201d, available at: https:\/\/health.gov\/healthliteracyonline\/2010\/Web_Guide_Health_Lit_Online.pdf (accessed November 2020)."},{"volume-title":"English Sentence Analysis: An Introductory Course","year":"2000","key":"key2022071415212390800_ref048"},{"issue":"1","key":"key2022071415212390800_ref049","doi-asserted-by":"crossref","first-page":"16","DOI":"10.1002\/asi.24210","article-title":"Assessing the quality of information on Wikipedia: a deep\u2010learning approach","volume":"71","year":"2020","journal-title":"Journal of the Association for Information Science and Technology"},{"issue":"4","key":"key2022071415212390800_ref050","doi-asserted-by":"crossref","first-page":"821","DOI":"10.1093\/heapro\/dau019","article-title":"Quality of online information on type 2 diabetes: a cross-sectional study","volume":"30","year":"2015","journal-title":"Health Promotion International"},{"issue":"2","key":"key2022071415212390800_ref051","doi-asserted-by":"crossref","first-page":"301","DOI":"10.1111\/j.1467-9868.2005.00503.x","article-title":"Regularization and variable selection via the elastic net","volume":"67","year":"2005","journal-title":"Journal of the Royal Statistical Society: Series B (Statistical Methodology)"}],"container-title":["Online Information Review"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.emerald.com\/insight\/content\/doi\/10.1108\/OIR-02-2021-0089\/full\/xml","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/www.emerald.com\/insight\/content\/doi\/10.1108\/OIR-02-2021-0089\/full\/html","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,7,24]],"date-time":"2025-07-24T22:42:18Z","timestamp":1753396938000},"score":1,"resource":{"primary":{"URL":"http:\/\/www.emerald.com\/oir\/article\/46\/4\/715-732\/314929"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,10,12]]},"references-count":51,"journal-issue":{"issue":"4","published-online":{"date-parts":[[2021,10,12]]},"published-print":{"date-parts":[[2022,7,18]]}},"alternative-id":["10.1108\/OIR-02-2021-0089"],"URL":"https:\/\/doi.org\/10.1108\/oir-02-2021-0089","relation":{},"ISSN":["1468-4527"],"issn-type":[{"type":"print","value":"1468-4527"}],"subject":[],"published":{"date-parts":[[2021,10,12]]}}}