{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,2]],"date-time":"2026-08-02T19:42:06Z","timestamp":1785699726312,"version":"3.56.0"},"reference-count":14,"publisher":"Frontiers Media SA","license":[{"start":{"date-parts":[[2023,3,14]],"date-time":"2023-03-14T00:00:00Z","timestamp":1678752000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100000780","name":"European Commission","doi-asserted-by":"publisher","id":[{"id":"10.13039\/501100000780","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["frontiersin.org"],"crossmark-restriction":true},"short-container-title":["Front. Artif. Intell."],"abstract":"<jats:sec><jats:title>Introduction<\/jats:title><jats:p>This study presents COVID-Twitter-BERT (CT-BERT), a transformer-based model that is pre-trained on a large corpus of COVID-19 related Twitter messages. CT-BERT is specifically designed to be used on COVID-19 content, particularly from social media, and can be utilized for various natural language processing tasks such as classification, question-answering, and chatbots. This paper aims to evaluate the performance of CT-BERT on different classification datasets and compare it with BERT-LARGE, its base model.<\/jats:p><\/jats:sec><jats:sec><jats:title>Methods<\/jats:title><jats:p>The study utilizes CT-BERT, which is pre-trained on a large corpus of COVID-19 related Twitter messages. The authors evaluated the performance of CT-BERT on five different classification datasets, including one in the target domain. The model's performance is compared to its base model, BERT-LARGE, to measure the marginal improvement. The authors also provide detailed information on the training process and the technical specifications of the model.<\/jats:p><\/jats:sec><jats:sec><jats:title>Results<\/jats:title><jats:p>The results indicate that CT-BERT outperforms BERT-LARGE with a marginal improvement of 10-30% on all five classification datasets. The largest improvements are observed in the target domain. The authors provide detailed performance metrics and discuss the significance of these results.<\/jats:p><\/jats:sec><jats:sec><jats:title>Discussion<\/jats:title><jats:p>The study demonstrates the potential of pre-trained transformer models, such as CT-BERT, for COVID-19 related natural language processing tasks. The results indicate that CT-BERT can improve the classification performance on COVID-19 related content, especially on social media. These findings have important implications for various applications, such as monitoring public sentiment and developing chatbots to provide COVID-19 related information. The study also highlights the importance of using domain-specific pre-trained models for specific natural language processing tasks. Overall, this work provides a valuable contribution to the development of COVID-19 related NLP models.<\/jats:p><\/jats:sec>","DOI":"10.3389\/frai.2023.1023281","type":"journal-article","created":{"date-parts":[[2023,3,14]],"date-time":"2023-03-14T05:28:32Z","timestamp":1678771712000},"update-policy":"https:\/\/doi.org\/10.3389\/crossmark-policy","source":"Crossref","is-referenced-by-count":130,"title":["COVID-Twitter-BERT: A natural language processing model to analyse COVID-19 content on Twitter"],"prefix":"10.3389","volume":"6","author":[{"given":"Martin","family":"M\u00fcller","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Marcel","family":"Salath\u00e9","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Per E.","family":"Kummervold","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1965","published-online":{"date-parts":[[2023,3,14]]},"reference":[{"key":"B1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.1903.10676","article-title":"Scibert: pretrained contextualized embeddings for scientific text","author":"Beltagy","year":"2019","journal-title":"arXiv preprint"},{"key":"B2","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.1810.04805","article-title":"Bert: pre-training of deep bidirectional transformers for language understanding","author":"Devlin","year":"2018","journal-title":"arXiv preprint"},{"key":"B3","first-page":"411","article-title":"spacy 2: natural language understanding with bloom embeddings, convolutional neural networks and incremental parsing","author":"Honnibal","year":"2017","journal-title":"To appear"},{"key":"B4","doi-asserted-by":"publisher","first-page":"e29584","DOI":"10.2196\/29584","article-title":"Categorizing vaccine confidence with a transformer-based machine learning model: analysis of nuances of vaccine sentiment in twitter discourse","volume":"9","author":"Kummervold","year":"2021","journal-title":"JMIR Med. Inform."},{"key":"B5","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.1909.11942","article-title":"Albert: a lite bert for self-supervised learning of language representations","author":"Lan","year":"2019","journal-title":"arXiv preprint"},{"key":"B6","doi-asserted-by":"publisher","first-page":"1234","DOI":"10.1093\/bioinformatics\/btz682","article-title":"Biobert: a pre-trained biomedical language representation model for biomedical text mining","volume":"36","author":"Lee","year":"2020","journal-title":"Bioinformatics"},{"key":"B7","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.1907.11692","article-title":"Roberta: a robustly optimized bert pretraining approach","author":"Liu","year":"2019","journal-title":"arXiv preprint"},{"key":"B8","doi-asserted-by":"publisher","first-page":"6627","DOI":"10.1016\/j.vaccine.2020.07.072","article-title":"\u201cvaccines for pregnant women...?! absurd\u201d\u2013mapping maternal vaccination discourse and stance on social media over six months","volume":"38","author":"Martin","year":"2020","journal-title":"Vaccine"},{"key":"B9","doi-asserted-by":"publisher","first-page":"81","DOI":"10.3389\/fpubh.2019.00081","article-title":"Crowdbreaks: tracking health trends using public social media data and crowdsourcing","volume":"7","author":"M\u00fcller","year":"2019","journal-title":"Front. Public Health"},{"key":"B10","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/S16-1001","article-title":"Semeval-2016 task 4: sentiment analysis in twitter","author":"Nakov","year":"2019","journal-title":"arXiv preprint"},{"key":"B11","doi-asserted-by":"publisher","first-page":"13762","DOI":"10.1073\/pnas.1704093114","article-title":"Critical dynamics in population vaccinating behavior","volume":"114","author":"Pananos","year":"2017","journal-title":"Proc. Natl. Acad. Sci. U.S.A."},{"key":"B12","first-page":"115","article-title":"\u201cSeeing stars: exploiting class relationships for sentiment categorization with respect to rating scales,\u201d","volume-title":"Proceedings of the 43rd Annual Meeting on Association for Computational Linguistics","author":"Pang","year":"2005"},{"key":"B13","first-page":"1631","article-title":"\u201cRecursive deep models for semantic compositionality over a sentiment treebank,\u201d","author":"Socher","year":"2013","journal-title":"Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing"},{"key":"B14","first-page":"5998","article-title":"\u201cAttention is all you need,\u201d","author":"Vaswani","year":"2017","journal-title":"Advances in Neural Information Processing Systems"}],"container-title":["Frontiers in Artificial Intelligence"],"original-title":[],"link":[{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/frai.2023.1023281\/full","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,3,14]],"date-time":"2023-03-14T05:28:40Z","timestamp":1678771720000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/frai.2023.1023281\/full"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,3,14]]},"references-count":14,"alternative-id":["10.3389\/frai.2023.1023281"],"URL":"https:\/\/doi.org\/10.3389\/frai.2023.1023281","relation":{},"ISSN":["2624-8212"],"issn-type":[{"value":"2624-8212","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,3,14]]},"article-number":"1023281"}}