{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,28]],"date-time":"2026-02-28T04:21:28Z","timestamp":1772252488327,"version":"3.50.1"},"reference-count":54,"publisher":"MDPI AG","issue":"1","license":[{"start":{"date-parts":[[2022,1,17]],"date-time":"2022-01-17T00:00:00Z","timestamp":1642377600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100006769","name":"Russian Science Foundation","doi-asserted-by":"publisher","award":["20-11-20246"],"award-info":[{"award-number":["20-11-20246"]}],"id":[{"id":"10.13039\/501100006769","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["BDCC"],"abstract":"<jats:p>Nowadays, the analysis of digital media aimed at prediction of the society\u2019s reaction to particular events and processes is a task of a great significance. Internet sources contain a large amount of meaningful information for a set of domains, such as marketing, author profiling, social situation analysis, healthcare, etc. In the case of healthcare, this information is useful for the pharmacovigilance purposes, including re-profiling of medications. The analysis of the mentioned sources requires the development of automatic natural language processing methods. These methods, in turn, require text datasets with complex annotation including information about named entities and relations between them. As the relevant literature analysis shows, there is a scarcity of datasets in the Russian language with annotated entity relations, and none have existed so far in the medical domain. This paper presents the first Russian-language textual corpus where entities have labels of different contexts within a single text, so that related entities share a common context. therefore this corpus is suitable for the task of belonging to the medical domain. Our second contribution is a method for the automated extraction of entity relations in Russian-language texts using the XLM-RoBERTa language model preliminarily trained on Russian drug review texts. A comparison with other machine learning methods is performed to estimate the efficiency of the proposed method. The method yields state-of-the-art accuracy of extracting the following relationship types: ADR\u2013Drugname, Drugname\u2013Diseasename, Drugname\u2013SourceInfoDrug, Diseasename\u2013Indication. As shown on the presented subcorpus from the Russian Drug Review Corpus, the method developed achieves a mean F1-score of 80.4% (estimated with cross-validation, averaged over the four relationship types). This result is 3.6% higher compared to the existing language model RuBERT, and 21.77% higher compared to basic ML classifiers.<\/jats:p>","DOI":"10.3390\/bdcc6010010","type":"journal-article","created":{"date-parts":[[2022,1,17]],"date-time":"2022-01-17T08:20:42Z","timestamp":1642407642000},"page":"10","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":4,"title":["Extraction of the Relations among Significant Pharmacological Entities in Russian-Language Reviews of Internet Users on Medications"],"prefix":"10.3390","volume":"6","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-6921-4133","authenticated-orcid":false,"given":"Alexander","family":"Sboev","sequence":"first","affiliation":[{"name":"National Research Centre \u201cKurchatov Institute\u201d, 123182 Moscow, Russia"},{"name":"Moscow Engineering Physics Institute, National Research Nuclear University, 115409 Moscow, Russia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5075-7229","authenticated-orcid":false,"given":"Anton","family":"Selivanov","sequence":"additional","affiliation":[{"name":"National Research Centre \u201cKurchatov Institute\u201d, 123182 Moscow, Russia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ivan","family":"Moloshnikov","sequence":"additional","affiliation":[{"name":"National Research Centre \u201cKurchatov Institute\u201d, 123182 Moscow, Russia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5595-6398","authenticated-orcid":false,"given":"Roman","family":"Rybka","sequence":"additional","affiliation":[{"name":"National Research Centre \u201cKurchatov Institute\u201d, 123182 Moscow, Russia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Artem","family":"Gryaznov","sequence":"additional","affiliation":[{"name":"National Research Centre \u201cKurchatov Institute\u201d, 123182 Moscow, Russia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Sanna","family":"Sboeva","sequence":"additional","affiliation":[{"name":"National Research Centre \u201cKurchatov Institute\u201d, 123182 Moscow, Russia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Gleb","family":"Rylkov","sequence":"additional","affiliation":[{"name":"National Research Centre \u201cKurchatov Institute\u201d, 123182 Moscow, Russia"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2022,1,17]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"288","DOI":"10.1016\/j.jbi.2015.11.001","article-title":"Pharmacovigilance through the development of text mining and natural language processing techniques","volume":"58","year":"2015","journal-title":"J. Biomed. Inform."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"012037","DOI":"10.1088\/1742-6596\/1686\/1\/012037","article-title":"A neural network algorithm for extracting pharmacological information from russian-language internet reviews on drugs","volume":"1686","author":"Sboev","year":"2020","journal-title":"J. Phys. Conf. Ser."},{"key":"ref_3","unstructured":"Sboev, A., Sboeva, S., Moloshnikov, I., Gryaznov, A., Rybka, R., Naumov, A., Selivanov, A., Rylkov, G., and Ilyin, V. (2021). An analysis of full-size Russian complexly NER labelled corpus of Internet user reviews on the drugs based on deep learning and language neural nets. arXiv."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"80","DOI":"10.37394\/232010.2020.17.10","article-title":"Artificial Intelligence: Learning and Limitations","volume":"17","author":"Oliveira","year":"2020","journal-title":"Wseas Trans. Adv. Eng. Educ."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"115","DOI":"10.37394\/23205.2020.19.16","article-title":"A Systemic Study of Pattern Recognition System Using Feedback Neural Networks","volume":"19","author":"Jebril","year":"2020","journal-title":"Wseas Trans. Comput."},{"key":"ref_6","first-page":"26","article-title":"POS-Tagging based Neural Machine Translation System for European Languages using Transformers\"","volume":"18","author":"Ganesh","year":"2021","journal-title":"Wseas Trans. Inf. Sci. Appl."},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Xu, H., Van Durme, B., and Murray, K. (2021, January 7\u201311). BERT, mBERT, or BiBERT? A Study on Contextualized Embeddings for Neural Machine Translation. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, Punta Cana, Dominican Republic.","DOI":"10.18653\/v1\/2021.emnlp-main.534"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Ge, Z., Sun, Y., and Smith, M. (2016, January 8\u201312). Authorship attribution using a neural network language model. Proceedings of the AAAI Conference on Artificial Intelligence, Burlingame, CA, USA.","DOI":"10.1609\/aaai.v30i1.9924"},{"key":"ref_9","unstructured":"Peters, M., Neumann, M., Iyyer, M., Gardner, M., Clark, C., and Lee, K. (2021, January 6\u201311). Deep contextualized word representations. Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Online."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Luong, M.T., Pham, H., and Manning, C.D. (2015). Effective approaches to attention-based neural machine translation. arXiv.","DOI":"10.18653\/v1\/D15-1166"},{"key":"ref_11","unstructured":"Portelli, B., Passabi, D., Serra, G., Santus, E., and Chersoni, E. (2021, January 8\u20139). Improving Adverse Drug Event Extraction with SpanBERT on Different Text Typologies. Proceedings of the 5th International Workshop on Health Intelligence (W3PHIAI-21), Palo Alto, CA, USA."},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Yan, H., Gui, T., Dai, J., Guo, Q., Zhang, Z., and Qiu, X. (2021). A Unified Generative Framework for Various NER Subtasks. arXiv.","DOI":"10.18653\/v1\/2021.acl-long.451"},{"key":"ref_13","unstructured":"Ge, S., Wu, F., Wu, C., Qi, T., Huang, Y., and Xie, X. (2021, October 30). FedNER: Privacy-Preserving Medical Named Entity Recognition with Federated Learning. Available online: https:\/\/arxiv.org\/abs\/2003.09288."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Wu, S., and He, Y. (2019, January 3\u20137). Enriching pre-trained language model with entity information for relation classification. Proceedings of the 28th ACM International Conference on Information and Knowledge Management, Beijing, China.","DOI":"10.1145\/3357384.3358119"},{"key":"ref_15","unstructured":"Giorgi, J., Wang, X., Sahar, N., Shin, W.Y., Bader, G.D., and Wang, B. (2019). End-to-end named entity recognition and relation extraction using pre-trained language models. arXiv."},{"key":"ref_16","unstructured":"Eberts, M., and Ulges, A. (2020). Span-Based Joint Entity and Relation Extraction with Transformer Pre-Training. ECAI 2020, IOS Press."},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"1234","DOI":"10.1093\/bioinformatics\/btz682","article-title":"BioBERT: A pre-trained biomedical language representation model for biomedical text mining","volume":"36","author":"Lee","year":"2019","journal-title":"Bioinformatics"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Gu, Y., Tinn, R., Cheng, H., Lucas, M., Usuyama, N., Liu, X., Naumann, T., Gao, J., and Poon, H. (2020). Domain-Specific Language Model Pretraining for Biomedical Natural Language Processing. arXiv.","DOI":"10.1145\/3458754"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Gordeev, D., Davletov, A., Rey, A., Akzhigitova, G., and Geymbukh, G. (2020). Relation extraction dataset for the russian language. Computational Linguistics and Intellectual Technologies: Proceedings of the International Conference \u201cDialog\u201d [Komp\u2019iuternaia Lingvistika i Intellektual\u2019nye Tehnologii: Trudy Mezhdunarodnoj Konferentsii \u201cDialog\u201d], Russian State University For The Humanities.","DOI":"10.28995\/2075-7182-2020-19-348-360"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Naseem, U., Dunn, A.G., Khushi, M., and Kim, J. (2021). Benchmarking for biomedical natural language processing tasks with a domain specific albert. arXiv.","DOI":"10.1186\/s12859-022-04688-w"},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"22","DOI":"10.1093\/jamia\/ocz075","article-title":"An ensemble of neural models for nested adverse drug events and medication extraction with subwords","volume":"27","author":"Ju","year":"2020","journal-title":"J. Am. Med. Inform. Assoc."},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"64","DOI":"10.1162\/tacl_a_00300","article-title":"Spanbert: Improving pre-training by representing and predicting spans","volume":"8","author":"Joshi","year":"2020","journal-title":"Trans. Assoc. Comput. Linguist."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Wang, J., and Lu, W. (2020, January 16\u201320). Two Are Better than One: Joint Entity and Relation Extraction with Table-Sequence Encoders. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), Online.","DOI":"10.18653\/v1\/2020.emnlp-main.133"},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"524","DOI":"10.1136\/jamia.2010.003939","article-title":"High accuracy information extraction of medication information from clinical notes: 2009 i2b2 medication extraction challenge","volume":"17","author":"Patrick","year":"2010","journal-title":"J. Am. Med. Inform. Assoc."},{"key":"ref_25","unstructured":"Anick, P., Hong, P., Xue, N., and Anick, D. (2010, January 12). I2B2 2010 challenge: Machine learning for information extraction from patient records. Proceedings of the 2010 i2b2\/VA Workshop on Challenges in Natural Language Processing for Clinical Data, Boston, MA, USA."},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"3","DOI":"10.1093\/jamia\/ocz166","article-title":"2018 n2c2 shared task on adverse drug events and medication extraction in electronic health records","volume":"27","author":"Henry","year":"2019","journal-title":"J. Am. Med. Inform. Assoc."},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"914","DOI":"10.1016\/j.jbi.2013.07.011","article-title":"The DDI corpus: An annotated corpus with pharmacological substances and drug\u2013drug interactions","volume":"46","author":"Declerck","year":"2013","journal-title":"J. Biomed. Inform."},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"1739","DOI":"10.1093\/bioinformatics\/btaa907","article-title":"Using Drug Descriptions and Molecular Structures for Drug-Drug Interaction Extraction from Literature","volume":"37","author":"Asada","year":"2020","journal-title":"Bioinformatics"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Beltagy, I., Lo, K., and Cohan, A. (2019). SciBERT: Pretrained Language Model for Scientific Text. arXiv.","DOI":"10.18653\/v1\/D19-1371"},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"885","DOI":"10.1016\/j.jbi.2012.04.008","article-title":"Development of a benchmark corpus to support the automatic extraction of drug-related adverse effects from medical case reports","volume":"45","author":"Gurulingappa","year":"2012","journal-title":"J. Biomed. Inform."},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Bruches, E., Pauls, A., Batura, T., and Isachenko, V. (2020, January 14\u201315). Entity Recognition and Relation Extraction from Scientific and Technical Texts in Russian. Proceedings of the 2020 Science and Artificial Intelligence Conference (SAI Ence), Novosibirsk, Russia.","DOI":"10.1109\/S.A.I.ence50533.2020.9303196"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Ivanin, V., Artemova, E., Batura, T., Ivanov, V., Sarkisyan, V., Tutubalina, E., and Smurov, I. (2020). Rurebus-2020 shared task: Russian relation extraction for business. Computational Linguistics and Intellectual Technologies, Russian State University for the Humanities.","DOI":"10.28995\/2075-7182-2020-19-416-431"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Bondarenko, I., Berezin, S., Pauls, A., Batura, T., Rubtsova, Y., and Tuchinov, B. (2020, January 14\u201315). Using Few-Shot Learning Techniques for Named Entity Recognition and Relation Extraction. Proceedings of the 2020 Science and Artificial Intelligence Conference (SAI Ence), Novosibirsk, Russia.","DOI":"10.1109\/S.A.I.ence50533.2020.9303192"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Loukachevitch, N., Artemova, E., Batura, T., Braslavski, P., Denisov, I., Ivanov, V., Manandhar, S., Pugachev, A., and Tutubalina, E. (2021). NEREL: A Russian Dataset with Nested Named Entities and Relations. arXiv.","DOI":"10.26615\/978-954-452-072-4_100"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Conneau, A., Khandelwal, K., Goyal, N., Chaudhary, V., Wenzek, G., Guzm\u00e1n, F., Grave, E., Ott, M., Zettlemoyer, L., and Stoyanov, V. (2019). Unsupervised cross-lingual representation learning at scale. arXiv.","DOI":"10.18653\/v1\/2020.acl-main.747"},{"key":"ref_36","first-page":"5998","article-title":"Attention is All you Need","volume":"Volume 30","author":"Vaswani","year":"2017","journal-title":"Advances in Neural Information Processing Systems"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Kudo, T., and Richardson, J. (2018). Sentencepiece: A simple and language independent subword tokenizer and detokenizer for neural text processing. arXiv.","DOI":"10.18653\/v1\/D18-2012"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Sboev, A., Selivanov, A., Rybka, R., Moloshnikov, I., and Rylkov, G. (2021, October 30). Evaluation of Machine Learning Methods for Relation Extraction Between Drug Adverse Effects and Medications in Russian Texts of Internet User Reviews. Available online: https:\/\/pos.sissa.it\/410\/006\/pdf.","DOI":"10.22323\/1.410.0006"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Smith, L.N. (2017, January 24\u201331). Cyclical learning rates for training neural networks. Proceedings of the 2017 IEEE Winter Conference on Applications of Computer Vision (WACV), Santa Rosa, CA, USA.","DOI":"10.1109\/WACV.2017.58"},{"key":"ref_40","first-page":"402","article-title":"Overfitting in neural nets: Backpropagation, conjugate gradient, and early stopping","volume":"13","author":"Caruana","year":"2000","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_41","doi-asserted-by":"crossref","first-page":"132502","DOI":"10.1109\/ACCESS.2020.3009733","article-title":"An evolutionary SVM model for DDOS attack detection in software defined networks","volume":"8","author":"Sahoo","year":"2020","journal-title":"IEEE Access"},{"key":"ref_42","doi-asserted-by":"crossref","first-page":"61","DOI":"10.1111\/mice.12564","article-title":"Automatic detection method of cracks from concrete surface imagery using two-step light gradient boosting machine","volume":"36","author":"Chun","year":"2021","journal-title":"Comput.-Aided Civil Infrastruct. Eng."},{"key":"ref_43","doi-asserted-by":"crossref","first-page":"102221","DOI":"10.1016\/j.ipm.2020.102221","article-title":"E-commerce product review sentiment classification based on a na\u00efve Bayes continuous learning framework","volume":"57","author":"Xu","year":"2020","journal-title":"Inf. Process. Manag."},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"Hosmer, D.W., Lemeshow, S., and Sturdivant, R.X. (2013). Applied Logistic Regression, John Wiley & Sons.","DOI":"10.1002\/9781118548387"},{"key":"ref_45","doi-asserted-by":"crossref","first-page":"293","DOI":"10.1023\/A:1018628609742","article-title":"Least squares support vector machine classifiers","volume":"9","author":"Suykens","year":"1999","journal-title":"Neural Process. Lett."},{"key":"ref_46","unstructured":"Rish, I. (2001, January 4). An empirical study of the naive Bayes classifier. Proceedings of the IJCAI 2001 workshop on empirical methods in artificial intelligence, Seattle, WA, USA."},{"key":"ref_47","unstructured":"Mason, L., Baxter, J., Bartlett, P., and Frean, M. (December, January 29). Boosting algorithms as gradient descent in function space. Proceedings of the NIPS, Denver, CO, USA."},{"key":"ref_48","unstructured":"Kuratov, Y., and Arkhipov, M. (2019). Adaptation of deep bidirectional multilingual transformers for Russian language. Komp\u2019juternaja Lingvistika i Intellektual\u2019nye Tehnologii, Russian State University For The Humanities."},{"key":"ref_49","unstructured":"Devlin, J., Chang, M.W., Lee, K., and Toutanova, K. (2018). Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv."},{"key":"ref_50","doi-asserted-by":"crossref","first-page":"357","DOI":"10.1038\/s41586-020-2649-2","article-title":"Array programming with NumPy","volume":"585","author":"Harris","year":"2020","journal-title":"Nature"},{"key":"ref_51","first-page":"2825","article-title":"Scikit-learn: Machine Learning in Python","volume":"12","author":"Pedregosa","year":"2011","journal-title":"J. Mach. Learn. Res."},{"key":"ref_52","unstructured":"Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., and Antiga, L. (2019, January 8\u201314). Pytorch: An imperative style, high-performance deep learning library. Proceedings of the 33rd Conference on Neural Information Processing Systems, Vancouver, BC, Canada."},{"key":"ref_53","unstructured":"Rajapakse, T.C. (2021, October 30). Simple Transformers. Available online: https:\/\/github.com\/ThilinaRajapakse\/simpletransformers."},{"key":"ref_54","doi-asserted-by":"crossref","unstructured":"raj Kanakarajan, K., Kundumani, B., and Sankarasubbu, M. (2021, January 11). BioELECTRA: Pretrained Biomedical text Encoder using Discriminators. Proceedings of the 20th Workshop on Biomedical Language Processing, Online.","DOI":"10.18653\/v1\/2021.bionlp-1.16"}],"container-title":["Big Data and Cognitive Computing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2504-2289\/6\/1\/10\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T22:02:39Z","timestamp":1760133759000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2504-2289\/6\/1\/10"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,1,17]]},"references-count":54,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2022,3]]}},"alternative-id":["bdcc6010010"],"URL":"https:\/\/doi.org\/10.3390\/bdcc6010010","relation":{"has-preprint":[{"id-type":"doi","id":"10.20944\/preprints202111.0344.v1","asserted-by":"object"}]},"ISSN":["2504-2289"],"issn-type":[{"value":"2504-2289","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,1,17]]}}}