{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,15]],"date-time":"2026-06-15T14:26:13Z","timestamp":1781533573741,"version":"3.54.5"},"reference-count":28,"publisher":"MDPI AG","issue":"5","license":[{"start":{"date-parts":[[2026,4,23]],"date-time":"2026-04-23T00:00:00Z","timestamp":1776902400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Ministry of Science and Higher Education of the Republic of Kazakhstan","award":["AP22787186"],"award-info":[{"award-number":["AP22787186"]}]},{"award":["AP22787186"],"award-info":[{"award-number":["AP22787186"]}],"id":[{"id":"https:\/\/ror.org\/05tyne317","id-type":"ROR","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Computers"],"abstract":"<jats:p>Learning meaningful representations of scientific documents is essential for information retrieval, knowledge discovery, and recommendation systems. Traditional methods such as TF-IDF rely on lexical matching and fail to capture deeper semantic relationships, while transformer-based approaches typically depend on limited supervision signals. In this work, we propose a Triple-Source automatic supervision framework for learning document embeddings from scientific corpora. The model integrates three types of supervision\u2013title\u2013abstract pairs, same-category document pairs, and document-level semantic relationships\u2014within a unified contrastive learning framework based on a multilingual XLM-RoBERTa encoder. Unlike prior approaches that rely on citation graphs or manual annotations, our method enables citation-free and annotation-free representation learning using only lightweight metadata. Experiments on a publicly available arXiv dataset consisting of 98,649 documents demonstrate improved semantic retrieval performance, achieving Recall@1 = 0.6181 for same-category retrieval and outperforming both TF-IDF and single-source transformer baselines. The learned embeddings also exhibit improved clustering of scientific domains, indicating more structured semantic representations.<\/jats:p>","DOI":"10.3390\/computers15050268","type":"journal-article","created":{"date-parts":[[2026,4,23]],"date-time":"2026-04-23T12:04:42Z","timestamp":1776945882000},"page":"268","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Learning Scientific Document Representations via Triple-Source Automatic Supervision Without Annotations or Citations"],"prefix":"10.3390","volume":"15","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-1470-3706","authenticated-orcid":false,"given":"Mussa","family":"Turdalyuly","sequence":"first","affiliation":[{"name":"Institute of Information and Computational Technologies, Almaty 050010, Kazakhstan"},{"name":"School of Engineering and Information Technology, META University, Almaty 050000, Kazakhstan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1332-3936","authenticated-orcid":false,"given":"Ainur","family":"Tursynkhan","sequence":"additional","affiliation":[{"name":"Software Engineering Department, International Engineering and Technological University, Almaty 050060, Kazakhstan"},{"name":"Department of Computer and Information Technology, Purdue University, West Lafayette, IN 47907, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2013-1513","authenticated-orcid":false,"given":"Aigerim","family":"Yerimbetova","sequence":"additional","affiliation":[{"name":"Institute of Information and Computational Technologies, Almaty 050010, Kazakhstan"},{"name":"School of Engineering and Information Technology, META University, Almaty 050000, Kazakhstan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Tolganay","family":"Turdalykyzy","sequence":"additional","affiliation":[{"name":"Institute of Information and Computational Technologies, Almaty 050010, Kazakhstan"},{"name":"School of Engineering and Information Technology, META University, Almaty 050000, Kazakhstan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9849-6176","authenticated-orcid":false,"given":"Bakzhan","family":"Sakenov","sequence":"additional","affiliation":[{"name":"Institute of Information and Computational Technologies, Almaty 050010, Kazakhstan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4835-5751","authenticated-orcid":false,"given":"Nurzhan","family":"Mukazhanov","sequence":"additional","affiliation":[{"name":"Institute of Information and Computational Technologies, Almaty 050010, Kazakhstan"},{"name":"School of Digital Technologies, Narxoz University, Almaty 050035, Kazakhstan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8134-0466","authenticated-orcid":false,"given":"Nazerke","family":"Baisholan","sequence":"additional","affiliation":[{"name":"Software Engineering Department, International Engineering and Technological University, Almaty 050060, Kazakhstan"},{"name":"Department of Computer and Information Technology, Purdue University, West Lafayette, IN 47907, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2026,4,23]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Reimers, N., and Gurevych, I. (2019, January 3\u20137). Sentence-BERT: Sentence Embeddings using Siamese BERT Networks. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, Hong Kong, China.","DOI":"10.18653\/v1\/D19-1410"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Cohan, A., Feldman, S., Beltagy, I., Downey, D., and Weld, D. (2020, January 5\u201310). SPECTER: Document-Level Representation Learning using Citation-Informed Transformers. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Online.","DOI":"10.18653\/v1\/2020.acl-main.207"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Beltagy, I., Lo, K., and Cohan, A. (2019, January 3\u20137). SciBERT: A Pretrained Language Model for Scientific Text. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, Hong Kong, China.","DOI":"10.18653\/v1\/D19-1371"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Batura, T., Yerimbetova, A., Mukazhanov, N., Shvarts, N., Sakenov, B., and Turdalyuly, M. (2025). Information Extraction from Multi-Domain Scientific Documents: Methods and Insights. Appl. Sci., 15.","DOI":"10.3390\/app15169086"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Gao, T., Yao, X., and Chen, D. (2021, January 7\u201311). SimCSE: Simple Contrastive Learning of Sentence Embeddings. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, Online.","DOI":"10.18653\/v1\/2021.emnlp-main.552"},{"key":"ref_6","first-page":"993","article-title":"Latent Dirichlet Allocation","volume":"3","author":"Blei","year":"2003","journal-title":"J. Mach. Learn. Res."},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Lo, K., Wang, L., Neumann, M., Kinney, R., and Weld, D. (2020, January 5\u201310). S2ORC: The Semantic Scholar Open Research Corpus. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Online.","DOI":"10.18653\/v1\/2020.acl-main.447"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Hu, J., Xia, W., Zhang, X., Fu, C., Wu, W., Huan, Z., Li, A., Tang, Z., and Zhou, J. (2024). Enhancing Sequential Recommendation via LLM-based Semantic Embedding Learning. Companion Proceedings of the ACM Web Conference 2024 (WWW \u201824), Association for Computing Machinery.","DOI":"10.1145\/3589335.3648307"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Rasool, A., Shahzad, M.I., Aslam, H., Chan, V., and Arshad, M.A. (2025). Emotion-Aware Embedding Fusion in Large Language Models (Flan-T5, Llama 2, DeepSeek-R1, and ChatGPT 4) for Intelligent Response Generation. AI, 6.","DOI":"10.3390\/ai6030056"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Tang, J., Zhang, J., Yao, L., Li, J., Zhang, L., and Su, Z. (2008, January 24\u201327). ArnetMiner: Extraction and Mining of Academic Social Networks. Proceedings of the 14th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, Las Vegas, NV, USA.","DOI":"10.1145\/1401890.1402008"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Colangelo, M.T., Meleti, M., Guizzardi, S., Calciolari, E., and Galli, C. (2025). A Comparative Analysis of Sentence Transformer Models for Automated Journal Recommendation Using PubMed Metadata. Big Data Cogn. Comput., 9.","DOI":"10.20944\/preprints202501.1334.v1"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Fahrudin, T.M., Funabiki, N., Brata, K.C., Naing, I., Aung, S.T., Muhaimin, A., and Prasetya, D.A. (2025). An Improved Reference Paper Collection System Using Web Scraping with Three Enhancements. Future Internet, 17.","DOI":"10.3390\/fi17050195"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Nosov, P., Melnyk, O., Malaksiano, M., Mamenko, P., Onyshko, D., Fomin, O., P\u00ed\u0161t\u011bk, V., and Ku\u010dera, P. (2025). Machine Learning-Based Semantic Analysis of Scientific Publications for Knowledge Extraction in Safety-Critical Domains. Mach. Learn. Knowl. Extr., 7.","DOI":"10.3390\/make7040150"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Malashin, I., Martysyuk, D., Tynchenko, V., Gantimurov, A., Nelyub, V., and Borodulin, A. (2026). Soft-Prompted Semantic Normalization for Unsupervised Analysis of the Scientific Literature. Mach. Learn. Knowl. Extr., 8.","DOI":"10.3390\/make8030063"},{"key":"ref_15","unstructured":"Neelakantan, A., Xu, T., Puri, R., Radford, A., Han, J.M., Tworek, J., Yuan, Q., Tezak, N., Kim, J.W., and Hallacy, C. (2022). Text and Code Embeddings by Contrastive Pre-Training. arXiv."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Ostendorff, M., Ruas, T., Zesch, T., Bourgonje, P., and Rehm, G. (2022, January 7\u201311). Neighborhood Contrastive Learning for Scientific Document Representations with Citation Embeddings. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, Abu Dhabi, United Arab Emirates.","DOI":"10.18653\/v1\/2022.emnlp-main.802"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Huang, Y., Zhu, S., Liu, W., Wang, J., and Wei, X. (2025). Addressing Asymmetry in Contrastive Learning: LLM-Driven Sentence Embeddings with Ranking and Label Smoothing. Symmetry, 17.","DOI":"10.3390\/sym17050646"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Oro, E., Granata, F.M., and Ruffolo, M. (2025). A Comprehensive Evaluation of Embedding Models and LLMs for IR and QA Across English and Italian. Big Data Cogn. Comput., 9.","DOI":"10.20944\/preprints202502.2143.v1"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Mukazhanov, N., Yerimbetova, A., Turdalyuly, M., and Sakenov, B. (2024, January 26\u201328). Named Entities Recognition in Kazakh Text by SpaCy NER Models. Proceedings of the 9th International Conference on Computer Science and Engineering (UBMK), Antalya, Turkey.","DOI":"10.1109\/UBMK63289.2024.10773475"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Mukazhanov, N., Batura, T., Yerimbetova, A., Turdalyuly, M., Sakenov, B., and Bayekeyeva, A. (2025, January 17\u201321). Kazakh Text Classification using Deep Learning Approaches. Proceedings of the 10th International Conference on Computer Science and Engineering (UBMK), Istanbul, Turkey.","DOI":"10.1109\/UBMK67458.2025.11207069"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Karpukhin, V., Oguz, B., Min, S., Lewis, P., and Wu, L. (2020, January 16\u201320). Dense Passage Retrieval for Open-Domain Question Answering. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, Online.","DOI":"10.18653\/v1\/2020.emnlp-main.550"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Sidi, M.L., and Gunal, S. (2023). A Purely Entity-Based Semantic Search Approach for Document Retrieval. Appl. Sci., 13.","DOI":"10.20944\/preprints202308.1279.v1"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Luo, H., Luo, X., Zhao, W., Peng, Q., Chen, K., Liu, Y., and Du, C. (2026). Domain-Specific Retrieval-Augmented Generation with Adaptive Embedding and Knowledge Distillation-Based Re-Ranking. Processes, 14.","DOI":"10.3390\/pr14010099"},{"key":"ref_24","unstructured":"Thakur, N., Reimers, N., R\u00fcckl\u00e9, A., Srivastava, A., and Gurevych, I. (2021, January 6\u201314). BEIR: A Heterogeneous Benchmark for Zero-Shot Evaluation of Information Retrieval Models. Proceedings of the 35th Conference on Neural Information Processing Systems, Online."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Conneau, A., Khandelwal, K., Goyal, N., Chaudhary, V., Wenzek, G., Guzm\u00e1n, F., Grave, E., Ott, M., Zettlemoyer, L., and Stoyanov, V. (2020, January 5\u201310). Unsupervised Cross-lingual Representation Learning at Scale. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Online.","DOI":"10.18653\/v1\/2020.acl-main.747"},{"key":"ref_26","unstructured":"van den Oord, A., Li, Y., and Vinyals, O. (2018). Representation Learning with Contrastive Predictive Coding. arXiv."},{"key":"ref_27","unstructured":"Clement, C.B., Bierbaum, M., O\u2019Keeffe, K.P., and Alemi, A.A. (2019). On the Use of ArXiv as a Dataset. arXiv."},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Muennighoff, N., Tazi, N., Magne, L., and Reimers, N. (2023, January 2\u20136). MTEB: Massive Text Embedding Benchmark. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, Dubrovnik, Croatia.","DOI":"10.18653\/v1\/2023.eacl-main.148"}],"container-title":["Computers"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2073-431X\/15\/5\/268\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,5,12]],"date-time":"2026-05-12T04:14:36Z","timestamp":1778559276000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2073-431X\/15\/5\/268"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,4,23]]},"references-count":28,"journal-issue":{"issue":"5","published-online":{"date-parts":[[2026,5]]}},"alternative-id":["computers15050268"],"URL":"https:\/\/doi.org\/10.3390\/computers15050268","relation":{},"ISSN":["2073-431X"],"issn-type":[{"value":"2073-431X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,4,23]]}}}