{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,8]],"date-time":"2026-04-08T05:02:52Z","timestamp":1775624572405,"version":"3.50.1"},"reference-count":42,"publisher":"MDPI AG","issue":"4","license":[{"start":{"date-parts":[[2026,4,5]],"date-time":"2026-04-05T00:00:00Z","timestamp":1775347200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["MAKE"],"abstract":"<jats:p>The modern information retrieval field increasingly relies on hybrid search systems combining sparse retrieval with dense neural models. However, most existing hybrid frameworks employ static mixing coefficients and independent component training, failing to account for the specific needs of individual queries and corpus heterogeneity. In this paper, we introduce an adaptive hybrid retrieval framework featuring query-driven alpha prediction that dynamically calibrates the mixing weights based on query latent representations instantiated in a lightweight low-latency configuration and a full-capacity encoder-scale predictor, enabling flexible trade-offs between computational efficiency and retrieval accuracy without relying on resource-inefficient LLM-based online evaluation. Furthermore, we propose antagonist negative sampling, a novel training paradigm that optimizes the dense encoder to resolve the systematic failures of the lexical retriever, prioritizing hard negatives where BM25 exhibits high uncertainty. Empirical evaluations on large-scale multilingual benchmarks (MLDR and MIRACL) indicate that our approach demonstrates superior average performance compared to state-of-the-art models such as BGE-M3 and mGTE, achieving an nDCG@10 of 74.3 on long-document retrieval. Notably, our framework recovers up to 92.5% of the theoretical oracle performance and yields significant improvements in nDCG@10 across 16 languages, particularly in challenging long-context scenarios.<\/jats:p>","DOI":"10.3390\/make8040091","type":"journal-article","created":{"date-parts":[[2026,4,6]],"date-time":"2026-04-06T05:30:56Z","timestamp":1775453456000},"page":"91","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Query-Adaptive Hybrid Search"],"prefix":"10.3390","volume":"8","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-9442-8021","authenticated-orcid":false,"given":"Pavel","family":"Posokhov","sequence":"first","affiliation":[{"name":"Information Technologies and Programming Faculty, ITMO University, 197101 Saint Petersburg, Russia"},{"name":"STC-Innovations Ltd., 194044 Saint Petersburg, Russia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0001-7557-7870","authenticated-orcid":false,"given":"Stepan","family":"Skrylnikov","sequence":"additional","affiliation":[{"name":"STC-Innovations Ltd., 194044 Saint Petersburg, Russia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9054-5252","authenticated-orcid":false,"given":"Sergei","family":"Masliukhin","sequence":"additional","affiliation":[{"name":"Information Technologies and Programming Faculty, ITMO University, 197101 Saint Petersburg, Russia"},{"name":"STC-Innovations Ltd., 194044 Saint Petersburg, Russia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0009-9560-3413","authenticated-orcid":false,"given":"Alina","family":"Zavgorodniaia","sequence":"additional","affiliation":[{"name":"STC-Innovations Ltd., 194044 Saint Petersburg, Russia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8992-9654","authenticated-orcid":false,"given":"Olesia","family":"Koroteeva","sequence":"additional","affiliation":[{"name":"Information Technologies and Programming Faculty, ITMO University, 197101 Saint Petersburg, Russia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7010-1585","authenticated-orcid":false,"given":"Yuri","family":"Matveev","sequence":"additional","affiliation":[{"name":"Information Technologies and Programming Faculty, ITMO University, 197101 Saint Petersburg, Russia"},{"name":"STC-Innovations Ltd., 194044 Saint Petersburg, Russia"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2026,4,5]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"2355","DOI":"10.3390\/make6040116","article-title":"Systematic Analysis of Retrieval-Augmented Generation-Based LLMs for Medical Chatbot Applications","volume":"6","author":"Bora","year":"2024","journal-title":"Mach. Learn. Knowl. Extr."},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Lakatos, R., Pollner, P., Hajdu, A., and Jo\u00f3, T. (2025). Investigating the Performance of Retrieval-Augmented Generation and Domain-Specific Fine-Tuning for the Development of AI-Driven Knowledge-Based Systems. Mach. Learn. Knowl. Extr., 7.","DOI":"10.3390\/make7010015"},{"key":"ref_3","first-page":"333","article-title":"The Probabilistic Relevance Framework: BM25 and Beyond","volume":"3","author":"Robertson","year":"2009","journal-title":"Found. Trends Inf. Retr."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Arabzadeh, N., Yan, X., and Clarke, C. (2021, January 1\u20135). Predicting Efficiency\/Effectiveness Trade-offs for Dense vs. Sparse Retrieval Strategy Selection. Proceedings of the CIKM \u201921: Proceedings of the 30th ACM International Conference on Information & Knowledge Management, Gold Coast, QLD, Australia.","DOI":"10.1145\/3459637.3482159"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Prasanna, S.R.M., Karpov, A., Samudravijaya, K., and Agrawal, S.S. (2022). Personalizing Retrieval-Based Dialogue Agents. Proceedings of the Speech and Computer, Springer.","DOI":"10.1007\/978-3-031-20980-2"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Karpukhin, V., Oguz, B., Min, S., Lewis, P., Wu, L., Edunov, S., Chen, D., and Yih, W.t. (2020, January 16\u201320). Dense Passage Retrieval for Open-Domain Question Answering. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), Online.","DOI":"10.18653\/v1\/2020.emnlp-main.550"},{"key":"ref_7","unstructured":"Dev, A., Sharma, A., Agrawal, S.S., and Rani, R. (2025). Hybrid Approach to the Personification of Dialogue Agents. Proceedings of the Artificial Intelligence and Speech Technology, Springer."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Matveev, Y., Makhnytkina, O., Posokhov, P., Matveev, A., and Skrylnikov, S. (2022). Personalizing Hybrid-Based Dialogue Agents. Mathematics, 10.","DOI":"10.3390\/math10244657"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Chen, J., Xiao, S., Zhang, P., Luo, K., Lian, D., and Liu, Z. (2024, January 11\u201316). M3-Embedding: Multi-Linguality, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation. Proceedings of the Findings of the Association for Computational Linguistics: ACL 2024, Bangkok, Thailand.","DOI":"10.18653\/v1\/2024.findings-acl.137"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Zhang, X., Zhang, Y., Long, D., Xie, W., Dai, Z., Tang, J., Lin, H., Yang, B., Xie, P., and Huang, F. (2024, January 12\u201316). mGTE: Generalized Long-Context Text Representation and Reranking Models for Multilingual Text Retrieval. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: Industry Track, Miami, FL, USA.","DOI":"10.18653\/v1\/2024.emnlp-industry.103"},{"key":"ref_11","unstructured":"Hsu, H.L., and Tzeng, J. (2025). DAT: Dynamic Alpha Tuning for Hybrid Retrieval in Retrieval-Augmented Generation. arXiv."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"1114","DOI":"10.1162\/tacl_a_00595","article-title":"MIRACL: A Multilingual Retrieval Dataset Covering 18 Diverse Languages","volume":"11","author":"Zhang","year":"2023","journal-title":"Trans. Assoc. Comput. Linguist."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Formal, T., Lassance, C., Piwowarski, B., and Clinchant, S. (2021). SPLADE v2: Sparse Lexical and Expansion Model for Information Retrieval. arXiv.","DOI":"10.1145\/3404835.3463098"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Formal, T., Piwowarski, B., and Clinchant, S. (2021, January 11\u201315). SPLADE: Sparse Lexical and Expansion Model for First Stage Ranking. Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval, New York, NY, USA.","DOI":"10.1145\/3404835.3463098"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Mallia, A., Khattab, O., Suel, T., and Tonellotto, N. (2021, January 11\u201315). Learning Passage Impacts for Inverted Indexes. Proceedings of the SIGIR \u201921: Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval, Virtual Event.","DOI":"10.1145\/3404835.3463030"},{"key":"ref_16","unstructured":"Lin, J.J., and Ma, X. (2021). A Few Brief Notes on DeepImpact, COIL, and a Conceptual Framework for Information Retrieval Techniques. arXiv."},{"key":"ref_17","unstructured":"Humeau, S., Shuster, K., Lachaux, M., and Weston, J. (2020, January 26\u201330). Poly-encoders: Architectures and Pre-training Strategies for Fast and Accurate Multi-sentence Scoring. Proceedings of the 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia."},{"key":"ref_18","unstructured":"Burstein, J., Doran, C., and Solorio, T. (2019). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), Association for Computational Linguistics."},{"key":"ref_19","unstructured":"Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V. (2019). RoBERTa: A Robustly Optimized BERT Pretraining Approach. arXiv."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Reimers, N., and Gurevych, I. (2019). Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. arXiv.","DOI":"10.18653\/v1\/D19-1410"},{"key":"ref_21","unstructured":"Cappellato, L., Eickhoff, C., Ferro, N., and N\u00e9v\u00e9ol, A. (2020). Hybrid First-stage Retrieval Models for Biomedical Literature. Proceedings of the Working Notes of CLEF 2020\u2014Conference and Labs of the Evaluation Forum, Thessaloniki, Greece, 22\u201325 September 2020, CEUR-WS.org. CEUR Workshop Proceedings."},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Sawarkar, K., Mangal, A., and Solanki, S.R. (2024, January 7\u20139). Blended RAG: Improving RAG (Retriever-Augmented Generation) Accuracy with Semantic Search and Hybrid Query-Based Retrievers. Proceedings of the 2024 IEEE 7th International Conference on Multimedia Information Processing and Retrieval (MIPR), San Jose, CA, USA.","DOI":"10.1109\/MIPR62202.2024.00031"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Khattab, O., and Zaharia, M. (2020, January 25\u201330). ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT. Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval, New York, NY, USA.","DOI":"10.1145\/3397271.3401075"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Santhanam, K., Khattab, O., Saad-Falcon, J., Potts, C., and Zaharia, M. (2022, January 10\u201315). ColBERTv2: Effective and Efficient Retrieval via Lightweight Late Interaction. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Seattle, WA, USA.","DOI":"10.18653\/v1\/2022.naacl-main.272"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Cormack, G.V., Clarke, C.L.A., and Buettcher, S. (2009, January 19\u201323). Reciprocal rank fusion outperforms condorcet and individual rank learning methods. Proceedings of the 32nd International ACM SIGIR Conference on Research and Development in Information Retrieval, New York, NY, USA.","DOI":"10.1145\/1571941.1572114"},{"key":"ref_26","first-page":"1","article-title":"An Analysis of Fusion Functions for Hybrid Retrieval","volume":"42","author":"Bruch","year":"2023","journal-title":"ACM Trans. Inf. Syst."},{"key":"ref_27","unstructured":"Ma, X., Sun, K., Pradeep, R., and Lin, J. (2021). A Replication Study of Dense Passage Retriever. arXiv."},{"key":"ref_28","unstructured":"Bernston, A. (2026, March 06). Azure AI Search: Outperforming Vector Search with Hybrid Retrieval and Reranking. Available online: https:\/\/techcommunity.microsoft.com\/blog\/azure-ai-foundry-blog\/azure-ai-search-outperforming-vector-search-with-hybrid-retrieval-and-reranking\/3929167."},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"132498","DOI":"10.1016\/j.neucom.2025.132498","article-title":"A differentiable and uncertainty-aware mutual information regularizer for bias mitigation","volume":"669","author":"Incremona","year":"2026","journal-title":"Neurocomputing"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Shaik, H., Villuri, G., and Doboli, A. (2025). An Overview of Large Language Models and a Novel, Large Language Model-Based Cognitive Architecture for Solving Open-Ended Problems. Mach. Learn. Knowl. Extr., 7.","DOI":"10.3390\/make7040134"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Matveev, A., Makhnytkina, O., Matveev, Y., Svischev, A., Korobova, P., Rybin, A., and Akulov, A. (2021). Virtual Dialogue Assistant for Remote Exams. Mathematics, 9.","DOI":"10.3390\/math9182229"},{"key":"ref_32","first-page":"1016","article-title":"Prompt-based multi-task learning for robust text retrieval","volume":"24","author":"Masliukhin","year":"2024","journal-title":"Sci. Tech. J. Inf. Technol. Mech. Opt."},{"key":"ref_33","unstructured":"Xiong, L., Xiong, C., Li, Y., Tang, K., Liu, J., Bennett, P.N., Ahmed, J., and Overwijk, A. (2020). Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text Retrieval. arXiv."},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Zhou, K., Gong, Y., Liu, X., Zhao, W.X., Shen, Y., Dong, A., Lu, J., Majumder, R., Wen, J.r., and Duan, N. (2022, January 7\u201311). SimANS: Simple Ambiguous Negatives Sampling for Dense Text Retrieval. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing: Industry Track, Abu Dhabi, United Arab Emirates.","DOI":"10.18653\/v1\/2022.emnlp-industry.56"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Hofst\u00e4tter, S., Lin, S.C., Yang, J.H., Lin, J., and Hanbury, A. (2021, January 11\u201315). Efficiently Teaching an Effective Dense Retriever with Balanced Topic Aware Sampling. Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval, New York, NY, USA.","DOI":"10.1145\/3404835.3462891"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Qu, Y., Ding, Y., Liu, J., Liu, K., Ren, R., Zhao, W.X., Dong, D., Wu, H., and Wang, H. (2021, January 6\u201311). RocketQA: An Optimized Training Approach to Dense Passage Retrieval for Open-Domain Question Answering. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Online.","DOI":"10.18653\/v1\/2021.naacl-main.466"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Ren, R., Qu, Y., Liu, J., Zhao, W.X., She, Q., Wu, H., Wang, H., and Wen, J.R. (2021, January 7\u201311). RocketQAv2: A Joint Training Method for Dense Passage Retrieval and Passage Re-ranking. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, Punta Cana, Dominican Republic.","DOI":"10.18653\/v1\/2021.emnlp-main.224"},{"key":"ref_38","doi-asserted-by":"crossref","first-page":"1389","DOI":"10.1162\/tacl_a_00433","article-title":"MKQA: A Linguistically Diverse Benchmark for Multilingual Open Domain Question Answering","volume":"9","author":"Longpre","year":"2021","journal-title":"Trans. Assoc. Comput. Linguist."},{"key":"ref_39","unstructured":"Su, J., Lu, Y., Pan, S., Wen, B., and Liu, Y. (2021). RoFormer: Enhanced Transformer with Rotary Position Embedding. arXiv."},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Posokhov, P., Masliukhin, S., Stepan, S., Tirskikh, D., and Makhnytkina, O. (2025, January 1\u201327). Relevance Scores Calibration for Ranked List Truncation via TMP Adapter. Proceedings of the Findings of the Association for Computational Linguistics: ACL 2025, Vienna, Austria.","DOI":"10.18653\/v1\/2025.findings-acl.402"},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Bahri, D., Tay, Y., Zheng, C., Metzler, D., and Tomkins, A. (2020, January 25\u201330). Choppy: Cut Transformer for Ranked List Truncation. Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval, New York, NY, USA.","DOI":"10.1145\/3397271.3401188"},{"key":"ref_42","unstructured":"van den Oord, A., Li, Y., and Vinyals, O. (2018). Representation Learning with Contrastive Predictive Coding. arXiv."}],"container-title":["Machine Learning and Knowledge Extraction"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2504-4990\/8\/4\/91\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,8]],"date-time":"2026-04-08T04:23:57Z","timestamp":1775622237000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2504-4990\/8\/4\/91"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,4,5]]},"references-count":42,"journal-issue":{"issue":"4","published-online":{"date-parts":[[2026,4]]}},"alternative-id":["make8040091"],"URL":"https:\/\/doi.org\/10.3390\/make8040091","relation":{},"ISSN":["2504-4990"],"issn-type":[{"value":"2504-4990","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,4,5]]}}}