{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,24]],"date-time":"2026-06-24T05:15:45Z","timestamp":1782278145137,"version":"3.54.5"},"reference-count":46,"publisher":"MDPI AG","issue":"2","license":[{"start":{"date-parts":[[2023,4,28]],"date-time":"2023-04-28T00:00:00Z","timestamp":1682640000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Ministry of Science and Higher Education of the Russian Federation","award":["075-15-2022-311"],"award-info":[{"award-number":["075-15-2022-311"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Informatics"],"abstract":"<jats:p>This paper provides new models of the attention-based random forests called LARF (leaf attention-based random forest). The first idea behind the models is to introduce a two-level attention, where one of the levels is the \u201cleaf\u201d attention, and the attention mechanism is applied to every leaf of trees. The second level is the tree attention depending on the \u201cleaf\u201d attention. The second idea is to replace the softmax operation in the attention with the weighted sum of the softmax operations with different parameters. It is implemented by applying a mixture of Huber\u2019s contamination models and can be regarded as an analog of the multi-head attention, with \u201cheads\u201d defined by selecting a value of the softmax parameter. Attention parameters are simply trained by solving the quadratic optimization problem. To simplify the tuning process of the models, it is proposed to convert the tuning contamination parameters into trainable parameters and to compute them by solving the quadratic optimization problem. Many numerical experiments with real datasets are performed for studying LARFs. The code of the proposed algorithms is available.<\/jats:p>","DOI":"10.3390\/informatics10020040","type":"journal-article","created":{"date-parts":[[2023,4,28]],"date-time":"2023-04-28T04:36:15Z","timestamp":1682656575000},"page":"40","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":1,"title":["LARF: Two-Level Attention-Based Random Forests with a Mixture of Contamination Models"],"prefix":"10.3390","volume":"10","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-1542-6480","authenticated-orcid":false,"given":"Andrei","family":"Konstantinov","sequence":"first","affiliation":[{"name":"Higher School of Artificial Intelligence, Peter the Great St.Petersburg Polytechnic University, Polytechnicheskaya, 29, 195251 St. Petersburg, Russia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5637-1420","authenticated-orcid":false,"given":"Lev","family":"Utkin","sequence":"additional","affiliation":[{"name":"Higher School of Artificial Intelligence, Peter the Great St.Petersburg Polytechnic University, Polytechnicheskaya, 29, 195251 St. Petersburg, Russia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3583-7324","authenticated-orcid":false,"given":"Vladimir","family":"Muliukha","sequence":"additional","affiliation":[{"name":"Higher School of Artificial Intelligence, Peter the Great St.Petersburg Polytechnic University, Polytechnicheskaya, 29, 195251 St. Petersburg, Russia"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2023,4,28]]},"reference":[{"key":"ref_1","unstructured":"Chaudhari, S., Mithal, V., Polatkan, G., and Ramanath, R. (2019). An attentive survey of attention models. arXiv."},{"key":"ref_2","unstructured":"Correia, A., and Colombini, E. (2021). Attention, please! A survey of neural attention models in deep learning. arXiv, Available online: https:\/\/arxiv.org\/abs\/2103.16775."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"6037","DOI":"10.1007\/s10462-022-10148-x","article-title":"Attention, please! A survey of neural attention models in deep learning","volume":"55","author":"Correia","year":"2022","journal-title":"Artif. Intell. Rev."},{"key":"ref_4","unstructured":"Lin, T., Wang, Y., Liu, X., and Qiu, X. (2021). A Survey of Transformers. arXiv."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"48","DOI":"10.1016\/j.neucom.2021.03.091","article-title":"A review on the attention mechanism of deep learning","volume":"452","author":"Niu","year":"2021","journal-title":"Neurocomputing"},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"5","DOI":"10.1023\/A:1010933404324","article-title":"Random forests","volume":"45","author":"Breiman","year":"2001","journal-title":"Mach. Learn."},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Konstantinov, A., Utkin, L., and Kirpichenko, S. (2022, January 27\u201329). AGBoost: Attention-based Modification of Gradient Boosting Machine. Proceedings of the 31st Conference of Open Innovations Association (FRUCT), Helsinki, Finland.","DOI":"10.23919\/FRUCT54823.2022.9770928"},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"346","DOI":"10.1016\/j.neunet.2022.07.029","article-title":"Attention-based Random Forest and Contamination Model","volume":"154","author":"Utkin","year":"2022","journal-title":"Neural Netw."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"1189","DOI":"10.1214\/aos\/1013203451","article-title":"Greedy function approximation: A gradient boosting machine","volume":"29","author":"Friedman","year":"2001","journal-title":"Ann. Stat."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"367","DOI":"10.1016\/S0167-9473(01)00065-2","article-title":"Stochastic gradient boosting","volume":"38","author":"Friedman","year":"2002","journal-title":"Comput. Stat. Data Anal."},{"key":"ref_11","unstructured":"Zhang, A., Lipton, Z., Li, M., and Smola, A. (2021). Dive into Deep Learning. arXiv."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"141","DOI":"10.1137\/1109020","article-title":"On estimating regression","volume":"9","author":"Nadaraya","year":"1964","journal-title":"Theory Probab. Appl."},{"key":"ref_13","first-page":"359","article-title":"Smooth regression analysis","volume":"26","author":"Watson","year":"1964","journal-title":"Sankhya Indian J. Stat. Ser. A"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Huber, P. (1981). Robust Statistics, Wiley.","DOI":"10.1002\/0471725250"},{"key":"ref_15","unstructured":"Utkin, L., and Konstantinov, A. (2022). Attention and Self-Attention in Random Forests. arXiv."},{"key":"ref_16","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A., Kaiser, L., and Polosukhin, I. (2017, January 4\u20139). Attention is all you need. Proceedings of the Advances in Neural Information Processing Systems, Long Beach, CA, USA."},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"3","DOI":"10.1007\/s10994-006-6226-1","article-title":"Extremely randomized trees","volume":"63","author":"Geurts","year":"2006","journal-title":"Mach. Learn."},{"key":"ref_18","unstructured":"Liu, F., Huang, X., Chen, Y., and Suykens, J. (2021). Random Features for Kernel Approximation: A Survey on Algorithms, Theory, and Beyond. arXiv."},{"key":"ref_19","unstructured":"Choromanski, K., Likhosherstov, V., Dohan, D., Song, X., Gane, A., Sarlos, T., Hawkins, P., Davis, J., Mohiuddin, A., and Kaiser, L. (2021, January 3\u20137). Rethinking Attention with Performers. Proceedings of the 2021 International Conference on Learning Representations, Vienna, Austria."},{"key":"ref_20","unstructured":"Choromanski, K., Chen, H., Lin, H., Ma, Y., Sehanobish, A., Jain, D., Ryoo, M., Varley, J., Zeng, A., and Likhosherstov, V. (2021). Hybrid Random Features. arXiv."},{"key":"ref_21","unstructured":"Ma, X., Kong, X., Wang, S., Zhou, C., May, J., Ma, H., and Zettlemoyer, L. (2021). Luna: Linear Unified Nested Attention. arXiv."},{"key":"ref_22","unstructured":"Schlag, I., Irie, K., and Schmidhuber, J. (2021, January 18\u201324). Linear transformers are secretly fast weight programmers. Proceedings of the International Conference on Machine Learning 2021. PMLR, Virtual."},{"key":"ref_23","unstructured":"Peng, H., Pappas, N., Yogatama, D., Schwartz, R., Smith, N., and Kong, L. (2021, January 3\u20137). Random Feature Attention. Proceedings of the International Conference on Learning Representations (ICLR 2021), Vienna, Austria."},{"key":"ref_24","unstructured":"Brauwers, G., and Frasincar, F. (2022). A General Survey on Attention Mechanisms in Deep Learning. arXiv."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Goncalves, T., Rio-Torto, I., Teixeira, L., and Cardoso, J. (2022). A survey on attention mechanisms for medical applications: Are we moving towards better algorithms?. arXiv.","DOI":"10.21203\/rs.3.rs-1594205\/v1"},{"key":"ref_26","unstructured":"Santana, A., and Colombini, E. (2021). Neural Attention Models in Deep Learning: Survey and Taxonomy. arXiv."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Soydaner, D. (2022). Attention Mechanism in Neural Networks: Where it Comes and Where it Goes. arXiv.","DOI":"10.1007\/s00521-022-07366-3"},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"33","DOI":"10.1007\/s41095-021-0247-3","article-title":"Transformers in computational visual media: A survey","volume":"8","author":"Xu","year":"2022","journal-title":"Comput. Vis. Media"},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"437","DOI":"10.1016\/j.jkss.2011.03.002","article-title":"A Weight-Adjusted Voting Algorithm for Ensemble of Classifiers","volume":"40","author":"Kim","year":"2011","journal-title":"J. Korean Stat. Soc."},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"36","DOI":"10.1007\/978-3-319-19369-4_4","article-title":"Random Forests with Weighted Voting for Anomalous Query Access Detection in Relational Databases","volume":"Volume 9120","author":"Ronao","year":"2015","journal-title":"Proceedings of the Artificial Intelligence and Soft Computing. ICAISC 2015"},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"963","DOI":"10.1142\/S0219622020500236","article-title":"A New Adaptive Weighted Deep Forest and its Modifications","volume":"19","author":"Utkin","year":"2020","journal-title":"Int. J. Inf. Technol. Decis."},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"496","DOI":"10.1002\/sam.11196","article-title":"A Weighted Random Forests Approach to Improve Predictive Performance","volume":"6","author":"Winham","year":"2013","journal-title":"Stat. Anal. Data Min."},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Xuan, S., Liu, G., and Li, Z. (2018, January 18\u201320). Refined Weighted Random Forest and Its Application to Credit Card Fraud Detection. Proceedings of the Computational Data and Social Networks, Shanghai, China.","DOI":"10.1007\/978-3-030-04648-4_29"},{"key":"ref_34","first-page":"1","article-title":"Weighted Random Forest Algorithm Based on Bayesian Algorithm","volume":"Volume 1924","author":"Zhang","year":"2021","journal-title":"Journal of Physics: Conference Series"},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"1215","DOI":"10.1002\/int.10143","article-title":"Building Classification Trees Using th Building Classification Trees Using the Total Uncertainty Criterion","volume":"18","author":"Abellan","year":"2003","journal-title":"Int. J. Intell. Syst."},{"key":"ref_36","doi-asserted-by":"crossref","first-page":"72","DOI":"10.1016\/j.knosys.2017.07.019","article-title":"A Random Forest approach using imprecise probabilities","volume":"134","author":"Abellan","year":"2017","journal-title":"Knowl.-Based Syst."},{"key":"ref_37","doi-asserted-by":"crossref","first-page":"228","DOI":"10.1016\/j.eswa.2017.12.029","article-title":"Increasing diversity in random forest learning algorithm via imprecise probabilities","volume":"97","author":"Abellan","year":"2018","journal-title":"Expert Syst. Appl."},{"key":"ref_38","doi-asserted-by":"crossref","first-page":"2514","DOI":"10.1016\/j.eswa.2013.09.050","article-title":"Analysis and extension of decision trees based on imprecise probabilities: Application on noisy data","volume":"41","author":"Mantas","year":"2014","journal-title":"Expert Syst. Appl."},{"key":"ref_39","first-page":"1","article-title":"Bagging of credal decision trees for imprecise classification","volume":"141","author":"Mantas","year":"2020","journal-title":"Expert Syst. Appl."},{"key":"ref_40","unstructured":"Cozman, F., Denoeux, T., Destercke, S., and Seidfenfeld, T. (2013, January 2\u20135). An imprecise boosting-like approach to regression. Proceedings of the ISIPTA \u201913, Proceedings of the Eighth International Symposium on Imprecise Probability: Theories and Applications, Compiegne, France."},{"key":"ref_41","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1016\/j.asoc.2020.106324","article-title":"Imprecise weighted extensions of random forests for classification and regression","volume":"92","author":"Utkin","year":"2020","journal-title":"Appl. Soft Comput."},{"key":"ref_42","unstructured":"Bahdanau, D., Cho, K., and Bengio, Y. (2014). Neural machine translation by jointly learning to align and translate. arXiv."},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Luong, T., Pham, H., and Manning, C. (2015, January 17\u201321). Effective approaches to attention-based neural machine translation. Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, Lisbon, Portugal.","DOI":"10.18653\/v1\/D15-1166"},{"key":"ref_44","unstructured":"Dua, D., and Graff, C. (2023, April 20). UCI Machine Learning Repository. Available online: http:\/\/archive.ics.uci.edu\/ml."},{"key":"ref_45","first-page":"1","article-title":"Statistical comparisons of classifiers over multiple data sets","volume":"7","author":"Demsar","year":"2006","journal-title":"J. Mach. Learn. Res."},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Walley, P. (1991). Statistical Reasoning with Imprecise Probabilities, Chapman and Hall.","DOI":"10.1007\/978-1-4899-3472-7"}],"container-title":["Informatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2227-9709\/10\/2\/40\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T19:25:26Z","timestamp":1760124326000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2227-9709\/10\/2\/40"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,4,28]]},"references-count":46,"journal-issue":{"issue":"2","published-online":{"date-parts":[[2023,6]]}},"alternative-id":["informatics10020040"],"URL":"https:\/\/doi.org\/10.3390\/informatics10020040","relation":{},"ISSN":["2227-9709"],"issn-type":[{"value":"2227-9709","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,4,28]]}}}