{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,30]],"date-time":"2026-07-30T11:05:10Z","timestamp":1785409510029,"version":"3.56.0"},"reference-count":16,"publisher":"Springer Science and Business Media LLC","issue":"7","license":[{"start":{"date-parts":[[2026,7,1]],"date-time":"2026-07-01T00:00:00Z","timestamp":1782864000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2026,7,11]],"date-time":"2026-07-11T00:00:00Z","timestamp":1783728000000},"content-version":"vor","delay-in-days":10,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"name":"Science Committee of the Ministry of Science and Higher Education of the Republic of Kazakhstan","award":["AP26198325"],"award-info":[{"award-number":["AP26198325"]}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Mach Learn"],"published-print":{"date-parts":[[2026,7]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>\n                    We show that for unconstrained Deep Linear Discriminant Analysis (LDA) classifiers, maximum-likelihood training admits pathological solutions in which class means drift together, covariances collapse, and the learned representation becomes almost non-discriminative. Conversely, cross-entropy training yields excellent accuracy but decouples the head from the underlying generative model, leading to highly inconsistent parameter estimates. To reconcile generative structure with discriminative performance, we introduce the\n                    <jats:italic>Discriminative Negative Log-Likelihood<\/jats:italic>\n                    (DNLL) loss, which augments the LDA log-likelihood with a simple penalty on the mixture density. DNLL can be interpreted as standard LDA NLL plus a term that explicitly discourages regions where several classes are simultaneously likely. Deep LDA trained with DNLL produces clean, well-separated latent spaces, matches the test accuracy of softmax classifiers on synthetic data and standard image benchmarks, and yields substantially better calibrated predictive probabilities.\n                  <\/jats:p>","DOI":"10.1007\/s10994-026-07097-9","type":"journal-article","created":{"date-parts":[[2026,7,11]],"date-time":"2026-07-11T09:15:53Z","timestamp":1783761353000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Deep Linear Discriminant Analysis Revisited"],"prefix":"10.1007","volume":"115","author":[{"given":"Maxat","family":"Tezekbayev","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Rustem","family":"Takhanov","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Arman","family":"Bolatov","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Zhenisbek","family":"Assylbekov","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2026,7,11]]},"reference":[{"key":"7097_CR1","doi-asserted-by":"publisher","DOI":"10.1016\/j.amc.2020.125578","volume":"389","author":"A-M Acu","year":"2021","unstructured":"Acu, A.-M., Ba\u015fcanbaz-Tunca, G., & Rasa, I. (2021). Information potential for some probability density functions. Applied Mathematics and Computation, 389, Article 125578. https:\/\/doi.org\/10.1016\/j.amc.2020.125578","journal-title":"Applied Mathematics and Computation"},{"key":"7097_CR2","doi-asserted-by":"publisher","unstructured":"Dorfer, M., Kelz, R.,& Widmer, G. (2016). Deep linear discriminant analysis. In: Bengio, Y., LeCun, Y. (eds.) 4th International Conference on Learning Representations, ICLR 2016, San Juan, Puerto Rico, May 2-4, 2016, Conference Track Proceedings. https:\/\/doi.org\/10.48550\/arXiv.1511.04707","DOI":"10.48550\/arXiv.1511.04707"},{"issue":"2","key":"7097_CR3","doi-asserted-by":"publisher","first-page":"179","DOI":"10.1111\/j.1469-1809.1936.tb02137.x","volume":"7","author":"RA Fisher","year":"1936","unstructured":"Fisher, R. A. (1936). The use of multiple measurements in taxonomic problems. Annals of eugenics, 7(2), 179\u2013188. https:\/\/doi.org\/10.1111\/j.1469-1809.1936.tb02137.x","journal-title":"Annals of eugenics"},{"key":"7097_CR4","doi-asserted-by":"publisher","unstructured":"Guo, C., Pleiss, G., Sun, Y.,& Weinberger, K.Q. (2017). On calibration of modern neural networks. In: Precup, D., Teh, Y.W. (eds.) Proceedings of the 34th International Conference on Machine Learning. Proceedings of Machine Learning Research, vol. 70, pp. 1321\u20131330 . https:\/\/doi.org\/10.48550\/arXiv.1706.04599","DOI":"10.48550\/arXiv.1706.04599"},{"issue":"1","key":"7097_CR5","doi-asserted-by":"publisher","first-page":"155","DOI":"10.1111\/j.2517-6161.1996.tb02073.x","volume":"58","author":"T Hastie","year":"1996","unstructured":"Hastie, T., & Tibshirani, R. (1996). Discriminant analysis by gaussian mixtures. Journal of the Royal Statistical Society Series B: Statistical Methodology, 58(1), 155\u2013176. https:\/\/doi.org\/10.1111\/j.2517-6161.1996.tb02073.x","journal-title":"Journal of the Royal Statistical Society Series B: Statistical Methodology"},{"key":"7097_CR6","doi-asserted-by":"publisher","unstructured":"Kingma, D.P.,& Ba, J. (2015). Adam: A method for stochastic optimization. In: Bengio, Y., LeCun, Y. (eds.) 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceedings. https:\/\/doi.org\/10.48550\/arXiv.1412.6980","DOI":"10.48550\/arXiv.1412.6980"},{"key":"7097_CR7","unstructured":"Krizhevsky, A.,& Hinton, G. (2009). Learning multiple layers of features from tiny images. Technical report, University of Toronto, Toronto, Ontario, Canada. https:\/\/www.cs.toronto.edu\/%7Ekriz\/learning-features-2009-TR.pdf"},{"key":"7097_CR8","doi-asserted-by":"publisher","unstructured":"Mukhoti, J., Kulharia, V., Sanyal, A., Golodetz, S., Torr, P., & Dokania, P. (2020). Calibrating deep neural networks using focal loss. In: Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M.F., Lin, H. (eds.) Advances in Neural Information Processing Systems, 33, 15288\u201315299. https:\/\/doi.org\/10.48550\/arXiv.2002.09437","DOI":"10.48550\/arXiv.2002.09437"},{"key":"7097_CR9","doi-asserted-by":"publisher","unstructured":"Mika, S., Ratsch, G., Weston, J., Scholkopf, B., & Mullers, K.-R. (1999). Fisher discriminant analysis with kernels. In: Neural Networks for Signal Processing IX: Proceedings of the 1999 IEEE Signal Processing Society Workshop (cat. No. 98th8468), 41\u201348. https:\/\/doi.org\/10.1109\/NNSP.1999.788121","DOI":"10.1109\/NNSP.1999.788121"},{"issue":"2","key":"7097_CR10","doi-asserted-by":"publisher","first-page":"159","DOI":"10.1111\/j.2517-6161.1948.tb00008.x","volume":"10","author":"CR Rao","year":"1948","unstructured":"Rao, C. R. (1948). The utilization of multiple measurements in problems of biological classification. Journal of the Royal Statistical Society. Series B (Methodological), 10(2), 159\u2013203. https:\/\/doi.org\/10.1111\/j.2517-6161.1948.tb00008.x","journal-title":"Journal of the Royal Statistical Society. Series B (Methodological)"},{"issue":"1","key":"7097_CR11","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1007\/BF02289451","volume":"31","author":"PH Sch\u00f6nemann","year":"1966","unstructured":"Sch\u00f6nemann, P. H. (1966). A generalized solution of the orthogonal procrustes problem. Psychometrika, 31(1), 1\u201310. https:\/\/doi.org\/10.1007\/BF02289451","journal-title":"Psychometrika"},{"key":"7097_CR12","doi-asserted-by":"publisher","unstructured":"Snell, J., Swersky, K.,& Zemel, R. (2017). Prototypical networks for few-shot learning. In: Guyon, I., Luxburg, U.V., Bengio, S., Wallach, H., Fergus, R., Vishwanathan, S., Garnett, R. (eds.) Advances in Neural Information Processing Systems, vol. 30, pp. 4080\u20134090. https:\/\/doi.org\/10.48550\/arXiv.1703.05175","DOI":"10.48550\/arXiv.1703.05175"},{"key":"7097_CR13","unstructured":"Vasilev, R.,& D\u2019yakonov, A. (2023). Calibration of neural networks. arXiv preprint arxiv:2303.10761"},{"key":"7097_CR14","unstructured":"Xiao, H., Rasul, K.,& Vollgraf, R. (2017). Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms. arXiv preprint arxiv:1708.07747"},{"issue":"11","key":"7097_CR15","doi-asserted-by":"publisher","first-page":"15101","DOI":"10.1007\/s11042-018-6855-y","volume":"78","author":"L Yan","year":"2019","unstructured":"Yan, L., Lu, H., Wang, C., Ye, Z., Chen, H., & Ling, H. (2019). Deep linear discriminant analysis hashing for image retrieval. Multimedia Tools and Applications, 78(11), 15101\u201315119. https:\/\/doi.org\/10.1007\/s11042-018-6855-y","journal-title":"Multimedia Tools and Applications"},{"key":"7097_CR16","doi-asserted-by":"publisher","unstructured":"Zhang, Z., Lu, W., Feng, X., Cao, J., & Xie, G. (2022). A discriminative feature learning approach with distinguishable distance metrics for remote sensing image classification and retrieval. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing,16, 889\u2013901. https:\/\/doi.org\/10.1109\/JSTARS.2022.3233032","DOI":"10.1109\/JSTARS.2022.3233032"}],"container-title":["Machine Learning"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10994-026-07097-9.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10994-026-07097-9","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10994-026-07097-9.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,7,30]],"date-time":"2026-07-30T10:25:49Z","timestamp":1785407149000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10994-026-07097-9"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,7]]},"references-count":16,"journal-issue":{"issue":"7","published-print":{"date-parts":[[2026,7]]}},"alternative-id":["7097"],"URL":"https:\/\/doi.org\/10.1007\/s10994-026-07097-9","relation":{},"ISSN":["0885-6125","1573-0565"],"issn-type":[{"value":"0885-6125","type":"print"},{"value":"1573-0565","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,7]]},"assertion":[{"value":"4 January 2026","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"27 May 2026","order":2,"name":"revised","label":"Revised","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"3 June 2026","order":3,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"11 July 2026","order":4,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"The authors declare no competing interests.","order":1,"name":"Ethics","label":"Competing Interests","group":{"name":"EthicsHeading","label":"Declarations"}}],"article-number":"175"}}