{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,6]],"date-time":"2026-05-06T16:15:11Z","timestamp":1778084111757,"version":"3.51.4"},"reference-count":57,"publisher":"Springer Science and Business Media LLC","issue":"3","license":[{"start":{"date-parts":[[2026,2,24]],"date-time":"2026-02-24T00:00:00Z","timestamp":1771891200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2026,2,24]],"date-time":"2026-02-24T00:00:00Z","timestamp":1771891200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Mach Learn"],"published-print":{"date-parts":[[2026,3]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>\n                    In this work, we present Learning to Learn with Optimal Transport for Unsupervised Scenarios (LOTUS), a simple yet effective method to perform model selection for multiple unsupervised machine learning (ML) tasks such as outlier detection and clustering. Our intuition behind this work is that a machine learning pipeline will perform well in a new dataset if it previously worked well on datasets with a\n                    <jats:italic>similar<\/jats:italic>\n                    underlying data distribution. We use Optimal Transport distances to find this similarity between unlabeled tabular datasets and recommend machine learning pipelines with one unified single method on two downstream unsupervised tasks: outlier detection and clustering. We present the effectiveness of our approach with experiments against strong baselines and show that LOTUS is a very promising first step toward model selection for multiple unsupervised ML tasks.\n                  <\/jats:p>","DOI":"10.1007\/s10994-025-06984-x","type":"journal-article","created":{"date-parts":[[2026,2,24]],"date-time":"2026-02-24T08:29:40Z","timestamp":1771921780000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Automated Machine Learning for Unsupervised Tabular Tasks"],"prefix":"10.1007","volume":"115","author":[{"given":"Prabhant","family":"Singh","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Pieter","family":"Gijsbers","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Elif Ceren Gok","family":"Yildirim","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Murat Onur","family":"Yildirim","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Joaquin","family":"Vanschoren","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2026,2,24]]},"reference":[{"key":"6984_CR1","unstructured":"Amos, B., Cohen, S., Luise, G. & Redko, I. (2023). Meta optimal transport. In ICML 2023."},{"key":"6984_CR2","doi-asserted-by":"crossref","unstructured":"Angiulli, F. & Pizzuti, C. (2002). Fast outlier detection in high dimensional spaces. In PKDD.","DOI":"10.1007\/3-540-45681-3_2"},{"issue":"2","key":"6984_CR3","doi-asserted-by":"publisher","first-page":"49","DOI":"10.1145\/304181.304187","volume":"28","author":"M Ankerst","year":"1999","unstructured":"Ankerst, M., Breunig, M. M., Kriegel, H.-P., & Sander, J. (1999). Optics: Ordering points to identify the clustering structure. ACM SIGMOD Record, 28(2), 49\u201360.","journal-title":"ACM SIGMOD Record"},{"key":"6984_CR4","unstructured":"Anonymous. (2024). HPOD: Hyperparameter optimization for unsupervised outlier detection. In AutoML 2024 methods track."},{"issue":"77","key":"6984_CR5","first-page":"1","volume":"18","author":"A Benavoli","year":"2017","unstructured":"Benavoli, A., Corani, G., Dem\u0161ar, J., & Zaffalon, M. (2017). Time for a change: A tutorial for comparing multiple classifiers through Bayesian analysis. Journal of Machine Learning Research, 18(77), 1\u201336.","journal-title":"Journal of Machine Learning Research"},{"issue":"null","key":"6984_CR6","first-page":"993","volume":"3","author":"DM Blei","year":"2003","unstructured":"Blei, D. M., Ng, A. Y., & Jordan, M. I. (2003). Latent Dirichlet allocation. Journal of Machine Learning Research, 3(null), 993\u20131022.","journal-title":"Journal of Machine Learning Research"},{"issue":"2","key":"6984_CR7","doi-asserted-by":"publisher","first-page":"93","DOI":"10.1145\/335191.335388","volume":"29","author":"MM Breunig","year":"2000","unstructured":"Breunig, M. M., Kriegel, H.-P., Ng, R. T., & Sander, J. (2000). LOF: Identifying density-based local outliers. SIGMOD Record, 29(2), 93\u2013104. https:\/\/doi.org\/10.1145\/335191.335388","journal-title":"SIGMOD Record"},{"issue":"1","key":"6984_CR8","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1080\/03610927408827101","volume":"3","author":"T Cali\u0144ski","year":"1974","unstructured":"Cali\u0144ski, T., & Harabasz, J. (1974). A dendrite method for cluster analysis. Communications in Statistics-Theory and Methods, 3(1), 1\u201327.","journal-title":"Communications in Statistics-Theory and Methods"},{"key":"6984_CR9","unstructured":"Cuturi, M. (2013). Sinkhorn distances: Lightspeed computation of optimal transport. In NIPS."},{"issue":"1","key":"6984_CR10","first-page":"1","volume":"7","author":"J Dem\u0161ar","year":"2006","unstructured":"Dem\u0161ar, J. (2006). Statistical comparisons of classifiers over multiple data sets. Journal of Machine Learning Research, 7(1), 1\u201330.","journal-title":"Journal of Machine Learning Research"},{"key":"6984_CR11","unstructured":"Ester, M., Kriegel, H.-P., Sander, J., & Xu, X. (1996). A density-based algorithm for discovering clusters in large spatial databases with noise. In Proceedings of the second international conference on knowledge discovery and data mining."},{"key":"6984_CR12","unstructured":"Feurer, M., Eggensperger, K., Falkner, S., Lindauer, M. & Hutter, F. (2020). Auto-Sklearn 2.0: Hands-free AutoML via meta-learning. arXiv:2007.04074 [cs.LG]."},{"key":"6984_CR13","unstructured":"Feurer, M., Klein, A., Eggensperger, K., Springenberg, J., Blum, M., & Hutter, F. (2015). Efficient and robust automated machine learning. In Advances in neural information processing systems."},{"issue":"5814","key":"6984_CR14","doi-asserted-by":"publisher","first-page":"972","DOI":"10.1126\/science.1136800","volume":"315","author":"BJ Frey","year":"2007","unstructured":"Frey, B. J., & Dueck, D. (2007). Clustering by passing messages between data points. Science, 315(5814), 972\u2013976. https:\/\/doi.org\/10.1126\/science.1136800","journal-title":"Science"},{"key":"6984_CR15","doi-asserted-by":"publisher","unstructured":"Gijsbers, P. & Vanschoren, J. (2021). GAMA: A general automated machine learning assistant. Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics) 12461 LNAI (pp. 560\u2013564). https:\/\/doi.org\/10.1007\/978-3-030-67670-4_39","DOI":"10.1007\/978-3-030-67670-4_39"},{"key":"6984_CR16","doi-asserted-by":"crossref","unstructured":"Gijsbers, P., & Vanschoren, J. (2021). GAMA, A general automated machine learning assistant. In Machine learning and knowledge discovery in databases. Applied Data Science and Demo Track.","DOI":"10.1007\/978-3-030-67670-4_39"},{"key":"6984_CR17","doi-asserted-by":"publisher","unstructured":"Goix, N. (2016). How to evaluate the quality of unsupervised anomaly detection algorithms? https:\/\/doi.org\/10.48550\/ARXIV.1607.01152. arXiv:1607.01152","DOI":"10.48550\/ARXIV.1607.01152"},{"key":"6984_CR18","unstructured":"Goldstein, M. & Dengel, A. R. (2012). Histogram-based outlier score (HBOS): A fast unsupervised anomaly detection algorithm."},{"key":"6984_CR19","doi-asserted-by":"crossref","unstructured":"Han, S., Hu, X., Huang, H., Jiang, M., & Zhao, Y. (2022). ADBench: Anomaly detection benchmark. In Thirty-sixth conference on neural information processing systems datasets and benchmarks track.","DOI":"10.52202\/068431-2329"},{"key":"6984_CR20","doi-asserted-by":"crossref","unstructured":"Hutter, F., Kotthoff, L. & Vanschoren, J. (2019). Automated machine learning: Methods, systems, challenges. Automated Machine Learning.","DOI":"10.1007\/978-3-030-05318-5"},{"issue":"4\u20135","key":"6984_CR21","doi-asserted-by":"publisher","first-page":"411","DOI":"10.1016\/S0893-6080(00)00026-5","volume":"13","author":"A Hyv\u00e4rinen","year":"2000","unstructured":"Hyv\u00e4rinen, A., & Oja, E. (2000). Independent component analysis: Algorithms and applications. Neural Networks, 13(4\u20135), 411\u201330.","journal-title":"Neural Networks"},{"key":"6984_CR22","doi-asserted-by":"publisher","unstructured":"Kriegel, H.-P., Schubert, M., & Zimek, A. (2008). Angle-based outlier detection in high-dimensional data. In Proceedings of the 14th ACM SIGKDD international conference on knowledge discovery and data mining (pp. 444\u2013452). https:\/\/doi.org\/10.1145\/1401890.1401946","DOI":"10.1145\/1401890.1401946"},{"key":"6984_CR23","unstructured":"Li, L., Jamieson, K., Rostamizadeh, A., Gonina, E., Ben-tzur, J., Hardt, M., Recht, B., & Talwalkar, A. (2020). A system for massively parallel hyperparameter tuning. In: Dhillon, I., Papailiopoulos, D., Sze, V. (Eds.), Proceedings of machine learning and systems (pp. 230\u2013246)."},{"key":"6984_CR24","doi-asserted-by":"crossref","unstructured":"Li, Y., Zha, D., Zou, N., & Hu, X. (2020). PyODDs: An end-to-end outlier detection system with automated machine learning. In Companion proceedings of the web conference 2020.","DOI":"10.1145\/3366424.3383530"},{"key":"6984_CR25","doi-asserted-by":"crossref","unstructured":"Liao, M., Li, Y., Kianifard, F., Obi, E. N., & Arcona, S. (2016). Cluster analysis and its application to healthcare claims data: A study of end-stage renal disease patients who initiated hemodialysis. BMC Nephrology, 17.","DOI":"10.1186\/s12882-016-0238-2"},{"key":"6984_CR26","doi-asserted-by":"crossref","unstructured":"Liu, Y., Li, S. & Tian, W. (2021). AutoCluster: Meta-learning based ensemble method for automated unsupervised clustering. In Pacific-Asia conference on knowledge discovery and data mining.","DOI":"10.1007\/978-3-030-75768-7_20"},{"key":"6984_CR27","doi-asserted-by":"crossref","unstructured":"Liu, Y., Li, S., & Tian, W. (2021). Autocluster. Meta-learning based ensemble method for automated unsupervised clustering. In Advances in knowledge discovery and data mining: 25th Pacific-Asia conference, PAKDD 2021 (pp. 246\u2013258). Springer.","DOI":"10.1007\/978-3-030-75768-7_20"},{"key":"6984_CR28","doi-asserted-by":"crossref","unstructured":"Liu, F. T., Ting, K. M., & Zhou, Z.-H. (2008). Isolation forest. In 2008 Eighth IEEE international conference on data mining (pp. 413\u2013422).","DOI":"10.1109\/ICDM.2008.17"},{"issue":"2","key":"6984_CR29","doi-asserted-by":"publisher","first-page":"129","DOI":"10.1109\/TIT.1982.1056489","volume":"28","author":"S Lloyd","year":"1982","unstructured":"Lloyd, S. (1982). Least squares quantization in PCM. IEEE Transactions on Information Theory, 28(2), 129\u2013137.","journal-title":"IEEE Transactions on Information Theory"},{"key":"6984_CR30","unstructured":"Ma, M. Q., Zhao, Y., Zhang, X. & Akoglu, L. (2021). A large-scale study on unsupervised outlier model selection: Do internal strategies suffice? CoRR arXiv:2104.01422"},{"key":"6984_CR31","doi-asserted-by":"crossref","unstructured":"Marques, H. O., Campello, R. J. G. B., Zimek, A., & Sander, J. (2015). On the internal evaluation of unsupervised outlier detection. In Proceedings of the 27th international conference on scientific and statistical database management.","DOI":"10.1145\/2791347.2791352"},{"key":"6984_CR32","unstructured":"M\u00e9moli, F., Sidiropoulos, A., & Singhal, K. (2018). Sketching and clustering metric measure spaces. arXiv:1801.00551"},{"key":"6984_CR33","doi-asserted-by":"publisher","first-page":"417","DOI":"10.1007\/s10208-011-9093-5","volume":"11","author":"F M\u00e9moli","year":"2011","unstructured":"M\u00e9moli, F. (2011). Gromov\u2013Wasserstein distances and the metric approach to object matching. Foundations of Computational Mathematics, 11, 417\u2013487.","journal-title":"Foundations of Computational Mathematics"},{"key":"6984_CR34","first-page":"2825","volume":"12","author":"F Pedregosa","year":"2011","unstructured":"Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., Vanderplas, J., Passos, A., Cournapeau, D., Brucher, M., Perrot, M., & Duchesnay, E. (2011). Scikit-learn: Machine learning in Python. Journal of Machine Learning Research, 12, 2825\u20132830.","journal-title":"Journal of Machine Learning Research"},{"key":"6984_CR35","doi-asserted-by":"publisher","first-page":"275","DOI":"10.1007\/s10994-015-5521-0","volume":"102","author":"T Pevn\u00fd","year":"2015","unstructured":"Pevn\u00fd, T. (2015). LODA: Lightweight on-line detector of anomalies. Machine Learning, 102, 275\u2013304.","journal-title":"Machine Learning"},{"key":"6984_CR36","unstructured":"Peyr\u00e9, G., Cuturi, M., & Solomon, J. (2016). Gromov\u2013Wasserstein averaging of kernel and distance matrices. In Proceedings of the 33rd international conference on machine learning."},{"key":"6984_CR37","doi-asserted-by":"crossref","unstructured":"Poulakis, Y., Doulkeridis, C., & Kyriazis, D. (2020). Autoclust. A framework for automated clustering based on cluster validity indices. In 2020 IEEE international conference on data mining (ICDM) (pp. 1220\u20131225). IEEE.","DOI":"10.1109\/ICDM50108.2020.00153"},{"key":"6984_CR38","unstructured":"Scetbon, M., & Cuturi, M. (2022). Low-rank optimal transport: Approximation, statistics and debiasing. NeurIPS 2022 arXiv:2205.12365."},{"key":"6984_CR39","unstructured":"Scetbon, M., Cuturi, M., & Peyr\u00e9, G. (2021). Low-rank Sinkhorn factorization. In Proceedings of the 38th international conference on machine learning."},{"key":"6984_CR40","unstructured":"Scetbon, M., Peyr\u00e9, G., & Cuturi, M. (2022). Linear-time Gromov Wasserstein distances using low rank couplings and costs. In Proceedings of the 39th international conference on machine learning."},{"key":"6984_CR41","unstructured":"Sch\u00f6lkopf, B., Williamson, R. C., Smola, A., Shawe-Taylor, J. & Platt, J. C. (1999). Support vector method for novelty detection. In NIPS."},{"key":"6984_CR42","doi-asserted-by":"crossref","unstructured":"Sculley, D. (2010). Web-scale k-means clustering. In Proceedings of the 19th international conference on world wide web. WWW \u201910 (pp. 1177\u20131178). Association for Computing Machinery.","DOI":"10.1145\/1772690.1772862"},{"key":"6984_CR43","unstructured":"Shchur, O., Turkmen, C., Erickson, N., Shen, H., Shirkov, A., Hu, T., & Wang, B. (2023). Autogluon-timeseries: AutoML for probabilistic time series forecasting. In AutoML conference."},{"key":"6984_CR44","doi-asserted-by":"crossref","unstructured":"Singh, P. & Vanschoren, J. (2023). AutoML for outlier detection with optimal transport distances. In Proceedings of the thirty-second international joint conference on artificial intelligence, IJCAI-23. Demo Track.","DOI":"10.24963\/ijcai.2023\/843"},{"key":"6984_CR45","doi-asserted-by":"crossref","unstructured":"Stern, D., Herbrich, R., Graepel, T., Samulowitz, H., Pulina, L., & Tacchella, A. (2010). Collaborative expert portfolio management. In Proceedings of the twenty-fourth AAAI conference on artificial intelligence AAAI-10(to appear).","DOI":"10.1609\/aaai.v24i1.7561"},{"key":"6984_CR46","doi-asserted-by":"crossref","unstructured":"Tang, J., Chen, Z., Fu, A.W.-C., & Cheung, D.W.-L. (2002). Enhancing effectiveness of outlier detections for low density patterns. In Pacific-Asia conference on knowledge discovery and data mining.","DOI":"10.1007\/3-540-47887-6_53"},{"key":"6984_CR47","doi-asserted-by":"publisher","DOI":"10.1145\/3589289","author":"D Treder-Tschechlov","year":"2023","unstructured":"Treder-Tschechlov, D., Fritz, M., Schwarz, H., & Mitschang, B. (2023). ML2DAC: Meta-learning to democratize AutoML for clustering analysis. Proceedings of the ACM on Management of Data. https:\/\/doi.org\/10.1145\/3589289","journal-title":"Proceedings of the ACM on Management of Data"},{"key":"6984_CR48","unstructured":"Tschechlov, D., Fritz, M., Schwarz, H., Velegrakis, Y., Zeinalipour-Yazti, D., Chrysanthis, P. & Guerra, F. (2021). AutoML4CLUST: Efficient AutoML for clustering analyses. In EDBT (pp. 343\u2013348)."},{"key":"6984_CR49","doi-asserted-by":"crossref","unstructured":"Vanschoren, J. (2018). Meta-learning: A survey. arXiv:1810.03548","DOI":"10.1007\/978-3-030-05318-5_2"},{"issue":"2","key":"6984_CR50","doi-asserted-by":"publisher","first-page":"49","DOI":"10.1145\/2641190.2641198","volume":"15","author":"J Vanschoren","year":"2013","unstructured":"Vanschoren, J., Rijn, J. N., Bischl, B., & Torgo, L. (2013). OpenML: Networked science in machine learning. SIGKDD Explorations, 15(2), 49\u201360. https:\/\/doi.org\/10.1145\/2641190.2641198","journal-title":"SIGKDD Explorations"},{"key":"6984_CR51","doi-asserted-by":"crossref","unstructured":"Villani, C. (2008). Optimal transport: Old and new. In Optimal transport.","DOI":"10.1007\/978-3-540-71050-9"},{"issue":"95","key":"6984_CR52","first-page":"2837","volume":"11","author":"NX Vinh","year":"2010","unstructured":"Vinh, N. X., Epps, J., & Bailey, J. (2010). Information theoretic measures for clusterings comparison: Variants, properties, normalization and correction for chance. Journal of Machine Learning Research, 11(95), 2837\u20132854.","journal-title":"Journal of Machine Learning Research"},{"issue":"95","key":"6984_CR53","first-page":"2837","volume":"11","author":"NX Vinh","year":"2010","unstructured":"Vinh, N. X., Epps, J., & Bailey, J. (2010). Information theoretic measures for clusterings comparison: Variants, properties, normalization and correction for chance. Journal of Machine Learning Research, 11(95), 2837\u20132854.","journal-title":"Journal of Machine Learning Research"},{"key":"6984_CR54","unstructured":"Vu, L., Kirchner, P., Aggarwal, C. C., & Samulowitz, H. (2024). Instance-level metalearning for outlier detection. In Proceedings of the thirty-third international joint conference on artificial intelligence, IJCAI-24."},{"key":"6984_CR55","unstructured":"Wang, C., Wu, Q., Weimer, M. & Zhu, E. (2021). FlaML: A fast and lightweight AutoML library. In MLSys."},{"key":"6984_CR56","unstructured":"Zhao, Y., Rossi, R., & Akoglu, L. (2021). Automatic unsupervised outlier model selection. In M. Ranzato, A. Beygelzimer, Y. Dauphin, P. S. Liang, J. W. Vaughan (Eds.), Advances in neural information processing systems."},{"key":"6984_CR57","first-page":"96","volume":"20","author":"Y Zhao","year":"2019","unstructured":"Zhao, Y., Nasrullah, Z., & Li, Z. (2019). PYOD: A python toolbox for scalable outlier detection. Journal of Machine Learning Research, 20, 96\u20131967.","journal-title":"Journal of Machine Learning Research"}],"container-title":["Machine Learning"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10994-025-06984-x.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10994-025-06984-x","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10994-025-06984-x.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,5,6]],"date-time":"2026-05-06T15:35:17Z","timestamp":1778081717000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10994-025-06984-x"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,2,24]]},"references-count":57,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2026,3]]}},"alternative-id":["6984"],"URL":"https:\/\/doi.org\/10.1007\/s10994-025-06984-x","relation":{},"ISSN":["0885-6125","1573-0565"],"issn-type":[{"value":"0885-6125","type":"print"},{"value":"1573-0565","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,2,24]]},"assertion":[{"value":"30 May 2025","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"23 September 2025","order":2,"name":"revised","label":"Revised","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"18 December 2025","order":3,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"24 February 2026","order":4,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare no Conflict of interest.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}}],"article-number":"39"}}