{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,20]],"date-time":"2026-05-20T23:32:14Z","timestamp":1779319934040,"version":"3.51.4"},"reference-count":38,"publisher":"Springer Science and Business Media LLC","issue":"11-12","license":[{"start":{"date-parts":[[2024,10,21]],"date-time":"2024-10-21T00:00:00Z","timestamp":1729468800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2024,10,21]],"date-time":"2024-10-21T00:00:00Z","timestamp":1729468800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"name":"EPFL Lausanne"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Mach Learn"],"published-print":{"date-parts":[[2024,12]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>The ability to collect and store ever more massive data, unlabeled in many cases, has been accompanied by the need to process them efficiently in order to extract relevant information and possibly design solutions based on the latter. In various situations, the vast majority of the observations exhibit the same behavior, while a small proportion deviates from it. Detecting these outlier observations (or equivalently defined as anomalies) is now one of the major challenges for machine learning applications (e.g. fraud detection or predictive maintenance). We propose here a novel methodology for outlier\/anomaly detection, by learning a scoring function defined on the feature space allowing for ranking the observations by degree of abnormality. The scoring function is built through maximization of an empirical performance criterion taking the form of a (two-sample) linear rank statistic. We show that bipartite ranking algorithms can thus be used to learn nearly optimal scoring function with provable theoretical guarantees. We illustrate our methodology with numerical experiments based on open access online code.<\/jats:p>","DOI":"10.1007\/s10994-024-06609-9","type":"journal-article","created":{"date-parts":[[2024,10,21]],"date-time":"2024-10-21T20:02:08Z","timestamp":1729540928000},"page":"8623-8653","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":2,"title":["Learning to rank anomalies: scalar performance criteria and maximization of rank statistics"],"prefix":"10.1007","volume":"113","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-9365-4801","authenticated-orcid":false,"given":"Myrto","family":"Limnios","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Nathan","family":"Noiry","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Stephan","family":"Cl\u00e9men\u00e7on","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2024,10,21]]},"reference":[{"issue":"49","key":"6609_CR2","first-page":"1653","volume":"15","author":"S Agarwal","year":"2014","unstructured":"Agarwal, S. (2014). Surrogate regret bounds for bipartite ranking via strongly proper losses. Journal of Machine Learning Research, 15(49), 1653\u20131674.","journal-title":"Journal of Machine Learning Research"},{"key":"6609_CR3","unstructured":"Ailon, N., Mohri, M. (2008). An efficient reduction of ranking to classification. In 21st Annual Conference on Learning Theory - COLT, 87\u201398"},{"key":"6609_CR4","unstructured":"Bergman, L., Hoshen, Y. (2020). Classification-Based Anomaly Detection for General Data. arXiv:2005.02359"},{"key":"6609_CR5","doi-asserted-by":"publisher","unstructured":"Bock, R. (2007). MAGIC Gamma Telescope. https:\/\/doi.org\/10.24432\/C52C8B","DOI":"10.24432\/C52C8B"},{"key":"6609_CR6","doi-asserted-by":"crossref","unstructured":"Breunig, M.M., Kriegel, H.P., Ng, R.T., Sander, J. (2000). Lof: identifying density-based local outliers. In ACM Sigmod Record, vol. 29, pp. 93\u2013104","DOI":"10.1145\/335191.335388"},{"key":"6609_CR7","doi-asserted-by":"crossref","unstructured":"Chandola, V., Banerjee, A., Kumar, V. (2009). Anomaly detection: A survey. ACM Computing Surveys 41(3)","DOI":"10.1145\/1541880.1541882"},{"key":"6609_CR8","unstructured":"Cl\u00e9men\u00e7on, S., Jakubowicz, J. (2013) Scoring anomalies: a M-estimation formulation. In: Proceedings of the Sixteenth International Conference on Artificial Intelligence and Statistics, 31, 659\u2013667"},{"key":"6609_CR9","unstructured":"Cl\u00e9men\u00e7on, S., Robbiano, S. (2014). Anomaly ranking as supervised bipartite ranking. In Proceedings of the 31st International Conference on International Conference on Machine Learning. ICML\u201914, 32, 343\u2013351"},{"issue":"2","key":"6609_CR10","doi-asserted-by":"publisher","first-page":"2806","DOI":"10.1214\/18-EJS1474","volume":"12","author":"S Cl\u00e9men\u00e7on","year":"2018","unstructured":"Cl\u00e9men\u00e7on, S., & Thomas, A. (2018). Mass volume curves and anomaly ranking. Electronic Journal of Statistics, 12(2), 2806\u20132872.","journal-title":"Electronic Journal of Statistics"},{"issue":"9","key":"6609_CR11","doi-asserted-by":"publisher","first-page":"4316","DOI":"10.1109\/TIT.2009.2025558","volume":"55","author":"S Cl\u00e9men\u00e7on","year":"2009","unstructured":"Cl\u00e9men\u00e7on, S., & Vayatis, N. (2009). Tree-based ranking methods. IEEE Transactions on Information Theory, 55(9), 4316\u20134336.","journal-title":"IEEE Transactions on Information Theory"},{"issue":"2","key":"6609_CR12","doi-asserted-by":"publisher","first-page":"844","DOI":"10.1214\/009052607000000910","volume":"36","author":"S Cl\u00e9men\u00e7on","year":"2008","unstructured":"Cl\u00e9men\u00e7on, S., Lugosi, G., & Vayatis, N. (2008). Ranking and empirical risk minimization of U-statistics. The Annals of Statistics, 36(2), 844\u2013874.","journal-title":"The Annals of Statistics"},{"key":"6609_CR13","first-page":"39","volume":"14","author":"S Cl\u00e9men\u00e7on","year":"2013","unstructured":"Cl\u00e9men\u00e7on, S., Depecker, M., & Vayatis, N. (2013). Ranking Forests. Journal of Machine Learning Research, 14, 39\u201373.","journal-title":"Journal of Machine Learning Research"},{"issue":"2","key":"6609_CR14","first-page":"39","volume":"14","author":"S Cl\u00e9men\u00e7on","year":"2013","unstructured":"Cl\u00e9men\u00e7on, S., Depecker, M., & Vayatis, N. (2013). Ranking forests. Journal of Machine Learning Research, 14(2), 39\u201373.","journal-title":"Journal of Machine Learning Research"},{"key":"6609_CR15","doi-asserted-by":"crossref","unstructured":"Cl\u00e9men\u00e7on, S., Baskiotis, N., Vayatis, N. (2016). Anomaly Ranking in a High Dimensional Space: The Unsupervised TreeRank Algorithm. In Unsupervised Learning Algorithms, pp. 33\u201354","DOI":"10.1007\/978-3-319-24211-8_2"},{"issue":"2","key":"6609_CR16","doi-asserted-by":"publisher","first-page":"4659","DOI":"10.1214\/21-EJS1907","volume":"15","author":"S Cl\u00e9men\u00e7on","year":"2021","unstructured":"Cl\u00e9men\u00e7on, S., Limnios, M., & Vayatis, N. (2021). Concentration inequalities for two-sample rank processes with application to bipartite ranking. Electronic Journal of Statistics, 15(2), 4659\u20134717.","journal-title":"Electronic Journal of Statistics"},{"key":"6609_CR17","unstructured":"Cl\u00e9men\u00e7on, S., Limnios, M., Vayatis, N. (2023). A bipartite ranking approach to the two-sample problem arXiv:2302.03592"},{"key":"6609_CR18","doi-asserted-by":"publisher","unstructured":"Cortez, P., Cerdeira, A., Almeida, F., Matos, T., Reis, J. (2009). Wine Quality. https:\/\/doi.org\/10.24432\/C56S3T","DOI":"10.24432\/C56S3T"},{"key":"6609_CR19","doi-asserted-by":"crossref","unstructured":"Craswell, N. (2009). Precision at n, pp. 2127\u20132128. Springer, Boston, MA","DOI":"10.1007\/978-0-387-39940-9_484"},{"key":"6609_CR20","doi-asserted-by":"crossref","unstructured":"Dang, X.H., Micenkov\u00e1, B., Assent, I., Ng, R.T. (2013). Local outlier detection with interpretation. In Machine Learning and Knowledge Discovery in Databases, pp. 304\u2013320. Springer, Berlin, Heidelberg","DOI":"10.1007\/978-3-642-40994-3_20"},{"key":"6609_CR21","doi-asserted-by":"crossref","unstructured":"Devroye, L., Gy\u00f6rfi, L., Lugosi, G. (1996). A probabilistic theory of pattern recognition. Springer New-York","DOI":"10.1007\/978-1-4612-0711-5"},{"key":"6609_CR22","doi-asserted-by":"publisher","first-page":"861","DOI":"10.1016\/j.patrec.2005.10.010","volume":"27","author":"T Fawcett","year":"2006","unstructured":"Fawcett, T. (2006). An Introduction to ROC Analysis. Pattern Recognition Letters, 27, 861\u2013874.","journal-title":"Pattern Recognition Letters"},{"key":"6609_CR23","doi-asserted-by":"crossref","unstructured":"Frery, J., Habrard, A., Sebban, M., Caelen, O., He-Guelton, L. (2017). Efficient top rank optimization with gradient boosting for supervised anomaly detection. In European Conference on Machine Learning and Principles and Practice of Knowledge Discovery in Databases (ECML\/PKDD\u201917)","DOI":"10.1007\/978-3-319-71249-9_2"},{"key":"6609_CR24","unstructured":"Goix, N., Sabourin, A., Cl\u00e9men\u00e7on, S. (2015). On anomaly ranking and excess-mass curves. In Proceedings of the 18th International Conference on Artificial Intelligence and Statistics"},{"issue":"3","key":"6609_CR25","doi-asserted-by":"publisher","first-page":"112","DOI":"10.1214\/aoms\/1177704476","volume":"33","author":"J H\u00e1jek","year":"1962","unstructured":"H\u00e1jek, J. (1962). Asymptotically most powerful rank-order tests. The Annals of Mathematical Statistics, 33(3), 112\u20131147.","journal-title":"The Annals of Mathematical Statistics"},{"key":"6609_CR26","doi-asserted-by":"crossref","unstructured":"Hastie, T., Tibshirani, R., Friedman, J. (2009). The Elements of Statistical Learning, 2nd edn. Springer Series in Statistics. Springer","DOI":"10.1007\/978-0-387-84858-7"},{"key":"6609_CR27","unstructured":"Limnios, M., Noiry, N., Cl\u00e9men\u00e7on, S. (2021). Learning to rank anomalies: Scalar performance criteria and maximization of two-sample rank statistics. In Proceedings of the Third International Workshop on Learning with Imbalanced Domains: Theory and Applications, 154, 63\u201375"},{"key":"6609_CR28","doi-asserted-by":"crossref","unstructured":"Liu, F.T., Ting, K.M., Zhou, Z.H. (2008). Isolation forest. In Data Mining, 2008. ICDM\u201908. Eighth IEEE International Conference On, pp. 413\u2013422","DOI":"10.1109\/ICDM.2008.17"},{"key":"6609_CR29","doi-asserted-by":"crossref","unstructured":"Liu, F.T., Ting, K.M., Zhou, Z.-H. (2008). Isolation forest. In 2008 Eighth IEEE International Conference on Data Mining, pp. 413\u2013422","DOI":"10.1109\/ICDM.2008.17"},{"key":"6609_CR30","doi-asserted-by":"crossref","unstructured":"M\u00fcller, E., Assent, I., Iglesias, P., M\u00fclle, Y., B\u00f6hm, K. (2012). Outlier ranking via subspace analysis in multiple views of the data. In 2012 IEEE 12th International Conference on Data Mining, pp. 529\u2013538","DOI":"10.1109\/ICDM.2012.112"},{"key":"6609_CR31","doi-asserted-by":"crossref","unstructured":"M\u00fcller, E., S\u00e1nchez, P.I., M\u00fclle, Y., B\u00f6hm, K. (2013). Ranking outlier nodes in subspaces of attributed graphs. In IEEE 29th International Conference on Data Engineering Workshops (ICDEW), pp. 216\u2013222","DOI":"10.1109\/ICDEW.2013.6547453"},{"issue":"3","key":"6609_CR32","doi-asserted-by":"publisher","first-page":"497","DOI":"10.1137\/1109069","volume":"9","author":"EA Nadaraya","year":"1964","unstructured":"Nadaraya, E. A. (1964). Some new estimates for distribution functions. Theory of Probability and its Applications, 9(3), 497\u2013500.","journal-title":"Theory of Probability and its Applications"},{"key":"6609_CR33","doi-asserted-by":"crossref","unstructured":"Sch\u00f6lkopf, B., Platt, J., Shawe-Taylor, A.J. J.\u00a0Smola, Williamson, R.C. (2001). Estimating the support of a high-dimensional distribution. Neural Computation 13(7)","DOI":"10.1162\/089976601750264965"},{"key":"6609_CR1","doi-asserted-by":"publisher","unstructured":"Statlog (Shuttle). https:\/\/doi.org\/10.24432\/C5WS31","DOI":"10.24432\/C5WS31"},{"issue":"8","key":"6609_CR34","first-page":"211","volume":"6","author":"I Steinwart","year":"2005","unstructured":"Steinwart, I., Hush, D., & Scovel, C. (2005). A classification framework for anomaly detection. Journal of Machine Learning Research, 6(8), 211\u2013232.","journal-title":"Journal of Machine Learning Research"},{"key":"6609_CR35","unstructured":"Thomas, A., Cl\u00e9mencon, S., Gramfort, A., Sabourin, A. (2017) Anomaly Detection in Extreme Regions via Empirical MV-sets on the Sphere. In Proceedings of the 20th International Conference on Artificial Intelligence and Statistics, 54, 1011\u20131019"},{"key":"6609_CR36","doi-asserted-by":"crossref","unstructured":"Van\u00a0der Vaart, A., Wellner, J. (1996). Weak convergence and empirical processes. Springer Verlag New-York","DOI":"10.1007\/978-1-4757-2545-2"},{"key":"6609_CR37","doi-asserted-by":"publisher","first-page":"80","DOI":"10.2307\/3001968","volume":"1","author":"F Wilcoxon","year":"1945","unstructured":"Wilcoxon, F. (1945). Individual comparisons by ranking methods. Biometrics, 1, 80\u201383.","journal-title":"Biometrics"},{"key":"6609_CR38","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1155\/2019\/2686378","volume":"2019","author":"X Xu","year":"2019","unstructured":"Xu, X., Liu, H., & Yao, M. (2019). Recent progress of anomaly detection. Complexity, 2019, 1\u201311.","journal-title":"Complexity"}],"container-title":["Machine Learning"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10994-024-06609-9.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10994-024-06609-9\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10994-024-06609-9.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,12,30]],"date-time":"2024-12-30T16:08:21Z","timestamp":1735574901000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10994-024-06609-9"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,10,21]]},"references-count":38,"journal-issue":{"issue":"11-12","published-print":{"date-parts":[[2024,12]]}},"alternative-id":["6609"],"URL":"https:\/\/doi.org\/10.1007\/s10994-024-06609-9","relation":{},"ISSN":["0885-6125","1573-0565"],"issn-type":[{"value":"0885-6125","type":"print"},{"value":"1573-0565","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,10,21]]},"assertion":[{"value":"21 June 2022","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"1 July 2024","order":2,"name":"revised","label":"Revised","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"4 August 2024","order":3,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"21 October 2024","order":4,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare no conflict of interest.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}},{"value":"Not applicable","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethical approval"}},{"value":"Not applicable.","order":4,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent to participate"}},{"value":"Not applicable.","order":5,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent for publication"}}]}}