{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,10]],"date-time":"2026-04-10T01:11:07Z","timestamp":1775783467694,"version":"3.50.1"},"reference-count":31,"publisher":"MIT Press","issue":"6","content-domain":{"domain":["direct.mit.edu"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2023,5,12]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>Pairwise learning is widely employed in ranking, similarity and metric learning, area under the ROC curve (AUC) maximization, and many other learning tasks involving sample pairs. Pairwise learning with deep neural networks was considered for ranking, but enough theoretical understanding about this topic is lacking. In this letter, we apply symmetric deep neural networks to pairwise learning for ranking with a hinge loss \u03d5h and carry out generalization analysis for this algorithm. A key step in our analysis is to characterize a function that minimizes the risk. This motivates us to first find the minimizer of \u03d5h-risk and then design our two-part deep neural networks with shared weights, which induces the antisymmetric property of the networks. We present convergence rates of the approximation error in terms of function smoothness and a noise condition and give an excess generalization error bound by means of properties of the hypothesis space generated by deep neural networks. Our analysis is based on tools from U-statistics and approximation theory.<\/jats:p>","DOI":"10.1162\/neco_a_01585","type":"journal-article","created":{"date-parts":[[2023,4,10]],"date-time":"2023-04-10T21:02:42Z","timestamp":1681160562000},"page":"1135-1158","update-policy":"https:\/\/doi.org\/10.1162\/mitpressjournals.corrections.policy","source":"Crossref","is-referenced-by-count":3,"title":["Generalization Analysis of Pairwise Learning for Ranking With Deep Neural Networks"],"prefix":"10.1162","volume":"35","author":[{"given":"Shuo","family":"Huang","sequence":"first","affiliation":[{"name":"Department of Mathematics, City University of Hong Kong, Kowloon, Hong Kong shuang56-c@my.cityu.edu.hk"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Junyu","family":"Zhou","sequence":"additional","affiliation":[{"name":"School of Data Science, City University of Hong Kong, Kowloon, Hong Kong junyuzhou4-c@my.cityu.edu.hk"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Han","family":"Feng","sequence":"additional","affiliation":[{"name":"Department of Mathematics, City University of Hong Kong, Kowloon, Hong Kong hanfeng@cityu.edu.hk"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ding-Xuan","family":"Zhou","sequence":"additional","affiliation":[{"name":"School of Mathematics and Statistics, University of Sydney, Sydney NSW 2006, Australia dingxuan.zhou@sydney.edu.au"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"281","published-online":{"date-parts":[[2023,5,12]]},"reference":[{"issue":"2","key":"2023050921091210600_B1","volume":"10","author":"Agarwal","year":"2009","journal-title":"Journal of Machine Learning Research"},{"issue":"2","key":"2023050921091210600_B2","doi-asserted-by":"publisher","first-page":"608","DOI":"10.1214\/009053606000001217","article-title":"Fast learning rates for plug-in classifiers","volume":"35","author":"Audibert","year":"2007","journal-title":"Annals of Statistics"},{"key":"2023050921091210600_B3","doi-asserted-by":"publisher","first-page":"259","DOI":"10.1016\/j.neucom.2014.09.044","article-title":"Robustness and generalization for metric learning","volume":"151","author":"Bellet","year":"2015","journal-title":"Neurocomputing"},{"key":"2023050921091210600_B4","first-page":"1533","article-title":"Semantic parsing on freebase from question-answer pairs","volume-title":"Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing","author":"Berant","year":"2013"},{"key":"2023050921091210600_B5","doi-asserted-by":"crossref","first-page":"2212","DOI":"10.1145\/3292500.3330745","article-title":"Fairness in recommendation ranking through pairwise comparisons","volume-title":"Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining","author":"Beutel","year":"2019"},{"issue":"1","key":"2023050921091210600_B6","doi-asserted-by":"publisher","first-page":"115","DOI":"10.1007\/s10994-015-5499-7","article-title":"Generalization bounds for metric and similarity learning","volume":"102","author":"Cao","year":"2016","journal-title":"Machine Learning"},{"issue":"12","key":"2023050921091210600_B7","doi-asserted-by":"publisher","first-page":"1513","DOI":"10.1016\/j.jat.2012.09.001","article-title":"The convergence rate of a regularized ranking algorithm","volume":"164","author":"Chen","year":"2012","journal-title":"Journal of Approximation Theory"},{"key":"2023050921091210600_B8","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1016\/j.jco.2016.07.001","article-title":"On the robustness of regularized pairwise learning methods based on kernels","volume":"37","author":"Christmann","year":"2016","journal-title":"Journal of Complexity"},{"issue":"2","key":"2023050921091210600_B9","doi-asserted-by":"crossref","first-page":"844","DOI":"10.1214\/009052607000000910","article-title":"Ranking and empirical minimization of U-statistics","volume":"36","author":"Cl\u00e9men\u00e7on","year":"2008","journal-title":"Annals of Statistics"},{"issue":"3","key":"2023050921091210600_B10","doi-asserted-by":"crossref","first-page":"273","DOI":"10.1007\/BF00994018","article-title":"Support-vector networks","volume":"20","author":"Cortes","year":"1995","journal-title":"Machine Learning"},{"key":"2023050921091210600_B11","doi-asserted-by":"crossref","DOI":"10.1017\/CBO9780511618796","volume-title":"Learning theory: An approximation theory viewpoint","author":"Cucker","year":"2007"},{"key":"2023050921091210600_B12","doi-asserted-by":"crossref","first-page":"111","DOI":"10.1007\/978-3-031-07155-3_5","article-title":"On the robustness of kernel-based pairwise learning","volume-title":"Artificial intelligence, big data and data science in statistics","author":"Gensler","year":"2022"},{"issue":"3","key":"2023050921091210600_B13","doi-asserted-by":"publisher","first-page":"497","DOI":"10.1162\/NECO_a_00556","article-title":"Guaranteed classification via regularized similarity learning","volume":"26","author":"Guo","year":"2014","journal-title":"Neural Computation"},{"issue":"4","key":"2023050921091210600_B14","doi-asserted-by":"publisher","first-page":"1853","DOI":"10.1109\/TPAMI.2020.3032422","article-title":"Depth selection for deep ReLU nets in feature extraction and generalization","volume":"44","author":"Han","year":"2020","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"2023050921091210600_B15","doi-asserted-by":"crossref","first-page":"409","DOI":"10.1007\/978-1-4612-0865-5_26","article-title":"Probability inequalities for sums of bounded random variables","volume-title":"The collected works of Wassily Hoeffding","author":"Hoeffding","year":"1994"},{"key":"2023050921091210600_B16","first-page":"377","article-title":"Learning theory approach to minimum error entropy criterion","volume-title":"Journal of Machine Learning Research","author":"Hu","year":"2013"},{"key":"2023050921091210600_B17","doi-asserted-by":"publisher","first-page":"26","DOI":"10.1016\/j.neunet.2015.08.012","article-title":"A linear functional strategy for regularized ranking","volume":"73","author":"Kriukova","year":"2016","journal-title":"Neural Networks"},{"key":"2023050921091210600_B18","first-page":"21236","article-title":"Sharper generalization bounds for pairwise learning","volume-title":"Advances in neural information processing systems","author":"Lei","year":"2020"},{"key":"2023050921091210600_B19","doi-asserted-by":"crossref","DOI":"10.1109\/TIT.2022.3151753","article-title":"Universal consistency of deep convolutional neural networks","volume-title":"IEEE Transactions on Information Theory","author":"Lin","year":"2022"},{"issue":"3","key":"2023050921091210600_B20","doi-asserted-by":"publisher","first-page":"225","DOI":"10.1561\/1500000016","article-title":"Learning to rank for information retrieval","volume":"3","author":"Liu","year":"2009","journal-title":"Foundations and Trends in Information Retrieval"},{"issue":"9","key":"2023050921091210600_B21","article-title":"Kernel analysis of deep networks","volume":"12","author":"Montavon","year":"2011","journal-title":"Journal of Machine Learning Research"},{"issue":"3","key":"2023050921091210600_B22","article-title":"Learning coordinate covariances via gradients","volume":"7","author":"Mukherjee","year":"2006","journal-title":"Journal of Machine Learning Research"},{"key":"2023050921091210600_B23","author":"Qin","year":"2013","journal-title":"Introducing LETOR 4.0 datasets"},{"issue":"5","key":"2023050921091210600_B24","first-page":"1373","article-title":"On ranking and generalization bounds","volume":"13","author":"Rejchel","year":"2012","journal-title":"Journal of Machine Learning Research"},{"issue":"4","key":"2023050921091210600_B25","doi-asserted-by":"publisher","first-page":"768","DOI":"10.1080\/10485252.2017.1369078","article-title":"Model selection consistency of U-statistics with convex loss and weighted lasso penalty","volume":"29","author":"Rejchel","year":"2017","journal-title":"Journal of Nonparametric Statistics"},{"key":"2023050921091210600_B26","volume-title":"Support vector machines","author":"Steinwart","year":"2008"},{"key":"2023050921091210600_B27","article-title":"Adaptivity of deep ReLU network for learning in Besov and mixed smooth Besov spaces: Optimal rate and curse of dimensionality","volume-title":"International Conference on Learning Representations","author":"Suzuki","year":"2018"},{"issue":"4","key":"2023050921091210600_B28","doi-asserted-by":"publisher","first-page":"743","DOI":"10.1162\/NECO_a_00817","article-title":"Online pairwise learning algorithms","volume":"28","author":"Ying","year":"2016","journal-title":"Neural Computation"},{"key":"2023050921091210600_B29","first-page":"233","article-title":"Online AUC maximization","volume-title":"Proceedings of the 28th International Conference on Machine Learning","author":"Zhao","year":"2011"},{"issue":"6","key":"2023050921091210600_B30","doi-asserted-by":"publisher","first-page":"895","DOI":"10.1142\/S0219530518500124","article-title":"Deep distributed convolutional neural networks: Universality","volume":"16","author":"Zhou","year":"2018","journal-title":"Analysis and Applications"},{"issue":"2","key":"2023050921091210600_B31","doi-asserted-by":"publisher","first-page":"787","DOI":"10.1016\/j.acha.2019.06.004","article-title":"Universality of deep convolutional neural networks","volume":"48","author":"Zhou","year":"2020","journal-title":"Applied and Computational Harmonic Analysis"}],"container-title":["Neural Computation"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/direct.mit.edu\/neco\/article-pdf\/35\/6\/1135\/2086331\/neco_a_01585.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/direct.mit.edu\/neco\/article-pdf\/35\/6\/1135\/2086331\/neco_a_01585.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,5,10]],"date-time":"2023-05-10T06:28:27Z","timestamp":1683700107000},"score":1,"resource":{"primary":{"URL":"https:\/\/direct.mit.edu\/neco\/article\/35\/6\/1135\/115599\/Generalization-Analysis-of-Pairwise-Learning-for"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,5,12]]},"references-count":31,"journal-issue":{"issue":"6","published-online":{"date-parts":[[2023,5,12]]},"published-print":{"date-parts":[[2023,5,12]]}},"URL":"https:\/\/doi.org\/10.1162\/neco_a_01585","relation":{},"ISSN":["0899-7667","1530-888X"],"issn-type":[{"value":"0899-7667","type":"print"},{"value":"1530-888X","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2023,6]]},"published":{"date-parts":[[2023,5,12]]}}}