{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,2]],"date-time":"2026-07-02T07:04:29Z","timestamp":1782975869683,"version":"3.54.5"},"reference-count":63,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2024,1,2]],"date-time":"2024-01-02T00:00:00Z","timestamp":1704153600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2024,1,2]],"date-time":"2024-01-02T00:00:00Z","timestamp":1704153600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J Glob Optim"],"published-print":{"date-parts":[[2026,5]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>We discuss two key problems related to learning and optimization of neural networks: the computation of the adversarial attack for adversarial robustness and approximate optimization of complex functions. We show that both problems can be cast as instances of DC-programming. We give an explicit decomposition of the corresponding functions as differences of convex functions (DC) and report the results of experiments demonstrating the effectiveness of the DCA algorithm applied to these problems.<\/jats:p>","DOI":"10.1007\/s10898-023-01344-2","type":"journal-article","created":{"date-parts":[[2024,1,2]],"date-time":"2024-01-02T01:03:11Z","timestamp":1704157391000},"page":"5-21","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":9,"title":["DC-programming for neural network optimizations"],"prefix":"10.1007","volume":"95","author":[{"given":"Pranjal","family":"Awasthi","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Anqi","family":"Mao","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Mehryar","family":"Mohri","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8461-1260","authenticated-orcid":false,"given":"Yutao","family":"Zhong","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2024,1,2]]},"reference":[{"key":"1344_CR1","doi-asserted-by":"crossref","unstructured":"Argyriou, A., Hauser, R., Micchelli, C.A., Pontil, M.: A DC-programming algorithm for kernel selection. In: International Conference on Machine Learning, pp. 41\u201348 (2006)","DOI":"10.1145\/1143844.1143850"},{"key":"1344_CR2","unstructured":"Awasthi, P., Frank, N., Mao, A., Mohri, M., Zhong, Y.: Calibration and consistency of adversarial surrogate losses. In: Advances in Neural Information Processing Systems, pp. 9804\u20139815 (2021)"},{"key":"1344_CR3","unstructured":"Awasthi, P., Mao, A., Mohri, M., Zhong, Y.: A finer calibration analysis for adversarial robustness. arXiv preprint arXiv:2105.01550 (2021)"},{"key":"1344_CR4","unstructured":"Awasthi, P., Mao, A., Mohri, M., Zhong, Y.: H-consistency bounds for surrogate loss minimizers. In: International Conference on Machine Learning, pp. 1117\u20131174 (2022)"},{"key":"1344_CR5","doi-asserted-by":"crossref","unstructured":"Awasthi, P., Mao, A., Mohri, M., Zhong, Y.: Multi-class H-consistency bounds. In: Advances in Neural Information Processing Systems, pp. 782\u2013795 (2022)","DOI":"10.52202\/068431-0057"},{"key":"1344_CR6","doi-asserted-by":"publisher","unstructured":"Awasthi, P., Cortes, C., Mohri, M.: Best-effort adaptation. https:\/\/doi.org\/10.48550\/arXiv.2305.05816 (2023)","DOI":"10.48550\/arXiv.2305.05816"},{"key":"1344_CR7","unstructured":"Awasthi, P., Mao, A., Mohri, M., Zhong, Y.: Theoretically grounded loss functions and algorithms for adversarial robustness. In: International Conference on Artificial Intelligence and Statistics, pp. 10077\u201310094 (2023)"},{"issue":"473","key":"1344_CR8","doi-asserted-by":"publisher","first-page":"138","DOI":"10.1198\/016214505000000907","volume":"101","author":"PL Bartlett","year":"2006","unstructured":"Bartlett, P.L., Jordan, M.I., McAuliffe, J.D.: Convexity, classification, and risk bounds. J. Am. Stat. Assoc. 101(473), 138\u2013156 (2006)","journal-title":"J. Am. Stat. Assoc."},{"key":"1344_CR9","doi-asserted-by":"crossref","unstructured":"Carlini, N., Wagner, D.: Towards evaluating the robustness of neural networks. In: IEEE Symposium on Security and Privacy (SP), pp. 39\u201357 (2017)","DOI":"10.1109\/SP.2017.49"},{"issue":"3","key":"1344_CR10","doi-asserted-by":"publisher","first-page":"273","DOI":"10.1023\/A:1022627411411","volume":"20","author":"C Cortes","year":"1995","unstructured":"Cortes, C., Vapnik, V.: Support-vector networks. Mach. Learn. 20(3), 273\u2013297 (1995)","journal-title":"Mach. Learn."},{"key":"1344_CR11","unstructured":"Cortes, C., Mohri, M., Suresh, A.T., Zhang, N.: A discriminative technique for multiple-source adaptation. In: International Conference on Machine Learning, pp. 2132\u20132143 (2021)"},{"key":"1344_CR12","unstructured":"Croce, F., Hein, M.: Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacks. In: International Conference on Machine Learning, pp. 2206\u20132216 (2020)"},{"issue":"3","key":"1344_CR13","first-page":"311","volume":"46","author":"F Wenlong","year":"2013","unstructured":"Wenlong, F., Tingsong, D., Zhai, J.: A new branch and bound algorithm based on DC-composition about nonconvex quadratic programming with box constrained. J. Math. Study 46(3), 311\u2013318 (2013)","journal-title":"J. Math. Study"},{"key":"1344_CR14","unstructured":"Goodfellow, I., Bengio, Y., Courville, A.: Deep Learning. MIT Press (2016)"},{"key":"1344_CR15","unstructured":"Goodfellow, I.J., Shlens, J., Szegedy, C.: Explaining and harnessing adversarial examples. arXiv preprint arXiv:1412.6572 (2014)"},{"key":"1344_CR16","unstructured":"Hiroshi, K., Thien, T.P., Hoang, T.: Optimization on Low Rank Nonconvex Structures, Volume\u00a015. Springer (2013)"},{"key":"1344_CR17","first-page":"1437","volume":"5","author":"T Hoang","year":"1964","unstructured":"Hoang, T.: Concave programming under linear constraints. Transl. Sov. Math. 5, 1437\u20131440 (1964)","journal-title":"Transl. Sov. Math."},{"key":"1344_CR18","unstructured":"Hoang, T.: Counter-examples to some results on D.C. optimization. Technical report, Institute of Mathematics, Hanoi, Vietnam (2002)"},{"key":"1344_CR19","unstructured":"Hoffman, J., Mohri, M., Zhang, N.: Algorithms and theory for multiple-source adaptation. In: Advances in Neural Information Processing Systems, pp. 8256\u20138266 (2018)"},{"issue":"3\u20134","key":"1344_CR20","first-page":"237","volume":"89","author":"J Hoffman","year":"2021","unstructured":"Hoffman, J., Mohri, M., Zhang, N.: Multiple-source adaptation theory and algorithms. Ann. Math. Artif. Intell. 89(3\u20134), 237\u2013270 (2021)","journal-title":"Ann. Math. Artif. Intell."},{"issue":"6","key":"1344_CR21","doi-asserted-by":"publisher","first-page":"569","DOI":"10.1007\/s10472-022-09791-5","volume":"90","author":"J Hoffman","year":"2022","unstructured":"Hoffman, J., Mohri, M., Zhang, N.: Multiple-source adaptation theory and algorithms - addendum. Ann. Math. Artif. Intell. 90(6), 569\u2013572 (2022)","journal-title":"Ann. Math. Artif. Intell."},{"issue":"5","key":"1344_CR22","doi-asserted-by":"publisher","first-page":"359","DOI":"10.1016\/0893-6080(89)90020-8","volume":"2","author":"K Hornik","year":"1989","unstructured":"Hornik, K., Stinchcombe, M., White, H.: Multilayer feedforward networks are universal approximators. Neural Networks 2(5), 359\u2013366 (1989)","journal-title":"Neural Networks"},{"key":"1344_CR23","unstructured":"Krizhevsky, A., Hinton, G., et\u00a0al.: Learning multiple layers of features from tiny images. Technical report (2009)"},{"key":"1344_CR24","unstructured":"Krizhevsky, A., Sutskever, I., Hinton, G.E.: Imagenet classification with deep convolutional neural networks. In: Advances in Neural Information Processing Systems, pp. 1097\u20131105 (2012)"},{"key":"1344_CR25","unstructured":"Kurakin, A., Goodfellow, I.J., Bengio, S.: Adversarial examples in the physical world. International Conference on Learning Representations workshop (2016)"},{"key":"1344_CR26","unstructured":"Kuznetsov, V., Mohri, M.: Learning theory and algorithms for forecasting non-stationary time series. In: Advances in Neural Information Processing Systems, pp. 541\u2013549 (2015)"},{"key":"1344_CR27","unstructured":"Kuznetsov, V., Mohri, M.: Time series prediction and online learning. In: Conference on Learning Theory, pp. 1190\u20131213 (2016)"},{"issue":"4","key":"1344_CR28","doi-asserted-by":"publisher","first-page":"367","DOI":"10.1007\/s10472-019-09683-1","volume":"88","author":"V Kuznetsov","year":"2020","unstructured":"Kuznetsov, V., Mohri, M.: Discrepancy-based theory and algorithms for forecasting non-stationary time series. Ann. Math. Artif. Intell. 88(4), 367\u2013399 (2020)","journal-title":"Ann. Math. Artif. Intell."},{"key":"1344_CR29","unstructured":"Kuznetsov, V., Mohri, M., Syed, U.: Multi-class deep boosting. In: Advances in Neural Information Processing Systems, pp. 2501\u20132509 (2014)"},{"key":"1344_CR30","doi-asserted-by":"crossref","unstructured":"Le Thi, H.A., Dinh, T.P.: A branch and bound method via dc optimization algorithms and ellipsoidal technique for box constrained nonconvex quadratic problems. J. Glob. Optim. 13(2), 171\u2013206 (1998)","DOI":"10.1023\/A:1008240227198"},{"key":"1344_CR31","doi-asserted-by":"crossref","unstructured":"Le Thi, H.A., Dinh, T.P.: The DC (difference of convex functions) programming and DCA revisited with DC models of real world nonconvex optimization problems. Ann. Oper. Research 133(1), 23\u201346 (2005)","DOI":"10.1007\/s10479-004-5022-1"},{"key":"1344_CR32","unstructured":"Le\u00a0Thi, H. A., Pham\u00a0Dinh, T.: DC Programming and DCA for Nonconvex Optimization: Theory, Algorithms and Applications. MAMERN09 (2009)"},{"key":"1344_CR33","doi-asserted-by":"crossref","unstructured":"Le\u00a0Thi, H.A., Dinh, T.P.: Recent advances in DC programming and DCA. Trans. Comput. Intell. 13:1\u201337 (2014)","DOI":"10.1007\/978-3-642-54455-2_1"},{"issue":"1","key":"1344_CR34","doi-asserted-by":"publisher","first-page":"5","DOI":"10.1007\/s10107-018-1235-y","volume":"169","author":"HA Le Thi","year":"2018","unstructured":"Le Thi, H.A., Dinh, T.P.: DC programming and DCA: thirty years of developments. Math. Program. 169(1), 5\u201368 (2018)","journal-title":"Math. Program."},{"issue":"1","key":"1344_CR35","doi-asserted-by":"publisher","first-page":"9","DOI":"10.1023\/A:1009777410170","volume":"2","author":"HA Le Thi","year":"1998","unstructured":"Le Thi, H.A., Dinh, T.P., Dung, M.L.: A combined DC optimization-ellipsoidal branch-and-bound algorithm for solving nonconvex quadratic programming problems. J. Combin. Optim. 2(1), 9\u201328 (1998)","journal-title":"J. Combin. Optim."},{"key":"1344_CR36","unstructured":"Long, P., Servedio, R.: Consistency versus realizable H-consistency for multiclass classification. In: International Conference on Machine Learning, pp. 801\u2013809 (2013)"},{"key":"1344_CR37","unstructured":"Madry, A., Makelov, A., Schmidt, L., Tsipras, D., Vladu, A.: Towards deep learning models resistant to adversarial attacks. arXiv preprint arXiv:1706.06083 (2017)"},{"key":"1344_CR38","doi-asserted-by":"crossref","unstructured":"Mao, A., Mohri, C., Mohri, M., Zhong, Y.: Two-stage learning to defer with multiple experts. In: Advances in Neural Information Processing Systems (2023)","DOI":"10.52202\/075280-0159"},{"key":"1344_CR39","doi-asserted-by":"crossref","unstructured":"Mao, A., Mohri, M., Zhong, Y.: H-consistency bounds: Characterization and extensions. In Advances in Neural Information Processing Systems (2023)","DOI":"10.52202\/075280-0199"},{"key":"1344_CR40","unstructured":"Mao, A., Mohri, M., Zhong, Y.: Theoretical analysis and applications. In: Cross-Entropy Loss Functions: International Conference on Machine Learning (2023)"},{"key":"1344_CR41","doi-asserted-by":"crossref","unstructured":"Mao, A., Mohri, M., Zhong, Y.: Structured prediction with stronger consistency guarantees. In: Advances in Neural Information Processing Systems (2023)","DOI":"10.52202\/075280-2032"},{"key":"1344_CR42","unstructured":"Mao, A., Mohri, M., Zhong, Y.: H-consistency bounds for pairwise misranking loss surrogates. In: International Conference on Machine Learning (2023)"},{"key":"1344_CR43","unstructured":"Mohri, M., Mu\u00f1oz Medina, A.: Learning algorithms for second-price auctions with reserve. J. Mach. Learn. Res. 17(1), 2632\u20132656 (2016)"},{"key":"1344_CR44","unstructured":"Mohri, M., Medina, A.M.: Learning theory and algorithms for revenue optimization in second price auctions with reserve. In: International Conference on Machine Learning, pp. 262\u2013270 (2014)"},{"key":"1344_CR45","unstructured":"Mohri, M., Rostamizadeh, A., Talwalkar, A.: Foundations of Machine Learning, 2nd Ed. MIT Press (2018)"},{"key":"1344_CR46","unstructured":"Pham Dinh, T., Le Thi, H.A.: Convex analysis approach to DC programming: theory, algorithms and applications. Acta mathematica vietnamica 22(1), 289\u2013355 (1997)"},{"key":"1344_CR47","doi-asserted-by":"crossref","unstructured":"Pham Dinh, T., Le Thi, H.A: A DC optimization algorithm for solving the trust-region subproblem. SIAM J. Optim. 8(2), 476\u2013505 (1998)","DOI":"10.1137\/S1052623494274313"},{"key":"1344_CR48","unstructured":"Reiner, H., Hoang, T.: Global Optimization: Deterministic Approaches. Springer, Berlin (2013)"},{"issue":"1","key":"1344_CR49","first-page":"1","volume":"103","author":"H Reiner","year":"1999","unstructured":"Reiner, H., Thoai, N.V.: DC Programming: Overview. J. Optim. Theory Appl. 103(1), 1\u201343 (1999)","journal-title":"J. Optim. Theory Appl."},{"key":"1344_CR50","unstructured":"Seck, I., Loosli, G., Canu, S., Niu, Y.-S.: Difference-of-convex algorithm applied to adversarial robustness verification (2019). https:\/\/ailab.criteo.com\/wp-content\/uploads\/2019\/10\/9.-Seck.pdf"},{"issue":"6","key":"1344_CR51","doi-asserted-by":"publisher","first-page":"1391","DOI":"10.1162\/NECO_a_00283","volume":"24","author":"BK Sriperumbudur","year":"2012","unstructured":"Sriperumbudur, B.K., Lanckriet, G.R.G.: A proof of convergence of the concave-convex procedure using Zangwill\u2019s theory. Neural Comput. 24(6), 1391\u20131407 (2012)","journal-title":"Neural Comput."},{"key":"1344_CR52","doi-asserted-by":"crossref","unstructured":"Sriperumbudur, B.K., Torres, D.A., Lanckriet, G.R.G.: Sparse Eigen Methods by D.C. Programming. In: International Conference on Machine Learning, pp. 831\u2013838 (2007)","DOI":"10.1145\/1273496.1273601"},{"issue":"2","key":"1344_CR53","doi-asserted-by":"publisher","first-page":"225","DOI":"10.1007\/s00365-006-0662-3","volume":"26","author":"I Steinwart","year":"2007","unstructured":"Steinwart, I.: How to compare different loss functions and their risks. Construct. Approxim. 26(2), 225\u2013287 (2007)","journal-title":"Construct. Approxim."},{"key":"1344_CR54","unstructured":"Sutskever, I., Vinyals, O., Le, Q.V.: Sequence to sequence learning with neural networks. In: Advances in Neural Information Processing Systems, pp. 3104\u20133112 (2014)"},{"key":"1344_CR55","unstructured":"Szegedy, C., Zaremba, W., Sutskever, I., Bruna, J., Erhan, D., Goodfellow, I., Fergus, R.: Intriguing properties of neural networks. arXiv preprint arXiv:1312.6199 (2013)"},{"issue":"36","key":"1344_CR56","first-page":"1007","volume":"8","author":"A Tewari","year":"2007","unstructured":"Tewari, A., Bartlett, P.L.: On the consistency of multiclass classification methods. J. Mach. Learn. Res. 8(36), 1007\u20131025 (2007)","journal-title":"J. Mach. Learn. Res."},{"key":"1344_CR57","unstructured":"Tsipras, D., Santurkar, S., Engstrom, L., Turner, A., Madry, A.: Robustness may be at odds with accuracy. arXiv preprint arXiv:1805.12152 (2018)"},{"key":"1344_CR58","unstructured":"Wong, E., Schmidt, F., Metzen, J.H., Zico Kolter, J.: Scaling provable adversarial defenses. Advances in Neural Information Processing Systems (2018)"},{"key":"1344_CR59","unstructured":"Yen, I.E.H., Peng, N., Wang, P.-O., Lin, S.-D.: On convergence rate of concave-convex procedure. In Proceedings of the NIPS 2012 Optimization Workshop, pp. 31\u201335 (2012)"},{"key":"1344_CR60","unstructured":"Yin, D., Ramchandran, K., Bartlett, P.L.: Rademacher complexity for adversarially robust generalization. In: International Conference of Machine Learning, pp. 7085\u20137094 (2019)"},{"issue":"4","key":"1344_CR61","doi-asserted-by":"publisher","first-page":"915","DOI":"10.1162\/08997660360581958","volume":"15","author":"AL Yuille","year":"2003","unstructured":"Yuille, A.L., Rangarajan, A.: The concave-convex procedure. Neural Comput. 15(4), 915\u2013936 (2003)","journal-title":"Neural Comput."},{"key":"1344_CR62","unstructured":"Zhang, M., Agarwal, S.: Bayes consistency vs. H-consistency: the interplay between surrogate loss functions and the scoring function class. In: Advances in Neural Information Processing Systems, pp. 16927\u201316936 (2020)"},{"issue":"1","key":"1344_CR63","doi-asserted-by":"publisher","first-page":"56","DOI":"10.1214\/aos\/1079120130","volume":"32","author":"T Zhang","year":"2004","unstructured":"Zhang, T.: Statistical behavior and consistency of classification methods based on convex risk minimization. Ann. Stat. 32(1), 56\u201385 (2004)","journal-title":"Ann. Stat."}],"container-title":["Journal of Global Optimization"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10898-023-01344-2.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10898-023-01344-2","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10898-023-01344-2.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,7,2]],"date-time":"2026-07-02T05:59:13Z","timestamp":1782971953000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10898-023-01344-2"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,1,2]]},"references-count":63,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2026,5]]}},"alternative-id":["1344"],"URL":"https:\/\/doi.org\/10.1007\/s10898-023-01344-2","relation":{},"ISSN":["0925-5001","1573-2916"],"issn-type":[{"value":"0925-5001","type":"print"},{"value":"1573-2916","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,1,2]]},"assertion":[{"value":"10 September 2022","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"12 November 2023","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"2 January 2024","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"The authors declare that they have no conflict of interest.","order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}}]}}