{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,24]],"date-time":"2026-06-24T14:58:05Z","timestamp":1782313085225,"version":"3.54.5"},"reference-count":53,"publisher":"Springer Science and Business Media LLC","issue":"17","license":[{"start":{"date-parts":[[2023,3,21]],"date-time":"2023-03-21T00:00:00Z","timestamp":1679356800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2023,3,21]],"date-time":"2023-03-21T00:00:00Z","timestamp":1679356800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100001659","name":"Deutsche Forschungsgemeinschaft","doi-asserted-by":"publisher","award":["ES434\/8-1"],"award-info":[{"award-number":["ES434\/8-1"]}],"id":[{"id":"10.13039\/501100001659","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001652","name":"Friedrich-Alexander-Universit\u00e4t Erlangen-N\u00fcrnberg","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100001652","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Appl Intell"],"published-print":{"date-parts":[[2023,9]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Progress in making neural networks more robust against adversarial attacks is mostly marginal, despite the great efforts of the research community. Moreover, the robustness evaluation is often imprecise, making it challenging to identify promising approaches. We do an observational study on the classification decisions of 19 different state-of-the-art neural networks trained to be robust against adversarial attacks. This analysis gives a new indication of the limits of the robustness of current models on a common benchmark. In addition, our findings suggest that current untargeted adversarial attacks induce misclassification toward only a limited amount of different classes. Similarly, we find that previous attacks under-explore the perturbation space during optimization. This leads to unsuccessful attacks for samples where the initial gradient direction is not a good approximation of the final adversarial perturbation direction. Additionally, we observe that both over- and under-confidence in model predictions result in an inaccurate assessment of model robustness. Based on these observations, we propose a novel loss function for adversarial attacks that consistently improves their efficiency and success rate compared to prior attacks for all 30 analyzed models.<\/jats:p>","DOI":"10.1007\/s10489-023-04532-5","type":"journal-article","created":{"date-parts":[[2023,3,21]],"date-time":"2023-03-21T11:02:51Z","timestamp":1679396571000},"page":"19843-19859","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":49,"title":["Exploring misclassifications of robust neural networks to enhance adversarial attacks"],"prefix":"10.1007","volume":"53","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-3967-2202","authenticated-orcid":false,"given":"Leo","family":"Schwinn","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ren\u00e9","family":"Raab","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"An","family":"Nguyen","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Dario","family":"Zanca","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Bjoern","family":"Eskofier","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2023,3,21]]},"reference":[{"key":"4532_CR1","unstructured":"Szegedy C, Zaremba W, Sutskever I, Bruna J, Erhan D, Goodfellow IJ, Fergus R (2014) Intriguing properties of neural networks. In: International conference on learning representations, ICLR"},{"key":"4532_CR2","unstructured":"Qin Y, Carlini N, Cottrell GW, Goodfellow IJ, Raffel C (2019) Imperceptible, robust, and targeted adversarial examples for automatic speech recognition. In: International conference on machine learning, ICML of proceedings of machine learning research, PMLR, vol 97, pp 5231\u20135240"},{"issue":"10","key":"4532_CR3","doi-asserted-by":"publisher","first-page":"120","DOI":"10.1109\/MCOM.2019.1900006","volume":"57","author":"S Hu","year":"2019","unstructured":"Hu S, Shang X, Qin Z, Li M, Wang Q, Wang C (2019) Adversarial examples for automatic speech recognition: attacks and countermeasures. IEEE Commun Mag 57(10):120\u2013126. https:\/\/doi.org\/10.1109\/MCOM.2019.1900006","journal-title":"IEEE Commun Mag"},{"key":"4532_CR4","doi-asserted-by":"crossref","unstructured":"Morris JX, Lifland E, Yoo JY, Grigsby J, Jin D, Qi Y (2020) Textattack: a framework for adversarial attacks, data augmentation, and adversarial training in NLP. In: Conference on empirical methods in natural language processing: system demonstrations, EMNLP, demo track, association for computational linguistics, pp 119\u2013126","DOI":"10.18653\/v1\/2020.emnlp-demos.16"},{"issue":"43","key":"4532_CR5","first-page":"1","volume":"21","author":"P Yang","year":"2020","unstructured":"Yang P, Chen J, Hsieh C-J, Wang J-L, Michael I, Jordan MI (2020) Greedy attack and gumbel attack: generating adversarial examples for discrete data. J Mach Learn Res JMLR 21(43):1\u201336","journal-title":"J Mach Learn Res JMLR"},{"issue":"3","key":"4532_CR6","doi-asserted-by":"publisher","first-page":"346","DOI":"10.1016\/j.eng.2019.12.012","volume":"6","author":"K Ren","year":"2020","unstructured":"Ren K, Zheng T, Qin Z, Liu X (2020) Adversarial attacks and defenses in deep learning. Engineering 6(3):346\u2013360. https:\/\/doi.org\/10.1016\/j.eng.2019.12.012. ISSN 2095-8099","journal-title":"Engineering"},{"key":"4532_CR7","unstructured":"Carmon Y, Raghunathan A, Schmidt L, Duchi JC, Liang P (2019) Unlabeled data improves adversarial robustness. In: Advances in neural information processing systems, NeurIPS, pp 11190\u2013 11201"},{"key":"4532_CR8","unstructured":"Ding GW, Sharma Y, Lui KYC, Huang R (2020) MMA training: direct input space margin maximization through adversarial training. In: International conference on learning representations, ICLR"},{"key":"4532_CR9","unstructured":"Hendrycks D, Lee K, Mazeika M (2019) Using pre-training can improve model robustness and uncertainty. In: International conference on machine learning, ICML of proceedings of machine learning research, PMLR, vol 97, pp 2712\u20132721"},{"key":"4532_CR10","unstructured":"Madry A, Makelov A, Schmidt L, Tsipras D, Vladu A (2018) Towards deep learning models resistant to adversarial attacks. In: 6th International conference on learning representations, ICLR"},{"key":"4532_CR11","doi-asserted-by":"crossref","unstructured":"Leon Bungert, Raab R, Roith T, Schwinn L, Tenbrinck D (2021) CLIP: cheap lipschitz training of neural networks. In: Scale space and variational methods in computer vision, SSVM of lecture notes in computer science. Springer, vol 12679, pp 307\u2013 319","DOI":"10.1007\/978-3-030-75549-2_25"},{"key":"4532_CR12","unstructured":"Schwinn L, Nguyen A, Raab R, Bungert L, Tenbrinck D, Zanca D, Burger M, Eskofier BM (2021a) Identifying untrustworthy predictions in neural networks by geometric gradient analysis. In: Conference on uncertainty in artificial intelligence, UAI of proceedings of machine learning research, AUAI Press, vol 161, pp 854\u2013864"},{"issue":"221","key":"4532_CR13","first-page":"1","volume":"22","author":"E Richardson","year":"2021","unstructured":"Richardson E, Weiss Y (2021) A bayes-optimal view on adversarial examples. J Mach Learn Res JMLR 22(221):1\u201328","journal-title":"J Mach Learn Res JMLR"},{"key":"4532_CR14","unstructured":"Jin C, Rinard M (2020) Manifold regularization for adversarial robustness. arXiv:2003.04286"},{"key":"4532_CR15","unstructured":"Pang T, Yang X, Dong Y, Xu T, Zhu J, Su H (2020) Boosting adversarial training with hypersphere embedding. In: Advances in neural information processing systems, NeurIPS"},{"key":"4532_CR16","unstructured":"Croce F, Hein M (2020) Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacks. In: International conference on machine learning, ICML, vol 119 of proceedings of machine learning research, pp 2206\u20132216. PMLR"},{"key":"4532_CR17","unstructured":"Athalye A, Carlini N, Wagner DA (2018) Obfuscated gradients give a false sense of security: circumventing defenses to adversarial examples. In: International conference on machine learning, ICML, vol 80 of proceedings of machine learning research, pp 274\u2013283. PMLR"},{"key":"4532_CR18","unstructured":"Tram\u00e8r F, Carlini N, Brendel W, Madry A (2020) On adaptive attacks to adversarial example defenses. In: Larochelle H, Marc\u2019Aurelio Ranzato RH, Balcan M-F, Lin H-T (eds) Advances in neural information processing systems, NeurIPS"},{"key":"4532_CR19","unstructured":"Uesato J, O\u2019Donoghue B, Kohli P, van den Oord A (2018) Adversarial risk and the dangers of evaluating against weak attacks. In: International conference on machine learning, ICML, vol 80 of proceedings of machine learning research, pp 5032\u20135041. PMLR"},{"key":"4532_CR20","unstructured":"Krizhevsky A (2009) Learning multiple layers of features from tiny images. Tech Rep"},{"key":"4532_CR21","doi-asserted-by":"crossref","unstructured":"Carlini N, Wagner DA (2017) Towards evaluating the robustness of neural networks. In: 2017 IEEE symposium on security and privacy, SP, pp 39\u201357. IEEE Computer Society","DOI":"10.1109\/SP.2017.49"},{"key":"4532_CR22","unstructured":"Brendel W, Rauber J, K\u00fcmmerer M, Ustyuzhaninov I, Bethge M (2019) Accurate, reliable and fast robustness evaluation. In: Advances in neural information processing systems, NeurIPS, pp 12841\u201312851"},{"key":"4532_CR23","unstructured":"Lin J, Song C, He K, Wang L, Hopcroft JE (2020) Nesterov accelerated gradient and scale invariance for adversarial attacks. In: International conference on learning representations, ICLR"},{"key":"4532_CR24","doi-asserted-by":"crossref","unstructured":"Schwinn L, Nguyen A, Raab R, Zanca D, Eskofier BM, Tenbrinck D, Burger M (2021) Dynamically sampled nonlocal gradients for stronger adversarial attacks. In: International joint conference on neural networks, IJCNN, pp 1\u20138. IEEE","DOI":"10.1109\/IJCNN52387.2021.9534190"},{"key":"4532_CR25","unstructured":"Pintor M, Roli F, Brendel W, Biggio B (2021) Fast minimum-norm adversarial attacks through adaptive norm constraints. In: Advances in neural information processing systems, NeurIPS, pp 20052\u201320062"},{"key":"4532_CR26","doi-asserted-by":"crossref","unstructured":"Mao X, Chen Y, Wang S, Su H, He Y, Xue H (2021) Composite adversarial attacks. In: Conference on artificial intelligence, AAAI, pp 8884\u20138892. AAAI Press","DOI":"10.1609\/aaai.v35i10.17075"},{"key":"4532_CR27","doi-asserted-by":"crossref","unstructured":"Mustafa A, Khan SH, Hayat M, Goecke R, Shen J, Shao L (2019) Adversarial defense by restricting the hidden space of deep neural networks. In: IEEE\/CVF international conference on computer vision, ICCV, pp 3384\u20133393. IEEE","DOI":"10.1109\/ICCV.2019.00348"},{"key":"4532_CR28","unstructured":"Ilyas A, Santurkar S, Tsipras D, Engstrom L, Tran B, Madry A (2019) Adversarial examples are not bugs, they are features. In: Advances in neural information processing systems, NeurIPS, pp 125\u2013136"},{"key":"4532_CR29","doi-asserted-by":"crossref","unstructured":"Deng J, Dong W, Socher R, Li LJ, Kai Li, Fei-Fei L (2009) Imagenet: a large-scale hierarchical image database. In: IEEE\/CVF computer society conference on computer vision and pattern recognition CVPR, pp 248\u2013255. IEEE","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"4532_CR30","unstructured":"Ozbulak U, Pintor M, Van Messem A, De Neve W (2021) Evaluating adversarial attacks on imagenet: a reality check on misclassification classes. In: Advances in neural information processing systems, NeurIPS, workshop on ImageNet: past, present, and future"},{"key":"4532_CR31","unstructured":"Croce F, Andriushchenko M, Sehwag V , Debenedetti E, Flammarion N, Chiang M, Mittal P, Hein M (2021) Robustbench: a standardized adversarial robustness benchmark. In: Advances in neural information processing systems, NeurIPS, track Datasets and benchmarks"},{"key":"4532_CR32","doi-asserted-by":"publisher","first-page":"46084","DOI":"10.1109\/ACCESS.2018.2866197","volume":"6","author":"H Kwon","year":"2018","unstructured":"Kwon H, Kim Y, Park KW, Yoon H, Choi D (2018) Multi-targeted adversarial example in evasion attack on deep neural network. IEEE Access 6:46084\u201346096","journal-title":"IEEE Access"},{"issue":"4598","key":"4532_CR33","doi-asserted-by":"publisher","first-page":"671","DOI":"10.1126\/science.220.4598.671","volume":"220","author":"S Kirkpatrick","year":"1983","unstructured":"Kirkpatrick S, Gelatt D, Vecchi M (1983) Optimization by simulated annealing. Science 220(4598):671\u2013680","journal-title":"Science"},{"key":"4532_CR34","doi-asserted-by":"publisher","first-page":"95","DOI":"10.1023\/A:1022602019183","volume":"3","author":"DE Goldberg","year":"1988","unstructured":"Goldberg DE, Holland JH (1988) Genetic algorithms and machine learning. Mach Learn 3:95\u201399","journal-title":"Mach Learn"},{"key":"4532_CR35","unstructured":"Hinton GE, Roweis ST (2002) Stochastic neighbor embedding. In: Advances in neural information processing systems, NeurIPS, pp 833\u2013840. MIT Press"},{"key":"4532_CR36","unstructured":"Neelakantan A, Vilnis L, Le QV, Sutskever I, Kaiser L, Kurach K, Martens J (2015) adding gradient noise improves learning for very deep networks. CoRR arXiv:1511.06807"},{"key":"4532_CR37","unstructured":"Wong E, Rice L, Kolter JZ (2020) Fast is better than free: revisiting adversarial training. In: International conference on learning representations, ICLR"},{"key":"4532_CR38","unstructured":"Zhang D, Zhang T, Lu Y, Zhu Z, Dong B (2019) You only propagate once: accelerating adversarial training via maximal principle. In: Advances in neural information processing systems, NeurIPS, pp 227\u2013238"},{"key":"4532_CR39","unstructured":"Engstrom L, Ilyas A, Salman H, Santurkar S, Tsipras D (2019) Robustness python library. https:\/\/github.com\/MadryLab\/robustness. [Accessed May 25th, 2021]"},{"key":"4532_CR40","unstructured":"Zhang H, Yu Y, Jiao J, Xing EP, El Ghaoui L, Jordan MI (2019) Theoretically principled trade-off between robustness and accuracy. In: International conference on machine learning, ICML, vol 97 of proceedings of machine learning research, pp 7472\u20137482. PMLR"},{"key":"4532_CR41","unstructured":"Huang L, Zhang C, Zhang H (2020) Self-adaptive training: beyond empirical risk minimization. In: Advances in neural information processing systems, NeurIPS"},{"key":"4532_CR42","unstructured":"Zhang J, Xu X, Han B, Niu G, Cui L, Sugiyama M, Kankanhalli MS (2020) Attacks which do not kill training make adversarial learning stronger. In: International conference on machine learning, ICML, vol 119 of proceedings of machine learning research, pp 11278\u201311287. PMLR"},{"key":"4532_CR43","unstructured":"Rice L, Wong E, Kolter JZ (2020) Overfitting in adversarially robust deep learning. In: International conference on machine learning, ICML, vol 119 of proceedings of machine learning research, pp 8093\u20138104. PMLR"},{"key":"4532_CR44","unstructured":"Sehwag V, Mahloujifar S, Handina T, Dai S, Xiang C, Chiang M, Mittal P (2021) Improving adversarial robustness using proxy distributions. CoRR arXiv:2104.09425"},{"key":"4532_CR45","unstructured":"Wu D, Xia S-T, Wang Y (2020) Adversarial weight perturbation helps robust generalization. In: Advances in neural information processing systems, NeurIPS"},{"key":"4532_CR46","unstructured":"Gowal S, Qin C, Uesato J, Mann TA, Kohli P (2020) Uncovering the limits of adversarial training against norm-bounded adversarial examples. CoRR arXiv:2010.03593"},{"key":"4532_CR47","unstructured":"Wang Y, Zou D, Yi J, Bailey J, Ma X, Gu Q (2020) Improving adversarial robustness requires revisiting misclassified examples. In: International conference on learning representations, ICLR"},{"key":"4532_CR48","unstructured":"Sehwag V, Wang S, Mittal P, Jana S (2020) HYDRA: pruning adversarially robust neural networks. In: Advances in neural information processing systems, NeurIPS"},{"key":"4532_CR49","unstructured":"Zhang J, Zhu J, Niu G, Han B, Sugiyama M, Kankanhalli MS (2021) Geometry-aware instance-reweighted adversarial training. In: International conference on learning representations, ICLR"},{"key":"4532_CR50","unstructured":"Sitawarin C, Chakraborty S, Wagner DA (2020) Improving adversarial robustness through progressive hardening. CoRR arXiv:2003.09347"},{"key":"4532_CR51","doi-asserted-by":"crossref","unstructured":"Chen J, Cheng Y, Gan Z, Gu Q, Liu J (2022) Efficient robust training via backward smoothing. In: Conference On artificial intelligence, AAAI, pp 6222\u20136230. AAAI Press","DOI":"10.1609\/aaai.v36i6.20571"},{"key":"4532_CR52","doi-asserted-by":"crossref","unstructured":"Cui J, Liu S, Wang L, Jia J (2021) Learnable boundary guided adversarial training. In: IEEE\/CVF international conference on computer vision, ICCV, pp 15701\u201315710. IEEE","DOI":"10.1109\/ICCV48922.2021.01543"},{"key":"4532_CR53","unstructured":"Zhu Z, Wu J, Yu B, Wu L, Ma J (2020) The anisotropic noise in stochastic gradient descent: its behavior of escaping from sharp minima and regularization effects. In: International conference on machine learning, ICML , vol 97 of proceedings of machine learning research, pp 7654\u20137663. PMLR"}],"container-title":["Applied Intelligence"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10489-023-04532-5.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10489-023-04532-5\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10489-023-04532-5.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,9,15]],"date-time":"2023-09-15T11:32:53Z","timestamp":1694777573000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10489-023-04532-5"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,3,21]]},"references-count":53,"journal-issue":{"issue":"17","published-print":{"date-parts":[[2023,9]]}},"alternative-id":["4532"],"URL":"https:\/\/doi.org\/10.1007\/s10489-023-04532-5","relation":{},"ISSN":["0924-669X","1573-7497"],"issn-type":[{"value":"0924-669X","type":"print"},{"value":"1573-7497","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,3,21]]},"assertion":[{"value":"15 February 2023","order":1,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"21 March 2023","order":2,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors have no competing interests to declare that are relevant to the content of this article. All the data used in the experiments is open source and freely available.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"<!--Emphasis Type='Bold' removed-->Competing interests"}}]}}