{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,13]],"date-time":"2026-01-13T09:01:14Z","timestamp":1768294874513,"version":"3.49.0"},"reference-count":61,"publisher":"Association for Computing Machinery (ACM)","issue":"2","content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["SIGKDD Explor. Newsl."],"published-print":{"date-parts":[[2025,12,30]]},"abstract":"<jats:p>This paper looks into the critical area of deep learning robustness and challenges the common belief that classification robustness and explanation robustness in image classification systems are inherently correlated. Through a novel evaluation approach leveraging clustering for e!cient assessment of explanation robustness, we demonstrate that enhancing explanation robustness does not necessarily flatten the input loss landscape with respect to explanation loss - contrary to flattened loss landscapes indicating better classification robustness. To further investigate this contradiction, a training method designed to adjust the loss landscape with respect to explanation loss is proposed. Through the new training method, we uncover that although such adjustments can impact the robustness of explanations, they do not have an influence on the robustness of classification. These findings not only challenge the previous assumption of a strong correlation between the two forms of robustness but also pave new pathways for understanding the relationship between loss landscape and explanation loss. Codes are provided in the supplement.<\/jats:p>","DOI":"10.1145\/3787470.3787476","type":"journal-article","created":{"date-parts":[[2026,1,1]],"date-time":"2026-01-01T00:46:21Z","timestamp":1767228381000},"page":"62-78","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Are Classification Robustness and Explanation Robustness Really Strongly Correlated? An Analysis Through Input Loss Landscape"],"prefix":"10.1145","volume":"27","author":[{"given":"Tiejin","family":"Chen","sequence":"first","affiliation":[{"name":"Arizona State University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Wenwang","family":"Huang","sequence":"additional","affiliation":[{"name":"Independent Researcher"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Linsey","family":"Pang","sequence":"additional","affiliation":[{"name":"Paypal AI"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Dongsheng","family":"Luo","sequence":"additional","affiliation":[{"name":"Florida International University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hua","family":"Wei","sequence":"additional","affiliation":[{"name":"Arizona State University"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2025,12,31]]},"reference":[{"key":"e_1_2_1_1_1","volume-title":"On the robustness of interpretability methods. arXiv preprint arXiv:1806.08049","author":"Alvarez-Melis D.","year":"2018","unstructured":"D. Alvarez-Melis and T. S. Jaakkola. On the robustness of interpretability methods. arXiv preprint arXiv:1806.08049, 2018."},{"key":"e_1_2_1_2_1","first-page":"274","volume-title":"International conference on machine learning","author":"Athalye A.","year":"2018","unstructured":"A. Athalye, N. Carlini, and D. Wagner. Obfuscated gradients give a false sense of security: Circumventing defenses to adversarial examples. In International conference on machine learning, pages 274--283. PMLR, 2018."},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1371\/journal.pone.0130140"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.5555\/1756006.1859912"},{"key":"e_1_2_1_5_1","first-page":"1014","volume-title":"International Conference on Machine Learning","author":"Boopathy A.","year":"2020","unstructured":"A. Boopathy, S. Liu, G. Zhang, C. Liu, P.-Y. Chen, S. Chang, and L. Daniel. Proper network interpretability helps adversarial robustness in classification. In International Conference on Machine Learning, pages 1014--1023. PMLR, 2020."},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-29135-8_5"},{"key":"e_1_2_1_7_1","volume-title":"(certified!!) adversarial robustness for free! arXiv preprint arXiv:2206.10550","author":"Carlini N.","year":"2022","unstructured":"N. Carlini, F. Tramer, K. D. Dvijotham, L. Rice, M. Sun, and J. Z. Kolter. (certified!!) adversarial robustness for free! arXiv preprint arXiv:2206.10550, 2022."},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/SP.2017.49"},{"key":"e_1_2_1_9_1","volume-title":"Unlabeled data improves adversarial robustness. Advances in neural information processing systems, 32","author":"Carmon Y.","year":"2019","unstructured":"Y. Carmon, A. Raghunathan, L. Schmidt, J. C. Duchi, and P. S. Liang. Unlabeled data improves adversarial robustness. Advances in neural information processing systems, 32, 2019."},{"key":"e_1_2_1_10_1","first-page":"32","article-title":"Robust attribution regularization","author":"Chen J.","year":"2019","unstructured":"J. Chen, X. Wu, V. Rastogi, Y. Liang, and S. Jha. Robust attribution regularization. Advances in Neural Information Processing Systems, 32, 2019.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP40776.2020.9054475"},{"key":"e_1_2_1_12_1","volume-title":"International Conference on Learning Representations","author":"Chen T.","year":"2020","unstructured":"T. Chen, Z. Zhang, S. Liu, S. Chang, and Z. Wang. Robust overfitting may be mitigated by properly learned smoothening. In International Conference on Learning Representations, 2020."},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"e_1_2_1_14_1","volume-title":"Explanations can be manipulated and geometry is to blame. Advances in neural information processing systems, 32","author":"Dombrowski A.-K.","year":"2019","unstructured":"A.-K. Dombrowski, M. Alber, C. Anders, M. Ackermann, K.-R. M\u00a8uller, and P. Kessel. Explanations can be manipulated and geometry is to blame. Advances in neural information processing systems, 32, 2019."},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v33i01.33013681"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00459"},{"key":"e_1_2_1_17_1","volume-title":"Explaining and harnessing adversarial examples. arXiv preprint arXiv:1412.6572","author":"Goodfellow I. J.","year":"2014","unstructured":"I. J. Goodfellow, J. Shlens, and C. Szegedy. Explaining and harnessing adversarial examples. arXiv preprint arXiv:1412.6572, 2014."},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_2_1_19_1","first-page":"2712","volume-title":"International conference on machine learning","author":"Hendrycks D.","year":"2019","unstructured":"D. Hendrycks, K. Lee, and M. Mazeika. Using pretraining can improve model robustness and uncertainty. In International conference on machine learning, pages 2712--2721. PMLR, 2019."},{"key":"e_1_2_1_20_1","volume-title":"Fooling neural network interpretations via adversarial model manipulation. Advances in neural information processing systems, 32","author":"Heo J.","year":"2019","unstructured":"J. Heo, S. Joo, and T. Moon. Fooling neural network interpretations via adversarial model manipulation. Advances in neural information processing systems, 32, 2019."},{"key":"e_1_2_1_21_1","volume-title":"Mobilenets: E!cient convolutional neural networks for mobile vision applications. arXiv preprint arXiv:1704.04861","author":"Howard A. G.","year":"2017","unstructured":"A. G. Howard, M. Zhu, B. Chen, D. Kalenichenko, W. Wang, T. Weyand, M. Andreetto, and H. Adam. Mobilenets: E!cient convolutional neural networks for mobile vision applications. arXiv preprint arXiv:1704.04861, 2017."},{"key":"e_1_2_1_22_1","first-page":"1988","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Huang W.","year":"2023","unstructured":"W. Huang, X. Zhao, G. Jin, and X. Huang. Safari: Versatile and e!cient evaluations for robustness of interpretability. In Proceedings of the IEEE\/CVF International Conference on Computer Vision, pages 1988-- 1998, 2023."},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.01578"},{"key":"e_1_2_1_24_1","volume-title":"Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980","author":"Kingma D. P.","year":"2014","unstructured":"D. P. Kingma and J. Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014."},{"key":"e_1_2_1_25_1","volume-title":"et al. Captum: A unified and generic model interpretability library for pytorch. arXiv preprint arXiv:2009.07896","author":"Kokhlikyan N.","year":"2020","unstructured":"N. Kokhlikyan, V. Miglani, M. Martin, E. Wang, B. Alsallakh, J. Reynolds, A. Melnikov, N. Kliushkina, C. Araya, S. Yan, et al. Captum: A unified and generic model interpretability library for pytorch. arXiv preprint arXiv:2009.07896, 2020."},{"key":"e_1_2_1_26_1","volume-title":"Learning multiple layers of features from tiny images","author":"Krizhevsky A.","year":"2009","unstructured":"A. Krizhevsky, G. Hinton, et al. Learning multiple layers of features from tiny images. 2009."},{"issue":"7","key":"e_1_2_1_27_1","first-page":"3","article-title":"Tiny imagenet visual recognition challenge","volume":"7","author":"Le Y.","year":"2015","unstructured":"Y. Le and X. Yang. Tiny imagenet visual recognition challenge. CS 231N, 7(7):3, 2015.","journal-title":"CS 231N"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1162\/neco.1989.1.4.541"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2022.109229"},{"key":"e_1_2_1_30_1","volume-title":"Defensive quantization: When e!ciency meets robustness. arXiv preprint arXiv:1904.08444","author":"Lin J.","year":"2019","unstructured":"J. Lin, C. Gan, and S. Han. Defensive quantization: When e!ciency meets robustness. arXiv preprint arXiv:1904.08444, 2019."},{"key":"e_1_2_1_31_1","first-page":"21476","article-title":"On the loss landscape of adversarial training: Identifying challenges and how to overcome them","volume":"33","author":"Liu C.","year":"2020","unstructured":"C. Liu, M. Salzmann, T. Lin, R. Tomioka, and S. S\u00a8usstrunk. On the loss landscape of adversarial training: Identifying challenges and how to overcome them. Advances in Neural Information Processing Systems, 33:21476--21487, 2020.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1145\/3422337.3447841"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIT.1982.1056489"},{"key":"e_1_2_1_34_1","volume-title":"Towards deep learning models resistant to adversarial attacks. arXiv preprint arXiv:1706.06083","author":"Madry A.","year":"2017","unstructured":"A. Madry, A. Makelov, L. Schmidt, D. Tsipras, and A. Vladu. Towards deep learning models resistant to adversarial attacks. arXiv preprint arXiv:1706.06083, 2017."},{"key":"e_1_2_1_35_1","volume-title":"Rethinking softmax cross-entropy loss for adversarial robustness. arXiv preprint arXiv:1905.10626","author":"Pang T.","year":"2019","unstructured":"T. Pang, K. Xu, Y. Dong, C. Du, N. Chen, and J. Zhu. Rethinking softmax cross-entropy loss for adversarial robustness. arXiv preprint arXiv:1905.10626, 2019."},{"key":"e_1_2_1_36_1","volume-title":"Bag of tricks for adversarial training. arXiv preprint arXiv:2010.00467","author":"Pang T.","year":"2020","unstructured":"T. Pang, X. Yang, Y. Dong, H. Su, and J. Zhu. Bag of tricks for adversarial training. arXiv preprint arXiv:2010.00467, 2020."},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1109\/SP.2016.41"},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00474"},{"key":"e_1_2_1_39_1","volume-title":"Grad-cam: Why did you say that? arXiv preprint arXiv:1611.07450","author":"Selvaraju R. R.","year":"2016","unstructured":"R. R. Selvaraju, A. Das, R. Vedantam, M. Cogswell, D. Parikh, and D. Batra. Grad-cam: Why did you say that? arXiv preprint arXiv:1611.07450, 2016."},{"key":"e_1_2_1_40_1","first-page":"32","article-title":"Adversarial training for free! Advances","author":"Shafahi A.","year":"2019","unstructured":"A. Shafahi, M. Najibi, M. A. Ghiasi, Z. Xu, J. Dickerson, C. Studer, L. S. Davis, G. Taylor, and T. Goldstein. Adversarial training for free! Advances in Neural Information Processing Systems, 32, 2019.","journal-title":"Neural Information Processing Systems"},{"key":"e_1_2_1_41_1","first-page":"3145","volume-title":"International conference on machine learning","author":"Shrikumar A.","year":"2017","unstructured":"A. Shrikumar, P. Greenside, and A. Kundaje. Learning important features through propagating activation differences. In International conference on machine learning, pages 3145--3153. PMLR, 2017."},{"key":"e_1_2_1_42_1","volume-title":"Deep inside convolutional networks: Visualising image classification models and saliency maps. arXiv preprint arXiv:1312.6034","author":"Simonyan K.","year":"2013","unstructured":"K. Simonyan, A. Vedaldi, and A. Zisserman. Deep inside convolutional networks: Visualising image classification models and saliency maps. arXiv preprint arXiv:1312.6034, 2013."},{"key":"e_1_2_1_43_1","volume-title":"Deep inside convolutional networks: Visualising image classification models and saliency maps. arxiv","author":"Simonyan K.","year":"2013","unstructured":"K. Simonyan, A. Vedaldi, and A. Zisserman. Deep inside convolutional networks: Visualising image classification models and saliency maps. arxiv 2013. arXiv preprint arXiv:1312.6034, 2019."},{"key":"e_1_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1145\/3375627.3375830"},{"key":"e_1_2_1_45_1","volume-title":"Pixeldefend: Leveraging generative models to understand and defend against adversarial examples. arXiv preprint arXiv:1710.10766","author":"Song Y.","year":"2017","unstructured":"Y. Song, T. Kim, S. Nowozin, S. Ermon, and N. Kushman. Pixeldefend: Leveraging generative models to understand and defend against adversarial examples. arXiv preprint arXiv:1710.10766, 2017."},{"key":"e_1_2_1_46_1","volume-title":"Striving for simplicity: The all convolutional net. arXiv preprint arXiv:1412.6806","author":"Springenberg J. T.","year":"2014","unstructured":"J. T. Springenberg, A. Dosovitskiy, T. Brox, and M. Riedmiller. Striving for simplicity: The all convolutional net. arXiv preprint arXiv:1412.6806, 2014."},{"key":"e_1_2_1_47_1","first-page":"3319","volume-title":"International conference on machine learning","author":"Sundararajan M.","year":"2017","unstructured":"M. Sundararajan, A. Taly, and Q. Yan. Axiomatic attribution for deep networks. In International conference on machine learning, pages 3319--3328. PMLR, 2017."},{"key":"e_1_2_1_48_1","volume-title":"Intriguing properties of neural networks. arXiv preprint arXiv:1312.6199","author":"Szegedy C.","year":"2013","unstructured":"C. Szegedy, W. Zaremba, I. Sutskever, J. Bruna, D. Erhan, I. Goodfellow, and R. Fergus. Intriguing properties of neural networks. arXiv preprint arXiv:1312.6199, 2013."},{"key":"e_1_2_1_49_1","volume-title":"Foiling explanations in deep neural networks. arXiv preprint arXiv:2211.14860","author":"Tamam S. V.","year":"2022","unstructured":"S. V. Tamam, R. Lapid, and M. Sipper. Foiling explanations in deep neural networks. arXiv preprint arXiv:2211.14860, 2022."},{"key":"e_1_2_1_50_1","volume-title":"Defense against explanation manipulation. Frontiers in big Data, 5:704203","author":"Tang R.","year":"2022","unstructured":"R. Tang, N. Liu, F. Yang, N. Zou, and X. Hu. Defense against explanation manipulation. Frontiers in big Data, 5:704203, 2022."},{"key":"e_1_2_1_51_1","volume-title":"Ensemble adversarial training: Attacks and defenses. arXiv preprint arXiv:1705.07204","author":"Tram'er F.","year":"2017","unstructured":"F. Tram'er, A. Kurakin, N. Papernot, I. Goodfellow, D. Boneh, and P. McDaniel. Ensemble adversarial training: Attacks and defenses. arXiv preprint arXiv:1705.07204, 2017."},{"key":"e_1_2_1_52_1","volume-title":"Robust models are more interpretable because attributions look normal. arXiv preprint arXiv:2103.11257","author":"Wang Z.","year":"2021","unstructured":"Z. Wang, M. Fredrikson, and A. Datta. Robust models are more interpretable because attributions look normal. arXiv preprint arXiv:2103.11257, 2021."},{"key":"e_1_2_1_53_1","volume-title":"Better di\"usion models further improve adversarial training. arXiv preprint arXiv:2302.04638","author":"Wang Z.","year":"2023","unstructured":"Z. Wang, T. Pang, C. Du, M. Lin, W. Liu, and S. Yan. Better di\"usion models further improve adversarial training. arXiv preprint arXiv:2302.04638, 2023."},{"key":"e_1_2_1_54_1","volume-title":"Robust explanation constraints for neural networks. arXiv preprint arXiv:2212.08507","author":"Wicker M.","year":"2022","unstructured":"M. Wicker, J. Heo, L. Costabello, and A. Weller. Robust explanation constraints for neural networks. arXiv preprint arXiv:2212.08507, 2022."},{"key":"e_1_2_1_55_1","first-page":"2958","article-title":"Adversarial weight perturbation helps robust generalization","volume":"33","author":"Wu D.","year":"2020","unstructured":"D. Wu, S.-T. Xia, and Y. Wang. Adversarial weight perturbation helps robust generalization. Advances in Neural Information Processing Systems, 33:2958--2969, 2020.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_2_1_56_1","volume-title":"Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms. arXiv preprint arXiv:1708.07747","author":"Xiao H.","year":"2017","unstructured":"H. Xiao, K. Rasul, and R. Vollgraf. Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms. arXiv preprint arXiv:1708.07747, 2017."},{"key":"e_1_2_1_57_1","volume-title":"Smooth adversarial training. arXiv preprint arXiv:2006.14536","author":"Xie C.","year":"2020","unstructured":"C. Xie, M. Tan, B. Gong, A. Yuille, and Q. V. Le. Smooth adversarial training. arXiv preprint arXiv:2006.14536, 2020."},{"key":"e_1_2_1_58_1","volume-title":"Feature squeezing: Detecting adversarial examples in deep neural networks. arXiv preprint arXiv:1704.01155","author":"Xu W.","year":"2017","unstructured":"W. Xu, D. Evans, and Y. Qi. Feature squeezing: Detecting adversarial examples in deep neural networks. arXiv preprint arXiv:1704.01155, 2017."},{"key":"e_1_2_1_59_1","volume-title":"Wide residual networks. arXiv preprint arXiv:1605.07146","author":"Zagoruyko S.","year":"2016","unstructured":"S. Zagoruyko and N. Komodakis. Wide residual networks. arXiv preprint arXiv:1605.07146, 2016."},{"key":"e_1_2_1_60_1","first-page":"7472","volume-title":"International conference on machine learning","author":"Zhang H.","year":"2019","unstructured":"H. Zhang, Y. Yu, J. Jiao, E. Xing, L. El Ghaoui, and M. Jordan. Theoretically principled trade-o\" between robustness and accuracy. In International conference on machine learning, pages 7472--7482. PMLR, 2019."},{"key":"e_1_2_1_61_1","volume-title":"29th {USENIX} Security Symposium ({USENIX} Security 20)","author":"Zhang X.","year":"2020","unstructured":"X. Zhang, N. Wang, H. Shen, S. Ji, X. Luo, and T. Wang. Interpretable deep learning under fire. In 29th {USENIX} Security Symposium ({USENIX} Security 20), 2020. 74"}],"container-title":["ACM SIGKDD Explorations Newsletter"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3787470.3787476","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,1,13]],"date-time":"2026-01-13T00:43:34Z","timestamp":1768265014000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3787470.3787476"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,12,30]]},"references-count":61,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2025,12,30]]}},"alternative-id":["10.1145\/3787470.3787476"],"URL":"https:\/\/doi.org\/10.1145\/3787470.3787476","relation":{},"ISSN":["1931-0145","1931-0153"],"issn-type":[{"value":"1931-0145","type":"print"},{"value":"1931-0153","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,12,30]]},"assertion":[{"value":"2025-12-31","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}