{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,12]],"date-time":"2026-07-12T14:13:07Z","timestamp":1783865587406,"version":"3.55.0"},"reference-count":235,"publisher":"Association for Computing Machinery (ACM)","issue":"9","license":[{"start":{"date-parts":[[2025,5,6]],"date-time":"2025-05-06T00:00:00Z","timestamp":1746489600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001659","name":"Deutsche Forschungsgemeinschaft","doi-asserted-by":"crossref","award":["453349072"],"award-info":[{"award-number":["453349072"]}],"id":[{"id":"10.13039\/501100001659","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Comput. Surv."],"published-print":{"date-parts":[[2025,9,30]]},"abstract":"<jats:p>\n            Measuring similarity of neural networks to understand and improve their behavior has become an issue of great importance and research interest. In this survey, we provide a comprehensive overview of two complementary perspectives of measuring neural network similarity: (i) representational similarity, which considers how\n            <jats:italic>activations<\/jats:italic>\n            of intermediate layers differ, and (ii) functional similarity, which considers how models differ in their\n            <jats:italic>outputs<\/jats:italic>\n            . In addition to providing detailed descriptions of existing measures, we summarize and discuss results on the properties of and relationships between these measures, and point to open research problems. We hope our work lays a foundation for more systematic research on the properties and applicability of similarity measures for neural network models.\n          <\/jats:p>","DOI":"10.1145\/3728458","type":"journal-article","created":{"date-parts":[[2025,4,8]],"date-time":"2025-04-08T11:44:54Z","timestamp":1744112694000},"page":"1-52","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":23,"title":["Similarity of Neural Network Models: A Survey of Functional and Representational Measures"],"prefix":"10.1145","volume":"57","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-7805-4725","authenticated-orcid":false,"given":"Max","family":"Klabunde","sequence":"first","affiliation":[{"name":"University of Passau, Passau, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3091-5095","authenticated-orcid":false,"given":"Tobias","family":"Schumacher","sequence":"additional","affiliation":[{"name":"University of Mannheim, Mannheim, Germany and RWTH Aachen University, Aachen, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5485-5720","authenticated-orcid":false,"given":"Markus","family":"Strohmaier","sequence":"additional","affiliation":[{"name":"University of Mannheim, Mannheim, Germany, GESIS - Leibniz Institute for the Social Sciences, Mannheim, Germany and Complexity Science Hub Vienna, Wien, Austria"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7620-1376","authenticated-orcid":false,"given":"Florian","family":"Lemmerich","sequence":"additional","affiliation":[{"name":"University of Passau, Passau, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,5,6]]},"reference":[{"issue":"1","key":"e_1_3_3_2_2","article-title":"Evaluating explainability for graph neural networks","volume":"10","author":"Agarwal Chirag","year":"2023","unstructured":"Chirag Agarwal, Owen Queen, Himabindu Lakkaraju, and Marinka Zitnik. 2023. Evaluating explainability for graph neural networks. Sci. Data 10, 1 (2023).","journal-title":"Sci. Data"},{"key":"e_1_3_3_3_2","volume-title":"NeurIPS","author":"Ansuini Alessio","year":"2019","unstructured":"Alessio Ansuini, Alessandro Laio, Jakob H. Macke, and Davide Zoccolan. 2019. Intrinsic dimension of data representations in deep neural networks. In NeurIPS."},{"issue":"1","key":"e_1_3_3_4_2","article-title":"An algorithm for computing the capacity of arbitrary discrete memoryless channels","volume":"18","author":"Arimoto Suguru","year":"1972","unstructured":"Suguru Arimoto. 1972. An algorithm for computing the capacity of arbitrary discrete memoryless channels. IEEE Trans. Intell. Transport. Syst. 18, 1 (1972).","journal-title":"IEEE Trans. Intell. Transport. Syst."},{"key":"e_1_3_3_5_2","article-title":"Do deep nets really need to be deep?","volume":"27","author":"Ba Jimmy","year":"2014","unstructured":"Jimmy Ba and Rich Caruana. 2014. Do deep nets really need to be deep? NeurIPS 27 (2014).","journal-title":"NeurIPS"},{"key":"e_1_3_3_6_2","volume-title":"ICML","author":"Balogh Andr\u00e1s","year":"2023","unstructured":"Andr\u00e1s Balogh and M\u00e1rk Jelasity. 2023. On the functional similarity of robust and non-robust neural representations. In ICML."},{"issue":"1","key":"e_1_3_3_7_2","article-title":"Beyond kappa: A review of interrater agreement measures","volume":"27","author":"Banerjee Mousumi","year":"1999","unstructured":"Mousumi Banerjee, Michelle Capozzoli, Laura McSweeney, and Debajyoti Sinha. 1999. Beyond kappa: A review of interrater agreement measures. Can. J. Stat. 27, 1 (1999).","journal-title":"Can. J. Stat."},{"key":"e_1_3_3_8_2","volume-title":"NeurIPS","author":"Bansal Yamini","year":"2021","unstructured":"Yamini Bansal, Preetum Nakkiran, and Boaz Barak. 2021. Revisiting model stitching to compare neural representations. In NeurIPS."},{"key":"e_1_3_3_9_2","volume-title":"ICML","author":"Barannikov Serguei","year":"2022","unstructured":"Serguei Barannikov, Ilya Trofimov, Nikita Balabin, and Evgeny Burnaev. 2022. Representation topology divergence: A method for comparing neural network representations. In ICML."},{"key":"e_1_3_3_10_2","unstructured":"Lorenzo Basile Santiago Acevedo Luca Bortolussi Fabio Anselmi and Alex Rodriguez. 2025. Intrinsic dimension correlation: uncovering nonlinear connections in multimodal representations. In The Thirteenth International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=Qj1KwBZaEI"},{"key":"e_1_3_3_11_2","volume-title":"ICLR","author":"Bau Anthony","year":"2019","unstructured":"Anthony Bau, Yonatan Belinkov, Hassan Sajjad, Nadir Durrani, Fahim Dalvi, and James R. Glass. 2019. Identifying and controlling important neurons in neural machine translation. In ICLR."},{"issue":"5","key":"e_1_3_3_12_2","article-title":"The intrinsic dimensionality of signal collections","volume":"15","author":"Bennett Robert","year":"1969","unstructured":"Robert Bennett. 1969. The intrinsic dimensionality of signal collections. IEEE Trans. Intell. Transport. Syst. 15, 5 (1969).","journal-title":"IEEE Trans. Intell. Transport. Syst."},{"key":"e_1_3_3_13_2","volume-title":"Positive Definite Matrices","author":"Bhatia Rajendra","year":"2007","unstructured":"Rajendra Bhatia. 2007. Positive Definite Matrices."},{"issue":"2","key":"e_1_3_3_14_2","article-title":"On the Bures-Wasserstein distance between positive definite matrices","volume":"37","author":"Bhatia Rajendra","year":"2019","unstructured":"Rajendra Bhatia, Tanvi Jain, and Yongdo Lim. 2019. On the Bures-Wasserstein distance between positive definite matrices. Exposit. Math. 37, 2 (2019).","journal-title":"Exposit. Math."},{"key":"e_1_3_3_15_2","article-title":"On the reproducibility of neural network predictions","author":"Bhojanapalli Srinadh","year":"2021","unstructured":"Srinadh Bhojanapalli, Kimberly Wilber, Andreas Veit, Ankit Singh Rawat, Seungyeon Kim, Aditya Menon, and Sanjiv Kumar. 2021. On the reproducibility of neural network predictions. arXiv:2102.03349. Retrieved from https:\/\/arxiv.org\/abs\/2102.03349","journal-title":"arXiv:2102.03349"},{"key":"e_1_3_3_16_2","doi-asserted-by":"crossref","unstructured":"Emily Black and Matt Fredrikson. 2021. Leave-one-out unfairness. In FAccT \u201921.","DOI":"10.1145\/3442188.3445894"},{"issue":"4","key":"e_1_3_3_17_2","article-title":"Computation of channel capacity and rate-distortion functions","volume":"18","author":"Blahut Richard","year":"1972","unstructured":"Richard Blahut. 1972. Computation of channel capacity and rate-distortion functions. IEEE Trans. Intell. Transport. Syst. 18, 4 (1972).","journal-title":"IEEE Trans. Intell. Transport. Syst."},{"key":"e_1_3_3_18_2","volume-title":"NeurIPS","author":"Boix-Adsera Enric","year":"2022","unstructured":"Enric Boix-Adsera, Hannah Lawrence, George Stepaniants, and Philippe Rigollet. 2022. GULP: A prediction-based metric between representations. In NeurIPS."},{"key":"e_1_3_3_19_2","article-title":"Pros and cons of GAN evaluation measures","volume":"179","author":"Borji Ali","year":"2019","unstructured":"Ali Borji. 2019. Pros and cons of GAN evaluation measures. Comput. Vis. Image Understand. 179 (2019).","journal-title":"Comput. Vis. Image Understand."},{"issue":"1","key":"e_1_3_3_20_2","article-title":"Diversity creation methods: A survey and categorisation","volume":"6","author":"Brown Gavin","year":"2005","unstructured":"Gavin Brown, Jeremy Wyatt, Rachel Harris, and Xin Yao. 2005. Diversity creation methods: A survey and categorisation. Inf. Fus. 6, 1 (2005).","journal-title":"Inf. Fus."},{"key":"e_1_3_3_21_2","article-title":"An extension of Kakutani\u2019s theorem on infinite product measures to the tensor product of semifinite w*-algebras","volume":"135","author":"Bures Donald","year":"1969","unstructured":"Donald Bures. 1969. An extension of Kakutani\u2019s theorem on infinite product measures to the tensor product of semifinite w*-algebras. Trans. Am. Math. Soc. 135 (1969).","journal-title":"Trans. Am. Math. Soc."},{"key":"e_1_3_3_22_2","article-title":"Intrinsic dimension estimation: Advances and open problems","volume":"328","author":"Camastra Francesco","year":"2016","unstructured":"Francesco Camastra and Antonino Staiano. 2016. Intrinsic dimension estimation: Advances and open problems. Inf. Sci. 328 (2016).","journal-title":"Inf. Sci."},{"key":"e_1_3_3_23_2","article-title":"Evaluation of text generation: A survey","author":"Celikyilmaz Asli","year":"2020","unstructured":"Asli Celikyilmaz, Elizabeth Clark, and Jianfeng Gao. 2020. Evaluation of text generation: A survey. arXiv:2006.14799 (2020).","journal-title":"arXiv:2006.14799"},{"issue":"4","key":"e_1_3_3_24_2","article-title":"Comprehensive survey on distance\/similarity measures between probability density functions","volume":"1","author":"Cha Sung-Hyuk","year":"2007","unstructured":"Sung-Hyuk Cha. 2007. Comprehensive survey on distance\/similarity measures between probability density functions. Int. J. Math. Models Methods Appl. Sci. 1, 4 (2007).","journal-title":"Int. J. Math. Models Methods Appl. Sci."},{"key":"e_1_3_3_25_2","article-title":"Graph-based similarity of neural network representations","author":"Chen Zuohui","year":"2021","unstructured":"Zuohui Chen, Yao Lu, Wen Yang, Qi Xuan, and Xiaoniu Yang. 2021. Graph-based similarity of neural network representations. arXiv:2111.11165. Retrieved from https:\/\/arxiv.org\/abs\/2111.11165","journal-title":"arXiv:2111.11165"},{"key":"e_1_3_3_26_2","volume-title":"NeurIPS","author":"Cianfarani Christian","year":"2022","unstructured":"Christian Cianfarani, Arjun Nitin Bhagoji, Vikash Sehwag, Ben Zhao, Heather Zheng, and Prateek Mittal. 2022. Understanding robust learning through the lens of representation similarities. In NeurIPS."},{"key":"e_1_3_3_27_2","volume-title":"UniReps: 2nd Edition of the Workshop on Unifying Representations in Neural Models","author":"Cloos Nathan","year":"2024","unstructured":"Nathan Cloos, Guangyu Robert Yang, and Christopher J. Cueva. 2024. A framework for standardizing similarity measures in a rapidly evolving field. In UniReps: 2nd Edition of the Workshop on Unifying Representations in Neural Models."},{"issue":"1","key":"e_1_3_3_28_2","article-title":"A coefficient of agreement for nominal scales","volume":"20","author":"Cohen Jacob","year":"1960","unstructured":"Jacob Cohen. 1960. A coefficient of agreement for nominal scales. Educ. Psychol. Meas. 20, 1 (1960).","journal-title":"Educ. Psychol. Meas."},{"issue":"1","key":"e_1_3_3_29_2","article-title":"Algorithms for learning kernels based on centered alignment","volume":"13","author":"Cortes Corinna","year":"2012","unstructured":"Corinna Cortes, Mehryar Mohri, and Afshin Rostamizadeh. 2012. Algorithms for learning kernels based on centered alignment. J. Mach. Learn. Res. 13, 1 (2012).","journal-title":"J. Mach. Learn. Res."},{"key":"e_1_3_3_30_2","volume-title":"NeurIPS","author":"Cristianini Nello","year":"2001","unstructured":"Nello Cristianini, John Shawe-Taylor, Andr\u00e9 Elisseeff, and Jaz S. Kandola. 2001. On kernel-target alignment. In NeurIPS."},{"key":"e_1_3_3_31_2","volume-title":"NeurIPS","author":"Csisz\u00e1rik Adri\u00e1n","year":"2021","unstructured":"Adri\u00e1n Csisz\u00e1rik, P\u00e9ter Kor\u00f6si-Szab\u00f3, \u00c1kos K. Matszangosz, Gergely Papp, and D\u00e1niel Varga. 2021. Similarity and matching of neural network representations. In NeurIPS."},{"key":"e_1_3_3_32_2","volume-title":"ACL","author":"Datta Arghya","year":"2023","unstructured":"Arghya Datta, Subhrangshu Nandi, Jingcheng Xu, Greg Ver Steeg, He Xie, Anoop Kumar, and Aram Galstyan. 2023. Measuring and mitigating local instability in deep neural networks. In ACL."},{"key":"e_1_3_3_33_2","first-page":"37","article-title":"Combinatorics and geometry of transportation polytopes: An update.","volume":"625","author":"Loera Jes\u00fas A. De","year":"2013","unstructured":"Jes\u00fas A. De Loera and Edward D. Kim. 2013. Combinatorics and geometry of transportation polytopes: An update. Discr. Geom. Algebr. Combin. 625 (2013), 37\u201376.","journal-title":"Discr. Geom. Algebr. Combin."},{"key":"e_1_3_3_34_2","volume-title":"NAACL-HLT","author":"Devlin Jacob","year":"2019","unstructured":"Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In NAACL-HLT."},{"key":"e_1_3_3_35_2","volume-title":"NeurIPS","author":"Ding Frances","year":"2021","unstructured":"Frances Ding, Jean-Stanislas Denain, and Jacob Steinhardt. 2021. Grounding representation similarity through statistical testing. In NeurIPS."},{"key":"e_1_3_3_36_2","article-title":"SimSMoE: Solving representational collapse via similarity measure","author":"Do Giang","year":"2024","unstructured":"Giang Do, Hung Le, and Truyen Tran. 2024. SimSMoE: Solving representational collapse via similarity measure. arXiv:2406.15883. Retrieved from https:\/\/arxiv.org\/abs\/2406.15883","journal-title":"arXiv:2406.15883"},{"key":"e_1_3_3_37_2","volume-title":"ACL","author":"Du Yupei","year":"2023","unstructured":"Yupei Du and Dong Nguyen. 2023. Measuring the instability of fine-tuning. In ACL."},{"key":"e_1_3_3_38_2","volume-title":"ICLR","author":"Duong Lyndon","year":"2023","unstructured":"Lyndon Duong, Jingyang Zhou, Josue Nassar, Jules Berman, Jeroen Olieslagers, and Alex H. Williams. 2023. Representational dissimilarity metric spaces for stochastic neural networks. In ICLR."},{"key":"e_1_3_3_39_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2013.57"},{"key":"e_1_3_3_40_2","volume-title":"NeurIPS","author":"Fard Mahdi Milani","year":"2016","unstructured":"Mahdi Milani Fard, Quentin Cormier, Kevin Robert Canini, and Maya R. Gupta. 2016. Launch and iterate: Reducing prediction churn. In NeurIPS."},{"issue":"1","key":"e_1_3_3_41_2","article-title":"Mistakes and how to avoid mistakes in using intercoder reliability indices.","volume":"11","author":"Feng Guangchao Charles","year":"2014","unstructured":"Guangchao Charles Feng. 2014. Mistakes and how to avoid mistakes in using intercoder reliability indices. Methodol: Eur. J. Res. Methods Behav. Soc. Sci. 11, 1 (2014).","journal-title":"Methodol: Eur. J. Res. Methods Behav. Soc. Sci."},{"key":"e_1_3_3_42_2","article-title":"Transferred discrepancy: Quantifying the difference between representations","author":"Feng Yunzhen","year":"2020","unstructured":"Yunzhen Feng, Runtian Zhai, Di He, Liwei Wang, and Bin Dong. 2020. Transferred discrepancy: Quantifying the difference between representations. arXiv:2007.12446. Retrieved from https:\/\/arxiv.org\/abs\/2007.12446","journal-title":"arXiv:2007.12446"},{"issue":"5","key":"e_1_3_3_43_2","article-title":"Measuring nominal scale agreement among many raters.","volume":"76","author":"Fleiss Joseph L.","year":"1971","unstructured":"Joseph L. Fleiss. 1971. Measuring nominal scale agreement among many raters. Psychol. Bull. 76, 5 (1971).","journal-title":"Psychol. Bull."},{"key":"e_1_3_3_44_2","article-title":"Deep ensembles: A loss landscape perspective","author":"Fort Stanislav","year":"2019","unstructured":"Stanislav Fort, Huiyi Hu, and Balaji Lakshminarayanan. 2019. Deep ensembles: A loss landscape perspective. arXiv:1912.02757. Retrieved from https:\/\/arxiv.org\/abs\/1912.02757","journal-title":"arXiv:1912.02757"},{"key":"e_1_3_3_45_2","volume-title":"NeurIPS","author":"Geirhos Robert","year":"2020","unstructured":"Robert Geirhos, Kristof Meding, and Felix A. Wichmann. 2020. Beyond accuracy: Quantifying trial-by-trial behaviour of CNNs and humans by measuring error consistency. In NeurIPS."},{"issue":"12","key":"e_1_3_3_46_2","article-title":"Community structure in social and biological networks","volume":"99","author":"Girvan Michelle","year":"2002","unstructured":"Michelle Girvan and Mark E. J. Newman. 2002. Community structure in social and biological networks. Natl. Acad. Sci. 99, 12 (2002).","journal-title":"Natl. Acad. Sci."},{"issue":"3","key":"e_1_3_3_47_2","article-title":"Interrater agreement and interrater reliability: Key concepts, approaches, and applications","volume":"9","author":"Gisev Natasa","year":"2013","unstructured":"Natasa Gisev, J. Simon Bell, and Timothy F. Chen. 2013. Interrater agreement and interrater reliability: Key concepts, approaches, and applications. Res. Soc. Admin. Pharm. 9, 3 (2013).","journal-title":"Res. Soc. Admin. Pharm."},{"key":"e_1_3_3_48_2","volume-title":"NeurIPS","author":"Godfrey Charles","year":"2022","unstructured":"Charles Godfrey, Davis Brown, Tegan Emerson, and Henry Kvinge. 2022. On the symmetries of deep learning models and their internal representations. In NeurIPS."},{"key":"e_1_3_3_49_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10710-017-9314-z"},{"key":"e_1_3_3_50_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-021-01453-z"},{"key":"e_1_3_3_51_2","volume-title":"Algorithmic Learning Theory","author":"Gretton Arthur","year":"2005","unstructured":"Arthur Gretton, Olivier Bousquet, Alex Smola, and Bernhard Sch\u00f6lkopf. 2005. Measuring statistical dependence with hilbert-schmidt norms. In Algorithmic Learning Theory."},{"key":"e_1_3_3_52_2","volume-title":"NeurIPS 2021 Workshop on Self-Supervised Learning","author":"Grigg Tom George","year":"2021","unstructured":"Tom George Grigg, Dan Busbridge, Jason Ramapuram, and Russ Webb. 2021. Do self-supervised and supervised methods learn similar visual representations? In NeurIPS 2021 Workshop on Self-Supervised Learning."},{"key":"e_1_3_3_53_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.00942"},{"key":"e_1_3_3_54_2","volume-title":"EMNLP","author":"Hamilton William L.","year":"2016","unstructured":"William L. Hamilton, Jure Leskovec, and Dan Jurafsky. 2016. Cultural shift or linguistic drift? Comparing two computational measures of semantic change. In EMNLP."},{"key":"e_1_3_3_55_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P16-1141"},{"key":"e_1_3_3_56_2","volume-title":"Proceedings of UniReps: The First Workshop on Unifying Representations in Neural Models","author":"Harvey Sarah E.","year":"2024","unstructured":"Sarah E. Harvey, Brett W. Larsen, and Alex H. Williams. 2024. Duality of bures and shape distances with implications for comparing neural representations. In Proceedings of UniReps: The First Workshop on Unifying Representations in Neural Models."},{"key":"e_1_3_3_57_2","volume-title":"Algebraic Topology","author":"Hatcher Allen","year":"2005","unstructured":"Allen Hatcher. 2005. Algebraic Topology."},{"key":"e_1_3_3_58_2","unstructured":"Lucas Hayne Heejung Jung and R. Carter. 2024. Does representation similarity capture function similarity? Transactions on Machine Learning Research (2024). Retrieved from https:\/\/openreview.net\/forum?id=YY2iA0hfia"},{"key":"e_1_3_3_59_2","volume-title":"NeurIPS","author":"Hermann Katherine","year":"2020","unstructured":"Katherine Hermann and Andrew Lampinen. 2020. What shapes feature representations? Exploring datasets, architectures, and training. In NeurIPS."},{"issue":"2","key":"e_1_3_3_60_2","article-title":"The most predictable criterion","volume":"26","author":"Hotelling Harald","year":"1935","unstructured":"Harald Hotelling. 1935. The most predictable criterion. J. Educ. Psychol. 26, 2 (1935).","journal-title":"J. Educ. Psychol."},{"issue":"3","key":"e_1_3_3_61_2","article-title":"Relations between two sets of variates","volume":"28","author":"Hotelling Harold","year":"1936","unstructured":"Harold Hotelling. 1936. Relations between two sets of variates. Biometrika 28, 3\/4 (1936).","journal-title":"Biometrika"},{"key":"e_1_3_3_62_2","article-title":"Inter-layer information similarity assessment of deep neural networks via topological similarity and persistence analysis of data neighbour dynamics","author":"Hryniowski Andrew","year":"2020","unstructured":"Andrew Hryniowski and Alexander Wong. 2020. Inter-layer information similarity assessment of deep neural networks via topological similarity and persistence analysis of data neighbour dynamics. arXiv:2012.03793. Retrieved from https:\/\/arxiv.org\/abs\/2012.03793","journal-title":"arXiv:2012.03793"},{"key":"e_1_3_3_63_2","volume-title":"NeurIPS","author":"Hsu Hsiang","year":"2022","unstructured":"Hsiang Hsu and Flavio Calmon. 2022. Rashomon capacity: A metric for predictive multiplicity in classification. In NeurIPS."},{"key":"e_1_3_3_64_2","article-title":"Similarity of neural architectures based on input gradient transferability","author":"Hwang Jaehui","year":"2023","unstructured":"Jaehui Hwang, Dongyoon Han, Byeongho Heo, Song Park, Sanghyuk Chun, and Jong-Seok Lee. 2023. Similarity of neural architectures based on input gradient transferability. arXiv:2210.11407. Retrieved from https:\/\/arxiv.org\/abs\/2210.11407","journal-title":"arXiv:2210.11407"},{"key":"e_1_3_3_65_2","volume-title":"UAI","author":"Jones Haydn T.","year":"2022","unstructured":"Haydn T. Jones, Jacob M. Springer, Garrett T. Kenyon, and Juston S. Moore. 2022. If you\u2019ve trained one you\u2019ve trained them all: Inter-architecture similarity increases with robustness. In UAI."},{"key":"e_1_3_3_66_2","volume-title":"Proceedings of UniReps: The First Workshop on Unifying Representations in Neural Models","author":"Khosla Meenakshi","year":"2024","unstructured":"Meenakshi Khosla and Alex H. Williams. 2024. Soft matching distance: A metric on neural representations that captures single-neuron tuning. In Proceedings of UniReps: The First Workshop on Unifying Representations in Neural Models."},{"key":"e_1_3_3_67_2","volume-title":"ICML","author":"Khrulkov Valentin","year":"2018","unstructured":"Valentin Khrulkov and Ivan Oseledets. 2018. Geometry score: A method for comparing generative adversarial networks. In ICML."},{"key":"e_1_3_3_68_2","volume-title":"ICLR","author":"Kingma Diederik P.","year":"2014","unstructured":"Diederik P. Kingma and Max Welling. 2014. Auto-encoding variational bayes. In ICLR."},{"key":"e_1_3_3_69_2","volume-title":"UniReps: The First Workshop on Unifying Representations in Neural Models","author":"Klabunde Max","year":"2023","unstructured":"Max Klabunde, Mehdi Ben Amor, Michael Granitzer, and Florian Lemmerich. 2023. Towards measuring representational similarity of large language models. In UniReps: The First Workshop on Unifying Representations in Neural Models."},{"key":"e_1_3_3_70_2","volume-title":"ECML PKDD","author":"Klabunde Max","year":"2022","unstructured":"Max Klabunde and Florian Lemmerich. 2022. On the prediction instability of graph neural networks. In ECML PKDD."},{"key":"e_1_3_3_71_2","unstructured":"Max Klabunde Tassilo Wald Tobias Schumacher Klaus Maier-Hein Markus Strohmaier and Florian Lemmerich. 2025. ReSi: A comprehensive benchmark for representational similarity measures. In The Thirteenth International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=PRvdO3nfFi"},{"key":"e_1_3_3_72_2","article-title":"Pointwise representational similarity","author":"Kolling Camila","year":"2023","unstructured":"Camila Kolling, Till Speicher, Vedant Nanda, Mariya Toneva, and Krishna P. Gummadi. 2023. Pointwise representational similarity. arXiv:2305.19294. Retrieved from https:\/\/arxiv.org\/abs\/2305.19294","journal-title":"arXiv:2305.19294"},{"key":"e_1_3_3_73_2","volume-title":"NeurIPS","author":"Kornblith Simon","year":"2021","unstructured":"Simon Kornblith, Ting Chen, Honglak Lee, and Mohammad Norouzi. 2021. Why do better loss functions lead to less transferable features? In NeurIPS."},{"key":"e_1_3_3_74_2","volume-title":"ICML","author":"Kornblith Simon","year":"2019","unstructured":"Simon Kornblith, Mohammad Norouzi, Honglak Lee, and Geoffrey E. Hinton. 2019. Similarity of neural network representations revisited. In ICML."},{"key":"e_1_3_3_75_2","article-title":"Representational similarity analysis - connecting the branches of systems neuroscience","volume":"2","author":"Kriegeskorte Nikolaus","year":"2008","unstructured":"Nikolaus Kriegeskorte, Marieke Mur, and Peter Bandettini. 2008. Representational similarity analysis - connecting the branches of systems neuroscience. Front. Syst. Neurosci. 2 (2008).","journal-title":"Front. Syst. Neurosci."},{"key":"e_1_3_3_76_2","volume-title":"EMNLP","author":"Kudugunta Sneha","year":"2019","unstructured":"Sneha Kudugunta, Ankur Bapna, Isaac Caswell, and Orhan Firat. 2019. Investigating multilingual NMT representations at scale. In EMNLP."},{"key":"e_1_3_3_77_2","article-title":"Measures of diversity in classifier ensembles and their relationship with the ensemble accuracy","volume":"51","author":"Kuncheva Ludmila I.","year":"2003","unstructured":"Ludmila I. Kuncheva and Christopher J. Whitaker. 2003. Measures of diversity in classifier ensembles and their relationship with the ensemble accuracy. Mach. Learn. 51 (2003).","journal-title":"Mach. Learn."},{"key":"e_1_3_3_78_2","unstructured":"Richard D. Lange David Rolnick and Konrad Kording. 2022. Clustering units in neural networks: Upstream vs downstream information. Transactions on Machine Learning Research (2022). Retrieved from https:\/\/openreview.net\/forum?id=Euf7KofunK"},{"issue":"1","key":"e_1_3_3_79_2","article-title":"A generalization of fisher\u2019s z test","volume":"30","author":"Lawley Derrik N.","year":"1938","unstructured":"Derrik N. Lawley. 1938. A generalization of fisher\u2019s z test. Biometrika 30, 1\/2 (1938).","journal-title":"Biometrika"},{"key":"e_1_3_3_80_2","volume-title":"ICLR","author":"Lee Yoonho","year":"2023","unstructured":"Yoonho Lee, Huaxiu Yao, and Chelsea Finn. 2023. Diversify and disambiguate: Out-of-distribution robustness via disagreement. In ICLR."},{"key":"e_1_3_3_81_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298701"},{"key":"e_1_3_3_82_2","volume-title":"ICLR","author":"Li Yixuan","year":"2016","unstructured":"Yixuan Li, Jason Yosinski, Jeff Clune, Hod Lipson, and John E. Hopcroft. 2016. Convergent learning: Do different neural networks learn the same representations? In ICLR."},{"key":"e_1_3_3_83_2","volume-title":"ISSTA","author":"Li Yuanchun","year":"2021","unstructured":"Yuanchun Li, Ziqi Zhang, Bingyan Liu, Ziyue Yang, and Yunxin Liu. 2021. ModelDiff: Testing-based DNN similarity comparison for model reuse detection. In ISSTA."},{"key":"e_1_3_3_84_2","doi-asserted-by":"publisher","unstructured":"Baihan Lin. 2022. Geometric and topological inference for deep representations of complex networks. In WWW\u201922. Association for Computing Machinery New York NY 334\u2013338. DOI:10.1145\/3487553.3524194","DOI":"10.1145\/3487553.3524194"},{"key":"e_1_3_3_85_2","article-title":"Adaptive geo-topological independence criterion","author":"Lin Baihan","year":"2018","unstructured":"Baihan Lin and Nikolaus Kriegeskorte. 2018. Adaptive geo-topological independence criterion. arXiv:1810.02923. Retrieved from https:\/\/arxiv.org\/abs\/1810.02923","journal-title":"arXiv:1810.02923"},{"issue":"1","key":"e_1_3_3_86_2","article-title":"Divergence measures based on the Shannon entropy","volume":"37","author":"Lin Jianhua","year":"1991","unstructured":"Jianhua Lin. 1991. Divergence measures based on the Shannon entropy. IEEE Trans. Inf. Theory 37, 1 (1991).","journal-title":"IEEE Trans. Inf. Theory"},{"key":"e_1_3_3_87_2","article-title":"Model stability with continuous data updates","author":"Liu Huiting","year":"2022","unstructured":"Huiting Liu, Avinesh P. V. S., Siddharth Patwardhan, Peter Grasch, and Sachin Agarwal. 2022. Model stability with continuous data updates. arXiv:2201.05692. Retrieved from https:\/\/arxiv.org\/abs\/2201.05692","journal-title":"arXiv:2201.05692"},{"key":"e_1_3_3_88_2","volume-title":"ECCV","author":"Lu Yao","year":"2022","unstructured":"Yao Lu, Wen Yang, Yunzhe Zhang, Zuohui Chen, Jinyin Chen, Qi Xuan, Zhen Wang, and Xiaoniu Yang. 2022. Understanding the dynamics of DNNs using graph modularity. In ECCV."},{"key":"e_1_3_3_89_2","volume-title":"ICML","author":"Ma Xingjun","year":"2018","unstructured":"Xingjun Ma, Yisen Wang, Michael E. Houle, Shuo Zhou, Sarah Erfani, Shutao Xia, Sudanthi Wijewickrema, and James Bailey. 2018. Dimensionality-driven learning with noisy labels. In ICML."},{"issue":"2","key":"e_1_3_3_90_2","article-title":"Misinterpretation and misuse of the kappa statistic","volume":"126","author":"Maclure Malcom","year":"1987","unstructured":"Malcom Maclure and Walter C. Willett. 1987. Misinterpretation and misuse of the kappa statistic. Am. J. Epidemiol. 126, 2 (1987).","journal-title":"Am. J. Epidemiol."},{"key":"e_1_3_3_91_2","volume-title":"NeurIPS","author":"Madani Omid","year":"2004","unstructured":"Omid Madani, David M. Pennock, and Gary William Flake. 2004. Co-Validation: Using model disagreement on unlabeled data to validate classification algorithms. In NeurIPS."},{"key":"e_1_3_3_92_2","volume-title":"ICLR","author":"Madry Aleksander","year":"2018","unstructured":"Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu. 2018. Towards deep learning models resistant to adversarial attacks. In ICLR."},{"key":"e_1_3_3_93_2","volume-title":"CVPR","author":"Maniparambil Mayug","year":"2024","unstructured":"Mayug Maniparambil, Raiymbek Akshulakov, Yasser Abdelaziz Dahou Djilali, Mohamed El Amine Seddik, Sanath Narayan, Karttikeya Mangalam, and Noel E. O\u2019Connor. 2024. Do vision and language encoders represent the world similarly? In CVPR."},{"key":"e_1_3_3_94_2","volume-title":"ICML","author":"Marx Charles T.","year":"2020","unstructured":"Charles T. Marx, Fl\u00e1vio P. Calmon, and Berk Ustun. 2020. Predictive multiplicity in classification. In ICML."},{"key":"e_1_3_3_95_2","volume-title":"WWW","author":"Mathew Binny","year":"2020","unstructured":"Binny Mathew, Sandipan Sikdar, Florian Lemmerich, and Markus Strohmaier. 2020. The POLAR framework: Polar opposites enable interpretability of pre-trained word embeddings. In WWW."},{"key":"e_1_3_3_96_2","volume-title":"NeurIPS","author":"May Avner","year":"2019","unstructured":"Avner May, Jian Zhang, Tri Dao, and Christopher R\u00e9. 2019. On the downstream performance of compressed word embeddings. In NeurIPS."},{"key":"e_1_3_3_97_2","volume-title":"Third BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP","author":"McCoy R. Thomas","year":"2020","unstructured":"R. Thomas McCoy, Junghyun Min, and Tal Linzen. 2020. BERTs of a feather do not generalize together: Large variability in generalization across models with similar test set performance. In Third BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP."},{"key":"e_1_3_3_98_2","article-title":"Exploring the interchangeability of CNN embedding spaces","author":"McNeely-White David","year":"2020","unstructured":"David McNeely-White, Benjamin Sattelberg, Nathaniel Blanchard, and Ross Beveridge. 2020. Exploring the interchangeability of CNN embedding spaces. arXiv:2010.02323. Retrieved from https:\/\/arxiv.org\/abs\/2010.02323","journal-title":"arXiv:2010.02323"},{"key":"e_1_3_3_99_2","volume-title":"Conference on Cognitive Computational Neuroscience","author":"Mehrer Johannes","year":"2018","unstructured":"Johannes Mehrer, Nikolaus Kriegeskorte, and Tim C. Kietzmann. 2018. Beware of the beginnings: Intermediate and higher-level representations in deep neural networks are strongly affected by weight initialization. In Conference on Cognitive Computational Neuroscience."},{"key":"e_1_3_3_100_2","article-title":"Individual differences among deep neural network models","volume":"11","author":"Mehrer Johannes","year":"2020","unstructured":"Johannes Mehrer, Courtney J. Spoerer, Nikolaus Kriegeskorte, and Tim C. Kietzmann. 2020. Individual differences among deep neural network models. Nat. Commun. 11 (2020).","journal-title":"Nat. Commun."},{"key":"e_1_3_3_101_2","volume-title":"Proceedings of the Third BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP","author":"Merchant Amil","year":"2020","unstructured":"Amil Merchant, Elahe Rahimtoroghi, Ellie Pavlick, and Ian Tenney. 2020. What happens to BERT embeddings during fine-tuning? In Proceedings of the Third BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP."},{"key":"e_1_3_3_102_2","volume-title":"NeurIPS","author":"Morcos Ari S.","year":"2018","unstructured":"Ari S. Morcos, Maithra Raghu, and Samy Bengio. 2018. Insights on representational similarity in neural networks with canonical correlation. In NeurIPS."},{"key":"e_1_3_3_103_2","volume-title":"ICLR","author":"Moschella Luca","year":"2023","unstructured":"Luca Moschella, Valentino Maiorca, Marco Fumero, Antonio Norelli, Francesco Locatello, and Emanuele Rodol\u00e0. 2023. Relative representations enable zero-shot latent space communication. In ICLR."},{"key":"e_1_3_3_104_2","volume-title":"ICLR 2024 Workshop on Representational Alignment","author":"Murphy Alex Graeme","year":"2024","unstructured":"Alex Graeme Murphy, Joel Zylberberg, and Alona Fyshe. 2024. Correcting biased centered kernel alignment measures in biological and artificial neural networks. In ICLR 2024 Workshop on Representational Alignment. https:\/\/openreview.net\/forum?id=E1NRrGtIHG"},{"key":"e_1_3_3_105_2","volume-title":"ICML","author":"Nanda Vedant","year":"2022","unstructured":"Vedant Nanda, Till Speicher, Camila Kolling, John P. Dickerson, Krishna P. Gummadi, and Adrian Weller. 2022. Measuring representational robustness of neural networks through shared invariances. In ICML."},{"issue":"23","key":"e_1_3_3_106_2","article-title":"Modularity and community structure in networks","volume":"103","author":"Newman Mark E. J.","year":"2006","unstructured":"Mark E. J. Newman. 2006. Modularity and community structure in networks. Proc. Natl. Acad. Sci. U.S.A. 103, 23 (2006).","journal-title":"Proc. Natl. Acad. Sci. U.S.A."},{"key":"e_1_3_3_107_2","article-title":"Finding and evaluating community structure in networks","volume":"69","author":"Newman Mark E. J.","year":"2004","unstructured":"Mark E. J. Newman and Michelle Girvan. 2004. Finding and evaluating community structure in networks. Phys. Rev. E 69 (2004).","journal-title":"Phys. Rev. E"},{"key":"e_1_3_3_108_2","volume-title":"ICLR","author":"Nguyen Thao","year":"2021","unstructured":"Thao Nguyen, Maithra Raghu, and Simon Kornblith. 2021. Do wide and deep networks learn the same things? Uncovering how neural network representations vary with width and depth. In ICLR."},{"key":"e_1_3_3_109_2","article-title":"On the origins of the block structure phenomenon in neural network representations","author":"Nguyen Thao","year":"2022","unstructured":"Thao Nguyen, Maithra Raghu, and Simon Kornblith. 2022. On the origins of the block structure phenomenon in neural network representations. Trans. Mach. Learn. Res. (2022).","journal-title":"Trans. Mach. Learn. Res."},{"key":"e_1_3_3_110_2","volume-title":"NeurIPS","author":"Ostrow Mitchell","year":"2023","unstructured":"Mitchell Ostrow, Adam Joseph Eisen, Leo Kozachkov, and Ila R. Fiete. 2023. Beyond geometry: Comparing the temporal structure of computation in neural circuits with dynamical similarity analysis. In NeurIPS."},{"key":"e_1_3_3_111_2","volume-title":"ICLR","author":"Pagliardini Matteo","year":"2023","unstructured":"Matteo Pagliardini, Martin Jaggi, Fran\u00e7ois Fleuret, and Sai Praneeth Karimireddy. 2023. Agree to disagree: Diversity through disagreement for better transferability. In ICLR."},{"key":"e_1_3_3_112_2","volume-title":"ICLR","author":"Park Namuk","year":"2022","unstructured":"Namuk Park and Songkuk Kim. 2022. How do vision transformers work? In ICLR."},{"key":"e_1_3_3_113_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00409"},{"issue":"1","key":"e_1_3_3_114_2","article-title":"Some new test criteria in multivariate analysis","volume":"26","author":"Pillai K. C. Sreedharan","year":"1955","unstructured":"K. C. Sreedharan Pillai. 1955. Some new test criteria in multivariate analysis. Ann. Math. Stat. 26, 1 (1955).","journal-title":"Ann. Math. Stat."},{"key":"e_1_3_3_115_2","volume-title":"NeurIPS","author":"Raghu Maithra","year":"2017","unstructured":"Maithra Raghu, Justin Gilmer, Jason Yosinski, and Jascha Sohl-Dickstein. 2017. SVCCA: Singular vector canonical correlation analysis for deep learning dynamics and interpretability. In NeurIPS."},{"key":"e_1_3_3_116_2","first-page":"12116","article-title":"Do vision transformers see like convolutional neural networks?","volume":"34","author":"Raghu Maithra","year":"2021","unstructured":"Maithra Raghu, Thomas Unterthiner, Simon Kornblith, Chiyuan Zhang, and Alexey Dosovitskiy. 2021. Do vision transformers see like convolutional neural networks? NeurIPS 34 (2021), 12116\u201312128.","journal-title":"NeurIPS"},{"key":"e_1_3_3_117_2","volume-title":"NAACL","author":"Rahamim Adir","year":"2024","unstructured":"Adir Rahamim and Yonatan Belinkov. 2024. ContraSim\u2014Analyzing neural representations based on contrastive learning. In NAACL."},{"key":"e_1_3_3_118_2","article-title":"Matrix correlation","volume":"49","author":"Ramsay James O.","year":"1984","unstructured":"James O. Ramsay, Jos ten Berge, and George P. H. Styan. 1984. Matrix correlation. Psychometrika 49 (1984).","journal-title":"Psychometrika"},{"issue":"2","key":"e_1_3_3_119_2","doi-asserted-by":"crossref","first-page":"182","DOI":"10.1038\/s41592-023-02150-0","article-title":"Understanding metric-related pitfalls in image analysis validation","volume":"21","author":"Reinke Annika","year":"2024","unstructured":"Annika Reinke, Minu D. Tizabi, Michael Baumgartner, Matthias Eisenmann, Doreen Heckmann-N\u00f6tzel, A. Emre Kavur, Tim R\u00e4dsch, Carole H. Sudre, Laura Acion, Michela Antonelli, et\u00a0al. 2024. Understanding metric-related pitfalls in image analysis validation. Nature Methods 21, 2 (2024), 182\u2013194.","journal-title":"Nature Methods"},{"issue":"3","key":"e_1_3_3_120_2","article-title":"A unifying tool for linear multivariate statistical methods: The RV-coefficient","volume":"25","author":"Robert Paul","year":"1976","unstructured":"Paul Robert and Yves Escoufier. 1976. A unifying tool for linear multivariate statistical methods: The RV-coefficient. J. Roy. Stat. Soc. Ser. C (Appl. Stat.) 25, 3 (1976).","journal-title":"J. Roy. Stat. Soc. Ser. C (Appl. Stat.)"},{"key":"e_1_3_3_121_2","doi-asserted-by":"publisher","DOI":"10.1109\/SaTML54575.2023.00039"},{"key":"e_1_3_3_122_2","volume-title":"BMVC\u201922","author":"Saha Aninda","year":"2022","unstructured":"Aninda Saha, Alina N. Bialkowski, and Sara Khalifa. 2022. Distilling representational similarity using centered kernel alignment (CKA). In BMVC\u201922. BMVA Press."},{"key":"e_1_3_3_123_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-86523-8_42"},{"key":"e_1_3_3_124_2","volume-title":"Machine Learning and Principles and Practice of Knowledge Discovery in Databases","author":"Schumacher Tobias","year":"2021","unstructured":"Tobias Schumacher, Hinrikus Wolf, Martin Ritzert, Florian Lemmerich, Martin Grohe, and Markus Strohmaier. 2021. The effects of randomness on the stability of node embeddings. In Machine Learning and Principles and Practice of Knowledge Discovery in Databases."},{"key":"e_1_3_3_125_2","article-title":"A generalized solution of the orthogonal procrustes problem","volume":"31","author":"Sch\u00f6nemann Peter H.","year":"1966","unstructured":"Peter H. Sch\u00f6nemann. 1966. A generalized solution of the orthogonal procrustes problem. Psychometrika 31 (1966).","journal-title":"Psychometrika"},{"key":"e_1_3_3_126_2","article-title":"Using distance on the Riemannian manifold to compare representations in brain and in models","volume":"239","author":"Shahbazi Mahdiyar","year":"2021","unstructured":"Mahdiyar Shahbazi, Ali Shirali, Hamid Aghajan, and Hamed Nili. 2021. Using distance on the Riemannian manifold to compare representations in brain and in models. NeuroImage 239 (2021).","journal-title":"NeuroImage"},{"key":"e_1_3_3_127_2","article-title":"Anti-Distillation: Improving reproducibility of deep networks","author":"Shamir Gil I.","year":"2020","unstructured":"Gil I. Shamir and Lorenzo Coviello. 2020. Anti-Distillation: Improving reproducibility of deep networks. arXiv:2010.09923. Retrieved from https:\/\/arxiv.org\/abs\/2010.09923","journal-title":"arXiv:2010.09923"},{"key":"e_1_3_3_128_2","article-title":"The analysis of proximities: Multidimensional scaling with an unknown distance function. I.","volume":"27","author":"Shepard Roger N.","year":"1962","unstructured":"Roger N. Shepard. 1962. The analysis of proximities: Multidimensional scaling with an unknown distance function. I. Psychometrika 27 (1962).","journal-title":"Psychometrika"},{"key":"e_1_3_3_129_2","volume-title":"NeurIPS","author":"Shi Jianghong","year":"2019","unstructured":"Jianghong Shi, Eric Shea-Brown, and Michael Buice. 2019. Comparison against task driven artificial neural networks reveals functional properties in mouse visual cortex. In NeurIPS."},{"issue":"3","key":"e_1_3_3_130_2","article-title":"The kappa statistic in reliability studies: Use, interpretation, and sample size requirements","volume":"85","author":"Sim Julius","year":"2005","unstructured":"Julius Sim and Chris C. Wright. 2005. The kappa statistic in reliability studies: Use, interpretation, and sample size requirements. Phys. Therapy 85, 3 (2005).","journal-title":"Phys. Therapy"},{"key":"e_1_3_3_131_2","volume-title":"Workshop at ICLR","author":"Simonyan Karen","year":"2013","unstructured":"Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman. 2013. Deep inside convolutional networks: Visualising image classification models and saliency maps. In Workshop at ICLR."},{"key":"e_1_3_3_132_2","volume-title":"AAAI Integrating Multiple Learned Models Workshop","author":"Skalak David B.","year":"1996","unstructured":"David B. Skalak et\u00a0al. 1996. The sources of increased accuracy for two proposed boosting algorithms. In AAAI Integrating Multiple Learned Models Workshop."},{"issue":"47","key":"e_1_3_3_133_2","first-page":"1393","article-title":"Feature selection via dependence maximization","volume":"13","author":"Song Le","year":"2012","unstructured":"Le Song, Alex Smola, Arthur Gretton, Justin Bedo, and Karsten Borgwardt. 2012. Feature selection via dependence maximization. J. Mach. Learn. Res. 13, 47 (2012), 1393\u20131434. http:\/\/jmlr.org\/papers\/v13\/song12a.html","journal-title":"J. Mach. Learn. Res."},{"key":"e_1_3_3_134_2","article-title":"Undivided attention: Are intermediate layers necessary for BERT?","author":"Sridhar Sharath Nittur","year":"2020","unstructured":"Sharath Nittur Sridhar and Anthony Sarah. 2020. Undivided attention: Are intermediate layers necessary for BERT? arXiv:2012.11881. Retrieved from https:\/\/arxiv.org\/abs\/2012.11881","journal-title":"arXiv:2012.11881"},{"key":"e_1_3_3_135_2","volume-title":"NeurIPS","author":"Stanton Samuel","year":"2021","unstructured":"Samuel Stanton, Pavel Izmailov, Polina Kirichenko, Alexander A. Alemi, and Andrew Gordon Wilson. 2021. Does knowledge distillation really work? In NeurIPS."},{"key":"e_1_3_3_136_2","article-title":"A comparison of consensus, consistency, and measurement approaches to estimating interrater reliability","volume":"9","author":"Stemler Steven E.","year":"2019","unstructured":"Steven E. Stemler. 2019. A comparison of consensus, consistency, and measurement approaches to estimating interrater reliability. Pract. Assess. Res. Eval. 9 (2019).","journal-title":"Pract. Assess. Res. Eval."},{"key":"e_1_3_3_137_2","article-title":"Getting aligned on representational alignment","author":"Sucholutsky Ilia","year":"2023","unstructured":"Ilia Sucholutsky, Lukas Muttenthaler, Adrian Weller, Andi Peng, Andreea Bobu, Been Kim, Bradley C. Love, Erin Grant, Jascha Achterberg, Joshua B Tenenbaum, et\u00a0al. 2023. Getting aligned on representational alignment. arXiv:2310.13018. Retrieved from https:\/\/arxiv.org\/abs\/2310.13018","journal-title":"arXiv:2310.13018"},{"key":"e_1_3_3_138_2","volume-title":"ICML","author":"Summers Cecilia","year":"2021","unstructured":"Cecilia Summers and Michael J. Dinneen. 2021. Nondeterminism and instability in neural network optimization. In ICML."},{"key":"e_1_3_3_139_2","article-title":"Deep intellectual property: A survey","author":"Sun Yuchen","year":"2023","unstructured":"Yuchen Sun, Tianpeng Liu, Panhe Hu, Qing Liao, Shouling Ji, Nenghai Yu, Deke Guo, and Li Liu. 2023. Deep intellectual property: A survey. arXiv:2304.14613.","journal-title":"arXiv:2304.14613"},{"issue":"6","key":"e_1_3_3_140_2","article-title":"Measuring and testing dependence by correlation of distances","volume":"35","author":"Sz\u00e9kely G\u00e1bor J.","year":"2007","unstructured":"G\u00e1bor J. Sz\u00e9kely, Maria L. Rizzo, and Nail K. Bakirov. 2007. Measuring and testing dependence by correlation of distances. Ann. Stat. 35, 6 (2007).","journal-title":"Ann. Stat."},{"key":"e_1_3_3_141_2","article-title":"An analysis of diversity measures","volume":"65","author":"Tang Ke","year":"2006","unstructured":"Ke Tang, Ponnuthurai N. Suganthan, and Xin Yao. 2006. An analysis of diversity measures. Mach. Learn. 65 (2006).","journal-title":"Mach. Learn."},{"key":"e_1_3_3_142_2","article-title":"Similarity of neural networks with gradients","author":"Tang Shuai","year":"2020","unstructured":"Shuai Tang, Wesley J. Maddox, Charlie Dickens, Tom Diethe, and Andreas Damianou. 2020. Similarity of neural networks with gradients. arXiv:2003.11498. Retrieved from https:\/\/arxiv.org\/abs\/2003.11498","journal-title":"arXiv:2003.11498"},{"key":"e_1_3_3_143_2","first-page":"4527","volume-title":"EMNLP\u201921","author":"Timkey William","year":"2021","unstructured":"William Timkey and Marten van Schijndel. 2021. All bark and no bite: Rogue dimensions in transformer language models obscure representational quality. In EMNLP\u201921. 4527\u20134546."},{"issue":"4","key":"e_1_3_3_144_2","article-title":"Interrater reliability and agreement of subjective judgments.","volume":"22","author":"Tinsley Howard E.","year":"1975","unstructured":"Howard E. Tinsley and David J. Weiss. 1975. Interrater reliability and agreement of subjective judgments. J. Counsel. Psychol. 22, 4 (1975).","journal-title":"J. Counsel. Psychol."},{"key":"e_1_3_3_145_2","volume-title":"ICLR","author":"Tsitsulin Anton","year":"2020","unstructured":"Anton Tsitsulin, Marina Munkhoeva, Davide Mottin, Panagiotis Karras, Alex Bronstein, Ivan Oseledets, and Emmanuel Mueller. 2020. The shape of data: Intrinsic distance for data distributions. In ICLR."},{"issue":"4","key":"e_1_3_3_146_2","article-title":"Fast estimation of $tr(f(A))$ via stochastic Lanczos quadrature","volume":"38","author":"Ubaru Shashanka","year":"2017","unstructured":"Shashanka Ubaru, Jie Chen, and Yousef Saad. 2017. Fast estimation of $tr(f(A))$ via stochastic Lanczos quadrature. SIAM J. Matrix Anal. Appl. 38, 4 (2017).","journal-title":"SIAM J. Matrix Anal. Appl."},{"issue":"6","key":"e_1_3_3_147_2","article-title":"A tutorial on canonical correlation methods","volume":"50","author":"Uurtio Viivi","year":"2017","unstructured":"Viivi Uurtio, Jo\u00e3o M. Monteiro, Jaz Kandola, John Shawe-Taylor, Delmiro Fernandez-Reyes, and Juho Rousu. 2017. A tutorial on canonical correlation methods. Comput. Surv. 50, 6 (2017).","journal-title":"Comput. Surv."},{"issue":"2","key":"e_1_3_3_148_2","article-title":"Canonical ridge and econometrics of joint production","volume":"4","author":"Vinod Hrishikesh D.","year":"1976","unstructured":"Hrishikesh D. Vinod. 1976. Canonical ridge and econometrics of joint production. J. Econometr. 4, 2 (1976).","journal-title":"J. Econometr."},{"key":"e_1_3_3_149_2","article-title":"Exploring new ways: Enforcing representational dissimilarity to learn new features and reduce error consistency","author":"Wald Tassilo","year":"2023","unstructured":"Tassilo Wald, Constantin Ulrich, Fabian Isensee, David Zimmerer, Gregor Koehler, Michael Baumgartner, and Klaus H Maier-Hein. 2023. Exploring new ways: Enforcing representational dissimilarity to learn new features and reduce error consistency. ICML Workshop SCIS (2023).","journal-title":"ICML Workshop SCIS"},{"issue":"2","key":"e_1_3_3_150_2","article-title":"Towards understanding the instability of network embedding","volume":"34","author":"Wang Chenxu","year":"2020","unstructured":"Chenxu Wang, Wei Rao, Wenna Guo, Pinghui Wang, Jun Liu, and Xiaohong Guan. 2020. Towards understanding the instability of network embedding. IEEE Trans. Knowl. Data Eng. 34, 2 (2020).","journal-title":"IEEE Trans. Knowl. Data Eng."},{"key":"e_1_3_3_151_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00252"},{"key":"e_1_3_3_152_2","article-title":"Understanding weight similarity of neural networks via chain normalization rule and hypothesis-training-testing","author":"Wang Guangcong","year":"2022","unstructured":"Guangcong Wang, Guangrun Wang, Wenqi Liang, and Jianhuang Lai. 2022. Understanding weight similarity of neural networks via chain normalization rule and hypothesis-training-testing. arXiv:2208.04369. Retrieved from https:\/\/arxiv.org\/abs\/2208.04369","journal-title":"arXiv:2208.04369"},{"key":"e_1_3_3_153_2","volume-title":"NeurIPS","author":"Wang Liwei","year":"2018","unstructured":"Liwei Wang, Lunjia Hu, Jiayuan Gu, Zhiqiang Hu, Yue Wu, Kun He, and John E. Hopcroft. 2018. Towards understanding learning representations: To what extent do different neural networks learn the same representation. In NeurIPS."},{"key":"e_1_3_3_154_2","volume-title":"ICML","author":"Wang Tongzhou","year":"2020","unstructured":"Tongzhou Wang and Phillip Isola. 2020. Understanding contrastive representation learning through alignment and uniformity on the hypersphere. In ICML."},{"issue":"3","key":"e_1_3_3_155_2","article-title":"Certain generalizations in the analysis of variance","volume":"24","author":"Wilks Samuel S.","year":"1932","unstructured":"Samuel S. Wilks. 1932. Certain generalizations in the analysis of variance. Biometrika 24, 3\/4 (1932).","journal-title":"Biometrika"},{"key":"e_1_3_3_156_2","volume-title":"NeurIPS","author":"Williams Alex H.","year":"2021","unstructured":"Alex H. Williams, Erin Kunz, Simon Kornblith, and Scott W. Linderman. 2021. Generalized shape metrics on neural representations. In NeurIPS."},{"key":"e_1_3_3_157_2","volume-title":"ACL","author":"Wu John","year":"2020","unstructured":"John Wu, Yonatan Belinkov, Hassan Sajjad, Nadir Durrani, Fahim Dalvi, and James Glass. 2020. Similarity analysis of contextual word representation models. In ACL."},{"issue":"1","key":"e_1_3_3_158_2","doi-asserted-by":"crossref","first-page":"2065","DOI":"10.1038\/s41467-021-22244-7","article-title":"Limits to visual representational correspondence between convolutional neural networks and the human brain","volume":"12","author":"Xu Yaoda","year":"2021","unstructured":"Yaoda Xu and Maryam Vaziri-Pashkam. 2021. Limits to visual representational correspondence between convolutional neural networks and the human brain. Nat. Commun. 12, 1 (2021), 2065.","journal-title":"Nat. Commun."},{"key":"e_1_3_3_159_2","volume-title":"ICCL","author":"Yadav Vikas","year":"2018","unstructured":"Vikas Yadav and Steven Bethard. 2018. A survey on recent advances in named entity recognition from deep learning models. In ICCL."},{"key":"e_1_3_3_160_2","article-title":"Unification of various techniques of multivariate analysis by means of generalized coefficient of determination","volume":"1","author":"Yanai Haruo","year":"1974","unstructured":"Haruo Yanai. 1974. Unification of various techniques of multivariate analysis by means of generalized coefficient of determination. Kodo Keiryogaku (Jpn. J. Behaviormetr.) 1 (1974).","journal-title":"Kodo Keiryogaku (Jpn. J. Behaviormetr.)"},{"issue":"6","key":"e_1_3_3_161_2","article-title":"A survey on canonical correlation analysis","volume":"33","author":"Yang Xinghao","year":"2021","unstructured":"Xinghao Yang, Weifeng Liu, Wei Liu, and Dacheng Tao. 2021. A survey on canonical correlation analysis. IEEE Trans. Knowl. Data Eng. 33, 6 (2021).","journal-title":"IEEE Trans. Knowl. Data Eng."},{"key":"e_1_3_3_162_2","volume-title":"NeurIPS","author":"Yang Zhuolin","year":"2021","unstructured":"Zhuolin Yang, Linyi Li, Xiaojun Xu, Shiliang Zuo, Qian Chen, Pan Zhou, Benjamin Rubinstein, Ce Zhang, and Bo Li. 2021. TRS: Transferability reduced ensemble via promoting gradient diversity and model smoothness. In NeurIPS."},{"key":"e_1_3_3_163_2","volume-title":"NeurIPS","author":"Yin Zi","year":"2018","unstructured":"Zi Yin and Yuanyuan Shen. 2018. On the dimensionality of word embedding. In NeurIPS."},{"key":"e_1_3_3_164_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE48307.2020.00014"},{"key":"e_1_3_3_165_2","doi-asserted-by":"publisher","unstructured":"Zikai Zhou Yunhang Shen Shitong Shao Linrui Gong and Shaohui Lin. 2024. Rethinking centered kernel alignment in knowledge distillation. In Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence (IJCAI\u201924) Jeju Korea. DOI:10.24963\/ijcai.2024\/628","DOI":"10.24963\/ijcai.2024\/628"},{"key":"e_1_3_3_166_2","volume-title":"ICLR","author":"Zong Martin","year":"2023","unstructured":"Martin Zong, Zengyu Qiu, Xinzhu Ma, Kunlin Yang, Chunya Liu, Jun Hou, Shuai Yi, and Wanli Ouyang. 2023. Better teacher better student: Dynamic prior knowledge for knowledge distillation. In ICLR."},{"key":"e_1_3_3_167_2","article-title":"SmolLM2: When smol goes big\u2013data-centric training of a small language model","author":"Allal Loubna Ben","year":"2025","unstructured":"Loubna Ben Allal, Anton Lozhkov, Elie Bakouch, Gabriel Mart\u00edn Bl\u00e1zquez, Guilherme Penedo, Lewis Tunstall, Andr\u00e9s Marafioti, Hynek Kydl\u00ed\u010dek, Agust\u00edn Piqueres Lajar\u00edn, Vaibhav Srivastav, et\u00a0al. 2025. SmolLM2: When smol goes big\u2013data-centric training of a small language model. arXiv:2502.02737. Retrieved from https:\/\/arxiv.org\/abs\/2502.02737","journal-title":"arXiv:2502.02737"},{"issue":"4","key":"e_1_3_3_168_2","article-title":"Detecting sequential patterns and determining their reliability with fallible observers","volume":"2","author":"Bakeman Roger","year":"1997","unstructured":"Roger Bakeman, Vicenq Quera, Duncan McArthur, and Byron F. Robinson. 1997. Detecting sequential patterns and determining their reliability with fallible observers. Psychol. Methods 2, 4 (1997).","journal-title":"Psychol. Methods"},{"key":"e_1_3_3_169_2","volume-title":"NeurIPS","author":"Barannikov Serguei","year":"2021","unstructured":"Serguei Barannikov, Ilya Trofimov, Grigorii Sotnikov, Ekaterina Trimbach, Alexander Korotin, Alexander Filippov, and Evgeny Burnaev. 2021. Manifold topology divergence: A framework for comparing data manifolds. In NeurIPS."},{"key":"e_1_3_3_170_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.354"},{"issue":"1","key":"e_1_3_3_171_2","article-title":"Probing classifiers: Promises, shortcomings, and advances","volume":"48","author":"Belinkov Yonatan","year":"2022","unstructured":"Yonatan Belinkov. 2022. Probing classifiers: Promises, shortcomings, and advances. Comput. Ling. 48, 1 (2022).","journal-title":"Comput. Ling."},{"key":"e_1_3_3_172_2","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/W14-3302"},{"key":"e_1_3_3_173_2","volume-title":"ICLR\u201924 Workshop on Representational Alignment","author":"Brown Davis","year":"2024","unstructured":"Davis Brown, Madelyn Ruth Shapiro, Alyson Bittner, Jackson Warley, and Henry Kvinge. 2024. Wild comparisons: A study of how representation similarity changes when input data is drawn from a shifted distribution. In ICLR\u201924 Workshop on Representational Alignment."},{"issue":"5","key":"e_1_3_3_174_2","article-title":"Bias, prevalence and kappa","volume":"46","author":"Byrt Ted","year":"1993","unstructured":"Ted Byrt, Janet Bishop, and John B. Carlin. 1993. Bias, prevalence and kappa. J. Clin. Epidemiol. 46, 5 (1993).","journal-title":"J. Clin. Epidemiol."},{"key":"e_1_3_3_175_2","volume-title":"NeurIPS","author":"Chen Wei","year":"2023","unstructured":"Wei Chen, Zichen Miao, and Qiang Qiu. 2023. Inner product-based neural network similarity. In NeurIPS."},{"key":"e_1_3_3_176_2","doi-asserted-by":"publisher","unstructured":"Alexis Conneau Kartikay Khandelwal Naman Goyal Vishrav Chaudhary Guillaume Wenzek Francisco Guzm\u00e1n Edouard Grave Myle Ott Luke Zettlemoyer and Veselin Stoyanov. 2020. Unsupervised cross-lingual representation learning at scale. Online. DOI:10.18653\/v1\/2020.acl-main.747","DOI":"10.18653\/v1\/2020.acl-main.747"},{"key":"e_1_3_3_177_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D18-1269"},{"key":"e_1_3_3_178_2","volume-title":"NeurIPS","author":"Cui Tianyu","year":"2022","unstructured":"Tianyu Cui, Yogesh Kumar, Pekka Marttinen, and Samuel Kaski. 2022. Deconfounded representation similarity for comparison of neural networks. In NeurIPS."},{"key":"e_1_3_3_179_2","volume-title":"ICLR\u201922 Workshop on Geometrical and Topological Representation Learning","author":"Davari MohammadReza","year":"2022","unstructured":"MohammadReza Davari, Stefan Horoi, Amine Natik, Guillaume Lajoie, Guy Wolf, and Eugene Belilovsky. 2022. On the inadequacy of CKA as a measure of similarity in deep learning. In ICLR\u201922 Workshop on Geometrical and Topological Representation Learning."},{"key":"e_1_3_3_180_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"e_1_3_3_181_2","volume-title":"ICLR","author":"Dosovitskiy Alexey","year":"2021","unstructured":"Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2021. An image is worth 16x16 words: Transformers for image recognition at scale. In ICLR."},{"key":"e_1_3_3_182_2","unstructured":"Marin Dujmovi\u0107 Jeffrey S. Bowers Federico Adolfi and Gaurav Malhotra. 2022. The pitfalls of measuring representational similarity using representational similarity analysis. bioRxiv (2022). https:\/\/www.biorxiv.org\/content\/early\/2022\/04\/07\/2022.04.05.487135"},{"key":"e_1_3_3_183_2","doi-asserted-by":"publisher","DOI":"10.1561\/0600000105"},{"key":"e_1_3_3_184_2","article-title":"Generalized procrustes analysis","volume":"40","author":"Gower John C.","year":"1975","unstructured":"John C. Gower. 1975. Generalized procrustes analysis. Psychometrika 40 (1975).","journal-title":"Psychometrika"},{"key":"e_1_3_3_185_2","volume-title":"ICLR\u201924 Workshop on Representational Alignment","author":"Guth Florentin","year":"2024","unstructured":"Florentin Guth and Brice M\u00e9nard. 2024. On the universality of neural encodings in CNNs. In ICLR\u201924 Workshop on Representational Alignment."},{"key":"e_1_3_3_186_2","volume-title":"NeurIPS","author":"Hamilton Will","year":"2017","unstructured":"Will Hamilton, Zhitao Ying, and Jure Leskovec. 2017. Inductive representation learning on large graphs. In NeurIPS."},{"key":"e_1_3_3_187_2","doi-asserted-by":"publisher","DOI":"10.18112\/openneuro.ds000105.v3.0.0"},{"key":"e_1_3_3_188_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_3_189_2","article-title":"Mobilenets: Efficient convolu-tional neural networks for mobile vision applications","author":"Howard A. G.","year":"2017","unstructured":"A. G. Howard. 2017. Mobilenets: Efficient convolu-tional neural networks for mobile vision applications. arXiv:1704.04861. Retrieved from https:\/\/arxiv.org\/abs\/1704.04861","journal-title":"arXiv:1704.04861"},{"key":"e_1_3_3_190_2","volume-title":"NeurIPS","author":"Hu Weihua","year":"2020","unstructured":"Weihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong, Hongyu Ren, Bowen Liu, Michele Catasta, and Jure Leskovec. 2020. Open graph benchmark: Datasets for machine learning on graphs. In NeurIPS, Vol. 33."},{"key":"e_1_3_3_191_2","volume-title":"ICML","author":"Ilyas Andrew","year":"2022","unstructured":"Andrew Ilyas, Sung Min Park, Logan Engstrom, Guillaume Leclerc, and Aleksander Madry. 2022. Datamodels: Predicting predictions from training data. In ICML."},{"issue":"1","key":"e_1_3_3_192_2","article-title":"A survey on contrastive self-supervised learning","volume":"9","author":"Jaiswal Ashish","year":"2020","unstructured":"Ashish Jaiswal, Ashwin Ramesh Babu, Mohammad Zaki Zadeh, Debapriya Banerjee, and Fillia Makedon. 2020. A survey on contrastive self-supervised learning. Technologies 9, 1 (2020).","journal-title":"Technologies"},{"key":"e_1_3_3_193_2","doi-asserted-by":"publisher","DOI":"10.3390\/sym11091066"},{"key":"e_1_3_3_194_2","first-page":"5583","volume-title":"ICML","author":"Kim Wonjae","year":"2021","unstructured":"Wonjae Kim, Bokyung Son, and Ildoo Kim. 2021. Vilt: Vision-and-language transformer without convolution or region supervision. In ICML. PMLR, 5583\u20135594."},{"key":"e_1_3_3_195_2","volume-title":"ICML","author":"Kim Wonjae","year":"2021","unstructured":"Wonjae Kim, Bokyung Son, and Ildoo Kim. 2021. ViLT: Vision-and-language transformer without convolution or region supervision. In ICML."},{"key":"e_1_3_3_196_2","volume-title":"ICLR","author":"Kipf Thomas N.","year":"2022","unstructured":"Thomas N. Kipf and Max Welling. 2022. Semi-supervised classification with graph convolutional networks. In ICLR."},{"key":"e_1_3_3_197_2","volume-title":"Artificial Neural Networks and Machine Learning\u2014ICANN\u201919: Workshop and Special Sessions","author":"Klocek Sylwester","year":"2019","unstructured":"Sylwester Klocek, \u0141ukasz Maziarka, Maciej Wo\u0142czyk, Jacek Tabor, Jakub Nowak, and Marek \u015amieja. 2019. Hypernetwork functional image representation. In Artificial Neural Networks and Machine Learning\u2014ICANN\u201919: Workshop and Special Sessions."},{"key":"e_1_3_3_198_2","unstructured":"Alex Krizhevsky Geoffrey Hinton and others. 2009. Learning multiple layers of features from tiny images. (2009)."},{"key":"e_1_3_3_199_2","volume-title":"NeurIPS","author":"Krizhevsky Alex","year":"2012","unstructured":"Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton. 2012. ImageNet classification with deep convolutional neural networks. In NeurIPS."},{"issue":"4","key":"e_1_3_3_200_2","article-title":"Metric learning: A survey","volume":"5","author":"Kulis Brian","year":"2013","unstructured":"Brian Kulis et\u00a0al. 2013. Metric learning: A survey. Found. Trends Mach. Learn. 5, 4 (2013).","journal-title":"Found. Trends Mach. Learn."},{"key":"e_1_3_3_201_2","volume-title":"NeurIPS","author":"Kynk\u00e4\u00e4nniemi Tuomas","year":"2019","unstructured":"Tuomas Kynk\u00e4\u00e4nniemi, Tero Karras, Samuli Laine, Jaakko Lehtinen, and Timo Aila. 2019. Improved precision and recall metric for assessing generative models. In NeurIPS."},{"key":"e_1_3_3_202_2","volume-title":"ICLR","author":"Lan Zhenzhong","year":"2020","unstructured":"Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2020. ALBERT: A lite BERT for self-supervised learning of language representations. In ICLR."},{"key":"e_1_3_3_203_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01167"},{"key":"e_1_3_3_204_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7299155"},{"issue":"3","key":"e_1_3_3_205_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3617833","article-title":"Multimodality representation learning: A survey on evolution, pretraining and its applications","volume":"20","author":"Manzoor Muhammad Arslan","year":"2023","unstructured":"Muhammad Arslan Manzoor, Sarah Albarri, Ziting Xian, Zaiqiao Meng, Preslav Nakov, and Shangsong Liang. 2023. Multimodality representation learning: A survey on evolution, pretraining and its applications. ACM Trans. Multimedia Comput. Commun. Appl. 20, 3 (2023), 1\u201334.","journal-title":"ACM Trans. Multimedia Comput. Commun. Appl."},{"key":"e_1_3_3_206_2","doi-asserted-by":"publisher","DOI":"10.5555\/972470.972475"},{"key":"e_1_3_3_207_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P19-1334"},{"key":"e_1_3_3_208_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.01395"},{"key":"e_1_3_3_209_2","volume-title":"ICLR","author":"Merity Stephen","year":"2017","unstructured":"Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. 2017. Pointer sentinel mixture models. In ICLR."},{"issue":"5","key":"e_1_3_3_210_2","article-title":"How to generate random matrices from the classical compact groups","volume":"54","author":"Mezzadri Francesco","year":"2007","unstructured":"Francesco Mezzadri. 2007. How to generate random matrices from the classical compact groups. Not. Am. Math. Soc. 54, 5 (2007).","journal-title":"Not. Am. Math. Soc."},{"key":"e_1_3_3_211_2","first-page":"4","volume-title":"NIPS Workshop on Deep Learning and Unsupervised Feature Learning","author":"Netzer Yuval","year":"2011","unstructured":"Yuval Netzer, Tao Wang, Adam Coates, Alessandro Bissacco, Baolin Wu, Andrew Y Ng, et\u00a0al. 2011. Reading digits in natural images with unsupervised feature learning. In NIPS Workshop on Deep Learning and Unsupervised Feature Learning, Vol. 2011. Granada, 4."},{"issue":"8","key":"e_1_3_3_212_2","first-page":"9","article-title":"Language models are unsupervised multitask learners","volume":"1","author":"Radford Alec","year":"2019","unstructured":"Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et\u00a0al. 2019. Language models are unsupervised multitask learners. OpenAI Blog 1, 8 (2019), 9.","journal-title":"OpenAI Blog"},{"key":"e_1_3_3_213_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.01044"},{"key":"e_1_3_3_214_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D19-1410"},{"key":"e_1_3_3_215_2","article-title":"What makes two models think alike?","author":"Salle Jeanne","year":"2024","unstructured":"Jeanne Salle, Louis Jalouzot, Nur Lan, Emmanuel Chemla, and Yair Lakretz. 2024. What makes two models think alike? arXiv:2406.12620. Retrieved from https:\/\/arxiv.org\/abs\/2406.12620","journal-title":"arXiv:2406.12620"},{"key":"e_1_3_3_216_2","volume-title":"ICML","author":"Shah Harshay","year":"2023","unstructured":"Harshay Shah, Sung Min Park, Andrew Ilyas, and Aleksander Madry. 2023. ModelDiff: A framework for comparing learning algorithms. In ICML."},{"key":"e_1_3_3_217_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P18-1238"},{"key":"e_1_3_3_218_2","volume-title":"ICLR","author":"Simonyan K.","year":"2015","unstructured":"K. Simonyan and A. Zisserman. 2015. Very deep convolutional networks for large-scale image recognition. In ICLR."},{"key":"e_1_3_3_219_2","volume-title":"NeurIPS","author":"Sitzmann Vincent","year":"2020","unstructured":"Vincent Sitzmann, Julien Martel, Alexander Bergman, David Lindell, and Gordon Wetzstein. 2020. Implicit neural representations with periodic activation functions. In NeurIPS."},{"key":"e_1_3_3_220_2","first-page":"1631","volume-title":"EMNLP\u201913","author":"Socher Richard","year":"2013","unstructured":"Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts. 2013. Recursive deep models for semantic compositionality over a sentiment treebank. In EMNLP\u201913, David Yarowsky, Timothy Baldwin, Anna Korhonen, Karen Livescu, and Steven Bethard (Eds.). Association for Computational Linguistics, 1631\u20131642."},{"key":"e_1_3_3_221_2","volume-title":"CVPR","author":"Somepalli Gowthami","year":"2022","unstructured":"Gowthami Somepalli, Liam Fowl, Arpit Bansal, Ping Yeh-Chiang, Yehuda Dar, Richard Baraniuk, Micah Goldblum, and Tom Goldstein. 2022. Can neural nets learn the same model twice? Investigating reproducibility and double descent from the decision boundary perspective. In CVPR."},{"key":"e_1_3_3_222_2","unstructured":"J. T. Springenberg A. Dosovitskiy T. Brox and M. Riedmiller. 2015. Striving for simplicity: The all convolutional net. In ICLR (Workshop Track). Retrieved from http:\/\/lmbweb.informatik.uni-freiburg.de\/Publications\/2015\/DB15a"},{"key":"e_1_3_3_223_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298594"},{"key":"e_1_3_3_224_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00293"},{"key":"e_1_3_3_225_2","volume-title":"ICML","author":"Tan Mingxing","year":"2019","unstructured":"Mingxing Tan and Quoc Le. 2019. EfficientNet: Rethinking model scaling for convolutional neural networks. In ICML."},{"key":"e_1_3_3_226_2","article-title":"A survey on self-supervised representation learning","author":"Uelwer Tobias","year":"2023","unstructured":"Tobias Uelwer, Jan Robine, Stefan Sylvius Wagner, Marc H\u00f6ftmann, Eric Upschulte, Sebastian Konietzny, Maike Behrendt, and Stefan Harmeling. 2023. A survey on self-supervised representation learning. arXiv:2308.11455. Retrieved from https:\/\/arxiv.org\/abs\/2308.11455","journal-title":"arXiv:2308.11455"},{"key":"e_1_3_3_227_2","volume-title":"NeurIPS","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In NeurIPS, Vol. 30."},{"key":"e_1_3_3_228_2","volume-title":"ICLR","author":"Veli\u010dkovi\u0107 Petar","year":"2018","unstructured":"Petar Veli\u010dkovi\u0107, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Li\u00f2, and Yoshua Bengio. 2018. Graph attention networks. In ICLR."},{"key":"e_1_3_3_229_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/W18-5446"},{"key":"e_1_3_3_230_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.00638"},{"key":"e_1_3_3_231_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2024.3367329"},{"key":"e_1_3_3_232_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/N18-1101"},{"key":"e_1_3_3_233_2","volume-title":"ICML","author":"Yang Zhilin","year":"2016","unstructured":"Zhilin Yang, William Cohen, and Ruslan Salakhudinov. 2016. Revisiting semi-supervised learning with graph embeddings. In ICML."},{"key":"e_1_3_3_234_2","volume-title":"ICML","author":"You Jiaxuan","year":"2019","unstructured":"Jiaxuan You, Rex Ying, and Jure Leskovec. 2019. Position-aware graph neural networks. In ICML."},{"key":"e_1_3_3_235_2","volume-title":"ICLR","author":"Zeng Hanqing","year":"2020","unstructured":"Hanqing Zeng, Hongkuan Zhou, Ajitesh Srivastava, Rajgopal Kannan, and Viktor Prasanna. 2020. GraphSAINT: Graph sampling based inductive learning method. In ICLR."},{"key":"e_1_3_3_236_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.463"}],"container-title":["ACM Computing Surveys"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3728458","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3728458","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T01:18:36Z","timestamp":1750295916000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3728458"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,5,6]]},"references-count":235,"journal-issue":{"issue":"9","published-print":{"date-parts":[[2025,9,30]]}},"alternative-id":["10.1145\/3728458"],"URL":"https:\/\/doi.org\/10.1145\/3728458","relation":{},"ISSN":["0360-0300","1557-7341"],"issn-type":[{"value":"0360-0300","type":"print"},{"value":"1557-7341","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,5,6]]},"assertion":[{"value":"2023-08-09","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-04-01","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-05-06","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}