{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,10]],"date-time":"2026-06-10T03:31:33Z","timestamp":1781062293017,"version":"3.54.1"},"reference-count":30,"publisher":"MIT Press","license":[{"start":{"date-parts":[[2023,8,16]],"date-time":"2023-08-16T00:00:00Z","timestamp":1692144000000},"content-version":"vor","delay-in-days":227,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["direct.mit.edu"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2023,8,15]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>In behavioral testing, system functionalities underrepresented in the standard evaluation setting (with a held-out test set) are validated through controlled input-output pairs. Optimizing performance on the behavioral tests during training (behavioral learning) would improve coverage of phenomena not sufficiently represented in the i.i.d. data and could lead to seemingly more robust models. However, there is the risk that the model narrowly captures spurious correlations from the behavioral test suite, leading to overestimation and misrepresentation of model performance\u2014one of the original pitfalls of traditional evaluation.<\/jats:p>\n               <jats:p>In this work, we introduce BeLUGA, an analysis method for evaluating behavioral learning considering generalization across dimensions of different granularity levels. We optimize behavior-specific loss functions and evaluate models on several partitions of the behavioral test suite controlled to leave out specific phenomena. An aggregate score measures generalization to unseen functionalities (or overfitting). We use BeLUGA to examine three representative NLP tasks (sentiment analysis, paraphrase identification, and reading comprehension) and compare the impact of a diverse set of regularization and domain generalization methods on generalization performance.1<\/jats:p>","DOI":"10.1162\/tacl_a_00590","type":"journal-article","created":{"date-parts":[[2023,8,16]],"date-time":"2023-08-16T19:47:17Z","timestamp":1692215237000},"page":"1066-1081","update-policy":"https:\/\/doi.org\/10.1162\/mitpressjournals.corrections.policy","source":"Crossref","is-referenced-by-count":3,"title":["Cross-functional Analysis of Generalization in Behavioral Learning"],"prefix":"10.1162","volume":"11","author":[{"given":"Pedro Henrique","family":"Luz de Araujo","sequence":"first","affiliation":[{"name":"Faculty of Computer Science, University of Vienna, Vienna, Austria"},{"name":"UniVie Doctoral School Computer Science, Vienna, Austria. pedro.henrique.luz.de.araujo@univie.ac.at"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Benjamin","family":"Roth","sequence":"additional","affiliation":[{"name":"Faculty of Computer Science, University of Vienna, Vienna, Austria"},{"name":"Faculty of Philological and Cultural Studies, University of Vienna, Vienna, Austria. benjamin.roth@univie.ac.at"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"281","published-online":{"date-parts":[[2023,8,15]]},"reference":[{"key":"2023081619465441900_bib1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.1907.02893","article-title":"Invariant risk minimization","author":"Arjovsky","year":"2019","journal-title":"CoRR"},{"key":"2023081619465441900_bib2","doi-asserted-by":"publisher","first-page":"49","DOI":"10.1162\/tacl_a_00254","article-title":"Analysis methods in neural language processing: A survey","volume":"7","author":"Belinkov","year":"2019","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"2023081619465441900_bib3","doi-asserted-by":"publisher","first-page":"4171","DOI":"10.18653\/v1\/N19-1423","article-title":"BERT: Pre-training of deep bidirectional transformers for language understanding","volume-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers)","author":"Devlin","year":"2019"},{"key":"2023081619465441900_bib4","first-page":"18212","article-title":"IRM\u2014when it works and when it doesn\u2019t: A test case of natural language inference","volume-title":"Advances in Neural Information Processing Systems","author":"Dranker","year":"2021"},{"key":"2023081619465441900_bib5","article-title":"First quora dataset release: Question pairs","author":"Iyer","year":"2017"},{"key":"2023081619465441900_bib6","doi-asserted-by":"publisher","first-page":"4110","DOI":"10.18653\/v1\/2021.naacl-main.324","article-title":"Dynabench: Rethinking benchmarking in NLP","volume-title":"Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Kiela","year":"2021"},{"key":"2023081619465441900_bib7","doi-asserted-by":"publisher","first-page":"1352","DOI":"10.18653\/v1\/2022.naacl-main.97","article-title":"Hatemoji: A test suite and adversarially-generated dataset for benchmarking and detecting emoji-based hate","volume-title":"Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Kirk","year":"2022"},{"key":"2023081619465441900_bib8","article-title":"Fine-tuning can distort pretrained features and underperform out-of-distribution","volume-title":"Proceedings of the 10th International Conference on Learning Representations","author":"Kumar","year":"2022"},{"key":"2023081619465441900_bib9","doi-asserted-by":"publisher","first-page":"5210","DOI":"10.18653\/v1\/2020.acl-main.465","article-title":"How can we accelerate progress towards human-like linguistic generalization?","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics","author":"Linzen","year":"2020"},{"key":"2023081619465441900_bib10","doi-asserted-by":"publisher","first-page":"2171","DOI":"10.18653\/v1\/N19-1225","article-title":"Inoculation by fine-tuning: A method for analyzing challenge datasets","volume-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers)","author":"Liu","year":"2019"},{"key":"2023081619465441900_bib11","article-title":"Decoupled weight decay regularization","volume-title":"Proceedings of the 7th International Conference on Learning Representations","author":"Loshchilov","year":"2019"},{"key":"2023081619465441900_bib12","doi-asserted-by":"publisher","first-page":"75","DOI":"10.18653\/v1\/2022.nlppower-1.8","article-title":"Checking HateCheck: A cross-functional analysis of behaviour-aware learning for hate speech detection","volume-title":"Proceedings of NLP Power! The First Workshop on Efficient Benchmarking in NLP","author":"Luz de Araujo","year":"2022"},{"key":"2023081619465441900_bib13","first-page":"10351","article-title":"Dynaboard: An evaluation-as-a-service platform for holistic next-generation benchmarking","volume-title":"Advances in Neural Information Processing Systems","author":"Ma","year":"2021"},{"key":"2023081619465441900_bib14","doi-asserted-by":"publisher","first-page":"79","DOI":"10.18653\/v1\/2022.deelio-1.8","article-title":"Fast few-shot debugging for NLU test suites","volume-title":"Proceedings of Deep Learning Inside Out (DeeLIO 2022): The 3rd Workshop on Knowledge Extraction and Integration for Deep Learning Architectures","author":"Malon","year":"2022"},{"key":"2023081619465441900_bib15","doi-asserted-by":"publisher","first-page":"3428","DOI":"10.18653\/v1\/P19-1334","article-title":"Right for the wrong reasons: Diagnosing syntactic heuristics in natural language inference","volume-title":"Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics","author":"McCoy","year":"2019"},{"key":"2023081619465441900_bib16","doi-asserted-by":"publisher","first-page":"4658","DOI":"10.18653\/v1\/P19-1459","article-title":"Probing neural network comprehension of natural language arguments","volume-title":"Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics","author":"Niven","year":"2019"},{"key":"2023081619465441900_bib17","doi-asserted-by":"publisher","first-page":"2383","DOI":"10.18653\/v1\/D16-1264","article-title":"SQuAD: 100,000+ questions for machine comprehension of text","volume-title":"Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing","author":"Rajpurkar","year":"2016"},{"key":"2023081619465441900_bib18","doi-asserted-by":"publisher","first-page":"4902","DOI":"10.18653\/v1\/2020.acl-main.442","article-title":"Beyond accuracy: Behavioral testing of NLP models with CheckList","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics","author":"Ribeiro","year":"2020"},{"key":"2023081619465441900_bib19","doi-asserted-by":"publisher","first-page":"41","DOI":"10.18653\/v1\/2021.acl-long.4","article-title":"HateCheck: Functional tests for hate speech detection models","volume-title":"Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers)","author":"R\u00f6ttger","year":"2021"},{"key":"2023081619465441900_bib20","doi-asserted-by":"publisher","first-page":"196","DOI":"10.18653\/v1\/K19-1019","article-title":"Diversify your datasets: Analyzing generalization via controlled variance in adversarial datasets","volume-title":"Proceedings of the 23rd Conference on Computational Natural Language Learning (CoNLL)","author":"Rozen","year":"2019"},{"key":"2023081619465441900_bib21","article-title":"Distributionally robust neural networks for group shifts: On the importance of regularization for worst-case generalization","volume-title":"Proceedings of the 8th International Conference on Learning Representations","author":"Sagawa","year":"2020"},{"key":"2023081619465441900_bib22","article-title":"Gradient matching for domain generalization","volume-title":"Proceedings of the 10th International Conference on Learning Representations","author":"Shi","year":"2022"},{"key":"2023081619465441900_bib23","first-page":"1631","article-title":"Recursive deep models for semantic compositionality over a sentiment treebank","volume-title":"Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing","author":"Socher","year":"2013"},{"key":"2023081619465441900_bib24","doi-asserted-by":"publisher","first-page":"9275","DOI":"10.18653\/v1\/2020.emnlp-main.746","article-title":"Dataset cartography: Mapping and diagnosing datasets with training dynamics","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)","author":"Swayamdipta","year":"2020"},{"key":"2023081619465441900_bib25","article-title":"Superglue: A stickier benchmark for general-purpose language understanding systems","volume-title":"Advances in Neural Information Processing Systems","author":"Wang","year":"2019"},{"key":"2023081619465441900_bib26","doi-asserted-by":"publisher","first-page":"353","DOI":"10.18653\/v1\/W18-5446","article-title":"GLUE: A multi-task benchmark and analysis platform for natural language understanding","volume-title":"Proceedings of the 2018 EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP","author":"Wang","year":"2018"},{"key":"2023081619465441900_bib27","doi-asserted-by":"publisher","first-page":"747","DOI":"10.18653\/v1\/P19-1073","article-title":"Errudite: Scalable, reproducible, and testable error analysis","volume-title":"Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics","author":"Tongshuang","year":"2019"},{"key":"2023081619465441900_bib28","doi-asserted-by":"crossref","DOI":"10.3115\/992730.992783","article-title":"More accurate tests for the statistical significance of result differences","volume-title":"COLING 2000 Volume 2: The 18th International Conference on Computational Linguistics","author":"Yeh","year":"2000"},{"key":"2023081619465441900_bib29","doi-asserted-by":"publisher","first-page":"1063","DOI":"10.18653\/v1\/2021.naacl-main.84","article-title":"Fine-tuning pre-trained language model with weak supervision: A contrastive-regularized self-training approach","volume-title":"Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Yue","year":"2021"},{"key":"2023081619465441900_bib30","doi-asserted-by":"publisher","first-page":"4791","DOI":"10.18653\/v1\/P19-1472","article-title":"HellaSwag: Can a machine really finish your sentence?","volume-title":"Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics","author":"Zellers","year":"2019"}],"container-title":["Transactions of the Association for Computational Linguistics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/direct.mit.edu\/tacl\/article-pdf\/doi\/10.1162\/tacl_a_00590\/2154470\/tacl_a_00590.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/direct.mit.edu\/tacl\/article-pdf\/doi\/10.1162\/tacl_a_00590\/2154470\/tacl_a_00590.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,8,16]],"date-time":"2023-08-16T19:47:27Z","timestamp":1692215247000},"score":1,"resource":{"primary":{"URL":"https:\/\/direct.mit.edu\/tacl\/article\/doi\/10.1162\/tacl_a_00590\/117216\/Cross-functional-Analysis-of-Generalization-in"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023]]},"references-count":30,"URL":"https:\/\/doi.org\/10.1162\/tacl_a_00590","relation":{},"ISSN":["2307-387X"],"issn-type":[{"value":"2307-387X","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2023]]},"published":{"date-parts":[[2023]]}}}