{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,12,20]],"date-time":"2025-12-20T22:04:52Z","timestamp":1766268292858,"version":"3.28.0"},"reference-count":78,"publisher":"MIT Press","license":[{"start":{"date-parts":[[2024,10,10]],"date-time":"2024-10-10T00:00:00Z","timestamp":1728518400000},"content-version":"vor","delay-in-days":283,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["direct.mit.edu"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2024,10,2]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>Recent models for natural language understanding are inclined to exploit simple patterns in datasets, commonly known as shortcuts. These shortcuts hinge on spurious correlations between labels and latent features existing in the training data. At inference time, shortcut-dependent models are likely to generate erroneous predictions under distribution shifts, particularly when some latent features are no longer correlated with the labels. To avoid this, previous studies have trained models to eliminate the reliance on shortcuts. In this study, we explore a different direction: pessimistically aggregating the predictions of a mixture-of-experts, assuming each expert captures relatively different latent features. The experimental results demonstrate that our post-hoc control over the experts significantly enhances the model\u2019s robustness to the distribution shift in shortcuts. Additionally, we show that our approach has some practical advantages. We also analyze our model and provide results to support the assumption.1<\/jats:p>","DOI":"10.1162\/tacl_a_00701","type":"journal-article","created":{"date-parts":[[2024,10,10]],"date-time":"2024-10-10T16:21:28Z","timestamp":1728577288000},"page":"1268-1289","update-policy":"http:\/\/dx.doi.org\/10.1162\/mitpressjournals.corrections.policy","source":"Crossref","is-referenced-by-count":2,"title":["Not Eliminate but Aggregate: Post-Hoc Control over Mixture-of-Experts to Address Shortcut Shifts in Natural Language Understanding"],"prefix":"10.1162","volume":"12","author":[{"given":"Ukyo","family":"Honda","sequence":"first","affiliation":[{"name":"CyberAgent, Japan. honda_ukyo@cyberagent.co.jp"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Tatsushi","family":"Oka","sequence":"additional","affiliation":[{"name":"Keio University, Japan. tatsushi.oka@keio.jp"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Peinan","family":"Zhang","sequence":"additional","affiliation":[{"name":"CyberAgent, Japan. zhang_peinan@cyberagent.co.jp"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Masato","family":"Mita","sequence":"additional","affiliation":[{"name":"CyberAgent, Japan. mita_masato@cyberagent.co.jp"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"281","published-online":{"date-parts":[[2024,10,2]]},"reference":[{"key":"2024112220403871500_bib1","article-title":"Invariant risk minimization","author":"Arjovsky","year":"2019","journal-title":"arXiv preprint arXiv:1907.02893v3"},{"key":"2024112220403871500_bib2","doi-asserted-by":"publisher","first-page":"877","DOI":"10.18653\/v1\/P19-1084","article-title":"Don\u2019t take the premise for granted: Mitigating artifacts in natural language inference","volume-title":"Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics","author":"Belinkov","year":"2019"},{"key":"2024112220403871500_bib3","doi-asserted-by":"publisher","DOI":"10.1007\/978-1-4757-4286-2","volume-title":"Statistical Decision Theory and Bayesian Analysis","author":"Berger","year":"1985"},{"key":"2024112220403871500_bib4","doi-asserted-by":"publisher","first-page":"5514","DOI":"10.1353\/dis.2023.a923671","article-title":"DISCO: Distilling counterfactuals with large language models","volume-title":"Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Chen","year":"2023"},{"key":"2024112220403871500_bib5","doi-asserted-by":"publisher","first-page":"4069","DOI":"10.18653\/v1\/D19-1418","article-title":"Don\u2019t take the easy way out: Ensemble based methods for avoiding known dataset biases","volume-title":"Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)","author":"Clark","year":"2019"},{"key":"2024112220403871500_bib6","doi-asserted-by":"publisher","first-page":"3031","DOI":"10.18653\/v1\/2020.findings-emnlp.272","article-title":"Learning to model and ignore dataset bias with mixed capacity ensembles","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2020","author":"Clark","year":"2020"},{"key":"2024112220403871500_bib7","article-title":"Electra: Pre-training text encoders as discriminators rather than generators","volume-title":"International Conference on Learning Representations","author":"Clark","year":"2020"},{"issue":"3","key":"2024112220403871500_bib8","doi-asserted-by":"publisher","first-page":"C95\u2013C127","DOI":"10.1111\/ectj.12068","article-title":"Using mixtures in econometric models: A brief review and some new results","volume":"19","author":"Compiani","year":"2016","journal-title":"The Econometrics Journal"},{"key":"2024112220403871500_bib9","first-page":"2189","article-title":"Environment inference for invariant learning","volume-title":"Proceedings of the 38th International Conference on Machine Learning","author":"Creager","year":"2021"},{"issue":"226","key":"2024112220403871500_bib10","first-page":"1","article-title":"Underspecification presents challenges for credibility in modern machine learning","volume":"23","author":"D\u2019Amour","year":"2022","journal-title":"Journal of Machine Learning Research"},{"key":"2024112220403871500_bib11","doi-asserted-by":"publisher","first-page":"4171","DOI":"10.18653\/v1\/N19-1423","article-title":"BERT: Pre-training of deep bidirectional transformers for language understanding","volume-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers)","author":"Devlin","year":"2019"},{"key":"2024112220403871500_bib12","first-page":"2278","article-title":"Decorrelate irrelevant, purify relevant: Overcome textual spurious correlations from a feature perspective","volume-title":"Proceedings of the 29th International Conference on Computational Linguistics","author":"Dou","year":"2022"},{"key":"2024112220403871500_bib13","doi-asserted-by":"publisher","first-page":"915","DOI":"10.18653\/v1\/2021.naacl-main.71","article-title":"Towards interpreting and mitigating shortcut learning behavior of NLU models","volume-title":"Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Mengnan","year":"2021"},{"key":"2024112220403871500_bib14","doi-asserted-by":"publisher","first-page":"1766","DOI":"10.18653\/v1\/2023.eacl-main.129","article-title":"Robustness challenges in model distillation and pruning for natural language understanding","volume-title":"Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics","author":"Mengnan","year":"2023"},{"key":"2024112220403871500_bib15","doi-asserted-by":"publisher","first-page":"4326","DOI":"10.18653\/v1\/2022.naacl-main.321","article-title":"Informativeness and invariance: Two perspectives on spurious correlations in natural language","volume-title":"Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Eisenstein","year":"2022"},{"key":"2024112220403871500_bib16","doi-asserted-by":"publisher","first-page":"1138","DOI":"10.1162\/tacl_a_00511","article-title":"Causal inference in natural language processing: Estimation, prediction, interpretation and beyond","volume":"10","author":"Feder","year":"2022","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"2024112220403871500_bib17","article-title":"A review of sparse expert models in deep learning","author":"Fedus","year":"2022","journal-title":"arXiv preprint arXiv:2209.01667v1"},{"issue":"120","key":"2024112220403871500_bib18","first-page":"1","article-title":"Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity","volume":"23","author":"Fedus","year":"2022","journal-title":"Journal of Machine Learning Research"},{"key":"2024112220403871500_bib19","doi-asserted-by":"publisher","first-page":"4112","DOI":"10.18653\/v1\/2022.emnlp-main.275","article-title":"Kernel-whitening: Overcome dataset bias with isotropic sentence embedding","volume-title":"Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing","author":"Gao","year":"2022"},{"key":"2024112220403871500_bib20","doi-asserted-by":"publisher","first-page":"1801","DOI":"10.18653\/v1\/2021.emnlp-main.135","article-title":"Competency problems: On finding and removing artifacts in language data","volume-title":"Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing","author":"Gardner","year":"2021"},{"key":"2024112220403871500_bib21","doi-asserted-by":"publisher","first-page":"1161","DOI":"10.18653\/v1\/D19-1107","article-title":"Are we modeling the task or the annotator? An investigation of annotator bias in natural language understanding datasets","volume-title":"Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)","author":"Geva","year":"2019"},{"key":"2024112220403871500_bib22","doi-asserted-by":"publisher","first-page":"1923","DOI":"10.18653\/v1\/2021.findings-acl.168","article-title":"End-to-end self-debiasing framework for robust NLU training","volume-title":"Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021","author":"Ghaddar","year":"2021"},{"key":"2024112220403871500_bib23","doi-asserted-by":"publisher","first-page":"107","DOI":"10.18653\/v1\/N18-2017","article-title":"Annotation artifacts in natural language inference data","volume-title":"Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers)","author":"Gururangan","year":"2018"},{"key":"2024112220403871500_bib24","doi-asserted-by":"publisher","first-page":"132","DOI":"10.18653\/v1\/D19-6115","article-title":"Unlearn dataset bias in natural language inference by fitting the residual","volume-title":"Proceedings of the 2nd Workshop on Deep Learning Approaches for Low-Resource NLP (DeepLo 2019)","author":"He","year":"2019"},{"key":"2024112220403871500_bib25","article-title":"DeBERTav3: Improving deBERTa using ELECTRA-style pre-training with gradient-disentangled embedding sharing","volume-title":"The Eleventh International Conference on Learning Representations","author":"He","year":"2023"},{"issue":"498","key":"2024112220403871500_bib26","doi-asserted-by":"publisher","first-page":"711","DOI":"10.1080\/01621459.2012.682541","article-title":"Mixture of regression models with varying mixing proportions: A semiparametric approach","volume":"107","author":"Huang","year":"2012","journal-title":"Journal of the American Statistical Association"},{"key":"2024112220403871500_bib27","first-page":"38516","article-title":"On feature learning in the presence of spurious correlations","volume-title":"Advances in Neural Information Processing Systems","author":"Izmailov","year":"2022"},{"issue":"1","key":"2024112220403871500_bib28","doi-asserted-by":"publisher","first-page":"79","DOI":"10.1162\/neco.1991.3.1.79","article-title":"Adaptive mixtures of local experts","volume":"3","author":"Jacobs","year":"1991","journal-title":"Neural computation"},{"key":"2024112220403871500_bib29","article-title":"Decoupling representation and classifier for long-tailed recognition","volume-title":"International Conference on Learning Representations","author":"Kang","year":"2020"},{"key":"2024112220403871500_bib30","article-title":"Explaining the efficacy of counterfactually augmented data","volume-title":"International Conference on Learning Representations","author":"Kaushik","year":"2021"},{"key":"2024112220403871500_bib31","article-title":"Last layer re-training is sufficient for robustness to spurious correlations","volume-title":"The Eleventh International Conference on Learning Representations","author":"Kirichenko","year":"2023"},{"key":"2024112220403871500_bib32","first-page":"5815","article-title":"Out-of-distribution generalization via risk extrapolation (rex)","volume-title":"Proceedings of the 38th International Conference on Machine Learning","author":"Krueger","year":"2021"},{"key":"2024112220403871500_bib33","article-title":"GShard: Scaling giant models with conditional computation and automatic sharding","volume-title":"International Conference on Learning Representations","author":"Lepikhin","year":"2021"},{"key":"2024112220403871500_bib34","article-title":"A structured self-attentive sentence embedding","volume-title":"International Conference on Learning Representations","author":"Lin","year":"2017"},{"key":"2024112220403871500_bib35","first-page":"6781","article-title":"Just train twice: Improving group robustness without training group information","volume-title":"Proceedings of the 38th International Conference on Machine Learning","author":"Liu","year":"2021"},{"key":"2024112220403871500_bib36","first-page":"19189","article-title":"A win-win deal: Towards sparse and robust pre-trained language models","volume-title":"Advances in Neural Information Processing Systems","author":"Liu","year":"2022"},{"key":"2024112220403871500_bib37","article-title":"Roberta: A robustly optimized bert pretraining approach","author":"Liu","year":"2019","journal-title":"arXiv preprint arXiv:1907.11692v1"},{"key":"2024112220403871500_bib38","doi-asserted-by":"publisher","first-page":"8706","DOI":"10.18653\/v1\/2020.acl-main.769","article-title":"End-to-end bias mitigation by modelling biases in corpora","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics","author":"Mahabadi","year":"2020"},{"key":"2024112220403871500_bib39","first-page":"739","article-title":"Causally motivated shortcut removal using auxiliary labels","volume-title":"Proceedings of The 25th International Conference on Artificial Intelligence and Statistics","author":"Makar","year":"2022"},{"key":"2024112220403871500_bib40","doi-asserted-by":"publisher","first-page":"3428","DOI":"10.18653\/v1\/P19-1334","article-title":"Right for the wrong reasons: Diagnosing syntactic heuristics in natural language inference","volume-title":"Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics","author":"McCoy","year":"2019"},{"key":"2024112220403871500_bib41","doi-asserted-by":"publisher","first-page":"7607","DOI":"10.18653\/v1\/2022.emnlp-main.517","article-title":"Debiasing masks: A new framework for shortcut mitigation in NLU","volume-title":"Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing","author":"Meissner","year":"2022"},{"key":"2024112220403871500_bib42","first-page":"20673","article-title":"Learning from failure: De-biasing classifier from biased classifier","volume-title":"Advances in Neural Information Processing Systems","author":"Nam","year":"2020"},{"issue":"1","key":"2024112220403871500_bib43","doi-asserted-by":"publisher","first-page":"1750861","DOI":"10.1080\/25742558.2020.1750861","article-title":"Approximation by finite mixtures of continuous density functions that vanish at infinity","volume":"7","author":"Tin Nguyen","year":"2020","journal-title":"Cogent Mathematics & Statistics"},{"key":"2024112220403871500_bib44","doi-asserted-by":"publisher","first-page":"12700","DOI":"10.1109\/CVPR46437.2021.01251","article-title":"Counterfactual VQA: A cause-effect look at language bias","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Niu","year":"2021"},{"key":"2024112220403871500_bib45","article-title":"Nuisances via negativa: Adjusting for spurious correlations via data augmentation","author":"Puli","year":"2022","journal-title":"arXiv preprint arXiv:2210.01302v2"},{"key":"2024112220403871500_bib46","article-title":"Out-of-distribution generalization in the presence of nuisance-induced spurious correlations","volume-title":"International Conference on Learning Representations","author":"Puli","year":"2022"},{"key":"2024112220403871500_bib47","doi-asserted-by":"publisher","first-page":"12350","DOI":"10.18653\/v1\/2023.findings-acl.781","article-title":"HuaSLIM: Human attention motivated shortcut learning identification and mitigation for large language models","volume-title":"Findings of the Association for Computational Linguistics: ACL 2023","author":"Ren","year":"2023"},{"key":"2024112220403871500_bib48","article-title":"Distributionally robust neural networks","volume-title":"International Conference on Learning Representations","author":"Sagawa","year":"2020"},{"key":"2024112220403871500_bib49","article-title":"Learning from others\u2019 mistakes: Avoiding dataset biases without modeling them","volume-title":"International Conference on Learning Representations","author":"Sanh","year":"2021"},{"key":"2024112220403871500_bib50","doi-asserted-by":"publisher","first-page":"3419","DOI":"10.18653\/v1\/D19-1341","article-title":"Towards debiasing fact verification models","volume-title":"Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)","author":"Schuster","year":"2019"},{"key":"2024112220403871500_bib51","article-title":"Outrageously large neural networks: The sparsely-gated mixture-of-experts layer","volume-title":"International Conference on Learning Representations","author":"Shazeer","year":"2017"},{"issue":"10","key":"2024112220403871500_bib52","doi-asserted-by":"publisher","first-page":"11349","DOI":"10.1609\/aaai.v36i10.21386","article-title":"Supervising model attention with human explanations for robust natural language inference","volume":"36","author":"Stacey","year":"2022","journal-title":"Proceedings of the AAAI Conference on Artificial Intelligence"},{"key":"2024112220403871500_bib53","doi-asserted-by":"publisher","first-page":"3809","DOI":"10.18653\/v1\/2022.emnlp-main.251","article-title":"Logical reasoning with span-level predictions for interpretable and robust NLI models","volume-title":"Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing","author":"Stacey","year":"2022"},{"key":"2024112220403871500_bib54","doi-asserted-by":"publisher","first-page":"8281","DOI":"10.18653\/v1\/2020.emnlp-main.665","article-title":"Avoiding the hypothesis- only bias in natural language inference via ensemble adversarial training","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)","author":"Stacey","year":"2020"},{"key":"2024112220403871500_bib55","doi-asserted-by":"publisher","first-page":"809","DOI":"10.18653\/v1\/N18-1074","article-title":"FEVER: A large-scale dataset for fact extraction and VERification","volume-title":"Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers)","author":"Thorne","year":"2018"},{"issue":"10","key":"2024112220403871500_bib56","doi-asserted-by":"publisher","first-page":"11376","DOI":"10.1609\/aaai.v36i10.21389","article-title":"Debiasing NLU models via causal intervention and counterfactual reasoning","volume":"36","author":"Tian","year":"2022","journal-title":"Proceedings of the AAAI Conference on Artificial Intelligence"},{"volume-title":"Statistical Analysis of Finite Mixture Distributions","year":"1985","author":"Titterington","key":"2024112220403871500_bib57"},{"key":"2024112220403871500_bib58","doi-asserted-by":"crossref","first-page":"8717","DOI":"10.18653\/v1\/2020.acl-main.770","article-title":"Mind the trade-off: Debiasing NLU models without degrading the in-distribution performance","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics","author":"Utama","year":"2020"},{"key":"2024112220403871500_bib59","doi-asserted-by":"publisher","first-page":"7597","DOI":"10.18653\/v1\/2020.emnlp-main.613","article-title":"Towards debiasing NLU models from unknown biases","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)","author":"Utama","year":"2020"},{"key":"2024112220403871500_bib60","first-page":"16196","article-title":"Counterfactual invariance to spurious correlations in text classification","volume-title":"Advances in Neural Information Processing Systems","author":"Veitch","year":"2021"},{"volume-title":"Statistical Decision Functions","year":"1950","author":"Wald","key":"2024112220403871500_bib61"},{"key":"2024112220403871500_bib62","doi-asserted-by":"publisher","DOI":"10.4337\/9780857930873.00015","volume-title":"Advances in Discrete Choice: Mixture Models","author":"Walker","year":"2011"},{"key":"2024112220403871500_bib63","doi-asserted-by":"publisher","first-page":"5037","DOI":"10.18653\/v1\/2022.naacl-main.371","article-title":"Robust (controlled) table-to-text generation with structure-aware equivariance learning","volume-title":"Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Wang","year":"2022"},{"key":"2024112220403871500_bib64","doi-asserted-by":"publisher","first-page":"3431","DOI":"10.18653\/v1\/2020.findings-emnlp.308","article-title":"Identifying spurious correlations for robust text classification","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2020","author":"Wang","year":"2020"},{"key":"2024112220403871500_bib65","doi-asserted-by":"publisher","first-page":"2302","DOI":"10.18653\/v1\/2022.findings-emnlp.170","article-title":"AutoCAD: Automatically generate counterfactuals for mitigating shortcut learning","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2022","author":"Wen","year":"2022"},{"key":"2024112220403871500_bib66","doi-asserted-by":"publisher","first-page":"1112","DOI":"10.18653\/v1\/N18-1101","article-title":"A broad-coverage challenge corpus for sentence understanding through inference","volume-title":"Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers)","author":"Williams","year":"2018"},{"key":"2024112220403871500_bib67","first-page":"1666","article-title":"Less is better: Recovering intended-feature subspace to robustify NLU models","volume-title":"Proceedings of the 29th International Conference on Computational Linguistics","author":"Ting","year":"2022"},{"key":"2024112220403871500_bib68","doi-asserted-by":"publisher","first-page":"2660","DOI":"10.18653\/v1\/2022.acl-long.190","article-title":"Generating data to mitigate spurious correlations in natural language inference datasets","volume-title":"Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Yuxiang","year":"2022"},{"issue":"3","key":"2024112220403871500_bib69","doi-asserted-by":"publisher","first-page":"391","DOI":"10.1214\/19-STS698","article-title":"An overview of semiparametric extensions of finite mixture models","volume":"34","author":"Xiang","year":"2019","journal-title":"Statistical Science"},{"key":"2024112220403871500_bib70","first-page":"13657","article-title":"Uncertainty calibration for ensemble-based debiasing methods","volume-title":"Advances in Neural Information Processing Systems","author":"Xiong","year":"2021"},{"key":"2024112220403871500_bib71","doi-asserted-by":"publisher","first-page":"3319","DOI":"10.18653\/v1\/2021.eacl-main.291","article-title":"Increasing robustness to spurious correlations using forgettable examples","volume-title":"Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume","author":"Yaghoobzadeh","year":"2021"},{"key":"2024112220403871500_bib72","first-page":"39584","article-title":"Change is hard: A closer look at subpopulation shift","volume-title":"Proceedings of the 40th International Conference on Machine Learning","author":"Yang","year":"2023"},{"key":"2024112220403871500_bib73","article-title":"Breaking the softmax bottleneck: A high-rank RNN language model","volume-title":"International Conference on Learning Representations","author":"Yang","year":"2018"},{"key":"2024112220403871500_bib74","article-title":"Xlnet: Generalized autoregressive pretraining for language understanding","volume-title":"Advances in Neural Information Processing Systems","author":"Yang","year":"2019"},{"key":"2024112220403871500_bib75","first-page":"25407","article-title":"Improving out-of-distribution robustness via selective augmentation","volume-title":"Proceedings of the 39th International Conference on Machine Learning","author":"Yao","year":"2022"},{"key":"2024112220403871500_bib76","doi-asserted-by":"publisher","first-page":"11627","DOI":"10.18653\/v1\/2022.emnlp-main.799","article-title":"Interventional training for out-of-distribution natural language understanding","volume-title":"Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing","author":"Sicheng","year":"2022"},{"key":"2024112220403871500_bib77","doi-asserted-by":"publisher","first-page":"3298","DOI":"10.18653\/v1\/2022.emnlp-main.217","article-title":"Making pretrained language models good long-tailed learners","volume-title":"Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing","author":"Zhang","year":"2022"},{"key":"2024112220403871500_bib78","first-page":"1298","article-title":"PAWS: Paraphrase adversaries from word scrambling","volume-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers)","author":"Zhang","year":"2019"}],"container-title":["Transactions of the Association for Computational Linguistics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/direct.mit.edu\/tacl\/article-pdf\/doi\/10.1162\/tacl_a_00701\/2480600\/tacl_a_00701.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/direct.mit.edu\/tacl\/article-pdf\/doi\/10.1162\/tacl_a_00701\/2480600\/tacl_a_00701.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,11,22]],"date-time":"2024-11-22T20:40:58Z","timestamp":1732308058000},"score":1,"resource":{"primary":{"URL":"https:\/\/direct.mit.edu\/tacl\/article\/doi\/10.1162\/tacl_a_00701\/124836\/Not-Eliminate-but-Aggregate-Post-Hoc-Control-over"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024]]},"references-count":78,"URL":"https:\/\/doi.org\/10.1162\/tacl_a_00701","relation":{},"ISSN":["2307-387X"],"issn-type":[{"type":"electronic","value":"2307-387X"}],"subject":[],"published-other":{"date-parts":[[2024]]},"published":{"date-parts":[[2024]]}}}