{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,18]],"date-time":"2026-05-18T13:43:12Z","timestamp":1779111792553,"version":"3.51.4"},"reference-count":84,"publisher":"MIT Press","license":[{"start":{"date-parts":[[2024,9,19]],"date-time":"2024-09-19T00:00:00Z","timestamp":1726704000000},"content-version":"vor","delay-in-days":262,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["direct.mit.edu"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2024,9,18]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>Despite the remarkable performance of generative large language models (LLMs) on abstractive summarization, they face two significant challenges: their considerable size and tendency to hallucinate. Hallucinations are concerning because they erode reliability and raise safety issues. Pruning is a technique that reduces model size by removing redundant weights, enabling more efficient sparse inference. Pruned models yield downstream task performance comparable to the original, making them ideal alternatives when operating on a limited budget. However, the effect that pruning has upon hallucinations in abstractive summarization with LLMs has yet to be explored. In this paper, we provide an extensive empirical study across five summarization datasets, two state-of-the-art pruning methods, and five instruction-tuned LLMs. Surprisingly, we find that hallucinations are less prevalent from pruned LLMs than the original models. Our analysis suggests that pruned models tend to depend more on the source document for summary generation. This leads to a higher lexical overlap between the generated summary and the source document, which could be a reason for the reduction in hallucination risk.1<\/jats:p>","DOI":"10.1162\/tacl_a_00695","type":"journal-article","created":{"date-parts":[[2024,9,19]],"date-time":"2024-09-19T19:44:48Z","timestamp":1726775088000},"page":"1163-1181","update-policy":"https:\/\/doi.org\/10.1162\/mitpressjournals.corrections.policy","source":"Crossref","is-referenced-by-count":16,"title":["Investigating Hallucinations in Pruned Large Language Models for Abstractive Summarization"],"prefix":"10.1162","volume":"12","author":[{"given":"George","family":"Chrysostomou","sequence":"first","affiliation":[{"name":"AstraZeneca, UK"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zhixue","family":"Zhao","sequence":"additional","affiliation":[{"name":"University of Sheffield, UK. zhixue.zhao@sheffield.ac.uk"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Miles","family":"Williams","sequence":"additional","affiliation":[{"name":"University of Sheffield, UK. mwilliams15@sheffield.ac.uk"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Nikolaos","family":"Aletras","sequence":"additional","affiliation":[{"name":"University of Sheffield, UK. n.aletras@sheffield.ac.uk"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"281","published-online":{"date-parts":[[2024,9,18]]},"reference":[{"key":"2024091919443610600_bib1","doi-asserted-by":"publisher","first-page":"68","DOI":"10.18653\/v1\/2023.newsum-1.7","article-title":"From sparse to dense: GPT-4 summarization with chain of density prompting","volume-title":"Proceedings of the 4th New Frontiers in Summarization Workshop","author":"Adams","year":"2023"},{"key":"2024091919443610600_bib2","article-title":"The Falcon series of open language models","volume":"arXiv:2311.16867","author":"Almazrouei","year":"2023","journal-title":"arXiv preprint"},{"key":"2024091919443610600_bib3","first-page":"129","article-title":"What is the state of neural network pruning?","volume-title":"Proceedings of Machine Learning and Systems","author":"Blalock","year":"2020"},{"key":"2024091919443610600_bib4","doi-asserted-by":"publisher","first-page":"6251","DOI":"10.18653\/v1\/2020.emnlp-main.506","article-title":"Factual error correction for abstractive summarization models","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)","author":"Cao","year":"2020"},{"key":"2024091919443610600_bib5","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v32i1.11912","article-title":"Faithful to the original: Fact-aware neural abstractive summarization","volume-title":"Proceedings of the Thirty-Second AAAI Conference on Artificial Intelligence and Thirtieth Innovative Applications of Artificial Intelligence Conference and Eighth AAAI Symposium on Educational Advances in Artificial Intelligence","author":"Cao","year":"2018"},{"key":"2024091919443610600_bib6","first-page":"24516","article-title":"Towards improving faithfulness in abstractive summarization","volume-title":"Advances in Neural Information Processing Systems","author":"Chen","year":"2022"},{"key":"2024091919443610600_bib7","doi-asserted-by":"publisher","first-page":"10755","DOI":"10.18653\/v1\/2023.findings-acl.685","article-title":"CaPE: Contrastive parameter ensembling for reducing hallucination in abstractive summarization","volume-title":"Findings of the Association for Computational Linguistics: ACL 2023","author":"Choubey","year":"2023"},{"key":"2024091919443610600_bib8","doi-asserted-by":"publisher","first-page":"6920","DOI":"10.18653\/v1\/2022.acl-long.477","article-title":"An empirical study on explanations in out-of-domain settings","volume-title":"Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Chrysostomou","year":"2022"},{"key":"2024091919443610600_bib9","doi-asserted-by":"publisher","first-page":"6104","DOI":"10.18653\/v1\/2021.emnlp-main.493","article-title":"Perhaps PTLMs should go to school \u2013 a task to assess open book and closed book QA","volume-title":"Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing","author":"Ciosici","year":"2021"},{"key":"2024091919443610600_bib10","doi-asserted-by":"publisher","first-page":"137","DOI":"10.3115\/1599081.1599099","article-title":"Sentence compression beyond word deletion","volume-title":"Proceedings of the 22nd International Conference on Computational Linguistics (Coling 2008)","author":"Cohn","year":"2008"},{"key":"2024091919443610600_bib11","article-title":"GPT3.int8(): 8-bit matrix multiplication for transformers at scale","volume-title":"Advances in Neural Information Processing Systems","author":"Dettmers","year":"2022"},{"key":"2024091919443610600_bib12","doi-asserted-by":"publisher","first-page":"774","DOI":"10.1162\/tacl_a_00397","article-title":"Towards question-answering as an automatic metric for evaluating the content quality of a summary","volume":"9","author":"Deutsch","year":"2021","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"2024091919443610600_bib13","article-title":"Sweeping heterogeneity with smart mops: Mixture of prompts for LLM task adaptation","author":"Dun","year":"2023","journal-title":"arXiv preprint"},{"key":"2024091919443610600_bib14","doi-asserted-by":"publisher","first-page":"5055","DOI":"10.18653\/v1\/2020.acl-main.454","article-title":"FEQA: A question answering evaluation framework for faithfulness assessment in abstractive summarization","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics","author":"Durmus","year":"2020"},{"key":"2024091919443610600_bib15","doi-asserted-by":"publisher","first-page":"1012","DOI":"10.1162\/tacl_a_00410","article-title":"Measuring and improving consistency in pretrained language models","volume":"9","author":"Elazar","year":"2021","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"2024091919443610600_bib16","doi-asserted-by":"publisher","first-page":"391","DOI":"10.1162\/tacl_a_00373","article-title":"SummEval: Re- evaluating summarization evaluation","volume":"9","author":"Fabbri","year":"2021","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"2024091919443610600_bib17","doi-asserted-by":"publisher","first-page":"2214","DOI":"10.18653\/v1\/P19-1213","article-title":"Ranking generated summaries by correctness: An interesting but challenging application for natural language inference","volume-title":"Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics","author":"Falke","year":"2019"},{"key":"2024091919443610600_bib18","doi-asserted-by":"publisher","first-page":"3046","DOI":"10.18653\/v1\/2022.findings-acl.240","article-title":"Factual consistency of multilingual pretrained language models","volume-title":"Findings of the Association for Computational Linguistics: ACL 2022","author":"Fierro","year":"2022"},{"key":"2024091919443610600_bib19","first-page":"4475","article-title":"Optimal brain compression: A framework for accurate post-training quantization and pruning","volume-title":"Advances in Neural Information Processing Systems","author":"Frantar","year":"2022"},{"key":"2024091919443610600_bib20","first-page":"10323","article-title":"SparseGPT: Massive language models can be accurately pruned in one-shot","volume-title":"Proceedings of the 40th International Conference on Machine Learning","author":"Frantar","year":"2023"},{"key":"2024091919443610600_bib21","article-title":"OPTQ: Accurate quantization for generative pre-trained transformers","volume-title":"The Eleventh International Conference on Learning Representations","author":"Frantar","year":"2023"},{"key":"2024091919443610600_bib22","doi-asserted-by":"publisher","first-page":"1061","DOI":"10.1162\/tacl_a_00413","article-title":"Compressing large-scale transformer- based models: A case study on BERT","volume":"9","author":"Ganesh","year":"2021","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"2024091919443610600_bib23","doi-asserted-by":"publisher","first-page":"13766","DOI":"10.18653\/v1\/2023.acl-long.770","article-title":"Optimal transport for unsupervised hallucination detection in neural machine translation","volume-title":"Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Guerreiro","year":"2023"},{"key":"2024091919443610600_bib24","doi-asserted-by":"publisher","first-page":"6098","DOI":"10.18653\/v1\/D19-1632","article-title":"The FLORES evaluation datasets for low-resource machine translation: Nepali\u2013English and Sinhala\u2013English","volume-title":"Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)","author":"Guzm\u00e1n","year":"2019"},{"issue":"2","key":"2024091919443610600_bib25","doi-asserted-by":"publisher","first-page":"207","DOI":"10.1016\/0925-2312(94)90055-8","article-title":"A simple and effective method for removal of hidden units and weights","volume":"6","author":"Hagiwara","year":"1994","journal-title":"Neurocomputing"},{"key":"2024091919443610600_bib26","article-title":"Learning both weights and connections for efficient neural network","volume-title":"Advances in Neural Information Processing Systems","author":"Han","year":"2015"},{"key":"2024091919443610600_bib27","article-title":"Pruning for protection: Increasing jailbreak resistance in aligned LLMs without fine-tuning","author":"Hasan","year":"2024","journal-title":"arXiv preprint"},{"key":"2024091919443610600_bib28","doi-asserted-by":"publisher","first-page":"293","DOI":"10.1109\/ICNN.1993.298572","article-title":"Optimal brain surgeon and general network pruning","volume-title":"IEEE International Conference on Neural Networks","author":"Hassibi","year":"1993"},{"key":"2024091919443610600_bib29","doi-asserted-by":"publisher","first-page":"446","DOI":"10.18653\/v1\/2020.emnlp-main.33","article-title":"What have we achieved on text summarization?","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)","author":"Huang","year":"2020"},{"key":"2024091919443610600_bib30","unstructured":"Srinivasan\n              Iyer\n            , XiVictoria Lin, RamakanthPasunuru, TodorMihaylov, DanielSimig, PingYu, KurtShuster, TianluWang, QingLiu, Punit SinghKoura, XianLi, BrianO\u2019Horo, GabrielPereyra, JeffWang, ChristopherDewan, AsliCelikyilmaz, LukeZettlemoyer, and VesStoyanov. 2023. OPT-IML: Scaling language model instruction meta learning through the lens of generalization. arXiv preprint, arXiv:2212.12017."},{"key":"2024091919443610600_bib31","article-title":"Compressing LLMs: The truth is rarely pure and never simple","volume-title":"The Twelfth International Conference on Learning Representations","author":"Jaiswal","year":"2024"},{"key":"2024091919443610600_bib32","doi-asserted-by":"publisher","first-page":"1827","DOI":"10.18653\/v1\/2023.findings-emnlp.123","article-title":"Towards mitigating LLM hallucination via self reflection","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2023","author":"Ji","year":"2023"},{"key":"2024091919443610600_bib33","article-title":"Mistral 7B","author":"Jiang","year":"2023","journal-title":"arXiv preprint"},{"key":"2024091919443610600_bib34","doi-asserted-by":"publisher","first-page":"555","DOI":"10.18653\/v1\/2022.gem-1.51","article-title":"Don\u2019t say what you don\u2019t know: Improving the consistency of abstractive summarization by constraining beam search","volume-title":"Proceedings of the 2nd Workshop on Natural Language Generation, Evaluation, and Metrics (GEM)","author":"King","year":"2022"},{"key":"2024091919443610600_bib35","doi-asserted-by":"publisher","first-page":"9332","DOI":"10.18653\/v1\/2020.emnlp-main.750","article-title":"Evaluating the factual consistency of abstractive text summarization","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)","author":"Kryscinski","year":"2020"},{"key":"2024091919443610600_bib36","doi-asserted-by":"publisher","first-page":"9662","DOI":"10.18653\/v1\/2023.emnlp-main.600","article-title":"SummEdits: Measuring LLM ability at factual reasoning through the lens of summarization","volume-title":"Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing","author":"Laban","year":"2023"},{"key":"2024091919443610600_bib37","doi-asserted-by":"publisher","first-page":"163","DOI":"10.1162\/tacl_a_00453","article-title":"SummaC: Re-visiting NLI-based models for inconsistency detection in summarization","volume":"10","author":"Laban","year":"2022","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"2024091919443610600_bib38","doi-asserted-by":"publisher","first-page":"2853","DOI":"10.18653\/v1\/2023.emnlp-main.172","article-title":"Critic-driven decoding for mitigating hallucinations in data-to-text generation","volume-title":"Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing","author":"Lango","year":"2023"},{"key":"2024091919443610600_bib39","doi-asserted-by":"publisher","first-page":"343","DOI":"10.18653\/v1\/2023.emnlp-industry.33","article-title":"Building real-world meeting summarization systems using large language models: A practical perspective","volume-title":"Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing: Industry Track","author":"Md","year":"2023"},{"key":"2024091919443610600_bib40","article-title":"Optimal brain damage","volume-title":"Advances in Neural Information Processing Systems","author":"LeCun","year":"1989"},{"key":"2024091919443610600_bib41","first-page":"74","article-title":"ROUGE: A package for automatic evaluation of summaries","volume-title":"Text Summarization Branches Out","author":"Lin","year":"2004"},{"issue":"01","key":"2024091919443610600_bib42","doi-asserted-by":"publisher","first-page":"9815","DOI":"10.1609\/aaai.v33i01.33019815","article-title":"Abstractive summarization: A survey of the state of the art","volume":"33","author":"Lin","year":"2019","journal-title":"Proceedings of the AAAI Conference on Artificial Intelligence"},{"key":"2024091919443610600_bib43","article-title":"LLM-Pruner: On the structural pruning of large language models","volume-title":"Thirty-seventh Conference on Neural Information Processing Systems","author":"Ma","year":"2023"},{"key":"2024091919443610600_bib44","doi-asserted-by":"publisher","first-page":"1","DOI":"10.18653\/v1\/W19-2201","article-title":"Plain English summarization of contracts","volume-title":"Proceedings of the Natural Legal Language Processing Workshop 2019","author":"Manor","year":"2019"},{"key":"2024091919443610600_bib45","doi-asserted-by":"publisher","first-page":"1906","DOI":"10.18653\/v1\/2020.acl-main.173","article-title":"On faithfulness and factuality in abstractive summarization","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics","author":"Maynez","year":"2020"},{"key":"2024091919443610600_bib46","article-title":"Pointer sentinel mixture models","volume-title":"International Conference on Learning Representations","author":"Merity","year":"2017"},{"key":"2024091919443610600_bib47","doi-asserted-by":"publisher","first-page":"529","DOI":"10.18653\/v1\/2023.clinicalnlp-1.56","article-title":"Calvados at MEDIQA-chat 2023: Improving clinical note generation with multi-task instruction finetuning","volume-title":"Proceedings of the 5th Clinical Natural Language Processing Workshop","author":"Milintsevich","year":"2023"},{"key":"2024091919443610600_bib48","article-title":"Accelerating sparse deep neural networks","author":"Mishra","year":"2021","journal-title":"arXiv preprint"},{"key":"2024091919443610600_bib49","first-page":"7197","article-title":"Up or down? Adaptive rounding for post-training quantization","volume-title":"Proceedings of the 37th International Conference on Machine Learning","author":"Nagel","year":"2020"},{"key":"2024091919443610600_bib50","doi-asserted-by":"publisher","first-page":"280","DOI":"10.18653\/v1\/K16-1028","article-title":"Abstractive text summarization using sequence-to-sequence RNNs and beyond","volume-title":"Proceedings of the 20th SIGNLL Conference on Computational Natural Language Learning","author":"Nallapati","year":"2016"},{"key":"2024091919443610600_bib51","doi-asserted-by":"publisher","first-page":"5255","DOI":"10.18653\/v1\/2023.findings-emnlp.349","article-title":"The cost of compression: Investigating the impact of compression on parametric knowledge in language models","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2023","author":"Namburi","year":"2023"},{"key":"2024091919443610600_bib52","doi-asserted-by":"publisher","first-page":"974","DOI":"10.1162\/tacl_a_00583","article-title":"Conditional generation with a question-answering blueprint","volume":"11","author":"Narayan","year":"2023","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"2024091919443610600_bib53","article-title":"DeepSparse","author":"Magic","year":"2021"},{"key":"2024091919443610600_bib54","unstructured":"OpenAI, JoshAchiam, StevenAdler, SandhiniAgarwal, LamaAhmad, IlgeAkkaya, Florencia LeoniAleman, DiogoAlmeida, JankoAltenschmidt, SamAltman, ShyamalAnadkat, RedAvila, IgorBabuschkin, SuchirBalaji, ValerieBalcom, PaulBaltescu, HaimingBao, MohammadBavarian, JeffBelgum, IrwanBello, JakeBerdine, GabrielBernadett-Shapiro, ChristopherBerner, LennyBogdonoff, OlegBoiko, MadelaineBoyd, Anna-LuisaBrakman, GregBrockman, TimBrooks, MilesBrundage, KevinButton, TrevorCai, RosieCampbell, AndrewCann, BrittanyCarey, ChelseaCarlson, RoryCarmichael, BrookeChan, CheChang, FotisChantzis, DerekChen, SullyChen, RubyChen, JasonChen, MarkChen, BenChess, ChesterCho, CaseyChu, Hyung WonChung, DaveCummings, JeremiahCurrier, YunxingDai, CoryDecareaux, ThomasDegry, NoahDeutsch, DamienDeville, ArkaDhar, DavidDohan, SteveDowling, SheilaDunning, AdrienEcoffet, AttyEleti, TynaEloundou, DavidFarhi, LiamFedus, NikoFelix, Sim\u00f3n PosadaFishman, JustonForte, IsabellaFulford, LeoGao, ElieGeorges, ChristianGibson, VikGoel, TarunGogineni, GabrielGoh, RaphaGontijo-Lopes, JonathanGordon, MorganGrafstein, ScottGray, RyanGreene, JoshuaGross, Shixiang ShaneGu, YufeiGuo, ChrisHallacy, JesseHan, JeffHarris, YuchenHe, MikeHeaton, JohannesHeidecke, ChrisHesse, AlanHickey, WadeHickey, PeterHoeschele, BrandonHoughton, KennyHsu, ShengliHu, XinHu, JoostHuizinga, ShantanuJain, ShawnJain, JoanneJang, AngelaJiang, RogerJiang, HaozhunJin, DennyJin, ShinoJomoto, BillieJonn, HeewooJun, TomerKaftan, \u0141ukaszKaiser, AliKamali, IngmarKanitscheider, Nitish ShirishKeskar, TabarakKhan, LoganKilpatrick, Jong WookKim, ChristinaKim, YongjikKim, Jan HendrikKirchner, JamieKiros, MattKnight, DanielKokotajlo, \u0141ukaszKondraciuk, AndrewKondrich, ArisKonstantinidis, KyleKosic, GretchenKrueger, VishalKuo, MichaelLampe, IkaiLan, TeddyLee, JanLeike, JadeLeung, DanielLevy, Chak MingLi, RachelLim, MollyLin, StephanieLin, MateuszLitwin, TheresaLopez, RyanLowe, PatriciaLue, AnnaMakanju, KimMalfacini, SamManning, TodorMarkov, YanivMarkovski, BiancaMartin, KatieMayer, AndrewMayne, BobMcGrew, Scott MayerMcKinney, ChristineMcLeavey, PaulMcMillan, JakeMcNeil, DavidMedina, AalokMehta, JacobMenick, LukeMetz, AndreyMishchenko, PamelaMishkin, VinnieMonaco, EvanMorikawa, DanielMossing, TongMu, MiraMurati, OlegMurk, DavidM\u00e9ly, AshvinNair, ReiichiroNakano, RajeevNayak, ArvindNeelakantan, RichardNgo, HyeonwooNoh, LongOuyang, CullenO\u2019Keefe, JakubPachocki, AlexPaino, JoePalermo, AshleyPantuliano, GiambattistaParascandolo, JoelParish, EmyParparita, AlexPassos, MikhailPavlov, AndrewPeng, AdamPerelman, Filipede Avila Belbute Peres, MichaelPetrov, Henrique Pondede Oliveira Pinto, MichaelPokorny, MichellePokrass, Vitchyr H.Pong, TollyPowell, AletheaPower, BorisPower, ElizabethProehl, RaulPuri, AlecRadford, JackRae, AdityaRamesh, CameronRaymond, FrancisReal, KendraRimbach, CarlRoss, BobRotsted, HenriRoussez, NickRyder, MarioSaltarelli, TedSanders, ShibaniSanturkar, GirishSastry, HeatherSchmidt, DavidSchnurr, JohnSchulman, DanielSelsam, KylaSheppard, TokiSherbakov, JessicaShieh, SarahShoker, PranavShyam, SzymonSidor, EricSigler, MaddieSimens, JordanSitkin, KatarinaSlama, IanSohl, BenjaminSokolowsky, YangSong, NatalieStaudacher, Felipe PetroskiSuch, NatalieSummers, IlyaSutskever, JieTang, NikolasTezak, Madeleine B.Thompson, PhilTillet, AminTootoonchian, ElizabethTseng, PrestonTuggle, NickTurley, JerryTworek, Juan Felipe Cer\u00f3nUribe, AndreaVallone, ArunVijayvergiya, ChelseaVoss, CarrollWainwright, Justin JayWang, AlvinWang, BenWang, JonathanWard, JasonWei, C. J.Weinmann, AkilaWelihinda, PeterWelinder, JiayiWeng, LilianWeng, MattWiethoff, DaveWillner, ClemensWinter, SamuelWolrich, HannahWong, LaurenWorkman, SherwinWu, JeffWu, MichaelWu, KaiXiao, TaoXu, SarahYoo, KevinYu, QimingYuan, WojciechZaremba, RowanZellers, ChongZhang, MarvinZhang, ShengjiaZhao, TianhaoZheng, JuntangZhuang, WilliamZhuk, and BarretZoph. 2024. GPT-4 technical report. arXiv preprint, arXiv:2303.08774."},{"key":"2024091919443610600_bib55","first-page":"27730","article-title":"Training language models to follow instructions with human feedback","volume":"35","author":"Ouyang","year":"2022","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2024091919443610600_bib56","doi-asserted-by":"publisher","first-page":"2463","DOI":"10.18653\/v1\/D19-1250","article-title":"Language models as knowledge bases?","volume-title":"Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)","author":"Petroni","year":"2019"},{"issue":"1","key":"2024091919443610600_bib57","first-page":"5485","article-title":"Exploring the limits of transfer learning with a unified text-to-text transformer","volume":"21","author":"Raffel","year":"2020","journal-title":"The Journal of Machine Learning Research"},{"key":"2024091919443610600_bib58","doi-asserted-by":"publisher","first-page":"1172","DOI":"10.18653\/v1\/2021.naacl-main.92","article-title":"The curious case of hallucinations in neural machine translation","volume-title":"Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Raunak","year":"2021"},{"key":"2024091919443610600_bib59","doi-asserted-by":"publisher","first-page":"4035","DOI":"10.18653\/v1\/D18-1437","article-title":"Object hallucination in image captioning","volume-title":"Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing","author":"Rohrbach","year":"2018"},{"key":"2024091919443610600_bib60","doi-asserted-by":"publisher","first-page":"3","DOI":"10.1007\/978-3-642-28569-1_1","volume-title":"Automatic Text Summarization: Past, Present and Future","author":"Saggion","year":"2013"},{"issue":"12","key":"2024091919443610600_bib61","doi-asserted-by":"publisher","first-page":"54","DOI":"10.1145\/3381831","article-title":"Green AI","volume":"63","author":"Schwartz","year":"2020","journal-title":"Communications of the ACM"},{"key":"2024091919443610600_bib62","first-page":"895","article-title":"HaRiM +: Evaluating summary quality with hallucination risk","volume-title":"Proceedings of the 2nd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 12th International Joint Conference on Natural Language Processing (Volume 1: Long Papers)","author":"Son","year":"2022"},{"key":"2024091919443610600_bib63","article-title":"A simple and effective pruning approach for large language models","volume-title":"The Twelfth International Conference on Learning Representations","author":"Sun","year":"2024"},{"key":"2024091919443610600_bib64","doi-asserted-by":"publisher","first-page":"11626","DOI":"10.18653\/v1\/2023.acl-long.650","article-title":"Understanding factual errors in summarization: Errors, summarizers, datasets, error detectors","volume-title":"Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Tang","year":"2023"},{"key":"2024091919443610600_bib65","doi-asserted-by":"publisher","first-page":"56","DOI":"10.18653\/v1\/2023.newsum-1.6","article-title":"In-context learning of large language models for controlled dialogue summarization: A holistic benchmark and empirical analysis","volume-title":"Proceedings of the 4th New Frontiers in Summarization Workshop","author":"Tang","year":"2023"},{"key":"2024091919443610600_bib66","article-title":"Llama 2: Open foundation and fine-tuned chat models","author":"Touvron","year":"2023","journal-title":"arXiv preprint"},{"key":"2024091919443610600_bib67","article-title":"A neural conversational model","author":"Vinyals","year":"2015","journal-title":"arXiv preprint"},{"key":"2024091919443610600_bib68","first-page":"605","article-title":"Generating (factual?) narrative summaries of RCTs: Experiments with neural multi-document summarization","volume":"2021","author":"Wallace","year":"2021","journal-title":"AMIA Summits on Translational Science Proceedings"},{"key":"2024091919443610600_bib69","doi-asserted-by":"publisher","first-page":"5008","DOI":"10.18653\/v1\/2020.acl-main.450","article-title":"Asking and answering questions to evaluate the factual consistency of summaries","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics","author":"Wang","year":"2020"},{"key":"2024091919443610600_bib70","doi-asserted-by":"publisher","first-page":"214","DOI":"10.1145\/3531146.3533088","article-title":"Taxonomy of risks posed by language models","volume-title":"Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency","author":"Weidinger","year":"2022"},{"key":"2024091919443610600_bib71","article-title":"How does calibration data affect the post-training pruning and quantization of large language models?","author":"Williams","year":"2023","journal-title":"arXiv preprint"},{"key":"2024091919443610600_bib72","doi-asserted-by":"publisher","first-page":"38","DOI":"10.18653\/v1\/2020.emnlp-demos.6","article-title":"Transformers: State-of-the-art natural language processing","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations","author":"Wolf","year":"2020"},{"key":"2024091919443610600_bib73","doi-asserted-by":"publisher","first-page":"2734","DOI":"10.18653\/v1\/2021.eacl-main.236","article-title":"On hallucination and predictive uncertainty in conditional language generation","volume-title":"Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume","author":"Xiao","year":"2021"},{"key":"2024091919443610600_bib74","article-title":"Can model compression improve NLP fairness","volume":"arXiv:2201.08542","author":"Guangxuan","year":"2022","journal-title":"arXiv preprint"},{"key":"2024091919443610600_bib75","doi-asserted-by":"publisher","first-page":"546","DOI":"10.1162\/tacl_a_00563","article-title":"Understanding and detecting hallucinations in neural machine translation via model introspection","volume":"11","author":"Weijia","year":"2023","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"2024091919443610600_bib76","first-page":"27263","article-title":"BARTScore: Evaluating generated text as text generation","volume-title":"Advances in Neural Information Processing Systems","author":"Yuan","year":"2021"},{"key":"2024091919443610600_bib77","article-title":"BERTScore: Evaluating text generation with BERT","volume-title":"International Conference on Learning Representations","author":"Zhang","year":"2020"},{"key":"2024091919443610600_bib78","doi-asserted-by":"publisher","first-page":"39","DOI":"10.1162\/tacl_a_00632","article-title":"Benchmarking Large Language Models for News Summarization","volume":"12","author":"Zhang","year":"2024","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"2024091919443610600_bib79","doi-asserted-by":"publisher","first-page":"4732","DOI":"10.18653\/v1\/2023.acl-long.261","article-title":"Incorporating attribution importance for improving faithfulness metrics","volume-title":"Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Zhao","year":"2023"},{"key":"2024091919443610600_bib80","doi-asserted-by":"publisher","first-page":"4039","DOI":"10.18653\/v1\/2022.findings-emnlp.298","article-title":"On the impact of temporal concept drift on model explanations","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2022","author":"Zhao","year":"2022"},{"key":"2024091919443610600_bib81","doi-asserted-by":"publisher","first-page":"2237","DOI":"10.18653\/v1\/2020.findings-emnlp.203","article-title":"Reducing quantity hallucinations in abstractive summarization","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2020","author":"Zhao","year":"2020"},{"key":"2024091919443610600_bib82","doi-asserted-by":"publisher","DOI":"10.22541\/au.170709121.16176681\/v1","article-title":"Reagent: A model-agnostic feature attribution method for generative language models","volume":"arXiv:2402.00794","author":"Zhao","year":"2024","journal-title":"arXiv preprint"},{"key":"2024091919443610600_bib83","doi-asserted-by":"publisher","first-page":"100205","DOI":"10.1016\/j.osnem.2022.100205","article-title":"Utilizing subjectivity level to mitigate identity term bias in toxic comments classification","volume":"29","author":"Zhao","year":"2022","journal-title":"Online Social Networks and Media"},{"key":"2024091919443610600_bib84","doi-asserted-by":"publisher","first-page":"1393","DOI":"10.18653\/v1\/2021.findings-acl.120","article-title":"Detecting hallucinated content in conditional neural sequence generation","volume-title":"Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021","author":"Zhou","year":"2021"}],"container-title":["Transactions of the Association for Computational Linguistics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/direct.mit.edu\/tacl\/article-pdf\/doi\/10.1162\/tacl_a_00695\/2470787\/tacl_a_00695.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/direct.mit.edu\/tacl\/article-pdf\/doi\/10.1162\/tacl_a_00695\/2470787\/tacl_a_00695.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,9,19]],"date-time":"2024-09-19T19:45:01Z","timestamp":1726775101000},"score":1,"resource":{"primary":{"URL":"https:\/\/direct.mit.edu\/tacl\/article\/doi\/10.1162\/tacl_a_00695\/124459\/Investigating-Hallucinations-in-Pruned-Large"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024]]},"references-count":84,"URL":"https:\/\/doi.org\/10.1162\/tacl_a_00695","relation":{},"ISSN":["2307-387X"],"issn-type":[{"value":"2307-387X","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2024]]},"published":{"date-parts":[[2024]]}}}