{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,6]],"date-time":"2026-07-06T18:13:59Z","timestamp":1783361639695,"version":"3.54.6"},"reference-count":83,"publisher":"MIT Press","license":[{"start":{"date-parts":[[2026,7,6]],"date-time":"2026-07-06T00:00:00Z","timestamp":1783296000000},"content-version":"vor","delay-in-days":186,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["direct.mit.edu"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2026,7,1]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>Language models (LMs) increasingly drive real-world applications that require world knowledge. However, the internal processes through which models turn data into representations of knowledge and beliefs about the world are poorly understood. To facilitate such studies, we present LMEnt, a suite including (1) a knowledge-rich pretraining corpus, fully annotated with entity mentions based on Wikipedia, (2) an entity-based retrieval method over pretraining data that outperforms existing tools by as much as 80.4%, and (3) 12 pretrained LMs with up to 1B parameters and 4K intermediate checkpoints, with comparable performance to popular open-source models on knowledge tasks. Together, these resources provide a controlled environment for analyzing connections between entity mentions in pretraining data and downstream performance. We show the utility of LMEnt by studying knowledge acquisition over training, finding that entity co-occurrence and mention forms\u2014which are difficult to study with existing tools\u2014affect learning trends. Moreover, as LMs form stronger associations between entities, their facts are harder to edit in-context, whereas inconsistencies in model predictions over training are indicative of editing success. We release LMEnt to support studies of knowledge in LMs, including knowledge representations, plasticity, editing, attribution, hallucinations, and learning dynamics.<\/jats:p>\n                  <jats:p>huggingface.co\/LMEnt<\/jats:p>\n                  <jats:p>github.com\/LMEnt<\/jats:p>","DOI":"10.1162\/tacl.a.746","type":"journal-article","created":{"date-parts":[[2026,7,6]],"date-time":"2026-07-06T17:15:00Z","timestamp":1783358100000},"page":"1654-1691","update-policy":"https:\/\/doi.org\/10.1162\/mitpressjournals.corrections.policy","source":"Crossref","is-referenced-by-count":0,"title":["LMEnt: A Suite for Analyzing Knowledge in Language Models from Pretraining Data to Representations"],"prefix":"10.1162","volume":"14","author":[{"given":"Daniela","family":"Gottesman","sequence":"first","affiliation":[{"name":"Tel Aviv University, gottesman3@mail.tau.ac.il"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Alon","family":"Gilaie-Dotan","sequence":"additional","affiliation":[{"name":"Tel Aviv University"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ido","family":"Cohen","sequence":"additional","affiliation":[{"name":"Tel Aviv University"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yoav","family":"Gur-Arieh","sequence":"additional","affiliation":[{"name":"Tel Aviv University"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Marius","family":"Mosbach","sequence":"additional","affiliation":[{"name":"Mila - Quebec AI Institute & McGill University"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ori","family":"Yoran","sequence":"additional","affiliation":[{"name":"Tel Aviv University"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Mor","family":"Geva","sequence":"additional","affiliation":[{"name":"Tel Aviv University"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"281","published-online":{"date-parts":[[2026,7,1]]},"reference":[{"key":"2026070613144920300_bib1","volume-title":"A review on language models as knowledge bases","author":"AlKhamissi","year":"2022"},{"key":"2026070613144920300_bib2","article-title":"SmolLM2: When smol goes big\u2013data-centric training of a small language model","volume-title":"arXiv preprint arXiv:2502.02737","author":"Allal","year":"2025"},{"key":"2026070613144920300_bib3","first-page":"1067","article-title":"Physics of language models: Part 3.1, knowledge storage and extraction","volume-title":"Forty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21\u201327, 2024","author":"Allen-Zhu","year":"2024"},{"key":"2026070613144920300_bib4","doi-asserted-by":"crossref","DOI":"10.2139\/ssrn.5250617","article-title":"Physics of language models: Part 3.3, knowledge capacity scaling laws","volume-title":"The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24\u201328, 2025","author":"Allen-Zhu","year":"2025"},{"key":"2026070613144920300_bib5","doi-asserted-by":"crossref","first-page":"5769","DOI":"10.18653\/v1\/2022.findings-emnlp.423","article-title":"Language models as agent models","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2022","author":"Andreas","year":"2022"},{"key":"2026070613144920300_bib6","doi-asserted-by":"crossref","first-page":"209","DOI":"10.18653\/v1\/2022.naacl-industry.24","article-title":"ReFinED: An efficient zero-shot-capable approach to end-to-end entity linking","volume-title":"Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies: Industry Track","author":"Ayoola","year":"2022"},{"key":"2026070613144920300_bib7","volume-title":"SmolLM3: Smol, Multilingual, Long-Context Reasoner","author":"Bakouch","year":"2025"},{"key":"2026070613144920300_bib8","article-title":"Lessons from studying two-hop latent reasoning","volume-title":"arXiv preprint arXiv:2411.16353","author":"Balesni","year":"2024"},{"key":"2026070613144920300_bib9","volume-title":"Elasticsearch: A distributed, RESTful search and analytics engine","author":"Banon","year":"2010"},{"key":"2026070613144920300_bib10","article-title":"Nvidia Nemotron Nano 2: An accurate and efficient hybrid mamba-transformer reasoning model","volume-title":"arXiv preprint arXiv:2508.14444","author":"Basant","year":"2025"},{"key":"2026070613144920300_bib11","article-title":"Establishing task scaling laws via compute-efficient model ladders","volume-title":"Second Conference on Language Modeling","author":"Bhagia","year":"2025"},{"key":"2026070613144920300_bib12","first-page":"2397","article-title":"Pythia: A suite for analyzing large language models across training and scaling","volume-title":"International conference on machine learning","author":"Biderman","year":"2023"},{"key":"2026070613144920300_bib13","doi-asserted-by":"crossref","DOI":"10.1109\/TKDE.2025.3554028","article-title":"A survey on mixture of experts in large language models","volume-title":"IEEE Transactions on Knowledge and Data Engineering","author":"Cai","year":"2025"},{"key":"2026070613144920300_bib14","doi-asserted-by":"crossref","first-page":"16051","DOI":"10.18653\/v1\/2025.acl-long.782","article-title":"The alternative annotator test for LLM-as-a-judge: How to statistically justify replacing human annotators with LLMs","volume-title":"Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Calderon","year":"2025"},{"key":"2026070613144920300_bib15","doi-asserted-by":"crossref","first-page":"60626","DOI":"10.52202\/079017-1939","article-title":"How do large language models acquire factual knowledge during pretraining?","volume":"37","author":"Chang","year":"2024","journal-title":"Advances in neural information processing systems"},{"key":"2026070613144920300_bib16","doi-asserted-by":"crossref","first-page":"283","DOI":"10.1162\/tacl_a_00644","article-title":"Evaluating the ripple effects of knowledge editing in language models","volume":"12","author":"Cohen","year":"2024","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"2026070613144920300_bib17","article-title":"Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities","volume-title":"ArXiv preprint: 2507.06261","author":"Comanici","year":"2025"},{"key":"2026070613144920300_bib18","doi-asserted-by":"crossref","first-page":"1280","DOI":"10.18653\/v1\/2024.acl-long.70","article-title":"DeepSeekMoE: Towards ultimate expert specialization in mixture-of-experts language models","volume-title":"Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Dai","year":"2024"},{"key":"2026070613144920300_bib19","first-page":"4171","article-title":"BERT: pre-training of deep bidirectional transformers for language understanding","volume-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019, Minneapolis, MN, USA, June 2\u20137, 2019, Volume 1 (Long and Short Papers)","author":"Devlin","year":"2019"},{"key":"2026070613144920300_bib20","first-page":"7735","article-title":"What\u2019s in my big data?","volume":"2024","author":"Elazar","year":"2024","journal-title":"International Conference on Learning Representations"},{"key":"2026070613144920300_bib21","article-title":"Measuring causal effects of data statistics on language model\u2019s factual predictions","volume-title":"arXiv preprint arXiv:2207.14251","author":"Elazar","year":"2022"},{"key":"2026070613144920300_bib22","article-title":"Tinystories: How small can language models be and still speak coherent english?","volume-title":"arXiv preprint arXiv:2305.07759","author":"Eldan","year":"2023"},{"key":"2026070613144920300_bib23","doi-asserted-by":"crossref","DOI":"10.63317\/2yu5nu76n7bu","article-title":"T-REx: A large scale alignment of natural language with knowledge base triples","volume-title":"Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018)","author":"Elsahar","year":"2018"},{"key":"2026070613144920300_bib24","doi-asserted-by":"crossref","first-page":"12216","DOI":"10.18653\/v1\/2023.emnlp-main.751","article-title":"Dissecting recall of factual associations in auto-regressive language models","volume-title":"Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing","author":"Geva","year":"2023"},{"key":"2026070613144920300_bib25","doi-asserted-by":"crossref","first-page":"136","DOI":"10.63317\/3yjtmzi2qxqk","article-title":"WikiCoref: An English coreference-annotated corpus of wikipedia articles","volume-title":"Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC\u201916)","author":"Ghaddar","year":"2016"},{"key":"2026070613144920300_bib26","doi-asserted-by":"crossref","first-page":"3994","DOI":"10.18653\/v1\/2024.emnlp-main.232","article-title":"Estimating knowledge in large language models without generating a single token","volume-title":"Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing","author":"Gottesman","year":"2024"},{"key":"2026070613144920300_bib27","doi-asserted-by":"crossref","first-page":"15789","DOI":"10.18653\/v1\/2024.acl-long.841","article-title":"OLMo: Accelerating the science of language models","volume-title":"Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Groeneveld","year":"2024"},{"key":"2026070613144920300_bib28","article-title":"A survey on LLM-as-a-judge","volume-title":"The Innovation","author":"Gu","year":"2024"},{"issue":"5","key":"2026070613144920300_bib29","doi-asserted-by":"crossref","first-page":"2351","DOI":"10.1007\/s10994-023-06495-7","article-title":"Training data influence analysis and estimation: A survey","volume":"113","author":"Hammoudeh","year":"2024","journal-title":"Machine Learning"},{"key":"2026070613144920300_bib30","doi-asserted-by":"crossref","first-page":"5553","DOI":"10.18653\/v1\/2020.acl-main.492","article-title":"Explaining black box predictions and unveiling data artifacts through influence functions","volume-title":"Proceedings of the 58th annual meeting of the association for computational linguistics","author":"Han","year":"2020"},{"key":"2026070613144920300_bib31","article-title":"DeBERTa: decodingenhanced BERT with disentangled attention","volume-title":"9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3\u20137, 2021","author":"He","year":"2021"},{"key":"2026070613144920300_bib32","doi-asserted-by":"crossref","DOI":"10.52202\/068431-2176","article-title":"Training compute-optimal large language models","volume-title":"arXiv preprint arXiv:2203.15556","author":"Hoffmann","year":"2022"},{"key":"2026070613144920300_bib33","article-title":"Training language models on the knowledge graph: Insights on hallucinations and their detectability","volume-title":"First Conference on Language Modeling","author":"Hron","year":"2024"},{"key":"2026070613144920300_bib34","doi-asserted-by":"crossref","first-page":"9691","DOI":"10.18653\/v1\/2025.acl-long.478","article-title":"Between circuits and chomsky: Pre-pretraining on formal languages imparts linguistic biases","volume-title":"Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Hu","year":"2025"},{"key":"2026070613144920300_bib35","article-title":"Qwen2.5-Coder technical report","volume-title":"arXiv preprint arXiv:2409.12186","author":"Hui","year":"2024"},{"key":"2026070613144920300_bib36","first-page":"15696","article-title":"Large language models struggle to learn long-tail knowledge","volume-title":"International Conference on Machine Learning, ICML 2023, 23\u201329 July 2023, Honolulu, Hawaii, USA","author":"Kandpal","year":"2023"},{"key":"2026070613144920300_bib37","doi-asserted-by":"crossref","first-page":"7721","DOI":"10.18653\/v1\/2023.findings-emnlp.518","article-title":"Impact of co-occurrence on factual knowledge of large language models","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2023","author":"Kang","year":"2023"},{"key":"2026070613144920300_bib38","doi-asserted-by":"crossref","first-page":"6769","DOI":"10.18653\/v1\/2020.emnlp-main.550","article-title":"Dense passage retrieval for open-domain question answering","volume-title":"Proceedings of the 2020 conference on empirical methods in natural language processing (EMNLP)","author":"Karpukhin","year":"2020"},{"key":"2026070613144920300_bib39","doi-asserted-by":"crossref","first-page":"552","DOI":"10.18653\/v1\/2020.conll-1.45","article-title":"Are pretrained language models symbolic reasoners over knowledge?","volume-title":"Proceedings of the 24th conference on computational natural language learning","author":"Kassner","year":"2020"},{"key":"2026070613144920300_bib40","doi-asserted-by":"crossref","first-page":"39","DOI":"10.1145\/3397271.3401075","article-title":"ColBERT: Efficient and effective passage search via contextualized late interaction over BERT","volume-title":"Proceedings of the 43rd International ACM SIGIR conference on research and development in Information Retrieval","author":"Khattab","year":"2020"},{"key":"2026070613144920300_bib41","first-page":"100232","article-title":"Knowledge entropy decay during language model pretraining hinders new knowledge acquisition","volume":"2025","author":"Kim","year":"2025","journal-title":"International Conference on Learning Representations"},{"key":"2026070613144920300_bib42","volume-title":"FlashDeBERTa","author":"Knowledgator","year":"2025"},{"key":"2026070613144920300_bib43","article-title":"NeMo: A toolkit for building AI applications using neural modules","volume-title":"arXiv preprint arXiv:1909.09577","author":"Kuchaiev","year":"2019"},{"key":"2026070613144920300_bib44","article-title":"LLM post-training: A deep dive into reasoning large language models","volume-title":"arXiv preprint arXiv:2502.21321","author":"Kumar","year":"2025"},{"key":"2026070613144920300_bib45","doi-asserted-by":"crossref","first-page":"1098","DOI":"10.1162\/tacl_a_00415","article-title":"PAQ: 65 million probably-asked questions and what you can do with them","volume":"9","author":"Lewis","year":"2021","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"2026070613144920300_bib46","first-page":"1813","article-title":"Implicit representations of meaning in neural language models","volume-title":"Proceedings of the 59th annual meeting of the Association for Computational Linguistics and the 11th international joint conference on natural language processing (Volume 1: Long papers)","author":"Li","year":"2021"},{"key":"2026070613144920300_bib47","first-page":"2757","article-title":"From generation to judgment: Opportunities and challenges of LLM-as-a-judge","volume-title":"Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, EMNLP 2025, Suzhou, China, November 4\u20139, 2025","author":"Li","year":"2025"},{"key":"2026070613144920300_bib48","doi-asserted-by":"crossref","first-page":"178","DOI":"10.18653\/v1\/2025.acl-demo.18","article-title":"OLMo-Trace: Tracing language model outputs back to trillions of training tokens","volume-title":"Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 3: System Demonstrations)","author":"Liu","year":"2025"},{"key":"2026070613144920300_bib49","article-title":"Infini-gram: Scaling unbounded n-gram language models to a trillion tokens","volume-title":"First Conference on Language Modeling","author":"Liu","year":"2024"},{"key":"2026070613144920300_bib50","doi-asserted-by":"crossref","DOI":"10.18653\/v1\/2025.findings-emnlp.113","article-title":"Tracing multilingual factual knowledge acquisition in pretraining","volume-title":"arXiv preprint arXiv:2505.14824","author":"Liu","year":"2025"},{"key":"2026070613144920300_bib51","article-title":"RoBERTa: A robustly optimized BERT pretraining approach","volume-title":"arXiv preprint arXiv:1907.11692","author":"Liu","year":"2019"},{"key":"2026070613144920300_bib52","article-title":"LLM360: Towards fully transparent open-source LLMs","volume-title":"First Conference on Language Modeling","author":"Liu","year":"2024"},{"key":"2026070613144920300_bib53","article-title":"Understanding and preventing capacity loss in reinforcement learning","volume-title":"The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25\u201329, 2022","author":"Lyle","year":"2022"},{"key":"2026070613144920300_bib54","article-title":"DataDecide: How to predict best pretraining data with small experiments","volume-title":"Forty-second International Conference on Machine Learning, ICML 2025, Vancouver, BC, Canada, July 13\u201319, 2025","author":"Magnusson","year":"2025"},{"key":"2026070613144920300_bib55","first-page":"9802","article-title":"When not to trust language models: Investigating effectiveness of parametric and non-parametric memories","volume-title":"Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2023, Toronto, Canada, July 9\u201314, 2023","author":"Mallen","year":"2023"},{"key":"2026070613144920300_bib56","doi-asserted-by":"crossref","first-page":"13380","DOI":"10.18653\/v1\/2024.acl-long.722","article-title":"Maverick: Efficient and accurate coreference resolution defying recent trends","volume-title":"Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Martinelli","year":"2024"},{"key":"2026070613144920300_bib57","doi-asserted-by":"crossref","first-page":"17359","DOI":"10.52202\/068431-1262","article-title":"Locating and editing factual associations in GPT","volume":"35","author":"Meng","year":"2022","journal-title":"Advances in neural information processing systems"},{"key":"2026070613144920300_bib58","first-page":"71963","article-title":"On linear representations and pretraining data frequency in language models","volume":"2025","author":"Merullo","year":"2025","journal-title":"International Conference on Learning Representations"},{"key":"2026070613144920300_bib59","first-page":"62061","article-title":"OLMoE: Open mixture-of-experts language models","volume":"2025","author":"Muennighoff","year":"2025","journal-title":"International Conference on Learning Representations"},{"key":"2026070613144920300_bib60","article-title":"2 OLMo 2 Furious","volume-title":"arXiv preprint arXiv:2501.00656","author":"OLMo","year":"2024"},{"key":"2026070613144920300_bib61","first-page":"48","article-title":"fairseq: A fast, extensible toolkit for sequence modeling","volume-title":"Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics (Demonstrations)","author":"Ott","year":"2019"},{"key":"2026070613144920300_bib62","doi-asserted-by":"crossref","first-page":"19889","DOI":"10.18653\/v1\/2025.findings-acl.1021","article-title":"How do LLMs acquire new knowledge? A knowledge circuits perspective on continual pre-training","volume-title":"Findings of the Association for Computational Linguistics: ACL 2025","author":"Ou","year":"2025"},{"key":"2026070613144920300_bib63","first-page":"2523","article-title":"KILT: a benchmark for knowledge intensive language tasks","volume-title":"Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Petroni","year":"2021"},{"key":"2026070613144920300_bib64","doi-asserted-by":"crossref","first-page":"2463","DOI":"10.18653\/v1\/D19-1250","article-title":"Language models as knowledge bases?","volume-title":"Proceedings of the 2019 conference on empirical methods in natural language processing and the 9th international joint conference on natural language processing (EMNLP-IJCNLP)","author":"Petroni","year":"2019"},{"key":"2026070613144920300_bib65","doi-asserted-by":"crossref","DOI":"10.52202\/079017-1139","article-title":"Dataset decomposition: Faster LLM training with variable sequence length curriculum","volume-title":"Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 10 \u2013 15, 2024","author":"Pouransari","year":"2024"},{"key":"2026070613144920300_bib66","doi-asserted-by":"crossref","first-page":"53728","DOI":"10.52202\/075280-2338","article-title":"Direct preference optimization: Your language model is secretly a reward model","volume":"36","author":"Rafailov","year":"2023","journal-title":"Advances in neural information processing systems"},{"key":"2026070613144920300_bib67","doi-asserted-by":"crossref","first-page":"840","DOI":"10.18653\/v1\/2022.findings-emnlp.59","article-title":"Impact of pretraining term frequencies on few-shot numerical reasoning","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2022","author":"Razeghi","year":"2022"},{"key":"2026070613144920300_bib68","first-page":"5418","article-title":"How much knowledge can you pack into the parameters of a language model?","volume-title":"Proceedings of the 2020 conference on empirical","author":"Roberts","year":"2020"},{"key":"2026070613144920300_bib69","doi-asserted-by":"crossref","first-page":"13295","DOI":"10.18653\/v1\/2025.findings-acl.688","article-title":"Compute optimal scaling of skills: Knowledge vs reasoning","volume-title":"Findings of the Association for Computational Linguistics: ACL 2025","author":"Roberts","year":"2025"},{"key":"2026070613144920300_bib70","first-page":"109","article-title":"Okapi at TREC-3","volume-title":"Proceedings of The Third Text REtrieval Conference, TREC 1994, Gaithersburg, Maryland, USA, November 2\u20134, 1994","author":"Robertson","year":"1994"},{"key":"2026070613144920300_bib71","article-title":"Code llama: Open foundation models for code","volume-title":"arXiv preprint arXiv:2308.12950","author":"Roziere","year":"2023"},{"key":"2026070613144920300_bib72","doi-asserted-by":"crossref","first-page":"6138","DOI":"10.18653\/v1\/2021.emnlp-main.496","article-title":"Simple entity-centric questions challenge dense retrievers","volume-title":"Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing","author":"Sciavolino","year":"2021"},{"key":"2026070613144920300_bib73","doi-asserted-by":"crossref","first-page":"15725","DOI":"10.18653\/v1\/2024.acl-long.840","article-title":"Dolma: An open corpus of three trillion tokens for language model pretraining research","volume-title":"Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Soldaini","year":"2024"},{"key":"2026070613144920300_bib74","article-title":"Overtrained language models are harder to fine-tune","volume-title":"Forty-second International Conference on Machine Learning, ICML 2025, Vancouver, BC, Canada, July 13\u201319, 2025","author":"Springer","year":"2025"},{"key":"2026070613144920300_bib75","first-page":"35159","article-title":"Why and how LLMs hallucinate: Connecting the dots with subsequence associations","volume":"38","author":"Sun","year":"2026","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2026070613144920300_bib76","article-title":"Llama: Open and efficient foundation language models","volume-title":"arXiv preprint arXiv:2302.13971","author":"Touvron","year":"2023"},{"key":"2026070613144920300_bib77","article-title":"Llama 2: Open foundation and fine-tuned chat models","volume-title":"arXiv preprint arXiv:2307.09288","author":"Touvron","year":"2023"},{"issue":"10","key":"2026070613144920300_bib78","doi-asserted-by":"crossref","first-page":"78","DOI":"10.1145\/2629489","article-title":"Wikidata: A free collaborative knowledgebase","volume":"57","author":"Vrande\u010di\u0107","year":"2014","journal-title":"Commun. ACM"},{"key":"2026070613144920300_bib79","article-title":"Generalization v.s. memorization: Tracing language models\u2019 capabilities back to pretraining data","volume-title":"The Thirteenth International Conference on Learning Representations","author":"Wang","year":"2025"},{"key":"2026070613144920300_bib80","doi-asserted-by":"crossref","first-page":"10210","DOI":"10.18653\/v1\/2024.acl-long.550","article-title":"Do large language models latently perform multi-hop reasoning?","volume-title":"Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Yang","year":"2024"},{"key":"2026070613144920300_bib81","doi-asserted-by":"crossref","first-page":"9924","DOI":"10.18653\/v1\/2023.emnlp-main.615","article-title":"Characterizing mechanisms for factual recall in language models","volume-title":"Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing","author":"Yu","year":"2023"},{"key":"2026070613144920300_bib82","doi-asserted-by":"crossref","first-page":"4862","DOI":"10.18653\/v1\/2023.emnlp-main.296","article-title":"Can we edit factual knowledge by in-context learning?","volume-title":"Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing","author":"Zheng","year":"2023"},{"key":"2026070613144920300_bib83","article-title":"How do language models learn facts? dynamics, curricula and hallucinations","volume-title":"Second Conference on Language Modeling","author":"Zucchet","year":"2025"}],"container-title":["Transactions of the Association for Computational Linguistics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/direct.mit.edu\/tacl\/article-pdf\/doi\/10.1162\/TACL.a.746\/2612499\/tacl.a.746.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/direct.mit.edu\/tacl\/article-pdf\/doi\/10.1162\/TACL.a.746\/2612499\/tacl.a.746.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,7,6]],"date-time":"2026-07-06T17:15:14Z","timestamp":1783358114000},"score":1,"resource":{"primary":{"URL":"https:\/\/direct.mit.edu\/tacl\/article\/doi\/10.1162\/TACL.a.746\/137435\/LMEnt-A-Suite-for-Analyzing-Knowledge-in-Language"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026]]},"references-count":83,"URL":"https:\/\/doi.org\/10.1162\/tacl.a.746","relation":{},"ISSN":["2307-387X"],"issn-type":[{"value":"2307-387X","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2026]]},"published":{"date-parts":[[2026]]}}}