{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,22]],"date-time":"2026-07-22T16:42:16Z","timestamp":1784738536452,"version":"3.55.0"},"reference-count":0,"publisher":"Association for the Advancement of Artificial Intelligence (AAAI)","issue":"1","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["AIES"],"abstract":"<jats:p>As language models (LMs) are increasingly deployed as\nautonomous agents, their robust adherence to human-assigned\nobjectives becomes crucial for safe operation. When these\nagents operate independently for extended periods without\nhuman oversight, even initially well-specified goals may\ngradually shift. Detecting and measuring goal drift - an\nagent's tendency to deviate from its original objective\nover time - presents significant challenges, as goals can\nshift gradually, causing only subtle behavioral changes.\nThis paper proposes a novel approach to analyzing goal\ndrift in LM agents. In our experiments, agents are first\nexplicitly given a goal through their system prompt, then\nexposed to competing objectives through environmental\npressures. We demonstrate that while the best-performing\nagent (a scaffolded version of Claude 3.5 Sonnet) maintains\nnearly perfect goal adherence for more than 100,000 tokens\nin our most difficult evaluation setting, all evaluated\nmodels exhibit some degree of goal drift. We also find that\ngoal drift correlates with models' increasing\nsusceptibility to pattern-matching behaviors as the context\nlength grows.<\/jats:p>","DOI":"10.1609\/aies.v8i1.36541","type":"journal-article","created":{"date-parts":[[2025,10,15]],"date-time":"2025-10-15T13:14:43Z","timestamp":1760534083000},"page":"192-203","source":"Crossref","is-referenced-by-count":5,"title":["Evaluating Goal Drift in Language Model Agents"],"prefix":"10.1609","volume":"8","author":[{"given":"Rauno","family":"Arike","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Elizabeth","family":"Donoway","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Henning","family":"Bartsch","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Marius","family":"Hobbhahn","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"9382","published-online":{"date-parts":[[2025,10,15]]},"container-title":["Proceedings of the AAAI\/ACM Conference on AI, Ethics, and Society"],"original-title":[],"link":[{"URL":"https:\/\/ojs.aaai.org\/index.php\/AIES\/article\/download\/36541\/38679","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/ojs.aaai.org\/index.php\/AIES\/article\/download\/36541\/38679","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,15]],"date-time":"2025-10-15T13:14:43Z","timestamp":1760534083000},"score":1,"resource":{"primary":{"URL":"https:\/\/ojs.aaai.org\/index.php\/AIES\/article\/view\/36541"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,10,15]]},"references-count":0,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2025,10,15]]}},"URL":"https:\/\/doi.org\/10.1609\/aies.v8i1.36541","relation":{},"ISSN":["3065-8365"],"issn-type":[{"value":"3065-8365","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,10,15]]}}}