{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,18]],"date-time":"2026-03-18T03:57:27Z","timestamp":1773806247397,"version":"3.50.1"},"reference-count":0,"publisher":"Association for the Advancement of Artificial Intelligence (AAAI)","issue":"39","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["AAAI"],"abstract":"<jats:p>The ability of Large Language Models (LLMs) to use ex\nternal tools unlocks powerful real-world interactions, mak\ning rigorous evaluation essential. However, current bench\nmarks primarily report final accuracy, revealing what mod\nels can do but obscuring the cognitive bottlenecks that define\n their true capability boundaries. To move from simple per\nformance scoring to a diagnostic tool, we introduce a frame\nworkgroundedinCognitive LoadTheory.Ourframeworkde\nconstructs task complexity into two quantifiable components:\n Intrinsic Load, the inherent structural complexity of the solu\ntion path, formalized with a novel Tool Interaction Graph; and\n Extraneous Load, the difficulty arising from ambiguous task\n presentation. To enable controlled experiments, we construct\n ToolLoad-Bench, the first benchmark with parametrically ad\njustable cognitive load. Our evaluation reveals distinct per\nformance cliffs as cognitive load increases, allowing us to\n precisely map each model\u2019s capability boundary. We validate\n that our framework\u2019s predictions are highly calibrated with\n empirical results, establishing a principled methodology for\n understanding an agent\u2019s limits and a practical foundation for\n building more efficient systems.<\/jats:p>","DOI":"10.1609\/aaai.v40i39.40650","type":"journal-article","created":{"date-parts":[[2026,3,18]],"date-time":"2026-03-18T02:57:20Z","timestamp":1773802640000},"page":"33611-33619","source":"Crossref","is-referenced-by-count":0,"title":["Beyond Accuracy: A Cognitive Load Framework for Mapping the Capability Boundaries of Tool-use Agents"],"prefix":"10.1609","volume":"40","author":[{"given":"Qihao","family":"Wang","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yue","family":"Hu","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Mingzhe","family":"Lu","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jiayue","family":"Wu","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yanbing","family":"Liu","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yuanmin","family":"Tang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"9382","published-online":{"date-parts":[[2026,3,14]]},"container-title":["Proceedings of the AAAI Conference on Artificial Intelligence"],"original-title":[],"link":[{"URL":"https:\/\/ojs.aaai.org\/index.php\/AAAI\/article\/download\/40650\/44611","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/ojs.aaai.org\/index.php\/AAAI\/article\/download\/40650\/44611","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,3,18]],"date-time":"2026-03-18T02:57:23Z","timestamp":1773802643000},"score":1,"resource":{"primary":{"URL":"https:\/\/ojs.aaai.org\/index.php\/AAAI\/article\/view\/40650"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,3,14]]},"references-count":0,"journal-issue":{"issue":"39","published-online":{"date-parts":[[2026,3,17]]}},"URL":"https:\/\/doi.org\/10.1609\/aaai.v40i39.40650","relation":{},"ISSN":["2374-3468","2159-5399"],"issn-type":[{"value":"2374-3468","type":"electronic"},{"value":"2159-5399","type":"print"}],"subject":[],"published":{"date-parts":[[2026,3,14]]}}}