{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,28]],"date-time":"2026-04-28T08:27:13Z","timestamp":1777364833967,"version":"3.51.4"},"reference-count":51,"publisher":"Springer Science and Business Media LLC","issue":"2","license":[{"start":{"date-parts":[[2025,8,21]],"date-time":"2025-08-21T00:00:00Z","timestamp":1755734400000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2025,8,21]],"date-time":"2025-08-21T00:00:00Z","timestamp":1755734400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100014736","name":"Lingnan University","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100014736","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["AI &amp; Soc"],"published-print":{"date-parts":[[2026,2]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>Many have argued, based on the Instrumental Convergence Thesis, that Artificial General Intelligences (AGIs) will exhibit power-seeking behavior. Such behavior, they warn, could harm human society and pose existential threats\u2014namely, the risk of human extinction or the permanent collapse of civilization. These arguments often rely on an implicit and underexamined assumption: that AGIs will develop world models\u2014internal representations of world dynamics\u2014that resemble those of humans. We challenge this assumption. We argue that once the anthropomorphic assumption\u2014that AGIs\u2019 world models will mirror our own\u2014is rejected, it becomes unclear whether AGIs would pursue the types of power commonly emphasized in the literature, or any familiar types of power at all. This analysis casts doubt on the strength of existing arguments linking the Instrumental Convergence Thesis to existential threats. Moreover, it reveals a deeper layer of uncertainty. AGIs with non-human world models may identify novel or unanticipated types of power that fall outside existing taxonomies, thereby posing underappreciated risks. We further argue that world model alignment\u2014an issue largely overlooked in comparison with value alignment\u2014should be recognized as a core dimension of AI alignment. We conclude by outlining several open questions to inform and guide future research.<\/jats:p>","DOI":"10.1007\/s00146-025-02572-8","type":"journal-article","created":{"date-parts":[[2025,8,21]],"date-time":"2025-08-21T13:29:07Z","timestamp":1755782947000},"page":"939-949","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":1,"title":["Will power-seeking AGIs harm human society?"],"prefix":"10.1007","volume":"41","author":[{"given":"Maomei","family":"Wang","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2025,8,21]]},"reference":[{"key":"2572_CR1","doi-asserted-by":"publisher","DOI":"10.1093\/pq\/pqae034","author":"A Bales","year":"2024","unstructured":"Bales A (2024) AI takeover and human disempowerment. Philos Q. https:\/\/doi.org\/10.1093\/pq\/pqae034","journal-title":"Philos Q"},{"key":"2572_CR2","doi-asserted-by":"crossref","unstructured":"Bobu A, Peng A, Agrawal P, Shah JA, Dragan AD (2024). Aligning human and robot representations. In: Proceedings of the 2024 ACM\/IEEE International Conference on Human-Robot Interaction, pp 42\u201354.","DOI":"10.1145\/3610977.3634987"},{"key":"2572_CR3","volume-title":"Superintelligence: paths, dangers, strategies","author":"N Bostrom","year":"2014","unstructured":"Bostrom N (2014) Superintelligence: paths, dangers, strategies. Oxford University Press"},{"key":"2572_CR4","unstructured":"Browne R (2025). Maths test stumps AI models trained by Google and OpenAI. Yahoo Finance. https:\/\/finance.yahoo.com\/news\/maths-test-stumps-ai-models-093000989.html. Accessed 31 July 2025"},{"key":"2572_CR5","unstructured":"Bubeck S, Chandrasekaran V, Eldan R, Gehrke J, Horvitz E, Kamar E, Lee P, Lee YT, Li Y, Lundberg S, Nori H, Palangi H, Ribeiro MT, Zhang Y (2023). Sparks of artificial general intelligence: Early experiments with GPT-4. arXiv preprint. https:\/\/arxiv.org\/abs\/2303.12712"},{"key":"2572_CR6","volume-title":"Essays on longtermism","author":"J Carlsmith","year":"2023","unstructured":"Carlsmith J (2023) Existential risk from power-seeking AI. In: Barrett J, Greaves H, Thorstad D (eds) Essays on longtermism. Oxford University Press"},{"key":"2572_CR7","unstructured":"Cho J, Puspitasari FD, Zheng S, Zheng J, Lee LH, Kim TH, Hong CS, Zhang C (2024). Sora as an agi world model? a complete survey on text-to-video generation. arXiv preprint. https:\/\/arxiv.org\/abs\/2403.05131"},{"key":"2572_CR8","unstructured":"Cotra A (2021). Why AI alignment could be hard with modern deep learning. Cold Takes. https:\/\/www.cold-takes.com\/why-ai-alignment-could-be-hard-with-modern-deep-learning\/. Accessed 10 Dec 2024"},{"key":"2572_CR9","doi-asserted-by":"publisher","DOI":"10.1145\/3746449","author":"J Ding","year":"2024","unstructured":"Ding J, Zhang Y, Shang Y, Zhang Y, Zong Z, Feng J, Yuan Y, Su H, Li N, Sukiennik N, Xu F (2024) Understanding world or predicting future? a comprehensive survey of world models. ACM Comput Surv. https:\/\/doi.org\/10.1145\/3746449","journal-title":"ACM Comput Surv"},{"key":"2572_CR10","doi-asserted-by":"publisher","first-page":"1195","DOI":"10.1007\/s00146-024-01930-2","volume":"40","author":"L Dung","year":"2024","unstructured":"Dung L (2024) The argument for near-term human disempowerment through AI. AI Soc 40:1195\u20131208","journal-title":"AI Soc"},{"key":"2572_CR12","doi-asserted-by":"crossref","unstructured":"Field S (2025). Why do experts disagree on existential risk and P (doom)? A survey of AI experts. arXiv preprint. https:\/\/arxiv.org\/abs\/2502.14870","DOI":"10.1007\/s43681-025-00762-0"},{"key":"2572_CR13","unstructured":"Fitzgerald M, Boddy A, Baum SD (2020). Survey of artificial general intelligence projects for ethics, risk, and policy: technical report 20\u20131. Global Catastrophic Risk Insitute."},{"key":"2572_CR14","doi-asserted-by":"publisher","first-page":"573","DOI":"10.1016\/j.neunet.2021.09.011","volume":"144","author":"K Friston","year":"2021","unstructured":"Friston K, Moran RJ, Nagai Y, Taniguchi T, Gomi H, Tenenbaum J (2021) World model learning and inference. Neural Netw 144:573\u2013590. https:\/\/doi.org\/10.1016\/j.neunet.2021.09.011","journal-title":"Neural Netw"},{"issue":"3","key":"2572_CR15","doi-asserted-by":"publisher","first-page":"411","DOI":"10.1007\/s11023-020-09539-2","volume":"30","author":"I Gabriel","year":"2020","unstructured":"Gabriel I (2020) Artificial intelligence, values, and alignment. Minds Mach 30(3):411\u2013437. https:\/\/doi.org\/10.1007\/s11023-020-09539-2","journal-title":"Minds Mach"},{"key":"2572_CR16","doi-asserted-by":"publisher","DOI":"10.1007\/s11098-024-02129-3","author":"JD Gallow","year":"2024","unstructured":"Gallow JD (2024) Instrumental divergence. Philos Stud. https:\/\/doi.org\/10.1007\/s11098-024-02129-3","journal-title":"Philos Stud"},{"key":"2572_CR17","doi-asserted-by":"publisher","DOI":"10.4324\/9781315802725","volume-title":"Mental models","author":"D Gentner","year":"2014","unstructured":"Gentner D, Stevens AL (2014) Mental models. Psychology Press"},{"issue":"2","key":"2572_CR18","doi-asserted-by":"publisher","first-page":"55","DOI":"10.55613\/jeet.v25i2.48","volume":"25","author":"B Goertzel","year":"2015","unstructured":"Goertzel B (2015) Superintelligence: fears, promises and potentials: reflections on Bostrom\u2019s superintelligence, Yudkowsky\u2019s from AI to Zombies, and Weaver and Veitas\u2019s \u201copen-ended Intelligence.\u201d J Ethic Emerg Technol 25(2):55\u201387","journal-title":"J Ethic Emerg Technol"},{"key":"2572_CR19","doi-asserted-by":"publisher","DOI":"10.1007\/s11098-024-02099-6","author":"S Goldstein","year":"2024","unstructured":"Goldstein S, Robinson P (2024) Shutdown-seeking AI. Philos Stud. https:\/\/doi.org\/10.1007\/s11098-024-02099-6","journal-title":"Philos Stud"},{"key":"2572_CR21","unstructured":"Ha D, Schmidhuber J (2018). World models. arXiv preprint. https:\/\/arxiv.org\/abs\/1803.10122"},{"key":"2572_CR22","unstructured":"Hadshar R (2023). A Review of the Evidence for Existential Risk from AI via Misaligned Power-Seeking. arXiv preprint. https:\/\/arxiv.org\/abs\/2310.18244"},{"key":"2572_CR23","unstructured":"Hafner D, Lillicrap T, Ba J, Norouzi M (2019) Dream to control: Learning behaviors by latent imagination. arXiv preprint. https:\/\/arxiv.org\/abs\/1912.01603"},{"key":"2572_CR24","doi-asserted-by":"publisher","DOI":"10.1038\/s41586-025-08744-2","author":"D Hafner","year":"2025","unstructured":"Hafner D, Pasukonis J, Ba J, Lillicrap T (2025) Mastering diverse control tasks through world models. Nature. https:\/\/doi.org\/10.1038\/s41586-025-08744-2","journal-title":"Nature"},{"key":"2572_CR25","doi-asserted-by":"crossref","unstructured":"Hao S, Gu Y, Ma H, Hong JJ, Wang Z, Wang DZ, Hu Z (2023). Reasoning with language model is planning with world model. arXiv preprint. https:\/\/arxiv.org\/abs\/2305.14992","DOI":"10.18653\/v1\/2023.emnlp-main.507"},{"key":"2572_CR26","unstructured":"Hendrycks D, Mazeika M (2022). X-risk analysis for AI research. arXiv preprint. https:\/\/arxiv.org\/abs\/2206.05862"},{"issue":"1","key":"2572_CR27","first-page":"1","volume":"35","author":"DA Herrmann","year":"2025","unstructured":"Herrmann DA, Levinstein BA (2025) Standards for belief representations in LLMs. Minds Mach 35(1):1\u201325","journal-title":"Minds Mach"},{"key":"2572_CR56","unstructured":"Ji J, Qiu T, Chen B, Zhang B, Lou H, Wang K, Duan Y, He Z, Zhou J, Zhang Z, Zeng F (2023). AI alignment: A comprehensive survey. arXiv preprint. https:\/\/arxiv.org\/abs\/2310.19852"},{"key":"2572_CR28","volume-title":"Thinking, fast and slow","author":"D Kahneman","year":"2011","unstructured":"Kahneman D (2011) Thinking, fast and slow. Macmillan"},{"key":"2572_CR29","doi-asserted-by":"publisher","DOI":"10.1613\/jair.1.12481","author":"TW Kim","year":"2021","unstructured":"Kim TW, Hooker J, Donaldson T (2021) Taking principles seriously: a hybrid approach to value alignment in artificial intelligence. J Artif Intell Res. https:\/\/doi.org\/10.1613\/jair.1.12481","journal-title":"J Artif Intell Res"},{"issue":"1","key":"2572_CR30","first-page":"1","volume":"62","author":"Y LeCun","year":"2022","unstructured":"LeCun Y (2022) A path towards autonomous machine intelligence (Version 0.9.2). OpenReview 62(1):1\u201362","journal-title":"OpenReview"},{"key":"2572_CR31","doi-asserted-by":"publisher","first-page":"391","DOI":"10.1007\/s11023-007-9079-x","volume":"17","author":"S Legg","year":"2007","unstructured":"Legg S, Hutter M (2007) Universal intelligence: a definition of machine intelligence. Minds Mach 17:391\u2013444. https:\/\/doi.org\/10.1007\/s11023-007-9079-x","journal-title":"Minds Mach"},{"key":"2572_CR32","doi-asserted-by":"crossref","unstructured":"Li H, Chong YQ, Stepputtis S, Campbell J, Hughes D, Lewis M, Sycara K (2023). Theory of mind for multi-agent collaboration via large language models. arXiv preprint. https:\/\/arxiv.org\/abs\/2310.10701","DOI":"10.18653\/v1\/2023.emnlp-main.13"},{"key":"2572_CR33","unstructured":"Liu B, Li X, Zhang J, Wang J, He T, Hong S, Liu H, Zhang S, Song K, Zhu K, Cheng Y (2025). Advances and challenges in foundation agents: From brain-inspired intelligence to evolutionary, collaborative, and safe systems. arXiv preprint. https:\/\/arxiv.org\/abs\/2504.01990"},{"key":"2572_CR34","unstructured":"Manvi R, Khanna S, Mai G, Burke M, Lobell D, Ermon S (2024). Geollm: Extracting geospatial knowledge from large language models. arXiv preprint. https:\/\/arxiv.org\/2310.06213"},{"issue":"5","key":"2572_CR35","doi-asserted-by":"publisher","first-page":"649","DOI":"10.1080\/0952813X.2021.1964003","volume":"35","author":"S McLean","year":"2023","unstructured":"McLean S, Read GJ, Thompson J, Baber C, Stanton NA, Salmon PM (2023) The risks associated with artificial general intelligence: a systematic review. J Exp Theor Artif Intell 35(5):649\u2013663. https:\/\/doi.org\/10.1080\/0952813X.2021.1964003","journal-title":"J Exp Theor Artif Intell"},{"issue":"6689","key":"2572_CR36","doi-asserted-by":"publisher","first-page":"eado7069","DOI":"10.1126\/science.ado7069","volume":"383","author":"M Mitchell","year":"2024","unstructured":"Mitchell M (2024) Debates on the nature of artificial general intelligence. Science 383(6689):eado7069","journal-title":"Science"},{"key":"2572_CR37","unstructured":"Morris MR, Sohl-Dickstein J, Fiedel N, Warkentin T, Dafoe A, Faust A, Farabet C, Legg S (2023). Levels of AGI for operationalizing progress on the path to AGI. arXiv preprint. https:\/\/arxiv.org\/abs\/2311.02462"},{"key":"2572_CR38","unstructured":"Ngo R, Chan L, Mindermann S (2023). The alignment problem from a deep learning perspective: a position paper. In: The Twelfth International Conference on Learning Representations."},{"key":"2572_CR39","first-page":"483","volume-title":"Artificial general intelligence","author":"SM Omohundro","year":"2008","unstructured":"Omohundro SM (2008) The basic AI drives. In: Wang P, Goertzel B, Franklin S (eds) Artificial general intelligence. IOS Press, pp 483\u2013492"},{"key":"2572_CR41","doi-asserted-by":"publisher","DOI":"10.1007\/s43681-025-00664-1","author":"E Riesen","year":"2025","unstructured":"Riesen E, Boespflug M (2025) Aligning with ideal values: a proposal for anchoring AI in moral expertise. AI Ethics. https:\/\/doi.org\/10.1007\/s43681-025-00664-1","journal-title":"AI Ethics"},{"key":"2572_CR42","volume-title":"Human compatible: artificial intelligence and the problem of control","author":"S Russell","year":"2019","unstructured":"Russell S (2019) Human compatible: artificial intelligence and the problem of control. Penguin"},{"key":"2572_CR43","doi-asserted-by":"publisher","first-page":"1253049","DOI":"10.3389\/frobt.2023.1253049","volume":"10","author":"R Sakagami","year":"2023","unstructured":"Sakagami R, Lay FS, D\u00f6mel A, Schuster MJ, Albu-Sch\u00e4ffer A, Stulp F (2023) Robotic world models\u2014conceptualization, review, and engineering best practices. Front Robot AI 10:1253049","journal-title":"Front Robot AI"},{"key":"2572_CR44","unstructured":"Shah R, Irpan A, Turner AM, Wang A, Conmy A, Lindner D, Brown-Cohen J, Ho L, Nanda N, Popa RA, Jain R (2025). An approach to technical agi safety and security. arXiv preprint. https:\/\/arxiv.org\/abs\/2504.01849"},{"key":"2572_CR46","unstructured":"Sucholutsky I, Muttenthaler L, Weller A, Peng A, Bobu A, Kim B, Love BC, Grant E, Groen I, Achterberg J, Tenenbaum JB (2023). Getting aligned on representational alignment. arXiv preprint. https:\/\/arxiv.org\/abs\/2310.13018"},{"issue":"1","key":"2572_CR47","doi-asserted-by":"publisher","first-page":"81","DOI":"10.1177\/1059712320962","volume":"30","author":"J Tani","year":"2022","unstructured":"Tani J, White J (2022) Cognitive neurorobotics and self in the shared world, a focused review of ongoing research. Adapt Behav 30(1):81\u2013100. https:\/\/doi.org\/10.1177\/1059712320962","journal-title":"Adapt Behav"},{"issue":"13","key":"2572_CR48","doi-asserted-by":"publisher","first-page":"780","DOI":"10.1016\/j.neunet.2021.09.011","volume":"37","author":"T Taniguchi","year":"2023","unstructured":"Taniguchi T, Murata S, Suzuki M, Ognibene D, Lanillos P, Ugur E, Pezzulo G (2023) World models and predictive coding for cognitive and developmental robotics: frontiers and challenges. Adv Robot 37(13):780\u2013806. https:\/\/doi.org\/10.1016\/j.neunet.2021.09.011","journal-title":"Adv Robot"},{"key":"2572_CR49","doi-asserted-by":"publisher","DOI":"10.1007\/s11098-024-02153-3","author":"E Thornley","year":"2024","unstructured":"Thornley E (2024) The shutdown problem: an AI engineering puzzle for decision theorists. Philos Stud. https:\/\/doi.org\/10.1007\/s11098-024-02153-3","journal-title":"Philos Stud"},{"key":"2572_CR50","unstructured":"Turner AM, Smith L, Shah R, Critch A, Tadepalli P (2019). Optimal policies tend to seek power. arXiv preprint. https:\/\/arxiv.org\/abs\/1912.01683"},{"key":"2572_CR51","unstructured":"Varshney KR, Ashktorab Z, Bouneffouf D, Riemer M, Weisz JD (2025). Scopes of alignment. arXiv preprint. https:\/\/arxiv.org\/abs\/2501.12405"},{"issue":"1","key":"2572_CR52","doi-asserted-by":"publisher","first-page":"012017","DOI":"10.1088\/1757-899X\/1321\/1\/012017","volume":"1321","author":"J White","year":"2024","unstructured":"White J (2024) Variable value alignment by design; averting risks with robot religion. IOP Conf Ser: Mater Sci Eng 1321(1):012017. https:\/\/doi.org\/10.1088\/1757-899X\/1321\/1\/012017","journal-title":"IOP Conf Ser: Mater Sci Eng"},{"key":"2572_CR53","unstructured":"Xing E, Deng M, Hou J, Hu Z (2025). Critiques of World Models. arXiv preprint. https:\/\/arxiv.org\/abs\/2507.05169"},{"key":"2572_CR54","unstructured":"Zhang W, Han J, Xu Z, Ni H, Liu H, Xiong H (2025). Towards urban general intelligence: a review and outlook of urban foundation models. arXiv preprint. https:\/\/arxiv.org\/abs\/2402.01749"}],"container-title":["AI &amp; SOCIETY"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s00146-025-02572-8.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s00146-025-02572-8","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s00146-025-02572-8.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,2,23]],"date-time":"2026-02-23T04:23:31Z","timestamp":1771820611000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s00146-025-02572-8"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,8,21]]},"references-count":51,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2026,2]]}},"alternative-id":["2572"],"URL":"https:\/\/doi.org\/10.1007\/s00146-025-02572-8","relation":{},"ISSN":["0951-5666","1435-5655"],"issn-type":[{"value":"0951-5666","type":"print"},{"value":"1435-5655","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,8,21]]},"assertion":[{"value":"14 January 2025","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"12 August 2025","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"21 August 2025","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare no competing interests.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}}]}}