{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T15:10:01Z","timestamp":1750173001206,"version":"3.41.0"},"reference-count":34,"publisher":"Springer Science and Business Media LLC","issue":"18","license":[{"start":{"date-parts":[[2025,4,2]],"date-time":"2025-04-02T00:00:00Z","timestamp":1743552000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2025,4,2]],"date-time":"2025-04-02T00:00:00Z","timestamp":1743552000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100001778","name":"Deakin University","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100001778","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Neural Comput &amp; Applic"],"published-print":{"date-parts":[[2025,6]]},"abstract":"<jats:title>Abstract<\/jats:title>\n          <jats:p>Large language models (LLMs) have recently demonstrated their impressive ability to provide context-aware responses via text. This ability could potentially be used to predict plausible solutions in sequential decision making tasks pertaining to pattern completion. For example, by observing a partial stack of cubes, LLMs can predict the correct sequence in which the remaining cubes should be stacked by extrapolating the observed patterns (e.g., cube sizes, colors or other attributes) in the partial stack. In this work, we introduce LaGR (<jats:italic>language-guided reinforcement learning<\/jats:italic>), which uses this predictive ability of LLMs to propose solutions to tasks that have been partially completed by a primary reinforcement learning (RL) agent, in order to subsequently guide the latter\u2019s training. However, as RL training is generally not sample-efficient, deploying this approach would inherently imply that the LLM be repeatedly queried for solutions; a process that can be expensive and infeasible. To address this issue, we introduce SEQ (<jats:italic>sample-efficient querying<\/jats:italic>), where we simultaneously train a secondary RL agent to decide when the LLM should be queried for solutions. Specifically, we use the quality of the solutions emanating from the LLM as the reward to train this agent. We show that our proposed framework LaGR-SEQ enables more efficient primary RL training, while simultaneously minimizing the number of queries to the LLM. We demonstrate our approach on a series of tasks and highlight the advantages of our approach, along with its limitations and potential future research directions.<\/jats:p>","DOI":"10.1007\/s00521-025-11156-y","type":"journal-article","created":{"date-parts":[[2025,4,4]],"date-time":"2025-04-04T13:53:44Z","timestamp":1743774824000},"page":"12447-12470","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["LaGR-SEQ: Language-guided reinforcement learning with sample-efficient querying"],"prefix":"10.1007","volume":"37","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-8918-3314","authenticated-orcid":false,"given":"Thommen Karimpanal","family":"George","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Buddhika Laknath","family":"Semage","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Santu","family":"Rana","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hung","family":"Le","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Truyen","family":"Tran","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Sunil","family":"Gupta","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Svetha","family":"Venkatesh","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2025,4,2]]},"reference":[{"key":"11156_CR1","unstructured":"Sutton R.S, Barto A.G, et al. (1998) Introduction to Reinforcement Learning vol. 135"},{"key":"11156_CR2","doi-asserted-by":"crossref","unstructured":"Mnih V, Kavukcuoglu K, Silver D, Rusu A.A, Veness J, Bellemare M.G, Graves A, Riedmiller M, Fidjeland A.K, Ostrovski G, et al. (2015) Human-level control through deep reinforcement learning. nature 518(7540), 529\u2013533","DOI":"10.1038\/nature14236"},{"key":"11156_CR3","unstructured":"Mnih V, Kavukcuoglu K, Silver D, Graves A, Antonoglou I, Wierstra D, Riedmiller M (2013) Playing atari with deep reinforcement learning. arXiv preprint arXiv:1312.5602"},{"key":"11156_CR4","doi-asserted-by":"crossref","unstructured":"Singh B, Kumar R, Singh V.P (2022) Reinforcement learning in robotic applications: a comprehensive survey. Artificial Intelligence Review, 1\u201346","DOI":"10.1007\/s10462-021-09997-9"},{"issue":"11","key":"11156_CR5","doi-asserted-by":"publisher","first-page":"1238","DOI":"10.1177\/0278364913495721","volume":"32","author":"J Kober","year":"2013","unstructured":"Kober J, Bagnell JA, Peters J (2013) Reinforcement learning in robotics: A survey. Int J Robot Res 32(11):1238\u20131274","journal-title":"Int J Robot Res"},{"issue":"2","key":"11156_CR6","doi-asserted-by":"publisher","first-page":"111","DOI":"10.1177\/1059712318818568","volume":"27","author":"T George Karimpanal","year":"2019","unstructured":"George Karimpanal T, Bouffanais R (2019) Self-organizing maps for storage and transfer of knowledge in reinforcement learning. Adapt Behav 27(2):111\u2013126","journal-title":"Adapt Behav"},{"key":"11156_CR7","first-page":"1877","volume":"33","author":"T Brown","year":"2020","unstructured":"Brown T, Mann B, Ryder N, Subbiah M, Kaplan JD, Dhariwal P, Neelakantan A, Shyam P, Sastry G, Askell A et al (2020) Language models are few-shot learners. Adv Neural Inf Process Syst 33:1877\u20131901","journal-title":"Adv Neural Inf Process Syst"},{"key":"11156_CR8","unstructured":"Touvron H, Martin L, Stone K, Albert P, Almahairi A, Babaei Y, Bashlykov N, Batra S, Bhargava P, Bhosale S, et al. (2023) Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288"},{"key":"11156_CR9","unstructured":"Llama\u00a0team A (2024) The llama3 herd of models"},{"key":"11156_CR10","unstructured":"Deepmind: Gemini: A family of highly capable multimodal models. Deepmind blog (2023)"},{"key":"11156_CR11","unstructured":"Anthropic: Introducing the next generation of claude. Anthropic blog (2024)"},{"key":"11156_CR12","unstructured":"Du Y, Watkins O, Wang Z, Colas C, Darrell T, Abbeel P, Gupta A, Andreas J (2023) Guiding pretraining in reinforcement learning with large language models. arXiv preprint arXiv:2302.06692"},{"key":"11156_CR13","unstructured":"Rivlin O, Hazan T, Karpas E (2020) Generalized planning with deep reinforcement learning. arXiv preprint arXiv:2005.02305"},{"key":"11156_CR14","unstructured":"Mirchandani S, Xia F, Florence P, Driess D, Arenas M.G, Rao K, Sadigh D, Zeng A, et al. (2023) Large language models as general pattern machines. In: 7th Annual Conference on Robot Learning"},{"key":"11156_CR15","unstructured":"Christiano P.F, Leike J, Brown T, Martic M, Legg S, Amodei D (2017) Deep reinforcement learning from human preferences. Advances in neural information processing systems 30"},{"key":"11156_CR16","doi-asserted-by":"crossref","unstructured":"Yuan A, Coenen A, Reif E, Ippolito D (2022) Wordcraft: story writing with large language models. In: 27th International Conference on Intelligent User Interfaces, pp. 841\u2013852","DOI":"10.1145\/3490099.3511105"},{"issue":"1","key":"11156_CR17","doi-asserted-by":"publisher","first-page":"20","DOI":"10.1186\/s13040-023-00339-9","volume":"16","author":"JG Meyer","year":"2023","unstructured":"Meyer JG, Urbanowicz RJ, Martin PC, O\u2019Connor K, Li R, Peng P-C, Bright TJ, Tatonetti N, Won KJ, Gonzalez-Hernandez G et al (2023) Chatgpt and large language models in academia: opportunities and challenges. BioData Mining 16(1):20","journal-title":"BioData Mining"},{"key":"11156_CR18","doi-asserted-by":"publisher","DOI":"10.1016\/j.lindif.2023.102274","volume":"103","author":"E Kasneci","year":"2023","unstructured":"Kasneci E, Se\u00dfler K, K\u00fcchemann S, Bannert M, Dementieva D, Fischer F, Gasser U, Groh G, G\u00fcnnemann S, H\u00fcllermeier E et al (2023) Chatgpt for good? on opportunities and challenges of large language models for education. Learn Individ Differ 103:102274","journal-title":"Learn Individ Differ"},{"key":"11156_CR19","unstructured":"Ahn M, Brohan A, Brown N, Chebotar Y, Cortes O, David B, Finn C, Fu C, Gopalakrishnan K, Hausman K, et al. (2022) Do as i can, not as i say: Grounding language in robotic affordances. arXiv preprint arXiv:2204.01691"},{"key":"11156_CR20","unstructured":"Driess D, Xia F, Sajjadi M.S, Lynch C, Chowdhery A, Ichter B, Wahid A, Tompson J, Vuong, Q, Yu T, et al. (2023) Palm-e: An embodied multimodal language model. arXiv preprint arXiv:2303.03378"},{"key":"11156_CR21","unstructured":"Huang W, Abbeel P, Pathak D, Mordatch I (2022) Language models as zero-shot planners: Extracting actionable knowledge for embodied agents. In: International Conference on Machine Learning, pp. 9118\u20139147 . PMLR"},{"key":"11156_CR22","first-page":"31199","volume":"35","author":"S Li","year":"2022","unstructured":"Li S, Puig X, Paxton C, Du Y, Wang C, Fan L, Chen T, Huang D-A, Aky\u00fcrek E, Anandkumar A et al (2022) Pre-trained language models for interactive decision-making. Adv Neural Inf Process Syst 35:31199\u201331212","journal-title":"Adv Neural Inf Process Syst"},{"key":"11156_CR23","doi-asserted-by":"crossref","unstructured":"Shridhar M, Thomason J, Gordon D, Bisk Y, Han W, Mottaghi R, Zettlemoyer L, Fox D (2020) Alfred: A benchmark for interpreting grounded instructions for everyday tasks. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pp. 10740\u201310749","DOI":"10.1109\/CVPR42600.2020.01075"},{"key":"11156_CR24","doi-asserted-by":"crossref","unstructured":"Lynch C, Sermanet P (2020) Language conditioned imitation learning over unstructured data. arXiv preprint arXiv:2005.07648","DOI":"10.15607\/RSS.2021.XVII.047"},{"key":"11156_CR25","doi-asserted-by":"crossref","unstructured":"Anderson P, Wu Q, Teney D, Bruce J, Johnson M, S\u00fcnderhauf N, Reid I, Gould S, Van Den\u00a0Hengel A (2018) Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 3674\u20133683","DOI":"10.1109\/CVPR.2018.00387"},{"key":"11156_CR26","unstructured":"Yuan H, Zhang C, Wang H, Xie F, Cai P, Dong H, Lu Z (2023) Plan4mc: Skill reinforcement learning and planning for open-world minecraft tasks. arXiv preprint arXiv:2303.16563"},{"key":"11156_CR27","unstructured":"Nottingham K, Ammanabrolu P, Suhr A, Choi Y, Hajishirzi H, Singh S, Fox R (2023) Do embodied agents dream of pixelated sheep? Embodied decision making using language guided world modelling. arXiv preprint arXiv:2301.12050"},{"key":"11156_CR28","unstructured":"Lin J, Du Y, Watkins O, Hafner D, Abbeel P, Klein D, Dragan A (2023) Learning to model the world with language. arXiv preprint arXiv:2308.01399"},{"key":"11156_CR29","unstructured":"Choi K, Cundy C, Srivastava S, Ermon S (2022) Lmpriors: Pre-trained language models as task-specific priors. arXiv preprint arXiv:2210.12530"},{"key":"11156_CR30","doi-asserted-by":"crossref","unstructured":"Kant Y, Ramachandran A, Yenamandra S, Gilitschenski I, Batra D, Szot A, Agrawal H (2022) Housekeep: Tidying virtual households using commonsense reasoning. In: European Conference on Computer Vision, pp. 355\u2013373. Springer","DOI":"10.1007\/978-3-031-19842-7_21"},{"key":"11156_CR31","unstructured":"Kwon M, Xie SM, Bullard K, Sadigh D (2023) Reward design with language models. arXiv preprint arXiv:2303.00001"},{"key":"11156_CR32","unstructured":"Puterman ML (2014) Markov Decision Processes: Discrete Stochastic Dynamic Programming"},{"key":"11156_CR33","doi-asserted-by":"crossref","unstructured":"Slivkins A et al. (2019) Introduction to multi-armed bandits. Foundations and Trends\u00ae in Machine Learning 12(1-2), 1\u2013286","DOI":"10.1561\/2200000068"},{"key":"11156_CR34","unstructured":"Karimpanal TG, Chamanbaz M, Li W, Jeruzalski T, Gupta A, Wilhelm E (2017) Adapting low-cost platforms for robotics research. arXiv preprint arXiv:1705.07231"}],"container-title":["Neural Computing and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s00521-025-11156-y.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s00521-025-11156-y\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s00521-025-11156-y.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T14:48:58Z","timestamp":1750171738000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s00521-025-11156-y"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,4,2]]},"references-count":34,"journal-issue":{"issue":"18","published-print":{"date-parts":[[2025,6]]}},"alternative-id":["11156"],"URL":"https:\/\/doi.org\/10.1007\/s00521-025-11156-y","relation":{},"ISSN":["0941-0643","1433-3058"],"issn-type":[{"type":"print","value":"0941-0643"},{"type":"electronic","value":"1433-3058"}],"subject":[],"published":{"date-parts":[[2025,4,2]]},"assertion":[{"value":"20 December 2023","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"26 February 2025","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"2 April 2025","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}]}}