{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,18]],"date-time":"2026-08-18T01:56:03Z","timestamp":1787018163650,"version":"build-2736575974"},"reference-count":47,"publisher":"MIT Press","license":[{"start":{"date-parts":[[2024,8,1]],"date-time":"2024-08-01T00:00:00Z","timestamp":1722470400000},"content-version":"vor","delay-in-days":213,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["direct.mit.edu"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2024,8,1]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>We describe a class of tasks called decision-oriented dialogues, in which AI assistants such as large language models (LMs) must collaborate with one or more humans via natural language to help them make complex decisions. We formalize three domains in which users face everyday decisions: (1) choosing an assignment of reviewers to conference papers, (2) planning a multi-step itinerary in a city, and (3) negotiating travel plans for a group of friends. In each of these settings, AI assistants and users have disparate abilities that they must combine to arrive at the best decision: Assistants can access and process large amounts of information, while users have preferences and constraints external to the system. For each task, we build a dialogue environment where agents receive a reward based on the quality of the final decision they reach. We evaluate LMs in self-play and in collaboration with humans and find that they fall short compared to human assistants, achieving much lower rewards despite engaging in longer dialogues. We highlight a number of challenges models face in decision-oriented dialogues, ranging from goal-directed behavior to reasoning and optimization, and release our environments as a testbed for future work.<\/jats:p>","DOI":"10.1162\/tacl_a_00679","type":"journal-article","created":{"date-parts":[[2024,8,1]],"date-time":"2024-08-01T16:13:33Z","timestamp":1722528813000},"page":"892-911","update-policy":"https:\/\/doi.org\/10.1162\/mitpressjournals.corrections.policy","source":"Crossref","is-referenced-by-count":32,"title":["Decision-Oriented Dialogue for Human-AI Collaboration"],"prefix":"10.1162","volume":"12","author":[{"given":"Jessy","family":"Lin","sequence":"first","affiliation":[{"name":"UC Berkeley, USA. jessy_lin@berkeley.edu"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Nicholas","family":"Tomlin","sequence":"additional","affiliation":[{"name":"UC Berkeley, USA. nicholas_tomlin@berkeley.edu"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jacob","family":"Andreas","sequence":"additional","affiliation":[{"name":"Microsoft Semantic Machines, USA. jaandrea@microsoft.com"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jason","family":"Eisner","sequence":"additional","affiliation":[{"name":"Microsoft Semantic Machines, USA. jason.eisner@microsoft.com"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"281","published-online":{"date-parts":[[2024,8,1]]},"reference":[{"key":"2024080120131506800_bib1","article-title":"Human-machine collaborative planning","volume-title":"Proceedings of the 2002 Workshop on Knowledge and Reasoning in Practical Dialogue Systems","author":"Allen","year":"2002"},{"issue":"1","key":"2024080120131506800_bib2","doi-asserted-by":"publisher","first-page":"7","DOI":"10.1080\/09528139508953799","article-title":"The TRAINS project: A case study in building a conversational planning agent","volume":"7","author":"Allen","year":"1995","journal-title":"Journal of Experimental & Theoretical Artificial Intelligence"},{"issue":"6624","key":"2024080120131506800_bib3","doi-asserted-by":"publisher","first-page":"1067","DOI":"10.1126\/science.ade9097","article-title":"Human-level play in the game of Diplomacy by combining language models with strategic reasoning","volume":"378","author":"Bakhtin","year":"2022","journal-title":"Science"},{"key":"2024080120131506800_bib4","doi-asserted-by":"publisher","first-page":"103216","DOI":"10.1016\/j.artint.2019.103216","article-title":"The Hanabi challenge: A new frontier for AI research","volume":"280","author":"Bard","year":"2020","journal-title":"Artificial Intelligence"},{"key":"2024080120131506800_bib5","first-page":"32","article-title":"The complexity of decentralized control of Markov decision processes","volume-title":"Proceedings of the Sixteenth Conference on Uncertainty in Artificial Intelligence (UAI)","author":"Bernstein","year":"2000"},{"key":"2024080120131506800_bib6","unstructured":"Greg\n              Brockman\n            , VickiCheung, LudwigPettersson, JonasSchneider, JohnSchulman, JieTang, and WojciechZaremba. 2016. OpenAI Gym."},{"key":"2024080120131506800_bib7","first-page":"1877","article-title":"Language models are few-shot learners","volume-title":"Advances in Neural Information Processing Systems (NeurIPS)","author":"Brown","year":"2020"},{"key":"2024080120131506800_bib8","doi-asserted-by":"publisher","first-page":"5016","DOI":"10.18653\/v1\/D18-1547","article-title":"MultiWOZ\u2014A large-scale multi-domain Wizard-of-Oz dataset for task-oriented dialogue modelling","volume-title":"Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing (EMNLP)","author":"Budzianowski","year":"2018"},{"issue":"11","key":"2024080120131506800_bib9","doi-asserted-by":"publisher","first-page":"925","DOI":"10.1016\/j.artint.2006.05.003","article-title":"Generating and evaluating evaluative arguments","volume":"170","author":"Carenini","year":"2006","journal-title":"Artificial Intelligence"},{"key":"2024080120131506800_bib10","article-title":"On the utility of learning about humans for human-AI coordination","volume-title":"Advances in Neural Information Processing Systems (NeurIPS)","author":"Carroll","year":"2019"},{"key":"2024080120131506800_bib11","article-title":"The Toronto paper matching system: An automated paper-reviewer assignment system","volume-title":"Proceedings of the ICML Workshop on Peer Reviewing and Publishing Models (PEER)","author":"Charlin","year":"2013"},{"issue":"240","key":"2024080120131506800_bib12","first-page":"1","article-title":"PaLM: Scaling language modeling with pathways","volume":"24","author":"Chowdhery","year":"2023","journal-title":"Journal of Machine Learning Research"},{"key":"2024080120131506800_bib13","article-title":"Open problems in cooperative AI","volume":"arXiv:2012.08630","author":"Dafoe","year":"2020","journal-title":"Computing Research Repository (CoRR)"},{"key":"2024080120131506800_bib14","doi-asserted-by":"publisher","first-page":"12619","DOI":"10.18653\/v1\/2023.findings-emnlp.840","article-title":"Pragmatics in language grounding: Phenomena, tasks, and modeling approaches","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2023","author":"Fried","year":"2023"},{"key":"2024080120131506800_bib15","article-title":"Cooperative inverse reinforcement learning","volume-title":"Advances in Neural Information Processing Systems (NeurIPS)","author":"Hadfield-Menell","year":"2016"},{"key":"2024080120131506800_bib16","doi-asserted-by":"publisher","first-page":"1766","DOI":"10.18653\/v1\/P17-1162","article-title":"Learning symmetric collaborative dialogue agents with dynamic knowledge graph embeddings","volume-title":"Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (ACL)","author":"He","year":"2017"},{"key":"2024080120131506800_bib17","doi-asserted-by":"publisher","first-page":"2333","DOI":"10.18653\/v1\/D18-1256","article-title":"Decoupling strategy and generation in negotiation dialogues","volume-title":"Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing (EMNLP)","author":"He","year":"2018"},{"key":"2024080120131506800_bib18","first-page":"805","article-title":"Fictitious self-play in extensive-form games","volume-title":"Proceedings of the 32nd International Conference on Machine Learning (ICML)","author":"Heinrich","year":"2015"},{"key":"2024080120131506800_bib19","article-title":"Measuring mathematical problem solving with the MATH dataset","volume-title":"Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks","author":"Hendrycks","year":"2021"},{"key":"2024080120131506800_bib20","doi-asserted-by":"publisher","first-page":"159","DOI":"10.1145\/302979.303030","article-title":"Principles of mixed-initiative user interfaces","volume-title":"Proceedings of the SIGCHI Conference on Human Factors in Computing Systems","author":"Horvitz","year":"1999"},{"issue":"6443","key":"2024080120131506800_bib21","doi-asserted-by":"publisher","first-page":"859","DOI":"10.1126\/science.aau6249","article-title":"Human-level performance in 3D multiplayer games with population-based reinforcement learning","volume":"364","author":"Jaderberg","year":"2019","journal-title":"Science"},{"key":"2024080120131506800_bib22","first-page":"4415","article-title":"Reward-rational (implicit) choice: A unifying formalism for reward learning","volume-title":"Advances in Neural Information Processing Systems (NeurIPS)","author":"Jeon","year":"2020"},{"issue":"3","key":"2024080120131506800_bib23","doi-asserted-by":"publisher","first-page":"192","DOI":"10.1016\/j.cogsys.2007.06.002","article-title":"Global vs. local information processing in visual\/spatial problem solving: The case of traveling salesman problem","volume":"8","author":"Kong","year":"2007","journal-title":"Cognitive Systems Research"},{"key":"2024080120131506800_bib24","doi-asserted-by":"publisher","first-page":"2443","DOI":"10.18653\/v1\/D17-1259","article-title":"Deal or no deal? End-to-end learning of negotiation dialogues","volume-title":"Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing (EMNLP)","author":"Lewis","year":"2017"},{"key":"2024080120131506800_bib25","doi-asserted-by":"publisher","first-page":"1813","DOI":"10.18653\/v1\/2021.acl-long.143","article-title":"Implicit representations of meaning in neural language models","volume-title":"Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers) (ACL-IJCNLP)","author":"Li","year":"2021"},{"key":"2024080120131506800_bib26","doi-asserted-by":"publisher","first-page":"1192","DOI":"10.18653\/v1\/D16-1127","article-title":"Deep reinforcement learning for dialogue generation","volume-title":"Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing (EMNLP)","author":"Li","year":"2016"},{"key":"2024080120131506800_bib27","doi-asserted-by":"publisher","first-page":"8546","DOI":"10.18653\/v1\/2022.acl-long.585","article-title":"Inferring rewards from language in context","volume-title":"Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (ACL)","author":"Lin","year":"2022"},{"key":"2024080120131506800_bib28","first-page":"12","article-title":"What is mixed-initiative interaction?","volume-title":"Proceedings of the AAAI Spring Symposium on Computational Models for Mixed Initiative Interaction","author":"Novick","year":"1997"},{"key":"2024080120131506800_bib29","article-title":"Show your work: Scratchpads for intermediate computation with language models","volume":"arXiv:2112.00114","author":"Nye","year":"2021","journal-title":"Computing Research Repository (CoRR)"},{"key":"2024080120131506800_bib30","first-page":"27730","article-title":"Training language models to follow instructions with human feedback","volume-title":"Advances in Neural Information Processing Systems","author":"Ouyang","year":"2022"},{"issue":"1\/3","key":"2024080120131506800_bib31","doi-asserted-by":"publisher","first-page":"243","DOI":"10.1023\/A:1007843931054","article-title":"Measuring constructed preferences: Towards a building code","volume":"19","author":"Payne","year":"1999","journal-title":"Journal of Risk and Uncertainty"},{"key":"2024080120131506800_bib32","first-page":"1","article-title":"Goal-driven answers in the Cards dialogue corpus","volume-title":"Proceedings of the 30th West Coast Conference on Formal Linguistics (WCCFL)","author":"Potts","year":"2012"},{"key":"2024080120131506800_bib33","doi-asserted-by":"crossref","DOI":"10.15607\/RSS.2016.XII.029","article-title":"Planning for autonomous cars that leverage effects on human actions","volume-title":"Proceedings of Robotics: Science and Systems (RSS)","author":"Sadigh","year":"2016"},{"key":"2024080120131506800_bib34","doi-asserted-by":"publisher","first-page":"3762","DOI":"10.18653\/v1\/2022.emnlp-main.248","article-title":"Neural theory-of-mind? on the limits of social intelligence in large LMs","volume-title":"Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing","author":"Sap","year":"2022"},{"key":"2024080120131506800_bib35","first-page":"68539","article-title":"Toolformer: Language models can teach themselves to use tools","volume-title":"Advances in Neural Information Processing Systems","author":"Schick","year":"2023"},{"key":"2024080120131506800_bib36","article-title":"Grounded agreement games: Emphasizing conversational grounding in visual dialogue settings","volume":"arXiv:1908.11279","author":"Schlangen","year":"2019","journal-title":"Computing Research Repository (CoRR)"},{"key":"2024080120131506800_bib37","doi-asserted-by":"publisher","first-page":"556","DOI":"10.1162\/tacl_a_00333","article-title":"Task-oriented dialogue as dataflow synthesis","volume":"8","author":"Machines","year":"2020","journal-title":"Transactions of the Association for Computational Linguistics (TACL)"},{"issue":"3","key":"2024080120131506800_bib38","doi-asserted-by":"publisher","first-page":"339","DOI":"10.1162\/089120100561737","article-title":"Dialogue act modeling for automatic tagging and recognition of conversational speech","volume":"26","author":"Stolcke","year":"2000","journal-title":"Computational Linguistics"},{"key":"2024080120131506800_bib39","first-page":"14502","article-title":"Collaborating with humans without human data","volume-title":"Advances in Neural Information Processing Systems (NeurIPS)","author":"Strouse","year":"2021"},{"key":"2024080120131506800_bib40","doi-asserted-by":"publisher","first-page":"2119","DOI":"10.18653\/v1\/D19-1218","article-title":"Executing instructions in situated collaborative interactions","volume-title":"Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)","author":"Suhr","year":"2019"},{"key":"2024080120131506800_bib41","article-title":"LLaMA: Open and efficient foundation language models","volume":"arXiv:2302.13971","author":"Touvron","year":"2023","journal-title":"Computing Research Repository (CoRR)"},{"issue":"01","key":"2024080120131506800_bib42","doi-asserted-by":"publisher","first-page":"7120","DOI":"10.1609\/aaai.v33i01.33017120","article-title":"A natural language corpus of common grounding under continuous and partially-observable context","volume":"33","author":"Udagawa","year":"2019","journal-title":"Proceedings of the AAAI Conference on Artificial Intelligence (AAAI)"},{"key":"2024080120131506800_bib43","first-page":"1072","article-title":"Emergence of Gricean maxims from multi-agent decision theory","volume-title":"Proceedings of the 2013 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT)","author":"Vogel","year":"2013"},{"key":"2024080120131506800_bib44","volume-title":"Commitment in Dialogue: Basic Concepts of Interpersonal Reasoning","author":"Walton","year":"1995"},{"key":"2024080120131506800_bib45","first-page":"24824","article-title":"Chain-of-thought prompting elicits reasoning in large language models","volume-title":"Advances in Neural Information Processing Systems","author":"Wei","year":"2022"},{"key":"2024080120131506800_bib46","doi-asserted-by":"publisher","first-page":"3844","DOI":"10.18653\/v1\/D18-1419","article-title":"AirDialogue: An environment for goal-oriented dialogue research","volume-title":"Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing (EMNLP)","author":"Wei","year":"2018"},{"key":"2024080120131506800_bib47","article-title":"ReAct: Synergizing reasoning and acting in language models","volume-title":"The Eleventh International Conference on Learning Representations (ICLR)","author":"Yao","year":"2023"}],"container-title":["Transactions of the Association for Computational Linguistics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/direct.mit.edu\/tacl\/article-pdf\/doi\/10.1162\/tacl_a_00679\/2464086\/tacl_a_00679.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/direct.mit.edu\/tacl\/article-pdf\/doi\/10.1162\/tacl_a_00679\/2464086\/tacl_a_00679.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,8,1]],"date-time":"2024-08-01T16:13:52Z","timestamp":1722528832000},"score":1,"resource":{"primary":{"URL":"https:\/\/direct.mit.edu\/tacl\/article\/doi\/10.1162\/tacl_a_00679\/123884\/Decision-Oriented-Dialogue-for-Human-AI"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024]]},"references-count":47,"URL":"https:\/\/doi.org\/10.1162\/tacl_a_00679","relation":{},"ISSN":["2307-387X"],"issn-type":[{"value":"2307-387X","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2024]]},"published":{"date-parts":[[2024]]}}}