{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2023,5,20]],"date-time":"2023-05-20T17:40:05Z","timestamp":1684604405999},"reference-count":31,"publisher":"Cambridge University Press (CUP)","issue":"1","license":[{"start":{"date-parts":[[2009,1,1]],"date-time":"2009-01-01T00:00:00Z","timestamp":1230768000000},"content-version":"unspecified","delay-in-days":0,"URL":"https:\/\/www.cambridge.org\/core\/terms"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Nat. Lang. Eng."],"published-print":{"date-parts":[[2009,1]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Policy learning is an active topic in dialogue systems research, but it has not been explored in relation to interactive question answering (IQA). We take a first step in learning adaptive interaction policies for question answering : we address the question of how to acquire enough reliable query constraints, how many database results to present to the user and when to present them, given the competing trade-offs between the length of the answer list, the length of the interaction, the type of database and the noise in the communication channel. The operating conditions are reflected in an objective function which we use to derive a hand-coded threshold-based policy and rewards to train a reinforcement learning policy. The same objective function is used for evaluation. We show that we can learn strategies for this complex trade-off problem which perform significantly better than a variety of hand-coded policies, for a wide range of noise conditions, user types, types of DB and turn-penalties. Our policy learning framework thus covers a wide spectrum of operating conditions. The learned policies produce an average<jats:italic>relative<\/jats:italic>increase in reward of 86.78% over the hand-coded policies. In 93% of the cases the learned policies perform significantly better than the hand-coded ones (<jats:italic>p<\/jats:italic>&lt; .001). Furthermore we show that the type of database has a significant effect on learning and we give qualitative descriptions of the learned IQA policies.<\/jats:p>","DOI":"10.1017\/s1351324908004907","type":"journal-article","created":{"date-parts":[[2008,10,22]],"date-time":"2008-10-22T09:50:56Z","timestamp":1224669056000},"page":"55-72","source":"Crossref","is-referenced-by-count":5,"title":["Does this list contain what you were searching for? Learning adaptive dialogue strategies for interactive question answering"],"prefix":"10.1017","volume":"15","author":[{"given":"V.","family":"RIESER","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"O.","family":"LEMON","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"56","published-online":{"date-parts":[[2009,1,1]]},"reference":[{"key":"S1351324908004907_ref2","unstructured":"Becker T. , Blaylock N. , Gerstenberger C. , Korthauer A. , Perera N. , Pitz M. , Poller P. , Schehl J. , Steffens F. , Stegmann R. , and Steigner J. 2007. In-car showcase based on TALK libraries. Technical Report, Deliverable 5.3, TALK Project."},{"key":"S1351324908004907_ref9","unstructured":"Lemon O. , and Liu X. 2007. Dialogue policy learning for combinations of noise and user simulations: transfer results. In Proceedings of 8th SIGdial Workshop, Association for Computational Linguistics, Antwerp, Belgium, pp. 55\u201358."},{"key":"S1351324908004907_ref25","doi-asserted-by":"publisher","DOI":"10.1007\/978-1-4020-4746-6"},{"key":"S1351324908004907_ref17","doi-asserted-by":"publisher","DOI":"10.1109\/TSA.2005.855836"},{"key":"S1351324908004907_ref26","volume-title":"Reinforcement Learning","author":"Sutton","year":"1998"},{"key":"S1351324908004907_ref31","volume-title":"Integrating a QA System with Dialogue Management for the Music Domain","author":"Zhao","year":"2007"},{"key":"S1351324908004907_ref19","doi-asserted-by":"crossref","unstructured":"Rieser V. , and Lemon O. 2006a. Cluster-based user simulations for learning dialogue strategies. In Proceedings of Interspeech\/ICSLP 2006, Pittsburgh, PA, USA.","DOI":"10.21437\/Interspeech.2006-489"},{"key":"S1351324908004907_ref30","doi-asserted-by":"crossref","unstructured":"Walker M. A. , Passonneau R. J. , and Boland J. E. 2001. Quantitative and qualitative evaluation of DARPA communicator spoken dialogue systems. In Proceedings of Association for Computational Linguistics (ACL), Toulouse, France.","DOI":"10.3115\/1073012.1073078"},{"key":"S1351324908004907_ref27","doi-asserted-by":"crossref","unstructured":"Varges S. , Weng F. , and Pon-Barry H. 2006. Interactive question answering and constraint relaxation in spoken dialogue systems. In Proceedings of 7th SIGdial Workshop, Sydney, Australia.","DOI":"10.3115\/1654595.1654601"},{"key":"S1351324908004907_ref8","doi-asserted-by":"crossref","unstructured":"Lemon O. , Georgila K. , and Henderson J. 2006. Evaluating effectiveness and portability of reinforcement learned dialogue strategies with real users: the TALK TownInfo evaluation. In Proceedings of Spoken Language Technology (SLT), Palm Beach, Aruba.","DOI":"10.1109\/SLT.2006.326774"},{"key":"S1351324908004907_ref1","first-page":"209","volume-title":"Handbook of Natural Language Processing","author":"Androutsopoulos","year":"2000"},{"key":"S1351324908004907_ref3","doi-asserted-by":"crossref","unstructured":"Chung G. 2004. Developing a flexible spoken dialog system using simulation. In Proceedings of Association for Computational Linguistics (ACL), Barcelona, Spain.","DOI":"10.3115\/1218955.1218964"},{"key":"S1351324908004907_ref5","doi-asserted-by":"crossref","unstructured":"Divi V. , Forlines C. , Gemert J. , van Raj B. , Schmidt-Nielsen B. , Wittenburg K. , Woelfel J. , Wolf P. , and Zhang F. F. 2004. A speech-in list-out approach to spoken user interfaces. In Proceedings of Human Language Technologies Conference (HLT), Boston, MA, USA.","DOI":"10.3115\/1613984.1614013"},{"key":"S1351324908004907_ref6","doi-asserted-by":"crossref","unstructured":"Dohsaka K. , Yasuda N. , and Aikawa K. 2003. Efficient spoken dialogue control depending on the speech recognition rate and system's database. In Proceedings of Eurospeech, Geneva, Switzerland.","DOI":"10.21437\/Eurospeech.2003-270"},{"key":"S1351324908004907_ref7","doi-asserted-by":"crossref","unstructured":"Georgila K. , Henderson J. , and Lemon O. 2006. User simulation for spoken dialogue systems: learning and evaluation. In Proceedings of Interspeech\/ICSLP 2006, Pittsburgh, PA, USA.","DOI":"10.21437\/Interspeech.2006-160"},{"key":"S1351324908004907_ref10","unstructured":"Lemon O. , Liu X. , Shapiro D. , and Tollander C. 2006. Hierarchical reinforcement learning of dialogue policies in a development environment for dialogue systems: REALL-DUDE. In Proceedings of 10th SEMdial Workshop on the Semantics and Pragmatics of Dialogue (BRANDIAL), Potsdam, Germany."},{"key":"S1351324908004907_ref20","doi-asserted-by":"crossref","unstructured":"Rieser V. , and Lemon O. 2006b. Using machine learning to explore human multimodal clarification strategies. In Proceedings of Association for Computational Linguistics (ACL), Sydney, Australia.","DOI":"10.3115\/1273073.1273158"},{"key":"S1351324908004907_ref11","doi-asserted-by":"crossref","unstructured":"Levin E. , and Pieraccini R. 1997. A stochastic model of computer-human interaction for learning dialogue strategies. In Proceedings of Eurospeech, Rhodos, Greece.","DOI":"10.21437\/Eurospeech.1997-380"},{"key":"S1351324908004907_ref12","doi-asserted-by":"publisher","DOI":"10.1109\/89.817450"},{"key":"S1351324908004907_ref13","doi-asserted-by":"crossref","unstructured":"Oviatt S. , Levow G.-A. , MacEarchern M. , and Kuhn K. 1996. Modelling hyperarticulate speech during human-computer error resolution. In Proceedings of the International Conference on Spoken Language Processing (ICSLP), Philadelphia, PA, USA.","DOI":"10.21437\/ICSLP.1996-208"},{"key":"S1351324908004907_ref14","unstructured":"Paek T. 2006. Reinforcement learning for spoken dialogue systems: comparing strengths and weaknesses for practical deployment. In Dialogue on Dialogues. Interspeech\/ICSLP Satellite Workshop, Pittsburgh, PA, USA."},{"key":"S1351324908004907_ref15","unstructured":"Paek T. , and Horvitz E. 2000. Conversation as action under uncertainty. In Proceedings of the 16th Conference on Uncertainty in Artificial Intelligence (UAI), Stanford, CA, USA."},{"key":"S1351324908004907_ref18","doi-asserted-by":"crossref","unstructured":"Rieser V. , Kruijff-Korbayov\u00e1 I. , and Lemon O. 2005. A corpus collection and annotation framework for learning multimodal clarification strategies. In Proceedings of 6th SIGdial Workshop, Lisbon, Portugal.","DOI":"10.3115\/1273073.1273158"},{"key":"S1351324908004907_ref4","unstructured":"Demberg V. , and Moore J. 2006. Information presentation in spoken dialogue systems. In Proceedings of European Association for Computational Linguistics (EACL), Trento, Italy."},{"key":"S1351324908004907_ref21","doi-asserted-by":"crossref","unstructured":"Scheffler K. , and Young S. J. 2002. Automatic learning of dialogue strategy using dialogue simulation and reinforcement learning. In Proceedings of Human Language Technology (HLT), San Diego, CA, USA.","DOI":"10.3115\/1289189.1289246"},{"key":"S1351324908004907_ref22","unstructured":"Shapiro D. , and Langley P. 2002. Separating skills from preference: using learning to program by reward. In Proceedings of 19th International Conference on Machine Learning, Sydney, Australia."},{"key":"S1351324908004907_ref23","doi-asserted-by":"crossref","first-page":"105","DOI":"10.1613\/jair.859","article-title":"Optimizing dialogue management with reinforcement learning: experiments with the NJFun system","volume":"16","author":"Singh","year":"2002","journal-title":"Journal of Artificial Intelligence Research (JAIR)"},{"key":"S1351324908004907_ref24","unstructured":"Skantze G. 2007. Making grounding decisions: data-driven estimation of dialogue costs and confidence thresholds. In Proceedings of 8th SIGdial, Antwerp, Belgium."},{"key":"S1351324908004907_ref29","doi-asserted-by":"publisher","DOI":"10.1017\/S1351324900002503"},{"key":"S1351324908004907_ref28","doi-asserted-by":"publisher","DOI":"10.1007\/s10579-005-2696-1"},{"key":"S1351324908004907_ref16","doi-asserted-by":"publisher","DOI":"10.1007\/11861461_19"}],"container-title":["Natural Language Engineering"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.cambridge.org\/core\/services\/aop-cambridge-core\/content\/view\/S1351324908004907","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,5,20]],"date-time":"2023-05-20T17:03:36Z","timestamp":1684602216000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.cambridge.org\/core\/product\/identifier\/S1351324908004907\/type\/journal_article"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2009,1]]},"references-count":31,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2009,1]]}},"alternative-id":["S1351324908004907"],"URL":"https:\/\/doi.org\/10.1017\/s1351324908004907","relation":{},"ISSN":["1351-3249","1469-8110"],"issn-type":[{"value":"1351-3249","type":"print"},{"value":"1469-8110","type":"electronic"}],"subject":[],"published":{"date-parts":[[2009,1]]}}}