{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,9]],"date-time":"2026-06-09T15:03:08Z","timestamp":1781017388861,"version":"3.54.1"},"reference-count":22,"publisher":"MIT Press","issue":"1","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Computational Linguistics"],"published-print":{"date-parts":[[2011,3]]},"abstract":"<jats:p>We present a new data-driven methodology for simulation-based dialogue strategy learning, which allows us to address several problems in the field of automatic optimization of dialogue strategies: learning effective dialogue strategies when no initial data or system exists, and determining a data-driven reward function. In addition, we evaluate the result with real users, and explore how results transfer between simulated and real interactions. We use Reinforcement Learning (RL) to learn multimodal dialogue strategies by interaction with a simulated environment which is \u201cbootstrapped\u201d from small amounts of Wizard-of-Oz (WOZ) data. This use of WOZ data allows data-driven development of optimal strategies for domains where no working prototype is available. Using simulation-based RL allows us to find optimal policies which are not (necessarily) present in the original data. Our results show that simulation-based RL significantly outperforms the average (human wizard) strategy as learned from the data by using Supervised Learning. The bootstrapped RL-based policy gains on average 50 times more reward when tested in simulation, and almost 18 times more reward when interacting with real users. Users also subjectively rate the RL-based policy on average 10% higher. We also show that results from simulated interaction do transfer to interaction with real users, and we explicitly evaluate the stability of the data-driven reward function.<\/jats:p>","DOI":"10.1162\/coli_a_00038","type":"journal-article","created":{"date-parts":[[2011,1,21]],"date-time":"2011-01-21T13:22:31Z","timestamp":1295616151000},"page":"153-196","source":"Crossref","is-referenced-by-count":22,"title":["Learning and Evaluation of Dialogue Strategies for New Applications: Empirical Methods for Optimization from Small Data Sets"],"prefix":"10.1162","volume":"37","author":[{"given":"Verena","family":"Rieser","sequence":"first","affiliation":[{"name":"School of GeoSciences\/University of Edinburgh"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Oliver","family":"Lemon","sequence":"additional","affiliation":[{"name":"School of Mathematical and Computer Sciences\/Heriot-Watt University"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"281","reference":[{"issue":"1","key":"p_4","first-page":"13","volume":"23","author":"Isard Amy","year":"1997","journal-title":"Computational Linguistics"},{"key":"p_9","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1111\/j.2517-6161.1977.tb01600.x","volume":"39","author":"Laird A., N.","year":"1977","journal-title":"Journal of Royal Statistical Society B"},{"key":"p_14","doi-asserted-by":"publisher","DOI":"10.1017\/S0269888909990166"},{"key":"p_15","doi-asserted-by":"publisher","DOI":"10.1016\/0885-2308(91)90019-M"},{"key":"p_19","doi-asserted-by":"publisher","DOI":"10.1017\/S1351324909005105"},{"key":"p_20","doi-asserted-by":"publisher","DOI":"10.1162\/coli.2006.32.2.263"},{"key":"p_23","doi-asserted-by":"publisher","DOI":"10.1162\/coli.2008.07-028-R2-05-82"},{"key":"p_28","doi-asserted-by":"publisher","DOI":"10.1016\/j.csl.2009.01.001"},{"issue":"2","key":"p_32","first-page":"210","volume":"2","year":"2011","journal-title":"Computer Speech and Language"},{"key":"p_38","first-page":"11","author":"Pieraccini E., R.","year":"2000","journal-title":"IEEE Transactions on Speech and Audio Processing, 8(1), pages"},{"key":"p_45","doi-asserted-by":"publisher","DOI":"10.1109\/TSA.2005.855836"},{"issue":"1","key":"p_52","first-page":"55","volume":"15","author":"Oliver Lemon Verena","year":"2008","journal-title":"Natural Language Engineering"},{"issue":"1","key":"p_54","first-page":"3","volume":"16","author":"Oliver Lemon Verena","year":"2009","journal-title":"Natural Language Engineering"},{"key":"p_60","doi-asserted-by":"publisher","DOI":"10.1017\/S0269888906000944"},{"key":"p_67","doi-asserted-by":"publisher","DOI":"10.1613\/jair.859"},{"issue":"3","key":"p_68","first-page":"325","volume":"43","year":"2005","journal-title":"Speech Communication"},{"key":"p_71","first-page":"363","author":"Kamm M., C.","year":"2000","journal-title":"Natural Language Engineering, 6(3), pages"},{"key":"p_72","doi-asserted-by":"publisher","DOI":"10.1613\/jair.713"},{"key":"p_73","doi-asserted-by":"publisher","DOI":"10.1007\/s10579-005-2696-1"},{"key":"p_76","doi-asserted-by":"publisher","DOI":"10.1016\/j.csl.2006.05.001"},{"key":"p_80","doi-asserted-by":"publisher","DOI":"10.1098\/rsta.2000.0593"},{"issue":"2","key":"p_81","first-page":"150","volume":"24","author":"Gasic M.","year":"2009","journal-title":"Computer Speech and Language"}],"container-title":["Computational Linguistics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mitpressjournals.org\/doi\/pdf\/10.1162\/coli_a_00038","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,3,2]],"date-time":"2025-03-02T00:47:47Z","timestamp":1740876467000},"score":1,"resource":{"primary":{"URL":"https:\/\/direct.mit.edu\/coli\/article\/37\/1\/153-196\/2091"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2011,3]]},"references-count":22,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2011,3]]}},"alternative-id":["10.1162\/coli_a_00038"],"URL":"https:\/\/doi.org\/10.1162\/coli_a_00038","relation":{},"ISSN":["0891-2017","1530-9312"],"issn-type":[{"value":"0891-2017","type":"print"},{"value":"1530-9312","type":"electronic"}],"subject":[],"published":{"date-parts":[[2011,3]]}}}