{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,29]],"date-time":"2026-06-29T13:35:02Z","timestamp":1782740102926,"version":"3.54.5"},"reference-count":44,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2021,8,17]],"date-time":"2021-08-17T00:00:00Z","timestamp":1629158400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"EPSRC Fellowship titled \u201cTask Based Information Retrieval\u201d","award":["EP\/P024289\/1"],"award-info":[{"award-number":["EP\/P024289\/1"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Inf. Syst."],"published-print":{"date-parts":[[2021,10,31]]},"abstract":"<jats:p>As conversational agents like Siri and Alexa gain in popularity and use, conversation is becoming a more and more important mode of interaction for search. Conversational search shares some features with traditional search, but differs in some important respects: conversational search systems are less likely to return ranked lists of results (a SERP), more likely to involve iterated interactions, and more likely to feature longer, well-formed user queries in the form of natural language questions. Because of these differences, traditional methods for search evaluation (such as the Cranfield paradigm) do not translate easily to conversational search. In this work, we propose a framework for offline evaluation of conversational search, which includes a methodology for creating test collections with relevance judgments, an evaluation measure based on a user interaction model, and an approach to collecting user interaction data to train the model. The framework is based on the idea of \u201csubtopics\u201d, often used to model novelty and diversity in search and recommendation, and the user model is similar to the geometric browsing model introduced by RBP and used in ERR. As far as we know, this is the first work to combine these ideas into a comprehensive framework for offline evaluation of conversational search.<\/jats:p>","DOI":"10.1145\/3451160","type":"journal-article","created":{"date-parts":[[2021,8,17]],"date-time":"2021-08-17T13:57:35Z","timestamp":1629208655000},"page":"1-22","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":37,"title":["How Am I Doing?: Evaluating Conversational Search Systems Offline"],"prefix":"10.1145","volume":"39","author":[{"given":"Aldo","family":"Lipani","sequence":"first","affiliation":[{"name":"University College London, London, United Kingdom"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ben","family":"Carterette","sequence":"additional","affiliation":[{"name":"Spotify, New York, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Emine","family":"Yilmaz","sequence":"additional","affiliation":[{"name":"University College London &amp; Amazon, London, United Kingdom"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2021,8,17]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1145\/1277741.1277902"},{"key":"e_1_2_1_2_1","unstructured":"AlexaPrize. 2020. The Alexa Prize the socialbot challenge. https:\/\/developer.amazon.com\/alexaprize\/ AlexaPrize. 2020. The Alexa Prize the socialbot challenge. https:\/\/developer.amazon.com\/alexaprize\/"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/3331184.3331265"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1145\/3290605.3300233"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/3209978.3210024"},{"key":"e_1_2_1_6_1","first-page":"34","article-title":"Conversational search (dagstuhl seminar 19461)","volume":"9","author":"Anand Avishek","year":"2020","unstructured":"Avishek Anand , Lawrence Cavedon , Hideo Joho , Mark Sanderson , and Benno Stein . 2020 . Conversational search (dagstuhl seminar 19461) . Dagstuhl Reports 9 , 11 (2020), 34 \u2013 83 . https:\/\/doi.org\/10.4230\/DagRep.9.11.34 10.4230\/DagRep.9.11.34 Avishek Anand, Lawrence Cavedon, Hideo Joho, Mark Sanderson, and Benno Stein. 2020. Conversational search (dagstuhl seminar 19461). Dagstuhl Reports 9, 11 (2020), 34\u201383. https:\/\/doi.org\/10.4230\/DagRep.9.11.34","journal-title":"Dagstuhl Reports"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/2532508.2532511"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/W17-5522"},{"key":"e_1_2_1_9_1","volume-title":"Zur Theorie der Frage: Kolloquium","author":"Bunt H.C.","year":"1981","unstructured":"H.C. Bunt . 1981 . Conversational principles in question-answer dialogues . In Zur Theorie der Frage: Kolloquium , 1978, Bad Homburg : Vortraege \/ hrsg. von D. Krallman und G. Stickel (Forschungsberichte des Instituts fuer Deutsche Sprache Mannheim). T\u00fcbingen, 119\u2013142. H.C. Bunt. 1981. Conversational principles in question-answer dialogues. In Zur Theorie der Frage: Kolloquium, 1978, Bad Homburg: Vortraege \/ hrsg. von D. Krallman und G. Stickel (Forschungsberichte des Instituts fuer Deutsche Sprache Mannheim). T\u00fcbingen, 119\u2013142."},{"key":"e_1_2_1_10_1","volume-title":"Overview of the TREC 2014 session track. In TREC.","author":"Carterette Ben","unstructured":"Ben Carterette , Evangelos Kanoulas , Mark M. Hall , and Paul D. Clough . 2014 . Overview of the TREC 2014 session track. In TREC. Ben Carterette, Evangelos Kanoulas, Mark M. Hall, and Paul D. Clough. 2014. Overview of the TREC 2014 session track. In TREC."},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.5555\/567363.567368"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/1645953.1646033"},{"key":"e_1_2_1_13_1","first-page":"18","volume-title":"Conference on Empirical Methods in Natural Language Processing\u201d. Association for Computational Linguistics","author":"Choi Eunsol","year":"2018","unstructured":"Eunsol Choi , He He , Mohit Iyyer , Mark Yatskar , Wen-tau Yih, Yejin Choi , Percy Liang , and Luke Zettlemoyer . 2018 . QuAC: Question answering in context. In \u201dProceedings of the 2018 Conference on Empirical Methods in Natural Language Processing\u201d. Association for Computational Linguistics , Brussels, Belgium, 2174\u20132184. https:\/\/doi.org\/10. 18653\/v1\/D 18 - 1241 10.18653\/v1 Eunsol Choi, He He, Mohit Iyyer, Mark Yatskar, Wen-tau Yih, Yejin Choi, Percy Liang, and Luke Zettlemoyer. 2018. QuAC: Question answering in context. In \u201dProceedings of the 2018 Conference on Empirical Methods in Natural Language Processing\u201d. Association for Computational Linguistics, Brussels, Belgium, 2174\u20132184. https:\/\/doi.org\/10.18653\/v1\/D18-1241"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/3357384.3358047"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/2939672.2939746"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/1390334.1390446"},{"key":"e_1_2_1_17_1","volume-title":"Overview of the TREC 2011 web track. In Proceedings of The Twentieth Text REtrieval Conference, TREC 2011","author":"Clarke Charles L. A.","year":"2011","unstructured":"Charles L. A. Clarke , Nick Craswell , Ian Soboroff , and Ellen M. Voorhees . 2011 . Overview of the TREC 2011 web track. In Proceedings of The Twentieth Text REtrieval Conference, TREC 2011 , Gaithersburg, Maryland, USA, November 15\u201318 , 2011 . http:\/\/trec.nist.gov\/pubs\/trec20\/papers\/WEB.OVERVIEW.pdf Charles L. A. Clarke, Nick Craswell, Ian Soboroff, and Ellen M. Voorhees. 2011. Overview of the TREC 2011 web track. In Proceedings of The Twentieth Text REtrieval Conference, TREC 2011, Gaithersburg, Maryland, USA, November 15\u201318, 2011. http:\/\/trec.nist.gov\/pubs\/trec20\/papers\/WEB.OVERVIEW.pdf"},{"key":"e_1_2_1_18_1","doi-asserted-by":"crossref","unstructured":"Jeffrey Dalton Chenyan Xiong and Jamie Callan. 2020. TREC CAsT 2019: The conversational assistance track overview. In TREC. Jeffrey Dalton Chenyan Xiong and Jamie Callan. 2020. TREC CAsT 2019: The conversational assistance track overview. In TREC.","DOI":"10.6028\/NIST.SP.1266.cast-overview"},{"key":"e_1_2_1_19_1","volume-title":"CLEF 2016: Working Notes of CLEF 2016-Conference and Labs of the Evaluation Forum (\u00c9vora, Portugal, 5\u20138","author":"Gebremeskel GG","year":"2016","unstructured":"GG Gebremeskel and AP de Vries . 2016 . Recommender systems evaluations: Offline, online, time and A\/A test . In CLEF 2016: Working Notes of CLEF 2016-Conference and Labs of the Evaluation Forum (\u00c9vora, Portugal, 5\u20138 , September, 2016). [Sl]: CEUR, 642\u2013656. GG Gebremeskel and AP de Vries. 2016. Recommender systems evaluations: Offline, online, time and A\/A test. In CLEF 2016: Working Notes of CLEF 2016-Conference and Labs of the Evaluation Forum (\u00c9vora, Portugal, 5\u20138, September, 2016). [Sl]: CEUR, 642\u2013656."},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1080\/08839514.2013.835230"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.5555\/2594602.2594604"},{"key":"e_1_2_1_22_1","volume-title":"NIPS Conversational AI Workshop.","author":"Guo Fenfei","year":"2017","unstructured":"Fenfei Guo , Angeliki Metallinou , Chandra Khatri , Anirudh Raju , Anu Venkatesh , and Ashwin Ram . 2017 . Topic-based evaluation for conversational bots . In NIPS Conversational AI Workshop. Fenfei Guo, Angeliki Metallinou, Chandra Khatri, Anirudh Raju, Anu Venkatesh, and Ashwin Ram. 2017. Topic-based evaluation for conversational bots. In NIPS Conversational AI Workshop."},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/N15-1086"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1145\/2009916.2010056"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/3341981.3344216"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D16-1230"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/2911451.2914737"},{"key":"e_1_2_1_28_1","volume-title":"Proceedings of the 10th International Workshop on Spoken Dialogue Systems Technology 2019 (IWSDS\u201919)","author":"Liu Xingkun","year":"2019","unstructured":"Xingkun Liu , Arash Eshghi , Pawel Swietojanski , and Verena Rieser . 2019 . Benchmarking natural language understanding services for building conversational agents . In Proceedings of the 10th International Workshop on Spoken Dialogue Systems Technology 2019 (IWSDS\u201919) . https:\/\/iwsds2019.unikore.it\/ Xingkun Liu, Arash Eshghi, Pawel Swietojanski, and Verena Rieser. 2019. Benchmarking natural language understanding services for building conversational agents. In Proceedings of the 10th International Workshop on Spoken Dialogue Systems Technology 2019 (IWSDS\u201919). https:\/\/iwsds2019.unikore.it\/"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/1416950.1416952"},{"key":"e_1_2_1_30_1","volume-title":"MS MARCO: A human generated machine reading comprehension dataset. In CoCo@NIPS","author":"Nguyen Tri","year":"2016","unstructured":"Tri Nguyen , Mir Rosenberg , Xia Song , Jianfeng Gao , Saurabh Tiwary , Rangan Majumder , and Li Deng . 2016 . MS MARCO: A human generated machine reading comprehension dataset. In CoCo@NIPS , Vol. abs\/ 1611 .09268. http:\/\/ceur-ws.org\/Vol-1773\/CoCoNIPS_2016_paper9.pdf Tri Nguyen, Mir Rosenberg, Xia Song, Jianfeng Gao, Saurabh Tiwary, Rangan Majumder, and Li Deng. 2016. MS MARCO: A human generated machine reading comprehension dataset. In CoCo@NIPS, Vol. abs\/1611.09268. http:\/\/ceur-ws.org\/Vol-1773\/CoCoNIPS_2016_paper9.pdf"},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.3115\/1073083.1073135"},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/W19-5941"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1145\/3020165.3020183"},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D16-1264"},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1162\/coli_a_00322"},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1145\/3331184.3331215"},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.5087\/dad.2018.101"},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1145\/3121050.3121069"},{"key":"e_1_2_1_39_1","volume-title":"Dynamic search\u2013optimizing the game of information seeking. arXiv preprint arXiv:1909.12425","author":"Tang Zhiwen","year":"2019","unstructured":"Zhiwen Tang and Grace Hui Yang . 2019. Dynamic search\u2013optimizing the game of information seeking. arXiv preprint arXiv:1909.12425 ( 2019 ). Zhiwen Tang and Grace Hui Yang. 2019. Dynamic search\u2013optimizing the game of information seeking. arXiv preprint arXiv:1909.12425 (2019)."},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.ipm.2019.102162"},{"key":"e_1_2_1_41_1","volume-title":"NIPS Conversational AI Workshop.","author":"Venkatesh Anu","year":"2017","unstructured":"Anu Venkatesh , Chandra Khatri , Ashwin Ram , Fenfei Guo , Raefer Gabriel , Ashish Nagar , Rohit Prasad , Ming Cheng , Behnam Hedayatnia , Angeliki Metallinou , Rahul Goel , Shaohua Yang , and Anirudh Raju . 2017 . On evaluating and comparing conversational agents . In NIPS Conversational AI Workshop. Anu Venkatesh, Chandra Khatri, Ashwin Ram, Fenfei Guo, Raefer Gabriel, Ashish Nagar, Rohit Prasad, Ming Cheng, Behnam Hedayatnia, Angeliki Metallinou, Rahul Goel, Shaohua Yang, and Anirudh Raju. 2017. On evaluating and comparing conversational agents. In NIPS Conversational AI Workshop."},{"key":"e_1_2_1_42_1","volume-title":"ICML Deep Learning Workshop. http:\/\/arxiv.org\/pdf\/1506","author":"Vinyals Oriol","unstructured":"Oriol Vinyals and Quoc V. Le . 2015. A neural conversational model . In ICML Deep Learning Workshop. http:\/\/arxiv.org\/pdf\/1506 .05869v3.pdf Oriol Vinyals and Quoc V. Le. 2015. A neural conversational model. In ICML Deep Learning Workshop. http:\/\/arxiv.org\/pdf\/1506.05869v3.pdf"},{"key":"e_1_2_1_43_1","volume-title":"TREC 2016 dynamic domain track overview. In TREC.","author":"Yang Grace Hui","year":"2016","unstructured":"Grace Hui Yang and Ian Soboroff . 2016 . TREC 2016 dynamic domain track overview. In TREC. Grace Hui Yang and Ian Soboroff. 2016. TREC 2016 dynamic domain track overview. In TREC."},{"key":"e_1_2_1_44_1","volume-title":"Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics","author":"Zhou Kangyan","unstructured":"Kangyan Zhou , Shrimai Prabhumoye , and Alan W. Black . 2018. A dataset for document grounded conversations . In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics , Brussels, Belgium, 708\u2013713. Kangyan Zhou, Shrimai Prabhumoye, and Alan W. Black. 2018. A dataset for document grounded conversations. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, Brussels, Belgium, 708\u2013713."}],"container-title":["ACM Transactions on Information Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3451160","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3451160","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T17:49:25Z","timestamp":1750268965000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3451160"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,8,17]]},"references-count":44,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2021,10,31]]}},"alternative-id":["10.1145\/3451160"],"URL":"https:\/\/doi.org\/10.1145\/3451160","relation":{},"ISSN":["1046-8188","1558-2868"],"issn-type":[{"value":"1046-8188","type":"print"},{"value":"1558-2868","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,8,17]]},"assertion":[{"value":"2020-05-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2021-02-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2021-08-17","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}