{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,12,8]],"date-time":"2025-12-08T22:36:24Z","timestamp":1765233384196,"version":"3.40.5"},"reference-count":65,"publisher":"Cambridge University Press (CUP)","issue":"1","license":[{"start":{"date-parts":[[2022,9,23]],"date-time":"2022-09-23T00:00:00Z","timestamp":1663891200000},"content-version":"unspecified","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["cambridge.org"],"crossmark-restriction":true},"short-container-title":["Nat. Lang. Eng."],"published-print":{"date-parts":[[2024,1]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Randomized prospective studies represent the gold standard for experimental design. In this paper, we present a randomized prospective study to validate the benefits of combining rule-based and data-driven natural language understanding methods in a virtual patient dialogue system. The system uses a rule-based pattern matching approach together with a machine learning (ML) approach in the form of a text-based convolutional neural network, combining the two methods with a simple logistic regression model to choose between their predictions for each dialogue turn. In an earlier, retrospective study, the hybrid system yielded a nearly 50% error reduction on our initial data, in part due to the differential performance between the two methods as a function of label frequency. Given these gains, and considering that our hybrid approach is unique among virtual patient systems, we compare the hybrid system to the rule-based system by itself in a randomized prospective study. We evaluate 110 unique medical student subjects interacting with the system over 5,296 conversation turns, to verify whether similar gains are observed in a deployed system. This prospective study broadly confirms the findings from the earlier one but also highlights important deficits in our training data. The hybrid approach still improves over either rule-based or ML approaches individually, even handling unseen classes with some success. However, we observe that live subjects ask more out-of-scope questions than expected. To better handle such questions, we investigate several modifications to the system combination component. These show significant overall accuracy improvements and modest F1 improvements on out-of-scope queries in an offline evaluation. We provide further analysis to characterize the difficulty of the out-of-scope problem that we have identified, as well as to suggest future improvements over the baseline we establish here.<\/jats:p>","DOI":"10.1017\/s1351324922000420","type":"journal-article","created":{"date-parts":[[2022,9,23]],"date-time":"2022-09-23T08:23:35Z","timestamp":1663921415000},"page":"31-72","update-policy":"https:\/\/doi.org\/10.1017\/policypage","source":"Crossref","is-referenced-by-count":2,"title":["A randomized prospective study of a hybrid rule- and data-driven virtual patient"],"prefix":"10.1017","volume":"30","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-5158-8508","authenticated-orcid":false,"given":"Adam","family":"Stiff","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Michael","family":"White","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Eric","family":"Fosler-Lussier","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Lifeng","family":"Jin","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Evan","family":"Jaffe","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Douglas","family":"Danforth","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"56","published-online":{"date-parts":[[2022,9,23]]},"reference":[{"key":"S1351324922000420_ref5","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/W17-2343"},{"key":"S1351324922000420_ref55","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-33197-8_25"},{"key":"S1351324922000420_ref23","first-page":"188","volume-title":"Irish Conference on Artificial Intelligence and Cognitive Science","author":"Khan","year":"2009"},{"key":"S1351324922000420_ref38","first-page":"2825","article-title":"Scikit-learn: Machine learning in Python","volume":"12","author":"Pedregosa","year":"2011","journal-title":"Journal of Machine Learning Research"},{"key":"S1351324922000420_ref57","unstructured":"Wilcox, B. (2019). Chatscript. Available at https:\/\/github.com\/ChatScript\/ChatScript (accessed 23 July 2019)."},{"key":"S1351324922000420_ref62","unstructured":"Yu, H. and Sable, C. (2005). Being erlang shen: Identifying answerable questions. In Proceedings of the Nineteenth International Joint Conference on Artificial Intelligence on Knowledge and Reasoning for Answering Questions. Citeseer, pp. 6\u201314."},{"key":"S1351324922000420_ref43","doi-asserted-by":"publisher","DOI":"10.1136\/amiajnl-2013-002544"},{"key":"S1351324922000420_ref6","doi-asserted-by":"publisher","DOI":"10.1007\/s10916-021-01737-4"},{"key":"S1351324922000420_ref18","first-page":"249","volume-title":"Proceedings of the 13th International Conference on Artificial Intelligence and Statistics (AISTATS)","author":"Glorot","year":"2010"},{"key":"S1351324922000420_ref12","unstructured":"DeVault, D. , Leuski, A. and Sagae, K. (2011a). An evaluation of alternative strategies for implementing dialogue policies using statistical classification and rules. In Proceedings of the 5th International Joint Conference on Natural Language Processing (IJCNLP). The Association for Computer Linguistics, pp. 1341\u20131345."},{"key":"S1351324922000420_ref8","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/W16-2922"},{"key":"S1351324922000420_ref34","doi-asserted-by":"publisher","DOI":"10.1007\/978-1-4614-8280-2_28"},{"key":"S1351324922000420_ref13","doi-asserted-by":"publisher","DOI":"10.5087\/dad.2011.107"},{"key":"S1351324922000420_ref15","first-page":"1871","article-title":"Liblinear: A library for large linear classification","volume":"9","author":"Fan","year":"2008","journal-title":"Journal of Machine Learning Research"},{"key":"S1351324922000420_ref40","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2019.2917862"},{"key":"S1351324922000420_ref7","first-page":"1","article-title":"Designing a virtual patient dialogue system based on terminology-rich resources: Challenges and evaluation","volume":"26","author":"Campillos-Llanos","year":"2019","journal-title":"Natural Language Engineering"},{"volume-title":"Proceedings of the Sixth International Conference on Language Resources and Evaluation (LREC\u201908)","year":"2008","author":"Robinson","key":"S1351324922000420_ref46"},{"key":"S1351324922000420_ref41","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/N18-1202"},{"key":"S1351324922000420_ref35","unstructured":"Morbini, F. , Forbell, E. , DeVault, D. , Sagae, K. , Traum, D. and Rizzo, A. (2012). A mixed-initiative conversational dialogue system for healthcare. In Proceedings of the 13th Annual Meeting of the Special Interest Group on Discourse and Dialogue. Association for Computational Linguistics, pp. 137\u2013139."},{"key":"S1351324922000420_ref54","doi-asserted-by":"publisher","DOI":"10.1016\/S0167-8655(99)00087-2"},{"key":"S1351324922000420_ref59","doi-asserted-by":"publisher","DOI":"10.1016\/S0893-6080(05)80023-1"},{"volume-title":"ADADELTA: An Adaptive Learning Rate Method","year":"2012","author":"Zeiler","key":"S1351324922000420_ref63"},{"key":"S1351324922000420_ref47","unstructured":"Sch\u00f6lkopf, B. , Williamson, R.C. , Smola, A.J. , Shawe-Taylor, J. and Platt, J.C. (2000). Support vector method for novelty detection. In Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 582\u2013588."},{"key":"S1351324922000420_ref27","doi-asserted-by":"publisher","DOI":"10.1145\/1357054.1357127"},{"key":"S1351324922000420_ref50","doi-asserted-by":"crossref","unstructured":"Stiff, A. , Song, Q. and Fosler-Lussier, E. (2020). How self-attention improves rare class performance in a question-answering dialogue agent. In Proceedings of the 21th Annual Meeting of the Special Interest Group on Discourse and Dialogue. Association for Computational Linguistics, pp. 196\u2013202.","DOI":"10.18653\/v1\/2020.sigdial-1.24"},{"key":"S1351324922000420_ref2","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2004.03.009"},{"volume-title":"Proceedings of the 1999 IEEE ASRU Workshop","year":"1999","author":"Glass","key":"S1351324922000420_ref17"},{"key":"S1351324922000420_ref33","first-page":"3111","volume-title":"Advances in Neural Information Processing Systems 26 (NIPS 2013)","author":"Mikolov","year":"2013"},{"key":"S1351324922000420_ref36","first-page":"807","volume-title":"Proceedings of the 27th International Conference on Machine Learning (ICML 2010)","volume":"3","author":"Nair","year":"2010"},{"key":"S1351324922000420_ref3","unstructured":"Brade\u0161ko, L. and Mladeni\u0107, D. (2012). A survey of chatbot systems through a loebner prize competition. In Proceedings of Slovenian Language Technologies Society Eighth Conference of Language Technologies. Slovenian Language Technologies Society, pp. 34\u201337."},{"key":"S1351324922000420_ref52","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP39728.2021.9413405"},{"key":"S1351324922000420_ref14","unstructured":"Devlin, J. , Chang, M.-W. , Lee, K. and Toutanova, K. (2019). Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers). Association for Computational Linguistics, pp. 4171\u20134186."},{"key":"S1351324922000420_ref64","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/W18-5001"},{"key":"S1351324922000420_ref53","doi-asserted-by":"publisher","DOI":"10.4018\/jgcms.2012070101"},{"key":"S1351324922000420_ref58","first-page":"1","article-title":"Making it real: Loebner-winning chatbot design","volume":"189","author":"Wilcox","year":"2013","journal-title":"ARBOR Ciencia, Pensamiento y Cultura"},{"volume-title":"Human Behavior and the Principle of Least Effort","year":"1949","author":"Zipf","key":"S1351324922000420_ref65"},{"key":"S1351324922000420_ref37","unstructured":"Paszke, A. , Gross, S. , Chintala, S. , Chanan, G. , Yang, E. , DeVito, Z. , Lin, Z. , Desmaison, A. , Antiga, L. , and Lerer, A. (2017). Automatic differentiation in PyTorch. In NIPS Autodiff Workshop."},{"key":"S1351324922000420_ref19","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P19-1267"},{"key":"S1351324922000420_ref32","doi-asserted-by":"publisher","DOI":"10.1080\/0142159X.2019.1616683"},{"key":"S1351324922000420_ref60","doi-asserted-by":"publisher","DOI":"10.1007\/s10994-005-1122-7"},{"key":"S1351324922000420_ref56","doi-asserted-by":"publisher","DOI":"10.1111\/j.1525-1497.2006.00421.x"},{"key":"S1351324922000420_ref21","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/W18-0502"},{"key":"S1351324922000420_ref48","first-page":"1929","article-title":"Dropout: A simple way to prevent neural networks from overfitting","volume":"15","author":"Srivastava","year":"2014","journal-title":"Journal of Machine Learning Research"},{"key":"S1351324922000420_ref61","first-page":"70","article-title":"Pebl: Web page classification without negative examples","volume":"1","author":"Yu","year":"2004","journal-title":"IEEE Transactions on Knowledge and Data Engineering"},{"key":"S1351324922000420_ref26","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/D14-1181"},{"key":"S1351324922000420_ref39","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/D14-1162"},{"key":"S1351324922000420_ref9","first-page":"2493","article-title":"Natural language processing (almost) from scratch","volume":"12","author":"Collobert","year":"2011","journal-title":"Journal of Machine Learning Research"},{"key":"S1351324922000420_ref44","unstructured":"Ram, A. , Prasad, R. , Khatri, C. , Venkatesh, A. , Gabriel, R. , Liu, Q. , Nunn, J. , Hedayatnia, B. , Cheng, M. , Nagar, A. , King, E. , Bland, K. , Wartick, A. , Pan, Y. , Song, H. , Jayadevan, S. , Hwang, G. and Pettigrue, A. (2018). Conversational AI: The science behind the alexa prize. arXiv preprint arXiv:1801.03604."},{"volume-title":"AISB Symposium on Questions, Discourse and Dialogue","year":"2014","author":"Stoyanchev","key":"S1351324922000420_ref51"},{"key":"S1351324922000420_ref4","unstructured":"Campillos-Llanos, L. , Bouamor, D. , Zweigenbaum, P. and Rosset, S. (2016). Managing linguistic and terminological variation in a medical dialogue system. In Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC\u201916). European Language Resources Association, pp. 3167\u20133173."},{"key":"S1351324922000420_ref42","unstructured":"Porter, M.F. (2001). Snowball: A language for stemming algorithms. Available at http:\/\/snowball.tartarus.org\/texts\/introduction.html (accessed 5 August 2022)."},{"key":"S1351324922000420_ref20","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/W15-0611"},{"key":"S1351324922000420_ref30","doi-asserted-by":"publisher","DOI":"10.1093\/jamia\/ocaa106"},{"volume-title":"Natural Language Processing with Python: Analyzing Text with the Natural Language Toolkit","year":"2009","author":"Bird","key":"S1351324922000420_ref1"},{"key":"S1351324922000420_ref10","unstructured":"Danforth, D. , Price, A. , Maicher, K. , Post, D. , Liston, B. , Clinchot, D. , Ledford, C. , Way, D. and Cronau, H. (2013). Can virtual standardized patients be used to assess communication skills in medical students. In Proceedings of the 17th Annual IAMSE Meeting, St. Andrews, Scotland. International Association of Medical Science Educators."},{"key":"S1351324922000420_ref22","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/W17-5002"},{"key":"S1351324922000420_ref45","doi-asserted-by":"crossref","unstructured":"Ravichandran, D. , Hovy, E. and Och, F.J. (2003). Statistical qa-classifier vs. re-ranker: What\u2019s the difference? In Proceedings of the ACL, 2003 Workshop on Multilingual Summarization and Question Answering-Volume 12. Association for Computational Linguistics, pp. 69\u201375.","DOI":"10.3115\/1119312.1119321"},{"key":"S1351324922000420_ref49","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2019.8682550"},{"key":"S1351324922000420_ref28","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D17-1230"},{"key":"S1351324922000420_ref31","first-page":"2579","article-title":"Visualizing data using t-sne","volume":"9","author":"Maaten","year":"2008","journal-title":"Journal of Machine Learning Research"},{"key":"S1351324922000420_ref11","unstructured":"DeVault, D. , Georgila, K. , Artstein, R. , Morbini, F. , Traum, D. , Scherer, S. , Rizzo, A.S. and Morency, L.-P. (2013). Verbal indicators of psychological distress in interactive dialogue with a virtual human. In Proceedings of the SIGDIAL, 2013 Conference. Association for Computational Linguistics, pp. 193\u2013202."},{"key":"S1351324922000420_ref29","unstructured":"Liu, B. , Lee, W.S. , Yu, P.S. and Li, X. (2002). Partially supervised classification of text documents. In ICML, vol. 2. Citeseer, pp. 387\u2013394."},{"key":"S1351324922000420_ref25","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2018-1581"},{"volume-title":"Alexa Prize SocialBot Grand Challenge 2 Proceedings","year":"2018","author":"Khatri","key":"S1351324922000420_ref24"},{"key":"S1351324922000420_ref16","unstructured":"Fielding, R.T. (2000). Rest: Architectural Styles and the Design of Network-Based Software Architectures. Doctoral dissertation, University of California."}],"container-title":["Natural Language Engineering"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.cambridge.org\/core\/services\/aop-cambridge-core\/content\/view\/S1351324922000420","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,4,25]],"date-time":"2024-04-25T12:50:45Z","timestamp":1714049445000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.cambridge.org\/core\/product\/identifier\/S1351324922000420\/type\/journal_article"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,9,23]]},"references-count":65,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2024,1]]}},"alternative-id":["S1351324922000420"],"URL":"https:\/\/doi.org\/10.1017\/s1351324922000420","relation":{},"ISSN":["1351-3249","1469-8110"],"issn-type":[{"type":"print","value":"1351-3249"},{"type":"electronic","value":"1469-8110"}],"subject":[],"published":{"date-parts":[[2022,9,23]]},"assertion":[{"value":"\u00a9 The Author(s), 2022. Published by Cambridge University Press","name":"copyright","label":"Copyright","group":{"name":"copyright_and_licensing","label":"Copyright and Licensing"}},{"value":"This is an Open Access article, distributed under the terms of the Creative Commons Attribution licence (https:\/\/creativecommons.org\/licenses\/by\/4.0\/), which permits unrestricted re-use, distribution, and reproduction in any medium, provided the original work is properly cited.","name":"license","label":"License","group":{"name":"copyright_and_licensing","label":"Copyright and Licensing"}},{"value":"This content has been made available to all.","name":"free","label":"Free to read"}]}}