{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,23]],"date-time":"2026-06-23T01:04:28Z","timestamp":1782176668847,"version":"3.54.5"},"reference-count":32,"publisher":"Frontiers Media SA","license":[{"start":{"date-parts":[[2023,1,9]],"date-time":"2023-01-09T00:00:00Z","timestamp":1673222400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/100000001","name":"National Science Foundation","doi-asserted-by":"publisher","award":["1761931"],"award-info":[{"award-number":["1761931"]}],"id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000001","name":"National Science Foundation","doi-asserted-by":"publisher","award":["2133842"],"award-info":[{"award-number":["2133842"]}],"id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["frontiersin.org"],"crossmark-restriction":true},"short-container-title":["Front. Big Data"],"abstract":"<jats:p>Virtual Mental Health Assistants (VMHAs) are utilized in health care to provide patient services such as counseling and suggestive care. They are not used for patient diagnostic assistance because they cannot adhere to safety constraints and specialized clinical process knowledge (<jats:sans-serif>ProKnow<\/jats:sans-serif>) used to obtain clinical diagnoses. In this work, we define <jats:sans-serif>ProKnow<\/jats:sans-serif> as an ordered set of information that maps to evidence-based guidelines or categories of conceptual understanding to experts in a domain. We also introduce a new dataset of diagnostic conversations guided by safety constraints and <jats:sans-serif>ProKnow<\/jats:sans-serif> that healthcare professionals use (<jats:sans-serif>ProKnow<\/jats:sans-serif>-<jats:bold>data<\/jats:bold>). We develop a method for natural language question generation (NLG) that collects diagnostic information from the patient interactively (<jats:sans-serif>ProKnow<\/jats:sans-serif>-<jats:bold>algo<\/jats:bold>). We demonstrate the limitations of using state-of-the-art large-scale language models (LMs) on this dataset. <jats:sans-serif>ProKnow<\/jats:sans-serif>-<jats:bold>algo<\/jats:bold> incorporates the process knowledge through explicitly modeling safety, knowledge capture, and explainability. As computational metrics for evaluation do not directly translate to clinical settings, we involve expert clinicians in designing evaluation metrics that test four properties: safety, logical coherence, and knowledge capture for explainability while minimizing the standard cross entropy loss to preserve distribution semantics-based similarity to the ground truth. LMs with <jats:sans-serif>ProKnow<\/jats:sans-serif>-<jats:bold>algo<\/jats:bold> generated 89% safer questions in the depression and anxiety domain (tested property: <jats:italic>safety<\/jats:italic>). Further, without <jats:sans-serif>ProKnow<\/jats:sans-serif>-<jats:bold>algo<\/jats:bold> generations question did not adhere to clinical process knowledge in <jats:sans-serif>ProKnow<\/jats:sans-serif>-<jats:bold>data<\/jats:bold> (tested property: <jats:italic>knowledge capture<\/jats:italic>). In comparison, <jats:sans-serif>ProKnow<\/jats:sans-serif>-<jats:bold>algo<\/jats:bold>-based generations yield a 96% reduction in our metrics to measure knowledge capture. The explainability of the generated question is assessed by computing similarity with concepts in depression and anxiety knowledge bases. Overall, irrespective of the type of LMs, <jats:sans-serif>ProKnow<\/jats:sans-serif>-<jats:bold>algo<\/jats:bold> achieved an averaged 82% improvement over simple pre-trained LMs on safety, explainability, and process-guided question generation. For reproducibility, we will make <jats:sans-serif>ProKnow<\/jats:sans-serif>-<jats:bold>data<\/jats:bold> and the code repository of <jats:sans-serif>ProKnow<\/jats:sans-serif>-<jats:bold>algo<\/jats:bold> publicly available upon acceptance.<\/jats:p>","DOI":"10.3389\/fdata.2022.1056728","type":"journal-article","created":{"date-parts":[[2023,1,10]],"date-time":"2023-01-10T16:04:27Z","timestamp":1673366667000},"update-policy":"https:\/\/doi.org\/10.3389\/crossmark-policy","source":"Crossref","is-referenced-by-count":16,"title":["ProKnow: Process knowledge for safety constrained and explainable question generation for mental health diagnostic assistance"],"prefix":"10.3389","volume":"5","author":[{"given":"Kaushik","family":"Roy","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Manas","family":"Gaur","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Misagh","family":"Soltani","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Vipula","family":"Rawte","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ashwin","family":"Kalyan","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Amit","family":"Sheth","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1965","published-online":{"date-parts":[[2023,1,9]]},"reference":[{"key":"B1","article-title":"Personalized prediction of suicide risk for web-based intervention,","author":"Alambo","year":"2018","journal-title":"NIMH Conference"},{"key":"B2","doi-asserted-by":"publisher","first-page":"463","DOI":"10.1162\/tacl_a_00111","article-title":"Large-scale analysis of counseling conversations: an application of natural language processing to mental health","volume":"4","author":"Althoff","year":"2016","journal-title":"Trans. Assoc. Comput. Linguist"},{"key":"B3","doi-asserted-by":"crossref","first-page":"94","DOI":"10.18653\/v1\/W17-1612","article-title":"Ethical research protocols for social media health research,","author":"Benton","year":"2017","journal-title":"Proceedings of the First ACL Workshop on Ethics in Natural Language Processing"},{"key":"B4","first-page":"1","article-title":"Towards augmenting crisis counselor training by improving message retrieval,","author":"Demasi","year":"2019","journal-title":"Proceedings of the Sixth Workshop on Computational Linguistics and Clinical Psychology"},{"key":"B5","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2105.06511","article-title":"Nlp is not enough-contextualization of user input in chatbots","author":"Dolbir","year":"2021","journal-title":"arXiv [Preprint] arXiv:"},{"key":"B6","first-page":"1606","article-title":"Retrofitting word vectors to semantic lexicons,","author":"Faruqui","year":"2015","journal-title":"Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies"},{"key":"B7","doi-asserted-by":"crossref","first-page":"514","DOI":"10.1145\/3308558.3313698","article-title":"Knowledge-aware assessment of severity of suicide risk for early intervention,","author":"Gaur","year":"2019","journal-title":"The World Wide Web Conference"},{"key":"B8","doi-asserted-by":"publisher","first-page":"51","DOI":"10.1109\/MIC.2020.3031769","article-title":"Semantics of the black-box: can knowledge graphs help make deep learning systems more interpretable and explainable?","volume":"25","author":"Gaur","year":"2021","journal-title":"IEEE Internet Comput"},{"key":"B9","doi-asserted-by":"publisher","first-page":"10672","DOI":"10.1609\/aaai.v36i10.21312","article-title":"Iseeq: Information seeking question generation using dynamic meta-information retrieval and knowledge graphs","volume":"36","author":"Gaur","year":"2022","journal-title":"Proc. AAAI Conf. Artif. Intell"},{"key":"B10","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.clpsych-1.12","article-title":"Learning to automate follow-up question generation using process knowledge for depression triage on reddit posts","author":"Gupta","year":"2022","journal-title":"arXiv [Preprint] arXiv:"},{"key":"B11","unstructured":"How AI and data could personalize higher educationHarv. Bus. Rev2019"},{"key":"B12","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.1503.02531","article-title":"Distilling the knowledge in a neural network","author":"Hinton","year":"2015","journal-title":"arXiv [Preprint] arXiv:"},{"key":"B13","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P18-1031","article-title":"Universal language model fine-tuning for text classification","author":"Howard","year":"2018","journal-title":"arXiv [Preprint] arXiv:"},{"key":"B14","volume-title":"Language use in teenage crisis intervention and the immediate outcome: A machine automated analysis of large scale text data","author":"Huang","year":"2015"},{"key":"B15","doi-asserted-by":"publisher","first-page":"509","DOI":"10.3928\/0048-5713-20020901-06","article-title":"The phq-9: a new depression diagnostic and severity measure","volume":"32","author":"Kroenke","year":"2002","journal-title":"Psychiatric Annals"},{"key":"B16","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2107.10410","article-title":"Evaluation of in-person counseling strategies to develop physical activity chatbot for women","author":"Liang","year":"2021","journal-title":"arXiv [Preprint] arXiv:"},{"key":"B17","doi-asserted-by":"crossref","first-page":"1106","DOI":"10.1145\/3308558.3313737","article-title":"Learning to generate questions by learningwhat not to generate,","author":"Liu","year":"2019","journal-title":"The World Wide Web Conference"},{"key":"B18","first-page":"142","article-title":"Counter-fitting word vectors to linguistic constraints,","author":"Mrk\u0161i\u0107","year":"2016","journal-title":"Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies"},{"key":"B19","doi-asserted-by":"publisher","first-page":"12350","DOI":"10.5210\/fm.v27i1.12350","article-title":"Spinning words as disguise: Shady services for ethical research?","volume":"27","author":"Reagle","year":"2022","journal-title":"First Monday"},{"key":"B20","doi-asserted-by":"publisher","DOI":"10.1145\/3477495.3531878","article-title":"Entity-conditioned question generation for robust attention distribution in neural information retrieval","author":"Reddy","year":"2022","journal-title":"arXiv [Preprint] arXiv:"},{"key":"B21","doi-asserted-by":"publisher","first-page":"113650","DOI":"10.1016\/j.eswa.2020.113650","article-title":"Towards integrated dialogue policy learning for multiple domains and .intents using hierarchical deep reinforcement learning","author":"Saha","year":"2020","journal-title":"Expert Syst Appl"},{"key":"B22","doi-asserted-by":"publisher","first-page":"e32875","DOI":"10.2196\/32875","article-title":"Operationalizing and implementing pretrained, large artificial intelligence linguistic models in the us health care system: outlook of generative pretrained transformer 3 (gpt-3) as a service model","volume":"10","author":"Sezgin","year":"2022","journal-title":"JMIR Med. Inform"},{"key":"B23","doi-asserted-by":"publisher","DOI":"10.1109\/MIC.2021.3101919","article-title":"Knowledge-intensive language understanding for explainable ai","author":"Sheth","year":"2021","journal-title":"arXiv [Preprint] arXiv:"},{"key":"B24","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D19-1441","article-title":"Patient knowledge distillation for bert model compression","author":"Sun","year":"2019","journal-title":"arXiv [Preprint] arXiv:"},{"key":"B25","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2201.08239","article-title":"Lamda: language models for dialog applications","author":"Thoppilan","year":"2022","journal-title":"arXiv [Preprint] arXiv:"},{"key":"B26","first-page":"5998","article-title":"Attention is all you need,","author":"Vaswani","year":"2017","journal-title":"Advances in Neural Information Processing Systems"},{"key":"B27","first-page":"353","article-title":"Glue: a multi-task benchmark and analysis platform for natural language understanding,","author":"Wang","year":"2018","journal-title":"Proceedings of the 2018 EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP"},{"key":"B28","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2112.04359","article-title":"Ethical and social risks of harm from language models","author":"Weidinger","year":"2021","journal-title":"arXiv [Preprint] arXiv:"},{"key":"B29","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.emnlp-demos.6","article-title":"Huggingface's transformers: State-of-the-art natural language processing","author":"Wolf","year":"2019","journal-title":"arXiv [Preprint] arXiv:"},{"key":"B30","doi-asserted-by":"publisher","first-page":"1191","DOI":"10.1145\/3110025.3123028","article-title":"Semi-supervised approach to monitoring clinical depressive symptoms in social media","volume":"2017","author":"Yazdavar","year":"2017","journal-title":"Proc. IEEE ACM Int. Conf. Adv. Soc. Netw. Anal. Min"},{"key":"B31","article-title":"On refining bert contextualized embeddings using semantic lexicons,","author":"Zervakis","year":"2021","journal-title":"Machine Learning with Symbolic Methods and Knowledge Graphs"},{"key":"B32","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D19-1253","article-title":"Addressing semantic drift in question generation for semi-supervised question answering","author":"Zhang","year":"2019","journal-title":"arXiv [Preprint] arXiv: 1909.06356"}],"container-title":["Frontiers in Big Data"],"original-title":[],"link":[{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/fdata.2022.1056728\/full","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,1,10]],"date-time":"2023-01-10T16:04:55Z","timestamp":1673366695000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/fdata.2022.1056728\/full"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,1,9]]},"references-count":32,"alternative-id":["10.3389\/fdata.2022.1056728"],"URL":"https:\/\/doi.org\/10.3389\/fdata.2022.1056728","relation":{},"ISSN":["2624-909X"],"issn-type":[{"value":"2624-909X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,1,9]]},"article-number":"1056728"}}