{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,30]],"date-time":"2026-07-30T14:02:56Z","timestamp":1785420176260,"version":"3.56.0"},"reference-count":37,"publisher":"MDPI AG","issue":"5","license":[{"start":{"date-parts":[[2026,4,27]],"date-time":"2026-04-27T00:00:00Z","timestamp":1777248000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100000038","name":"Natural Sciences and Engineering Research Council of Canada","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100000038","id-type":"DOI","asserted-by":"crossref"}]},{"id":[{"id":"https:\/\/ror.org\/01h531d29","id-type":"ROR","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100000038","name":"Alberta Machine Intelligence Institute","doi-asserted-by":"publisher","award":["DGECR-2022-00369"],"award-info":[{"award-number":["DGECR-2022-00369"]}],"id":[{"id":"10.13039\/501100000038","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100000038","name":"Alberta Machine Intelligence Institute","doi-asserted-by":"publisher","award":["RGPIN-2022-0346"],"award-info":[{"award-number":["RGPIN-2022-0346"]}],"id":[{"id":"10.13039\/501100000038","id-type":"DOI","asserted-by":"publisher"}]},{"award":["DGECR-2022-00369"],"award-info":[{"award-number":["DGECR-2022-00369"]}],"id":[{"id":"https:\/\/ror.org\/03msd8560","id-type":"ROR","asserted-by":"publisher"}]},{"award":["RGPIN-2022-0346"],"award-info":[{"award-number":["RGPIN-2022-0346"]}],"id":[{"id":"https:\/\/ror.org\/03msd8560","id-type":"ROR","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Computers"],"abstract":"<jats:p>Large Language Models (LLMs) used for clinical decision support must not only make accurate predictions but also generate rationales that are consistent with, and sufficient for, those predictions. Building on Reason2Decide, a two-stage rationale-driven multi-task framework, we propose Reason2Decide-C (R2D-C, where C denotes cycle consistency), which augments Reason2Decide\u2019s stage 2 training with confidence-adaptive scheduled sampling and cycle-consistent rationale-to-label training. In stage 1, we pretrain our model on rationale generation. In stage 2, we jointlytrain on label prediction and rationale generation, gradually replacing gold labels with model-predicted labels based on confidence. Simultaneously, we feed the rationale logits back into the model to recover the label, thus enforcing explanation sufficiency. We evaluate R2D-C on one proprietary triage dataset, as well as public biomedical QA and reasoning datasets. Across model sizes, R2D-C substantially improves rationale\u2013prediction consistency (where stage 1 and stage 2 predictions agree) and sufficiency (where the rationale alone recovers the ground-truth label) over other baselines while matching or modestly improving predictive performance (F1); in several settings R2D-C surpasses 40\u00d7 larger foundation models. Ablations confirm that the full combination is optimal, maximizing alignment and LLM-as-a-Judge rationale quality. These results demonstrate that confidence-adaptive scheduled sampling and cycle-consistent rationale-to-label training substantially enhance explanation alignment without sacrificing accuracy.<\/jats:p>","DOI":"10.3390\/computers15050279","type":"journal-article","created":{"date-parts":[[2026,4,28]],"date-time":"2026-04-28T08:12:31Z","timestamp":1777363951000},"page":"279","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Reason2Decide-C: Adaptive Cycle-Consistent Training for Clinical Rationales"],"prefix":"10.3390","volume":"15","author":[{"ORCID":"https:\/\/orcid.org\/0009-0009-0789-6421","authenticated-orcid":false,"given":"H M Quamran","family":"Hasan","sequence":"first","affiliation":[{"name":"Department of Computing Science, Alberta Machine Intelligence Institute, University of Alberta, Edmonton, AB T6G 2R3, Canada"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Housam Khalifa Bashier","family":"Babiker","sequence":"additional","affiliation":[{"name":"Department of Computing Science, Alberta Machine Intelligence Institute, University of Alberta, Edmonton, AB T6G 2R3, Canada"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-4486-9738","authenticated-orcid":false,"given":"Mi-Young","family":"Kim","sequence":"additional","affiliation":[{"name":"Department of Science, Augustana Faculty, University of Alberta, Camrose, AB T4V 2R3, Canada"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0739-2946","authenticated-orcid":false,"given":"Randy","family":"Goebel","sequence":"additional","affiliation":[{"name":"Department of Computing Science, Alberta Machine Intelligence Institute, University of Alberta, Edmonton, AB T6G 2R3, Canada"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2026,4,27]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Gurrapu, S., Kulkarni, A., Huang, L., Lourentzou, I., and Batarseh, F.A. (2023). Rationalization for explainable NLP: A survey. Front. Artif. Intell., 6.","DOI":"10.3389\/frai.2023.1225093"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Cross, J.L., Choma, M.A., and Onofrey, J.A. (2024). Bias in medical AI: Implications for clinical decision-making. PLoS Digit. Health, 3.","DOI":"10.1371\/journal.pdig.0000651"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Lyu, Q., Apidianaki, M., and Callison-Burch, C. (2024). Towards Faithful Model Explanation in NLP: A Survey. arXiv.","DOI":"10.1162\/coli_a_00511"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Yang, J., Glockner, M., Rocha, A., and Gurevych, I. (2025). Self-Rationalization in the Wild: A Large Scale Out-of-Distribution Evaluation on NLI-related tasks. arXiv.","DOI":"10.1162\/tacl_a_00741"},{"key":"ref_5","unstructured":"Bhan, M., Vittaut, J.N., Chesneau, N., Chandar, S., and Lesot, M.J. (2026). NeuroFaith: Evaluating LLM Self-Explanation Faithfulness via Internal Representation Alignment. arXiv."},{"key":"ref_6","unstructured":"Ku, L.W., Martins, A., and Srikumar, V. (2024). Are self-explanations from Large Language Models faithful?. Proceedings of the Findings of the Association for Computational Linguistics: ACL 2024, Bangkok, Thailand, 11\u201316 August 2024, Association for Computational Linguistics."},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Hsieh, C.Y., Li, C.L., Yeh, C.K., Nakhost, H., Fujii, Y., Ratner, A., Krishna, R., Lee, C.Y., and Pfister, T. (2023). Distilling Step-by-Step! Outperforming Larger Language Models with Less Training Data and Smaller Model Sizes. arXiv.","DOI":"10.18653\/v1\/2023.findings-acl.507"},{"key":"ref_8","unstructured":"Hasan, H.M.Q., Bashier, H.K., Dai, J., Kim, M.Y., and Goebel, R. (2025). Reason2Decide: Rationale-Driven Multi-Task Learning. arXiv."},{"key":"ref_9","unstructured":"Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., Chi, E., Le, Q., and Zhou, D. (2023). Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. arXiv."},{"key":"ref_10","first-page":"1171","article-title":"Scheduled Sampling for Sequence Prediction with Recurrent Neural Networks","volume":"Volume 1","author":"Bengio","year":"2015","journal-title":"Proceedings of the 29th International Conference on Neural Information Processing Systems (NIPS\u201915), Montreal, QC, Canada, 7\u201312 December 2015"},{"key":"ref_11","unstructured":"Birch, A., Finch, A., Hayashi, H., Konstas, I., Luong, T., Neubig, G., Oda, Y., and Sudoh, K. (2019). Generalization in Generation: A closer look at Exposure Bias. Proceedings of the 3rd Workshop on Neural Generation and Translation, Hong Kong, China, 4 November 2019, Association for Computational Linguistics."},{"key":"ref_12","unstructured":"Jang, E., Gu, S., and Poole, B. (2017). Categorical Reparameterization with Gumbel-Softmax. arXiv."},{"key":"ref_13","unstructured":"Blodgett, S.L., Cercas Curry, A., Dev, S., Madaio, M., Nenkova, A., Yang, D., and Xiao, Z. (2024). Properties and Challenges of LLM-Generated Explanations. Proceedings of the Third Workshop on Bridging Human\u2013Computer Interaction and Natural Language Processing, Mexico City, Mexico, 21 June 2024, Association for Computational Linguistics."},{"key":"ref_14","unstructured":"Atakishiyev, S., Babiker, H.K.B., Dai, J., Farruque, N., Hayashi, T., Hriti, N.S., Rahman, M.A., Smith, I., Kim, M.Y., and Za\u00efane, O.R. (2025). Explainability of Large Language Models: Opportunities and Challenges toward Generating Trustworthy Explanations. arXiv."},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"41","DOI":"10.1023\/A:1007379606734","article-title":"Multitask Learning","volume":"28","author":"Caruana","year":"1997","journal-title":"Mach. Learn."},{"key":"ref_16","unstructured":"Ruder, S. (2017). An Overview of Multi-Task Learning in Deep Neural Networks. arXiv."},{"key":"ref_17","unstructured":"Hinton, G., Vinyals, O., and Dean, J. (2015). Distilling the Knowledge in a Neural Network. arXiv."},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"1789","DOI":"10.1007\/s11263-021-01453-z","article-title":"Knowledge Distillation: A Survey","volume":"129","author":"Gou","year":"2021","journal-title":"Int. J. Comput. Vis."},{"key":"ref_19","first-page":"5485","article-title":"Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer","volume":"21","author":"Raffel","year":"2023","journal-title":"J. Mach. Learn. Res."},{"key":"ref_20","unstructured":"Zong, C., Xia, F., Li, W., and Navigli, R. (2021). Confidence-Aware Scheduled Sampling for Neural Machine Translation. Proceedings of the Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, Online, 1\u20136 August 2021, Association for Computational Linguistics."},{"key":"ref_21","unstructured":"Burstein, J., Doran, C., and Solorio, T. (2019). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), Minneapolis, MN, USA, 2\u20137 June 2019, Association for Computational Linguistics."},{"key":"ref_22","unstructured":"Rumshisky, A., Roberts, K., Bethard, S., and Naumann, T. (2020). Pretrained Language Models for Biomedical and Clinical Tasks: Understanding and Extending the State-of-the-Art. Proceedings of the 3rd Clinical Natural Language Processing Workshop, Online, 19 November 2020, Association for Computational Linguistics."},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"1234","DOI":"10.1093\/bioinformatics\/btz682","article-title":"BioBERT: A pre-trained biomedical language representation model for biomedical text mining","volume":"36","author":"Lee","year":"2019","journal-title":"Bioinformatics"},{"key":"ref_24","unstructured":"Yang, A., Li, A., Yang, B., Zhang, B., Hui, B., Zheng, B., Yu, B., Gao, C., Huang, C., and Lv, C. (2025). Qwen3 Technical Report. arXiv."},{"key":"ref_25","unstructured":"Inui, K., Jiang, J., Ng, V., and Wan, X. (2019). PubMedQA: A Dataset for Biomedical Research Question Answering. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), Hong Kong, China, 3\u20137 November 2019, Association for Computational Linguistics."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Tchango, A.F., Goel, R., Wen, Z., Martel, J., and Ghosn, J. (2022). DDXPlus: A New Dataset For Automatic Medical Diagnosis. arXiv.","DOI":"10.52202\/068431-2270"},{"key":"ref_27","unstructured":"Wu, J., Deng, W., Li, X., Liu, S., Mi, T., Peng, Y., Xu, Z., Liu, Y., Cho, H., and Choi, C.I. (2025). MedReason: Eliciting Factual Medical Reasoning Steps in LLMs via Knowledge Graphs. arXiv."},{"key":"ref_28","unstructured":"Loshchilov, I., and Hutter, F. (2019). Decoupled Weight Decay Regularization. arXiv."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., and Funtowicz, M. (2020). HuggingFace\u2019s Transformers: State-of-the-art Natural Language Processing. arXiv.","DOI":"10.18653\/v1\/2020.emnlp-demos.6"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Labrak, Y., Bazoge, A., Morin, E., Gourraud, P.A., Rouvier, M., and Dufour, R. (2024). BioMistral: A Collection of Open-Source Pretrained Large Language Models for Medical Domains. arXiv.","DOI":"10.18653\/v1\/2024.findings-acl.348"},{"key":"ref_31","unstructured":"Ankit Pal, M.S. (2026, February 27). OpenBioLLMs: Advancing Open-Source Large Language Models for Healthcare and Life Sciences. Available online: https:\/\/huggingface.co\/aaditya\/Llama3-OpenBioLLM-70B."},{"key":"ref_32","unstructured":"Zhang, T., Kishore, V., Wu, F., Weinberger, K.Q., and Artzi, Y. (2020). BERTScore: Evaluating Text Generation with BERT. arXiv."},{"key":"ref_33","unstructured":"Isabelle, P., Charniak, E., and Lin, D. (2002). Bleu: A Method for Automatic Evaluation of Machine Translation. Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, Philadelphia, PA, USA, 7\u201312 July 2002, Association for Computational Linguistics."},{"key":"ref_34","first-page":"4794","article-title":"Framework for Evaluating Faithfulness of Local Explanations","volume":"Volume 162","author":"Chaudhuri","year":"2022","journal-title":"Proceedings of the 39th International Conference on Machine Learning, Baltimore, MD, USA, 7\u201323 July 2022"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"DeYoung, J., Jain, S., Rajani, N.F., Lehman, E., Xiong, C., Socher, R., and Wallace, B.C. (2020). ERASER: A Benchmark to Evaluate Rationalized NLP Models. arXiv.","DOI":"10.18653\/v1\/2020.acl-main.408"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Zheng, L., Chiang, W.L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., and Xing, E.P. (2023). Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. arXiv.","DOI":"10.52202\/075280-2020"},{"key":"ref_37","unstructured":"Che, W., Nabende, J., Shutova, E., and Pilehvar, M.T. (2025). Knowledge-Augmented Multimodal Clinical Rationale Generation for Disease Diagnosis with Small Language Models. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Vienna, Austria, 27 July\u20131 August 2025, Association for Computational Linguistics."}],"container-title":["Computers"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2073-431X\/15\/5\/279\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,29]],"date-time":"2026-04-29T04:28:34Z","timestamp":1777436914000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2073-431X\/15\/5\/279"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,4,27]]},"references-count":37,"journal-issue":{"issue":"5","published-online":{"date-parts":[[2026,5]]}},"alternative-id":["computers15050279"],"URL":"https:\/\/doi.org\/10.3390\/computers15050279","relation":{},"ISSN":["2073-431X"],"issn-type":[{"value":"2073-431X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,4,27]]}}}