{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,2]],"date-time":"2026-05-02T04:41:33Z","timestamp":1777696893492,"version":"3.51.4"},"reference-count":44,"publisher":"SAGE Publications","issue":"1","license":[{"start":{"date-parts":[[2025,4,1]],"date-time":"2025-04-01T00:00:00Z","timestamp":1743465600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/journals.sagepub.com\/page\/policies\/text-and-data-mining-license"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["No. 62372189"],"award-info":[{"award-number":["No. 62372189"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Research Grants Council of the Hong Kong Special Administrative Region, China","award":["UGC\/FDS16\/E09\/22"],"award-info":[{"award-number":["UGC\/FDS16\/E09\/22"]}]}],"content-domain":{"domain":["journals.sagepub.com"],"crossmark-restriction":true},"short-container-title":["Intelligent Data Analysis: An International Journal"],"published-print":{"date-parts":[[2026,1]]},"abstract":"<jats:p>\n                    Clinical trials are essential for discovering new treatments and advancing medical knowledge. However, the high uncertainty of carrying out clinical trials often ends with ineffective results. Therefore, the accurate prediction of clinical trial outcomes has become a significant challenge. Numerous publicly accessible clinical trial reports have been discovered to be beneficial in alleviating this challenge but lack necessary annotations to be formal datasets for deep model training. To address the issue, this paper proposes to construct a new clinical trial dataset by extracting publicly available clinical trial reports from\n                    <jats:italic toggle=\"yes\">ClinicalTrials.gov<\/jats:italic>\n                    and PubMed. In addition, a new two-stage method is proposed for the prediction of clinical trial outcomes across all trial phases. Specifically, our method first employs a prompt template combined with each clinical trial report to prompt a large language model to generate a concise summarization text containing essential information related to the clinical trial outcomes. Subsequently, this summarization text is utilized to train a classifier to predict the outcomes. Extensive experiments were conducted on the dataset, and our method was compared with several state-of-the-art classification models. The results showed that our method achieved the best performance in predicting clinical trial outcomes, especially using small amounts of training data under a data imbalance difficulty.\n                  <\/jats:p>","DOI":"10.1177\/1088467x251330470","type":"journal-article","created":{"date-parts":[[2025,4,2]],"date-time":"2025-04-02T03:48:52Z","timestamp":1743565732000},"page":"112-126","update-policy":"https:\/\/doi.org\/10.1177\/sage-journals-update-policy","source":"Crossref","is-referenced-by-count":0,"title":["A two-stage framework by leveraging large language model for predicting clinical trial outcomes"],"prefix":"10.1177","volume":"30","author":[{"given":"Baoshuo","family":"Kan","sequence":"first","affiliation":[{"name":"School of Computer Science, South China Normal University, Guangzhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hengdong","family":"Zhu","sequence":"additional","affiliation":[{"name":"School of Computer Science, South China Normal University, Guangzhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Heng","family":"Weng","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Dampness Syndrome of Chinese Medicine, The Second Affiliated Hospital of Guangzhou University of Chinese Medicine, Guangzhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Kun","family":"Zeng","sequence":"additional","affiliation":[{"name":"School of Computer Science, Sun Yat-Sen University, Guangzhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Fu Lee","family":"Wang","sequence":"additional","affiliation":[{"name":"School of Science and Technology, Hong Kong Metropolitan University, Hong Kong, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0001-5802-9274","authenticated-orcid":false,"given":"Tianyong","family":"Hao","sequence":"additional","affiliation":[{"name":"School of Computer Science, South China Normal University, Guangzhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"179","published-online":{"date-parts":[[2025,4]]},"reference":[{"key":"e_1_3_4_2_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.patter.2022.100445"},{"key":"e_1_3_4_3_2","doi-asserted-by":"crossref","unstructured":"Luo J Qiao Z Glass L et al. Clinicalrisk: a new therapy-related clinical trial dataset for predicting trial status and failure reasons. In: Proceedings of the 32nd ACM international conference on information and knowledge management 2023 pp.5356\u20135360.","DOI":"10.1145\/3583780.3615113"},{"key":"e_1_3_4_4_2","volume-title":"Basic principles of drug discovery and development","author":"Blass BE","year":"2015","unstructured":"Blass BE. Basic principles of drug discovery and development. Philadelphia, USA: Elsevier, 2015."},{"key":"e_1_3_4_5_2","doi-asserted-by":"publisher","DOI":"10.1177\/1740774518820060"},{"key":"e_1_3_4_6_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-18539-2"},{"key":"e_1_3_4_7_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.drudis.2016.07.008"},{"key":"e_1_3_4_8_2","article-title":"Improving the prediction of clinical success using machine learning","author":"Munos B","unstructured":"Munos B, Niederreiter J, Riccaboni M. Improving the prediction of clinical success using machine learning. medRxiv 2021.","journal-title":"medRxiv"},{"key":"e_1_3_4_9_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.ipm.2018.11.009"},{"key":"e_1_3_4_10_2","doi-asserted-by":"publisher","DOI":"10.1186\/s12911-019-0973-y"},{"key":"e_1_3_4_11_2","doi-asserted-by":"crossref","unstructured":"Jin Q Tan C Chen M et al. Predicting clinical trial results by implicit evidence integration. In: Proceedings of the 2020 conference on empirical methods in natural language processing (EMNLP) 2020 pp.1461\u20131477.","DOI":"10.18653\/v1\/2020.emnlp-main.114"},{"key":"e_1_3_4_12_2","doi-asserted-by":"crossref","unstructured":"Katsimpras G Paliouras G. Predicting intervention approval in clinical trials through multi-document summarization. In: Proceedings of the 60th annual meeting of the association for computational linguistics (Volume 1: Long Papers) 2022 pp.1947\u20131957.","DOI":"10.18653\/v1\/2022.acl-long.137"},{"key":"e_1_3_4_13_2","doi-asserted-by":"crossref","unstructured":"Lehman E DeYoung J Barzilay R et al. Inferring which medical treatments work from reports of clinical trials. In: Proceedings of the 2019 conference of the north American chapter of the association for computational linguistics: Human language technologies Volume 1 (Long and Short Papers) 2019 pp.3705\u20133717.","DOI":"10.18653\/v1\/N19-1371"},{"key":"e_1_3_4_14_2","doi-asserted-by":"crossref","unstructured":"Marshall IJ Kuiper J Banner E et al. Automating biomedical evidence synthesis: Robotreviewer. In: Proceedings of the conference. Association for computational linguistics. Meeting volume 2017 2017 p.7. NIH Public Access.","DOI":"10.18653\/v1\/P17-4002"},{"key":"e_1_3_4_15_2","doi-asserted-by":"crossref","unstructured":"Wang Z Sun J. Trial2vec: zero-shot clinical trial document similarity search using self-supervision. In: 2022 Findings of the association for computational linguistics: EMNLP 2022 2022.","DOI":"10.18653\/v1\/2022.findings-emnlp.476"},{"key":"e_1_3_4_16_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.chembiol.2016.07.023"},{"key":"e_1_3_4_17_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.patter.2021.100312"},{"key":"e_1_3_4_18_2","doi-asserted-by":"publisher","DOI":"10.1186\/1471-2164-13-S8-S21"},{"key":"e_1_3_4_19_2","doi-asserted-by":"crossref","unstructured":"Gao J Xiao C Glass LM et al. Compose: cross-modal pseudo-siamese network for patient trial matching. In: Proceedings of the 26th ACM SIGKDD international conference on knowledge discovery & data mining 2020 pp.803\u2013812.","DOI":"10.1145\/3394486.3403123"},{"key":"e_1_3_4_20_2","doi-asserted-by":"crossref","unstructured":"Zhang X Xiao C Glass LM et al. Deepenroll: patient-trial matching with deep embedding and entailment prediction. In: Proceedings of the web conference 2020 2020b pp.1029\u20131037.","DOI":"10.1145\/3366423.3380181"},{"key":"e_1_3_4_21_2","unstructured":"Qi Y Tang Q. Predicting phase 3 clinical trial results by modeling phase 2 clinical trial subject level data using deep learning. In: Machine learning for healthcare conference 2019 pp.288\u2013303. PMLR."},{"key":"e_1_3_4_22_2","unstructured":"Lo AW Siah KW Wong CH. Machine learning with statistical imputation for predicting drug approvals. Available at SSRN 2973611."},{"key":"e_1_3_4_23_2","first-page":"46","article-title":"A novel system for extractive clinical note summarization using EHR data","author":"Liang J","year":"2019","unstructured":"Liang J, Tsou C-H. A novel system for extractive clinical note summarization using EHR data. NAACL HLT 2019 2019: 46\u201354.","journal-title":"NAACL HLT 2019"},{"key":"e_1_3_4_24_2","doi-asserted-by":"crossref","unstructured":"Lins RD Oliveira H Cabral L et al. The CNN-Corpus: a large textual corpus for single-document extractive summarization. In: Proceedings of the ACM symposium on document engineering 2019 2019 pp.1\u201310.","DOI":"10.1145\/3342558.3345388"},{"key":"e_1_3_4_25_2","doi-asserted-by":"crossref","unstructured":"Sosea T Zhan H Li JJ et al. Unsupervised extractive summarization of emotion triggers. In: Proceedings of the 61st annual meeting of the association for computational linguistics (Volume 1: Long Papers) 2023 pp.9550\u20139569. Association for Computational Linguistics.","DOI":"10.18653\/v1\/2023.acl-long.531"},{"key":"e_1_3_4_26_2","first-page":"397","article-title":"Text summarization techniques: a brief survey","volume":"8","author":"Allahyari M","year":"2017","unstructured":"Allahyari M, Pouriyeh S, Assefi M, et al. Text summarization techniques: a brief survey. Int J Adv Comput Sci Appl 2017; 8: 397\u2013405.","journal-title":"Int J Adv Comput Sci Appl"},{"key":"e_1_3_4_27_2","doi-asserted-by":"crossref","unstructured":"Chern Ic Wang Z Das S et al. Improving factuality of abstractive summarization via contrastive reward learning. In: Proceedings of the 3rd workshop on trustworthy natural language processing (TrustNLP 2023) 2023 pp.55\u201360.","DOI":"10.18653\/v1\/2023.trustnlp-1.6"},{"key":"e_1_3_4_28_2","doi-asserted-by":"crossref","unstructured":"Pu D Wang Y Demberg V. Incorporating distributions of discourse structure for long document abstractive summarization. In: Proceedings of the 61st annual meeting of the association for computational linguistics (Volume 1: Long Papers) 2023 pp.5574\u20135590. Association for Computational Linguistics.","DOI":"10.18653\/v1\/2023.acl-long.306"},{"key":"e_1_3_4_29_2","doi-asserted-by":"crossref","unstructured":"Lewis M Liu Y Goyal N et al. BART: denoising sequence-to-sequence pre-training for natural language generation translation and comprehension. In: Proceedings of the 58th annual meeting of the association for computational linguistics 2020 pp.7871\u20137880.","DOI":"10.18653\/v1\/2020.acl-main.703"},{"key":"e_1_3_4_30_2","first-page":"9","article-title":"Language models are unsupervised multitask learners","volume":"1","author":"Radford A","year":"2019","unstructured":"Radford A, Wu J, Child R, et al. Language models are unsupervised multitask learners. OpenAI Blog 2019; 1: 9.","journal-title":"OpenAI Blog"},{"key":"e_1_3_4_31_2","doi-asserted-by":"crossref","unstructured":"Nentidis A. Overview of bioasq 2021: the ninth bioasq challenge on large-scale biomedical semantic indexing and question answering. In: Experimental IR meets multilinguality multimodality and interaction: 12th International conference of the CLEF association CLEF 2021 virtual event September 21\u201324 2021 proceedings volume 12880 2021 pp.239. Springer Nature.","DOI":"10.1007\/978-3-030-85251-1_18"},{"key":"e_1_3_4_32_2","first-page":"605","article-title":"Generating (factual?) narrative summaries of rcts: experiments with neural multi-document summarization","volume":"2021","author":"Wallace BC","year":"2021","unstructured":"Wallace BC, Saha S, Soboczenski F, et al. Generating (factual?) narrative summaries of rcts: experiments with neural multi-document summarization. AMIA Summit Transl Sci Proc 2021; 2021: 605.","journal-title":"AMIA Summit Transl Sci Proc"},{"key":"e_1_3_4_33_2","unstructured":"Team G Anil R Borgeaud S et al. Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805 2023."},{"key":"e_1_3_4_34_2","doi-asserted-by":"crossref","unstructured":"Lee JY Dernoncourt F. Sequential short-text classification with recurrent and convolutional neural networks. In: Proceedings of NAACL-HLT 2016 pp.515\u2013520.","DOI":"10.18653\/v1\/N16-1062"},{"key":"e_1_3_4_35_2","doi-asserted-by":"crossref","unstructured":"Wang J Wang Z Zhang D et al. Combining knowledge with deep convolutional neural networks for short text classification. In: Proceedings of the 26th international joint conference on artificial intelligence 2017 pp.2915\u20132921.","DOI":"10.24963\/ijcai.2017\/406"},{"key":"e_1_3_4_36_2","unstructured":"Kenton JDM -WC Toutanova LK. BERT: Pre-training of deep bidirectional transformers for language understanding. In: Proceedings of NAACL-HLT 2019 pp.4171\u20134186."},{"key":"e_1_3_4_37_2","unstructured":"Liu Y Ott M Goyal N et al. ROBERTA: a robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692 2019."},{"key":"e_1_3_4_38_2","doi-asserted-by":"crossref","unstructured":"Sun C Qiu X Xu Y et al. How to fine-tune bert for text classification? In: Chinese computational linguistics: 18th China national conference CCL 2019 kunming China October 18\u201320 2019 proceedings 18 2019 pp.194\u2013206. Springer.","DOI":"10.1007\/978-3-030-32381-3_16"},{"key":"e_1_3_4_39_2","doi-asserted-by":"crossref","unstructured":"Wang Z Ng P Ma X et al. Multi-passage bert: A globally normalized bert model for open-domain question answering. In: Proceedings of the 2019 conference on empirical methods in natural language processing and the 9th international joint conference on natural language processing (EMNLP-IJCNLP) 2019 pp.5878\u20135882.","DOI":"10.18653\/v1\/D19-1599"},{"key":"e_1_3_4_40_2","unstructured":"Beltagy I Peters ME Cohan A. Longformer: the long-document transformer. arXiv preprint arXiv:2004.05150 2020."},{"key":"e_1_3_4_41_2","first-page":"5998","article-title":"Attention is all you need","volume":"30","author":"Vaswani A","year":"2017","unstructured":"Vaswani A, Shazeer N, Parmar N, et al. Attention is all you need. Adv Neural Inf Process Syst 2017; 30: 5998\u20136008.","journal-title":"Adv Neural Inf Process Syst"},{"key":"e_1_3_4_42_2","unstructured":"He P Liu X Gao J et al. Deberta: decoding-enhanced bert with disentangled attention. In: International conference on learning representations 2020."},{"key":"e_1_3_4_43_2","unstructured":"Zhang J Zhao Y Saleh M et al. Pegasus: Pre-training with extracted gap-sentences for abstractive summarization. In: International conference on machine learning 2020a pp.11328\u201311339. PMLR."},{"key":"e_1_3_4_44_2","first-page":"8024","article-title":"Pytorch: an imperative style, high-performance deep learning library","volume":"32","author":"Paszke A","year":"2019","unstructured":"Paszke A, Gross S, Massa F, et al. Pytorch: an imperative style, high-performance deep learning library. Adv Neural Inf Process Syst 2019; 32: 8024\u20138035.","journal-title":"Adv Neural Inf Process Syst"},{"key":"e_1_3_4_45_2","unstructured":"Loshchilov I Hutter F. Fixing weight decay regularization in adam 2018."}],"container-title":["Intelligent Data Analysis: An International Journal"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/1088467X251330470","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/full-xml\/10.1177\/1088467X251330470","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/1088467X251330470","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,29]],"date-time":"2026-04-29T09:21:15Z","timestamp":1777454475000},"score":1,"resource":{"primary":{"URL":"https:\/\/journals.sagepub.com\/doi\/10.1177\/1088467X251330470"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,4]]},"references-count":44,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2026,1]]}},"alternative-id":["10.1177\/1088467X251330470"],"URL":"https:\/\/doi.org\/10.1177\/1088467x251330470","relation":{},"ISSN":["1088-467X","1571-4128"],"issn-type":[{"value":"1088-467X","type":"print"},{"value":"1571-4128","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,4]]}}}