{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,4]],"date-time":"2026-06-04T20:12:51Z","timestamp":1780603971743,"version":"3.54.1"},"reference-count":0,"publisher":"IOS Press","isbn-type":[{"value":"9781643682648","type":"print"},{"value":"9781643682655","type":"electronic"}],"license":[{"start":{"date-parts":[[2022,6,6]],"date-time":"2022-06-06T00:00:00Z","timestamp":1654473600000},"content-version":"unspecified","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by-nc\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2022,6,6]]},"abstract":"<jats:p>Sample size is an important indicator of the power of randomized controlled trials (RCTs). In this paper, we designed a total sample size extractor using a combination of syntactic and machine learning methods, and evaluated it on 300 Covid-19 abstracts (Covid-Set) and 100 generic RCT abstracts (General-Set). To improve the performance, we applied transfer learning from a large public corpus of annotated abstracts. We achieved an average F1 score of 0.73 on the Covid-Set testing set, and 0.60 on the General-Set using exact matches. The F1 scores for loose matches on both datasets were over 0.74. Compared with the state-of-the-art tool, our extractor reports total sample sizes directly and improved F1 scores by at least 4% without transfer learning. We demonstrated that transfer learning improved the sample size extraction accuracy and minimized human labor on annotations.<\/jats:p>","DOI":"10.3233\/shti220151","type":"book-chapter","created":{"date-parts":[[2022,6,7]],"date-time":"2022-06-07T09:33:04Z","timestamp":1654594384000},"source":"Crossref","is-referenced-by-count":2,"title":["A Sample Size Extractor for RCT Reports"],"prefix":"10.3233","author":[{"given":"Fengyang","family":"Lin","sequence":"first","affiliation":[{"name":"Department of Biomedical Informatics, Columbia University, New York, NY, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Hao","family":"Liu","sequence":"additional","affiliation":[{"name":"Department of Biomedical Informatics, Columbia University, New York, NY, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Paul","family":"Moon","sequence":"additional","affiliation":[{"name":"College of Physicians and Surgeons: Institute of Human Nutrition, Columbia University, New York, NY, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Chunhua","family":"Weng","sequence":"additional","affiliation":[{"name":"Department of Biomedical Informatics, Columbia University, New York, NY, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"7437","container-title":["Studies in Health Technology and Informatics","MEDINFO 2021: One World, One Health \u2013 Global Partnership for Digital Innovation"],"original-title":[],"link":[{"URL":"https:\/\/ebooks.iospress.nl\/pdf\/doi\/10.3233\/SHTI220151","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,6,7]],"date-time":"2022-06-07T09:33:05Z","timestamp":1654594385000},"score":1,"resource":{"primary":{"URL":"https:\/\/ebooks.iospress.nl\/doi\/10.3233\/SHTI220151"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,6,6]]},"ISBN":["9781643682648","9781643682655"],"references-count":0,"URL":"https:\/\/doi.org\/10.3233\/shti220151","relation":{},"ISSN":["0926-9630","1879-8365"],"issn-type":[{"value":"0926-9630","type":"print"},{"value":"1879-8365","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,6,6]]}}}