{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,5]],"date-time":"2026-08-05T15:23:49Z","timestamp":1785943429766,"version":"3.56.0"},"reference-count":85,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2025,3,10]],"date-time":"2025-03-10T00:00:00Z","timestamp":1741564800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Knowl. Discov. Data"],"published-print":{"date-parts":[[2025,4,30]]},"abstract":"<jats:p>This research pioneers the use of fine-tuned Large Language Models (LLMs) to automate Systematic Literature Reviews (SLRs), presenting a significant and novel contribution in integrating AI to enhance academic research methodologies. Our study employed advanced fine-tuning methodologies on open sourced LLMs, applying textual data mining techniques to automate the knowledge discovery and synthesis phases of an SLR process, thus demonstrating a practical and efficient approach for extracting and analyzing high-quality information from large academic datasets. The results maintained high fidelity in factual accuracy in LLM responses, and were validated through the replication of an existing PRISMA-conforming SLR. Our research proposed solutions for mitigating LLM hallucination and proposed mechanisms for tracking LLM responses to their sources of information, thus demonstrating how this approach can meet the rigorous demands of scholarly research. The findings ultimately confirmed the potential of fine-tuned LLMs in streamlining various labor-intensive processes of conducting literature reviews. As a scalable proof-of-concept, this study highlights the broad applicability of our approach across multiple research domains. The potential demonstrated here advocates for updates to PRISMA reporting guidelines, incorporating AI-driven processes to ensure methodological transparency and reliability in future SLRs. This study broadens the appeal of AI-enhanced tools across various academic and research fields, demonstrating how to conduct comprehensive and accurate literature reviews with more efficiency in the face of ever-increasing volumes of academic studies while maintaining high standards.<\/jats:p>","DOI":"10.1145\/3715964","type":"journal-article","created":{"date-parts":[[2025,1,31]],"date-time":"2025-01-31T11:24:36Z","timestamp":1738322676000},"page":"1-39","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":54,"title":["Automating Research Synthesis with Domain-Specific Large Language Model Fine-Tuning"],"prefix":"10.1145","volume":"19","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-9416-1435","authenticated-orcid":false,"given":"Teo","family":"Susnjak","sequence":"first","affiliation":[{"name":"Department of Computer Science and IT, Massey University College of Sciences, Auckland, New Zealand"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-9477-2265","authenticated-orcid":false,"given":"Peter","family":"Hwang","sequence":"additional","affiliation":[{"name":"Department of Computer Science and IT, Massey University College of Sciences, Auckland, New Zealand"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0683-436X","authenticated-orcid":false,"given":"Napoleon","family":"Reyes","sequence":"additional","affiliation":[{"name":"Department of Computer Science and IT, Massey University College of Sciences, Auckland, New Zealand"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7648-285X","authenticated-orcid":false,"given":"Andre L. C.","family":"Barczak","sequence":"additional","affiliation":[{"name":"Bond University School of Information Technology, Gold Coast, Australia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0836-4266","authenticated-orcid":false,"given":"Timothy","family":"McIntosh","sequence":"additional","affiliation":[{"name":"Cyberoo Pty Ltd., Surry Hills, Australia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0701-0204","authenticated-orcid":false,"given":"Surangika","family":"Ranathunga","sequence":"additional","affiliation":[{"name":"Department of Computer Science and IT, Massey University College of Sciences, Auckland, New Zealand"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,3,10]]},"reference":[{"key":"e_1_3_2_2_2","unstructured":"Renat Aksitov Chung-Ching Chang David Reitter Siamak Shakeri and Yunhsuan Sung. 2023. Characterizing attribution and fluency tradeoffs for retrieval-augmented large language models. arXiv:2302.05578. Retrieved from https:\/\/arxiv.org\/abs\/2302.05578"},{"key":"e_1_3_2_3_2","doi-asserted-by":"publisher","DOI":"10.3390\/systems11070351"},{"key":"e_1_3_2_4_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.jksuci.2020.04.020"},{"key":"e_1_3_2_5_2","doi-asserted-by":"publisher","unstructured":"Akari Asai Zeqiu Wu Yizhong Wang Avirup Sil and Hannaneh Hajishirzi. 2023. Self-RAG: Learning to retrieve generate and critique through self-reflection. arXiv:2310.11511. DOI: 10.48550\/arXiv.2310.11511","DOI":"10.48550\/arXiv.2310.11511"},{"key":"e_1_3_2_6_2","article-title":"Cheap, quick, and rigorous: Artificial intelligence and the systematic literature review","author":"Atkinson Cameron F.","year":"2023","unstructured":"Cameron F. Atkinson. 2023. Cheap, quick, and rigorous: Artificial intelligence and the systematic literature review. Social Science Computer Review 42 (2023), 08944393231196281.","journal-title":"Social Science Computer Review"},{"key":"e_1_3_2_7_2","doi-asserted-by":"publisher","DOI":"10.1145\/3485127"},{"key":"e_1_3_2_8_2","doi-asserted-by":"crossref","unstructured":"Francisco Bolanos Angelo Salatino Francesco Osborne and Enrico Motta. 2024. Artificial intelligence for literature reviews: Opportunities and challenges. arXiv:2402.08565. Retrieved from https:\/\/arxiv.org\/abs\/2402.08565","DOI":"10.1007\/s10462-024-10902-3"},{"key":"e_1_3_2_9_2","unstructured":"Tom B. Brown Benjamin Mann Nick Ryder Melanie Subbiah Jared Kaplan Prafulla Dhariwal Arvind Neelakantan Pranav Shyam Girish Sastry Amanda Askell et al. 2020. Language models are few-shot learners. arXiv:2005.14165. Retrieved from https:\/\/arxiv.org\/abs\/2005.14165"},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","DOI":"10.1002\/jrsm.1399"},{"key":"e_1_3_2_11_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.jbi.2016.10.014"},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.1109\/WEEF-GEDC59520.2023.10344098"},{"key":"e_1_3_2_13_2","doi-asserted-by":"publisher","unstructured":"Yupeng Chang Xu Wang Jindong Wang Yuan Wu Linyi Yang Kaijie Zhu Hao Chen Xiaoyuan Yi Cunxiang Wang Yidong Wang et al. 2024. A survey on evaluation of large language models. ACM Transactions on Intelligent Systems and Technology 15 3 Article 39 (Mar. 2024) 45 pages. DOI: 10.1145\/3641289","DOI":"10.1145\/3641289"},{"key":"e_1_3_2_14_2","doi-asserted-by":"publisher","DOI":"10.1371\/journal.pone.0223994"},{"key":"e_1_3_2_15_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v38i16.29728"},{"key":"e_1_3_2_16_2","doi-asserted-by":"publisher","DOI":"10.47909\/ijsmc.52"},{"key":"e_1_3_2_17_2","doi-asserted-by":"publisher","DOI":"10.1007\/s00607-023-01181-x"},{"key":"e_1_3_2_18_2","unstructured":"Tim Dettmers Artidoro Pagnoni Ari Holtzman and Luke Zettlemoyer. 2023. Qlora: Efficient finetuning of quantized LLMs. arXiv:2305.14314. Retrieved from https:\/\/arxiv.org\/abs\/2305.14314"},{"key":"e_1_3_2_19_2","doi-asserted-by":"publisher","DOI":"10.1177\/135581960501000110"},{"key":"e_1_3_2_20_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.jclinepi.2008.10.016"},{"key":"e_1_3_2_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/APSEC.2017.10"},{"key":"e_1_3_2_22_2","doi-asserted-by":"publisher","DOI":"10.14778\/3611479.3611527"},{"key":"e_1_3_2_23_2","doi-asserted-by":"publisher","DOI":"10.1136\/bmjgh-2018-000882"},{"key":"e_1_3_2_24_2","doi-asserted-by":"publisher","DOI":"10.1186\/1748-5908-5-56"},{"key":"e_1_3_2_25_2","doi-asserted-by":"publisher","DOI":"10.1111\/j.1471-1842.2009.00848.x"},{"key":"e_1_3_2_26_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.bjps.2023.03.004"},{"key":"e_1_3_2_27_2","doi-asserted-by":"publisher","DOI":"10.1038\/nature25753"},{"key":"e_1_3_2_28_2","doi-asserted-by":"crossref","unstructured":"Suchin Gururangan Ana Marasovi\u0107 Swabha Swayamdipta Kyle Lo Iz Beltagy Doug Downey and Noah A. Smith. 2020. Don\u2019t stop pretraining: Adapt language models to domains and tasks. arXiv:2004.10964. Retrieved from https:\/\/arxiv.org\/abs\/2004.10964","DOI":"10.18653\/v1\/2020.acl-main.740"},{"key":"e_1_3_2_29_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.envint.2018.02.018"},{"key":"e_1_3_2_30_2","doi-asserted-by":"publisher","DOI":"10.1136\/bmjgh-2018-000858"},{"key":"e_1_3_2_31_2","article-title":"Methods for using Bing\u2019s AI-powered search engine for data extraction for a systematic review","author":"Hill James Edward","year":"2023","unstructured":"James Edward Hill, Catherine Harris, and Andrew Clegg. 2023. Methods for using Bing\u2019s AI-powered search engine for data extraction for a systematic review. Research Synthesis Methods (2023).","journal-title":"Research Synthesis Methods"},{"key":"e_1_3_2_32_2","doi-asserted-by":"publisher","DOI":"10.1186\/s13643-017-0454-2"},{"key":"e_1_3_2_33_2","doi-asserted-by":"publisher","unstructured":"Xinying Hou Yanjie Zhao Yue Liu Zhou Yang Kailong Wang Li Li Xiapu Luo David Lo John C. Grundy and Haoyu Wang. 2023. Large language models for software engineering: A systematic literature review. arXiv:2308.10620. DOI: 10.48550\/arXiv.2308.10620","DOI":"10.48550\/arXiv.2308.10620"},{"key":"e_1_3_2_34_2","first-page":"2790","volume-title":"International Conference on Machine Learning","author":"Houlsby Neil","year":"2019","unstructured":"Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. 2019. Parameter-efficient transfer learning for NLP. In International Conference on Machine Learning. PMLR, 2790\u20132799."},{"key":"e_1_3_2_35_2","doi-asserted-by":"crossref","unstructured":"Jeremy Howard and Sebastian Ruder. 2018. Universal language model fine-tuning for text classification. arXiv:1801.06146. Retrieved from https:\/\/arxiv.org\/abs\/1801.06146","DOI":"10.18653\/v1\/P18-1031"},{"key":"e_1_3_2_36_2","unstructured":"Edward J. Hu Yelong Shen Phillip Wallis Zeyuan Allen-Zhu Yuanzhi Li Shean Wang Lu Wang and Weizhu Chen. 2021. LoRA: Low-rank adaptation of large language models. arXiv:2106.09685. Retrieved from https:\/\/arxiv.org\/abs\/2106.09685"},{"key":"e_1_3_2_37_2","unstructured":"Neel Jain Ping-yeh Chiang Yuxin Wen John Kirchenbauer Hong-Min Chu Gowthami Somepalli Brian R. Bartoldson Bhavya Kailkhura Avi Schwarzschild Aniruddha Saha et al. 2023. NEFTune: Noisy embeddings improve instruction finetuning. arXiv:2310.05914. Retrieved from https:\/\/arxiv.org\/abs\/2310.05914"},{"key":"e_1_3_2_38_2","unstructured":"Albert Q. Jiang Alexandre Sablayrolles Arthur Mensch Chris Bamford Devendra Singh Chaplot Diego de las Casas Florian Bressand Gianna Lengyel Guillaume Lample Lucile Saulnier et al. 2023. Mistral 7B. arXiv:2310.06825. Retrieved from https:\/\/arxiv.org\/abs\/2310.06825"},{"key":"e_1_3_2_39_2","doi-asserted-by":"publisher","DOI":"10.1186\/s13643-015-0066-7"},{"key":"e_1_3_2_40_2","unstructured":"Adam Tauman Kalai and Santosh S. Vempala. 2023. Calibrated language models must hallucinate. arXiv:2311.14648. Retrieved from https:\/\/arxiv.org\/abs\/2311.14648"},{"key":"e_1_3_2_41_2","doi-asserted-by":"publisher","unstructured":"Qusai Khraisha Sophie Put Johanna Kappenberg Azza Warraitch and Kristin Hadfield. 2023. Can large language models replace humans in the systematic review process? Evaluating GPT-4\u2019s efficacy in screening and extracting data from peer-reviewed and grey literature in multiple languages. arXiv:2310.17526. DOI: 10.48550\/arXiv.2310.17526","DOI":"10.48550\/arXiv.2310.17526"},{"key":"e_1_3_2_42_2","doi-asserted-by":"publisher","DOI":"10.5530\/bems.9.1.5"},{"key":"e_1_3_2_43_2","unstructured":"Patrick Lewis Ethan Perez Aleksandara Piktus Fabio Petroni Vladimir Karpukhin Naman Goyal Heinrich Kuttler M. Lewis Wen tau Yih Tim Rockt\u00e4schel et al. 2020. Retrieval-augmented generation for knowledge-intensive NLP tasks. arXiv:2005.11401. Retrieved from https:\/\/arxiv.org\/abs\/2005.11401"},{"key":"e_1_3_2_44_2","doi-asserted-by":"publisher","unstructured":"Yifan Li Yifan Du Kun Zhou Jinpeng Wang Wayne Xin Zhao and Ji-Rong Wen. 2023. Evaluating object hallucination in large vision-language models. arXiv:2305.10355. DOI: 10.48550\/arXiv.2305.10355","DOI":"10.48550\/arXiv.2305.10355"},{"key":"e_1_3_2_45_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.jclinepi.2009.06.006"},{"key":"e_1_3_2_46_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.acl-long.229"},{"key":"e_1_3_2_47_2","doi-asserted-by":"publisher","DOI":"10.1002\/jrsm.1106"},{"key":"e_1_3_2_48_2","doi-asserted-by":"publisher","DOI":"10.1186\/s13643-019-1074-9"},{"key":"e_1_3_2_49_2","doi-asserted-by":"publisher","DOI":"10.1177\/136140969900400111"},{"key":"e_1_3_2_50_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.cose.2023.103424"},{"key":"e_1_3_2_51_2","doi-asserted-by":"publisher","DOI":"10.1109\/TAI.2023.3332837"},{"key":"e_1_3_2_52_2","article-title":"The inadequacy of reinforcement learning from human feedback-radicalizing large language models via semantic vulnerabilities","author":"McIntosh Timothy R.","year":"2024","unstructured":"Timothy R. McIntosh, Teo Susnjak, Tong Liu, Paul Watters, and Malka N. Halgamuge. 2024. The inadequacy of reinforcement learning from human feedback-radicalizing large language models via semantic vulnerabilities. IEEE Transactions on Cognitive and Developmental Systems\u00a016, 4 (2024), 1561\u20131574.","journal-title":"IEEE Transactions on Cognitive and Developmental Systems\u00a016, 4 (2024), 1561\u20131574"},{"key":"e_1_3_2_53_2","doi-asserted-by":"publisher","DOI":"10.1136\/ebmental-2019-300129"},{"key":"e_1_3_2_54_2","doi-asserted-by":"publisher","DOI":"10.1145\/3605943"},{"key":"e_1_3_2_55_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11096-016-0289-2"},{"key":"e_1_3_2_56_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICAPAI55158.2022.9801564"},{"key":"e_1_3_2_57_2","doi-asserted-by":"publisher","DOI":"10.1093\/asj\/sjad093"},{"key":"e_1_3_2_58_2","doi-asserted-by":"publisher","DOI":"10.7759\/cureus.43023"},{"key":"e_1_3_2_59_2","doi-asserted-by":"publisher","DOI":"10.17705\/1CAIS.03743"},{"key":"e_1_3_2_60_2","doi-asserted-by":"publisher","DOI":"10.1111\/jebm.12265"},{"key":"e_1_3_2_61_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.ijsu.2021.105906"},{"key":"e_1_3_2_62_2","doi-asserted-by":"publisher","DOI":"10.1038\/s41591-023-02366-9"},{"key":"e_1_3_2_63_2","doi-asserted-by":"publisher","DOI":"10.1186\/s13643-023-02243-z"},{"key":"e_1_3_2_64_2","doi-asserted-by":"publisher","DOI":"10.5195\/jmla.2021.962"},{"key":"e_1_3_2_65_2","doi-asserted-by":"publisher","DOI":"10.1146\/annurev-psych-010418-102803"},{"key":"e_1_3_2_66_2","doi-asserted-by":"publisher","DOI":"10.1002\/asi.24851"},{"key":"e_1_3_2_67_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-981-99-3243-6_7"},{"key":"e_1_3_2_68_2","unstructured":"Teo Susnjak. 2023. PRISMA-DFLLM: An extension of PRISMA for systematic literature reviews using domain-specific finetuned large language models. arXiv:2306.14905. Retrieved from https:\/\/arxiv.org\/abs\/2306.14905"},{"key":"e_1_3_2_69_2","doi-asserted-by":"publisher","DOI":"10.1186\/s41239-021-00313-7"},{"key":"e_1_3_2_70_2","doi-asserted-by":"publisher","unstructured":"James Thorne Andreas Vlachos Christos Christodoulopoulos and Arpit Mittal. 2018. FEVER: A large-scale dataset for Fact Extraction and VERification. arXiv:1803.05355. DOI: 10.18653\/v1\/N18-1074","DOI":"10.18653\/v1\/N18-1074"},{"key":"e_1_3_2_71_2","doi-asserted-by":"publisher","DOI":"10.1177\/1534484305278283"},{"key":"e_1_3_2_72_2","doi-asserted-by":"publisher","DOI":"10.1177\/1534484316671606"},{"key":"e_1_3_2_73_2","doi-asserted-by":"publisher","DOI":"10.1186\/2046-4053-3-74"},{"key":"e_1_3_2_74_2","doi-asserted-by":"publisher","DOI":"10.1002\/jrsm.1335"},{"key":"e_1_3_2_75_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.infsof.2021.106589"},{"key":"e_1_3_2_76_2","article-title":"Attention is all you need","volume":"30","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems, Vol. 30.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_77_2","doi-asserted-by":"publisher","DOI":"10.1177\/02683962211048201"},{"key":"e_1_3_2_78_2","doi-asserted-by":"publisher","DOI":"10.1016\/J.EMJ.2020.09.007"},{"key":"e_1_3_2_79_2","doi-asserted-by":"publisher","DOI":"10.1177\/0739456X17723971"},{"key":"e_1_3_2_80_2","doi-asserted-by":"publisher","DOI":"10.1016\/S1976-1317(08)60041-9"},{"key":"e_1_3_2_81_2","doi-asserted-by":"publisher","DOI":"10.1145\/3649506"},{"key":"e_1_3_2_82_2","unstructured":"Shunyu Yao Jeffrey Zhao Dian Yu Nan Du I. Shafran Karthik Narasimhan and Yuan Cao. 2022. ReAct: Synergizing reasoning and acting in language models. arXiv:2210.03629. Retrieved from https:\/\/arxiv.org\/abs\/2210.03629"},{"key":"e_1_3_2_83_2","unstructured":"Muru Zhang Ofir Press William Merrill Alisa Liu and Noah A. Smith. 2023. How language model hallucinations can snowball. arXiv:2305.13534. Retrieved from https:\/\/arxiv.org\/abs\/2305.13534"},{"key":"e_1_3_2_84_2","doi-asserted-by":"publisher","DOI":"10.1145\/3468889"},{"key":"e_1_3_2_85_2","doi-asserted-by":"publisher","DOI":"10.1111\/J.1365-2648.2006.03721.X"},{"key":"e_1_3_2_86_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-48858-0_25"}],"container-title":["ACM Transactions on Knowledge Discovery from Data"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3715964","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3715964","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,9,30]],"date-time":"2025-09-30T16:21:57Z","timestamp":1759249317000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3715964"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,3,10]]},"references-count":85,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2025,4,30]]}},"alternative-id":["10.1145\/3715964"],"URL":"https:\/\/doi.org\/10.1145\/3715964","relation":{},"ISSN":["1556-4681","1556-472X"],"issn-type":[{"value":"1556-4681","type":"print"},{"value":"1556-472X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,3,10]]},"assertion":[{"value":"2024-05-09","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-01-25","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-03-10","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}