{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,24]],"date-time":"2026-07-24T04:14:49Z","timestamp":1784866489965,"version":"3.55.0"},"reference-count":35,"publisher":"Oxford University Press (OUP)","license":[{"start":{"date-parts":[[2025,2,5]],"date-time":"2025-02-05T00:00:00Z","timestamp":1738713600000},"content-version":"vor","delay-in-days":35,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/100010665","name":"H2020 Marie Sklodowska-Curie Actions","doi-asserted-by":"publisher","award":["945405"],"award-info":[{"award-number":["945405"]}],"id":[{"id":"10.13039\/100010665","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100010269","name":"Wellcome Trust","doi-asserted-by":"publisher","award":["218302\/Z\/19\/Z"],"award-info":[{"award-number":["218302\/Z\/19\/Z"]}],"id":[{"id":"10.13039\/100010269","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100010665","name":"H2020 Marie Sklodowska-Curie Actions","doi-asserted-by":"publisher","award":["945405"],"award-info":[{"award-number":["945405"]}],"id":[{"id":"10.13039\/100010665","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100010269","name":"Wellcome Trust","doi-asserted-by":"publisher","award":["218302\/Z\/19\/Z"],"award-info":[{"award-number":["218302\/Z\/19\/Z"]}],"id":[{"id":"10.13039\/100010269","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2025,2,5]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>Curation of literature in life sciences is a growing challenge. The continued increase in the rate of publication, coupled with the relatively fixed number of curators worldwide, presents a major challenge to developers of biomedical knowledgebases. Very few knowledgebases have resources to scale to the whole relevant literature and all have to prioritize their efforts.<\/jats:p>\n               <jats:p>In this work, we take a first step to alleviating the lack of curator time in RNA science by generating summaries of literature for noncoding RNAs using large language models (LLMs). We demonstrate that high-quality, factually accurate summaries with accurate references can be automatically generated from the literature using a commercial LLM and a chain of prompts and checks. Manual assessment was carried out for a subset of summaries, with the majority being rated extremely high quality.<\/jats:p>\n               <jats:p>We apply our tool to a selection of &amp;gt;4600 ncRNAs and make the generated summaries available via the RNAcentral resource. We conclude that automated literature summarization is feasible with the current generation of LLMs, provided that careful prompting and automated checking are applied.<\/jats:p>\n               <jats:p>Database URL: https:\/\/rnacentral.org\/<\/jats:p>","DOI":"10.1093\/database\/baaf006","type":"journal-article","created":{"date-parts":[[2025,1,17]],"date-time":"2025-01-17T15:16:31Z","timestamp":1737126991000},"source":"Crossref","is-referenced-by-count":12,"title":["LitSumm: large language models for literature summarization of noncoding RNAs"],"prefix":"10.1093","volume":"2025","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-8297-0953","authenticated-orcid":false,"given":"Andrew","family":"Green","sequence":"first","affiliation":[{"name":"European Molecular Biology Laboratory, European Bioinformatics Institute (EMBL-EBI), Wellcome Genome Campus , Hinxton CB10 1SD,","place":["UK"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9572-273X","authenticated-orcid":false,"given":"Carlos Eduardo","family":"Ribas","sequence":"additional","affiliation":[{"name":"European Molecular Biology Laboratory, European Bioinformatics Institute (EMBL-EBI), Wellcome Genome Campus , Hinxton CB10 1SD,","place":["UK"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8457-4455","authenticated-orcid":false,"given":"Nancy","family":"Ontiveros-Palacios","sequence":"additional","affiliation":[{"name":"European Molecular Biology Laboratory, European Bioinformatics Institute (EMBL-EBI), Wellcome Genome Campus , Hinxton CB10 1SD,","place":["UK"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6043-807X","authenticated-orcid":false,"given":"Sam","family":"Griffiths-Jones","sequence":"additional","affiliation":[{"name":"School of Biological Sciences, Faculty of Medicine, Biology and Health, Michael Smith Building, The University of Manchester , Manchester M13 9NT,","place":["UK"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7279-2682","authenticated-orcid":false,"given":"Anton I","family":"Petrov","sequence":"additional","affiliation":[{"name":"Riboscope Ltd , 23 King St, Cambridge CB1 1AH,","place":["UK"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6982-4660","authenticated-orcid":false,"given":"Alex","family":"Bateman","sequence":"additional","affiliation":[{"name":"European Molecular Biology Laboratory, European Bioinformatics Institute (EMBL-EBI), Wellcome Genome Campus , Hinxton CB10 1SD,","place":["UK"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6497-2883","authenticated-orcid":false,"given":"Blake","family":"Sweeney","sequence":"additional","affiliation":[{"name":"European Molecular Biology Laboratory, European Bioinformatics Institute (EMBL-EBI), Wellcome Genome Campus , Hinxton CB10 1SD,","place":["UK"]}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"286","published-online":{"date-parts":[[2025,2,5]]},"reference":[{"key":"2025092510524536000_R1","doi-asserted-by":"crossref","DOI":"10.1371\/journal.pbio.2002846","article-title":"Biocuration: distilling data into knowledge","volume":"16","author":"International Society for Biocuration","year":"2018","journal-title":"PLoS Biol"},{"key":"2025092510524536000_R2","doi-asserted-by":"crossref","first-page":"D523","DOI":"10.1093\/nar\/gkac1052","article-title":"UniProt: the Universal Protein Knowledgebase in 2023","volume":"51","author":"Bateman","year":"2023","journal-title":"Nucleic Acids Res"},{"key":"2025092510524536000_R3","doi-asserted-by":"crossref","DOI":"10.1093\/genetics\/iyac191","article-title":"Saccharomyces genome database update: server architecture, pan-genome nomenclature, and external resources","volume":"224","author":"Wong","year":"2023","journal-title":"Genetics"},{"key":"2025092510524536000_R4","doi-asserted-by":"crossref","first-page":"D899","DOI":"10.1093\/nar\/gkaa1026","article-title":"FlyBase: updates to the Drosophila melanogaster knowledge base","volume":"49","author":"Larkin","year":"2021","journal-title":"Nucleic Acids Res"},{"key":"2025092510524536000_R5","article-title":"Gene set summarization using large language models","volume-title":"ArXiv","author":"Joachimiak","year":"2023"},{"key":"2025092510524536000_R6","article-title":"Generative artificial intelligence GPT-4 accelerates knowledge mining and machine learning for synthetic biology. Generative artificial intelligence GPT-4 accelerates knowledge mining and machine learning for synthetic biology","volume-title":"bioRxiv","author":"Xiao","year":"2023"},{"key":"2025092510524536000_R7","article-title":"A comprehensive benchmark study on biomedical text generation and mining with ChatGPT. A comprehensive benchmark study on biomedical text generation and mining with ChatGPT","volume-title":"bioRxiv","author":"Chen","year":"2023"},{"key":"2025092510524536000_R8","doi-asserted-by":"crossref","first-page":"D192","DOI":"10.1093\/nar\/gkaa1047","article-title":"Rfam 14: expanded coverage of metagenomic, viral and microRNA families","volume":"49","author":"Kalvari","year":"2021","journal-title":"Nucleic Acids Res"},{"key":"2025092510524536000_R9","doi-asserted-by":"crossref","first-page":"D212","DOI":"10.1093\/nar\/gkaa921","article-title":"RNAcentral 2021: secondary structure integration, improved sequence search and new member databases","volume":"49","author":"RNAcentral Consortium","year":"2021","journal-title":"Nucleic Acids Res"},{"key":"2025092510524536000_R10","doi-asserted-by":"crossref","first-page":"D1003","DOI":"10.1093\/nar\/gkac888","article-title":"Genenames.org: the HGNC resources in 2023","volume":"51","author":"Seal","year":"2023","journal-title":"Nucleic Acids Res"},{"key":"2025092510524536000_R11","doi-asserted-by":"crossref","first-page":"D155","DOI":"10.1093\/nar\/gky1141","article-title":"miRBase: from microRNA sequences to function","volume":"47","author":"Kozomara","year":"2019","journal-title":"Nucleic Acids Res"},{"key":"2025092510524536000_R12","doi-asserted-by":"crossref","first-page":"D204","DOI":"10.1093\/nar\/gkab1101","article-title":"MirGeneDB 2.1: toward a complete sampling of all major animal phyla","volume":"50","author":"Fromm","year":"2022","journal-title":"Nucleic Acids Res"},{"key":"2025092510524536000_R13","doi-asserted-by":"crossref","first-page":"D291","DOI":"10.1093\/nar\/gkac835","article-title":"snoDB 2.0: an enhanced interactive database, specializing in human snoRNAs","volume":"51","author":"Bergeron","year":"2023","journal-title":"Nucleic Acids Res"},{"key":"2025092510524536000_R14","article-title":"Attention is all you need. Attention is all you need","volume-title":"arXiv [cs.CL]","author":"Vaswani","year":"2017"},{"key":"2025092510524536000_R15","article-title":"Language models are few-shot learners. Language models are few-shot learners","volume-title":"arXiv [cs.CL]","author":"Brown","year":"2020"},{"key":"2025092510524536000_R16","article-title":"Survey of hallucination in natural language generation","volume-title":"arXiv [cs.CL]","author":"Ji","year":"2022"},{"key":"2025092510524536000_R17","doi-asserted-by":"crossref","first-page":"D121","DOI":"10.1093\/nar\/gkac1051","article-title":"The European Nucleotide Archive in 2022","volume":"51","author":"Burgin","year":"2023","journal-title":"Nucleic Acids Res"},{"key":"2025092510524536000_R18","article-title":"BERTopic: neural topic modeling with a class-based TF-IDF procedure. BERTopic: neural topic modeling with a class-based TF-IDF procedure","volume-title":"arXiv [cs.CL]","author":"Grootendorst","year":"2022"},{"key":"2025092510524536000_R19","doi-asserted-by":"crossref","DOI":"10.18653\/v1\/D19-1410","article-title":"Sentence-BERT: sentence embeddings using siamese BERT-Networks","volume-title":"arXiv [cs.CL]","author":"Reimers","year":"2019"},{"key":"2025092510524536000_R20","article-title":"UMAP: uniform manifold approximation and projection for dimension reduction","volume-title":"arXiv [stat.ML]","author":"McInnes","year":"2018"},{"key":"2025092510524536000_R21","first-page":"33","article-title":"Accelerated hierarchical density based clustering","author":"McInnes"},{"key":"2025092510524536000_R22","first-page":"317","volume-title":"Lecture Notes in Computer Science, Lecture Notes in Computer Science","author":"Allaoui","year":"2020"},{"key":"2025092510524536000_R23","article-title":"Check your facts and try again: improving large language models with external knowledge and automated feedback. Check your facts and try again: improving large language models with external knowledge and automated feedback","volume-title":"arXiv [cs.CL]","author":"Peng","year":"2023"},{"key":"2025092510524536000_R24","article-title":"Chain-of-thought prompting elicits reasoning in large language models. Chain-of-thought prompting elicits reasoning in large language models","volume-title":"arXiv [cs.CL]","author":"Wei","year":"2022"},{"key":"2025092510524536000_R25","article-title":"Do multi-document summarization models synthesize? Do multi-document summarization models synthesize?","volume-title":"arXiv [cs.CL]","author":"DeYoung","year":"2023"},{"key":"2025092510524536000_R26","article-title":"Automated metrics for medical multi-document summarization disagree with human evaluations. Automated metrics for medical multi-document summarization disagree with human evaluations","volume-title":"arXiv [cs.CL]","author":"Wang","year":"2023"},{"key":"2025092510524536000_R27","article-title":"Exploring the limits of transfer learning with a unified text-to-text transformer. Exploring the limits of transfer learning with a unified text-to-text transformer","volume-title":"arXiv [cs.LG]","author":"Raffel","year":"2019"},{"key":"2025092510524536000_R28","article-title":"Summarizing, simplifying, and synthesizing medical evidence using GPT-3 (with varying success). Summarizing, simplifying, and synthesizing medical evidence using GPT-3 (with varying success)","volume-title":"arXiv [cs.CL]","author":"Shaib","year":"2023"},{"key":"2025092510524536000_R29","article-title":"Visconde: multi-document QA with GPT-3 and neural reranking","volume-title":"arXiv [cs.CL]","author":"Pereira","year":"2022"},{"key":"2025092510524536000_R30","doi-asserted-by":"crossref","DOI":"10.1371\/journal.pone.0065390","article-title":"The SPECIES and ORGANISMS resources for fast and accurate identification of taxonomic names in text","volume":"8","author":"Pafilis","year":"2013","journal-title":"PLoS One"},{"key":"2025092510524536000_R31","article-title":"Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena","volume-title":"arXiv [cs.CL]","author":"Zheng","year":"2023"},{"key":"2025092510524536000_R32","article-title":"Prometheus: inducing fine-grained evaluation capability in language models. Prometheus: inducing fine-grained evaluation capability in language models","volume-title":"arXiv [cs.CL]","author":"Kim","year":"2023"},{"key":"2025092510524536000_R33","article-title":"Lost in the middle: how language models use long contexts. Lost in the middle: how language models use long contexts","volume-title":"arXiv [cs.CL]","author":"Liu","year":"2023"},{"key":"2025092510524536000_R34","article-title":"Effective long-context scaling of foundation models. Effective long-context scaling of foundation models","volume-title":"arXiv [cs.CL]","author":"Xiong","year":"2023"},{"key":"2025092510524536000_R35","article-title":"Retrieval-augmented generation for knowledge-intensive NLP tasks. Retrieval-augmented generation for knowledge-intensive NLP tasks","volume-title":"arXiv [cs.CL]","author":"Lewis","year":"2020"}],"container-title":["Database"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/database\/article-pdf\/doi\/10.1093\/database\/baaf006\/61769336\/baaf006.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/database\/article-pdf\/doi\/10.1093\/database\/baaf006\/61769336\/baaf006.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,9,25]],"date-time":"2025-09-25T14:52:58Z","timestamp":1758811978000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/database\/article\/doi\/10.1093\/database\/baaf006\/8002547"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025]]},"references-count":35,"URL":"https:\/\/doi.org\/10.1093\/database\/baaf006","relation":{},"ISSN":["1758-0463"],"issn-type":[{"value":"1758-0463","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2025]]},"published":{"date-parts":[[2025]]},"article-number":"baaf006"}}