{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,4]],"date-time":"2026-06-04T22:15:02Z","timestamp":1780611302782,"version":"3.54.1"},"reference-count":34,"publisher":"Oxford University Press (OUP)","issue":"6","license":[{"start":{"date-parts":[[2025,11,6]],"date-time":"2025-11-06T00:00:00Z","timestamp":1762387200000},"content-version":"vor","delay-in-days":5,"URL":"https:\/\/creativecommons.org\/licenses\/by-nc\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["62402376"],"award-info":[{"award-number":["62402376"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["62572389"],"award-info":[{"award-number":["62572389"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["72293581"],"award-info":[{"award-number":["72293581"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["72274152"],"award-info":[{"award-number":["72274152"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2025,11,1]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>Reconstructing bioinformatics workflows from the literature is the foundation of scientific analysis. However, the required details\u2014processing steps, software tools, versions, and parameter settings\u2014are dispersed across narrative text, tables, figure captions, and supplemental files. Manual reconstruction typically takes hours per paper and is error-prone, while existing question-answering (QA) and retrieval systems focus on local passages and lack the full-text, multimodal capabilities needed to automatically rebuild complete workflows. We introduce BioWorkflow, a large language model (LLM)-based, retrieval-augmented framework that automates end-to-end workflow extraction from publications by (i) parsing PDFs and building a unified index over text, tables, and figures with chunk-level summaries and embeddings; (ii) hierarchically decomposing queries with dynamic reformulation when new entities or ambiguities emerge; (iii) performing iterative, context-aware retrieval and assembling a directed workflow that captures steps, tools, versions, and parameters; and (iv) linking each predicted element to its cited evidence and running automated consistency checks to suppress hallucinations and ensure traceability. Evaluated on 100 expert-annotated papers, BioWorkflow recovers ~80% of workflow steps (versus ~20% for existing tools), improves reproducibility, completeness, and accuracy by &amp;gt;20% over strong LLM baselines, and reduces curation time to 3\u20135\u00a0minutes per paper, enabling rapid and reliable reuse of published pipelines.<\/jats:p>","DOI":"10.1093\/bib\/bbaf571","type":"journal-article","created":{"date-parts":[[2025,11,9]],"date-time":"2025-11-09T05:47:44Z","timestamp":1762667264000},"source":"Crossref","is-referenced-by-count":2,"title":["BioWorkflow: Retrieving comprehensive bioinformatics workflows from publications"],"prefix":"10.1093","volume":"26","author":[{"given":"Yidan","family":"Wang","sequence":"first","affiliation":[{"name":"School of Computer Science and Technology, Faculty of Electronics and Information Engineering, Xi'an Jiaotong University , No. 28 Xianning West Road, Beilin District, Xi'an, Shaanxi 710049 ,","place":["China"]},{"name":"Shaanxi Engineering Research Center of Medical and Health Big Data, Xi'an Jiaotong University , No. 28 Xianning West Road, Beilin District, Xi'an 710049 ,","place":["China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3862-6557","authenticated-orcid":false,"given":"Jiayin","family":"Wang","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Faculty of Electronics and Information Engineering, Xi'an Jiaotong University , No. 28 Xianning West Road, Beilin District, Xi'an, Shaanxi 710049 ,","place":["China"]},{"name":"Shaanxi Engineering Research Center of Medical and Health Big Data, Xi'an Jiaotong University , No. 28 Xianning West Road, Beilin District, Xi'an 710049 ,","place":["China"]},{"name":"The Second Affiliated Hospital of Xi'an Jiaotong University , No. 157 Xiwu Road, Beilin District, Xi'an 710003 , Shaanxi,","place":["China"]}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"286","published-online":{"date-parts":[[2025,11,6]]},"reference":[{"key":"2025110900473534500_ref1","doi-asserted-by":"publisher","first-page":"1981","DOI":"10.1093\/bib\/bby063","article-title":"A brief history of bioinformatics","volume":"20","author":"Gauthier","year":"2019","journal-title":"Brief Bioinform"},{"key":"2025110900473534500_ref2","article-title":"A data-supported history of bioinformatics tools","author":"Levin","journal-title":"arXiv preprint arXiv:1807.06808"},{"key":"2025110900473534500_ref3","doi-asserted-by":"publisher","article-title":"Documenting bioinformatics software via reverse engineering","author":"Marques","DOI":"10.48550\/arXiv.2305.04349"},{"key":"2025110900473534500_ref4","doi-asserted-by":"publisher","article-title":"Lost in the middle: how language models use long contexts","author":"Liu","DOI":"10.48550\/arXiv.2307.03172"},{"key":"2025110900473534500_ref5","doi-asserted-by":"publisher","article-title":"SciDaSynth: interactive structured knowledge extraction and synthesis from scientific literature with large language model","author":"Wang","DOI":"10.48550\/arXiv.2404.13765"},{"key":"2025110900473534500_ref6","doi-asserted-by":"publisher","article-title":"Retrieval-augmented generation for large language models: a survey","author":"Gao","DOI":"10.48550\/arXiv.2312.10997"},{"key":"2025110900473534500_ref7","doi-asserted-by":"publisher","article-title":"Scaling laws for neural language models","author":"Kaplan","DOI":"10.48550\/arXiv.2001.08361"},{"key":"2025110900473534500_ref8","doi-asserted-by":"publisher","article-title":"Misinforming LLMs: vulnerabilities, challenges and opportunities","author":"Zhou","DOI":"10.48550\/arXiv.2408.01168"},{"key":"2025110900473534500_ref9","first-page":"9459","article-title":"Retrieval-augmented generation for knowledge-intensive NLP tasks","volume":"33","author":"Lewis","year":"2020","journal-title":"Advances in neural information processing systems"},{"key":"2025110900473534500_ref10","doi-asserted-by":"publisher","article-title":"Generate rather than retrieve: large language models are strong context generators","author":"Yu","DOI":"10.48550\/arXiv.2209.10063"},{"key":"2025110900473534500_ref11","doi-asserted-by":"publisher","article-title":"MultiHOP-RAG: benchmarking retrieval-augmented generation for multi-hop queries","author":"Tang","DOI":"10.48550\/arXiv.2401.15391"},{"key":"2025110900473534500_ref12","first-page":"6769","volume-title":"EMNLP","author":"Karpukhin","year":"2020"},{"key":"2025110900473534500_ref13","doi-asserted-by":"publisher","article-title":"PubmedQA: a dataset for biomedical research question answering","author":"Jin","DOI":"10.48550\/arXiv.1909.06146"},{"key":"2025110900473534500_ref14","doi-asserted-by":"publisher","article-title":"Improving small language models on PubMedQA via generative data augmentation","author":"Guo","DOI":"10.48550\/arXiv.2305.07804"},{"key":"2025110900473534500_ref15","doi-asserted-by":"publisher","article-title":"A systematic survey of prompt engineering in large language models: techniques and applications","author":"Sahoo","DOI":"10.48550\/arXiv.2402.07927"},{"key":"2025110900473534500_ref16","doi-asserted-by":"publisher","article-title":"Enhancing scientific reproducibility through automated BioCompute object creation using retrieval-augmented generation from publications","author":"Kim","DOI":"10.48550\/arXiv.2409.15076"},{"key":"2025110900473534500_ref17","doi-asserted-by":"crossref","first-page":"136","DOI":"10.5731\/pdajpst.2016.006734","article-title":"Biocompute objects\u2014a step towards evaluation and validation of biomedical scientific computations","volume":"71","author":"Simonyan","year":"2016","journal-title":"PDA J Pharm Sci Technol"},{"key":"2025110900473534500_ref18","volume-title":"IEEE P2791 BioCompute Working Group (BCOWG): Standard for Bioinformatics Computations and Analyses Generated by High-Throughput Sequencing (HTS) to Facilitate Communication","author":"Mazumder","year":"2020"},{"key":"2025110900473534500_ref19","doi-asserted-by":"publisher","first-page":"1108","DOI":"10.1016\/j.drudis.2022.01.007","article-title":"Communicating regulatory high-throughput sequencing data using BioCompute objects","volume":"27","author":"King","year":"2022","journal-title":"Drug Discov Today"},{"key":"2025110900473534500_ref20","doi-asserted-by":"publisher","first-page":"e3000099","DOI":"10.1371\/journal.pbio.3000099","article-title":"Enabling precision medicine via standard communication of HTS provenance, analysis, and results","volume":"16","author":"Alterovitz","year":"2018","journal-title":"PLoS Biol"},{"key":"2025110900473534500_ref21","doi-asserted-by":"publisher","first-page":"103884","DOI":"10.1016\/j.drudis.2024.103884","article-title":"Communicating computational workflows in a regulatory environment","volume":"29","author":"Keeney","year":"2024","journal-title":"Drug Discov Today"},{"key":"2025110900473534500_ref22","doi-asserted-by":"publisher","first-page":"855","DOI":"10.1093\/glycob\/cwac046","article-title":"Modeling and integration of N-glycan biomarkers in a comprehensive biomarker data model","volume":"32","author":"Lyman","year":"2022","journal-title":"Glycobiology"},{"key":"2025110900473534500_ref23","doi-asserted-by":"publisher","first-page":"5640","DOI":"10.1038\/s41467-024-49777-x","article-title":"A data science roadmap for open science organizations engaged in early-stage drug discovery","volume":"15","author":"Edfeldt","year":"2024","journal-title":"Nat Commun"},{"key":"2025110900473534500_ref24"},{"key":"2025110900473534500_ref25","doi-asserted-by":"publisher","article-title":"HotpotQA: a dataset for diverse, explainable multi-hop question answering","author":"Yang","DOI":"10.48550\/arXiv.1809.09600"},{"key":"2025110900473534500_ref26","author":"Guti\u00e9rrez"},{"key":"2025110900473534500_ref27","doi-asserted-by":"publisher","article-title":"SiReRAG: indexing similar and related information for multihop reasoning","author":"Zhang","DOI":"10.48550\/arXiv.2412.06206"},{"key":"2025110900473534500_ref28","doi-asserted-by":"publisher","article-title":"Vendi-RAG: adaptively trading-off diversity and quality significantly improves retrieval augmented generation with LLMs","author":"Rezaei","DOI":"10.48550\/arXiv.2502.11228"},{"key":"2025110900473534500_ref29","author":"Xu"},{"key":"2025110900473534500_ref30","doi-asserted-by":"publisher","article-title":"Collab-RAG: boosting retrieval-augmented generation for complex question answering via white-box and Black-box LLM collaboration","author":"Xu","DOI":"10.48550\/arXiv.2504.04915"},{"key":"2025110900473534500_ref31","doi-asserted-by":"publisher","article-title":"ReaRAG: knowledge-guided reasoning enhances factuality of large reasoning models with iterative retrieval augmented generation","author":"Lee","DOI":"10.48550\/arXiv.2503.21729"},{"key":"2025110900473534500_ref32","doi-asserted-by":"publisher","article-title":"What external knowledge is preferred by LLMs? Characterizing and exploring chain of evidence in imperfect context","author":"Chang","DOI":"10.48550\/arXiv.2412.12632"},{"key":"2025110900473534500_ref33"},{"key":"2025110900473534500_ref34"}],"container-title":["Briefings in Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bib\/article-pdf\/26\/6\/bbaf571\/65235086\/bbaf571.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bib\/article-pdf\/26\/6\/bbaf571\/65235086\/bbaf571.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,11,9]],"date-time":"2025-11-09T05:47:48Z","timestamp":1762667268000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bib\/article\/doi\/10.1093\/bib\/bbaf571\/8315884"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,11,1]]},"references-count":34,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2025,11,1]]}},"URL":"https:\/\/doi.org\/10.1093\/bib\/bbaf571","relation":{},"ISSN":["1467-5463","1477-4054"],"issn-type":[{"value":"1467-5463","type":"print"},{"value":"1477-4054","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2025,11]]},"published":{"date-parts":[[2025,11,1]]},"article-number":"bbaf571"}}