{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,22]],"date-time":"2026-07-22T19:55:04Z","timestamp":1784750104217,"version":"3.55.0"},"reference-count":59,"publisher":"Frontiers Media SA","license":[{"start":{"date-parts":[[2025,8,20]],"date-time":"2025-08-20T00:00:00Z","timestamp":1755648000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/100000026","name":"National Institute on Drug Abuse","doi-asserted-by":"publisher","award":["R01 DA053028"],"award-info":[{"award-number":["R01 DA053028"]}],"id":[{"id":"10.13039\/100000026","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["frontiersin.org"],"crossmark-restriction":true},"short-container-title":["Front. Neuroinform."],"abstract":"<jats:p>We show that recent (mid-to-late 2024) commercial large language models (LLMs) are capable of good quality metadata extraction and annotation with very little work on the part of investigators for several exemplar real-world annotation tasks in the neuroimaging literature. We investigated the GPT-4o LLM from OpenAI which performed comparably with several groups of specially trained and supervised human annotators. The LLM achieves similar performance to humans, between 0.91 and 0.97 on zero-shot prompts without feedback to the LLM. Reviewing the disagreements between LLM and gold standard human annotations we note that actual LLM errors are comparable to human errors in most cases, and in many cases these disagreements are not errors. Based on the specific types of annotations we tested, with exceptionally reviewed gold-standard correct values, the LLM performance is usable for metadata annotation at scale. We encourage other research groups to develop and make available more specialized \u201cmicro-benchmarks,\u201d like the ones we provide here, for testing both LLMs, and more complex agent systems annotation performance in real-world metadata annotation tasks.<\/jats:p>","DOI":"10.3389\/fninf.2025.1609077","type":"journal-article","created":{"date-parts":[[2025,8,20]],"date-time":"2025-08-20T05:34:01Z","timestamp":1755668041000},"update-policy":"https:\/\/doi.org\/10.3389\/crossmark-policy","source":"Crossref","is-referenced-by-count":7,"title":["Large language models can extract metadata for annotation of human neuroimaging publications"],"prefix":"10.3389","volume":"19","author":[{"given":"Matthew D.","family":"Turner","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Abhishek","family":"Appaji","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Nibras","family":"Ar Rakib","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Pedram","family":"Golnari","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Arcot K.","family":"Rajasekar","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Anitha Rathnam","family":"K V","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Satya S.","family":"Sahoo","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yue","family":"Wang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Lei","family":"Wang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jessica A.","family":"Turner","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1965","published-online":{"date-parts":[[2025,8,20]]},"reference":[{"key":"B1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2402.01788","article-title":"LitLLM: A toolkit for scientific literature review.","author":"Agarwal","year":"2024","journal-title":"arXiv [Preprint]"},{"key":"B2","doi-asserted-by":"publisher","first-page":"602","DOI":"10.1109\/ICMLA58977.2023.00089","article-title":"ChatGPT vs. Human annotators: A comprehensive analysis of ChatGPT for text annotation","author":"Aldeen","year":"2023","journal-title":"Proceedings of the 2023 International Conference on Machine Learning and Applications (ICMLA)"},{"key":"B3","doi-asserted-by":"publisher","first-page":"884","DOI":"10.1016\/j.neuroimage.2018.08.075","article-title":"Polygenic risk score for schizophrenia and structural brain connectivity in older age: A longitudinal connectome and tractography study.","volume":"183","author":"Alloza","year":"2018","journal-title":"NeuroImage"},{"key":"B4","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2312.03740","article-title":"A survey on prompting techniques in LLMs.","author":"Bhandari","year":"2024","journal-title":"arXiv [Preprint]"},{"key":"B5","doi-asserted-by":"publisher","DOI":"10.17487\/RFC8259","author":"Bray","year":"2017","journal-title":"The JavaScript Object Notation (JSON) Data Interchange Format."},{"key":"B6","doi-asserted-by":"publisher","first-page":"28","DOI":"10.18653\/v1\/2024.hcinlp-1.3","article-title":"This reference does not exist: An exploration of LLM citation accuracy and relevance","author":"Byun","year":"2024","journal-title":"Proceedings of the Third Workshop on Bridging Human\u2013Computer Interaction and Natural Language Processing"},{"key":"B7","doi-asserted-by":"publisher","first-page":"145","DOI":"10.1186\/s12883-016-0659-3","article-title":"Regional homogeneity changes between heroin relapse and non-relapse patients under methadone maintenance treatment: A resting-state fMRI study.","volume":"16","author":"Chang","year":"2016","journal-title":"BMC Neurol."},{"key":"B8","doi-asserted-by":"publisher","first-page":"371","DOI":"10.1007\/7854_2024_462","article-title":"Large-scale neuroimaging of mental illness","author":"Ching","year":"2024","journal-title":"Principles and Advances in Population Neuroscience"},{"key":"B9","doi-asserted-by":"publisher","first-page":"1170","DOI":"10.1111\/acer.14048","article-title":"Alterations in white matter microstructure and connectivity in young adults with alcohol use disorder.","volume":"43","author":"Chumin","year":"2019","journal-title":"Alcohol. Clin. Exp. Res."},{"key":"B10","doi-asserted-by":"publisher","first-page":"3533","DOI":"10.1093\/bioinformatics\/btz070","article-title":"PMC text mining subset in BioC: About three million full-text articles and growing.","volume":"35","author":"Comeau","year":"2019","journal-title":"Bioinformatics"},{"key":"B11","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2212.10450","article-title":"Is GPT-3 a good data annotator?","author":"Ding","year":"2023","journal-title":"arXiv [Preprint]"},{"key":"B12","doi-asserted-by":"publisher","first-page":"3302","DOI":"10.14778\/3611479.3611527","article-title":"How large language models will disrupt data management.","volume":"16","author":"Fernandez","year":"2023","journal-title":"Proc. VLDB Endow."},{"key":"B13","doi-asserted-by":"publisher","first-page":"52","DOI":"10.1016\/j.drugalcdep.2016.08.626","article-title":"A preliminary study of longitudinal neuroadaptation associated with recovery from addiction.","volume":"168","author":"Forster","year":"2016","journal-title":"Drug Alcohol Depend."},{"key":"B14","doi-asserted-by":"publisher","first-page":"425","DOI":"10.1016\/j.bpsc.2024.10.015","article-title":"ENIGMA-Meditation: Worldwide Consortium for Neuroscientific Investigations of Meditation Practices.","volume":"10","author":"Ganesan","year":"2024","journal-title":"Biol. Psychiatry Cogn. Neurosci. Neuroimaging"},{"key":"B15","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2305.14627","article-title":"Enabling large language models to generate text with citations.","author":"Gao","year":"2023","journal-title":"arXiv [Preprint]"},{"key":"B16","doi-asserted-by":"publisher","first-page":"190","DOI":"10.1016\/j.schres.2017.01.001","article-title":"Load-dependent hyperdeactivation of the default mode network in people with schizophrenia.","volume":"185","author":"Hahn","year":"2017","journal-title":"Schizophr. Res."},{"key":"B17","doi-asserted-by":"publisher","first-page":"233","DOI":"10.3163\/1536-5050.101.3.020","article-title":"OpenRefine (version 2.5). http:\/\/openrefine.org. Free, open-source tool for cleaning and transforming data.","volume":"101","author":"Ham","year":"2013","journal-title":"J. Med. Libr. Assoc. JMLA"},{"key":"B18","author":"Hisaharo","year":"2024","journal-title":"Optimizing LLM Inference Clusters for Enhanced Performance and Energy Efficiency."},{"key":"B19","doi-asserted-by":"publisher","first-page":"775","DOI":"10.1016\/j.nicl.2018.06.003","article-title":"Abnormal degree centrality in chronic users of codeine-containing cough syrups: A resting-state functional magnetic resonance imaging study.","volume":"19","author":"Hua","year":"2018","journal-title":"NeuroImage Clin."},{"key":"B20","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2307.02185","article-title":"Citation: A key to building responsible and accountable large language models.","author":"Huang","year":"2024","journal-title":"arXiv [Preprint]"},{"key":"B21","doi-asserted-by":"publisher","first-page":"75","DOI":"10.1016\/j.drugalcdep.2016.07.021","article-title":"Dorsal anterior cingulate glutamate is associated with engagement of the default mode network during exposure to smoking cues.","volume":"167","author":"Janes","year":"2016","journal-title":"Drug Alcohol Depend."},{"key":"B22","doi-asserted-by":"publisher","first-page":"202","DOI":"10.1016\/j.eng.2024.04.002","article-title":"Preventing the immense increase in the life-cycle energy and carbon footprints of LLM-powered intelligent chatbots.","volume":"40","author":"Jiang","year":"2024","journal-title":"Engineering"},{"key":"B23","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2212.14024","article-title":"Demonstrate-Search-Predict: Composing retrieval and language models for knowledge-intensive NLP.","author":"Khattab","year":"","journal-title":"arXiv [Preprint]"},{"key":"B24","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2310.03714","article-title":"DSPy: Compiling declarative language model calls into self-improving pipelines.","author":"Khattab","year":"","journal-title":"arXiv [Preprint]"},{"key":"B25","doi-asserted-by":"publisher","first-page":"252","DOI":"10.1002\/hbm.24369","article-title":"Aberrant structural\u2013functional coupling in adult cannabis users.","volume":"40","author":"Kim","year":"2019","journal-title":"Hum. Brain Mapp."},{"key":"B26","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2311.05769","article-title":"Chatbots are not reliable text annotators.","author":"Kristensen-McLachlan","year":"2023","journal-title":"arXiv [Preprint]"},{"key":"B27","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2312.07559","article-title":"PaperQA: Retrieval-augmented generative agent for scientific research.","author":"L\u00e1la","year":"2023","journal-title":"arXiv [Preprint]"},{"key":"B28","doi-asserted-by":"publisher","first-page":"487","DOI":"10.3389\/fpsyt.2018.00487","article-title":"From affective science to psychiatric disorder: Ontology as a semantic bridge.","volume":"9","author":"Larsen","year":"2018","journal-title":"Front. Psychiatry"},{"key":"B29","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2307.11760","article-title":"Large language models understand and can be enhanced by emotional stimuli.","author":"Li","year":"2023","journal-title":"arXiv [Preprint]"},{"key":"B30","doi-asserted-by":"publisher","first-page":"641","DOI":"10.26615\/978-954-452-092-2_069","article-title":"A practical survey on zero-shot prompt design for in-context learning","author":"Li","year":"2023","journal-title":"Proceedings of the Conference Recent Advances in Natural Language Processing - Large Language Models for Natural Language Processings"},{"key":"B31","doi-asserted-by":"publisher","first-page":"101959","DOI":"10.1016\/j.nicl.2019.101959","article-title":"Examining resting-state functional connectivity in first-episode schizophrenia with 7T fMRI and MEG.","volume":"24","author":"Lottman","year":"2019","journal-title":"NeuroImage Clin."},{"key":"B32","doi-asserted-by":"publisher","first-page":"1233","DOI":"10.1039\/D3DD00113J","article-title":"14 examples of how LLMs can transform materials science and chemistry: A reflection on a large language model hackathon.","volume":"2","author":"Maik Jablonka","year":"2023","journal-title":"Digit. Discov."},{"key":"B33","doi-asserted-by":"publisher","first-page":"333","DOI":"10.1016\/j.jbi.2015.06.026","article-title":"An ontology for Autism Spectrum Disorder (ASD) to infer ASD phenotypes from Autism Diagnostic Interview-Revised data.","volume":"56","author":"Mugzach","year":"2015","journal-title":"J. Biomed. Inform."},{"key":"B34","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2410.18889","article-title":"Are LLMs better than reported? Detecting label errors and mitigating their effect on model performance.","author":"Nahum","year":"2024","journal-title":"arXiv [Preprint]"},{"key":"B35","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2311.16733","article-title":"LLMs for Science: Usage for code generation and data analysis.","author":"Nejjar","year":"2024","journal-title":"arXiv [Preprint]"},{"key":"B36","doi-asserted-by":"publisher","first-page":"299","DOI":"10.1038\/nn.4500","article-title":"Best practices in data analysis and sharing in neuroimaging using MRI.","volume":"23","author":"Nichols","year":"2016","journal-title":"Nat. Neurosci."},{"key":"B37","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2405.00492","article-title":"Is temperature the creativity parameter of large language models?","author":"Peeperkorn","year":"2024","journal-title":"arXiv [Preprint]"},{"key":"B38","doi-asserted-by":"publisher","first-page":"32","DOI":"10.1007\/978-3-030-58793-2_3","article-title":"Data cleaning: A case study with openrefine and trifacta wrangler","author":"Petrova-Antonova","year":"2020","journal-title":"Quality of Information and Communications Technology. QUATIC 2020. Communications in Computer and Information Science"},{"key":"B39","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2402.05201","article-title":"The effect of sampling temperature on problem solving in large language models.","author":"Renze","year":"2024","journal-title":"arXiv [Preprint]"},{"key":"B40","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2102.07350","article-title":"Prompt programming for large language models: Beyond the few-shot paradigm.","author":"Reynolds","year":"2021","journal-title":"arXiv [Preprint]"},{"key":"B41","doi-asserted-by":"publisher","first-page":"2114","DOI":"10.1093\/jamia\/ocae074","article-title":"Large language models for biomedicine: Foundations, opportunities, challenges, and best practices.","volume":"31","author":"Sahoo","year":"2024","journal-title":"J. Am. Med. Inform. Assoc."},{"key":"B42","doi-asserted-by":"publisher","first-page":"1216443","DOI":"10.3389\/fninf.2023.1216443","article-title":"NeuroBridge ontology: Computable provenance metadata to give the long tail of neuroimaging data a FAIR chance for secondary use.","volume":"17","author":"Sahoo","year":"2023","journal-title":"Front. Neuroinformatics"},{"key":"B43","doi-asserted-by":"publisher","first-page":"e41723","DOI":"10.7554\/eLife.41723","article-title":"Alcoholism gender differences in brain responsivity to emotional stimuli.","volume":"8","author":"Sawyer","year":"2019","journal-title":"eLife."},{"key":"B44","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2406.06608","article-title":"The Prompt Report: A systematic survey of prompt engineering techniques.","author":"Schulhoff","year":"2025","journal-title":"arXiv [Preprint]"},{"key":"B45","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2407.10930","article-title":"Fine-tuning and prompt optimization: Two great steps that work better together.","author":"Soylu","year":"2024","journal-title":"arXiv [Preprint]"},{"key":"B46","doi-asserted-by":"publisher","first-page":"3645","DOI":"10.18653\/v1\/P19-1355","article-title":"Energy and policy considerations for deep learning in NLP","author":"Strubell","year":"2019","journal-title":"Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics"},{"key":"B47","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2402.13446","article-title":"Large language models for data annotation and synthesis: A survey.","author":"Tan","year":"2024","journal-title":"arXiv [Preprint]"},{"key":"B48","doi-asserted-by":"publisher","first-page":"1930","DOI":"10.1038\/s41591-023-02448-8","article-title":"Large language models in medicine.","volume":"29","author":"Thirunavukarasu","year":"2023","journal-title":"Nat. Med."},{"key":"B49","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1038\/s41398-020-0705-1","article-title":"ENIGMA and global neuroscience: A decade of large-scale studies of the brain in health and disease across more than 40 countries.","volume":"10","author":"Thompson","year":"2020","journal-title":"Transl. Psychiatry"},{"key":"B50","first-page":"99","article-title":"A review of multi-label classification methods","author":"Tsoumakas","year":"2006","journal-title":"Proceedings of the 2nd ADBIS Workshop on Data Mining and Knowledge Discovery (ADMKD 2006)"},{"key":"B51","doi-asserted-by":"publisher","first-page":"29","DOI":"10.1186\/2047-217X-3-29","article-title":"The rise of large-scale imaging studies in psychiatry.","volume":"3","author":"Turner","year":"2014","journal-title":"GigaScience"},{"key":"B52","doi-asserted-by":"publisher","first-page":"3991","DOI":"10.1093\/brain\/awz330","article-title":"Evolutionary modifications in human brain connectivity associated with schizophrenia.","volume":"142","author":"van den Heuvel","year":"2019","journal-title":"Brain"},{"key":"B53","doi-asserted-by":"publisher","first-page":"2418","DOI":"10.1017\/S0033291718000041","article-title":"Differential neural reward mechanisms in treatment-responsive and treatment-resistant schizophrenia.","volume":"48","author":"Vanes","year":"2018","journal-title":"Psychol. Med."},{"key":"B54","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2305.05003","article-title":"Revisiting relation extraction in the era of large language models.","author":"Wadhwa","year":"2024","journal-title":"arXiv [Preprint]"},{"key":"B55","doi-asserted-by":"publisher","first-page":"1215261","DOI":"10.3389\/fninf.2023.1215261","article-title":"NeuroBridge: A prototype platform for discovery of the long-tail neuroimaging data.","volume":"17","author":"Wang","year":"2023","journal-title":"Front. Neuroinformatics"},{"key":"B56","first-page":"1135","article-title":"Enabling scientific reproducibility through FAIR data management: An ontology-driven deep learning approach in the neurobridge project.","volume":"2022","author":"Wang","year":"2023","journal-title":"AMIA Annu. Symp. Proc."},{"key":"B57","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2402.02008","article-title":"How well do LLMs cite relevant medical references? An evaluation framework and analyses.","author":"Wu","year":"2024","journal-title":"arXiv [Preprint]"},{"key":"B58","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2407.04130","article-title":"Towards automating text annotation: A case study on semantic proximity annotation using GPT-4.","author":"Yadav","year":"2024","journal-title":"arXiv [Preprint]"},{"key":"B59","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2303.18223","article-title":"A survey of large language models.","author":"Zhao","year":"2023","journal-title":"arXiv [Preprint]"}],"container-title":["Frontiers in Neuroinformatics"],"original-title":[],"link":[{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/fninf.2025.1609077\/full","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,8,20]],"date-time":"2025-08-20T05:34:06Z","timestamp":1755668046000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/fninf.2025.1609077\/full"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,8,20]]},"references-count":59,"alternative-id":["10.3389\/fninf.2025.1609077"],"URL":"https:\/\/doi.org\/10.3389\/fninf.2025.1609077","relation":{},"ISSN":["1662-5196"],"issn-type":[{"value":"1662-5196","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,8,20]]},"article-number":"1609077"}}