{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,16]],"date-time":"2026-07-16T07:07:14Z","timestamp":1784185634273,"version":"3.55.0"},"reference-count":36,"publisher":"Springer Science and Business Media LLC","issue":"5","license":[{"start":{"date-parts":[[2026,7,16]],"date-time":"2026-07-16T00:00:00Z","timestamp":1784160000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2026,7,16]],"date-time":"2026-07-16T00:00:00Z","timestamp":1784160000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/100007195","name":"Universit\u00e0 degli Studi di Napoli Federico II","doi-asserted-by":"crossref","id":[{"id":"10.13039\/100007195","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Data Min Knowl Disc"],"published-print":{"date-parts":[[2026,10]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>\n                    Named Entity Recognition (NER) in specialized domains like biomedicine suffers from acute data scarcity, requiring expensive expert annotations. While data augmentation offers a promising solution, it inevitably introduces noisy and mislabeled samples that can degrade model performance. This problem is amplified in few-shot scenarios where every training example matters. We introduce PALAUNER (Policy-based Active Learning to Augment Named Entity Recognition), a reinforcement learning framework that learns to select high-quality samples from augmented data pools. Using a deep Q-network, our agent evaluates samples based on content features and model predictions, deciding which examples will improve NER performance. Experiments across five BioNER benchmarks demonstrate that PALAUNER consistently enhances diverse augmentation methods, from simple perturbations to GPT-based generation. Average F1 improvements are of 0.5\u22127.1 points in few-shot settings. PALAUNER\u2019s modular design enables seamless integration with emerging augmentation techniques, providing a generalizable solution for training data quality enhancement. We publicly release our code on GitHub: (\n                    <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"https:\/\/github.com\/picuslab\/palauner\" ext-link-type=\"uri\">https:\/\/github.com\/picuslab\/palauner<\/jats:ext-link>\n                    ).\n                  <\/jats:p>","DOI":"10.1007\/s10618-026-01239-2","type":"journal-article","created":{"date-parts":[[2026,7,16]],"date-time":"2026-07-16T06:53:49Z","timestamp":1784184829000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Palauner: policy-based active learning to augment named entity recognition datasets"],"prefix":"10.1007","volume":"40","author":[{"given":"Marco","family":"Postiglione","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Andrea","family":"Vignali","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Giancarlo","family":"Sperl\u00ed","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Guido Maria","family":"Secondulfo","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Vincenzo","family":"Moscato","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2026,7,16]]},"reference":[{"key":"1239_CR1","doi-asserted-by":"publisher","first-page":"161","DOI":"10.1186\/1471-2105-13-161","volume":"13","author":"M Bada","year":"2012","unstructured":"Bada M, Eckert M, Evans D et al (2012) Concept annotation in the craft corpus. BMC Bioinformatics 13:161. https:\/\/doi.org\/10.1186\/1471-2105-13-161","journal-title":"BMC Bioinformatics"},{"key":"1239_CR2","doi-asserted-by":"publisher","unstructured":"Bartolini I, Moscato V, Postiglione M, et\u00a0al (2022) COSINER: context similarity data augmentation for named entity recognition. In: Similarity Search and Applications - 15th International Conference, SISAP 2022, Bologna, Italy, October 5-7, 2022, Proceedings, Lecture Notes in Computer Science, vol 13590. Springer, pp 11\u201324, https:\/\/doi.org\/10.1007\/978-3-031-17849-8_2,","DOI":"10.1007\/978-3-031-17849-8_2"},{"key":"1239_CR3","doi-asserted-by":"publisher","DOI":"10.1016\/j.is.2023.102291","volume":"119","author":"I Bartolini","year":"2023","unstructured":"Bartolini I, Moscato V, Postiglione M et al (2023) Data augmentation via context similarity: An application to biomedical named entity recognition. Inf Syst 119:102291. https:\/\/doi.org\/10.1016\/j.is.2023.102291 (https:\/\/www.sciencedirect.com\/science\/article\/pii\/S0306437923001278)","journal-title":"Inf Syst"},{"key":"1239_CR4","doi-asserted-by":"publisher","first-page":"221","DOI":"10.1186\/s12911-024-02624-x","volume":"24","author":"H Chen","year":"2024","unstructured":"Chen H, Dan L, Lu Y et al (2024) An improved data augmentation approach and its application in medical named entity recognition. BMC Med Inform Decis Mak 24:221","journal-title":"BMC Med Inform Decis Mak"},{"key":"1239_CR5","doi-asserted-by":"publisher","unstructured":"Chen S, Aguilar G, Neves L, et\u00a0al (2021) Data augmentation for cross-domain named entity recognition. In: Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, Online and Punta Cana, Dominican Republic, pp 5346\u20135356, https:\/\/doi.org\/10.18653\/v1\/2021.emnlp-main.434, https:\/\/aclanthology.org\/2021.emnlp-main.434","DOI":"10.18653\/v1\/2021.emnlp-main.434"},{"key":"1239_CR6","doi-asserted-by":"publisher","first-page":"11","DOI":"10.1016\/J.JBI.2015.09.010","volume":"58","author":"Y Chen","year":"2015","unstructured":"Chen Y, Lasko TA, Mei Q et al (2015) A study of active learning methods for named entity recognition in clinical text. J Biomed Inform 58:11\u201318. https:\/\/doi.org\/10.1016\/J.JBI.2015.09.010","journal-title":"J Biomed Inform"},{"key":"1239_CR7","doi-asserted-by":"publisher","first-page":"129","DOI":"10.1613\/jair.295","volume":"4","author":"DA Cohn","year":"1996","unstructured":"Cohn DA, Ghahramani Z, Jordan MI (1996) Active learning with statistical models. J Artif Intell Res 4:129\u2013145","journal-title":"J Artif Intell Res"},{"key":"1239_CR8","unstructured":"Culotta A, McCallum A (2005) Reducing labeling effort for structured prediction tasks. In: Veloso MM, Kambhampati S (eds) Proceedings, The Twentieth National Conference on Artificial Intelligence and the Seventeenth Innovative Applications of Artificial Intelligence Conference, July 9-13, 2005, Pittsburgh, Pennsylvania, USA. AAAI Press \/ The MIT Press, pp 746\u2013751, http:\/\/www.aaai.org\/Library\/AAAI\/2005\/aaai05-117.php"},{"key":"1239_CR9","doi-asserted-by":"publisher","unstructured":"Dai X, Adel H (2020) An analysis of simple data augmentation for named entity recognition. In: Scott D, Bel N, Zong C (eds) Proceedings of the 28th International Conference on Computational Linguistics, COLING 2020, Barcelona, Spain (Online), December 8-13, 2020. International Committee on Computational Linguistics, pp 3861\u20133867, https:\/\/doi.org\/10.18653\/v1\/2020.coling-main.343,","DOI":"10.18653\/v1\/2020.coling-main.343"},{"key":"1239_CR10","first-page":"1578","volume":"2024","author":"B Ding","year":"2024","unstructured":"Ding B et al (2024) Data augmentation using llms: data perspectives, learning paradigms and challenges. Find Assoc Comput Linguist ACL 2024:1578\u20131596","journal-title":"Find Assoc Comput Linguist ACL"},{"key":"1239_CR11","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1016\/j.jbi.2013.12.006","volume":"47","author":"RI Dogan","year":"2014","unstructured":"Dogan RI, Leaman R, Lu Z (2014) NCBI disease corpus: a resource for disease name recognition and concept normalization. J Biomed Informatics 47:1\u201310. https:\/\/doi.org\/10.1016\/j.jbi.2013.12.006","journal-title":"J Biomed Informatics"},{"key":"1239_CR12","doi-asserted-by":"publisher","unstructured":"Fang M, Li Y, Cohn T (2017) Learning how to active learn: a deep reinforcement learning approach. In: Palmer M, Hwa R, Riedel S (eds) Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, EMNLP 2017, Copenhagen, Denmark, September 9-11, 2017. Assoc Comput Linguist , pp 595\u2013605, https:\/\/doi.org\/10.18653\/v1\/d17-1063,","DOI":"10.18653\/v1\/d17-1063"},{"key":"1239_CR13","doi-asserted-by":"publisher","unstructured":"Gerner M, Nenadic G, Bergman CM (2010) Linnaeus: a species name identification system for biomedical literature. BMC Bioinformatics 11(85). https:\/\/doi.org\/10.1186\/1471-2105-11-85","DOI":"10.1186\/1471-2105-11-85"},{"key":"1239_CR14","first-page":"6842","volume":"2024","author":"J Kim","year":"2024","unstructured":"Kim J et al (2024) Pushing the limits of low-resource ner using llm artificial data. Find Assoc Comput Linguist ACL 2024:6842\u20136854","journal-title":"Find Assoc Comput Linguist ACL"},{"key":"1239_CR15","doi-asserted-by":"publisher","unstructured":"Kim Y (2014) Convolutional neural networks for sentence classification. In: Moschitti A, Pang B, Daelemans W (eds) Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing, EMNLP 2014, October 25-29, 2014, Doha, Qatar, A meeting of SIGDAT, a Special Interest Group of the ACL. ACL, pp 1746\u20131751, https:\/\/doi.org\/10.3115\/v1\/d14-1181,","DOI":"10.3115\/v1\/d14-1181"},{"key":"1239_CR16","doi-asserted-by":"crossref","unstructured":"Labrak Y, Rouvier M, Dufour R (2024) A zero-shot and few-shot study of instruction-finetuned large language models applied to clinical and biomedical tasks. In: Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024). pp 2049\u20132066","DOI":"10.63317\/2fcovvb4zao3"},{"key":"1239_CR17","doi-asserted-by":"publisher","first-page":"1234","DOI":"10.1093\/bioinformatics\/btz682","volume":"36","author":"J Lee","year":"2020","unstructured":"Lee J, Yoon W, Kim S et al (2020) Biobert: a pre-trained biomedical language representation model for biomedical text mining. Bioinformatics 36:1234\u20131240. https:\/\/doi.org\/10.1093\/bioinformatics\/btz682","journal-title":"Bioinformatics"},{"key":"1239_CR18","doi-asserted-by":"publisher","unstructured":"Li J, Sun Y, Johnson RJ et al (2016) Biocreative V CDR task corpus: a resource for chemical disease relation extraction. Database J Biol Databases Curation 2016. https:\/\/doi.org\/10.1093\/database\/baw068","DOI":"10.1093\/database\/baw068"},{"issue":"1","key":"1239_CR19","doi-asserted-by":"publisher","first-page":"50","DOI":"10.1109\/TKDE.2020.2981314","volume":"34","author":"J Li","year":"2022","unstructured":"Li J, Sun A, Han J et al (2022) A survey on deep learning for named entity recognition. IEEE Trans Knowl Data Eng 34(1):50\u201370. https:\/\/doi.org\/10.1109\/TKDE.2020.2981314","journal-title":"IEEE Trans Knowl Data Eng"},{"issue":"7540","key":"1239_CR20","doi-asserted-by":"publisher","first-page":"529","DOI":"10.1038\/nature14236","volume":"518","author":"V Mnih","year":"2015","unstructured":"Mnih V, Kavukcuoglu K, Silver D et al (2015) Human-level control through deep reinforcement learning. Nat 518(7540):529\u2013533. https:\/\/doi.org\/10.1038\/nature14236","journal-title":"Nat"},{"issue":"5","key":"1239_CR21","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3609483","volume":"14","author":"V Moscato","year":"2023","unstructured":"Moscato V, Postiglione M, Sperl\u00ed G (2023) Few-shot named entity recognition: definition, taxonomy and research directions. ACM Trans Intell Syst Technol 14(5):1\u201346","journal-title":"ACM Trans Intell Syst Technol"},{"key":"1239_CR22","unstructured":"Mou G (2024) Reducing and exploiting data augmentation noise through meta reweighting contrastive learning for text classification. arXiv preprint arXiv:2409.17474"},{"key":"1239_CR23","doi-asserted-by":"crossref","unstructured":"Munnangi M, Feldman S, Wallace B, et\u00a0al (2024) On-the-fly definition augmentation of llms for biomedical ner. In: Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pp 3833\u20133854","DOI":"10.18653\/v1\/2024.naacl-long.212"},{"key":"1239_CR24","doi-asserted-by":"crossref","unstructured":"Nagar A, Schlegel V, Nguyen TT, et\u00a0al (2024) Llms are not zero-shot reasoners for biomedical information extraction. arXiv preprint arXiv:2408.12249","DOI":"10.18653\/v1\/2025.insights-1.11"},{"key":"1239_CR25","doi-asserted-by":"publisher","unstructured":"Reimers N, Gurevych I (2019) Sentence-bert: sentence embeddings using siamese bert-networks. In: Inui K, Jiang J, Ng V, et\u00a0al (eds) Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, EMNLP-IJCNLP 2019, Hong Kong, China, November 3-7, 2019. Association for Computational Linguistics, pp 3980\u20133990, https:\/\/doi.org\/10.18653\/V1\/D19-1410,","DOI":"10.18653\/V1\/D19-1410"},{"key":"1239_CR26","doi-asserted-by":"publisher","unstructured":"Seung HS, Opper M, Sompolinsky H (1992) Query by committee. In: Haussler D (ed) Proceedings of the Fifth Annual ACM Conference on Computational Learning Theory, COLT 1992, Pittsburgh, PA, USA, July 27-29, 1992. ACM, pp 287\u2013294, https:\/\/doi.org\/10.1145\/130385.130417,","DOI":"10.1145\/130385.130417"},{"key":"1239_CR27","unstructured":"Shen Y, Yun H, Lipton ZC, et\u00a0al (2018) Deep active learning for named entity recognition. In: 6th International Conference on Learning Representations, ICLR 2018, Vancouver, BC, Canada, April 30 - May 3, 2018, Conference Track Proceedings. OpenReview.net, https:\/\/openreview.net\/forum?id=ry018WZAZ"},{"key":"1239_CR28","doi-asserted-by":"publisher","first-page":"S2","DOI":"10.1186\/gb-2008-9-s2-s2","volume":"9","author":"LL Smith","year":"2008","unstructured":"Smith LL, Tanabe LK, Ando R et al (2008) Overview of biocreative ii gene mention recognition. Genome Biol 9:S2\u2013S2","journal-title":"Genome Biol"},{"key":"1239_CR29","doi-asserted-by":"crossref","unstructured":"Song S, Shen F, Zhao J (2023) Ropda: Robust prompt-based data augmentation for low-resource named entity recognition. arXiv preprint arXiv:2307.07417","DOI":"10.1609\/aaai.v38i17.29868"},{"issue":"6","key":"1239_CR30","doi-asserted-by":"publisher","DOI":"10.1007\/s11704-024-40555-y","volume":"18","author":"D Xu","year":"2024","unstructured":"Xu D, Chen W, Peng W et al (2024) Large language models for generative information extraction: a survey. Front Comp Sci 18(6):186357","journal-title":"Front Comp Sci"},{"key":"1239_CR31","doi-asserted-by":"publisher","DOI":"10.1051\/wujns\/2023284299","author":"H Yu","year":"2023","unstructured":"Yu H, Ni K, Xu R et al (2023) Ept: data augmentation with embedded prompt tuning for low-resource named entity recognition. WUJNS. https:\/\/doi.org\/10.1051\/wujns\/2023284299","journal-title":"WUJNS"},{"key":"1239_CR32","unstructured":"Zhang L, Tian Z, Zhou W, et\u00a0al (2024) Learning from long-tailed noisy data with sample selection and balanced loss. In: Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence, IJCAI 2024, Jeju, South Korea, August 3-9, 2024. ijcai.org, pp 5471\u20135480, https:\/\/www.ijcai.org\/proceedings\/2024\/605"},{"issue":"9","key":"1239_CR33","doi-asserted-by":"publisher","DOI":"10.1016\/j.heliyon.2024.e30053","volume":"10","author":"T Zhang","year":"2024","unstructured":"Zhang T, Tsai CY (2024) Evolution and emerging trends of named entity recognition: bibliometric analysis from 2000 to 2023. Heliyon 10(9):e30053","journal-title":"Heliyon"},{"key":"1239_CR34","doi-asserted-by":"crossref","unstructured":"Zhao o (2025) Few-shot biomedical ner empowered by llms-assisted data augmentation and multi-scale feature extraction. BioData Mining","DOI":"10.1186\/s13040-025-00443-y"},{"key":"1239_CR35","doi-asserted-by":"publisher","unstructured":"Zhou R, Li X, He R, et\u00a0al (2022a) Melm: Data augmentation with masked entity language modeling for low-resource ner. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp 2251\u20132262, https:\/\/doi.org\/10.18653\/v1\/2022.acl-long.160","DOI":"10.18653\/v1\/2022.acl-long.160"},{"key":"1239_CR36","doi-asserted-by":"publisher","unstructured":"Zhou R, Li X, He R, et\u00a0al (2022b) MELM: data augmentation with masked entity language modeling for low-resource NER. In: Muresan S, Nakov P, Villavicencio A (eds) Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2022, Dublin, Ireland, May 22-27, 2022. Association for Computational Linguistics, pp 2251\u20132262, https:\/\/doi.org\/10.18653\/V1\/2022.ACL-LONG.160,","DOI":"10.18653\/V1\/2022.ACL-LONG.160"}],"container-title":["Data Mining and Knowledge Discovery"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10618-026-01239-2.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10618-026-01239-2","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10618-026-01239-2.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,7,16]],"date-time":"2026-07-16T06:53:55Z","timestamp":1784184835000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10618-026-01239-2"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,7,16]]},"references-count":36,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2026,10]]}},"alternative-id":["1239"],"URL":"https:\/\/doi.org\/10.1007\/s10618-026-01239-2","relation":{},"ISSN":["1384-5810","1573-756X"],"issn-type":[{"value":"1384-5810","type":"print"},{"value":"1573-756X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,7,16]]},"assertion":[{"value":"15 January 2026","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"20 June 2026","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"16 July 2026","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"The authors declare no conflict of interest.","order":1,"name":"Ethics","label":"Conflict of interest","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"This research has not involved human participants and\/or animals.","order":2,"name":"Ethics","label":"Ethical Approval","group":{"name":"EthicsHeading","label":"Declarations"}}],"article-number":"75"}}