{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,24]],"date-time":"2026-07-24T09:08:11Z","timestamp":1784884091561,"version":"3.55.0"},"reference-count":49,"publisher":"Springer Science and Business Media LLC","issue":"3","license":[{"start":{"date-parts":[[2024,6,26]],"date-time":"2024-06-26T00:00:00Z","timestamp":1719360000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2024,6,26]],"date-time":"2024-06-26T00:00:00Z","timestamp":1719360000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100002347","name":"Bundesministerium f\u00fcr Bildung und Forschung","doi-asserted-by":"publisher","award":["01IS17044"],"award-info":[{"award-number":["01IS17044"]}],"id":[{"id":"10.13039\/501100002347","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Inf Syst Front"],"published-print":{"date-parts":[[2026,6]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>\n                    Despite achieving state-of-the-art results in nearly all Natural Language Processing applications, fine-tuning Transformer-encoder based language models still requires a significant amount of labeled data to achieve satisfying work. A well known technique to reduce the amount of human effort in acquiring a labeled dataset is\n                    <jats:italic>Active Learning<\/jats:italic>\n                    (AL): an iterative process in which only the minimal amount of samples is labeled. AL strategies require access to a quantified confidence measure of the model predictions. A common choice is the softmax activation function for the final Neural Network layer. In this paper, we compare eight alternatives on seven datasets and show that the softmax function provides misleading probabilities. Our finding is that most of the methods primarily identify hard-to-learn-from samples (commonly called outliers), resulting in worse than random performance, instead of samples, which actually reduce the uncertainty of the learned language model. As a solution, this paper proposes Uncertainty-Clipping, a heuristic to systematically exclude samples, which results in improvements for most methods compared to the softmax function.\n                  <\/jats:p>","DOI":"10.1007\/s10796-024-10503-z","type":"journal-article","created":{"date-parts":[[2024,6,26]],"date-time":"2024-06-26T02:02:35Z","timestamp":1719367355000},"page":"901-917","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":4,"title":["Comparing and Improving Active Learning Uncertainty Measures for Transformer Models by Discarding Outliers"],"prefix":"10.1007","volume":"28","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-5985-4348","authenticated-orcid":false,"given":"Julius","family":"Gonsior","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Christian","family":"Falkenberg","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Silvio","family":"Magino","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2537-9841","authenticated-orcid":false,"given":"Anja","family":"Reusch","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5334-059X","authenticated-orcid":false,"given":"Claudio","family":"Hartmann","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1665-977X","authenticated-orcid":false,"given":"Maik","family":"Thiele","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8107-2775","authenticated-orcid":false,"given":"Wolfgang","family":"Lehner","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2024,6,26]]},"reference":[{"key":"10503_CR1","unstructured":"Baram Y, Yaniv RE, Luz K (2004). Online choice of active learning algorithms. Journal of Machine Learning Research 5(Mar):255\u2013291"},{"key":"10503_CR2","unstructured":"Blundell C, Cornebise J, Kavukcuoglu K, Wierstra D (2015) Weight uncertainty in neural network. In: ICML, PMLR, pp 1613\u20131622"},{"key":"10503_CR3","unstructured":"Coleman C, Yeh C, Mussmann S, Mirzasoleiman B, Bailis P, Liang P, Leskovec J, Zaharia M (2020). Selection via proxy: Efficient data selection for deep learning. ICLR"},{"key":"10503_CR4","unstructured":"D\u2019Arcy M, Downey D (2022). Limitations of active learning with deep transformer language models"},{"key":"10503_CR5","doi-asserted-by":"crossref","unstructured":"Devlin J, Chang MW, Lee K, Toutanova K (2019) Bert: Pre-training of deep bidirectional transformers for language understanding. In: NAACL, Association for Computational Linguistics, pp 4171\u20134186","DOI":"10.18653\/v1\/N19-1423"},{"key":"10503_CR6","unstructured":"Gal Y, Ghahramani Z (2016) Dropout as a bayesian approximation: Representing model uncertainty in deep learning. In: ICML, PMLR, pp 1050\u20131059"},{"key":"10503_CR7","unstructured":"Gal, Y., Islam, R., Ghahramani, Z. (2017). Deep bayesian active learning with image data. In: International Conference on Machine Learning, PMLR, pp 1183\u20131192"},{"key":"10503_CR8","unstructured":"Gawlikowski, J., Tassi, C.R.N., Ali, M., Lee, J., Humt, M., Feng, J., Kruspe, A., Triebel, R., Jung, P., Roscher, R., et\u00a0al. (2021). A survey of uncertainty in deep neural networks. arXiv preprint arXiv:2107.03342"},{"key":"10503_CR9","unstructured":"Gleave, A., Irving, G. (2022). Uncertainty estimation for language reward models. arXiv preprint arXiv:2203.07472"},{"key":"10503_CR10","unstructured":"Gonsior, J., Rehak, J., Thiele, M., Koci, E., G\u00fcnther, M., Lehner, W. (2020). Active learning for spreadsheet cell classification. In: EDBT\/ICDT Workshops"},{"key":"10503_CR11","doi-asserted-by":"publisher","first-page":"47","DOI":"10.1007\/978-3-031-18840-4_4","volume-title":"Discovery Science","author":"J Gonsior","year":"2022","unstructured":"Gonsior, J., Thiele, M., & Lehner, W. (2022). Imital: Learned active learning strategy on synthetic data. In P. Pascal & D. Ienco (Eds.), Discovery Science (pp. 47\u201356). Cham: Springer Nature Switzerland."},{"key":"10503_CR12","doi-asserted-by":"crossref","unstructured":"Hein, M., Andriushchenko, M., Bitterwolf, J. (2019). Why relu networks yield high-confidence predictions far away from the training data and how to mitigate the problem. In: CVPR, IEEE, pp 41\u201350","DOI":"10.1109\/CVPR.2019.00013"},{"key":"10503_CR13","unstructured":"Houlsby N, Husz\u00e1r F, Ghahramani Z, Lengyel M (2011) Bayesian active learning for classification and preference learning. 1112.5745"},{"key":"10503_CR14","doi-asserted-by":"crossref","unstructured":"Hsu, W.N., Lin, H.T. (2015). Active learning by learning. In: Proceedings of the Twenty-Ninth AAAI Conference on Artificial Intelligence, AAAI Press, AAAI\u201915, pp 2659\u20132665","DOI":"10.1609\/aaai.v29i1.9597"},{"key":"10503_CR15","unstructured":"Huang, S.J., Jin, R., Zhou, Z.H. (2010) Active learning by querying informative and representative examples. In: Lafferty J, Williams C, Shawe-Taylor J, Zemel R, Culotta A (eds) NeurIPS, Curran Associates, Inc., vol\u00a023, pp 892\u2013900"},{"key":"10503_CR16","unstructured":"Jiang, H., Kim, B., Guan, M., Gupta, M. (2018). To trust or not to trust a classifier. NeurIPS 31"},{"key":"10503_CR17","doi-asserted-by":"crossref","unstructured":"Karamcheti S, Krishna R, Fei-Fei L, Manning C (2021). Mind your outliers! investigating the negative impact of outliers on active learning for visual question answering. In: ACL-IJCNLP, Association for Computational Linguistics, pp 7265\u20137281","DOI":"10.18653\/v1\/2021.acl-long.564"},{"key":"10503_CR18","unstructured":"Kirsch, A., & v Amersfoort, J., Gal, Y. (2019). Batchbald: Efficient and diverse batch acquisition for deep bayesian active learning. NeurIPS, Curran Associates Inc, 32, 7026\u20137037."},{"key":"10503_CR19","doi-asserted-by":"crossref","unstructured":"Konyushkova, K., Sznitman, R., Fua, P. (2015). Introducing geometry in active learning for image segmentation. In: Guyon I, Luxburg UV, Bengio S, Wallach H, Fergus R, Vishwanathan S, Garnett R (eds) 2015 IEEE International Conference on Computer Vision (ICCV), IEEE, vol\u00a030, pp 4225\u20134235","DOI":"10.1109\/ICCV.2015.340"},{"key":"10503_CR20","unstructured":"Konyushkova, K., Sznitman, R., Fua, P. (2018). Discovering general-purpose active learning strategies. arXiv preprint arXiv:1810.04114"},{"key":"10503_CR21","unstructured":"Lakshminarayanan B, Pritzel A, Blundell C (2017) Simple and scalable predictive uncertainty estimation using deep ensembles. NeurIPS 30"},{"issue":"7553","key":"10503_CR22","doi-asserted-by":"publisher","first-page":"436","DOI":"10.1038\/nature14539","volume":"521","author":"Y LeCun","year":"2015","unstructured":"LeCun, Y., Bengio, Y., & Hinton, G. (2015). Deep learning. Nature, 521(7553), 436\u2013444.","journal-title":"Nature"},{"key":"10503_CR23","doi-asserted-by":"crossref","unstructured":"Lewis, D. D., & Gale, W. A. (1994). A sequential algorithm for training text classifiers. SIGIR \u201994 (pp. 3\u201312). London: Springer.","DOI":"10.1007\/978-1-4471-2099-5_1"},{"key":"10503_CR24","doi-asserted-by":"crossref","unstructured":"Lhoest, Q., del Moral, A.V., Jernite, Y., Thakur, A., von Platen, P., Patil, S., Chaumond, J., Drame, M., Plu, J., Tunstall, L., et\u00a0al (2021). Datasets: A community library for natural language processing. arXiv preprint arXiv:2109.02846","DOI":"10.18653\/v1\/2021.emnlp-demo.21"},{"key":"10503_CR25","unstructured":"Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., Stoyanov, V. (2019). Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692"},{"key":"10503_CR26","doi-asserted-by":"crossref","unstructured":"Lowell, D., Lipton, Z.C., Wallace, B.C. (2019). Practical obstacles to deploying active learning. EMNLP-IJCNLP pp 21\u201330","DOI":"10.18653\/v1\/D19-1003"},{"key":"10503_CR27","unstructured":"McCallumzy, A.K., Nigamy, K. (1998). Employing em and pool-based active learning for text classification. In: ICML, Citeseer, pp 359\u2013367"},{"key":"10503_CR28","unstructured":"Mo\u017cejko, M., Susik, M., Karczewski, R. (2018). Inhibited softmax for uncertainty estimation in neural networks. arXiv preprint arXiv:1810.01861"},{"key":"10503_CR29","unstructured":"Pearce, T., Brintrup, A., Zhu, J. (2021). Understanding softmax confidence and uncertainty. arXiv preprint arXiv:2106.04972"},{"key":"10503_CR30","unstructured":"Sankararaman, K.A., Wang, S., Fang, H. (2022). Bayesformer: Transformer with uncertainty estimation. arXiv preprint arXiv:2206.00826"},{"key":"10503_CR31","first-page":"309","volume-title":"ICDM","author":"T Scheffer","year":"2001","unstructured":"Scheffer, T., Decomain, C., & Wrobel, S. (2001). Mining the web with active hidden markov models. In F. Hoffmann, D. J. Hand, N. Adams, D. Fisher, & G. Guimaraes (Eds.), ICDM (pp. 309\u2013318). Soc: IEEE Comput."},{"key":"10503_CR32","unstructured":"Schr\u00f6der, C., Niekler, A. (2020). A survey of active learning for text classification using deep neural networks. arXiv preprint arXiv:2008.07267"},{"key":"10503_CR33","unstructured":"Schr\u00f6der, C., M\u00fcller, L., Niekler, A., Potthast, M. (2021). Small-text: Active learning for text classification in python. arXiv preprint arXiv:2107.10314"},{"key":"10503_CR34","doi-asserted-by":"crossref","unstructured":"Schr\u00f6der, C., Niekler, A., Potthast, M. (2022). Revisiting uncertainty-based query strategies for active learning with transformers. In: ACL, Association for Computational Linguistics, pp 2194\u20132203","DOI":"10.18653\/v1\/2022.findings-acl.172"},{"key":"10503_CR35","unstructured":"Sener, O., Savarese, S. (2018). Active learning for convolutional neural networks: A core-set approach. International Conference on Learning Representations"},{"key":"10503_CR36","unstructured":"Sensoy, M., Kaplan, L., Kandemir, M. (2018). Evidential deep learning to quantify classification uncertainty. NeurIPS 31"},{"issue":"1","key":"10503_CR37","first-page":"1","volume":"6","author":"B Settles","year":"2012","unstructured":"Settles, B. (2012). Active learning. Artificial Intelligence and Machine Learning, 6(1), 1\u2013114.","journal-title":"Artificial Intelligence and Machine Learning"},{"key":"10503_CR38","doi-asserted-by":"crossref","unstructured":"Seung, H.S., Opper, M., Sompolinsky, H. (1992). Query by committee. In: Proceedings of the fifth annual workshop on Computational learning theory, ACM, COLT \u201992, pp 287\u2013294","DOI":"10.1145\/130385.130417"},{"issue":"3","key":"10503_CR39","doi-asserted-by":"publisher","first-page":"379","DOI":"10.1002\/j.1538-7305.1948.tb01338.x","volume":"27","author":"CE Shannon","year":"1948","unstructured":"Shannon, C. E. (1948). A mathematical theory of communication. Bell System Technical Journal, 27(3), 379\u2013423.","journal-title":"Bell System Technical Journal"},{"key":"10503_CR40","doi-asserted-by":"crossref","unstructured":"Siddhant, A., Lipton, Z.C. (2018). Deep bayesian active learning for natural language processing: Results of a large-scale empirical study. arXiv preprint arXiv:1808.05697","DOI":"10.18653\/v1\/D18-1318"},{"key":"10503_CR41","doi-asserted-by":"crossref","unstructured":"Szegedy, C., Vanhoucke, V., Ioffe, S., Shlens, J., Wojna, Z. (2016). Rethinking the inception architecture for computer vision. In: CVPR, IEEE, pp 2818\u20132826","DOI":"10.1109\/CVPR.2016.308"},{"issue":"01","key":"10503_CR42","doi-asserted-by":"publisher","first-page":"5117","DOI":"10.1609\/aaai.v33i01.33015117","volume":"33","author":"YP Tang","year":"2019","unstructured":"Tang, Y. P., & Huang, S. J. (2019). Self-paced active learning: Query the right thing at the right time. Proceedings of the AAAI Conference on Artificial Intelligence, 33(01), 5117\u20135124.","journal-title":"Proceedings of the AAAI Conference on Artificial Intelligence"},{"key":"10503_CR43","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, L., Polosukhin, I. (2017). Attention is all you need. NeurIPS 30"},{"key":"10503_CR44","doi-asserted-by":"crossref","unstructured":"Wang, Z., & Ye, J. (2015). Querying discriminative and representative samples for batch mode active learning. ACM Transactions on Knowledge Discovery from Data, 9(3), 1\u201323.","DOI":"10.1145\/2700408"},{"key":"10503_CR45","doi-asserted-by":"crossref","unstructured":"Weiss, M., Tonella, P. (2022). Simple techniques work surprisingly well for neural network test prioritization and active learning (replicability study). arXiv preprint arXiv:2205.00664","DOI":"10.1145\/3533767.3534375"},{"key":"10503_CR46","doi-asserted-by":"crossref","unstructured":"Yoo, D., Kweon, I.S. (2019). Learning loss for active learning. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp 93\u2013102","DOI":"10.1109\/CVPR.2019.00018"},{"key":"10503_CR47","doi-asserted-by":"crossref","unstructured":"Zhan, X., Liu, H., Li, Q., Chan, A.B. (2021). A comparative survey: Benchmarking for pool-based active learning. IJCAI pp 4679\u20134686, survey Track","DOI":"10.24963\/ijcai.2021\/634"},{"key":"10503_CR48","unstructured":"Zhang, J., Kailkhura, B., Han, T.Y.J. (2020). Mix-n-match: Ensemble and compositional methods for uncertainty calibration in deep learning. In: ICML, PMLR, pp 11117\u201311128"},{"key":"10503_CR49","doi-asserted-by":"crossref","unstructured":"Zhang, M., Plank, B. (2021). Cartography active learning. arXiv preprint arXiv:2109.04282","DOI":"10.18653\/v1\/2021.findings-emnlp.36"}],"updated-by":[{"DOI":"10.1007\/s10796-024-10519-5","type":"correction","label":"Correction","source":"publisher","updated":{"date-parts":[[2024,7,15]],"date-time":"2024-07-15T00:00:00Z","timestamp":1721001600000}}],"container-title":["Information Systems Frontiers"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10796-024-10503-z.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10796-024-10503-z","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10796-024-10503-z.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,7,24]],"date-time":"2026-07-24T08:57:35Z","timestamp":1784883455000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10796-024-10503-z"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,6,26]]},"references-count":49,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2026,6]]}},"alternative-id":["10503"],"URL":"https:\/\/doi.org\/10.1007\/s10796-024-10503-z","relation":{},"ISSN":["1387-3326","1572-9419"],"issn-type":[{"value":"1387-3326","type":"print"},{"value":"1572-9419","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,6,26]]},"assertion":[{"value":"10 June 2024","order":1,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"26 June 2024","order":2,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"15 July 2024","order":4,"name":"change_date","label":"Change Date","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"Correction","order":5,"name":"change_type","label":"Change Type","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"A Correction to this paper has been published:","order":6,"name":"change_details","label":"Change Details","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"https:\/\/doi.org\/10.1007\/s10796-024-10519-5","URL":"https:\/\/doi.org\/10.1007\/s10796-024-10519-5","order":7,"name":"change_details","label":"Change Details","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"Not applicable","order":1,"name":"Ethics","label":"Ethics approval and consent to participate","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"Not applicable","order":2,"name":"Ethics","label":"Consent for publication","group":{"name":"EthicsHeading","label":"Declarations"}}]}}