{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,26]],"date-time":"2026-02-26T15:24:47Z","timestamp":1772119487088,"version":"3.50.1"},"reference-count":78,"publisher":"Springer Science and Business Media LLC","issue":"8","license":[{"start":{"date-parts":[[2025,5,8]],"date-time":"2025-05-08T00:00:00Z","timestamp":1746662400000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2025,5,8]],"date-time":"2025-05-08T00:00:00Z","timestamp":1746662400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100004004","name":"Universit\u00e0 degli Studi di Trento","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100004004","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Artif Intell Rev"],"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>In this paper, we argue that the prevailing approach to training and evaluating machine learning models often fails to consider their real-world application within organizational or societal contexts, where they are intended to create beneficial value for people. We propose a shift in perspective, redefining model assessment and selection to emphasize integration into workflows that combine machine predictions with human expertise, particularly in scenarios requiring human intervention for low-confidence predictions. Traditional metrics like accuracy and f-score fail to capture the beneficial value of models in such hybrid settings. To address this, we introduce a simple yet theoretically sound \u201cvalue\u201d metric that incorporates task-specific costs for correct predictions, errors, and rejections, offering a practical framework for real-world evaluation. Through extensive experiments, we show that existing metrics fail to capture real-world needs, often leading to suboptimal choices in terms of value when used to rank classifiers. Furthermore, we emphasize the critical role of calibration in determining model value, showing that simple, well-calibrated models can often outperform more complex models that are challenging to calibrate.<\/jats:p>","DOI":"10.1007\/s10462-025-11242-6","type":"journal-article","created":{"date-parts":[[2025,5,8]],"date-time":"2025-05-08T00:36:23Z","timestamp":1746664583000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":4,"title":["Rethinking and recomputing the value of machine learning models"],"prefix":"10.1007","volume":"58","author":[{"given":"Burcu","family":"Sayin","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jie","family":"Yang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xinyue","family":"Chen","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Andrea","family":"Passerini","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Fabio","family":"Casati","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2025,5,8]]},"reference":[{"key":"11242_CR1","doi-asserted-by":"crossref","unstructured":"Agrawal S, Awekar A (2018) Deep learning for detecting cyberbullying across multiple social media platforms. In: Pasi G, Piwowarski B, Azzopardi L, Hanbury A (eds) Advances in information retrieval. Springer, Cham, pp 141\u2013153","DOI":"10.1007\/978-3-319-76941-7_11"},{"key":"11242_CR2","doi-asserted-by":"publisher","unstructured":"Arango A, P\u00e9rez J, Poblete B (2019) Hate speech detection is not as easy as you may think: A closer look at model validation. In: Proceedings of the 42nd International ACM SIGIR Conference on research and development in information retrieval. SIGIR\u201919, pp. 45\u201354. Association for Computing Machinery New York, NY, USA.https:\/\/doi.org\/10.1145\/3331184.3331262","DOI":"10.1145\/3331184.3331262"},{"key":"11242_CR3","doi-asserted-by":"publisher","unstructured":"Badjatiya P, Gupta S, Gupta M, Varma V (2017) Deep learning for hate speech detection in tweets. In: Proceedings of the 26th International Conference on World Wide Web Companion. WWW \u201917 Companion, pp. 759\u2013760. International World Wide Web Conferences Steering Committee Republic and Canton of Geneva, CHE.https:\/\/doi.org\/10.1145\/3041021.3054223","DOI":"10.1145\/3041021.3054223"},{"key":"11242_CR4","doi-asserted-by":"publisher","unstructured":"Bahat Y, Shakhnarovich G (2020) Classification confidence estimation with test-time data-augmentation. ArXiv abs\/2006.16705. https:\/\/doi.org\/10.48550\/ARXIV.2006.16705","DOI":"10.48550\/ARXIV.2006.16705"},{"key":"11242_CR5","doi-asserted-by":"publisher","unstructured":"Balda E, Behboodi A, Mathar R (2020) Adversarial examples in deep neural networks: An overview. In: Deep learning: algorithms and applications, pp. 31\u201365. https:\/\/doi.org\/10.1007\/978-3-030-31760-7_2","DOI":"10.1007\/978-3-030-31760-7_2"},{"key":"11242_CR6","doi-asserted-by":"publisher","first-page":"394","DOI":"10.1007\/BF00379115","volume":"78","author":"R Bendel","year":"1989","unstructured":"Bendel R, Higgins S, Teberg J, Pyke D (1989) Comparison of skewness coefficient, coefficient of variation, and Gini coefficient as inequality measures within populations. Oecologia 78:394\u2013400. https:\/\/doi.org\/10.1007\/BF00379115","journal-title":"Oecologia"},{"key":"11242_CR7","unstructured":"Bragg J, Mausam Weld D.S (2016) Optimal testing for crowd workers. In: Proceedings of the 2016 international conference on autonomous agents & multiagent systems. AAMAS \u201916, pp. 966\u2013974. International foundation for autonomous agents and multiagent systems Richland, SC"},{"key":"11242_CR8","doi-asserted-by":"publisher","unstructured":"Brown T.B, Mann B, Ryder N, Subbiah M, Kaplan J, Dhariwal P, Neelakantan A, Shyam P, Sastry G, Askell A, Agarwal S, Herbert-Voss A, Krueger G, Henighan T, Child R, Ramesh A, Ziegler D.M, Wu J, Winter C, Hesse C, Chen M, Sigler E, Litwin M, Gray S, Chess B, Clark J, Berner C, McCandlish S, Radford A, Sutskever I, Amodei D (2020) Language models are few-shot learners. ArXiv abs\/2005.14165. https:\/\/doi.org\/10.48550\/ARXIV.2005.14165","DOI":"10.48550\/ARXIV.2005.14165"},{"key":"11242_CR9","doi-asserted-by":"publisher","first-page":"3834","DOI":"10.3390\/s21113834","volume":"21","author":"M Bukowski","year":"2021","unstructured":"Bukowski M, Kurek J, Antoniuk I, Jegorowa A (2021) Decision confidence assessment in multi-class classification. Sensors 21:3834. https:\/\/doi.org\/10.3390\/s21113834","journal-title":"Sensors"},{"key":"11242_CR10","doi-asserted-by":"publisher","unstructured":"Callaghan W, Goh J, Mohareb M, Lim A, Law E (2018) Mechanicalheart: a human-machine framework for the classification of phonocardiograms. In: CSCW\u201918, vol. 2, pp. 28\u201312817. https:\/\/doi.org\/10.1145\/3274297","DOI":"10.1145\/3274297"},{"key":"11242_CR11","doi-asserted-by":"publisher","unstructured":"Casati F, Noel P, Yang J (2021) On the value of ml models. In: Neurips workshop on human decisions. https:\/\/doi.org\/10.48550\/ARXIV.2112.06775","DOI":"10.48550\/ARXIV.2112.06775"},{"key":"11242_CR12","doi-asserted-by":"publisher","unstructured":"Chai X, Deng L, Yang Q, Ling CX (2004) Test-cost sensitive naive bayes classification. In: Fourth IEEE international conference on data mining (ICDM\u201904), pp. 51\u201358. https:\/\/doi.org\/10.1109\/ICDM.2004.10092","DOI":"10.1109\/ICDM.2004.10092"},{"key":"11242_CR13","unstructured":"Charoenphakdee N, Cui Z, Zhang Y, Sugiyama M (2021) Classification with rejection based on cost-sensitive classification. In: Proceedings of the 38th international conference on machine learning, vol. 139, pp. 1507\u20131517. https:\/\/proceedings.mlr.press\/v139\/charoenphakdee21a.html"},{"key":"11242_CR14","doi-asserted-by":"publisher","unstructured":"Cheng J, Bernstein M.S (2015) Flock: Hybrid crowd-machine learning classifiers. In: Proceedings of the 18th Acm conference on computer supported cooperative work & social computing. https:\/\/doi.org\/10.1145\/2675133.2675214","DOI":"10.1145\/2675133.2675214"},{"issue":"5","key":"11242_CR15","doi-asserted-by":"publisher","first-page":"1140","DOI":"10.1109\/72.410358","volume":"6","author":"LP Cordella","year":"1995","unstructured":"Cordella LP, De Stefano C, Tortorella F, Vento M (1995) A method for improving classification reliability of multilayer perceptrons. IEEE Trans Neural Netw 6(5):1140\u20131147. https:\/\/doi.org\/10.1109\/72.410358","journal-title":"IEEE Trans Neural Netw"},{"issue":"1","key":"11242_CR16","doi-asserted-by":"publisher","first-page":"84","DOI":"10.1109\/5326.827457","volume":"30","author":"C De Stefano","year":"2000","unstructured":"De Stefano C, Sansone C, Vento M (2000) To reject or not to reject: that is the question-an answer in case of neural classifiers. IEEE Trans Syst Man Cybern 30(1):84\u201394. https:\/\/doi.org\/10.1109\/5326.827457","journal-title":"IEEE Trans Syst Man Cybern"},{"key":"11242_CR17","doi-asserted-by":"publisher","first-page":"637","DOI":"10.1007\/s12599-019-00595-2","volume":"61","author":"D Dellermann","year":"2019","unstructured":"Dellermann D, Ebel P, S\u00f6llner M, Leimeister JM (2019) Hybrid intelligence. Business Inform Syst Eng 61:637\u2013643. https:\/\/doi.org\/10.1007\/s12599-019-00595-2","journal-title":"Business Inform Syst Eng"},{"key":"11242_CR18","doi-asserted-by":"crossref","unstructured":"Dellermann D, Calma A, Lipusch N, Weber T, Weigel S, Ebel PA (2019) The future of human-ai collaboration: a taxonomy of design knowledge for hybrid intelligence systems. ArXiv abs\/2105.03354","DOI":"10.24251\/HICSS.2019.034"},{"key":"11242_CR19","doi-asserted-by":"publisher","unstructured":"Domingos P (1999) Metacost: A general method for making classifiers cost-sensitive. In: Proceedings of the Fifth ACM SIGKDD international conference on knowledge discovery and data mining. KDD \u201999, pp. 155\u2013164. Association for computing machinery New York, NY, USA. https:\/\/doi.org\/10.1145\/312129.312220","DOI":"10.1145\/312129.312220"},{"key":"11242_CR20","unstructured":"Elkan C (2001) The foundations of cost-sensitive learning. In: Proceedings of the 17th International Joint Conference on Artificial Intelligence - Volume 2. IJCAI\u201901, pp. 973\u2013978. Morgan Kaufmann Publishers Inc. San Francisco, CA, USA"},{"key":"11242_CR21","unstructured":"Elkan C (2001) The foundations of cost-sensitive learning. In: Proceedings of the 17th international joint conference on artificial intelligence, pp. 973\u2013978"},{"key":"11242_CR22","doi-asserted-by":"publisher","unstructured":"Fumera G, Roli F (2002) Support vector machines with embedded reject option. In: Proceedings of the first international workshop on pattern recognition with support vector machines. SVM \u201902, pp. 68\u201382. Springer Berlin, Heidelberg. https:\/\/doi.org\/10.1007\/3-540-45665-1_6","DOI":"10.1007\/3-540-45665-1_6"},{"key":"11242_CR23","doi-asserted-by":"publisher","unstructured":"Gadiraju U, Yang J, Bozzon A (2017) Clarity is a worthwhile quality: On the role of task clarity in microtask crowdsourcing. In: Proceedings of the 28th ACM conference on hypertext and social media. HT \u201917, pp. 5\u201314. Association for computing machinery New York, NY, USA. https:\/\/doi.org\/10.1145\/3078714.3078715","DOI":"10.1145\/3078714.3078715"},{"key":"11242_CR24","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.1705.08500","author":"Y Geifman","year":"2017","unstructured":"Geifman Y, El-Yaniv R (2017) Selective classification for deep neural networks. Adv Neural Inform Proc Syst. https:\/\/doi.org\/10.48550\/ARXIV.1705.08500","journal-title":"Adv Neural Inform Proc Syst"},{"key":"11242_CR25","unstructured":"Gunel B.S (2022) Towards reliable hybrid human-machine classifiers. PhD thesis at University of Trento. https:\/\/hdl.handle.net\/11572\/349843"},{"key":"11242_CR26","doi-asserted-by":"publisher","unstructured":"Guo C, Pleiss G, Sun Y, Weinberger K.Q (2017) On calibration of modern neural networks. In: proceedings of the 34th international conference on machine learning - Volume 70. ICML\u201917, pp. 1321\u20131330.https:\/\/doi.org\/10.48550\/ARXIV.1706.04599","DOI":"10.48550\/ARXIV.1706.04599"},{"issue":"5","key":"11242_CR27","doi-asserted-by":"publisher","first-page":"2266","DOI":"10.1109\/TKDE.2019.2948168","volume":"33","author":"L Han","year":"2021","unstructured":"Han L, Roitero K, Gadiraju U, Sarasua C, Checco A, Maddalena E, Demartini G (2021) The impact of task abandonment in crowdsourcing. IEEE Trans Knowl Data Eng 33(5):2266\u20132279. https:\/\/doi.org\/10.1109\/TKDE.2019.2948168","journal-title":"IEEE Trans Knowl Data Eng"},{"key":"11242_CR28","doi-asserted-by":"publisher","unstructured":"Han L, Maddalena E, Checco A, Sarasua C, Gadiraju U, Roitero K, Demartini G (2020) Crowd worker strategies in relevance judgment tasks. In: Proceedings of the 13th International conference on web search and data mining. WSDM \u201920, pp. 241\u2013249. Association for computing machinery New York, NY, USA. https:\/\/doi.org\/10.1145\/3336191.3371857","DOI":"10.1145\/3336191.3371857"},{"key":"11242_CR29","doi-asserted-by":"publisher","unstructured":"Han L, Roitero K, Gadiraju U, Sarasua C, Checco A, Maddalena E, Demartini G (2019) All those wasted hours: On task abandonment in crowdsourcing. In: Proceedings of the twelfth ACM international conference on web search and data mining. WSDM \u201919, pp. 321\u2013329. Association for Computing Machinery New York, NY, USA. https:\/\/doi.org\/10.1145\/3289600.3291035","DOI":"10.1145\/3289600.3291035"},{"key":"11242_CR30","unstructured":"Heitmann M, Siebert C, Hartmann J, Schamp C (2020) More than a feeling: benchmarks for sentiment analysis accuracy. In: Communication & Computational Methods eJournal"},{"issue":"3","key":"11242_CR31","doi-asserted-by":"publisher","first-page":"179","DOI":"10.1109\/TSSC.1970.300339","volume":"6","author":"ME Hellman","year":"1970","unstructured":"Hellman ME (1970) The nearest neighbor classification rule with a reject option. IEEE Trans Syst Sci Cybern 6(3):179\u2013185. https:\/\/doi.org\/10.1109\/TSSC.1970.300339","journal-title":"IEEE Trans Syst Sci Cybern"},{"key":"11242_CR32","doi-asserted-by":"crossref","unstructured":"He H, Ma Y (2013) Imbalanced learning: foundations, algorithms, and applications","DOI":"10.1002\/9781118646106"},{"key":"11242_CR33","unstructured":"Hendrickx K, Perini L, Plas D, Meert W, Davis J (2021) Machine learning with a reject option: a survey. arXiv"},{"key":"11242_CR34","doi-asserted-by":"publisher","first-page":"962","DOI":"10.1162\/tacl_a_00407","volume":"9","author":"Z Jiang","year":"2021","unstructured":"Jiang Z, Araki J, Ding H, Neubig G (2021) How can we know when language models know? On the calibration of language models for question answering. Trans Assoc Comput Linguist 9:962\u2013977. https:\/\/doi.org\/10.1162\/tacl_a_00407","journal-title":"Trans Assoc Comput Linguist"},{"key":"11242_CR35","doi-asserted-by":"publisher","unstructured":"Jiang H, Kim B, Guan M.Y, Gupta M (2018) To trust or not to trust a classifier. In: Proceedings of the 32nd international conference on neural information processing systems. NIPS\u201918, pp. 5546\u20135557. Curran Associates Inc. Red Hook, NY, USA. https:\/\/doi.org\/10.48550\/ARXIV.1805.11783","DOI":"10.48550\/ARXIV.1805.11783"},{"key":"11242_CR36","unstructured":"Kamar E, Hacker S, Horvitz E (2012) Combining human and machine intelligence in large-scale crowdsourcing. In: AAMAS\u201912 - Volume 1, pp. 467\u2013474"},{"key":"11242_CR37","doi-asserted-by":"publisher","unstructured":"Krivosheev E, Casati F, Benatallah B (2018) Crowd-based multi-predicate screening of papers in literature reviews. In: Proceedings of the 2018 World wide web conference. WWW \u201918, pp. 55\u201364. International world wide web conferences steering committee republic and canton of Geneva, CHE. https:\/\/doi.org\/10.1145\/3178876.3186036","DOI":"10.1145\/3178876.3186036"},{"key":"11242_CR38","first-page":"5052","volume":"11","author":"M Kull","year":"2017","unstructured":"Kull M, Silva Filho T, Flach PA (2017) Beyond sigmoids: How to obtain well-calibrated probabilities from binary classifiers with beta calibration. Electr J Stat 11:5052\u20135080","journal-title":"Electr J Stat"},{"key":"11242_CR39","unstructured":"Kull M, Perello\u00a0Nieto M, K\u00e4ngsepp M, Silva\u00a0Filho T, Song H, Flach P.: Beyond temperature scaling: Obtaining well-calibrated multi-class probabilities with dirichlet calibration. In: Wallach H, Larochelle H, Beygelzimer A, Alch\u00e9-Buc F, Fox E, Garnett R. (eds.) (2019) Advances in neural information processing systems, vol. 32. https:\/\/proceedings.neurips.cc\/paper\/2019\/file\/8ca01ea920679a0fe3728441494041b9-Paper.pdf"},{"key":"11242_CR40","unstructured":"Li H (2013) Error rate analysis of labeling by crowdsourcing. In: International conference on machine learning (ICML2013), Workshop on machine learning meets crowdsourcing"},{"key":"11242_CR41","doi-asserted-by":"crossref","unstructured":"Ling C, Sheng V (2010) Cost-sensitive learning and the class imbalance problem. Encycl Mach Learn","DOI":"10.1007\/978-0-387-30164-8_110"},{"key":"11242_CR42","unstructured":"Liu Q, Ihler A.T, Steyvers M (2013) Scoring workers in crowdsourcing: How many control questions are enough? In: Burges C.J, Bottou L, Welling M, Ghahramani Z, Weinberger K.Q. (eds.) Advances in neural information processing systems, vol. 26. https:\/\/proceedings.neurips.cc\/paper\/2013\/file\/cc1aa436277138f61cda703991069eaf-Paper.pdf"},{"key":"11242_CR43","doi-asserted-by":"publisher","unstructured":"Liu Y, Ott M, Goyal N, Du J, Joshi M, Chen D, Levy O, Lewis M, Zettlemoyer L, Stoyanov V (2019) Roberta: a robustly optimized bert pretraining approach. ArXiv abs\/1907.11692. https:\/\/doi.org\/10.48550\/ARXIV.1907.11692","DOI":"10.48550\/ARXIV.1907.11692"},{"key":"11242_CR44","unstructured":"Maas A.L, Daly R.E, Pham P.T, Huang D, Ng A.Y, Potts C (2011) Learning word vectors for sentiment analysis. In: Proceedings of the 49th Annual meeting of the association for computational linguistics: human language technologies, pp. 142\u2013150. Association for computational linguistics Portland, Oregon, USA. https:\/\/aclanthology.org\/P11-1015"},{"key":"11242_CR45","unstructured":"Nagar Y, Malone TW (2011) Making business predictions by combining human and machine intelligence in prediction markets. In: International conference on interaction sciences"},{"key":"11242_CR46","unstructured":"Nagar Y, Malone T.W (2012) Improving predictions with hybrid markets. In: AAAI fall symposium: machine aggregation of human judgment"},{"key":"11242_CR47","doi-asserted-by":"publisher","unstructured":"Ng AY (2004) Feature selection, l1 vs. l2 regularization, and rotational invariance. In: Proceedings of the Twenty-First International Conference on Machine Learning. ICML \u201904, p. 78. Association for Computing Machinery New York, NY, USA. https:\/\/doi.org\/10.1145\/1015330.1015435","DOI":"10.1145\/1015330.1015435"},{"key":"11242_CR48","unstructured":"Nu\u00f1ez A.C (2022) Combining diverse forms of human and machine intelligence. In: PhD thesis at Massachusetts institute of technology"},{"key":"11242_CR49","unstructured":"Qarout R, Checco A, Bontcheva K (2018) Investigating stability and reliability of crowdsourcing output. In: Proceedings of the 1st workshop on disentangling the relation between crowdsourcing and bias management (CrowdBias 2018) Co-located the 6th AAAI conference on human computation and crowdsourcing (HCOMP 2018)"},{"key":"11242_CR50","doi-asserted-by":"publisher","unstructured":"Qiu S, Gadiraju U, Bozzon A (2020) Improving worker engagement through conversational microtask crowdsourcing. In: Proceedings of the 2020 CHI conference on human factors in computing systems. CHI \u201920, pp. 1\u201312. Association for computing machinery New York, NY, USA. https:\/\/doi.org\/10.1145\/3313831.3376403","DOI":"10.1145\/3313831.3376403"},{"key":"11242_CR51","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.1910.10683","author":"C Raffel","year":"2020","unstructured":"Raffel C, Shazeer N, Roberts A, Lee K, Narang S, Matena M, Zhou Y, Li W, Liu PJ (2020) Exploring the limits of transfer learning with a unified text-to-text transformer. J Mach Learn Res. https:\/\/doi.org\/10.48550\/ARXIV.1910.10683","journal-title":"J Mach Learn Res"},{"key":"11242_CR52","unstructured":"Raghu M, Blumer K, Corrado G, Kleinberg J.M, Obermeyer Z, Mullainathan S (2019) The algorithmic automation problem: prediction, triage, and human effort. CoRR abs\/1903.12220. arXiv:1903.12220"},{"key":"11242_CR53","doi-asserted-by":"publisher","unstructured":"Rodriguez C, Daniel F, Casati F (2014) Crowd-based mining of reusable process model patterns. In: Business process management, pp. 51\u201366. https:\/\/doi.org\/10.1007\/978-3-319-10172-9_4","DOI":"10.1007\/978-3-319-10172-9_4"},{"key":"11242_CR54","doi-asserted-by":"publisher","unstructured":"Ruder S, Plank B (2018) Strong baselines for neural semi-supervised learning under domain shift. In: The 56th annual meeting of the association for computational linguistics (ACL 2018), pp. 1044\u20131054. https:\/\/doi.org\/10.18653\/v1\/P18-1096","DOI":"10.18653\/v1\/P18-1096"},{"key":"11242_CR55","doi-asserted-by":"publisher","first-page":"5283","DOI":"10.1007\/s10462-021-10021-3","volume":"54","author":"B Sayin","year":"2021","unstructured":"Sayin B, Krivosheev E, Passerini JYA, Casati F (2021) A review and experimental analysis of active learning over crowdsourced data. Artif Intel Rev 54:5283\u20135305. https:\/\/doi.org\/10.1007\/s10462-021-10021-3","journal-title":"Artif Intel Rev"},{"key":"11242_CR56","unstructured":"Sayin B, Casati F, Passerini A, Yang J, Chen X (2022) Rethinking and recomputing the value of ml models. arXiv preprint arXiv:2209.15157"},{"key":"11242_CR57","doi-asserted-by":"publisher","unstructured":"Sayin B, Krivosheev E, Ram\u00edrez J, Casati F, Taran E, Malanina V, Yang J (2021) Crowd-powered hybrid classification services: Calibration is all you need. In: 2021 IEEE International conference on web services (ICWS), pp. 42\u201350. https:\/\/doi.org\/10.1109\/ICWS53863.2021.00019","DOI":"10.1109\/ICWS53863.2021.00019"},{"key":"11242_CR58","doi-asserted-by":"publisher","unstructured":"Sayin B, Yang J, Passerini A, Casati F (2021) The science of rejection: a research area for human computation. In: The 9th AAAI conference on human computation and crowdsourcing. HCOMP 2021. https:\/\/doi.org\/10.48550\/ARXIV.2111.06736","DOI":"10.48550\/ARXIV.2111.06736"},{"key":"11242_CR59","doi-asserted-by":"crossref","unstructured":"Sayin B, Yang J, Passerini A, Casati F (2023) Value-aware active learning. In: Frontiers in artificial intelligence and applications. Volume 368: HHAI 2023: Augmenting Human Intellect, pp. 215\u2013223","DOI":"10.3233\/FAIA230085"},{"key":"11242_CR60","doi-asserted-by":"crossref","unstructured":"Sayin B, Yang J, Passerini A, Casati F (2023) Value-based hybrid intelligence. In: Frontiers in artificial intelligence and applications. Volume 368: HHAI 2023: Augmenting Human Intellect, pp. 366\u2013370","DOI":"10.3233\/FAIA230100"},{"issue":"3","key":"11242_CR61","doi-asserted-by":"publisher","first-page":"379","DOI":"10.1002\/j.1538-7305.1948.tb01338.x","volume":"27","author":"CE Shannon","year":"1948","unstructured":"Shannon CE (1948) A mathematical theory of communication. Bell Syst Tech J 27(3):379\u2013423. https:\/\/doi.org\/10.1002\/j.1538-7305.1948.tb01338.x","journal-title":"Bell Syst Tech J"},{"key":"11242_CR62","unstructured":"Sheng V.S, Ling C.X (2006) Thresholding for making classifiers cost-sensitive. In: Proceedings of the 21st National conference on artificial intelligence - Volume 1. AAAI\u201906, pp. 476\u2013481"},{"key":"11242_CR63","unstructured":"Silva\u00a0Filho T, Song H, Perell\u00f3-Nieto M, Santos-Rodr\u00edguez R, Kull M, Flach P.A (2021) Classifier calibration: How to assess and improve predicted class probabilities: a survey. CoRR abs\/2112.10327"},{"key":"11242_CR64","doi-asserted-by":"publisher","unstructured":"Suri M (2022) PiCkLe at SemEval-2022 task 4: Boosting pre-trained language models with task specific metadata and cost sensitive learning. In: Proceedings of the 16th International Workshop on Semantic Evaluation (SemEval-2022), pp. 464\u2013472. Association for Computational Linguistics Seattle, United States. https:\/\/doi.org\/10.18653\/v1\/2022.semeval-1.63","DOI":"10.18653\/v1\/2022.semeval-1.63"},{"key":"11242_CR65","doi-asserted-by":"crossref","unstructured":"Sutton R.T, Pincock D, Baumgart D.C, Sadowski D, Fedorak R, Kroeker K (2020) An overview of clinical decision support systems: benefits, risks, and strategies for success. npj Digital Medicine 3","DOI":"10.1038\/s41746-020-0221-y"},{"key":"11242_CR66","doi-asserted-by":"publisher","unstructured":"Teerapittayanon S, McDanel B, Kung H.T (2017) Branchynet: fast inference via early exiting from deep neural networks. ArXiv abs\/1709.01686. https:\/\/doi.org\/10.48550\/ARXIV.1709.01686","DOI":"10.48550\/ARXIV.1709.01686"},{"key":"11242_CR67","doi-asserted-by":"publisher","unstructured":"Thai-Nghe N, Gantner Z, Schmidt-Thieme L (2010) Cost-sensitive learning methods for imbalanced data. In: The 2010 international joint conference on neural networks (IJCNN), pp. 1\u20138. https:\/\/doi.org\/10.1109\/IJCNN.2010.5596486","DOI":"10.1109\/IJCNN.2010.5596486"},{"key":"11242_CR68","doi-asserted-by":"publisher","unstructured":"Ting KM (1998) Inducing cost-sensitive trees via instance weighting. In: Proceedings of the Second European Symposium on Principles of Data Mining and Knowledge Discovery. PKDD \u201998, pp. 139\u2013147. Springer Berlin, Heidelberg. https:\/\/doi.org\/10.1007\/BFb0094814","DOI":"10.1007\/BFb0094814"},{"key":"11242_CR69","unstructured":"Tomani C, Buettner F (2019) Towards trustworthy predictions from deep neural networks with fast adversarial calibration. In: AAAI conference on artificial intelligence"},{"key":"11242_CR70","doi-asserted-by":"publisher","unstructured":"Tu C.-Y, Lin H.-T (2020) Cost learning network for imbalanced classification. In: 2020 international conference on technologies and applications of artificial intelligence (TAAI), pp. 47\u201351. https:\/\/doi.org\/10.1109\/TAAI51410.2020.00017","DOI":"10.1109\/TAAI51410.2020.00017"},{"key":"11242_CR71","doi-asserted-by":"publisher","unstructured":"Waseem Z, Hovy D (2016) Hateful symbols or hateful people? predictive features for hate speech detection on Twitter. In: Proceedings of the NAACL Student Research Workshop, pp. 88\u201393. Association for Computational Linguistics San Diego, California. https:\/\/doi.org\/10.18653\/v1\/N16-2013. https:\/\/aclanthology.org\/N16-2013","DOI":"10.18653\/v1\/N16-2013"},{"key":"11242_CR72","unstructured":"Whitehill J, Wu T.-f, Bergsma J, Movellan J, Ruvolo P (2009) Whose vote should count more: Optimal integration of labels from labelers of unknown expertise. In: Bengio Y, Schuurmans D, Lafferty J, Williams C, Culotta A. (eds.) Advances in neural information processing systems, vol. 22. https:\/\/proceedings.neurips.cc\/paper\/2009\/file\/f899139df5e1059396431415e770c6dd-Paper.pdf"},{"key":"11242_CR73","doi-asserted-by":"publisher","unstructured":"Wilder B, Horvitz E, Kamar E (2021) Learning to complement humans. In: Proceedings of the Twenty-Ninth international joint conference on artificial intelligence. IJCAI\u201920. https:\/\/doi.org\/10.48550\/ARXIV.2005.00582","DOI":"10.48550\/ARXIV.2005.00582"},{"key":"11242_CR74","doi-asserted-by":"crossref","unstructured":"Wu M.-H, Quinn A.J (2017) Confusing the crowd: Task instruction quality on amazon mechanical turk. In: AAAI Conference on human computation & crowdsourcing","DOI":"10.1609\/hcomp.v5i1.13317"},{"key":"11242_CR75","unstructured":"Wu Y, Zeng Z, He K, Mou Y, Wang P, Xu W (2022) Distribution calibration for out-of-domain detection with Bayesian approximation. In: Proceedings of the 29th International Conference on Computational Linguistics, pp. 608\u2013615. International committee on computational linguistics Gyeongju, Republic of Korea. https:\/\/aclanthology.org\/2022.coling-1.50"},{"key":"11242_CR76","doi-asserted-by":"crossref","unstructured":"Yang J, Redi J, Demartini G, Bozzon A (2016) Modeling task complexity in crowdsourcing. In: AAAI conference on human computation & crowdsourcing","DOI":"10.1609\/hcomp.v4i1.13283"},{"key":"11242_CR77","doi-asserted-by":"publisher","unstructured":"Zadrozny B, Langford J, Abe N (2003) Cost-sensitive learning by cost-proportionate example weighting. In: Proceedings of the Third IEEE international conference on data mining. ICDM \u201903, p. 435. IEEE Computer Society USA. https:\/\/doi.org\/10.1109\/ICDM.2003.1250950","DOI":"10.1109\/ICDM.2003.1250950"},{"key":"11242_CR78","unstructured":"Zhou D, Basu S, Mao Y, Platt J (2012) Learning from the wisdom of crowds by minimax entropy. In: Pereira F, Burges C.J, Bottou L, Weinberger K.Q. (eds.) Advances in neural information processing systems, vol. 25. https:\/\/proceedings.neurips.cc\/paper\/2012\/file\/46489c17893dfdcf028883202cefd6d1-Paper.pdf"}],"container-title":["Artificial Intelligence Review"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10462-025-11242-6.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10462-025-11242-6\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10462-025-11242-6.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,23]],"date-time":"2025-06-23T06:35:47Z","timestamp":1750660547000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10462-025-11242-6"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,5,8]]},"references-count":78,"journal-issue":{"issue":"8","published-online":{"date-parts":[[2025,8]]}},"alternative-id":["11242"],"URL":"https:\/\/doi.org\/10.1007\/s10462-025-11242-6","relation":{"has-preprint":[{"id-type":"doi","id":"10.21203\/rs.3.rs-4833578\/v1","asserted-by":"object"}]},"ISSN":["1573-7462"],"issn-type":[{"value":"1573-7462","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,5,8]]},"assertion":[{"value":"17 April 2025","order":1,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"8 May 2025","order":2,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential Conflict of interest.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}}],"article-number":"238"}}