{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,17]],"date-time":"2026-07-17T14:59:51Z","timestamp":1784300391016,"version":"3.55.0"},"reference-count":51,"publisher":"Springer Science and Business Media LLC","issue":"4","license":[{"start":{"date-parts":[[2024,4,30]],"date-time":"2024-04-30T00:00:00Z","timestamp":1714435200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2024,4,30]],"date-time":"2024-04-30T00:00:00Z","timestamp":1714435200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"name":"Federal Ministry of Education and Research of Germany and the state of North Rhine-Westphalia","award":["Lamarr Institute for Machine Learning and Artificial Intelligence"],"award-info":[{"award-number":["Lamarr Institute for Machine Learning and Artificial Intelligence"]}]},{"DOI":"10.13039\/501100016378","name":"Technische Universit\u00e4t Dortmund","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100016378","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Data Min Knowl Disc"],"published-print":{"date-parts":[[2024,7]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>With machine learning (ML) becoming a popular tool across all domains, practitioners are in dire need of comprehensive reporting on the state-of-the-art. Benchmarks and open databases provide helpful insights for many tasks, however suffer from several phenomena: Firstly, they overly focus on prediction quality, which is problematic considering the demand for more sustainability in ML. Depending on the use case at hand, interested users might also face tight resource constraints and thus should be allowed to interact with reporting frameworks, in order to prioritize certain reported characteristics. Furthermore, as some practitioners might not yet be well-skilled in ML, it is important to convey information on a more abstract, comprehensible level. Usability and extendability are key for moving with the state-of-the-art and in order to be trustworthy, frameworks should explicitly address reproducibility. In this work, we analyze established reporting systems under consideration of the aforementioned issues. Afterwards, we propose STREP, our novel framework that aims at overcoming these shortcomings and paves the way towards more sustainable and trustworthy reporting. We use STREP\u2019s (publicly available) implementation to investigate various existing report databases. Our experimental results unveil the need for making reporting more resource-aware and demonstrate our framework\u2019s capabilities of overcoming current reporting limitations. With our work, we want to initiate a paradigm shift in reporting and help with making ML advances more considerate of sustainability and trustworthiness.<\/jats:p>","DOI":"10.1007\/s10618-024-01020-3","type":"journal-article","created":{"date-parts":[[2024,4,30]],"date-time":"2024-04-30T07:02:19Z","timestamp":1714460539000},"page":"1909-1928","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":11,"title":["Towards more sustainable and trustworthy reporting in machine learning"],"prefix":"10.1007","volume":"38","author":[{"given":"Raphael","family":"Fischer","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Thomas","family":"Liebig","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Katharina","family":"Morik","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2024,4,30]]},"reference":[{"issue":"4\/5","key":"1020_CR1","doi-asserted-by":"publisher","first-page":"6","DOI":"10.1147\/JRD.2019.2942288","volume":"63","author":"M Arnold","year":"2019","unstructured":"Arnold M, Bellamy RK, Hind M, Houde S, Mehta S, Mojsilovi\u0107 A, Nair R, Ramamurthy KN, Olteanu A, Piorkowski D et al (2019) Factsheets: Increasing trust in ai services through supplier\u2019s declarations of conformity. IBM J Res Dev 63(4\/5):6\u20131","journal-title":"IBM J Res Dev"},{"key":"1020_CR2","doi-asserted-by":"crossref","unstructured":"Avin S, Belfield H, Brundage M, Krueger G, Wang J et\u00a0al (2021) Filling gaps in trustworthy development of AI. Science 374(6573):1327\u20131329. American Association for the Advancement of Science","DOI":"10.1126\/science.abi7176"},{"issue":"1","key":"1020_CR3","doi-asserted-by":"publisher","first-page":"12","DOI":"10.1007\/s13347-022-00510-w","volume":"35","author":"K Baum","year":"2022","unstructured":"Baum K, Mantel S, Schmidt E, Speith T (2022) From responsibility to reason-giving explainable artificial intelligence. Philos Technol 35(1):12","journal-title":"Philos Technol"},{"key":"1020_CR4","doi-asserted-by":"crossref","unstructured":"Beckh K, M\u00fcller S, Jakobs M, Toborek V, Tan H, Fischer R, Welke P, Houben S, Rueden L (2023) Harnessing prior knowledge for explainable machine learning: An overview. In: First IEEE conference on secure and trustworthy machine learning","DOI":"10.1109\/SaTML54575.2023.00038"},{"key":"1020_CR5","doi-asserted-by":"publisher","unstructured":"Bender EM, Gebru T, McMillan-Major A, Shmitchell S (2021) On the dangers of stochastic parrots: can language models be too big? In: Conference on fairness, accountability, and transparency, pp 610\u2013623. https:\/\/doi.org\/10.1145\/3442188.3445922","DOI":"10.1145\/3442188.3445922"},{"key":"1020_CR6","doi-asserted-by":"crossref","unstructured":"Buschj\u00e4ger S, Pfahler L, Buss J, Morik K, Rhode W (2020) On-site gamma-hadron separation with deep learning on fpgas. In: European conference on machine learning and knowledge discovery in databases, pp 478\u2013493","DOI":"10.1007\/978-3-030-67667-4_29"},{"key":"1020_CR7","doi-asserted-by":"crossref","unstructured":"Casta\u00f1o J, Mart\u00ednez-Fern\u00e1ndez S, Franch X, Bogner J (2023) Exploring the carbon footprint of hugging face\u2019s ML models: a repository mining study. _eprint: arXiv:2305.11164","DOI":"10.1109\/ESEM56168.2023.10304801"},{"key":"1020_CR8","doi-asserted-by":"crossref","unstructured":"Chatila R, Dignum V, Fisher M, Giannotti F, Morik K, Russell S, Yeung K (2021) Trustworthy ai. Reflections on artificial intelligence for humanity, pp 13\u201339. Springer","DOI":"10.1007\/978-3-030-69128-8_2"},{"key":"1020_CR9","unstructured":"Croce F, Andriushchenko M, Sehwag V, Debenedetti E, Flammarion N, Chiang M, Mittal P, Hein M (2020) Robustbench: a standardized adversarial robustness benchmark. Preprint arXiv:2010.09670"},{"key":"1020_CR10","doi-asserted-by":"publisher","first-page":"81555","DOI":"10.1109\/ACCESS.2019.2923736","volume":"7","author":"W Cui","year":"2019","unstructured":"Cui W (2019) Visual analytics: a comprehensive overview. IEEE Access 7:81555\u201381573. https:\/\/doi.org\/10.1109\/ACCESS.2019.2923736","journal-title":"IEEE Access"},{"key":"1020_CR11","unstructured":"Dabbas E (2021) Interactive dashboards and data apps with plotly and dash"},{"key":"1020_CR12","unstructured":"Dems\u0306ar J (2006) Statistical comparisons of classifiers over multiple data sets. J Mach Learn Res 7:1\u201330. JMLR. org"},{"key":"1020_CR13","unstructured":"Dems\u0306ar J, Curk T, Erjavec A, Gorup U, Hoc\u0306evar T et\u00a0al (2013) Orange: data mining toolbox in Python. J Mach Learn Res 14(1):2349\u20132353. JMLR. org"},{"key":"1020_CR14","doi-asserted-by":"publisher","unstructured":"Dignum V (2019) Responsible artificial intelligence: how to develop and use AI in a responsible way. https:\/\/doi.org\/10.1007\/978-3-030-30371-6","DOI":"10.1007\/978-3-030-30371-6"},{"key":"1020_CR15","unstructured":"EU AI HLEG (2020) Assessment List for Trustworthy Artificial Intelligence (ALTAI) for self-assessment. https:\/\/futurium.ec.europa.eu\/en\/european-ai-alliance\/pages\/altai-assessment-list-trustworthy-artificial-intelligence"},{"key":"#cr-split#-1020_CR16.1","unstructured":"European Commission (2019) Commission Delegated Regulation"},{"key":"#cr-split#-1020_CR16.2","unstructured":"(EU) 2019\/2014 with regard to energy labelling of household washing machines and household washer-dryers. https:\/\/eur-lex.europa.eu\/legal-content\/EN\/ALL\/?uri=CELEX:32019R2014"},{"key":"1020_CR17","unstructured":"European Parliament (2023) A step closer to the first rules on artificial intelligence. European Parliament News. https:\/\/www.europarl.europa.eu\/news\/en\/press-room\/20230505IPR84904\/ai-act-a-step-closer-to-the-first-rules-on-artificial-intelligence"},{"key":"1020_CR18","unstructured":"Feurer M, Rijn JNv, Kadra A, Gijsbers P, Mallik N et al (2021) OpenML-Python: an extensible Python API for OpenML. J Mach Learn Res 22(100):1\u20135"},{"key":"1020_CR19","first-page":"39","volume-title":"Machine learning and principles and practice of knowledge discovery in databases","author":"R Fischer","year":"2022","unstructured":"Fischer R, Jakobs M, M\u00fccke S, Morik K (2022) A unified framework for assessing energy efficiency of machine learning. Machine learning and principles and practice of knowledge discovery in databases. Springer, Cham, pp 39\u201354"},{"key":"1020_CR20","doi-asserted-by":"crossref","unstructured":"Fischer R, Pauly A, Wilking R, Kini A, Graurock D (2023) Prioritization of identified data science use cases in industrial manufacturing via C-EDIF scoring. In: IEEE international conference on data science and advanced analytics, pp 1\u20134","DOI":"10.1109\/DSAA60987.2023.10302632"},{"key":"1020_CR21","unstructured":"Fischer R, Saadallah A (2023) AutoXPCR: Automated multi-objective model selection for time series forecasting. Preprint arXiv:2312.13038"},{"key":"1020_CR22","doi-asserted-by":"publisher","unstructured":"Fischer R, van der Staay A, Buschj\u00e4ger S (2024) Stress-testing USB accelerators for efficient edge inference. Research Square preprint. https:\/\/doi.org\/10.21203\/rs.3.rs-3793927","DOI":"10.21203\/rs.3.rs-3793927"},{"key":"1020_CR23","unstructured":"Godahewa R, Bergmeir C, Webb GI, Hyndman RJ, Montero-Manso P (2021) Monash time series forecasting archive. In: Neural information processing systems track on datasets and benchmarks. forthcoming"},{"key":"1020_CR24","doi-asserted-by":"crossref","unstructured":"Hauer MP, Krafft TD, Zweig K (2023) Overview of transparency and inspectability mechanisms to achieve accountability of artificial intelligence systems. Data Policy 5:36. Cambridge University Press","DOI":"10.1017\/dap.2023.30"},{"key":"1020_CR25","unstructured":"Howard AG, Zhu M, Chen B, Kalenichenko D, Wang W et\u00a0al (2017) Mobilenets: Efficient convolutional neural networks for mobile vision applications. Preprint arXiv:1704.04861"},{"key":"1020_CR26","doi-asserted-by":"publisher","unstructured":"Hutson M (2018) Artificial intelligence faces reproducibility crisis. Science 359(6377):725\u2013726. https:\/\/doi.org\/10.1126\/science.359.6377.725. _eprint: https:\/\/www.science.org\/doi\/pdf\/10.1126\/science.359.6377.725","DOI":"10.1126\/science.359.6377.725"},{"key":"1020_CR27","unstructured":"Ismail-Fawaz A, Dempster A, Tan CW, Herrmann M, Miller L et\u00a0al (2023) An approach to multiple comparison benchmark evaluations that is stable under manipulation of the comparate set. Preprint arXiv:2305.11921"},{"key":"1020_CR28","doi-asserted-by":"publisher","unstructured":"Jain S (2022) Hugging face, pp 51\u201367. https:\/\/doi.org\/10.1007\/978-1-4842-8844-3_4","DOI":"10.1007\/978-1-4842-8844-3_4"},{"issue":"6","key":"1020_CR29","doi-asserted-by":"publisher","first-page":"103477","DOI":"10.1016\/j.ipm.2023.103477","volume":"60","author":"D Kang","year":"2023","unstructured":"Kang D, Kang T, Jang J (2023) Papers with code or without code? Impact of GitHub repository usability on the diffusion of machine learning research. Inf Process Manag 60(6):103477. https:\/\/doi.org\/10.1016\/j.ipm.2023.103477","journal-title":"Inf Process Manag"},{"key":"1020_CR30","doi-asserted-by":"crossref","unstructured":"Kar AK, Choudhary SK, Singh VK (2022) How can artificial intelligence impact sustainability: A systematic literature review. J Clean Prod 134120. Elsevier","DOI":"10.1016\/j.jclepro.2022.134120"},{"key":"1020_CR31","unstructured":"Lacoste A, Luccioni A, Schmidt V, Dandres T (2019) Quantifying the carbon emissions of machine learning. Preprint arXiv:1910.09700"},{"key":"1020_CR32","doi-asserted-by":"publisher","unstructured":"Marwedel P, Morik K (2022) Machine learning under resource constraints - volume 1: fundamentals. https:\/\/doi.org\/10.1515\/9783110785944","DOI":"10.1515\/9783110785944"},{"key":"1020_CR33","doi-asserted-by":"crossref","unstructured":"Mierswa I, Wurst M, Klinkenberg R, Scholz M, Euler T (2006) YALE: rapid prototyping for complex data mining tasks. In: ACM SIGKDD international conference on knowledge discovery and data mining (KDD 2006), pp 935\u2013940. ACM Press, New York, USA. ACM. http:\/\/rapid-i.com\/component\/option,com_docman\/task,doc_download\/gid,25\/Itemid,62\/","DOI":"10.1145\/1150402.1150531"},{"key":"1020_CR34","doi-asserted-by":"publisher","unstructured":"Mitchell M, Wu S, Zaldivar A, Barnes P, Vasserman L et\u00a0al (2019) Model cards for model reporting. In: Proceedings of the conference on fairness, accountability, and transparency, FAT* 2019, pp 220\u2013229. Association for Computing Machinery, New York, NY, USA. https:\/\/doi.org\/10.1145\/3287560.3287596","DOI":"10.1145\/3287560.3287596"},{"key":"1020_CR35","doi-asserted-by":"publisher","unstructured":"Morik KJ, Kotthaus H, Fischer R, M\u00fccke S, Jakobs M, Piatkowski N, Pauly A, Heppe L, Heinrich D (2022) Yes we care!-certification for machine learning methods through the care label framework. Front Artif Intell 5. https:\/\/doi.org\/10.3389\/frai.2022.975029","DOI":"10.3389\/frai.2022.975029"},{"issue":"1","key":"1020_CR36","doi-asserted-by":"publisher","first-page":"11","DOI":"10.1007\/s42484-023-00099-z","volume":"5","author":"S M\u00fccke","year":"2023","unstructured":"M\u00fccke S, Heese R, M\u00fcller S, Wolter M, Piatkowski N (2023) Feature selection on quantum computers. Quantum Mach Intell 5(1):11","journal-title":"Quantum Mach Intell"},{"key":"1020_CR37","unstructured":"Patterson D, Gonzalez J, Le Q, Liang C, Munguia L-M, Rothchild D, So D, Texier M, Dean J (2021) Carbon emissions and large neural network training. Preprint arXiv:2104.10350"},{"key":"1020_CR38","unstructured":"Pineau J, Vincent-Lamarre P, Sinha K, Larivi\u00e8re V, Beygelzimer A et\u00a0al (2021) Improving reproducibility in machine learning research (a report from the neurips 2019 reproducibility program). J Mach Learn Res 22(1):7459\u20137478. JMLRORG"},{"key":"1020_CR39","doi-asserted-by":"crossref","unstructured":"Piorkowski D, Park S, Wang AY, Wang D, Muller M, Portnoy F (2021) How ai developers overcome communication challenges in a multidisciplinary team: A case study. Proceedings of the ACM on human-computer interaction 5(CSCW1), pp 1\u201325. ACM New York, NY, USA","DOI":"10.1145\/3449205"},{"key":"1020_CR40","doi-asserted-by":"crossref","unstructured":"Sakaguchi K, Bras RL, Bhagavatula C, Choi Y (2021) Winogrande: An adversarial winograd schema challenge at scale. Commun ACM 64(9):99\u2013106. ACM New York, NY, USA","DOI":"10.1145\/3474381"},{"issue":"12","key":"1020_CR41","doi-asserted-by":"publisher","first-page":"54","DOI":"10.1145\/3381831","volume":"63","author":"R Schwartz","year":"2020","unstructured":"Schwartz R, Dodge J, Smith NA, Etzioni O (2020) Green AI. Commun ACM 63(12):54\u201363","journal-title":"Commun ACM"},{"key":"1020_CR42","unstructured":"Srivastava A, Rastogi A, Rao A, Shoeb AAM, Abid A et\u00a0al (2022) Beyond the imitation game: quantifying and extrapolating the capabilities of language models. Preprint arXiv:2206.04615"},{"key":"1020_CR43","unstructured":"Stojnic R, Taylor R, Kardas M, Saravia E, Cucurull G, Westbury A, Scialom T (2018) Papers With Code - The latest in Machine Learning. https:\/\/paperswithcode.com\/"},{"key":"1020_CR44","doi-asserted-by":"crossref","unstructured":"Strubell E, Ganesh A, McCallum A (2020) Energy and Policy Considerations for Modern Deep Learning Research. In: AAAI conference on artificial intelligence, pp 13693\u201313696","DOI":"10.1609\/aaai.v34i09.7123"},{"key":"1020_CR45","doi-asserted-by":"publisher","unstructured":"Sun X, Zhou T, Li G, Hu J, Yang H, Li B (2017) An Empirical Study on Real Bugs for Machine Learning Programs. In: 2017 24th Asia-Pacific software engineering conference (APSEC), pp 348\u2013357. https:\/\/doi.org\/10.1109\/APSEC.2017.41","DOI":"10.1109\/APSEC.2017.41"},{"key":"1020_CR46","doi-asserted-by":"publisher","unstructured":"The pandas development team (2022) pandas-dev\/pandas: Pandas 1.4.1. Zenodo. https:\/\/doi.org\/10.5281\/zenodo.6053272","DOI":"10.5281\/zenodo.6053272"},{"issue":"2","key":"1020_CR47","doi-asserted-by":"publisher","first-page":"49","DOI":"10.1145\/2641190.2641198","volume":"15","author":"J Vanschoren","year":"2014","unstructured":"Vanschoren J, Van Rijn JN, Bischl B, Torgo L (2014) Openml: networked science in machine learning. ACM SIGKDD Explor Newslett 15(2):49\u201360","journal-title":"ACM SIGKDD Explor Newslett"},{"key":"1020_CR48","unstructured":"Wang A, Pruksachatkun Y, Nangia N, Singh A, Michael J, Hill F, Levy O, Bowman S (2019) Superglue: A stickier benchmark for general-purpose language understanding systems. Advances in neural information processing systems 32"},{"issue":"3","key":"1020_CR49","doi-asserted-by":"publisher","first-page":"213","DOI":"10.1007\/s43681-021-00043-6","volume":"1","author":"A Wynsberghe","year":"2021","unstructured":"Wynsberghe A (2021) Sustainable AI: AI for sustainability and the sustainability of AI. AI Ethics 1(3):213\u2013218. https:\/\/doi.org\/10.1007\/s43681-021-00043-6","journal-title":"AI Ethics"},{"issue":"4","key":"1020_CR50","first-page":"39","volume":"41","author":"M Zaharia","year":"2018","unstructured":"Zaharia M, Chen A, Davidson A, Ghodsi A, Hong SA et al (2018) Accelerating the machine learning lifecycle with MLflow. IEEE Data Eng Bull 41(4):39\u201345","journal-title":"IEEE Data Eng Bull"}],"container-title":["Data Mining and Knowledge Discovery"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10618-024-01020-3.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10618-024-01020-3\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10618-024-01020-3.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,7,30]],"date-time":"2024-07-30T10:35:22Z","timestamp":1722335722000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10618-024-01020-3"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,4,30]]},"references-count":51,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2024,7]]}},"alternative-id":["1020"],"URL":"https:\/\/doi.org\/10.1007\/s10618-024-01020-3","relation":{},"ISSN":["1384-5810","1573-756X"],"issn-type":[{"value":"1384-5810","type":"print"},{"value":"1573-756X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,4,30]]},"assertion":[{"value":"6 December 2023","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"25 March 2024","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"30 April 2024","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare no competing interests.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}]}}