{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,11,7]],"date-time":"2025-11-07T15:59:17Z","timestamp":1762531157071,"version":"build-2065373602"},"reference-count":62,"publisher":"Springer Science and Business Media LLC","issue":"11","license":[{"start":{"date-parts":[[2025,8,28]],"date-time":"2025-08-28T00:00:00Z","timestamp":1756339200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2025,8,28]],"date-time":"2025-08-28T00:00:00Z","timestamp":1756339200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"name":"Forschungszentrum J\u00fclich GmbH"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Knowl Inf Syst"],"published-print":{"date-parts":[[2025,11]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>\n                    In deep learning, using larger training datasets usually leads to more accurate models. However, simply adding more but redundant data may be inefficient, as some training samples may be more informative than others. We propose Focal Sampling, a method that biases SGD (Stochastic Gradient Descent) towards samples that are found to be more important after a few training epochs, by sampling them more often for the rest of the training. In contrast to state-of-the-art, our approach requires less computational overhead to estimate sample importance, as it computes estimates once during training using the prediction probabilities, and does not require restarting training. In the experimental evaluation, we see that our learning technique trains faster than state-of-the-art and can achieve higher test accuracy, especially when datasets are not well balanced or when using multiple data augmentations. Lastly, results suggest that our approach has intrinsic balancing properties and that balancing datasets based on class importance, rather than by number of samples, can achieve higher test accuracy. Code is available at\n                    <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"https:\/\/jugit.fz-juelich.de\/ias-8\/sgd_biased\" ext-link-type=\"uri\">https:\/\/jugit.fz-juelich.de\/ias-8\/sgd_biased<\/jats:ext-link>\n                    .\n                  <\/jats:p>","DOI":"10.1007\/s10115-025-02563-7","type":"journal-article","created":{"date-parts":[[2025,8,28]],"date-time":"2025-08-28T11:47:18Z","timestamp":1756381638000},"page":"11161-11191","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Focal Sampling: SGD biased towards early important samples for efficient image classification with augmentation selection"],"prefix":"10.1007","volume":"67","author":[{"given":"Alessio","family":"Quercia","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Fernanda","family":"Nader","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Abigail","family":"Morrison","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hanno","family":"Scharr","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ira","family":"Assent","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2025,8,28]]},"reference":[{"key":"2563_CR1","unstructured":"Schuhmann C, Beaumont R, Vencu R, Gordon CW, Wightman R, Cherti M, Coombes T, Katta A, Mullis C, Wortsman M, Schramowski P, Kundurthy SR, Crowson K, Schmidt L, Kaczmarczyk R, Jitsev J (2022) LAION-5b: an open large-scale dataset for training next generation image-text models, In: Thirty-sixth Conference on Neural Information Processing Systems Datasets and Benchmarks Track. [Online]. Available: https:\/\/openreview.net\/forum?id=M3Y74vmsMcY"},{"key":"2563_CR2","doi-asserted-by":"crossref","unstructured":"Cherti M, Beaumont R, Wightman R, Wortsman M, Ilharco G, Gordon C, Schuhmann C, Schmidt L, Jitsev J (2022) Reproducible scaling laws for contrastive language-image learning. [Online]. Available: https:\/\/arxiv.org\/abs\/2212.07143","DOI":"10.1109\/CVPR52729.2023.00276"},{"key":"2563_CR3","unstructured":"Toneva M, Sordoni A, Combes R\u00a0T\u00a0d, Trischler A, Bengio Y, Gordon GJ (2018) An empirical study of example forgetting during deep neural network learning, arXiv:1812.05159"},{"key":"2563_CR4","first-page":"2881","volume":"33","author":"V Feldman","year":"2020","unstructured":"Feldman V, Zhang C (2020) What neural networks memorize and why: discovering the long tail via influence estimation. Adv Neural Inf Process Syst 33:2881\u20132891","journal-title":"Adv Neural Inf Process Syst"},{"key":"2563_CR5","doi-asserted-by":"crossref","unstructured":"Feldman V (2020) Does learning require memorization? a short tale about a long tail, In: Proceedings of the 52nd Annual ACM SIGACT Symposium on Theory of Computing, pp. 954\u2013959","DOI":"10.1145\/3357713.3384290"},{"key":"2563_CR6","unstructured":"Jiang Z, Zhang C, Talwar K, Mozer MC (2020) Characterizing structural regularities of labeled data in overparameterized models, arXiv:2002.03206"},{"key":"2563_CR7","first-page":"19\u00a0920","volume":"33","author":"G Pruthi","year":"2020","unstructured":"Pruthi G, Liu F, Kale S, Sundararajan M (2020) Estimating training data influence by tracing gradient descent. Adv Neural Inf Process Syst 33:19\u00a0920-19\u00a0930","journal-title":"Adv Neural Inf Process Syst"},{"key":"2563_CR8","unstructured":"Paul M, Ganguli S, Dziugaite GK (2021) Deep learning on a data diet: finding important examples early in training, Advances in Neural Information Processing Systems, 34"},{"key":"2563_CR9","unstructured":"Sorscher B, Geirhos R, Shekhar S, Ganguli S, Morcos AS (2022) Beyond neural scaling laws: beating power law scaling via data pruning, arXiv:2206.14486"},{"key":"2563_CR10","unstructured":"Hacohen G, Weinshall D (2019) On the power of curriculum learning in training deep networks, In: International Conference on Machine Learning. PMLR, pp. 2535\u20132544"},{"key":"2563_CR11","unstructured":"Mindermann S, Brauner JM, Razzak MT, Sharma M, Kirsch A, Xu W, H\u00f6ltgen B, Gomez AN, Morisot A, Farquhar S et\u00a0al. (2022) Prioritized training on points that are learnable, worth learning, and not yet learnt, In: International Conference on Machine Learning. PMLR, pp. 15630\u201315649"},{"issue":"8","key":"2563_CR12","first-page":"6723","volume":"35","author":"S Banerjee","year":"2021","unstructured":"Banerjee S, Chakraborty S (2021) Deterministic mini-batch sequencing for training deep neural networks. Proceed AAAI Conf Artif Intell 35(8):6723\u20136731","journal-title":"Proceed AAAI Conf Artif Intell"},{"key":"2563_CR13","unstructured":"Needell D, Ward R, Srebro N (2014) Stochastic gradient descent, weighted sampling, and the randomized kaczmarz algorithm, Advances in neural information processing systems, 27"},{"key":"2563_CR14","unstructured":"Alain G, Lamb A, Sankar C, Courville A, Bengio Y (2015) Variance reduction in SGD by distributed importance sampling, arXiv:1511.06481"},{"key":"2563_CR15","unstructured":"El\u00a0Hanchi A, Stephens D, Maddison C (2022) Stochastic reweighted gradient descent, In: International Conference on Machine Learning. PMLR, pp. 8359\u20138374"},{"key":"2563_CR16","unstructured":"Zhao P, Zhang T (2015) Stochastic optimization with importance sampling for regularized loss minimization. In: international conference on machine learning. PMLR, pp. 1\u20139"},{"key":"2563_CR17","unstructured":"Schaul T, Quan J, Antonoglou I, Silver D (2015) Prioritized experience replay, arXiv:1511.05952"},{"key":"2563_CR18","unstructured":"Katharopoulos A, Fleuret F (2017) Biased importance sampling for deep neural network training, arXiv:1706.00043"},{"key":"2563_CR19","unstructured":"Johnson TB, Guestrin C (2018) Training deep models faster with robust, approximate importance sampling, Advances in Neural Information Processing Systems, 31"},{"key":"2563_CR20","unstructured":"Katharopoulos A, Fleuret F (2018) Not all samples are created equal: Deep learning with importance sampling, In: International conference on machine learning. PMLR, pp. 2525\u20132534"},{"key":"2563_CR21","unstructured":"Ganapathiraman V, Rodriguez FC, Joshi A (2022) Impon: efficient importance sampling with online regression for rapid neural network training"},{"key":"2563_CR22","unstructured":"Loshchilov I, Hutter F (2015) Online batch selection for faster training of neural networks, arXiv:1511.06343"},{"key":"2563_CR23","unstructured":"Simpson AJ (2015) oddball sgd: Novelty driven stochastic gradient descent for training deep neural networks, arXiv:1509.05765"},{"key":"2563_CR24","doi-asserted-by":"crossref","unstructured":"Shrivastava A, Gupta A, Girshick R (2016) Training region-based object detectors with online hard example mining, In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 761\u2013769","DOI":"10.1109\/CVPR.2016.89"},{"key":"2563_CR25","doi-asserted-by":"crossref","unstructured":"Lin T-Y, Goyal P, Girshick R, He K, Doll\u00e1r P (2017) Focal loss for dense object detection In: Proceedings of the IEEE international conference on computer vision, pp. 2980\u20132988","DOI":"10.1109\/ICCV.2017.324"},{"key":"2563_CR26","unstructured":"Chang H-S, Learned-Miller E, McCallum A (2017) Active bias: training more accurate neural networks by emphasizing high variance samples, Advances in Neural Information Processing Systems, 30"},{"key":"2563_CR27","unstructured":"Jiang A\u00a0H, Wong D\u00a0L-K, Zhou G, Andersen D\u00a0G, Dean J, Ganger G\u00a0R, Joshi G, Kaminksy M, Kozuch M, Lipton ZC, et\u00a0al. (2019) Accelerating deep learning by focusing on the biggest losers, arXiv:1910.00762"},{"key":"2563_CR28","unstructured":"Kawaguchi K, Lu H (2020) Ordered sgd: a new stochastic optimization framework for empirical risk minimization, In: International Conference on Artificial Intelligence and Statistics. PMLR, pp. 669\u2013679"},{"key":"2563_CR29","unstructured":"Dong C, Jin X, Gao W, Wang Y, Zhang H, Wu X, Yang J, Liu X (2021) One backward from ten forward, subsampling for large-scale deep learning arXiv:2104.13114"},{"key":"2563_CR30","unstructured":"Lu Y\u00a0S, Zamoshchin D, Chen Z (2022) Diet selective-backprop: Accelerating training in deep learning by pruning examples, Stanford University, Tech. Rep. [Online]. Available: http:\/\/cs231n.stanford.edu\/reports\/2022\/pdfs\/93p.pdf"},{"key":"2563_CR31","doi-asserted-by":"crossref","unstructured":"Bengio Y, Louradour J, Collobert R, Weston J (2009) Curriculum learning, In: Proceedings of the 26th annual international conference on machine learning, pp. 41\u201348","DOI":"10.1145\/1553374.1553380"},{"key":"2563_CR32","unstructured":"Jiang L, Zhou Z, Leung T, Li L-J, Fei-Fei L (2018) Mentornet: learning data-driven curriculum for very deep neural networks on corrupted labels, In: International conference on machine learning. PMLR, pp. 2304\u20132313"},{"key":"2563_CR33","first-page":"8602","volume":"33","author":"T Zhou","year":"2020","unstructured":"Zhou T, Wang S, Bilmes J (2020) Curriculum learning by dynamic instance hardness. Adv Neural Inf Process Syst 33:8602\u20138613","journal-title":"Adv Neural Inf Process Syst"},{"key":"2563_CR34","unstructured":"Zhou T, Wang S, Bilmes J (2021) Curriculum learning by optimizing learning dynamics, In: International conference on artificial intelligence and statistics. PMLR, pp. 433\u2013441"},{"key":"2563_CR35","unstructured":"Coleman C, Yeh C, Mussmann S, Mirzasoleiman B, Bailis P, Liang P, Leskovec J, Zaharia M (2019) Selection via proxy: efficient data selection for deep learning, arXiv:1906.11829"},{"key":"2563_CR36","unstructured":"Mirzasoleiman B, Bilmes J, Leskovec J (2020) Coresets for data-efficient training of machine learning models, In: International Conference on Machine Learning. PMLR, pp. 6950\u20136960"},{"key":"2563_CR37","doi-asserted-by":"crossref","unstructured":"Killamsetty K, Sivasubramanian D, Ramakrishnan G, Iyer R (2020) Glister: generalization based data subset selection for efficient and robust learning, arXiv:2012.10630","DOI":"10.1609\/aaai.v35i9.16988"},{"key":"2563_CR38","unstructured":"Killamsetty K, Durga S, Ramakrishnan G, De A, Iyer R (2021) Grad-match: gradient matching based data subset selection for efficient deep model training, In: International Conference on Machine Learning. PMLR, pp. 5464\u20135474"},{"key":"2563_CR39","doi-asserted-by":"crossref","unstructured":"Quercia A, Morrison A, Scharr H, Assent I (2023) Sgd biased towards early important samples for efficient training, In: 2023 IEEE International Conference on Data Mining (ICDM). IEEE","DOI":"10.1109\/ICDM58522.2023.00163"},{"key":"2563_CR40","first-page":"5850","volume":"33","author":"S Fort","year":"2020","unstructured":"Fort S, Dziugaite GK, Paul M, Kharaghani S, Roy DM, Ganguli S (2020) Deep learning versus kernel learning: an empirical study of loss landscape geometry and the time evolution of the neural tangent kernel. Adv Neural Inf Process Syst 33:5850\u20135861","journal-title":"Adv Neural Inf Process Syst"},{"key":"2563_CR41","unstructured":"Arpit D, Jastrzkebski S, Ballas N, Krueger D, Bengio E, Kanwal MS, Maharaj T, Fischer A, Courville A, Bengio Y et\u00a0al. (2017) A closer look at memorization in deep networks, In: International conference on machine learning. PMLR, pp. 233\u2013242"},{"key":"2563_CR42","unstructured":"Mangalam K, Prabhu VU (2019) Do deep neural networks learn shallow learnable examples first? in: ICML 2019 Workshop on Identifying and Understanding Deep Learning Phenomena. [Online]. Available: https:\/\/openreview.net\/forum?id=HkxHv4rn24"},{"key":"2563_CR43","first-page":"4308","volume":"33","author":"T Castells","year":"2020","unstructured":"Castells T, Weinzaepfel P, Revaud J (2020) Superloss: a generic loss for robust curriculum learning. Adv Neural Inf Process Syst 33:4308\u20134319","journal-title":"Adv Neural Inf Process Syst"},{"key":"2563_CR44","unstructured":"AdeelH (2020) Pytorch-multi-class-focal-loss, https:\/\/github.com\/AdeelH\/pytorch-multi-class-focal-loss\/blob\/master\/focal_loss.py"},{"key":"2563_CR45","first-page":"15\u00a0288","volume":"33","author":"J Mukhoti","year":"2020","unstructured":"Mukhoti J, Kulharia V, Sanyal A, Golodetz S, Torr P, Dokania P (2020) Calibrating deep neural networks using focal loss. Adv Neural Inf Process Syst 33:15\u00a0288-15\u00a0299","journal-title":"Adv Neural Inf Process Syst"},{"key":"2563_CR46","unstructured":"Krizhevsky A, Hinton G. et\u00a0al. (2009) Learning multiple layers of features from tiny images"},{"key":"2563_CR47","doi-asserted-by":"crossref","unstructured":"Van\u00a0Horn G, Cole E, Beery S, Wilber K, Belongie S, Mac\u00a0Aodha O (2021) Benchmarking representation learning for natural world image collections, In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp. 12884\u201312893","DOI":"10.1109\/CVPR46437.2021.01269"},{"key":"2563_CR48","unstructured":"Wightman R, Touvron H, J\u00e9gou H (2021) Resnet strikes back: an improved training procedure in timm, arXiv:2110.00476"},{"issue":"3","key":"2563_CR49","doi-asserted-by":"publisher","first-page":"211","DOI":"10.1007\/s11263-015-0816-y","volume":"115","author":"O Russakovsky","year":"2015","unstructured":"Russakovsky O, Deng J, Su H, Krause J, Satheesh S, Ma S, Huang Z, Karpathy A, Khosla A, Bernstein M et al (2015) Imagenet large scale visual recognition challenge. Int J Comput Vision 115(3):211\u2013252","journal-title":"Int J Comput Vision"},{"key":"2563_CR50","doi-asserted-by":"crossref","unstructured":"Cubuk ED, Zoph B, Mane D, Vasudevan V, Le QV (2019) Autoaugment: learning augmentation strategies from data, In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp. 113\u2013123","DOI":"10.1109\/CVPR.2019.00020"},{"key":"2563_CR51","doi-asserted-by":"crossref","unstructured":"M\u00fcller SG, Hutter F (2021) Trivialaugment: tuning-free yet state-of-the-art data augmentation, In: Proceedings of the IEEE\/CVF international conference on computer vision, pp. 774\u2013782","DOI":"10.1109\/ICCV48922.2021.00081"},{"key":"2563_CR52","doi-asserted-by":"publisher","first-page":"321","DOI":"10.1613\/jair.953","volume":"16","author":"NV Chawla","year":"2002","unstructured":"Chawla NV, Bowyer KW, Hall LO, Kegelmeyer WP (2002) Smote: synthetic minority over-sampling technique. J Artif Intell Res 16:321\u2013357","journal-title":"J Artif Intell Res"},{"key":"2563_CR53","doi-asserted-by":"crossref","unstructured":"Cubuk ED, Zoph B, Shlens J, Le QV (2020) Randaugment: practical automated data augmentation with a reduced search space. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition workshops, pp. 702\u2013703","DOI":"10.1109\/CVPRW50498.2020.00359"},{"issue":"4","key":"2563_CR54","doi-asserted-by":"publisher","first-page":"541","DOI":"10.1162\/neco.1989.1.4.541","volume":"1","author":"Y LeCun","year":"1989","unstructured":"LeCun Y, Boser B, Denker JS, Henderson D, Howard RE, Hubbard W, Jackel LD (1989) Backpropagation applied to handwritten zip code recognition. Neural Comput 1(4):541\u2013551","journal-title":"Neural Comput"},{"key":"2563_CR55","doi-asserted-by":"crossref","unstructured":"Szegedy C, Liu W, Jia Y, Sermanet P, Reed S, Anguelov D, Erhan D, Vanhoucke V, Rabinovich A (2015) Going deeper with convolutions. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 1\u20139","DOI":"10.1109\/CVPR.2015.7298594"},{"key":"2563_CR56","doi-asserted-by":"crossref","unstructured":"Han D, Kim J, Kim J (2017) Deep pyramidal residual networks. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 5927\u20135935","DOI":"10.1109\/CVPR.2017.668"},{"key":"2563_CR57","doi-asserted-by":"crossref","unstructured":"Han D, Kim J, Kim J (2017) Deep pyramidal residual networks, IEEE CVPR","DOI":"10.1109\/CVPR.2017.668"},{"key":"2563_CR58","unstructured":"Paszke A, Gross S, Chintala S, Chanan G, Yang E, DeVito Z, Lin Z, Desmaison A, Antiga L, Lerer A (2017) Automatic differentiation in pytorch"},{"key":"2563_CR59","unstructured":"Falcon W and The PyTorch Lightning team (2019) PyTorch Lightning, 3. [Online]. Available: https:\/\/github.com\/PyTorchLightning\/pytorch-lightning"},{"key":"2563_CR60","unstructured":"Yadan O (2019) Hydra - a framework for elegantly configuring complex applications, Github. [Online]. Available: https:\/\/github.com\/facebookresearch\/hydra"},{"key":"2563_CR61","doi-asserted-by":"crossref","unstructured":"Kesselheim S, Herten A, Krajsek K, Ebert J, Jitsev J, Cherti M, Langguth M, Gong B, Stadtler S, Mozaffari A et\u00a0al. (2021) Juwels booster\u2013a supercomputer for large-scale ai research, In: International conference on high performance computing. Springer, pp. 453\u2013468","DOI":"10.1007\/978-3-030-90539-2_31"},{"key":"2563_CR62","doi-asserted-by":"publisher","first-page":"A182","DOI":"10.17815\/jlsrf-7-182","volume":"7","author":"P Th\u00f6rnig","year":"2021","unstructured":"Th\u00f6rnig P (2021) Jureca: data centric and booster modules implementing the modular supercomputing architecture at j\u00fclich supercomputing centre. J Large-scale Res Facil JLSRF 7:A182\u2013A182","journal-title":"J Large-scale Res Facil JLSRF"}],"container-title":["Knowledge and Information Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10115-025-02563-7.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10115-025-02563-7\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10115-025-02563-7.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,11,7]],"date-time":"2025-11-07T15:51:52Z","timestamp":1762530712000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10115-025-02563-7"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,8,28]]},"references-count":62,"journal-issue":{"issue":"11","published-print":{"date-parts":[[2025,11]]}},"alternative-id":["2563"],"URL":"https:\/\/doi.org\/10.1007\/s10115-025-02563-7","relation":{},"ISSN":["0219-1377","0219-3116"],"issn-type":[{"type":"print","value":"0219-1377"},{"type":"electronic","value":"0219-3116"}],"subject":[],"published":{"date-parts":[[2025,8,28]]},"assertion":[{"value":"9 July 2024","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"22 July 2025","order":2,"name":"revised","label":"Revised","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"24 July 2025","order":3,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"28 August 2025","order":4,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare that they have no conflict of interest.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}},{"value":"The main effect of our work is to speed up the learning process and automatically balance classes. Speedup reduces energy consumption and potentially allows for using less powerful hardware for training, a tiny step towards democratizing AI. Balancing classes may help to emphasize underrepresented samples and thus may raise visibility of minorities in data. It depends on the target application if this results in highly desired fair treatment of minorities or an unfair deviation from underlying distributions.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethical approval"}},{"value":"All presented results are reproducible using our code, which is available as an anonymous repository at\n                      \n                      .","order":4,"name":"Ethics","group":{"name":"EthicsHeading","label":"Reproducibility Statement"}}]}}