{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,7,30]],"date-time":"2025-07-30T16:31:28Z","timestamp":1753893088176,"version":"3.41.2"},"reference-count":34,"publisher":"Frontiers Media SA","license":[{"start":{"date-parts":[[2024,9,4]],"date-time":"2024-09-04T00:00:00Z","timestamp":1725408000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["frontiersin.org"],"crossmark-restriction":true},"short-container-title":["Front. Artif. Intell."],"abstract":"<jats:p>Training Deep Neural Networks (DNNs) places immense compute requirements on the underlying hardware platforms, expending large amounts of time and energy. An important factor contributing to the long training times is the increasing dataset complexity required to reach state-of-the-art performance in real-world applications. To address this challenge, we explore the use of input mixing, where multiple inputs are combined into a single composite input with an associated composite label for training. The goal is for training on the mixed input to achieve a similar effect as training separately on each the constituent inputs that it represents. This results in a lower number of inputs (or mini-batches) to be processed in each epoch, proportionally reducing training time. We find that naive input mixing leads to a considerable drop in learning performance and model accuracy due to interference between the forward\/backward propagation of the mixed inputs. We propose two strategies to address this challenge and realize training speedups from input mixing with minimal impact on accuracy. First, we reduce the impact of inter-input interference by exploiting the spatial separation between the features of the constituent inputs in the network's intermediate representations. We also adaptively vary the mixing ratio of constituent inputs based on their loss in previous epochs. Second, we propose heuristics to automatically identify the subset of the training dataset that is subject to mixing in each epoch. Across ResNets of varying depth, MobileNetV2 and two Vision Transformer networks, we obtain upto 1.6 \u00d7 and 1.8 \u00d7 speedups in training for the ImageNet and Cifar10 datasets, respectively, on an Nvidia RTX 2080Ti GPU, with negligible loss in classification accuracy.<\/jats:p>","DOI":"10.3389\/frai.2024.1387936","type":"journal-article","created":{"date-parts":[[2024,9,4]],"date-time":"2024-09-04T05:24:50Z","timestamp":1725427490000},"update-policy":"https:\/\/doi.org\/10.3389\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["MixTrain: accelerating DNN training via input mixing"],"prefix":"10.3389","volume":"7","author":[{"given":"Sarada","family":"Krithivasan","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Sanchari","family":"Sen","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Swagath","family":"Venkataramani","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Anand","family":"Raghunathan","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1965","published-online":{"date-parts":[[2024,9,4]]},"reference":[{"key":"B1","article-title":"\u201cExtremely large minibatch SGD: training resnet-50 on imagenet in 15 minutes,\u201d","author":"Akiba","year":"2017","journal-title":"CoRR, abs\/1711.04325"},{"journal-title":"Deep Neural Network Training Costs","year":"2018","author":"Amodei","key":"B2"},{"key":"B3","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2009.5206848","article-title":"\u201cImageNet: a large-scale hierarchical image database,\u201d","author":"Deng","year":"2009"},{"key":"B4","article-title":"\u201cAn image is worth 16x16 words: transformers for image recognition at scale,\u201d","author":"Dosovitskiy","year":"2020","journal-title":"CoRR, abs\/2010.11929"},{"key":"B5","article-title":"\u201cCPT: efficient deep neural network training via cyclic precision,\u201d","volume-title":"CoRR, abs\/2101.09868","author":"Fu","year":"2021"},{"key":"B6","article-title":"\u201cFractrain: Fractionally squeezing bit savings both temporally and spatially for efficient dnn training,\u201d","author":"Fu","year":"2020","journal-title":"Advances in Neural Information Processing Systems"},{"key":"B7","article-title":"\u201cAccurate, large minibatch SGD: training imagenet in 1 hour,\u201d","author":"Goyal","year":"2017","journal-title":"CoRR, abs\/1706.02677"},{"key":"B8","first-page":"181","article-title":"\u201cDeepcore: a comprehensive library forcoreset selection indeep learning,\u201d","author":"Guo","year":"2022","journal-title":"Database and Expert Systems Applications"},{"key":"B9","article-title":"\u201cDeep residual learning for image recognition,\u201d","author":"He","year":"2015","journal-title":"CoRR, abs\/1512.03385"},{"key":"B10","doi-asserted-by":"publisher","first-page":"1","DOI":"10.5555\/3546258.3546499","article-title":"Sparsity in deep learning: pruning and growth for efficient inference and training in neural networks","volume":"22","author":"Hoefler","year":"2021","journal-title":"J. Mach. Learn. Res"},{"key":"B11","article-title":"\u201cSubmodular combinatorial information measures with applications in machine learning,\u201d","author":"Iyer","year":"2020","journal-title":"CoRR, abs\/2006.15412"},{"key":"B12","article-title":"\u201cAccelerating deep learning by focusing on the biggest losers,\u201d","author":"Jiang","year":"2019","journal-title":"CoRR, abs\/1910.00762"},{"key":"B13","first-page":"5464","article-title":"\u201cGRAD-MATCH: gradient matching based data subset selection for efficient deep model training,\u201d","author":"Killamsetty","year":"2021","journal-title":"Proceedings of the 38th International Conference on Machine Learning, ICML 2021, 18-24 July 2021, Virtual Event"},{"key":"B14","article-title":"\u201cGLISTER: generalization based data subset selection for efficient and robust learning,\u201d","author":"Killamsetty","year":"2020","journal-title":"CoRR, abs\/2012.10630"},{"key":"B15","article-title":"\u201cAdam: a method for stochastic optimization,\u201d","author":"Kingma","year":"2015","journal-title":"3rd International Conference on Learning Representations, ICLR 2015"},{"journal-title":"Cifar-10 (Canadian Institute for Advanced Research","year":"2010","author":"Krizhevsky","key":"B16"},{"key":"B17","article-title":"\u201cPrunetrain: Gradual structured pruning from scratch for faster neural network training,\u201d","volume-title":"CoRR, abs\/1901.09290","author":"Lym","year":"2019"},{"key":"B18","article-title":"\u201cActive learning by acquiring contrastive examples,\u201d","author":"Margatina","year":"2021","journal-title":"CoRR, abs\/2109.03764"},{"key":"B19","first-page":"6950","article-title":"\u201cCoresets for data-efficient training of machine learning models,\u201d","author":"Mirzasoleiman","year":"2020","journal-title":"Proceedings of the 37th International Conference on Machine Learning"},{"key":"B20","first-page":"20596","article-title":"\u201cDeep learning on a data diet: finding important examples early in training,\u201d","author":"Paul","year":"2021","journal-title":"Advances in Neural Information Processing Systems"},{"key":"B21","article-title":"\u201cInverted residuals and linear bottlenecks: Mobile networks for classification, detection and segmentation,\u201d","author":"Sandler","year":"2018","journal-title":"CoRR, abs\/1801.04381"},{"key":"B22","article-title":"\u201cDomain-independent dominance of adaptive methods,\u201d","volume-title":"CoRR, abs\/1912.01823","author":"Savarese","year":"2019"},{"journal-title":"Imagenet Leaderboard","year":"2023","author":"Stojnic","key":"B23"},{"journal-title":"Energy and Policy Considerations for Deep Learning in NLP","year":"2019","author":"Strubell","key":"B24"},{"key":"B25","article-title":"\u201cHybrid 8-bit floating point (hfp8) training and inference for deep neural networks,\u201d","volume-title":"NeurIPS","author":"Sun","year":"2019"},{"key":"B26","article-title":"\u201cOn the importance of initialization and momentum in deep learning,\u201d","author":"Sutskever","year":"2013","journal-title":"Proceedings of the 30th International Conference on Machine Learning"},{"key":"B27","article-title":"\u201cEfficientnetv2: Smaller models and faster training,\u201d","author":"Tan","year":"2021","journal-title":"CoRR, abs\/2104.00298"},{"key":"B28","article-title":"\u201cFixing the train-test resolution discrepancy,\u201d","volume-title":"CoRR, abs\/1906.06423","author":"Touvron","year":"2019"},{"key":"B29","doi-asserted-by":"publisher","first-page":"3569","DOI":"10.1007\/s10994-023-06480-0","article-title":"Better schedules for low precision training of deep neural networks","volume":"113","author":"Wolfe","year":"2024","journal-title":"Mach Learn"},{"key":"B30","article-title":"\u201cScaling SGD batch size to 32k for imagenet training,\u201d","author":"You","year":"2017","journal-title":"CoRR, abs\/1708.03888"},{"key":"B31","article-title":"\u201cGrowing efficient deep networks by structured continuous sparsification,\u201d","author":"Yuan","year":"2020","journal-title":"CoRR, abs\/2007.15353"},{"key":"B32","article-title":"\u201cCutmix: Regularization strategy to train strong classifiers with localizable features,\u201d","author":"Yun","year":"2019","journal-title":"CoRR, abs\/1905.04899"},{"key":"B33","article-title":"\u201cmixup: Beyond empirical risk minimization,\u201d","author":"Zhang","year":"2017","journal-title":"CoRR, abs\/1710.09412"},{"key":"B34","article-title":"\u201cAutoassist: A framework to accelerate training of deep neural networks,\u201d","volume-title":"CoRR, abs\/1905.03381","author":"Zhang","year":"2019"}],"container-title":["Frontiers in Artificial Intelligence"],"original-title":[],"link":[{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/frai.2024.1387936\/full","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,10,1]],"date-time":"2024-10-01T12:59:29Z","timestamp":1727787569000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/frai.2024.1387936\/full"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,9,4]]},"references-count":34,"alternative-id":["10.3389\/frai.2024.1387936"],"URL":"https:\/\/doi.org\/10.3389\/frai.2024.1387936","relation":{},"ISSN":["2624-8212"],"issn-type":[{"type":"electronic","value":"2624-8212"}],"subject":[],"published":{"date-parts":[[2024,9,4]]},"article-number":"1387936"}}