{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,1,29]],"date-time":"2025-01-29T05:42:36Z","timestamp":1738129356536,"version":"3.33.0"},"reference-count":79,"publisher":"Springer Science and Business Media LLC","issue":"10","license":[{"start":{"date-parts":[[2025,1,28]],"date-time":"2025-01-28T00:00:00Z","timestamp":1738022400000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2025,1,28]],"date-time":"2025-01-28T00:00:00Z","timestamp":1738022400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Front. Comput. Sci."],"published-print":{"date-parts":[[2025,10]]},"abstract":"<jats:title>Abstract<\/jats:title>\n          <jats:p>Low-precision training has emerged as a practical approach, saving the cost of time, memory, and energy during deep neural networks (DNNs) training. Typically, the use of lower precision introduces quantization errors that need to be minimized to maintain model performance, often neglecting to consider the potential benefits of reducing training precision. This paper rethinks low-precision training, highlighting the potential benefits of lowering precision: (1) low precision can serve as a form of regularization in DNN training by constraining excessive variance in the model; (2) layer-wise low precision can be seen as an alternative dimension of sparsity, orthogonal to pruning, contributing to improved generalization in DNNs. Based on these analyses, we propose a simple yet powerful technique\u2013DPC (Decreasing Precision with layer Capacity), which directly assigns different bit-widths to model layers, without the need for an exhaustive analysis of the training process or any delicate low-precision criteria. Thorough extensive experiments on five datasets and fourteen models across various applications consistently demonstrate the effectiveness of the proposed DPC technique in saving computational cost (\u221216.21%\u2013\u221244.37%) while achieving comparable or even superior accuracy (up to +0.68%, +0.21% on average). Furthermore, we offer feature embedding visualizations and conduct further analysis with experiments to investigate the underlying mechanisms behind DPC\u2019s effectiveness, enhancing our understanding of low-precision training. Our source code will be released upon paper acceptance.<\/jats:p>","DOI":"10.1007\/s11704-024-40669-3","type":"journal-article","created":{"date-parts":[[2025,1,28]],"date-time":"2025-01-28T13:14:54Z","timestamp":1738070094000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Efficient deep neural network training via decreasing precision with layer capacity"],"prefix":"10.1007","volume":"19","author":[{"given":"Ao","family":"Shen","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zhiquan","family":"Lai","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Tao","family":"Sun","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Shengwei","family":"Li","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Keshi","family":"Ge","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Weijie","family":"Liu","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Dongsheng","family":"Li","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2025,1,28]]},"reference":[{"key":"40669_CR1","first-page":"279","volume-title":"Proceedings of the 38th IEEE International Conference on Computer Design (ICCD)","author":"J Du","year":"2020","unstructured":"Du J, Shen M, Du Y. A distributed in-situ CNN inference system for IOT applications. In: Proceedings of the 38th IEEE International Conference on Computer Design (ICCD). 2020, 279\u2013287"},{"key":"40669_CR2","volume-title":"Proceedings of the 10th International Conference on Learning Representations","author":"C Yang","year":"2022","unstructured":"Yang C, Wu Z, Chee J, De Sa C, Udell M. How low can we go: trading memory for error in low-precision training. In: Proceedings of the 10th International Conference on Learning Representations. 2022"},{"key":"40669_CR3","first-page":"1966","volume-title":"Proceedings of 2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"F Zhu","year":"2020","unstructured":"Zhu F, Gong R, Yu F, Liu X, Wang Y, Li Z, Yang X, Yan J. Towards unified int8 training for convolutional neural network. In: Proceedings of 2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 2020, 1966\u20131976"},{"key":"40669_CR4","volume-title":"Proceedings of the 6th International Conference on Learning Representations","author":"P Micikevicius","year":"2018","unstructured":"Micikevicius P, Narang S, Alben J, Diamos G, Elsen E, Garcia D, Ginsburg B, Houston M, Kuchaiev O, Venkatesh G, Wu H. Mixed precision training. In: Proceedings of the 6th International Conference on Learning Representations. 2018"},{"key":"40669_CR5","volume-title":"DoReFa-Net: training low bitwidth convolutional neural networks with low bitwidth gradients","author":"S Zhou","year":"2016","unstructured":"Zhou S, Wu Y, Ni Z, Zhou X, Wen H, Zou Y. DoReFa-Net: training low bitwidth convolutional neural networks with low bitwidth gradients. 2016, arXiv preprint arXiv: 1606.06160"},{"key":"40669_CR6","first-page":"10612","volume-title":"Proceedings of the 35th AAAI Conference on Artificial Intelligence","author":"L Yang","year":"2021","unstructured":"Yang L, Jin Q. FracBits: mixed precision quantization via fractional bit-widths. In: Proceedings of the 35th AAAI Conference on Artificial Intelligence. 2021, 10612\u201310620"},{"issue":"5","key":"40669_CR7","doi-asserted-by":"publisher","first-page":"37","DOI":"10.1109\/MM.2020.3009475","volume":"40","author":"A T Elthakeb","year":"2020","unstructured":"Elthakeb A T, Pilligundla P, Mireshghallah F, Yazdanbakhsh A, Esmaeilzadeh H. ReLeQ: a reinforcement learning approach for automatic deep quantization of neural networks. IEEE Micro, 2020, 40(5): 37\u201345.","journal-title":"IEEE Micro"},{"key":"40669_CR8","volume-title":"Proceedings of the 9th International Conference on Learning Representations","author":"H Yang","year":"2021","unstructured":"Yang H, Duan L, Chen Y, Li H. BSQ: exploring bit-level sparsity for mixed-precision neural network quantization. In: Proceedings of the 9th International Conference on Learning Representations, 2021"},{"key":"40669_CR9","doi-asserted-by":"publisher","first-page":"192","DOI":"10.1145\/3503221.3508417","volume-title":"Proceedings of the 27th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming","author":"Z Ma","year":"2022","unstructured":"Ma Z, He J, Qiu J, Cao H, Wang Y, Sun Z, Zheng L, Wang H, Tang S, Zheng T, Lin J, Feng G, Huang Z, Gao J, Zeng A, Zhang J, Zhong R, Shi T, Liu S, Zheng W, Tang J, Yang H, Liu X, Zhai J, Chen W. BaGuaLu: targeting brain scale pretrained models with over 37 million cores. In: Proceedings of the 27th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming. 2022, 192\u2013204"},{"key":"40669_CR10","volume-title":"Proceedings of the 10th International Conference on Learning Representations","author":"S Lee","year":"2021","unstructured":"Lee S, Park J, Jeon D. Toward efficient low-precision training: data format optimization and hysteresis quantization. In: Proceedings of the 10th International Conference on Learning Representations. 2021"},{"key":"40669_CR11","first-page":"1796","volume":"33","author":"X Sun","year":"2020","unstructured":"Sun X, Wang N, Chen C Y, Ni J M, Agrawal A, Cui X, Venkataramani S, El Maghraoui K, Srinivasan V V, Gopalakrishnan K. Ultra-low precision 4-bit training of deep neural networks. In: Proceedings of the 34th International Conference on Neural Information Processing Systems. 2020, 33: 1796\u20131807","journal-title":"Proceedings of the 34th International Conference on Neural Information Processing Systems"},{"key":"40669_CR12","first-page":"1017","volume-title":"Proceedings of the 34th International Conference on Neural Information Processing Systems","author":"Y Fu","year":"2020","unstructured":"Fu Y, You H, Zhao Y, Wang Y, Li C, Gopalakrishnan K, Wang Z, Lin Y. FracTrain: fractionally squeezing bit savings both temporally and spatially for efficient DNN training. In: Proceedings of the 34th International Conference on Neural Information Processing Systems. 2020, 1017"},{"key":"40669_CR13","first-page":"1740","volume-title":"Proceedings of the 31st International Conference on Neural Information Processing Systems","author":"U Koster","year":"2017","unstructured":"Koster U, Webb T J, Wang X, Nassar M, Bansal A K, Constable W H, Elibol O H, Gray S, Hall S, Hornof L, Khosrowshahi A, Kloss C, Pai R J, Rao N. Flexpoint: an adaptive numerical format for efficient training of deep neural networks. In: Proceedings of the 31st International Conference on Neural Information Processing Systems. 2017, 1740\u20131750"},{"key":"40669_CR14","volume-title":"Proceedings of the 6th International Conference on Learning Representations","author":"D Das","year":"2018","unstructured":"Das D, Mellempudi N, Mudigere D, Kalamkar D, Avancha S, Banerjee K, Sridharan S, Vaidyanathan K, Kaul B, Georganas E, Heinecke A, Dubey P, Corbal J, Shustrov N, Dubtsov R, Fomenko E, Pirogov V. Mixed precision training of convolutional neural networks using integer operations. In: Proceedings of the 6th International Conference on Learning Representations. 2018"},{"key":"40669_CR15","volume-title":"Proceedings of the 9th International Conference on Learning Representations","author":"S Fox","year":"2020","unstructured":"Fox S, Rasoulinezhad S, Faraone J, Leong P, Leong P. A block minifloat representation for training deep neural networks. In: Proceedings of the 9th International Conference on Learning Representations. 2020"},{"key":"40669_CR16","first-page":"441","volume-title":"Proceedings of the 33rd International Conference on Neural Information Processing Systems","author":"X Sun","year":"2019","unstructured":"Sun X, Choi J, Chen C Y, Wang N, Venkataramani S, Srinivasan V V, Cui X, Zhang W, Gopalakrishnan K. Hybrid 8-bit floating point (HFP8) training and inference for deep neural networks. In: Proceedings of the 33rd International Conference on Neural Information Processing Systems. 2019, 441"},{"key":"40669_CR17","doi-asserted-by":"publisher","first-page":"373","DOI":"10.1007\/978-3-030-01237-3_23","volume-title":"Proceedings of the 15th European Conference on Computer Vision-ECCV 2018","author":"D Zhang","year":"2018","unstructured":"Zhang D, Yang J, Ye D, Hua G. LQ-Nets: learned quantization for highly accurate and compact deep neural networks. In: Proceedings of the 15th European Conference on Computer Vision-ECCV 2018. 2018, 373\u2013390"},{"key":"40669_CR18","volume-title":"Proceedings of the 6th International Conference on Learning Representations","author":"J Choi","year":"2018","unstructured":"Choi J, Wang Z, Venkataramani S, Chuang P I J, Srinivasan V, Gopalakrishnan K. PACT: parameterized clipping activation for quantized neural networks. In: Proceedings of the 6th International Conference on Learning Representations, 2018"},{"key":"40669_CR19","first-page":"75","volume-title":"Proceedings of the 34th International Conference on Neural Information Processing Systems","author":"J Chen","year":"2020","unstructured":"Chen J, Gai Y, Yao Z, Mahoney M W, Gonzalez J E. A statistical framework for low-bitwidth training of deep neural networks. In: Proceedings of the 34th International Conference on Neural Information Processing Systems. 2020, 75"},{"key":"40669_CR20","first-page":"309","volume-title":"Proceedings of the 37th International Conference on Machine Learning","author":"F Fu","year":"2020","unstructured":"Fu F, Hu Y, He Y, Jiang J, Shao Y, Zhang C, Cui B. Don\u2019t waste your bits! Squeeze activations and gradients for deep neural networks via TINYSCRIPT. In: Proceedings of the 37th International Conference on Machine Learning. 2020, 309"},{"key":"40669_CR21","volume-title":"Proceedings of the 6th International Conference on Learning Representations","author":"S Wu","year":"2018","unstructured":"Wu S, Li G, Chen F, Shi L. Training and inference with integers in deep neural networks. In: Proceedings of the 6th International Conference on Learning Representations, 2018"},{"key":"40669_CR22","doi-asserted-by":"publisher","first-page":"70","DOI":"10.1016\/j.neunet.2019.12.027","volume":"125","author":"Y Yang","year":"2020","unstructured":"Yang Y, Deng L, Wu S, Yan T, Xie Y, Li G. Training highperformance and large-scale deep neural networks with full 8-bit integers. Neural Networks, 2020, 125: 70\u201382.","journal-title":"Neural Networks"},{"key":"40669_CR23","first-page":"7686","volume-title":"Proceedings of the 32nd International Conference on Neural Information Processing Systems","author":"N Wang","year":"2018","unstructured":"Wang N, Choi J, Brand D, Chen C Y, Gopalakrishnan K. Training deep neural networks with 8-bit floating point numbers. In: Proceedings of the 32nd International Conference on Neural Information Processing Systems. 2018, 7686\u20137695"},{"issue":"5","key":"40669_CR24","doi-asserted-by":"publisher","first-page":"1093","DOI":"10.1162\/neco.1997.9.5.1093","volume":"9","author":"Y Grandvalet","year":"1997","unstructured":"Grandvalet Y, Canu S, Boucheron S. Noise injection: theoretical prospects. Neural Computation, 1997, 9(5): 1093\u20131108.","journal-title":"Neural Computation"},{"key":"40669_CR25","volume-title":"Proceedings of the 5th International Conference on Learning Representations","author":"A Neelakantan","year":"2017","unstructured":"Neelakantan A, Vilnis L, Le Q V, Sutskever I, Kaiser L, Kurach K, Martens J. Adding gradient noise improves learning for very deep networks. In: Proceedings of the 5th International Conference on Learning Representations. 2017"},{"key":"40669_CR26","volume-title":"Proceedings of the 9th International Conference on Learning Representations","author":"Y Fu","year":"2021","unstructured":"Fu Y, Guo H, Li M, Yang X, Ding Y, Chandra V, Lin Y. CPT: efficient deep neural network training via cyclic precision. In: Proceedings of the 9th International Conference on Learning Representations. 2021"},{"issue":"1","key":"40669_CR27","doi-asserted-by":"publisher","first-page":"2383","DOI":"10.1038\/s41467-018-04316-3","volume":"9","author":"D C Mocanu","year":"2018","unstructured":"Mocanu D C, Mocanu E, Stone P, Nguyen P H, Gibescu M, Liotta A. Scalable training of artificial neural networks with adaptive sparse connectivity inspired by network science. Nature Communications, 2018, 9(1): 2383.","journal-title":"Nature Communications"},{"key":"40669_CR28","first-page":"2943","volume-title":"Proceedings of the 37th International Conference on Machine Learning","author":"U Evci","year":"2020","unstructured":"Evci U, Gale T, Menick J, Castro P S, Elsen E. Rigging the lottery: making all tickets winners. In: Proceedings of the 37th International Conference on Machine Learning. 2020, 2943\u20132952"},{"key":"40669_CR29","first-page":"21","volume-title":"Proceedings of the 10th International Conference on Learning Representations, ICLR 2022","author":"S Liu","year":"2022","unstructured":"Liu S, Chen T, Chen X, Shen L, Mocanu D C, Wang Z, Pechenizkiy M. The unreasonable effectiveness of random pruning: return of the most naive baseline for sparse training. In: Proceedings of the 10th International Conference on Learning Representations, ICLR 2022. 2022, 21"},{"key":"40669_CR30","volume-title":"Proceedings of the 5th International Conference on Learning Representations","author":"A Zhou","year":"2017","unstructured":"Zhou A, Yao A, Guo Y, Xu L, Chen Y. Incremental network quantization: towards lossless CNNs with low-precision weights. In: Proceedings of the 5th International Conference on Learning Representations, 2017"},{"key":"40669_CR31","volume-title":"Quantizing deep convolutional networks for efficient inference: a whitepaper","author":"R Krishnamoorthi","year":"2018","unstructured":"Krishnamoorthi R. Quantizing deep convolutional networks for efficient inference: a whitepaper. 2018, arXiv preprint arXiv: 1806.08342"},{"key":"40669_CR32","first-page":"714","volume-title":"Proceedings of the 33rd International Conference on Neural Information Processing Systems","author":"R Banner","year":"2019","unstructured":"Banner R, Nahshan Y, Soudry D. Post training 4-bit quantization of convolutional networks for rapid-deployment. In: Proceedings of the 33rd International Conference on Neural Information Processing Systems. 2019, 714"},{"key":"40669_CR33","first-page":"8604","volume-title":"Proceedings of 2019 IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"K Wang","year":"2019","unstructured":"Wang K, Liu Z, Lin Y, Lin J, Han S. HAQ: hardware-aware automated quantization with mixed precision. In: Proceedings of 2019 IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 2019, 8604\u20138612"},{"key":"40669_CR34","first-page":"3","volume-title":"Proceedings of the 50th Annual International Symposium on Computer Architecture","author":"C Guo","year":"2023","unstructured":"Guo C, Tang J, Hu W, Leng J, Zhang C, Yang F, Liu Y, Guo M, Zhu Y. OliVe: accelerating large language models via hardware-friendly outlier-victim pair quantization. In: Proceedings of the 50th Annual International Symposium on Computer Architecture. 2023, 3"},{"key":"40669_CR35","doi-asserted-by":"publisher","first-page":"1414","DOI":"10.1109\/MICRO56248.2022.00095","volume-title":"Proceedings of the 55th IEEE\/ACM International Symposium on Microarchitecture (MICRO)","author":"C Guo","year":"2022","unstructured":"Guo C, Zhang C, Leng J, Liu Z, Yang F, Liu Y, Guo M, Zhu Y. ANT: exploiting adaptive numerical data type for low-bit deep neural network quantization. In: Proceedings of the 55th IEEE\/ACM International Symposium on Microarchitecture (MICRO). 2022, 1414\u20131433"},{"issue":"4","key":"40669_CR36","doi-asserted-by":"publisher","first-page":"814","DOI":"10.1016\/j.ipm.2017.02.008","volume":"53","author":"A Onan","year":"2017","unstructured":"Onan A, Korukoglu S, Bulut H. A hybrid ensemble pruning approach based on consensus clustering and multi-objective evolutionary algorithm for sentiment classification. Information Processing & Management, 2017, 53(4): 814\u2013833.","journal-title":"Information Processing & Management"},{"key":"40669_CR37","volume-title":"Proceedings of the 7th International Conference on Learning Representations","author":"A Achille","year":"2018","unstructured":"Achille A, Rovere M, Soatto S. Critical learning periods in deep networks. In: Proceedings of the 7th International Conference on Learning Representations. 2018"},{"issue":"1","key":"40669_CR38","first-page":"1947","volume":"19","author":"A Achille","year":"2018","unstructured":"Achille A, Soatto S. Emergence of invariance and disentanglement in deep representations. Journal of Machine Learning Research, 2018, 19(1): 1947\u20131980.","journal-title":"Journal of Machine Learning Research"},{"key":"40669_CR39","first-page":"2847","volume-title":"Proceedings of the 34th International Conference on Machine Learning","author":"M Raghu","year":"2017","unstructured":"Raghu M, Poole B, Kleinberg J, Ganguli S, Sohl Dickstein J. On the expressive power of deep neural networks. In: Proceedings of the 34th International Conference on Machine Learning. 2017, 2847\u20132854"},{"key":"40669_CR40","volume-title":"Proceedings of the 8th International Conference on Learning Representations","author":"B Martinez","year":"2020","unstructured":"Martinez B, Yang J, Bulat A, Tzimiropoulos G. Training binary neural networks with real-to-binary convolutions. In: Proceedings of the 8th International Conference on Learning Representations. 2020"},{"key":"40669_CR41","first-page":"5151","volume-title":"Proceedings of the 32nd International Conference on Neural Information Processing systems","author":"R Banner","year":"2018","unstructured":"Banner R, Hubara I, Hoffer E, Soudry D. Scalable methods for 8-bit training of neural networks. In: Proceedings of the 32nd International Conference on Neural Information Processing systems. 2018, 5151\u20135159"},{"key":"40669_CR42","first-page":"430","volume-title":"Proceedings of the 16th European Conference on Computer Vision","author":"E Park","year":"2020","unstructured":"Park E, Yoo S. PROFIT: a novel training method for sub-4-bit MobileNet models. In: Proceedings of the 16th European Conference on Computer Vision. 2020, 430\u2013446"},{"key":"40669_CR43","first-page":"7015","volume-title":"Proceedings of the 36th International Conference on Machine Learning","author":"G Yang","year":"2019","unstructured":"Yang G, Zhang T, Kirichenko P, Bai J, Wilson A G, De Sa C. SWALP: stochastic weight averaging in low precision training. In: Proceedings of the 36th International Conference on Machine Learning. 2019, 7015\u20137024"},{"key":"40669_CR44","doi-asserted-by":"publisher","first-page":"480","DOI":"10.1145\/3437801.3441624","volume-title":"Proceedings of the 26th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming","author":"R Han","year":"2021","unstructured":"Han R, Si M, Demmel J, You Y. Dynamic scaling for low-precision learning. In: Proceedings of the 26th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming. 2021, 480\u2013482"},{"key":"40669_CR45","first-page":"1","volume-title":"Proceedings of the SC21: International Conference for High Performance Computing, Networking, Storage and Analysis","author":"B Feng","year":"2021","unstructured":"Feng B, Wang Y, Geng T, Li A, Ding Y. APNN-TC: accelerating arbitrary precision neural networks on ampere GPU tensor cores. In: Proceedings of the SC21: International Conference for High Performance Computing, Networking, Storage and Analysis. 2021, 1\u201314"},{"key":"40669_CR46","first-page":"1","volume-title":"SC21: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis","author":"F Knorr","year":"2021","unstructured":"Knorr F, Thoman P, Fahringer T. ndzip-gpu: efficient lossless compression of scientific floating-point data on GPUs. In: SC21: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis. 2021, 1\u201313"},{"key":"40669_CR47","first-page":"35769","volume-title":"Proceedings of the 36th International Conference on Neural Information Processing Systems","author":"P Savarese","year":"2022","unstructured":"Savarese P, Yuan X, Li Y, Maire M. Not all bits have equal value: heterogeneous precisions via trainable noise. In: Proceedings of the 36th International Conference on Neural Information Processing Systems. 2022, 35769\u201335782"},{"key":"40669_CR48","volume-title":"Proceedings of the 37th Conference on Neural Information Processing Systems (NeurIPS 2023)","author":"H Xi","year":"2023","unstructured":"Xi H, Li C, Chen J, Zhu J. Training transformers with 4-bit integers. In: Proceedings of the 37th Conference on Neural Information Processing Systems (NeurIPS 2023). 2023"},{"key":"40669_CR49","first-page":"9295","volume-title":"Proceedings of the 39th International Conference on Machine Learning","author":"X Huang","year":"2022","unstructured":"Huang X, Shen Z, Li S, Liu Z, Hu X, Wicaksana J, Xing E, Cheng K T. SDQ: stochastic differentiable quantization with mixed precision. In: Proceedings of the 39th International Conference on Machine Learning. 2022, 9295\u20139309"},{"key":"40669_CR50","first-page":"218","volume-title":"Proceedings of the 33rd International Conference on Neural Information Processing Systems","author":"A Chakrabarti","year":"2019","unstructured":"Chakrabarti A, Moseley B. Backprop with approximate activations for memory-efficient network training. In: Proceedings of the 33rd International Conference on Neural Information Processing Systems. 2019, 218"},{"key":"40669_CR51","doi-asserted-by":"publisher","first-page":"485","DOI":"10.1145\/3437801.3441597","volume-title":"Proceedings of the 26th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming","author":"S Jin","year":"2021","unstructured":"Jin S, Li G, Song S L, Tao D. A novel memory-efficient deep learning training framework via error-bounded lossy compression. In: Proceedings of the 26th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming. 2021, 485\u2013487"},{"key":"40669_CR52","first-page":"1803","volume-title":"Proceedings of the 38th International Conference on Machine Learning","author":"J Chen","year":"2021","unstructured":"Chen J, Zheng L, Yao Z, Wang D, Stoica I, Mahoney M, Gonzalez J. ActNN: reducing training memory footprint via 2-bit activation compressed training. In: Proceedings of the 38th International Conference on Machine Learning. 2021, 1803\u20131813"},{"key":"40669_CR53","doi-asserted-by":"publisher","first-page":"7920","DOI":"10.1109\/CVPR.2018.00826","volume-title":"Proceedings of 2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"B Zhuang","year":"2018","unstructured":"Zhuang B, Shen C, Tan M, Liu L, Reid I. Towards effective low-bitwidth convolutional neural networks. In: Proceedings of 2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 2018, 7920\u20137928"},{"key":"40669_CR54","first-page":"1225","volume-title":"Proceedings of the 33rd International Conference on Machine Learning","author":"M Hardt","year":"2016","unstructured":"Hardt M, Recht B, Singer Y. Train faster, generalize better: Stability of stochastic gradient descent. In: Proceedings of the 33rd International Conference on Machine Learning. 2016, 1225\u20131234"},{"issue":"3","key":"40669_CR55","doi-asserted-by":"publisher","first-page":"107","DOI":"10.1145\/3446776","volume":"64","author":"C Zhang","year":"2021","unstructured":"Zhang C, Bengio S, Hardt M, Recht B, Vinyals O. Understanding deep learning (still) requires rethinking generalization. Communications of the ACM, 2021, 64(3): 107\u2013115.","journal-title":"Communications of the ACM"},{"key":"40669_CR56","volume-title":"Towards understanding the role of over-parametrization in generalization of neural networks","author":"B Neyshabur","year":"2018","unstructured":"Neyshabur B, Li Z, Bhojanapalli S, LeCun Y, Srebro N. Towards understanding the role of over-parametrization in generalization of neural networks. 2018, arXiv preprint arXiv: 180512076"},{"key":"40669_CR57","first-page":"638","volume-title":"Readings in Computer Vision: Issues, Problem, Principles, and Paradigms","author":"T Poggio","year":"1987","unstructured":"Poggio T, Torre V, Koch C. Computational vision and regularization theory. In: Fischler M A, Firschein Q, eds. Readings in Computer Vision: Issues, Problem, Principles, and Paradigms. Los Altos: Morgan Kaufmann, 1987, 638\u2013643"},{"key":"40669_CR58","first-page":"370","volume-title":"Proceedings of the 40th International Conference on Machine Learning","author":"T Dettmers","year":"2023","unstructured":"Dettmers T, Zettlemoyer L. The case for 4-bit precision: k-bit inference scaling laws. In: Proceedings of the 40th International Conference on Machine Learning. 2023, 370"},{"key":"40669_CR59","first-page":"1712","volume":"33","author":"J Su","year":"2020","unstructured":"Su J, Chen Y, Cai T, Wu T, Gao R, Wang L, Lee J D. Sanity-checking pruning methods: random tickets can win the jackpot. In: Proceedings of the 34th International Conference on Neural Information Processing Systems. 2020, 33: 1712","journal-title":"Proceedings of the 34th International Conference on Neural Information Processing Systems"},{"key":"40669_CR60","volume-title":"Proceedings of the 7th International Conference on Learning Representations","author":"J Frankle","year":"2019","unstructured":"Frankle J, Carbin M. The lottery ticket hypothesis: finding sparse, trainable neural networks. In: Proceedings of the 7th International Conference on Learning Representations. 2019"},{"key":"40669_CR61","first-page":"4646","volume-title":"Proceedings of the 36th International Conference on Machine Learning","author":"H Mostafa","year":"2019","unstructured":"Mostafa H, Wang X. Parameter efficient training of deep convolutional neural networks by dynamic sparse reparameterization. In: Proceedings of the 36th International Conference on Machine Learning. 2019, 4646\u20134655"},{"key":"40669_CR62","first-page":"770","volume-title":"Proceedings of 2016 IEEE Conference on Computer Vision and Pattern Recognition","author":"K He","year":"2016","unstructured":"He K, Zhang X, Ren S, Sun J. Deep residual learning for image recognition. In: Proceedings of 2016 IEEE Conference on Computer Vision and Pattern Recognition. 2016, 770\u2013778"},{"key":"40669_CR63","doi-asserted-by":"publisher","first-page":"4510","DOI":"10.1109\/CVPR.2018.00474","volume-title":"Proceedings of 2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"M Sandler","year":"2018","unstructured":"Sandler M, Howard A, Zhu M, Zhmoginov A, Chen L C. MobileNetV2: inverted residuals and linear bottlenecks. In: Proceedings of 2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 2018, 4510\u20134520"},{"key":"40669_CR64","first-page":"462","volume-title":"Proceedings of the 33rd International Conference on Neural Information Processing Systems","author":"Y Wang","year":"2019","unstructured":"Wang Y, Jiang Z, Chen X, Xu P, Zhao Y, Lin Y, Wang Z. E2-train: training state-of-the-art CNNs with over 80% energy savings. In: Proceedings of the 33rd International Conference on Neural Information Processing Systems. 2019, 462"},{"key":"40669_CR65","volume-title":"Learning multiple layers of features from tiny images","author":"A Krizhevsky","year":"2012","unstructured":"Krizhevsky A. Learning multiple layers of features from tiny images. University of Toronto, 2012"},{"key":"40669_CR66","doi-asserted-by":"publisher","first-page":"248","DOI":"10.1109\/CVPR.2009.5206848","volume-title":"Proceedings of 2009 IEEE Conference on Computer Vision and Pattern Recognition","author":"J Deng","year":"2009","unstructured":"Deng J, Dong W, Socher R, Li L J, Li K, Fei-Fei L. ImageNet: a large-scale hierarchical image database. In: Proceedings of 2009 IEEE Conference on Computer Vision and Pattern Recognition. 2009, 248\u2013255"},{"key":"40669_CR67","volume-title":"Proceedings of the 7th International Conference on Learning Representations","author":"A Baevski","year":"2019","unstructured":"Baevski A, Auli M. Adaptive input representations for neural language modeling. In: Proceedings of the 7th International Conference on Learning Representations. 2019"},{"key":"40669_CR68","volume-title":"Proceedings of the 6th International Conference on Learning Representations","author":"S Merity","year":"2018","unstructured":"Merity S, Keskar N S, Socher R. Regularizing and optimizing LSTM language models. In: Proceedings of the 6th International Conference on Learning Representations. 2018"},{"key":"40669_CR69","doi-asserted-by":"publisher","first-page":"420","DOI":"10.1007\/978-3-030-01261-8_25","volume-title":"Proceedings of the 15th European Conference on Computer Vision-ECCV 2018","author":"X Wang","year":"2018","unstructured":"Wang X, Yu F, Dou Z Y, Darrell T, Gonzalez J E. SkipNet: learning dynamic routing in convolutional networks. In: Proceedings of the 15th European Conference on Computer Vision-ECCV 2018. 2018, 420\u2013436"},{"key":"40669_CR70","volume-title":"MobileNets: efficient convolutional neural networks for mobile vision applications","author":"A G Howard","year":"2017","unstructured":"Howard A G, Zhu M, Chen B, Kalenichenko D, Wang W, Weyand T, Andreetto M, Adam H. MobileNets: efficient convolutional neural networks for mobile vision applications. 2017, arXiv preprint arXiv: 1704.04861"},{"key":"40669_CR71","first-page":"660","volume-title":"Proceedings of the 33rd Neural Information Processing Systems","author":"L Hou","year":"2019","unstructured":"Hou L, Zhu J, Kwok J, Gao F, Qin T, Liu T Y. Normalization helps training of quantized LSTM. In: Proceedings of the 33rd Neural Information Processing Systems. 2019, 660"},{"issue":"6","key":"40669_CR72","doi-asserted-by":"publisher","first-page":"84","DOI":"10.1145\/3065386","volume":"60","author":"A Krizhevsky","year":"2017","unstructured":"Krizhevsky A, Sutskever I, Hinton G E. ImageNet classification with deep convolutional neural networks. Communications of the ACM, 2017, 60(6): 84\u201390.","journal-title":"Communications of the ACM"},{"key":"40669_CR73","volume-title":"Very deep convolutional networks for large-scale image recognition","author":"K Simonyan","year":"2014","unstructured":"Simonyan K, Zisserman A. Very deep convolutional networks for large-scale image recognition. 2014, arXiv preprint arXiv: 1409.1556"},{"key":"40669_CR74","volume-title":"PipeTransformer: automated elastic pipelining for distributed training of transformers","author":"C He","year":"2021","unstructured":"He C, Li S, Soltanolkotabi M, Avestimehr S. PipeTransformer: automated elastic pipelining for distributed training of transformers. 2021, arXiv preprint arXiv: 2102.03161"},{"key":"40669_CR75","first-page":"6078","volume-title":"Proceedings of the 31st International Conference on Neural Information Processing Systems","author":"M Raghu","year":"2017","unstructured":"Raghu M, Gilmer J, Yosinski J, Sohl-Dickstein J. SVCCA: singular vector canonical correlation analysis for deep learning dynamics and interpretability. In: Proceedings of the 31st International Conference on Neural Information Processing Systems. 2017, 6078\u20136087"},{"key":"40669_CR76","volume-title":"Neural Networks and Deep Learning","author":"M A Nielsen","year":"2015","unstructured":"Nielsen M A. Neural Networks and Deep Learning. San Francisco, CA, USA: Determination Press, 2015"},{"key":"40669_CR77","volume-title":"Proceedings of the 11th International Conference on Learning Representations","author":"E Frantar","year":"2023","unstructured":"Frantar E, Ashkboos S, Hoefler T, Alistarh D. GPTQ: accurate posttraining quantization for generative pre-trained transformers. In: Proceedings of the 11th International Conference on Learning Representations. 2023"},{"key":"40669_CR78","first-page":"441","volume-title":"Proceedings of the 37th International Conference on Neural Information Processing Systems","author":"T Dettmers","year":"2024","unstructured":"Dettmers T, Pagnoni A, Holtzman A, Zettlemoyer L. QLORA: efficient finetuning of quantized LLMs. In: Proceedings of the 37th International Conference on Neural Information Processing Systems. 2024, 441"},{"key":"40669_CR79","first-page":"2261","volume-title":"Proceedings of 2017 IEEE Conference on Computer Vision and Pattern Recognition","author":"G Huang","year":"2017","unstructured":"Huang G, Liu Z, Van Der Maaten L, Weinberger K Q. Densely connected convolutional networks. In: Proceedings of 2017 IEEE Conference on Computer Vision and Pattern Recognition. 2017, 2261\u20132269"}],"container-title":["Frontiers of Computer Science"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11704-024-40669-3.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11704-024-40669-3\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11704-024-40669-3.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,1,28]],"date-time":"2025-01-28T13:49:27Z","timestamp":1738072167000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11704-024-40669-3"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,1,28]]},"references-count":79,"journal-issue":{"issue":"10","published-print":{"date-parts":[[2025,10]]}},"alternative-id":["40669"],"URL":"https:\/\/doi.org\/10.1007\/s11704-024-40669-3","relation":{},"ISSN":["2095-2228","2095-2236"],"issn-type":[{"value":"2095-2228","type":"print"},{"value":"2095-2236","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,1,28]]},"assertion":[{"value":"3 July 2024","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"30 October 2024","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"28 January 2025","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"Competing interests The authors declare that they have no competing interests or financial conflicts to disclose.","order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethics"}}],"article-number":"1910355"}}