{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,25]],"date-time":"2026-06-25T16:46:27Z","timestamp":1782405987326,"version":"3.54.5"},"reference-count":51,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2026,6,25]],"date-time":"2026-06-25T00:00:00Z","timestamp":1782345600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"funder":[{"DOI":"10.13039\/501100012166","name":"National Key R&D Program of China","doi-asserted-by":"crossref","award":["2022YFB4501604"],"award-info":[{"award-number":["2022YFB4501604"]}],"id":[{"id":"10.13039\/501100012166","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Archit. Code Optim."],"published-print":{"date-parts":[[2026,6,30]]},"abstract":"<jats:p>As deep neural networks (DNNs) continue to grow in scale and complexity, GPU memory limitations have become a significant challenge for DNN model training, especially on resource-constrained commercial GPUs. While model quantization facilitates memory-efficient training, it often necessitates a tradeoff between quantization granularity and model accuracy. And quantization imposes additional computational overhead, which adversely affects the training throughput and apportions out the performance gains it brings. In this article, we propose FDSR, an adaptive tensor quantization method that leverages frequency domain division and similarity-based data reuse to break the memory bottleneck in visual model training. FDSR leverages the frequency-domain characteristics of tensors in terms of memory consumption and model accuracy, and proposes a fine-grained tensor quantization with different quantization bit-widths. It adaptively optimizes the quantization parameters according to model accuracy during training while employing sparsification according to data frequency-domain features, minimizing memory consumption and accuracy loss. To counteract the computational cost, FDSR incorporates a novel similarity-based reuse strategy that avoids redundant quantization\/dequantization computations, further enhanced by a tailored Locality-Sensitive Hashing (LSH) mechanism and optimized kernels. Experimental results demonstrate that FDSR achieves an average of 10.20\u00d7 activation memory compression with only 1.10% average accuracy loss across various models on the commercial GPU. Compared to the state-of-the-art quantization methods, FDSR improves memory optimization by up to 68.6% and increases throughput by up to 25.55%, with consistent performance improvements on different GPU architectures.<\/jats:p>","DOI":"10.1145\/3802593","type":"journal-article","created":{"date-parts":[[2026,4,13]],"date-time":"2026-04-13T11:06:09Z","timestamp":1776078369000},"page":"1-26","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["FDSR: Efficient Model Training via Adaptive Tensor Quantization Based on Frequency Domain Division and Similarity Data Reuse"],"prefix":"10.1145","volume":"23","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-7120-894X","authenticated-orcid":false,"given":"Song","family":"Liu","sequence":"first","affiliation":[{"name":"School of Computer Science and Technology, Xi'an Jiaotong University","place":["Xi'an, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0005-5490-3998","authenticated-orcid":false,"given":"Fei","family":"Li","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Xi'an Jiaotong University","place":["Xi'an, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-5013-0972","authenticated-orcid":false,"given":"Chenyu","family":"Zhao","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Xi'an Jiaotong University","place":["Xi'an, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0004-6546-9421","authenticated-orcid":false,"given":"Qin","family":"Xia","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Xi'an Jiaotong University","place":["Xi'an, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2215-7159","authenticated-orcid":false,"given":"Shiqiang","family":"Nie","sequence":"additional","affiliation":[{"name":"Xi'an Jiaotong University School of Computer Science and Technology","place":["Xi'An, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9449-3453","authenticated-orcid":false,"given":"Jinyu","family":"Wang","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Xi'an Jiaotong University","place":["Xi'an, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1179-3435","authenticated-orcid":false,"given":"Weiguo","family":"Wu","sequence":"additional","affiliation":[{"name":"Electronics and Information Engineering, Xi'an Jiaotong University","place":["Xi'an, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,6,25]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.1145\/3437984.3458839"},{"key":"e_1_3_1_3_2","unstructured":"Harshavardhan Adepu Zhanpeng Zeng Li Zhang and Vikas Singh. 2024. Framequant: Flexible low-bit quantization for transformers. In Proceedings of the 41st International Conference on Machine Learning (Vienna Austria) (ICML\u201924). JMLR.org. Article 8 (2024) 25."},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.1167\/jov.21.10.14"},{"key":"e_1_3_1_5_2","unstructured":"Ayan Chakrabarti and Benjamin Moseley. 2019. Backprop with approximate activations for memory-efficient network training. In Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019 NeurIPS 2019 December 8-14 2019 Vancouver BC Canada Hanna M. Wallach Hugo Larochelle Alina Beygelzimer Florence d\u2019Alch\u00e9-Buc Emily B. Fox and Roman Garnett (Eds.). 2426\u20132435. https:\/\/proceedings.neurips.cc\/paper\/2019\/hash\/30c8e1ca872524fbf7ea5c519ca397ee-Abstract.html"},{"key":"e_1_3_1_6_2","first-page":"1803","volume-title":"Proceeding of International Conference on Machine Learning","author":"Chen Jianfei","year":"2021","unstructured":"Jianfei Chen, Lianmin Zheng, Zhewei Yao, Dequan Wang, Ion Stoica, Michael Mahoney, and Joseph Gonzalez. 2021. ActNN: Reducing training memory footprint via 2-bit activation compressed training. In Proceeding of International Conference on Machine Learning. PMLR, 1803\u20131813."},{"key":"e_1_3_1_7_2","unstructured":"Jungwook Choi Zhuo Wang Swagath Venkataramani Pierce I-Jen Chuang Vijayalakshmi Srinivasan and Kailash Gopalakrishnan. 2018. Pact: Parameterized clipping activation for quantized neural networks. CoRR abs\/1805.06085 (2018). arXiv:1805.06085. Retrieved from http:\/\/arxiv.org\/abs\/1805.06085"},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"e_1_3_1_9_2","unstructured":"Alexey Dosovitskiy Lucas Beyer Alexander Kolesnikov Dirk Weissenborn Xiaohua Zhai Thomas Unterthiner Mostafa Dehghani Matthias Minderer Georg Heigold Sylvain Gelly Jakob Uszkoreit and Neil Houlsby. 2021. An image is worth 16x16 words: Transformers for image recognition at scale. In 9th International Conference on Learning Representations (ICLR\u201921). Virtual Event Austria. (May 3\u20137 2021). OpenReview.net. https:\/\/openreview.net\/forum?id=YicbFdNTTy"},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","unstructured":"Dayou Du Gu Gong and Xiaowen Chu. 2024. Model quantization and hardware acceleration for vision transformers: A comprehensive survey. CoRR abs\/2405.00314 (2024). DOI:10.48550\/ARXIV.2405.00314","DOI":"10.48550\/ARXIV.2405.00314"},{"key":"e_1_3_1_11_2","unstructured":"Steven K. Esser Jeffrey L. McKinstry Deepika Bablani Rathinakumar Appuswamy and Dharmendra S. Modha. 2020. Learned step size quantization. In 8th International Conference on Learning Representations (ICLR\u201920). Addis Ababa Ethiopia. (April 26\u201330 2020). OpenReview.net. https:\/\/openreview.net\/forum?id=rkgO66VKDS"},{"key":"e_1_3_1_12_2","unstructured":"R. David Evans and Tor M. Aamodt. 2021. AC-GC: Lossy activation compression with guaranteed convergence. In Advances in Neural Information Processing Systems 34: Annual Conference on Neural Information Processing Systems 2021 NeurIPS 2021 December 6-14 2021 Virtual Marc\u2019Aurelio Ranzato Alina Beygelzimer Yann N. Dauphin Percy Liang and Jennifer Wortman Vaughan (Eds.). 27434\u201327448. Retrieved from https:\/\/proceedings.neurips.cc\/paper\/2021\/hash\/e655c7716a4b3ea67f48c6322fc42ed6-Abstract.html"},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA45697.2020.00075"},{"key":"e_1_3_1_14_2","unstructured":"Angela Fan Pierre Stock Benjamin Graham Edouard Grave R\u00e9mi Gribonval Herve Jegou and Armand Joulin. 2021. Training with quantization noise for extreme model compression. In 9th International Conference on Learning Representations (ICLR 2021). Virtual Event Austria. (May 3\u20137 2021). Retrieved from OpenReview.net. https:\/\/openreview.net\/forum?id=dV19Yyi1fS3"},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58536-5_5"},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.14778\/3489496.3489500"},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1117\/12.20700"},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1201\/9781003162810-13"},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00501"},{"key":"e_1_3_1_20_2","unstructured":"Song Han Huizi Mao and William J. Dally. 2016. Deep compression: Compressing deep neural network with pruning trained quantization and huffman coding. In 4th International Conference on Learning Representations (ICLR\u201916). Yoshua Bengio and Yann LeCun (Eds.). Conference Track Proceedings. San Juan Puerto Rico. (May 2\u20134 2016). http:\/\/arxiv.org\/abs\/1510.00149"},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA.2018.00062"},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICDSP.1997.628094"},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.243"},{"key":"e_1_3_1_25_2","unstructured":"Omid Jafari Preeti Maurya Parth Nagarkar Khandker Mushfiqul Islam and Chidambaram Crushev. 2021. A survey on locality sensitive hashing algorithms and their applications. CoRR abs\/2102.08942 (2021). arXiv:2102.08942. Retrieved from https:\/\/arxiv.org\/abs\/2102.08942"},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA56546.2023.10071051"},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO50266.2020.00065"},{"key":"e_1_3_1_28_2","volume-title":"Learning Multiple Layers of Features from Tiny Images","author":"Krizhevsky Alex","year":"2009","unstructured":"Alex Krizhevsky, Vinod Nair, and Geoffrey Hinton. 2009. Learning Multiple Layers of Features from Tiny Images. Technical Report. University of Toronto. Retrieved from https:\/\/www.cs.toronto.edu\/kriz\/learning-features-2009-TR.pdf"},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.1371\/journal.pcbi.1010932"},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.324"},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"e_1_3_1_32_2","unstructured":"Xiaofan Lin Cong Zhao and Wei Pan. 2017. Towards accurate binary convolutional neural network. In Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017 December 4-9 2017 Long Beach CA USA Isabelle Guyon Ulrike von Luxburg Samy Bengio Hanna M. Wallach Rob Fergus S. V. N. Vishwanathan and Roman Garnett (Eds.). 345\u2013353. Retrieved from https:\/\/proceedings.neurips.cc\/paper\/2017\/hash\/b1a59b315fc9a3002ce38bbe070ec3f5-Abstract.html"},{"key":"e_1_3_1_33_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v35i10.17054"},{"key":"e_1_3_1_34_2","first-page":"14139","volume-title":"Proceeding of International Conference on Machine Learning","author":"Liu Xiaoxuan","year":"2022","unstructured":"Xiaoxuan Liu, Lianmin Zheng, Dequan Wang, Yukuo Cen, Weize Chen, Xu Han, Jianfei Chen, Zhiyuan Liu, Jie Tang, Joey Gonzalez, et\u00a0al. 2022. GACT: Activation compressed training for generic network architectures. In Proceeding of International Conference on Machine Learning. PMLR, 14139\u201314152."},{"key":"e_1_3_1_35_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00986"},{"key":"e_1_3_1_36_2","unstructured":"Daisuke Miyashita Edward H. Lee and Boris Murmann. 2016. Convolutional neural networks using logarithmic data representation. CoRR abs\/1603.01025 (2016). arXiv:1603.01025. Retrieved from http:\/\/arxiv.org\/abs\/1603.01025"},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.01527"},{"key":"e_1_3_1_38_2","unstructured":"Markus Nagel Marios Fournarakis Rana Ali Amjad Yelysei Bondarenko Mart Van Baalen and Tijmen Blankevoort. 2021. A white paper on neural network quantization. CoRR abs\/2106.08295 (2021). arXiv:2106.08295. Retrieved from https:\/\/arxiv.org\/abs\/2106.08295"},{"key":"e_1_3_1_39_2","first-page":"205","volume-title":"Proceedings of the Chinese Conference on Pattern Recognition and Computer Vision (PRCV)","author":"Nan Gongrui","year":"2023","unstructured":"Gongrui Nan and Fei Chao. 2023. Frequency domain distillation for data-free quantization of vision transformer. In Proceedings of the Chinese Conference on Pattern Recognition and Computer Vision (PRCV). Springer, 205\u2013216."},{"key":"e_1_3_1_40_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE.2019.00138"},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","DOI":"10.1145\/3330345.3330384"},{"key":"e_1_3_1_42_2","volume-title":"Discrete Cosine Transform: Algorithms, Advantages, Applications","author":"Rao K. Ramamohan","year":"2014","unstructured":"K. Ramamohan Rao and Ping Yip. 2014. Discrete Cosine Transform: Algorithms, Advantages, Applications. Academic Press."},{"key":"e_1_3_1_43_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2016.2577031"},{"key":"e_1_3_1_44_2","doi-asserted-by":"publisher","DOI":"10.1145\/3623402"},{"key":"e_1_3_1_45_2","volume-title":"Deep Learning with PyTorch","author":"Stevens Eli","year":"2020","unstructured":"Eli Stevens, Luca Antiga, and Thomas Viehmann. 2020. Deep Learning with PyTorch. Manning Publications."},{"key":"e_1_3_1_46_2","first-page":"36036","volume-title":"Proceeding of International Conference on Machine Learning","author":"Wang Guanchu","year":"2023","unstructured":"Guanchu Wang, Zirui Liu, Zhimeng Jiang, Ninghao Liu, Na Zou, and Xia Hu. 2023. DIVISION: Memory efficient training via dual activation precision. In Proceeding of International Conference on Machine Learning. PMLR, 36036\u201336057."},{"key":"e_1_3_1_47_2","doi-asserted-by":"publisher","DOI":"10.1145\/3617688"},{"key":"e_1_3_1_48_2","doi-asserted-by":"publisher","DOI":"10.1145\/3485447.3511985"},{"key":"e_1_3_1_49_2","first-page":"38087","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Xiao Guangxuan","year":"2023","unstructured":"Guangxuan Xiao, Ji Lin, Mickael Seznec, Hao Wu, Julien Demouth, and Song Han. 2023. Smoothquant: Accurate and efficient post-training quantization for large language models. In Proceedings of the International Conference on Machine Learning. PMLR, 38087\u201338099."},{"key":"e_1_3_1_50_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58610-2_1"},{"key":"e_1_3_1_51_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00748"},{"key":"e_1_3_1_52_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v32i1.11623"}],"container-title":["ACM Transactions on Architecture and Code Optimization"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3802593","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,25]],"date-time":"2026-06-25T15:55:15Z","timestamp":1782402915000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3802593"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6,25]]},"references-count":51,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2026,6,30]]}},"alternative-id":["10.1145\/3802593"],"URL":"https:\/\/doi.org\/10.1145\/3802593","relation":{},"ISSN":["1544-3566","1544-3973"],"issn-type":[{"value":"1544-3566","type":"print"},{"value":"1544-3973","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,6,25]]},"assertion":[{"value":"2025-10-21","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-03-09","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-06-25","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}