{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,15]],"date-time":"2026-05-15T08:12:59Z","timestamp":1778832779191,"version":"3.51.4"},"reference-count":52,"publisher":"Association for Computing Machinery (ACM)","issue":"5","license":[{"start":{"date-parts":[[2026,5,15]],"date-time":"2026-05-15T00:00:00Z","timestamp":1778803200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"funder":[{"name":"special Deep Space Exploration Cultivation Fund of USTC","award":["YD2390000601"],"award-info":[{"award-number":["YD2390000601"]}]},{"name":"Fund of Robot Technology","award":["25kftk03"],"award-info":[{"award-number":["25kftk03"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2026,5,31]]},"abstract":"<jats:p>Recently, learned image compression models have achieved better compression performance than traditional non-learning image compression standards. Those learned models usually utilize spatial self-attention and Convolutional Neural Network (CNN) to extract non-local and local features and generate the latent representation. However, previous methods adopt a linear layer to fuse non-local and local features and lack the flexibility to adaptively adjust feature weights and capture complex non-linear interactions between distinct feature representations. Additionally, how to more effectively compress the latent representation based on its channel similarity characteristics remains unexplored. To solve the above issues, we propose a novel image compression method with frequency feature interaction and non-local cross-similarity prior. More specifically, we extend the previous spatial self-attention module and alternately use spatial and channel self-attention modules to extract non-local spatial and channel features, respectively, and depth-wise convolution is utilized to extract local features. As local features focus on high-frequency detail information and non-local features concentrate on low-frequency structural information, we propose a frequency interaction module (FIM) that generates two weight maps to dynamically fuse non-local and local features. Moreover, we observe the non-local cross-similarity in different channels of the latent representation, which indicates that different channels share similar non-local semantic and structural information, but have distinct local detail information. So, we design a dual transformer entropy model to emphasize non-local features and remove local features. Experiment results validate our method achieves promising compression performance on the Kodak, CLIC, and Tecnick datasets.<\/jats:p>","DOI":"10.1145\/3803542","type":"journal-article","created":{"date-parts":[[2026,3,25]],"date-time":"2026-03-25T14:25:31Z","timestamp":1774448731000},"page":"1-22","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Learned Image Compression with Frequency Feature Interaction and Non-local Cross-similarity Prior"],"prefix":"10.1145","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-9741-2252","authenticated-orcid":false,"given":"Jian","family":"Wang","sequence":"first","affiliation":[{"name":"Department of Automation, University of Science and Technology of China, Hefei, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5688-4130","authenticated-orcid":false,"given":"Qiang","family":"Ling","sequence":"additional","affiliation":[{"name":"Department of Automation, University of Science and Technology of China, Hefei, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2026,5,15]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"crossref","first-page":"30147","DOI":"10.1109\/CVPR52734.2025.02806","volume-title":"2025 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Ahmed Sabbir","year":"2025","unstructured":"Sabbir Ahmed, Abdullah Al Arafat, Deniz Najafi, Akhlak Mahmood, Mamshad Nayeem Rizve, Mohaiminul Al Nahian, Ranyang Zhou, Shaahin Angizi, and Adnan Siraj Rakin. 2025. DeepCompress-ViT: Rethinking model compression to enhance efficiency of vision transformers at the edge. In 2025 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 30147\u201330156."},{"key":"e_1_3_1_3_2","first-page":"2691","volume-title":"International Conference on Learning Representations (ICLR)","author":"Ball\u00e9 Johannes","year":"2017","unstructured":"Johannes Ball\u00e9, Valero Laparra, and Eero P. Simoncelli. 2017. End-to-end optimized image compression. In International Conference on Learning Representations (ICLR), 2691\u20132717."},{"key":"e_1_3_1_4_2","first-page":"4961","volume-title":"International Conference on Learning Representations (ICLR)","author":"Ball\u00e9 Johannes","year":"2018","unstructured":"Johannes Ball\u00e9, David Minnen, Saurabh Singh, Sung Jin Hwang, and Nick Johnston. 2018. Variational image compression with a scale hyperprior. In International Conference on Learning Representations (ICLR), 4961\u20134983."},{"key":"e_1_3_1_5_2","author":"Bellard Fabrice","year":"2015","unstructured":"Fabrice Bellard. 2015. BPG Image Format. Retrieved from https:\/\/bellard.org\/bpg","journal-title":"BPG Image Format"},{"key":"e_1_3_1_6_2","unstructured":"Gisle Bjontegaard. 2001. Calculation of Average PSNR Differences between RD-Curves. ITU SG16 Doc. VCEG-M33."},{"issue":"10","key":"e_1_3_1_7_2","doi-asserted-by":"crossref","first-page":"3736","DOI":"10.1109\/TCSVT.2021.3101953","article-title":"Overview of the versatile video coding (VVC) standard and its applications","volume":"31","author":"Bross Benjamin","year":"2021","unstructured":"Benjamin Bross, Ye-Kui Wang, Yan Ye, Shan Liu, Jianle Chen, Gary J. Sullivan, and Jens-Rainer Ohm. 2021. Overview of the versatile video coding (VVC) standard and its applications. IEEE Transactions on Circuits and Systems for Video Technology 31, 10 (2021), 3736\u20133764.","journal-title":"IEEE Transactions on Circuits and Systems for Video Technology"},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2021.3058615"},{"key":"e_1_3_1_9_2","first-page":"22367","volume-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Chen Xiangyu","year":"2023","unstructured":"Xiangyu Chen, Xintao Wang, Jiantao Zhou, Yu Qiao, and Chao Dong. 2023. Activating more pixels in image super-resolution transformer. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 22367\u201322377."},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.01131"},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","DOI":"10.1109\/PCS.2018.8456308"},{"key":"e_1_3_1_12_2","first-page":"7939","volume-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Cheng Zhengxue","year":"2020","unstructured":"Zhengxue Cheng, Heming Sun, Masaru Takeuchi, and Jiro Katto. 2020. Learned image compression with discretized Gaussian mixture likelihoods and attention modules. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 7939\u20137948."},{"key":"e_1_3_1_13_2","first-page":"53","volume-title":"European Conference on Computer Vision","author":"Chu Xiaojie","year":"2022","unstructured":"Xiaojie Chu, Liangyu Chen, Chengpeng Chen, and Xin Lu. 2022. Improving image restoration by revisiting global information aggregation. In European Conference on Computer Vision, 53\u201371."},{"key":"e_1_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2023.3237274"},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2023.3263099"},{"key":"e_1_3_1_16_2","first-page":"14677","volume-title":"IEEE\/CVF International Conference on Computer Vision (CVPR)","author":"Gao Ge","year":"2021","unstructured":"Ge Gao, Pei You, Rong Pan, Shunyuan Han, Yuanyuan Zhang, Yuchao Dai, and Hojae Lee. 2021. Neural image compression via attentional multi-scale back projection and frequency decomposition. In IEEE\/CVF International Conference on Computer Vision (CVPR), 14677\u201314686."},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2021.3089491"},{"key":"e_1_3_1_18_2","first-page":"5718","volume-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"He Dailan","year":"2022","unstructured":"Dailan He, Ziming Yang, Weikun Peng, Rui Ma, Hongwei Qin, and Yan Wang. 2022. ELIC: Efficient learned image compression with unevenly grouped space-channel contextual adaptive coding. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 5718\u20135727."},{"key":"e_1_3_1_19_2","first-page":"14771","volume-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"He Dailan","year":"2021","unstructured":"Dailan He, Yaoyan Zheng, Baocheng Sun, Yan Wang, and Hongwei Qin. 2021. Checkerboard context model for efficient learned image compression. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 14771\u201314780."},{"key":"e_1_3_1_20_2","first-page":"1","volume-title":"International Conference on Learning Representations (ICLR)","author":"Johannes Ball\u00e9","year":"2016","unstructured":"Ball\u00e9 Johannes, Valero Laparra, and Eero P. Simoncelli. 2016. Density modeling of images using a generalized normalization transformation. In International Conference on Learning Representations (ICLR), 1\u201314."},{"key":"e_1_3_1_21_2","first-page":"5992","volume-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Kim Jun-Hyuk","year":"2022","unstructured":"Jun-Hyuk Kim, Byeongho Heo, and Jong-Seok Lee. 2022. Joint global and local hierarchical priors for learned image compression. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 5992\u20136001."},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2024.3371686"},{"key":"e_1_3_1_23_2","doi-asserted-by":"crossref","first-page":"276","DOI":"10.1016\/j.neunet.2023.12.018","article-title":"Large-scale cross-modal hashing with unified learning and multi-object regional correlation reasoning","volume":"171","author":"Li Bo","year":"2024","unstructured":"Bo Li and Zhixin Li. 2024. Large-scale cross-modal hashing with unified learning and multi-object regional correlation reasoning. Neural Networks: The Official Journal of the International Neural Network Society 171 (2024), 276\u2013292.","journal-title":"Neural Networks: The Official Journal of the International Neural Network Society"},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.1145\/3652583.3658001"},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.knosys.2020.106719"},{"key":"e_1_3_1_26_2","first-page":"3214","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Li Mu","year":"2018","unstructured":"Mu Li, Wangmeng Zuo, Shuhang Gu, Debin Zhao, and David Zhang. 2018. Learning convolutional networks for content-weighted image compression. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 3214\u20133223."},{"key":"e_1_3_1_27_2","first-page":"1833","volume-title":"IEEE\/CVF International Conference on Computer Vision (CVPR)","author":"Liang Jingyun","year":"2021","unstructured":"Jingyun Liang, Jiezhang Cao, Guolei Sun, Kai Zhang, Luc Van Gool, and Radu Timofte. 2021. SWINIR: Image restoration using Swin Transformer. In IEEE\/CVF International Conference on Computer Vision (CVPR), 1833\u20131844."},{"key":"e_1_3_1_28_2","first-page":"14388","volume-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Liu Jinming","year":"2023","unstructured":"Jinming Liu, Heming Sun, and Jiro Katto. 2023. Learned image compression with mixed transformer-CNN architectures. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 14388\u201314397."},{"key":"e_1_3_1_29_2","first-page":"10012","volume-title":"IEEE\/CVF International Conference on Computer Vision (CVPR)","author":"Liu Ze","year":"2021","unstructured":"Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. 2021. Swin Transformer: Hierarchical vision transformer using shifted windows. In IEEE\/CVF International Conference on Computer Vision (CVPR), 10012\u201310022."},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.knosys.2021.107254"},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2020.3026003"},{"key":"e_1_3_1_32_2","first-page":"1","volume-title":"Advances in Neural Information Processing Systems (NIPS)","volume":"31","author":"Minnen David","year":"2018","unstructured":"David Minnen, Johannes Ball\u00e9, and George D. Toderici. 2018. Joint autoregressive and hierarchical priors for learned image compression. In Advances in Neural Information Processing Systems (NIPS), Vol. 31, 1\u201310."},{"key":"e_1_3_1_33_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICIP40778.2020.9190935"},{"key":"e_1_3_1_34_2","doi-asserted-by":"publisher","DOI":"10.1145\/3551389"},{"key":"e_1_3_1_35_2","first-page":"4287","volume-title":"International Conference on Learning Representations (ICLR)","author":"Park Namuk","year":"2022","unstructured":"Namuk Park and Songkuk Kim. 2022. How do vision transformers work? In International Conference on Learning Representations (ICLR), 4287\u20134312."},{"key":"e_1_3_1_36_2","first-page":"15787","volume-title":"International Conference on Learning Representations (ICLR)","author":"Qian Yichen","year":"2021","unstructured":"Yichen Qian, Zhiyu Tan, Xiuyu Sun, Ming Lin, Dongyang Li, Zhenhong Sun, Hao Li, and Rong Jin. 2021. Learning accurate entropy model with global reference for image compression. In International Conference on Learning Representations (ICLR), 15787\u201315803."},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neunet.2023.03.037"},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2012.2221191"},{"key":"e_1_3_1_39_2","unstructured":"Lucas Theis Wenzhe Shi Andrew Cunningham and Ferenc Husz\u00e1r. 2017. Lossy image compression with compressive autoencoders. arXiv:1703.00395. Retrieved from https:\/\/arxiv.org\/abs\/1703.00395"},{"key":"e_1_3_1_40_2","doi-asserted-by":"crossref","unstructured":"Dimitrios Tsolakis George E. Tsekouras Antonios D. Niros and Anastasios Rigos. 2012. On the systematic development of fast fuzzy vector quantization for grayscale image compression. Neural Networks: The Official Journal of the International Neural Network Society 36 (2012) 83\u201396.","DOI":"10.1016\/j.neunet.2012.09.009"},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","DOI":"10.1109\/30.125072"},{"key":"e_1_3_1_42_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACSSC.2003.1292216"},{"key":"e_1_3_1_43_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2003.815165"},{"key":"e_1_3_1_44_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01234-2_1"},{"key":"e_1_3_1_45_2","first-page":"1589","article-title":"DTSNet: Dynamic transformer slimming for efficient vision recognition","author":"Xiao Wenjing","year":"2025","unstructured":"Wenjing Xiao, Xianzhi Li, Long Hu, Yixue Hao, and Min Chen. 2025. DTSNet: Dynamic transformer slimming for efficient vision recognition. IEEE Transactions on Multimedia 28 (2025), 1589\u20131600.","journal-title":"IEEE Transactions on Multimedia"},{"key":"e_1_3_1_46_2","doi-asserted-by":"publisher","DOI":"10.1145\/3474085.3475213"},{"key":"e_1_3_1_47_2","doi-asserted-by":"publisher","DOI":"10.1145\/3650034"},{"key":"e_1_3_1_48_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.00564"},{"issue":"1","key":"e_1_3_1_49_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3580499","article-title":"A universal optimization framework for learning-based image codec","volume":"20","author":"Zhao Jing","year":"2023","unstructured":"Jing Zhao, Bin Li, Jiahao Li, Ruiqin Xiong, and Yan Lu. 2023. A universal optimization framework for learning-based image codec. ACM Transactions on Multimedia Computing, Communications and Applications 20, 1 (2023), 1\u201319.","journal-title":"ACM Transactions on Multimedia Computing, Communications and Applications"},{"key":"e_1_3_1_50_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.knosys.2021.107284"},{"key":"e_1_3_1_51_2","first-page":"17612","volume-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Zhu Xiaosu","year":"2022","unstructured":"Xiaosu Zhu, Jingkuan Song, Lianli Gao, Feng Zheng, and Heng Tao Shen. 2022. Unified multivariate gaussian mixture for efficient neural image compression. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 17612\u201317621."},{"key":"e_1_3_1_52_2","first-page":"12188","volume-title":"International Conference on Learning Representations (ICLR)","author":"Zhu Yinhao","year":"2022","unstructured":"Yinhao Zhu, Yang Yang, and Taco Cohen. 2022. Transformer-based transform coding. In International Conference on Learning Representations (ICLR), 12188\u201312222."},{"key":"e_1_3_1_53_2","first-page":"17492","volume-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Zou Renjie","year":"2022","unstructured":"Renjie Zou, Chunfeng Song, and Zhaoxiang Zhang. 2022. The devil is in the details: Window-based attention for image compression. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 17492\u201317501."}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3803542","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,5,15]],"date-time":"2026-05-15T07:51:58Z","timestamp":1778831518000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3803542"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,5,15]]},"references-count":52,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2026,5,31]]}},"alternative-id":["10.1145\/3803542"],"URL":"https:\/\/doi.org\/10.1145\/3803542","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,5,15]]},"assertion":[{"value":"2024-06-23","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-03-14","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-05-15","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}