{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T05:02:46Z","timestamp":1750309366582,"version":"3.41.0"},"reference-count":55,"publisher":"Association for Computing Machinery (ACM)","issue":"7","license":[{"start":{"date-parts":[[2024,5,15]],"date-time":"2024-05-15T00:00:00Z","timestamp":1715731200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["No. 62325109, No. U21B2013, and No. 61971277"],"award-info":[{"award-number":["No. 62325109, No. U21B2013, and No. 61971277"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Ministry of Higher Education (MOHE) Malaysia FRGS Scheme","award":["No. FRGS\/1\/2022\/ICT02\/HWUM\/02\/1"],"award-info":[{"award-number":["No. FRGS\/1\/2022\/ICT02\/HWUM\/02\/1"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2024,7,31]]},"abstract":"<jats:p>The rapid advancement of multimedia and imaging technologies has resulted in increasingly diverse visual and semantic data. A large range of applications such as remote-assisted driving requires the amalgamated storage and transmission of various visual and semantic data. However, existing works suffer from the limitation of insufficiently exploiting the redundancy between different types of data. In this article, we propose a unified framework to jointly compress a diverse spectrum of visual and semantic data, including images, point clouds, segmentation maps, object attributes, and relations. We develop a unifying process that embeds the representations of these data into a joint embedding graph according to their categories, which enables flexible handling of joint compression tasks for various visual and semantic data. To fully leverage the redundancy between different data types, we further introduce an embedding-based adaptive joint encoding process and a Semantic Adaptation Module to efficiently encode diverse data based on the learned embeddings in the joint embedding graph. Experiments on the Cityscapes, MSCOCO, and KITTI datasets demonstrate the superiority of our framework, highlighting promising steps toward scalable multimedia processing.<\/jats:p>","DOI":"10.1145\/3654800","type":"journal-article","created":{"date-parts":[[2024,3,28]],"date-time":"2024-03-28T13:13:12Z","timestamp":1711631592000},"page":"1-24","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["A Unified Framework for Jointly Compressing Visual and Semantic Data"],"prefix":"10.1145","volume":"20","author":[{"ORCID":"https:\/\/orcid.org\/0009-0000-8353-8727","authenticated-orcid":false,"given":"Shizhan","family":"Liu","sequence":"first","affiliation":[{"name":"Shanghai Jiao Tong University, Shanghai, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8307-7107","authenticated-orcid":false,"given":"Weiyao","family":"Lin","sequence":"additional","affiliation":[{"name":"Shanghai Jiao Tong University, Shanghai,  China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1127-1570","authenticated-orcid":false,"given":"Yihang","family":"Chen","sequence":"additional","affiliation":[{"name":"Shanghai Jiao Tong University, Shanghai,  China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9570-155X","authenticated-orcid":false,"given":"Yufeng","family":"Zhang","sequence":"additional","affiliation":[{"name":"Shanghai Jiao Tong University, Shanghai,  China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2522-5778","authenticated-orcid":false,"given":"Wenrui","family":"Dai","sequence":"additional","affiliation":[{"name":"Shanghai Jiao Tong University, Shanghai,  China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3005-4109","authenticated-orcid":false,"given":"John","family":"See","sequence":"additional","affiliation":[{"name":"Heriot-Watt University Malaysia, Putrajaya,  Malaysia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4552-0029","authenticated-orcid":false,"given":"Hong-Kai","family":"Xiong","sequence":"additional","affiliation":[{"name":"Shanghai Jiao Tong University, Shanghai,  China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2024,5,15]]},"reference":[{"key":"e_1_3_1_2_2","first-page":"2042","volume-title":"Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP\u201919)","author":"Akbari Mohammad","year":"2019","unstructured":"Mohammad Akbari, Jie Liang, and Jingning Han. 2019. DSSLIC: Deep semantic segmentation-based layered image compression. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP\u201919). IEEE, 2042\u20132046."},{"key":"e_1_3_1_3_2","first-page":"4342","volume-title":"Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP\u201920)","author":"Alvar Saeed Ranjbar","year":"2020","unstructured":"Saeed Ranjbar Alvar and Ivan V. Baji\u0107. 2020. Bit allocation for multi-task collaborative intelligence. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP\u201920). IEEE, 4342\u20134346."},{"key":"e_1_3_1_4_2","unstructured":"Johannes Ball\u00e9 Valero Laparra and Eero P. Simoncelli. 2016. End-to-end optimized image compression. Retrieved from https:\/\/arXiv:1611.01704"},{"key":"e_1_3_1_5_2","unstructured":"Johannes Ball\u00e9 David Minnen Saurabh Singh Sung Jin Hwang and Nick Johnston. 2018. Variational image compression with a scale hyperprior. Retrieved from https:\/\/arXiv:1802.01436"},{"key":"e_1_3_1_6_2","unstructured":"F. Bellard. 2017. BPG image format. Retrieved from http:\/\/bellard.org\/bpg\/"},{"key":"e_1_3_1_7_2","unstructured":"Gisle Bj\u00f8ntegaard. 2001. Calculation of average PSNR differences between RD-curves. Retrieved from https:\/\/www.itu.int\/wftp3\/av-arch\/video-site\/0104_Aus\/VCEG-M33.doc"},{"key":"e_1_3_1_8_2","first-page":"694","volume-title":"Proceedings of the IEEE International Conference on Image Processing (ICIP\u201919)","author":"Chang Jianhui","year":"2019","unstructured":"Jianhui Chang, Qi Mao, Zhenghui Zhao, Shanshe Wang, Shiqi Wang, Hong Zhu, and Siwei Ma. 2019. Layered conceptual image compression via deep semantic synthesis. In Proceedings of the IEEE International Conference on Image Processing (ICIP\u201919). IEEE, 694\u2013698."},{"key":"e_1_3_1_9_2","first-page":"1","article-title":"Semantic-aware visual decomposition for image coding","author":"Chang Jianhui","year":"2023","unstructured":"Jianhui Chang, Jian Zhang, Jiguo Li, Shiqi Wang, Qi Mao, Chuanmin Jia, Siwei Ma, and Wen Gao. 2023. Semantic-aware visual decomposition for image coding. Int. J. Comput. Vision 131, 9 (2023), 1\u201323.","journal-title":"Int. J. Comput. Vision"},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2022.3159477"},{"key":"e_1_3_1_11_2","first-page":"1","volume-title":"Proceedings of the IEEE International Conference on Multimedia and Expo (ICME\u201921)","author":"Chang Jianhui","year":"2021","unstructured":"Jianhui Chang, Zhenghui Zhao, Lingbo Yang, Chuanmin Jia, Jian Zhang, and Siwei Ma. 2021. Thousand to one: Semantic prior modeling for conceptual coding. In Proceedings of the IEEE International Conference on Multimedia and Expo (ICME\u201921). IEEE, 1\u20136."},{"key":"e_1_3_1_12_2","first-page":"3094","volume-title":"Proceedings of the IEEE International Conference on Image Processing (ICIP\u201920)","author":"Chen Zhuo","year":"2020","unstructured":"Zhuo Chen, Ling-Yu Duan, Shiqi Wang, Weisi Lin, and Alex C. Kot. 2020. Data representation in hybrid coding framework for feature maps compression. In Proceedings of the IEEE International Conference on Image Processing (ICIP\u201920). IEEE, 3094\u20133098."},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1145\/3343031.3350849"},{"key":"e_1_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00796"},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.350"},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.1145\/3419635.3419668"},{"key":"e_1_3_1_17_2","first-page":"625","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","volume":"36","author":"Fu Chunyang","year":"2022","unstructured":"Chunyang Fu, Ge Li, Rui Song, Wei Gao, and Shan Liu. 2022. Octattention: Octree-based large-scale contexts model for point cloud compression. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 36. 625\u2013633."},{"key":"e_1_3_1_18_2","unstructured":"Jean-loup Gailly and Mark Adler. 2004. Zlib compression library. Retrieved from http:\/\/www.dspace.cam.ac.uk\/handle\/1810\/3486"},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2021.3104305"},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2012.6248074"},{"key":"e_1_3_1_21_2","unstructured":"Google. 2017. Draco 3D graphics compression. Retrieved from https:\/\/github.com\/google\/draco"},{"key":"e_1_3_1_22_2","first-page":"1","volume-title":"Proceedings of the Picture Coding Symposium (PCS\u201919)","author":"Guarda Andr\u00e9 F. R.","year":"2019","unstructured":"Andr\u00e9 F. R. Guarda, Nuno M. M. Rodrigues, and Fernando Pereira. 2019. Point cloud coding: Adopting a deep learning-based approach. In Proceedings of the Picture Coding Symposium (PCS\u201919). IEEE, 1\u20135."},{"key":"e_1_3_1_23_2","first-page":"5718","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"He Dailan","year":"2022","unstructured":"Dailan He, Ziming Yang, Weikun Peng, Rui Ma, Hongwei Qin, and Yan Wang. 2022. Elic: Efficient learned image compression with unevenly grouped space-channel contextual adaptive coding. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 5718\u20135727."},{"key":"e_1_3_1_24_2","first-page":"160","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops","author":"Hoang Trinh Man","year":"2020","unstructured":"Trinh Man Hoang, Jinjia Zhou, and Yibo Fan. 2020. Image compression with encoder-decoder matched semantic segmentation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops. 160\u2013161."},{"key":"e_1_3_1_25_2","first-page":"475","volume-title":"Proceedings of the IEEE International Conference on Visual Communications and Image Processing (VCIP\u201920)","author":"Hu Yuzhang","year":"2020","unstructured":"Yuzhang Hu, Sifeng Xia, Wenhan Yang, and Jiaying Liu. 2020. Sensitivity-aware bit allocation for intermediate deep feature compression. In Proceedings of the IEEE International Conference on Visual Communications and Image Processing (VCIP\u201920). IEEE, 475\u2013478."},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2021.3120720"},{"key":"e_1_3_1_27_2","unstructured":"Min Lin Qiang Chen and Shuicheng Yan. 2013. Network in network. Retrieved from https:\/\/arXiv:1312.4400"},{"key":"e_1_3_1_28_2","first-page":"740","volume-title":"Proceedings of the 13th European Conference on Computer Vision (ECCV\u201914)","author":"Lin Tsung-Yi","year":"2014","unstructured":"Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll\u00e1r, and C. Lawrence Zitnick. 2014. Microsoft COCO: Common objects in context. In Proceedings of the 13th European Conference on Computer Vision (ECCV\u201914). Springer, 740\u2013755."},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.1109\/MMUL.2020.2990863"},{"key":"e_1_3_1_30_2","unstructured":"Haojie Liu Tong Chen Peiyao Guo Qiu Shen Xun Cao Yao Wang and Zhan Ma. 2019. Non-local attention optimized deep image compression. Retrieved from https:\/\/arXiv:1904.09757"},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-021-01491-7"},{"key":"e_1_3_1_32_2","first-page":"10629","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Mentzer Fabian","year":"2019","unstructured":"Fabian Mentzer, Eirikur Agustsson, Michael Tschannen, Radu Timofte, and Luc Van Gool. 2019. Practical full resolution learned lossless image compression. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 10629\u201310638."},{"key":"e_1_3_1_33_2","article-title":"Joint autoregressive and hierarchical priors for learned image compression","volume":"31","author":"Minnen David","year":"2018","unstructured":"David Minnen, Johannes Ball\u00e9, and George D. Toderici. 2018. Joint autoregressive and hierarchical priors for learned image compression. Adv. Neural Info. Process. Syst. 31 (2018).","journal-title":"Adv. Neural Info. Process. Syst."},{"key":"e_1_3_1_34_2","unstructured":"MPEGGroup. 2021. Mpeg g-pcc tmc13. Retrieved from https:\/\/github.com\/MPEGGroup\/mpeg-pcc-tmc13"},{"key":"e_1_3_1_35_2","volume-title":"Proceedings of the Picture Coding Symposium","volume":"2018","author":"Ohm Jens-Rainer","year":"2018","unstructured":"Jens-Rainer Ohm and Gary J. Sullivan. 2018. Versatile video coding\u2013towards the next generation of video compression. In Proceedings of the Picture Coding Symposium, Vol. 2018."},{"issue":"1","key":"e_1_3_1_36_2","article-title":"Lossless data compression algorithm\u2013a review","volume":"5","author":"Parekar P. M.","year":"2014","unstructured":"P. M. Parekar and S. S. Thakare. 2014. Lossless data compression algorithm\u2013a review. Int. J. Comput. Sci. Info. Technol. 5, 1 (2014).","journal-title":"Int. J. Comput. Sci. Info. Technol."},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.1145\/3551389"},{"key":"e_1_3_1_38_2","article-title":"JPEG-compatible joint image compression and encryption algorithm with file size preservation","author":"Peng Yuxiang","year":"2023","unstructured":"Yuxiang Peng, Chong Fu, Guixing Cao, Wei Song, Junxin Chen, and Chiu-Wing Sham. 2023. JPEG-compatible joint image compression and encryption algorithm with file size preservation. ACM Trans. Multimedia Comput. Commun. Appl. 20, 4 (2023).","journal-title":"ACM Trans. Multimedia Comput. Commun. Appl."},{"key":"e_1_3_1_39_2","first-page":"652","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Qi Charles R.","year":"2017","unstructured":"Charles R. Qi, Hao Su, Kaichun Mo, and Leonidas J. Guibas. 2017. Pointnet: Deep learning on point sets for 3D classification and segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 652\u2013660."},{"key":"e_1_3_1_40_2","first-page":"3349","volume-title":"Proceedings of the IEEE International Conference on Image Processing (ICIP\u201920)","author":"Singh Saurabh","year":"2020","unstructured":"Saurabh Singh, Sami Abu-El-Haija, Nick Johnston, Johannes Ball\u00e9, Abhinav Shrivastava, and George Toderici. 2020. End-to-end learning of compressible features. In Proceedings of the IEEE International Conference on Image Processing (ICIP\u201920). IEEE, 3349\u20133353."},{"key":"e_1_3_1_41_2","first-page":"66","volume-title":"Proceedings of the IEEE International Conference on Image Processing (ICIP\u201916)","author":"Sneyers Jon","year":"2016","unstructured":"Jon Sneyers and Pieter Wuille. 2016. FLIF: Free lossless image format based on MANIAC compression. In Proceedings of the IEEE International Conference on Image Processing (ICIP\u201916). IEEE, 66\u201370."},{"key":"e_1_3_1_42_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-020-01305-2"},{"key":"e_1_3_1_43_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2012.2221191"},{"key":"e_1_3_1_44_2","first-page":"3099","volume-title":"Proceedings of the IEEE International Conference on Image Processing (ICIP\u201920)","author":"Suzuki Satoshi","year":"2020","unstructured":"Satoshi Suzuki, Motohiro Takagi, Shoichiro Takeda, Ryuichi Tanida, and Hideaki Kimata. 2020. Deep feature compression with spatio-temporal arranging for collaborative intelligence. In Proceedings of the IEEE International Conference on Image Processing (ICIP\u201920). IEEE, 3099\u20133103."},{"key":"e_1_3_1_45_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00377"},{"key":"e_1_3_1_46_2","doi-asserted-by":"publisher","DOI":"10.1145\/103085.103089"},{"key":"e_1_3_1_47_2","unstructured":"Jianqiang Wang Hao Zhu Zhan Ma Tong Chen Haojie Liu and Qiu Shen. 2019. Learned point cloud geometry compression. Retrieved from https:\/\/arXiv:1909.12037"},{"key":"e_1_3_1_48_2","unstructured":"Yuxin Wu Alexander Kirillov Francisco Massa Wan-Yen Lo and Ross Girshick. 2019. Detectron2. Retrieved from https:\/\/github.com\/facebookresearch\/detectron2"},{"key":"e_1_3_1_49_2","volume-title":"Proceedings of the ACM Multimedia Conference","author":"Yang Shuyu","year":"2023","unstructured":"Shuyu Yang, Yinan Zhou, Yaxiong Wang, Yujiao Wu, Li Zhu, and Zhedong Zheng. 2023. Towards unified text-based person retrieval: A large-scale multi-attribute and language search benchmark. In Proceedings of the ACM Multimedia Conference."},{"key":"e_1_3_1_50_2","article-title":"Divide-and-conquer-based RDO-free CU partitioning for 8K video compression","author":"Yuan Hang","year":"2023","unstructured":"Hang Yuan, Wei Gao, Siwei Ma, and Yiqiang Yan. 2023. Divide-and-conquer-based RDO-free CU partitioning for 8K video compression. ACM Trans. Multimedia Comput. Commun. Appl. 20, 4 (2023).","journal-title":"ACM Trans. Multimedia Comput. Commun. Appl."},{"key":"e_1_3_1_51_2","doi-asserted-by":"publisher","DOI":"10.1109\/JIOT.2020.3039359"},{"key":"e_1_3_1_52_2","volume-title":"Proceedings of the IEEE International Conference on Computer Vision (ICCV\u201923)","author":"Zhang Lvmin","year":"2023","unstructured":"Lvmin Zhang, Anyi Rao, and Maneesh Agrawala. 2023. Adding conditional control to text-to-image diffusion models. In Proceedings of the IEEE International Conference on Computer Vision (ICCV\u201923)."},{"key":"e_1_3_1_53_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-021-01505-4"},{"issue":"12","key":"e_1_3_1_54_2","doi-asserted-by":"crossref","first-page":"11347","DOI":"10.1109\/JIOT.2020.3028766","article-title":"Toward automated vehicle teleoperation: Vision, opportunities, and challenges","volume":"7","author":"Zhang Tao","year":"2020","unstructured":"Tao Zhang. 2020. Toward automated vehicle teleoperation: Vision, opportunities, and challenges. IEEE Internet Things J. 7, 12 (2020), 11347\u201311354.","journal-title":"IEEE Internet Things J."},{"key":"e_1_3_1_55_2","unstructured":"Yufeng Zhang Weiyao Lin Wenrui Dai Huabin Liu and Hongkai Xiong. 2023. Scene graph lossless compression with adaptive prediction for objects and relations. Retrieved from https:\/\/arXiv:2304.13359"},{"key":"e_1_3_1_56_2","first-page":"1","volume-title":"Proceedings of the IEEE International Conference on Multimedia and Expo (ICME\u201921)","author":"Zhang Zhicong","year":"2021","unstructured":"Zhicong Zhang, Mengyang Wang, Mengyao Ma, Jiahui Li, and Xiaopeng Fan. 2021. MSFC: Deep feature compression in multi-task network. In Proceedings of the IEEE International Conference on Multimedia and Expo (ICME\u201921). IEEE, 1\u20136."}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3654800","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3654800","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T00:06:09Z","timestamp":1750291569000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3654800"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,5,15]]},"references-count":55,"journal-issue":{"issue":"7","published-print":{"date-parts":[[2024,7,31]]}},"alternative-id":["10.1145\/3654800"],"URL":"https:\/\/doi.org\/10.1145\/3654800","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"type":"print","value":"1551-6857"},{"type":"electronic","value":"1551-6865"}],"subject":[],"published":{"date-parts":[[2024,5,15]]},"assertion":[{"value":"2023-12-11","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-03-21","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-05-15","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}