{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,25]],"date-time":"2026-04-25T14:44:10Z","timestamp":1777128250114,"version":"3.51.4"},"reference-count":56,"publisher":"Association for Computing Machinery (ACM)","issue":"7","license":[{"start":{"date-parts":[[2024,3,27]],"date-time":"2024-03-27T00:00:00Z","timestamp":1711497600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2024,7,31]]},"abstract":"<jats:p>The scene graph is a novel data structure describing objects and their pairwise relationship within image scenes. As the size of scene graphs in vision and multimedia applications increases, the need for lossless storage and transmission of such data becomes more critical. However, the compression of scene graphs is less studied because of the complicated data structures involved and complex distributions. Existing solutions usually involve general-purpose compressors or graph structure compression methods, which are weak at reducing the redundancy in scene graph data. This article introduces a novel lossless compression framework with adaptive predictors for the joint compression of objects and relations in scene graph data. The proposed framework comprises a unified prior extractor and specialized element predictors to adapt to different data elements. Furthermore, to exploit the context information within and between graph elements, Graph Context Convolution is proposed to support different graph context modeling schemes for different graph elements. Finally, an overarching framework incorporates the learned distribution model to predict numerical data under complicated conditional constraints. Experiments conducted on labeled or generated scene graphs demonstrate the effectiveness of the proposed framework for scene graph lossless compression.<\/jats:p>","DOI":"10.1145\/3649503","type":"journal-article","created":{"date-parts":[[2024,2,26]],"date-time":"2024-02-26T12:35:33Z","timestamp":1708950933000},"page":"1-23","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":2,"title":["Scene Graph Lossless Compression with Adaptive Prediction for Objects and Relations"],"prefix":"10.1145","volume":"20","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-8307-7107","authenticated-orcid":false,"given":"Weiyao","family":"Lin","sequence":"first","affiliation":[{"name":"Department of Electronic Engineering, Shanghai Jiao Tong University, Shanghai, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9570-155X","authenticated-orcid":false,"given":"Yufeng","family":"Zhang","sequence":"additional","affiliation":[{"name":"Department of Electronic Engineering, Shanghai Jiao Tong University, Shanghai, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2522-5778","authenticated-orcid":false,"given":"Wenrui","family":"Dai","sequence":"additional","affiliation":[{"name":"Department of Computer Science and Engineering, Shanghai Jiao Tong University, Shanghai, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9174-1696","authenticated-orcid":false,"given":"Huabin","family":"Liu","sequence":"additional","affiliation":[{"name":"Department of Electronic Engineering, Shanghai Jiao Tong University, Shanghai, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3005-4109","authenticated-orcid":false,"given":"John","family":"See","sequence":"additional","affiliation":[{"name":"School of Mathematical and Computer Sciences, Heriot-Watt University Malaysia, Putrajaya, Malaysia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4552-0029","authenticated-orcid":false,"given":"Hongkai","family":"Xiong","sequence":"additional","affiliation":[{"name":"Department of Electronic Engineering, Shanghai Jiao Tong University, Shanghai, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2024,3,27]]},"reference":[{"key":"e_1_3_3_2_2","doi-asserted-by":"publisher","unstructured":"Jyrki Alakuijala Andrea Farruggia Paolo Ferragina Eugene Kliuchnikov Robert Obryk Zoltan Szabadka and Lode Vandevenne. 2018. Brotli: A general-purpose data compressor. ACM Trans. Inf. Syst. 37 1 Article 4 (January 2019) 30. 10.1145\/3231935","DOI":"10.1145\/3231935"},{"key":"e_1_3_3_3_2","first-page":"382","volume-title":"ECCV","author":"Anderson Peter","year":"2016","unstructured":"Peter Anderson, Basura Fernando, Mark Johnson, and Stephen Gould. 2016. SPICE: Semantic propositional image caption evaluation. In ECCV. 382\u2013398."},{"key":"e_1_3_3_4_2","volume-title":"ICLR.","author":"Ball\u00e9 Johannes","year":"2017","unstructured":"Johannes Ball\u00e9, Valero Laparra, and Eero P. Simoncelli. 2017. End-to-end optimized image compression. In ICLR."},{"key":"e_1_3_3_5_2","volume-title":"ICLR.","author":"Ball\u00e9 Johannes","year":"2018","unstructured":"Johannes Ball\u00e9, David Minnen, Saurabh Singh, Sung Jin Hwang, and Nick Johnston. 2018. Variational image compression with a scale hyperprior. In ICLR."},{"key":"e_1_3_3_6_2","article-title":"Survey and taxonomy of lossless graph compression and space-efficient graph representations","volume":"1806","author":"Besta Maciej","year":"2018","unstructured":"Maciej Besta and Torsten Hoefler. 2018. Survey and taxonomy of lossless graph compression and space-efficient graph representations. CoRR abs\/1806.01799 (2018).","journal-title":"CoRR"},{"key":"e_1_3_3_7_2","doi-asserted-by":"crossref","first-page":"595","DOI":"10.1145\/988672.988752","volume-title":"WWW","author":"Boldi Paolo","year":"2004","unstructured":"Paolo Boldi and Sebastiano Vigna. 2004. The webgraph framework I: compression techniques. In WWW. 595\u2013602."},{"key":"e_1_3_3_8_2","first-page":"18","volume-title":"SPIRE","author":"Brisaboa Nieves R.","year":"2009","unstructured":"Nieves R. Brisaboa, Susana Ladra, and Gonzalo Navarro. 2009. k2-trees for compact web graph representation. In SPIRE. 18\u201330."},{"key":"e_1_3_3_9_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2021.3137605"},{"key":"e_1_3_3_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2021.3058615"},{"key":"e_1_3_3_11_2","doi-asserted-by":"publisher","DOI":"10.1145\/3499027"},{"key":"e_1_3_3_12_2","first-page":"7936","volume-title":"CVPR","author":"Cheng Zhengxue","year":"2020","unstructured":"Zhengxue Cheng, Heming Sun, Masaru Takeuchi, and Jiro Katto. 2020. Learned image compression with discretized Gaussian mixture likelihoods and attention modules. In CVPR. 7936\u20137945."},{"key":"e_1_3_3_13_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCOM.1984.1096090"},{"key":"e_1_3_3_14_2","article-title":"Zstandard\u2014Real-time compression algorithm","author":"Collet Yann","year":"2021","unstructured":"Yann Collet. 2021. Zstandard\u2014Real-time compression algorithm. Retrieved from https:\/\/facebook.github.io\/zstd\/","journal-title":"R"},{"key":"e_1_3_3_15_2","article-title":"Asymmetric numeral systems as close to capacity low state entropy coders","volume":"1311","author":"Duda Jarek","year":"2013","unstructured":"Jarek Duda. 2013. Asymmetric numeral systems as close to capacity low state entropy coders. CoRR abs\/1311.2540 (2013).","journal-title":"CoRR"},{"key":"e_1_3_3_16_2","unstructured":"Haisheng Fu Feng Liang Jianping Lin Bing Li Mohammad Akbari Jie Liang Guohe Zhang Dong Liu Chengjie Tu and Jingning Han. 2021. Learned Image compression with discretized gaussian-laplacian-logistic mixture model and concatenated residual modules. CoRR abs\/2107.06463 (2021)."},{"key":"e_1_3_3_17_2","first-page":"457","volume-title":"EMNLP","author":"Fukui Akira","year":"2016","unstructured":"Akira Fukui, Dong Huk Park, Daylen Yang, Anna Rohrbach, Trevor Darrell, and Marcus Rohrbach. 2016. Multimodal compact bilinear pooling for visual question answering and visual grounding. In EMNLP. 457\u2013468."},{"key":"e_1_3_3_18_2","first-page":"372","volume-title":"DCC","author":"Goyal Mohit","year":"2020","unstructured":"Mohit Goyal, Kedar Tatwawadi, Shubham Chandak, and Idoia Ochoa. 2020. DZip: Improved general-purpose lossless compression based on novel neural network modeling. In DCC. 372."},{"key":"e_1_3_3_19_2","article-title":"zlib Home Site","author":"Roelofs Jean-Loup Gailly Greg","year":"2021","unstructured":"Jean-Loup Gailly Greg Roelofs and Mark Adler. 2021. zlib Home Site. Retrieved from https:\/\/www.zlib.net\/","journal-title":"R"},{"key":"e_1_3_3_20_2","unstructured":"Xiaotian Han Jianwei Yang Houdong Hu Lei Zhang Jianfeng Gao and Pengchuan Zhang. 2021. Image Scene Graph Generation (SGG) Benchmark. arxiv:2107.12604 [cs.CV]."},{"key":"e_1_3_3_21_2","first-page":"14766","article-title":"Checkerboard context model for efficient learned image compression","author":"He Dailan","year":"2021","unstructured":"Dailan He, Yaoyan Zheng, Baochen Sun, Yan Wang, and Hongwei Qin. 2021. Checkerboard context model for efficient learned image compression. In CVPR. 14766\u201314775. Retrieved from https:\/\/api.semanticscholar.org\/CorpusID:232404386","journal-title":"CVPR"},{"key":"e_1_3_3_22_2","first-page":"12134","volume-title":"Adv. Neural Inf. Process. Syst. 32","author":"Hoogeboom Emiel","year":"2019","unstructured":"Emiel Hoogeboom, Jorn W. T. Peters, Rianne van den Berg, and Max Welling. 2019. Integer discrete flows and lossless compression. In Adv. Neural Inf. Process. Syst. 32. 12134\u201312144."},{"key":"e_1_3_3_23_2","doi-asserted-by":"publisher","DOI":"10.1109\/JRPROC.1952.273898"},{"key":"e_1_3_3_24_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.01025"},{"key":"e_1_3_3_25_2","doi-asserted-by":"publisher","DOI":"10.1145\/3460474"},{"key":"e_1_3_3_26_2","first-page":"3408","volume-title":"ICML.","author":"Kingma Friso H.","year":"2019","unstructured":"Friso H. Kingma, Pieter Abbeel, and Jonathan Ho. 2019. Bit-Swap: Recursive bits-back coding for lossless compression with hierarchical latent variables. In ICML.3408\u20133417."},{"key":"e_1_3_3_27_2","article-title":"Variational graph auto-encoders","volume":"1611","author":"Kipf Thomas N.","year":"2016","unstructured":"Thomas N. Kipf and Max Welling. 2016. Variational graph auto-encoders. CoRR abs\/1611.07308 (2016).","journal-title":"CoRR"},{"key":"e_1_3_3_28_2","volume-title":"ICLR.","author":"Kipf Thomas N.","year":"2017","unstructured":"Thomas N. Kipf and Max Welling. 2017. Semi-supervised classification with graph convolutional networks. In ICLR."},{"key":"e_1_3_3_29_2","article-title":"CMIX","author":"Knoll Byron","year":"2022","unstructured":"Byron Knoll. 2022. CMIX. Retrieved from http:\/\/www.byronknoll.com\/cmix.html","journal-title":"R"},{"key":"e_1_3_3_30_2","unstructured":"Alina Kuznetsova Hassan Rom Neil Alldrin Jasper R. R. Uijlings Ivan Krasin Jordi Pont-Tuset Shahab Kamali Stefan Popov Matteo Malloci Tom Duerig and Vittorio Ferrari. 2018. The open images dataset V4: unified image classification object detection and visual relationship detection at scale. CoRR abs\/1811.00982 (2018)."},{"key":"e_1_3_3_31_2","doi-asserted-by":"crossref","unstructured":"Ranjay Krishna Yuke Zhu Oliver Groth Justin Johnson Kenji Hata Joshua Kravitz Stephanie Chen Yannis Kalantidis Li-Jia Li David A. Shamma Michael S. Bernstein and Li Fei-Fei. 2017. Visual genome: Connecting language and vision using crowdsourced dense image annotations. Int. J. Comput. Vis. 123 1 (2017) 32\u201373.","DOI":"10.1007\/s11263-016-0981-7"},{"key":"e_1_3_3_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2020.2985225"},{"key":"e_1_3_3_33_2","doi-asserted-by":"crossref","unstructured":"Tsung-Yi Lin Michael Maire Serge J. Belongie James Hays Pietro Perona Deva Ramanan Piotr Doll\u00e1r and C. Lawrence Zitnick. 2014. Microsoft COCO: Common objects in context. ECCV 5 (2014) 740\u2013755.","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"e_1_3_3_34_2","doi-asserted-by":"crossref","unstructured":"Weiyao Lin Huabin Liu Shizhan Liu Yuxi Li Hongkai Xiong Guojun Qi and Nicu Sebe. 2023. HiEve: A large-scale benchmark for human-centric video analysis in complex events. Int. J. Comput. Vis. 131 11 (2023) 2994\u20133018.","DOI":"10.1007\/s11263-023-01842-6"},{"key":"e_1_3_3_35_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.image.2019.115659"},{"key":"e_1_3_3_36_2","first-page":"11541","article-title":"Fully convolutional scene graph generation","author":"Liu Hengyue","year":"2021","unstructured":"Hengyue Liu, Ning Yan, Masood S. Mortazavi, and Bir Bhanu. 2021. Fully convolutional scene graph generation. In CVPR. 11541\u201311551. Retrieved from https:\/\/api.semanticscholar.org\/CorpusID:232417534","journal-title":"CVPR"},{"key":"e_1_3_3_37_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2019.2957990"},{"key":"e_1_3_3_38_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2020.3026003"},{"key":"e_1_3_3_39_2","unstructured":"Matthew V. Mahoney. 2005. Adaptive weighing of context models for lossless data compression. Retrieved from https:\/\/cs.fit.edu\/mmahoney\/compression\/cs200516.pdf"},{"key":"e_1_3_3_40_2","first-page":"14111","volume-title":"CVPR","author":"Marino Kenneth","year":"2021","unstructured":"Kenneth Marino, Xinlei Chen, Devi Parikh, Abhinav Gupta, and Marcus Rohrbach. 2021. KRISP: Integrating implicit and symbolic knowledge for open-domain knowledge-based VQA. In CVPR. 14111\u201314121."},{"key":"e_1_3_3_41_2","first-page":"10629","volume-title":"CVPR.","author":"Mentzer Fabian","year":"2019","unstructured":"Fabian Mentzer, Eirikur Agustsson, Michael Tschannen, Radu Timofte, and Luc Van Gool. 2019. Practical full resolution learned lossless image compression. In CVPR.10629\u201310638."},{"key":"e_1_3_3_42_2","first-page":"10794","volume-title":"NeurIPS","author":"Minnen David","year":"2018","unstructured":"David Minnen, Johannes Ball\u00e9, and George Toderici. 2018. Joint autoregressive and hierarchical priors for learned image compression. In NeurIPS. 10794\u201310803."},{"key":"e_1_3_3_43_2","unstructured":"Igor Pavlov. 2021. LZMA SDK (Software Development Kit). Retrieved from https:\/\/www.7-zip.org\/sdk.html"},{"key":"e_1_3_3_44_2","first-page":"593","volume-title":"ESWC","author":"Schlichtkrull Michael Sejr","year":"2018","unstructured":"Michael Sejr Schlichtkrull, Thomas N. Kipf, Peter Bloem, Rianne van den Berg, Ivan Titov, and Max Welling. 2018. Modeling relational data with graph convolutional networks. In ESWC. 593\u2013607."},{"key":"e_1_3_3_45_2","volume-title":"*SEM","author":"Schluter Natalie","year":"2015","unstructured":"Natalie Schluter. 2015. The complexity of finding the maximum spanning DAG and other restrictions for DAG parsing of natural language. In *SEM."},{"key":"e_1_3_3_46_2","doi-asserted-by":"publisher","DOI":"10.1016\/0898-1221(81)90008-0"},{"key":"e_1_3_3_47_2","first-page":"13931","article-title":"Energy-based learning for scene graph generation","author":"Suhail Mohammed","year":"2021","unstructured":"Mohammed Suhail, Abhay Mittal, Behjat Siddiquie, Chris Broaddus, Jayan Eledath, G\u00e9rard G. Medioni, and Leonid Sigal. 2021. Energy-based learning for scene graph generation. In CVPR. 13931\u201313940. Retrieved from https:\/\/api.semanticscholar.org\/CorpusID:232104772","journal-title":"CVPR"},{"key":"e_1_3_3_48_2","first-page":"3713","volume-title":"CVPR","author":"Tang Kaihua","year":"2020","unstructured":"Kaihua Tang, Yulei Niu, Jianqiang Huang, Jiaxin Shi, and Hanwang Zhang. 2020. Unbiased scene graph generation from biased training. In CVPR. 3713\u20133722."},{"key":"e_1_3_3_49_2","doi-asserted-by":"publisher","DOI":"10.1109\/WACV45572.2020.9093614"},{"key":"e_1_3_3_50_2","first-page":"3097","volume-title":"CVPR","author":"Xu Danfei","year":"2017","unstructured":"Danfei Xu, Yuke Zhu, Christopher B. Choy, and Li Fei-Fei. 2017. Scene graph generation by iterative message passing. In CVPR. 3097\u20133106."},{"key":"e_1_3_3_51_2","first-page":"5449","volume-title":"ICML","author":"Xu Keyulu","year":"2018","unstructured":"Keyulu Xu, Chengtao Li, Yonglong Tian, Tomohiro Sonobe, Ken-ichi Kawarabayashi, and Stefanie Jegelka. 2018. Representation learning on graphs with jumping knowledge networks. In ICML. 5449\u20135458."},{"key":"e_1_3_3_52_2","volume-title":"ECCV","author":"Yao Ting","year":"2018","unstructured":"Ting Yao, Yingwei Pan, Yehao Li, and Tao Mei. 2018. Exploring visual relationship for image captioning. In ECCV. Retrieved from https:\/\/api.semanticscholar.org\/CorpusID:52304560"},{"key":"e_1_3_3_53_2","first-page":"10718","volume-title":"AAAI","author":"Yoon Sangwoong","year":"2021","unstructured":"Sangwoong Yoon, Woo-Young Kang, Sungwook Jeon, SeongEun Lee, Changjin Han, Jonghun Park, and Eun-Sol Kim. 2021. Image-to-image retrieval by learning similarity between scene graphs. In AAAI. AAAI Press, 10718\u201310726."},{"key":"e_1_3_3_54_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2020.3040074"},{"key":"e_1_3_3_55_2","first-page":"1","article-title":"Boosting scene graph generation with visual relation saliency","volume":"19","author":"Zhang Yong Hong","year":"2022","unstructured":"Yong Hong Zhang, Yingwei Pan, Ting Yao, Rui Huang, Tao Mei, and Chang Wen Chen. 2022. Boosting scene graph generation with visual relation saliency. ACM Trans. Multim. Comput., Commun. Applic 19 (2022), 1\u201317. Retrieved from https:\/\/api.semanticscholar.org\/CorpusID:247500307","journal-title":"ACM Trans. Multim. Comput., Commun. Applic"},{"key":"e_1_3_3_56_2","doi-asserted-by":"crossref","unstructured":"Jing Zhao Bin Li Jiahao Li Ruiqin Xiong and Yan Lu. 2023. A universal optimization framework for learning-based image codec. ACM Trans. Multim. Comput. Commun. Applic. 20 1 (2023) 1\u201319.","DOI":"10.1145\/3580499"},{"key":"e_1_3_3_57_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIT.1977.1055714"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3649503","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3649503","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T23:56:53Z","timestamp":1750291013000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3649503"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,3,27]]},"references-count":56,"journal-issue":{"issue":"7","published-print":{"date-parts":[[2024,7,31]]}},"alternative-id":["10.1145\/3649503"],"URL":"https:\/\/doi.org\/10.1145\/3649503","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,3,27]]},"assertion":[{"value":"2023-07-19","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-02-15","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-03-27","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}