{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,24]],"date-time":"2026-06-24T14:53:21Z","timestamp":1782312801248,"version":"3.54.5"},"reference-count":70,"publisher":"Association for Computing Machinery (ACM)","issue":"7","funder":[{"name":"National Key Research and Development Program of China","award":["2022YFB4500600"],"award-info":[{"award-number":["2022YFB4500600"]}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62276112"],"award-info":[{"award-number":["62276112"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Jilin Province Science and Technology Development Plan Key RD Project","award":["20230201088GX"],"award-info":[{"award-number":["20230201088GX"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2026,7,31]]},"abstract":"<jats:p>\n                    Recently, Remote Sensing Visual Question Answering (RSVQA) has attracted increasing attention from both academia and industry, which is the basis for understanding the underlying correspondence between remote sensing imagery and text descriptions. However, current methods are still insufficient in learning discriminative visual and textual representations for answer reasoning, mainly due to two reasons: (1) the remote sensing image environment is complex and changeable, and the target scales vary significantly, making it difficult to extract discriminative visual features; and (2) there is a lack of effective guidance from remote sensing domain knowledge to learn discriminative features. To this end, we propose a\n                    <jats:italic toggle=\"yes\">D<\/jats:italic>\n                    iscriminative\n                    <jats:italic toggle=\"yes\">R<\/jats:italic>\n                    epresentation\n                    <jats:italic toggle=\"yes\">L<\/jats:italic>\n                    earning (DRL) method that includes two key strategies: visual feature enhancement and prior knowledge guidance. Specifically, we employ the Fourier transform to simulate the diverse visual environment and force the model to mine discriminative visual representations by imposing consistency constraints with the original features. In addition, we leverage the Remote Sensing Multimodal Large Language Model (RSMLLM) to generate captions rich in remote sensing domain-specific prior knowledge. These captions, derived from RSMLLM\u2019s powerful knowledge integration and summarization capabilities, can then be compared and fused with visual representations to generate more discriminative representations, which are ultimately used for answer reasoning. Finally, recognizing that most existing RSVQA methods rely solely on static remote sensing images, we introduce RSVideoQA, a novel satellite video question answering dataset. This dataset is designed to facilitate the exploration of the rich spatio-temporal dynamics inherent in video sequences. Experimental results across three distinct datasets validate the effectiveness of our proposed method. Our dataset and code will be released at\n                    <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" ext-link-type=\"uri\" xlink:href=\"https:\/\/github.com\/chill-han\/DRL\">https:\/\/github.com\/chill-han\/DRL<\/jats:ext-link>\n                    .\n                  <\/jats:p>","DOI":"10.1145\/3777473","type":"journal-article","created":{"date-parts":[[2025,11,19]],"date-time":"2025-11-19T16:05:03Z","timestamp":1763568303000},"page":"1-21","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["Discriminative Representation Learning for Remote Sensing Visual Question Answering"],"prefix":"10.1145","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-2037-6692","authenticated-orcid":false,"given":"Yingda","family":"Lyu","sequence":"first","affiliation":[{"name":"College of Computer Science and Technology, Center for Public Education Research, Jilin University, Changchun, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-3797-6763","authenticated-orcid":false,"given":"Han","family":"Yan","sequence":"additional","affiliation":[{"name":"College of Computer Science and Technology, Jilin University, Changchun, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3023-3027","authenticated-orcid":false,"given":"Yu","family":"Liu","sequence":"additional","affiliation":[{"name":"College of Computer Science and Technology, Jilin University, Changchun, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9410-4120","authenticated-orcid":false,"given":"Haipeng","family":"Chen","sequence":"additional","affiliation":[{"name":"College of Computer Science and Technology, Jilin University, Changchun, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,6,24]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.279"},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1109\/TGRS.2020.2988782"},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.1109\/TGRS.2022.3192460"},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.1109\/TGRS.2022.3173811"},{"key":"e_1_3_1_6_2","doi-asserted-by":"publisher","DOI":"10.1109\/TGRS.2023.3312479"},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v38i6.28357"},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.1109\/TGRS.2023.3237606"},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.24963\/ijcai.2025\/647"},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.isprsjprs.2024.03.013"},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","DOI":"10.1109\/TGRS.2022.3140809"},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.670"},{"key":"e_1_3_1_13_2","volume-title":"Proceedings of the IEEE International Conference on Computer Vision (ICCV)","author":"Kafle Kushal","year":"2017","unstructured":"Kushal Kafle and Christopher Kanan. 2017. An analysis of visual question answering algorithms. In Proceedings of the IEEE International Conference on Computer Vision (ICCV)."},{"key":"e_1_3_1_14_2","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Agrawal Aishwarya","year":"2018","unstructured":"Aishwarya Agrawal, Dhruv Batra, Devi Parikh, and Aniruddha Kembhavi. 2018. Don\u2019t just assume; look and answer: Overcoming priors for visual question answering. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)."},{"key":"e_1_3_1_15_2","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Johnson Justin","year":"2017","unstructured":"Justin Johnson, Bharath Hariharan, Laurens van der Maaten, Li Fei-Fei, C. Lawrence Zitnick, and Ross Girshick. 2017. CLEVR: A diagnostic dataset for compositional language and elementary visual reasoning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)."},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00686"},{"key":"e_1_3_1_17_2","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Noh Hyeonwoo","year":"2016","unstructured":"Hyeonwoo Noh, Paul Hongsuck Seo, and Bohyung Han. 2016. Image question answering using convolutional neural network with dynamic parameter prediction. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)."},{"key":"e_1_3_1_18_2","volume-title":"Advances in Neural Information Processing Systems","volume":"29","author":"Kim Jin-Hwa","year":"2016","unstructured":"Jin-Hwa Kim, Sang-Woo Lee, Donghyun Kwak, Min-Oh Heo, Jeonghee Kim, Jung-Woo Ha, and Byoung-Tak Zhang. 2016. Multimodal residual learning for visual QA. In Advances in Neural Information Processing Systems. D. Lee, M. Sugiyama, U. Luxburg, I. Guyon, and R. Garnett (Eds.), Vol. 29, Curran Associates, Inc. Retrieved from https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2016\/file\/9b04d152845ec0a378394003c96da594-Paper.pdf"},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICME.2017.8019436"},{"key":"e_1_3_1_20_2","volume-title":"International Conference on Learning Representations","author":"Kim Jin-Hwa","year":"2017","unstructured":"Jin-Hwa Kim, Kyoung-Woon On, Woosang Lim, Jeonghee Kim, Jung-Woo Ha, and Byoung-Tak Zhang. 2017. Hadamard product for low-rank bilinear pooling. In International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=r1rhWnZkg"},{"key":"e_1_3_1_21_2","volume-title":"Proceedings of the IEEE International Conference on Computer Vision (ICCV)","author":"Yu Zhou","year":"2017","unstructured":"Zhou Yu, Jun Yu, Jianping Fan, and Dacheng Tao. 2017. Multi-modal factorized bilinear pooling with co-attention learning for visual question answering. In Proceedings of the IEEE International Conference on Computer Vision (ICCV)."},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D16-1044"},{"key":"e_1_3_1_23_2","volume-title":"Proceedings of the IEEE International Conference on Computer Vision (ICCV)","author":"Ben-Younes Hedi","year":"2017","unstructured":"Hedi Ben-Younes, Remi Cadene, Matthieu Cord, and Nicolas Thome. 2017. MUTAN: Multimodal tucker fusion for visual question answering. In Proceedings of the IEEE International Conference on Computer Vision (ICCV)."},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.1109\/TNNLS.2018.2817340"},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.10"},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00636"},{"key":"e_1_3_1_27_2","unstructured":"Jingkuan Song Pengpeng Zeng Lianli Gao and Heng Tao Shen. 2022. From pixels to objects: Cubic visual attention for visual question answering. arXiv:2206.01923. Retrieved from https:\/\/arxiv.org\/abs\/2206.01923"},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2023.109339"},{"key":"e_1_3_1_29_2","unstructured":"Lu Jiasen Dhruv Batra Devi Parikh and Stefan Lee. 2019. ViLBERT: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. In Advances in Neural Information Processing Systems. H. Wallach H. Larochelle A. Beygelzimer F. d'Alch\u00e9-Buc E. Fox and R. Garnett (Eds.) Vol. 32 Curran Associates Inc. Retrieved from https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2019\/file\/c74d97b01eae257e44aa9d5bade97baf-Paper.pdf"},{"key":"e_1_3_1_30_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Su Weijie","year":"2020","unstructured":"Weijie Su, Xizhou Zhu, Yue Cao, Bin Li, Lewei Lu, Furu Wei, and Jifeng Dai. 2020. VL-BERT: Pre-training of generic visual-linguistic representations. In Proceedings of the International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=SygXPaEYvH"},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00644"},{"key":"e_1_3_1_32_2","first-page":"3784","volume-title":"Advances in Neural Information Processing Systems","volume":"34","author":"Wen Zhiquan","year":"2021","unstructured":"Zhiquan Wen, Guanghui Xu, Mingkui Tan, Qingyao Wu, and Qi Wu. 2021. Debiased visual question answering from feature and sample perspectives. In Advances in Neural Information Processing Systems. M. Ranzato, A. Beygelzimer, Y. Dauphin, P.S. Liang, and J. Wortman Vaughan (Eds.), Vol. 34, Curran Associates, Inc., 3784\u20133796. Retrieved from https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2021\/file\/1f4477bad7af3616c1f933a02bfabe4e-Paper.pdf"},{"key":"e_1_3_1_33_2","first-page":"19730","volume-title":"Proceedings of the 40th International Conference on Machine Learning","volume":"202","author":"Li Junnan","year":"2023","unstructured":"Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. 2023. BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In Proceedings of the 40th International Conference on Machine Learning. Andreas Krause, Emma Brunskill, Kyunghyun Cho, Barbara Engelhardt, Sivan Sabato, and Jonathan Scarlett (Eds.), Proceedings of Machine Learning Research, Vol. 202, PMLR, 19730\u201319742. Retrieved from https:\/\/proceedings.mlr.press\/v202\/li23q.html"},{"key":"e_1_3_1_34_2","doi-asserted-by":"publisher","DOI":"10.52202\/075280-2142"},{"key":"e_1_3_1_35_2","first-page":"10560","volume-title":"Advances in Neural Information Processing Systems","volume":"35","author":"Lin Yuanze","year":"2022","unstructured":"Yuanze Lin, Yujia Xie, Dongdong Chen, Yichong Xu, Chenguang Zhu, and Lu Yuan. 2022. REVIVE: Regional visual representation matters in knowledge-based visual question answering. In Advances in Neural Information Processing Systems. S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh (Eds.), Vol. 35, Curran Associates, Inc., 10560\u201310571. Retrieved from https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2022\/file\/44956951349095f74492a5471128a7e0-Paper-Conference.pdf"},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.00501"},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.1109\/TNNLS.2020.3017530"},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v38i2.27888"},{"key":"e_1_3_1_39_2","first-page":"14974","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Shao Zhenwei","year":"2023","unstructured":"Zhenwei Shao, Zhou Yu, Meng Wang, and Jun Yu. 2023. Prompting large language models with answer heuristics for knowledge-based visual question answering. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 14974\u201314983."},{"key":"e_1_3_1_40_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2021.3120867"},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2022.3205212"},{"key":"e_1_3_1_42_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.emnlp-demo.49"},{"key":"e_1_3_1_43_2","doi-asserted-by":"publisher","DOI":"10.1109\/IGARSS47720.2021.9553307"},{"key":"e_1_3_1_44_2","doi-asserted-by":"publisher","DOI":"10.1109\/TGRS.2021.3079918"},{"key":"e_1_3_1_45_2","doi-asserted-by":"publisher","DOI":"10.1109\/JSTARS.2023.3261361"},{"key":"e_1_3_1_46_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.isprsjprs.2024.06.002"},{"key":"e_1_3_1_47_2","doi-asserted-by":"publisher","DOI":"10.1109\/TGRS.2022.3203314"},{"key":"e_1_3_1_48_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2021.3090981"},{"key":"e_1_3_1_49_2","first-page":"8748","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Radford Alec","year":"2021","unstructured":"Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. In Proceedings of the International Conference on Machine Learning. PMLR, 8748\u20138763."},{"key":"e_1_3_1_50_2","first-page":"1372","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops","author":"Chappuis Christel","year":"2022","unstructured":"Christel Chappuis, Val\u00e9rie Zermatten, Sylvain Lobry, Bertrand Le Saux, and Devis Tuia. 2022. Prompt-RSVQA: Prompting visual context to a language model for remote sensing visual question answering. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, 1372\u20131381."},{"key":"e_1_3_1_51_2","doi-asserted-by":"publisher","DOI":"10.1109\/TGRS.2024.3502800"},{"key":"e_1_3_1_52_2","doi-asserted-by":"publisher","DOI":"10.1109\/LGRS.2024.3490534"},{"key":"e_1_3_1_53_2","doi-asserted-by":"publisher","DOI":"10.1109\/TGRS.2024.3413174"},{"key":"e_1_3_1_54_2","first-page":"27831","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Kuckreja Kartik","year":"2024","unstructured":"Kartik Kuckreja, Muhammad Sohail Danish, Muzammal Naseer, Abhijit Das, Salman Khan, and Fahad Shahbaz Khan. 2024. GeoChat: Grounded large vision-language model for remote sensing. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 27831\u201327840."},{"key":"e_1_3_1_55_2","unstructured":"Junwei Luo Zhen Pang Yongjun Zhang Tingzhu Wang Linlin Wang Bo Dang Jiangwei Lao Jian Wang Jingdong Chen Yihua Tan and Yansheng Li. 2024. SkySenseGPT: A fine-grained instruction tuning dataset and model for remote sensing vision-language understanding. arXiv:2406.10100. Retrieved from https:\/\/arxiv.org\/abs\/2406.10100"},{"issue":"5","key":"e_1_3_1_56_2","first-page":"6","article-title":"STAR: A first-ever dataset and a large-scale benchmark for scene graph generation in large-size satellite imagery","volume":"2","author":"Li Yansheng","year":"2024","unstructured":"Yansheng Li, Linlin Wang, Tingzhu Wang, Xue Yang, Junwei Luo, Qi Wang, Youming Deng, Wenbin Wang, Xian Sun, Haifeng Li, et al. 2024. STAR: A first-ever dataset and a large-scale benchmark for scene graph generation in large-size satellite imagery. IEEE Transactions on Pattern Analysis and Machine Intelligence 2, 5 (2024), 6.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"e_1_3_1_57_2","doi-asserted-by":"crossref","first-page":"440","DOI":"10.1007\/978-3-031-72904-1_26","volume-title":"Computer Vision\u2014ECCV 2024","author":"Muhtar Dilxat","year":"2025","unstructured":"Dilxat Muhtar, Zhenshi Li, Feng Gu, Xueliang Zhang, and Pengfeng Xiao. 2025. LHRS-Bot: Empowering remote sensing with VGI-enhanced large multimodal language model. In Computer Vision\u2014ECCV 2024. Ale\u0161 Leonardis, Elisa Ricci, Stefan Roth, Olga Russakovsky, Torsten Sattler, and G\u00fcl Varol (Eds.), Springer Nature, Cham, 440\u2013457."},{"key":"e_1_3_1_58_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.isprsjprs.2025.01.020"},{"key":"e_1_3_1_59_2","doi-asserted-by":"publisher","DOI":"10.3390\/rs16091477"},{"key":"e_1_3_1_60_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v39i6.32683"},{"key":"e_1_3_1_61_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.01415"},{"key":"e_1_3_1_62_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v38i12.29314"},{"key":"e_1_3_1_63_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.00788"},{"key":"e_1_3_1_64_2","doi-asserted-by":"publisher","DOI":"10.1145\/3700596"},{"key":"e_1_3_1_65_2","doi-asserted-by":"publisher","DOI":"10.1109\/TGRS.2024.3390838"},{"key":"e_1_3_1_66_2","doi-asserted-by":"publisher","DOI":"10.1109\/TGRS.2024.3449154"},{"key":"e_1_3_1_67_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00175"},{"key":"e_1_3_1_68_2","unstructured":"Huaishao Luo Lei Ji Ming Zhong Yang Chen Wen Lei Nan Duan and Tianrui Li. 2021. CLIP4Clip: An empirical study of CLIP for end to end video clip retrieval. arXiv:2104.08860. Retrieved from https:\/\/arxiv.org\/abs\/2104.08860"},{"key":"e_1_3_1_69_2","unstructured":"Yi Wang Yinan He Yizhuo Li Kunchang Li Jiashuo Yu Xin Ma Xinyuan Chen Yaohui Wang Ping Luo Ziwei Liu Yali Wang Limin Wang and Yu Qiao. 2023. InternVid: A large-scale video-text dataset for multimodal understanding and generation. arXiv:2307.06942. Retrieved from https:\/\/arxiv.org\/abs\/2307.06942"},{"key":"e_1_3_1_70_2","unstructured":"Yi Wang Kunchang Li Yizhuo Li Yinan He Bingkun Huang Zhiyu Zhao Hongjie Zhang Jilan Xu Yi Liu Zun Wang et al. 2022. InternVideo: General video foundation models via generative and discriminative learning. arXiv:2212.03191. Retrieved from https:\/\/arxiv.org\/abs\/2212.03191"},{"key":"e_1_3_1_71_2","doi-asserted-by":"crossref","unstructured":"Jiapeng Wang Chengyu Wang Kunzhe Huang Jun Huang and Lianwen Jin. 2024. VideoCLIP-XL: Advancing long description understanding for video CLIP models. arXiv:2410.00741. Retrieved from https:\/\/arxiv.org\/abs\/2410.00741","DOI":"10.18653\/v1\/2024.emnlp-main.898"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3777473","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,24]],"date-time":"2026-06-24T14:38:26Z","timestamp":1782311906000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3777473"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6,24]]},"references-count":70,"journal-issue":{"issue":"7","published-print":{"date-parts":[[2026,7,31]]}},"alternative-id":["10.1145\/3777473"],"URL":"https:\/\/doi.org\/10.1145\/3777473","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,6,24]]},"assertion":[{"value":"2025-06-19","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-10-31","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-06-24","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}