{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,24]],"date-time":"2026-06-24T14:54:04Z","timestamp":1782312844254,"version":"3.54.5"},"reference-count":115,"publisher":"Association for Computing Machinery (ACM)","issue":"7","license":[{"start":{"date-parts":[[2026,6,24]],"date-time":"2026-06-24T00:00:00Z","timestamp":1782259200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62276155, 62576195, and 62576194"],"award-info":[{"award-number":["62276155, 62576195, and 62576194"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"China National University Student Innovation & Entrepreneurship Development Program","award":["2025282 and 2025283"],"award-info":[{"award-number":["2025282 and 2025283"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2026,7,31]]},"abstract":"<jats:p>\n                    Composed Video Retrieval (CVR) is a novel video retrieval paradigm. Unlike traditional single-modal video retrieval paradigms (e.g., text to video or video to video), CVR employs multi-modal queries (including both a reference video and a natural language modification) to retrieve the target video that best matches the modified reference video. Existing CVR methods primarily rely on generalized knowledge from vision-language pretrained models or utilize caption expansions to enhance video comprehension. However, these approaches overlook the benefits offered by the\n                    <jats:italic toggle=\"yes\">shareability<\/jats:italic>\n                    and\n                    <jats:italic toggle=\"yes\">variability<\/jats:italic>\n                    of videos for multi-modal query understanding. To overcome this limitation, we introduce a novel CVR framework named shaREd and diFferential semantIcs eNhancement nEtwork (\n                    <jats:italic toggle=\"yes\">REFINE<\/jats:italic>\n                    ). REFINE is the first framework to exploit the\n                    <jats:italic toggle=\"yes\">shareability<\/jats:italic>\n                    and\n                    <jats:italic toggle=\"yes\">variability<\/jats:italic>\n                    of videos to improve multi-modal query comprehension. Specifically, REFINE leverages learnable tokens to achieve enhanced shared feature representation. Moreover, it introduces a carefully designed Differential Block to disentangle differential semantics between frames and employs modification associations to guide multi-modal query feature fusion. Additionally, REFINE has been extended to the Composed Image Retrieval task, making it effectively generalize across existing composed multi-modal retrieval scenarios and outperform existing methods. Extensive qualitative and quantitative evaluations on four benchmark datasets validate the superiority of the proposed REFINE framework.\n                  <\/jats:p>","DOI":"10.1145\/3796712","type":"journal-article","created":{"date-parts":[[2026,2,14]],"date-time":"2026-02-14T14:27:41Z","timestamp":1771079261000},"page":"1-24","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["REFINE: Composed Video Retrieval via Shared and Differential Semantics Enhancement"],"prefix":"10.1145","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-5653-8286","authenticated-orcid":false,"given":"Yupeng","family":"Hu","sequence":"first","affiliation":[{"name":"School of Software, Shandong University, Jinan, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0001-5136-159X","authenticated-orcid":false,"given":"Zixu","family":"Li","sequence":"additional","affiliation":[{"name":"School of Software, Shandong University, Jinan, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0003-0365-8553","authenticated-orcid":false,"given":"Zhiwei","family":"Chen","sequence":"additional","affiliation":[{"name":"School of Software, Shandong University, Jinan, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0009-0318-7599","authenticated-orcid":false,"given":"Qinlei","family":"Huang","sequence":"additional","affiliation":[{"name":"School of Software, Shandong University, Jinan, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-7724-5662","authenticated-orcid":false,"given":"Zhiheng","family":"Fu","sequence":"additional","affiliation":[{"name":"School of Software, Shandong University, Jinan, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1492-0970","authenticated-orcid":false,"given":"Mingzhu","family":"Xu","sequence":"additional","affiliation":[{"name":"School of Software, Shandong University, Jinan, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1476-0273","authenticated-orcid":false,"given":"Liqiang","family":"Nie","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Harbin Institute of Technology (Shenzhen), Shenzhen, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,6,24]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.1145\/3620669"},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2021.3090521"},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2023.3278131"},{"key":"e_1_3_1_5_2","unstructured":"Zhenlong Yuan Xiangyan Qu Chengxuan Qian Rui Chen Jing Tang Lei Sun Xiangxiang Chu Dapeng Zhang Yiwei Wang Yujun Cai et al. 2025. Video-STAR: Reinforcing open-vocabulary action recognition with tools. arXiv:2510.08480. Retrieved from https:\/\/arxiv.org\/abs\/2510.08480"},{"key":"e_1_3_1_6_2","first-page":"9753","volume-title":"AAAI","volume":"39","author":"Yuan Zhenlong","year":"2025","unstructured":"Zhenlong Yuan, Cong Liu, Fei Shen, Zhaoxin Li, Jinguo Luo, Tianlu Mao, and Zhaoqi Wang. 2025. MSP-MVS: Multi-granularity segmentation prior guided multi-view stereo. In AAAI, Vol. 39, 9753\u20139762."},{"key":"e_1_3_1_7_2","unstructured":"Xinlei Yu Chengming Xu Guibin Zhang Yongbo He Zhangquan Chen Zhucun Xue Jiangning Zhang Yue Liao Xiaobin Hu Yu-Gang Jiang et al. 2025. Visual multi-agent system: Mitigating hallucination snowballing via visual flow. arXiv:2509.21789. Retrieved from https:\/\/arxiv.org\/abs\/2509.21789"},{"key":"e_1_3_1_8_2","unstructured":"Shilin Lu Zhuming Lian Zihan Zhou Shaocong Zhang Chen Zhao and Adams Wai-Kin Kong. 2025. Does FLUX already know how to perform physically plausible image composition? arXiv:2509.21278. Retrieved from https:\/\/arxiv.org\/abs\/2509.21278"},{"key":"e_1_3_1_9_2","first-page":"9743","volume-title":"AAAI","volume":"39","author":"Yuan Zhenlong","year":"2025","unstructured":"Zhenlong Yuan, Jinguo Luo, Fei Shen, Zhaoxin Li, Cong Liu, Tianlu Mao, and Zhaoqi Wang. 2025. DVP-MVS: Synergize depth-edge and visibility prior for multi-view stereo. In AAAI, Vol. 39, 9743\u20139752."},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2021.3073867"},{"key":"e_1_3_1_11_2","first-page":"1","article-title":"DUDB: Deep unfolding-based dual-branch feature fusion network for pan-sharpening remote sensing images","volume":"62","author":"Tao Hailin","year":"2023","unstructured":"Hailin Tao, Jinjiang Li, Zhen Hua, and Fan Zhang. 2023. DUDB: Deep unfolding-based dual-branch feature fusion network for pan-sharpening remote sensing images. IEEE Transactions on Geoscience and Remote Sensing 62 (2023), 1\u201317.","journal-title":"IEEE Transactions on Geoscience and Remote Sensing"},{"key":"e_1_3_1_12_2","first-page":"32967","volume-title":"ACL","author":"Tian Yang","year":"2025","unstructured":"Yang Tian, Fan Liu, Jingyuan Zhang, V. W. Yupeng Hu, and Liqiang Nie. 2025. CoRe-MMRAG: Cross-source knowledge reconciliation for multimodal RAG. In ACL, 32967\u201332982."},{"key":"e_1_3_1_13_2","unstructured":"Zihan Zhou Shilin Lu Shuli Leng Shaocong Zhang Zhuming Lian Xinlei Yu and Adams Wai-Kin Kong. 2025. DragFlow: Unleashing DiT priors with region based supervision for drag editing. arXiv:2510.02253. Retrieved from https:\/\/arxiv.org\/abs\/2510.02253"},{"key":"e_1_3_1_14_2","unstructured":"Changshi Zhou Haichuan Xu Ningquan Gu Zhipeng Wang Bin Cheng Pengpeng Zhang Yanchao Dong Mitsuhiro Hayashibe Yanmin Zhou and Bin He. 2025. Language-guided long horizon manipulation with LLM-based planning and visual perception. arXiv:2509.02324. Retrieved from https:\/\/arxiv.org\/abs\/2509.02324"},{"key":"e_1_3_1_15_2","unstructured":"Yang Tian Fan Liu Jingyuan Zhang Wei Bi Yupeng Hu and Liqiang Nie. 2025. Open multimodal retrieval-augmented factual image generation. arXiv:2510.22521. Retrieved from https:\/\/arxiv.org\/abs\/2510.22521"},{"key":"e_1_3_1_16_2","unstructured":"Shilin Lu Zihan Zhou Jiayou Lu Yuanzhi Zhu and Adams Wai-Kin Kong. 2024. Robust watermarking using generative priors against image editing: From benchmarking to advances. arXiv:2410.18775. Retrieved from https:\/\/arxiv.org\/abs\/2410.18775"},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2024.3508752"},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2025.3587485"},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","DOI":"10.1109\/TGRS.2025.3574962"},{"key":"e_1_3_1_20_2","doi-asserted-by":"crossref","first-page":"673","DOI":"10.1109\/JSAC.2025.3610398","article-title":"SIMAC: A semantic-driven integrated multimodal sensing and communication framework","volume":"44","author":"Peng Yubo","year":"2025","unstructured":"Yubo Peng, Luping Xiang, Kun Yang, Feibo Jiang, Kezhi Wang, and Dapeng Oliver Wu. 2025. SIMAC: A semantic-driven integrated multimodal sensing and communication framework. IEEE Journal on Selected Areas in Communications 44 (2025), 673\u2013688.","journal-title":"IEEE Journal on Selected Areas in Communications"},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.2196\/68205"},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.17265\/2160-6579\/2016.12.002"},{"key":"e_1_3_1_23_2","unstructured":"O. U. Yifan. 2018. Participating in Chinese Social Question and Answer Communities: A Case Study of Zhihu.com. Chinese Journal of Communication (2025) 1\u201317."},{"key":"e_1_3_1_24_2","doi-asserted-by":"crossref","unstructured":"Xu Liu Yibo Lu Xinxian Wang and Xinyu Wu. 2025. Training-free multi-style fusion through reference-based adaptive modulation. arXiv:250918602. Retrieved from https:\/\/arxiv.org\/abs\/2509.18602","DOI":"10.1007\/978-981-95-4398-4_11"},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2023.3272169"},{"key":"e_1_3_1_26_2","volume-title":"ICML","author":"Pu Ruitao","year":"2025","unstructured":"Ruitao Pu, Yang Qin, Xiaomin Song, Dezhong Peng, Zhenwen Ren, and Yuan Sun. 2025. SHE: Streaming-media hashing retrieval. In ICML."},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v39i19.34199"},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2024.3354928"},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2023.3237615"},{"key":"e_1_3_1_30_2","doi-asserted-by":"crossref","first-page":"130845","DOI":"10.1016\/j.neucom.2025.130845","article-title":"Conditional variational underwater image enhancement with kernel decomposition and adaptive hybrid normalization","volume":"650","author":"Zhang Haopeng","year":"2025","unstructured":"Haopeng Zhang, Hongli Xu, Hao Liu, Xiaosheng Yu, Xiangyue Zhang, and Chengdong Wu. 2025. Conditional variational underwater image enhancement with kernel decomposition and adaptive hybrid normalization. Neurocomputing 650 (2025), 130845.","journal-title":"Neurocomputing"},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v39i8.32881"},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2024.3395969"},{"key":"e_1_3_1_33_2","doi-asserted-by":"publisher","DOI":"10.1145\/3581783.3611864"},{"key":"e_1_3_1_34_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2022.3141255"},{"issue":"11","key":"e_1_3_1_35_2","doi-asserted-by":"crossref","first-page":"5332","DOI":"10.1109\/TCYB.2025.3571913","article-title":"Cross-model nested fusion network for salient object detection in optical remote sensing images","volume":"55","author":"Xu Mingzhu","year":"2025","unstructured":"Mingzhu Xu, Sen Wang, Yupeng Hu, Haoyu Tang, Runmin Cong, and Liqiang Nie. 2025. Cross-model nested fusion network for salient object detection in optical remote sensing images. IEEE Transactions on Cybernetics 55, 11 (2025), 5332\u20135345.","journal-title":"IEEE Transactions on Cybernetics"},{"key":"e_1_3_1_36_2","article-title":"Energy-constrained motion planning and scheduling for autonomous robots in complex environments","author":"Ma Zhichao","year":"2025","unstructured":"Zhichao Ma, Zheyu Zhang, Zijun Gao, Aijia Sun, Yinuo Yang, and Hao Liu. 2025. Energy-constrained motion planning and scheduling for autonomous robots in complex environments. Preprints. Retrieved from https:\/\/doi.org\/10.20944\/preprints202509.1316.v1","journal-title":"Preprints"},{"key":"e_1_3_1_37_2","first-page":"128084","volume-title":"ICRA","author":"Liu Hao","year":"2025","unstructured":"Hao Liu, Xiang Li, Xiang Zhang, Gang Liu, and Mingquan Lu. 2025. In-pipe navigation development environment and a smooth path planning method on pipeline surface. In ICRA. IEEE, 128084\u2013128090."},{"key":"e_1_3_1_38_2","unstructured":"Zhilin Zhang Naveed Ahmed Saleem Janvekar Pengbin Feng and Nitika Bhaskar. 2025. Graph-based detection of abusive computational nodes (2025). US Patent 12223056."},{"key":"e_1_3_1_39_2","doi-asserted-by":"crossref","first-page":"126292","DOI":"10.1016\/j.eswa.2024.126292","article-title":"GLC: A dual-perspective approach for identifying influential nodes in complex networks","volume":"268","author":"Ruan Yirun","year":"2025","unstructured":"Yirun Ruan, Sizheng Liu, Jun Tang, Yanming Guo, and Tianyuan Yu. 2025. GLC: A dual-perspective approach for identifying influential nodes in complex networks. Expert Systems with Applications 268 (2025), 126292.","journal-title":"Expert Systems with Applications"},{"key":"e_1_3_1_40_2","unstructured":"Yun Ting and Cloud Listening. 2024. When Radio Become a Broadcasting Application. Retrieved from https:\/\/doi.org\/revista-hermes-la-revue-2023-2-page-173?lang=es"},{"key":"e_1_3_1_41_2","unstructured":"Yichen Wu Xu Liu Chenxuan Zhao and Xinyu Wu. 2025. Prompt-guided dual latent steering for inversion problems. arXiv:250918619. Retrieved from https:\/\/arxiv.org\/abs\/2509.18619"},{"key":"e_1_3_1_42_2","first-page":"4569","volume-title":"IJCAI","author":"Liu Kaiming","year":"2024","unstructured":"Kaiming Liu, Yunhong Gong, Yu Cao, Zhenwen Ren, Dezhong Peng, and Yuan Sun. 2024. Dual semantic fusion hashing for multi-label cross-modal retrieval. In IJCAI, 4569\u20134577."},{"key":"e_1_3_1_43_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIFS.2025.3608672"},{"key":"e_1_3_1_44_2","first-page":"3046","article-title":"Post-distillation via neural resuscitation","volume":"26","author":"Bao Zhiqiang","year":"2023","unstructured":"Zhiqiang Bao, Zihao Chen, Chang-Dong Wang, Wei-Shi Zheng, Zhenhua Huang, and Yunwen Chen. 2023. Post-distillation via neural resuscitation. IEEE Transactions on Multimedia 26 (2023), 3046\u20133060.","journal-title":"IEEE Transactions on Multimedia"},{"key":"e_1_3_1_45_2","unstructured":"Changshi Zhou Feng Luan Jiarui Hu Shaoqiang Meng Zhipeng Wang Yanchao Dong Yanmin Zhou and Bin He. 2025. Learning efficient robotic garment manipulation with standardization. arXiv:2506.22769. Retrieved from https:\/\/arxiv.org\/abs\/2506.22769"},{"key":"e_1_3_1_46_2","first-page":"5845","volume-title":"Transactions on Mechatronics","volume":"30","author":"Zhou Changshi","year":"2025","unstructured":"Changshi Zhou, Rong Jiang, Feng Luan, Shaoqiang Meng, Zhipeng Wang, Yanchao Dong, Yanmin Zhou, and Bin He. 2025. Dual-arm robotic fabric manipulation with quasi-static and dynamic primitives for rapid garment flattening. Transactions on Mechatronics 30 (2025), 5845\u20135855."},{"key":"e_1_3_1_47_2","doi-asserted-by":"crossref","first-page":"14448","DOI":"10.1109\/TASE.2025.3560680","article-title":"SSFold: Learning to fold arbitrary crumpled cloth using graph dynamics from human demonstration","volume":"22","author":"Zhou Changshi","year":"2025","unstructured":"Changshi Zhou, Haichuan Xu, Jiarui Hu, Feng Luan, Zhipeng Wang, Yanchao Dong, Yanmin Zhou, and Bin He. 2025. SSFold: Learning to fold arbitrary crumpled cloth using graph dynamics from human demonstration. IEEE Transactions on Automation Science and Engineering 22 (2025), 14448\u201314460.","journal-title":"IEEE Transactions on Automation Science and Engineering"},{"key":"e_1_3_1_48_2","doi-asserted-by":"crossref","first-page":"5270","DOI":"10.1609\/aaai.v38i6.28334","article-title":"CoVR: Learning composed video retrieval from web video captions","volume":"38","author":"Ventura Lucas","year":"2024","unstructured":"Lucas Ventura, Antoine Yang, Cordelia Schmid, and G\u00fcl Varol. 2024. CoVR: Learning composed video retrieval from web video captions. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 38, 5270\u20135279.","journal-title":"Proceedings of the AAAI Conference on Artificial Intelligence"},{"key":"e_1_3_1_49_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2024.3463799"},{"key":"e_1_3_1_50_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.02540"},{"key":"e_1_3_1_51_2","first-page":"1","volume-title":"ECCV","author":"Hummel Thomas","year":"2024","unstructured":"Thomas Hummel, Shyamgopal Karthik, Mariana-Iuliana Georgescu, and Zeynep Akata. 2024. EgoCVR: An egocentric benchmark for fine-grained composed video retrieval. In ECCV. Springer, 1\u201317."},{"key":"e_1_3_1_52_2","first-page":"12888","volume-title":"ICML","author":"Li Junnan","year":"2022","unstructured":"Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. 2022. BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In ICML. PMLR, 12888\u201312900."},{"key":"e_1_3_1_53_2","first-page":"19730","volume-title":"ICML","author":"Li Junnan","year":"2023","unstructured":"Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. 2023. BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In ICML. PMLR, 19730\u201319742."},{"key":"e_1_3_1_54_2","doi-asserted-by":"publisher","DOI":"10.1109\/TGRS.2023.3300317"},{"key":"e_1_3_1_55_2","unstructured":"Zhenlong Yuan Jing Tang Jinguo Luo Rui Chen Chengxuan Qian Lei Sun Xiangxiang Chu Yujun Cai Dapeng Zhang and Shuo Li. 2025. AutoDrive-R2: Incentivizing reasoning and self-reflection capacity for VLA model in autonomous driving. arXiv:2509.01944. Retrieved from https:\/\/arxiv.org\/abs\/2509.01944"},{"key":"e_1_3_1_56_2","first-page":"1298","volume-title":"ACM MM","author":"Yu Xinlei","year":"2025","unstructured":"Xinlei Yu, Changmiao Wang, Hui Jin, Ahmed Elazab, Gangyong Jia, Xiang Wan, Changqing Zou, and Ruiquan Ge. 2025. CRISP-SAM2: SAM2 with cross-modal interaction and semantic prompting for multi-organ segmentation. In ACM MM, 1298\u20131307."},{"key":"e_1_3_1_57_2","first-page":"1506","volume-title":"EMBC","author":"Zhao Haochen","year":"2022","unstructured":"Haochen Zhao, Jianwei Niu, Hui Meng, Yong Wang, Qingfeng Li, and Ziniu Yu. 2022. Focal U-Net: A focal self-attention based U-Net for breast lesion segmentation in ultrasound images. In EMBC. IEEE, 1506\u20131511."},{"key":"e_1_3_1_58_2","unstructured":"Xinlei Yu Zhangquan Chen Yudong Zhang Shilin Lu Ruolin Shen Jiangning Zhang Xiaobin Hu Yanwei Fu and Shuicheng Yan. 2025. Visual document understanding and question answering: A multi-agent collaboration framework with test-time scaling. arXiv:2508.03404. Retrieved from https:\/\/arxiv.org\/abs\/2508.03404"},{"key":"e_1_3_1_59_2","first-page":"2294","volume-title":"CVPR","author":"Lu Shilin","year":"2023","unstructured":"Shilin Lu, Yanzhu Liu, and Adams Wai-Kin Kong. 2023. TF-ICON: Diffusion-based training-free cross-domain image composition. In CVPR, 2294\u20132305."},{"key":"e_1_3_1_60_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.00615"},{"key":"e_1_3_1_61_2","volume-title":"IEEE TCSVT","author":"Yuan Zhenlong","year":"2025","unstructured":"Zhenlong Yuan, Zhidong Yang, Yujun Cai, Kuangxin Wu, Mufan Liu, Dapeng Zhang, Hao Jiang, Zhaoxin Li, and Zhaoqi Wang. 2025. SED-MVS: Segmentation-driven and edge-aligned deformation multi-view stereo with depth restoration and occlusion constraint. In IEEE TCSVT."},{"key":"e_1_3_1_62_2","doi-asserted-by":"publisher","DOI":"10.1007\/s41095-023-0369-x"},{"key":"e_1_3_1_63_2","doi-asserted-by":"publisher","DOI":"10.1109\/TGRS.2024.3439401"},{"key":"e_1_3_1_64_2","unstructured":"Yujun Wang Jinhe Bi Yunpu Ma and Soeren Pirk. 2025. ASCD: Attention-steerable contrastive decoding for reducing hallucination in MLLM. arXiv:2506.14766. Retrieved from https:\/\/arxiv.org\/abs\/2506.14766"},{"key":"e_1_3_1_65_2","unstructured":"Jinhe Bi Yifan Wang Danqi Yan Wenke Aniri Zengjie Huang Xiaowen Jin Artur Ma Mang Hecker Xun Xiao Ye Hinrich Schuetze et al. 2025. PRISM: Self-pruning intrinsic selection method for training-free multimodal data selection. arXiv:250212119. Retrieved from https:\/\/arxiv.org\/abs\/2502.12119"},{"key":"e_1_3_1_66_2","unstructured":"Jinhe Bi Danqi Yan Yifan Wang Wenke Huang Haokun Chen Guancheng Wan Mang Ye Xun Xiao Hinrich Schuetze Volker Tresp et al. 2025. CoT-kinetics: A theoretical modeling assessing LRM reasoning process. arXiv:2505.13408. Retrieved from https:\/\/arxiv.org\/abs\/2505.13408"},{"key":"e_1_3_1_67_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2025.acl-long.739"},{"key":"e_1_3_1_68_2","first-page":"3593","volume-title":"CVPR","author":"Hosseinzadeh Mehrdad","year":"2020","unstructured":"Mehrdad Hosseinzadeh and Yang Wang. 2020. Composed query image retrieval using locally bounded features. In CVPR. IEEE, 3593\u20133602."},{"key":"e_1_3_1_69_2","doi-asserted-by":"publisher","DOI":"10.1145\/3394171.3413917"},{"key":"e_1_3_1_70_2","doi-asserted-by":"publisher","DOI":"10.1145\/3474085.3475483"},{"key":"e_1_3_1_71_2","first-page":"1","volume-title":"ICLR","author":"Delmas Ginger","year":"2022","unstructured":"Ginger Delmas, Rafael Sampaio de Rezende, Gabriela Csurka, and Diane Larlus. 2022. ARTEMIS: Attention-based retrieval with text-explicit matching and implicit similarity. In ICLR. OpenReview.net, 1\u201312."},{"key":"e_1_3_1_72_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2022.3204213"},{"key":"e_1_3_1_73_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v40i8.37608"},{"key":"e_1_3_1_74_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v40i25.39181"},{"key":"e_1_3_1_75_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP49660.2025.10888153"},{"key":"e_1_3_1_76_2","doi-asserted-by":"publisher","DOI":"10.1145\/3617597"},{"key":"e_1_3_1_77_2","article-title":"ENCODER: Entity mining and modification relation binding for composed image retrieval","author":"Li Zixu","year":"2025","unstructured":"Zixu Li, Zhiwei Chen, Haokun Wen, Zhiheng Fu, Yupeng Hu, and Weili Guan. 2025. ENCODER: Entity mining and modification relation binding for composed image retrieval. In AAAI.","journal-title":"AAAI"},{"key":"e_1_3_1_78_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP49660.2025.10890642"},{"key":"e_1_3_1_79_2","unstructured":"Zixu Li Zhiheng Fu Yupeng Hu Zhiwei Chen Haokun Wen and Liqiang Nie. 2025. FineCIR: Explicit parsing of fine-grained modification semantics for composed image retrieval. arXiv:2503.21309. Retrieved from https:\/\/arxiv.org\/abs\/2503.21309"},{"key":"e_1_3_1_80_2","doi-asserted-by":"publisher","DOI":"10.1145\/3746027.3755366"},{"key":"e_1_3_1_81_2","doi-asserted-by":"publisher","DOI":"10.1145\/3477495.3532047"},{"key":"e_1_3_1_82_2","first-page":"8748","volume-title":"ICML","author":"Radford Alec","year":"2021","unstructured":"Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. In ICML. PMLR, 8748\u20138763."},{"key":"e_1_3_1_83_2","first-page":"1","volume-title":"ACM Transactions on Multimedia Computing, Communications and Applications","volume":"19","author":"Zhu Hongguang","year":"2023","unstructured":"Hongguang Zhu, Yunchao Wei, Yao Zhao, Chunjie Zhang, and Shujuan Huang. 2023. AMC: Adaptive multi-expert collaborative network for text-guided image retrieval. ACM Transactions on Multimedia Computing, Communications and Applications 19 (2023), 1\u201322."},{"key":"e_1_3_1_84_2","doi-asserted-by":"publisher","DOI":"10.24963\/ijcai.2023\/68"},{"key":"e_1_3_1_85_2","first-page":"886","volume-title":"ACM MM","author":"Zhao Haochen","year":"2024","unstructured":"Haochen Zhao, Hui Meng, Deqian Yang, Xiaozheng Xie, Xiaoze Wu, Qingfeng Li, and Jianwei Niu. 2024. GuidedNet: Semi-supervised multi-organ segmentation via labeled data guide unlabeled data. In ACM MM, 886\u2013895."},{"key":"e_1_3_1_86_2","first-page":"1566","volume-title":"ICASSP","author":"Chen Yinda","year":"2024","unstructured":"Yinda Chen, Wei Huang, Xiaoyu Liu, Shiyu Deng, Qi Chen, and Zhiwei Xiong. 2024. Learning multiscale consistency for self-supervised electron microscopy instance segmentation. In ICASSP. IEEE, 1566\u20131570."},{"key":"e_1_3_1_87_2","first-page":"21832","volume-title":"ICCV","author":"Zhao Haochen","year":"2025","unstructured":"Haochen Zhao, Jianwei Niu, Xuefeng Liu, Xiaozheng Xie, Li Kuang, Haotian Yang, Bin Dai, Hui Meng, and Yong Wang. 2025. Keep your friends close, and your enemies farther: Distance-aware voxel-wise contrastive learning for semi-supervised multi-organ segmentation. In ICCV, 21832\u201321842."},{"key":"e_1_3_1_88_2","first-page":"124","volume-title":"MICCAI","author":"Chen Yinda","year":"2024","unstructured":"Yinda Chen, Che Liu, Xiaoyu Liu, Rossella Arcucci, and Zhiwei Xiong. 2024. BIMCV-R: A landmark dataset for 3D CT text-image retrieval. In MICCAI. Springer Nature, 124\u2013134."},{"key":"e_1_3_1_89_2","doi-asserted-by":"publisher","DOI":"10.1145\/3746027.3755445"},{"key":"e_1_3_1_90_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v40i28.39507"},{"key":"e_1_3_1_91_2","doi-asserted-by":"publisher","DOI":"10.1145\/3581783.3611817"},{"key":"e_1_3_1_92_2","volume-title":"IEEE TPAMI","author":"Wen Haokun","year":"2023","unstructured":"Haokun Wen, Xuemeng Song, Jianhua Yin, Jianlong Wu, Weili Guan, and Liqiang Nie. 2023. Self-training boosted multi-factor matching network for composed image retrieval. In IEEE TPAMI."},{"key":"e_1_3_1_93_2","first-page":"1228","volume-title":"AAAI","volume":"38","author":"Chen Yanzhe","year":"2024","unstructured":"Yanzhe Chen, Huasong Zhong, Xiangteng He, Yuxin Peng, Jiahuan Zhou, and Lele Cheng. 2024. FashionERN: Enhance-and-refine network for composed fashion image retrieval. In AAAI, Vol. 38, 1228\u20131236."},{"key":"e_1_3_1_94_2","first-page":"15028","volume-title":"CVPR","author":"Han Yunpeng","year":"2023","unstructured":"Yunpeng Han, Lisai Zhang, Qingcai Chen, Zhijian Chen, Zhonghua Li, Jianxin Yang, and Zhao Cao. 2023. FashionSAP: Symbols and attributes prompt for fine-grained fashion vision-language pre-training. In CVPR, 15028\u201315038."},{"key":"e_1_3_1_95_2","first-page":"14105","volume-title":"CVPR","author":"Goenka Sonam","year":"2022","unstructured":"Sonam Goenka, Zhaoheng Zheng, Ayush Jaiswal, Rakesh Chada, Yue Wu, Varsha Hedau, and Pradeep Natarajan. 2022. FashionVLP: Vision language transformer for fashion retrieval with feedback. In CVPR, 14105\u201314115."},{"key":"e_1_3_1_96_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.01115"},{"key":"e_1_3_1_97_2","first-page":"676","volume-title":"NeurIPS","author":"Guo Xiaoxiao","year":"2018","unstructured":"Xiaoxiao Guo, Hui Wu, Yu Cheng, Steven Rennie, Gerald Tesauro, and Rog\u00e9rio Schmidt Feris. 2018. Dialog-based interactive image retrieval. In NeurIPS. MIT Press, 676\u2013686."},{"key":"e_1_3_1_98_2","first-page":"15338","volume-title":"ICCV","author":"Baldrati Alberto","year":"2023","unstructured":"Alberto Baldrati, Lorenzo Agnolucci, Marco Bertini, and Alberto Del Bimbo. 2023. Zero-shot composed image retrieval with textual inversion. In ICCV, 15338\u201315347."},{"key":"e_1_3_1_99_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00213"},{"key":"e_1_3_1_100_2","doi-asserted-by":"crossref","unstructured":"Max Bain Arsha Nagrani G\u00fcl Varol and Andrew Zisserman. 2021. Frozen in time: A joint video and image encoder for end-to-end retrieval. arXiv:2104.00650. Retrieved from https:\/\/arxiv.org\/abs\/2104.00650","DOI":"10.1109\/ICCV48922.2021.00175"},{"key":"e_1_3_1_101_2","doi-asserted-by":"publisher","DOI":"10.1109\/WACV57701.2024.00565"},{"key":"e_1_3_1_102_2","first-page":"2991","volume-title":"AAAI","volume":"38","author":"Levy Matan","year":"2024","unstructured":"Matan Levy, Rami Ben-Ari, Nir Darshan, and Dani Lischinski. 2024. Data roaming and quality assessment for composed image retrieval. In AAAI, Vol. 38, 2991\u20132999."},{"key":"e_1_3_1_103_2","first-page":"6576","volume-title":"AAAI","volume":"38","author":"Yang Xingyu","year":"2024","unstructured":"Xingyu Yang, Daqing Liu, Heng Zhang, Yong Luo, Chaoyue Wang, and Jing Zhang. 2024. Decomposing semantic shifts for composed image retrieval. In AAAI, Vol. 38, 6576\u20136584."},{"key":"e_1_3_1_104_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_1_105_2","volume-title":"ICLR","author":"Yue W. U.","year":"2025","unstructured":"W. U. Yue, Zhaobo Qi, Yiling Wu, Junshu Sun, Yaowei Wang, and Shuhui Wang. 2025. Learning fine-grained representations through textual token disentanglement in composed video retrieval. In ICLR."},{"key":"e_1_3_1_106_2","first-page":"6439","volume-title":"CVPR","author":"Vo Nam","year":"2019","unstructured":"Nam Vo, Lu Jiang, Chen Sun, Kevin Murphy, Li-Jia Li, Li Fei-Fei, and James Hays. 2019. Composing text and image for image retrieval\u2014An empirical odyssey. In CVPR. IEEE, 6439\u20136448."},{"key":"e_1_3_1_107_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00307"},{"key":"e_1_3_1_108_2","first-page":"1369","volume-title":"ACM SIGIR","author":"Wen Haokun","year":"2021","unstructured":"Haokun Wen, Xuemeng Song, Xin Yang, Yibing Zhan, and Liqiang Nie. 2021. Comprehensive linguistic-visual composition network for image retrieval. In ACM SIGIR. ACM, 1369\u20131378."},{"key":"e_1_3_1_109_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2023.3299791"},{"key":"e_1_3_1_110_2","volume-title":"ICLR","author":"Chen Yiyang","year":"2024","unstructured":"Yiyang Chen, Zhedong Zheng, Wei Ji, Leigang Qu, and Tat-Seng Chua. 2024. Composed image retrieval with text feedback via multi-grained uncertainty regularization. In ICLR."},{"key":"e_1_3_1_111_2","doi-asserted-by":"publisher","DOI":"10.1145\/3699715"},{"key":"e_1_3_1_112_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.02080"},{"key":"e_1_3_1_113_2","doi-asserted-by":"publisher","DOI":"10.1145\/3664647.3681493"},{"key":"e_1_3_1_114_2","volume-title":"TMLR","author":"Gu Geonmo","year":"2024","unstructured":"Geonmo Gu, Sanghyuk Chun, Wonjae Kim, HeeJae Jun, Yoohoon Kang, and Sangdoo Yun. 2024. CompoDiff: Versatile composed image retrieval with latent diffusion. In TMLR."},{"key":"e_1_3_1_115_2","unstructured":"Zheyuan Liu Weixuan Sun Damien Teney and Stephen Gould. 2024. Candidate set re-ranking for composed image retrieval with dual multi-modal encoder. arXiv:2305.16304. Retrieved from https:\/\/arxiv.org\/abs\/2305.16304"},{"key":"e_1_3_1_116_2","volume-title":"IEEE TMM","author":"Xu Yahui","year":"2023","unstructured":"Yahui Xu, Yi Bin, Jiwei Wei, Yang Yang, Guoqing Wang, and Heng Tao Shen. 2023. Multi-modal transformer with global-local alignment for composed query image retrieval. In IEEE TMM."}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3796712","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,24]],"date-time":"2026-06-24T14:44:46Z","timestamp":1782312286000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3796712"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6,24]]},"references-count":115,"journal-issue":{"issue":"7","published-print":{"date-parts":[[2026,7,31]]}},"alternative-id":["10.1145\/3796712"],"URL":"https:\/\/doi.org\/10.1145\/3796712","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,6,24]]},"assertion":[{"value":"2025-03-31","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-01-04","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-06-24","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}