{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,16]],"date-time":"2026-06-16T15:10:36Z","timestamp":1781622636803,"version":"3.54.5"},"reference-count":73,"publisher":"Association for Computing Machinery (ACM)","issue":"6","license":[{"start":{"date-parts":[[2023,5,30]],"date-time":"2023-05-30T00:00:00Z","timestamp":1685404800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100012166","name":"National Key Research and Development of China","doi-asserted-by":"crossref","award":["2018AAA0102100"],"award-info":[{"award-number":["2018AAA0102100"]}],"id":[{"id":"10.13039\/501100012166","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62120106009 and 62072026"],"award-info":[{"award-number":["62120106009 and 62072026"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Beijing Natural Science Foundation","award":["JQ20022"],"award-info":[{"award-number":["JQ20022"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2023,11,30]]},"abstract":"<jats:p>\n            Text-guided image retrieval integrates reference image and text feedback as a multimodal query to search the image corresponding to user intention. Recent approaches employ multi-level matching, multiple accesses, or multiple subnetworks for better performance regardless of the heavy burden of storage and computation in the deployment. Additionally, these models not only rely on expert knowledge to handcraft image-text composing modules but also do inference by the static computational graph. It limits the representation capability and generalization ability of networks in the face of challenges from complex and varied combinations of reference image and text feedback. To break the shackles of the static network concept, we introduce the dynamic router mechanism to achieve data-dependent expert activation and flexible collaboration of multiple experts to explore more implicit multimodal fusion patterns. Specifically, we construct AMC, our\n            <jats:italic>A<\/jats:italic>\n            daptive\n            <jats:italic>M<\/jats:italic>\n            ulti-expert\n            <jats:italic>C<\/jats:italic>\n            ollaborative network, by using the proposed router to activate the different experts with different levels of image-text interaction. Since routers can dynamically adjust the activation of experts for the current samples, AMC can achieve the adaptive fusion mode for the different reference image and text combinations and generate dynamic computational graphs according to varied multimodal queries. Extensive experiments on two benchmark datasets demonstrate that due to benefits from the image-text composing representation produced by an adaptive multi-expert collaboration mechanism, AMC has better retrieval performance and zero-shot generalization ability than the state-of-the-art method while keeping the lightweight model and fast retrieval speed. Moreover, we analyze the visualization of path activation, attention map, and retrieval results to further understand the routing decisions and semantic localization ability of AMC. The codes and pretrained models are available at\n            <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" ext-link-type=\"url\" xlink:href=\"https:\/\/github.com\/KevinLight831\/AMC\">https:\/\/github.com\/KevinLight831\/AMC<\/jats:ext-link>\n            .\n          <\/jats:p>","DOI":"10.1145\/3584703","type":"journal-article","created":{"date-parts":[[2023,2,20]],"date-time":"2023-02-20T11:48:12Z","timestamp":1676893692000},"page":"1-22","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":11,"title":["AMC: Adaptive Multi-expert Collaborative Network for Text-guided Image Retrieval"],"prefix":"10.1145","volume":"19","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-1356-5153","authenticated-orcid":false,"given":"Hongguang","family":"Zhu","sequence":"first","affiliation":[{"name":"Beijing Jiaotong University, Beijing Key Laboratory of Advanced Information Science and Network Technology, and Peng Cheng Laboratory"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2812-8781","authenticated-orcid":false,"given":"Yunchao","family":"Wei","sequence":"additional","affiliation":[{"name":"Beijing Jiaotong University, Beijing Key Laboratory of Advanced Information Science and Network Technology, and Peng Cheng Laboratory"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8581-9554","authenticated-orcid":false,"given":"Yao","family":"Zhao","sequence":"additional","affiliation":[{"name":"Beijing Jiaotong University, Beijing Key Laboratory of Advanced Information Science and Network Technology, and Peng Cheng Laboratory"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1161-8995","authenticated-orcid":false,"given":"Chunjie","family":"Zhang","sequence":"additional","affiliation":[{"name":"Beijing Jiaotong University and Beijing Key Laboratory of Advanced Information Science and Network Technology"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6045-8652","authenticated-orcid":false,"given":"Shujuan","family":"Huang","sequence":"additional","affiliation":[{"name":"Beijing Jiaotong University and Beijing Key Laboratory of Advanced Information Science and Network Technology"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2023,5,30]]},"reference":[{"key":"e_1_3_1_2_2","first-page":"7708","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Ak Kenan E.","year":"2018","unstructured":"Kenan E. Ak, Ashraf A. Kassim, Joo Hwee Lim, and Jo Yew Tham. 2018. Learning attribute representations with localization for flexible fashion search. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 7708\u20137717."},{"key":"e_1_3_1_3_2","article-title":"Layer normalization","author":"Ba Jimmy Lei","year":"2016","unstructured":"Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton. 2016. Layer normalization. arXiv preprint arXiv:1607.06450 (2016).","journal-title":"arXiv preprint arXiv:1607.06450"},{"key":"e_1_3_1_4_2","first-page":"663","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Berg Tamara L.","year":"2010","unstructured":"Tamara L. Berg, Alexander C. Berg, and Jonathan Shih. 2010. Automatic attribute discovery and characterization from noisy web data. In Proceedings of the European Conference on Computer Vision. 663\u2013676."},{"key":"e_1_3_1_5_2","first-page":"3588","volume-title":"Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision","author":"Cai Shaofeng","year":"2021","unstructured":"Shaofeng Cai, Yao Shu, and Wei Wang. 2021. Dynamic routing networks. In Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision. 3588\u20133597."},{"key":"e_1_3_1_6_2","first-page":"3001","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Chen Yanbei","year":"2020","unstructured":"Yanbei Chen, S. Gong, and L. Bazzani. 2020. Image search with text feedback by visiolinguistic attention learning. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 3001\u20133011."},{"key":"e_1_3_1_7_2","first-page":"23","article-title":"Cross-modal graph matching network for image-text retrieval","volume":"18","author":"Cheng Yuhao","year":"2022","unstructured":"Yuhao Cheng, Xiaoguang Zhu, Jiuchao Qian, Fei Wen, and Peilin Liu. 2022. Cross-modal graph matching network for image-text retrieval. ACM Transactions on Multimedia Computing, Communications and Applications 18, 4 (March 2022), Article 95, 23 pages.","journal-title":"ACM Transactions on Multimedia Computing, Communications and Applications"},{"key":"e_1_3_1_8_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Delmas Ginger","year":"2021","unstructured":"Ginger Delmas, Rafael S. Rezende, Gabriela Csurka, and Diane Larlus. 2021. ARTEMIS: Attention-based retrieval with text-explicit matching and implicit similarity. In Proceedings of the International Conference on Learning Representations."},{"key":"e_1_3_1_9_2","doi-asserted-by":"crossref","first-page":"248","DOI":"10.1109\/CVPR.2009.5206848","volume-title":"Proceedings of the 2009 IEEE Conference on Computer Vision and Pattern Recognition","author":"Deng Jia","year":"2009","unstructured":"Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. 2009. ImageNet: A large-scale hierarchical image database. In Proceedings of the 2009 IEEE Conference on Computer Vision and Pattern Recognition. IEEE, Los Alamitos, CA, 248\u2013255."},{"key":"e_1_3_1_10_2","article-title":"Bert: Pre-training of deep bidirectional transformers for language understanding","author":"Devlin Jacob","year":"2018","unstructured":"Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805 (2018).","journal-title":"arXiv preprint arXiv:1810.04805"},{"key":"e_1_3_1_11_2","article-title":"Modality-agnostic attention fusion for visual search with text feedback","author":"Dodds Eric","year":"2020","unstructured":"Eric Dodds, Jack Culpepper, Simao Herdade, Yang Zhang, and Kofi Boakye. 2020. Modality-agnostic attention fusion for visual search with text feedback. arXiv preprint arXiv:2007.00145 (2020).","journal-title":"arXiv preprint arXiv:2007.00145"},{"key":"e_1_3_1_12_2","volume-title":"Proceedings of the British Machine Vision Conference.","author":"Faghri Fartash","year":"2017","unstructured":"Fartash Faghri, David J. Fleet, Jamie Ryan Kiros, and Sanja Fidler. 2017. VSE++: Improving visual-semantic embeddings with hard negatives. In Proceedings of the British Machine Vision Conference."},{"issue":"11","key":"e_1_3_1_13_2","doi-asserted-by":"crossref","first-page":"665","DOI":"10.1038\/s42256-020-00257-z","article-title":"Shortcut learning in deep neural networks","volume":"2","author":"Geirhos Robert","year":"2020","unstructured":"Robert Geirhos, J\u00f6rn-Henrik Jacobsen, Claudio Michaelis, Richard Zemel, Wieland Brendel, Matthias Bethge, and Felix A. Wichmann. 2020. Shortcut learning in deep neural networks. Nature Machine Intelligence 2, 11 (2020), 665\u2013673.","journal-title":"Nature Machine Intelligence"},{"key":"e_1_3_1_14_2","doi-asserted-by":"crossref","first-page":"2436","DOI":"10.1109\/TIP.2020.3046921","article-title":"Gated path selection network for semantic segmentation","volume":"30","author":"Geng Qichuan","year":"2021","unstructured":"Qichuan Geng, Hong Zhang, Xiaojuan Qi, Gao Huang, Ruigang Yang, and Zhong Zhou. 2021. Gated path selection network for semantic segmentation. IEEE Transactions on Image Processing 30 (2021), 2436\u20132449.","journal-title":"IEEE Transactions on Image Processing"},{"key":"e_1_3_1_15_2","first-page":"1171","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Ghosh Arnab","year":"2019","unstructured":"Arnab Ghosh, Richard Zhang, Puneet K. Dokania, Oliver Wang, Alexei A. Efros, Philip H. S. Torr, and Eli Shechtman. 2019. Interactive sketch & fill: Multiclass sketch-to-image translation. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 1171\u20131180."},{"key":"e_1_3_1_16_2","first-page":"241","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Gordo Albert","year":"2016","unstructured":"Albert Gordo, Jon Almaz\u00e1n, Jerome Revaud, and Diane Larlus. 2016. Deep image retrieval: Learning global representations for image search. In Proceedings of the European Conference on Computer Vision. 241\u2013257."},{"key":"e_1_3_1_17_2","first-page":"4600","volume-title":"Proceedings of the 29th ACM International Conference on Multimedia","author":"Gu Chunbin","year":"2021","unstructured":"Chunbin Gu, Jiajun Bu, Zhen Zhang, Zhi Yu, Dongfang Ma, and Wei Wang. 2021. Image search with text feedback by deep hierarchical attention mutual information maximization. In Proceedings of the 29th ACM International Conference on Multimedia. 4600\u20134609."},{"key":"e_1_3_1_18_2","article-title":"Dialog-based interactive image retrieval","author":"Guo Xiaoxiao","year":"2018","unstructured":"Xiaoxiao Guo, Hui Wu, Yu Cheng, Steven Rennie, Gerald Tesauro, and Rogerio Feris. 2018. Dialog-based interactive image retrieval. In Proceedings of the 32nd Conference on Neural Information Processing Systems (NeurIPS\u201918). 1\u201311.","journal-title":"Proceedings of the 32nd Conference on Neural Information Processing Systems (NeurIPS\u201918)."},{"key":"e_1_3_1_19_2","article-title":"The Fashion IQ dataset: Retrieving images by combining side information and relative natural language feedback","author":"Guo Xiaoxiao","year":"2019","unstructured":"Xiaoxiao Guo, Hui Wu, Yupeng Gao, Steven Rennie, and Rogerio Feris. 2019. The Fashion IQ dataset: Retrieving images by combining side information and relative natural language feedback. arXiv preprint arXiv:1905.12794 (2019).","journal-title":"arXiv preprint arXiv:1905.12794"},{"key":"e_1_3_1_20_2","first-page":"1463","volume-title":"Proceedings of the IEEE International Conference on Computer Vision","author":"Han Xintong","year":"2017","unstructured":"Xintong Han, Zuxuan Wu, Phoenix X. Huang, Xiao Zhang, Menglong Zhu, Yuan Li, Yang Zhao, and Larry S. Davis. 2017. Automatic spatially-aware fashion concept discovery. In Proceedings of the IEEE International Conference on Computer Vision. 1463\u20131471."},{"key":"e_1_3_1_21_2","article-title":"Dynamic neural networks: A survey","author":"Han Yizeng","year":"2022","unstructured":"Yizeng Han, Gao Huang, Shiji Song, Le Yang, Honghui Wang, and Yulin Wang. 2022. Dynamic neural networks: A survey. IEEE Transactions on Pattern Analysis and Machine Intelligence 44 (2022), 7436\u20137456.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"e_1_3_1_22_2","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201916)","author":"He Kaiming","year":"2016","unstructured":"Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201916)."},{"key":"e_1_3_1_23_2","first-page":"241","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Herrmann Charles","year":"2020","unstructured":"Charles Herrmann, Richard Strong Bowen, and Ramin Zabih. 2020. Channel selection using Gumbel Softmax. In Proceedings of the European Conference on Computer Vision. 241\u2013257."},{"issue":"8","key":"e_1_3_1_24_2","doi-asserted-by":"crossref","first-page":"1735","DOI":"10.1162\/neco.1997.9.8.1735","article-title":"Long short-term memory","volume":"9","author":"Hochreiter Sepp","year":"1997","unstructured":"Sepp Hochreiter and J\u00fcrgen Schmidhuber. 1997. Long short-term memory. Neural Computation 9, 8 (1997), 1735\u20131780.","journal-title":"Neural Computation"},{"key":"e_1_3_1_25_2","first-page":"1501","volume-title":"Proceedings of the IEEE International Conference on Computer Vision","author":"Huang Xun","year":"2017","unstructured":"Xun Huang and Serge Belongie. 2017. Arbitrary style transfer in real-time with adaptive instance normalization. In Proceedings of the IEEE International Conference on Computer Vision. 1501\u20131510."},{"key":"e_1_3_1_26_2","first-page":"448","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Ioffe Sergey","year":"2015","unstructured":"Sergey Ioffe and Christian Szegedy. 2015. Batch normalization: Accelerating deep network training by reducing internal covariate shift. In Proceedings of the International Conference on Machine Learning. 448\u2013456."},{"key":"e_1_3_1_27_2","first-page":"4021","volume-title":"Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision","author":"Jandial Surgan","year":"2022","unstructured":"Surgan Jandial, Pinkesh Badjatiya, Pranit Chawla, Ayush Chopra, Mausoom Sarkar, and Balaji Krishnamurthy. 2022. SAC: Semantic attention composition for text-conditioned image retrieval. In Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision. 4021\u20134030."},{"key":"e_1_3_1_28_2","article-title":"Categorical reparameterization with Gumbel-Softmax","author":"Jang Eric","year":"2016","unstructured":"Eric Jang, Shixiang Gu, and Ben Poole. 2016. Categorical reparameterization with Gumbel-Softmax. arXiv preprint arXiv:1611.01144 (2016).","journal-title":"arXiv preprint arXiv:1611.01144"},{"key":"e_1_3_1_29_2","first-page":"1","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence (AAAI\u201921)","author":"Kim Jongseok","year":"2021","unstructured":"Jongseok Kim, Youngjae Yu, Hoeseong Kim, and Gunhee Kim. 2021. Dual compositional learning in interactive image retrieval. In Proceedings of the AAAI Conference on Artificial Intelligence (AAAI\u201921). 1\u20139."},{"key":"e_1_3_1_30_2","article-title":"Adam: A method for stochastic optimization","author":"Kingma Diederik P.","year":"2014","unstructured":"Diederik P. Kingma and Jimmy Ba. 2014. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980 (2014).","journal-title":"arXiv preprint arXiv:1412.6980"},{"key":"e_1_3_1_31_2","volume-title":"Proceedings of the European Conference on Computer Vision.","author":"Lee Kuang-Huei","year":"2018","unstructured":"Kuang-Huei Lee, Xi Chen, Gang Hua, Houdong Hu, and Xiaodong He. 2018. Stacked cross attention for image-text matching. In Proceedings of the European Conference on Computer Vision."},{"key":"e_1_3_1_32_2","first-page":"802","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Lee Seungmin","year":"2021","unstructured":"Seungmin Lee, Dongwan Kim, and Bohyung Han. 2021. Cosmo: Content-style modulation for image retrieval with text feedback. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 802\u2013812."},{"key":"e_1_3_1_33_2","first-page":"121","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Li Xiujun","year":"2020","unstructured":"Xiujun Li, Xi Yin, Chunyuan Li, Pengchuan Zhang, Xiaowei Hu, Lei Zhang, Lijuan Wang, et\u00a0al. 2020. Oscar: Object-semantics aligned pre-training for vision-language tasks. In Proceedings of the European Conference on Computer Vision. 121\u2013137."},{"key":"e_1_3_1_34_2","article-title":"Unified loss of pair similarity optimization for vision-language retrieval","author":"Li Zheng","year":"2022","unstructured":"Zheng Li, Caili Guo, Xin Wang, Zerun Feng, Jenq-Neng Hwang, and Zhongtian Du. 2022. Unified loss of pair similarity optimization for vision-language retrieval. arXiv preprint arXiv:2209.13869 (2022).","journal-title":"arXiv preprint arXiv:2209.13869"},{"key":"e_1_3_1_35_2","first-page":"19","article-title":"Modality-invariant image-text embedding for image-sentence matching","volume":"15","author":"Liu Ruoyu","year":"2019","unstructured":"Ruoyu Liu, Yao Zhao, Shikui Wei, Liang Zheng, and Yi Yang. 2019. Modality-invariant image-text embedding for image-sentence matching. ACM Transactions on Multimedia Computing, Communications and Applications 15, 1 (Feb. 2019), Article 27, 19 pages.","journal-title":"ACM Transactions on Multimedia Computing, Communications and Applications"},{"key":"e_1_3_1_36_2","first-page":"2125","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Liu Zheyuan","year":"2021","unstructured":"Zheyuan Liu, Cristian Rodriguez-Opazo, Damien Teney, and Stephen Gould. 2021. Image retrieval on real-life images with pre-trained vision-and-language models. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 2125\u20132134."},{"key":"e_1_3_1_37_2","first-page":"2503","volume-title":"Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision","author":"Ma Jin","year":"2020","unstructured":"Jin Ma, Shanmin Pang, Bo Yang, Jihua Zhu, and Yaochen Li. 2020. Spatial-content image search in complex scenes. In Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision. 2503\u20132511."},{"key":"e_1_3_1_38_2","first-page":"1930","volume-title":"Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining","author":"Ma Jiaqi","year":"2018","unstructured":"Jiaqi Ma, Zhe Zhao, Xinyang Yi, Jilin Chen, Lichan Hong, and Ed H. Chi. 2018. Modeling task relationships in multi-task learning with multi-gate mixture-of-experts. In Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. 1930\u20131939."},{"key":"e_1_3_1_39_2","first-page":"1930","volume-title":"Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining","author":"Ma Jiaqi","year":"2018","unstructured":"Jiaqi Ma, Zhe Zhao, Xinyang Yi, Jilin Chen, Lichan Hong, and Ed H. Chi. 2018. Modeling task relationships in multi-task learning with multi-gate mixture-of-experts. In Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. 1930\u20131939."},{"key":"e_1_3_1_40_2","first-page":"11741","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","volume":"34","author":"Ma Zhe","year":"2020","unstructured":"Zhe Ma, Jianfeng Dong, Zhongzi Long, Yao Zhang, Yuan He, Hui Xue, and Shouling Ji. 2020. Fine-grained fashion similarity learning by attribute-specific embedding network. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 34. 11741\u201311748."},{"key":"e_1_3_1_41_2","first-page":"4718","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Mai Long","year":"2017","unstructured":"Long Mai, Hailin Jin, Zhe Lin, Chen Fang, Jonathan Brandt, and Feng Liu. 2017. Spatial-semantic image search by visual feature synthesis. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 4718\u20134727."},{"key":"e_1_3_1_42_2","first-page":"23","article-title":"Fine-grained visual textual alignment for cross-modal retrieval using transformer encoders","volume":"17","author":"Messina Nicola","year":"2021","unstructured":"Nicola Messina, Giuseppe Amato, Andrea Esuli, Fabrizio Falchi, Claudio Gennaro, and St\u00e9phane Marchand-Maillet. 2021. Fine-grained visual textual alignment for cross-modal retrieval using transformer encoders. ACM Transactions on Multimedia Computing, Communications and Applications 17, 4 (Nov. 2021), Article 128, 23 pages.","journal-title":"ACM Transactions on Multimedia Computing, Communications and Applications"},{"key":"e_1_3_1_43_2","first-page":"8080","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Mullapudi Ravi Teja","year":"2018","unstructured":"Ravi Teja Mullapudi, William R. Mark, Noam Shazeer, and Kayvon Fatahalian. 2018. HydraNets: Specialized dynamic architectures for efficient inference. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 8080\u20138089."},{"key":"e_1_3_1_44_2","first-page":"59","volume-title":"Proceedings of the 20th ACM International Conference on Multimedia (MM\u201912)","author":"Nie Liqiang","year":"2012","unstructured":"Liqiang Nie, Shuicheng Yan, Meng Wang, Richang Hong, and Tat-Seng Chua. 2012. Harvesting visual concepts for image search with complex queries. In Proceedings of the 20th ACM International Conference on Multimedia (MM\u201912). 59\u201368."},{"key":"e_1_3_1_45_2","first-page":"3456","volume-title":"Proceedings of the IEEE International Conference on Computer Vision","author":"Noh Hyeonwoo","year":"2017","unstructured":"Hyeonwoo Noh, Andre Araujo, Jack Sim, Tobias Weyand, and Bohyung Han. 2017. Large-scale image retrieval with attentive deep local features. In Proceedings of the IEEE International Conference on Computer Vision. 3456\u20133465."},{"key":"e_1_3_1_46_2","first-page":"2337","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Park Taesung","year":"2019","unstructured":"Taesung Park, Ming-Yu Liu, Ting-Chun Wang, and Jun-Yan Zhu. 2019. Semantic image synthesis with spatially-adaptive normalization. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 2337\u20132346."},{"key":"e_1_3_1_47_2","volume-title":"Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP\u201914)","author":"Pennington Jeffrey","year":"2014","unstructured":"Jeffrey Pennington, Richard Socher, and Christopher D. Manning. 2014. GloVe: Global vectors for word representation. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP\u201914)."},{"key":"e_1_3_1_48_2","volume-title":"Proceedings of the 32nd AAAI Conference on Artificial Intelligence","author":"Perez Ethan","year":"2018","unstructured":"Ethan Perez, Florian Strub, Harm De Vries, Vincent Dumoulin, and Aaron Courville. 2018. Film: Visual reasoning with a general conditioning layer. In Proceedings of the 32nd AAAI Conference on Artificial Intelligence."},{"key":"e_1_3_1_49_2","first-page":"1104","volume-title":"Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval","author":"Qu Leigang","year":"2021","unstructured":"Leigang Qu, Meng Liu, Jianlong Wu, Zan Gao, and Liqiang Nie. 2021. Dynamic modality interaction modeling for image-text retrieval. In Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval. 1104\u20131113."},{"key":"e_1_3_1_50_2","article-title":"Learning transferable visual models from natural language supervision","author":"Radford Alec","year":"2021","unstructured":"Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, et\u00a0al. 2021. Learning transferable visual models from natural language supervision. arXiv preprint arXiv:2103.00020 (2021).","journal-title":"arXiv preprint arXiv:2103.00020"},{"key":"e_1_3_1_51_2","article-title":"Fashion-Gen: The generative fashion dataset and challenge","author":"Rostamzadeh Negar","year":"2018","unstructured":"Negar Rostamzadeh, Seyedarian Hosseini, Thomas Boquet, Wojciech Stokowiec, Ying Zhang, Christian Jauvin, and Chris Pal. 2018. Fashion-Gen: The generative fashion dataset and challenge. arXiv preprint arXiv:1806.08317 (2018).","journal-title":"arXiv preprint arXiv:1806.08317"},{"issue":"5","key":"e_1_3_1_52_2","doi-asserted-by":"crossref","first-page":"644","DOI":"10.1109\/76.718510","article-title":"Relevance feedback: A power tool for interactive content-based image retrieval","volume":"8","author":"Rui Yong","year":"1998","unstructured":"Yong Rui, Thomas S. Huang, Michael Ortega, and Sharad Mehrotra. 1998. Relevance feedback: A power tool for interactive content-based image retrieval. IEEE Transactions on Circuits and Systems for Video Technology 8, 5 (1998), 644\u2013655.","journal-title":"IEEE Transactions on Circuits and Systems for Video Technology"},{"key":"e_1_3_1_53_2","first-page":"4967","volume-title":"Advances in Neural Information Processing Systems 30","author":"Santoro Adam","year":"2017","unstructured":"Adam Santoro, David Raposo, David G. Barrett, Mateusz Malinowski, Razvan Pascanu, Peter Battaglia, and Timothy Lillicrap. 2017. A simple neural network module for relational reasoning. In Advances in Neural Information Processing Systems 30, I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett (Eds.). Curran Associates, 4967\u20134976."},{"key":"e_1_3_1_54_2","first-page":"618","volume-title":"Proceedings of the IEEE International Conference on Computer Vision","author":"Selvaraju Ramprasaath R.","year":"2017","unstructured":"Ramprasaath R. Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra. 2017. Grad-CAM: Visual explanations from deep networks via gradient-based localization. In Proceedings of the IEEE International Conference on Computer Vision. 618\u2013626."},{"key":"e_1_3_1_55_2","article-title":"Fashion-IQ 2020 challenge 2nd place team\u2019s solution","author":"Shin Minchul","year":"2020","unstructured":"Minchul Shin, Yoonjae Cho, and Seongwuk Hong. 2020. Fashion-IQ 2020 challenge 2nd place team\u2019s solution. arXiv e-printsarXiv:2007.06404 (2020).","journal-title":"arXiv e-prints"},{"key":"e_1_3_1_56_2","article-title":"Representation learning with contrastive predictive coding","author":"Oord Aaron Van den","year":"2018","unstructured":"Aaron Van den Oord, Yazhe Li, and Oriol Vinyals. 2018. Representation learning with contrastive predictive coding. arXiv e-printsarXiv:1807.03748 (2018).","journal-title":"arXiv e-prints"},{"key":"e_1_3_1_57_2","unstructured":"Ashish Vaswani Noam Shazeer Niki Parmar Jakob Uszkoreit Llion Jones Aidan N. Gomez \u0141ukasz Kaiser and Illia Polosukhin. 2017. Attention is all you need. In Proceedings of the 31st Conference on Neural Information Processing Systems (NIPS\u201917) . 1\u201311."},{"key":"e_1_3_1_58_2","first-page":"6439","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Vo Nam","year":"2019","unstructured":"Nam Vo, Lu Jiang, Chen Sun, Kevin Murphy, Li-Jia Li, Li Fei-Fei, and James Hays. 2019. Composing text and image for image retrieval-an empirical odyssey. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 6439\u20136448."},{"key":"e_1_3_1_59_2","first-page":"1369","volume-title":"Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval","author":"Wen Haokun","year":"2021","unstructured":"Haokun Wen, Xuemeng Song, Xin Yang, Yibing Zhan, and Liqiang Nie. 2021. Comprehensive linguistic-visual composition network for image retrieval. In Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval. 1369\u20131378."},{"key":"e_1_3_1_60_2","first-page":"8817","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Wu Zuxuan","year":"2018","unstructured":"Zuxuan Wu, Tushar Nagarajan, Abhishek Kumar, Steven Rennie, Larry S. Davis, Kristen Grauman, and Rogerio Feris. 2018. BlockDrop: Dynamic inference paths in residual networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 8817\u20138826."},{"key":"e_1_3_1_61_2","first-page":"10524","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Xiong Ruibin","year":"2020","unstructured":"Ruibin Xiong, Yunchang Yang, Di He, Kai Zheng, Shuxin Zheng, Chen Xing, Huishuai Zhang, Yanyan Lan, Liwei Wang, and Tieyan Liu. 2020. On layer normalization in the transformer architecture. In Proceedings of the International Conference on Machine Learning. 10524\u201310533."},{"key":"e_1_3_1_62_2","first-page":"17","article-title":"Interactive re-ranking via object entropy-guided question answering for cross-modal image retrieval","volume":"18","author":"Yanagi Rintaro","year":"2022","unstructured":"Rintaro Yanagi, Ren Togo, Takahiro Ogawa, and Miki Haseyama. 2022. Interactive re-ranking via object entropy-guided question answering for cross-modal image retrieval. ACM Transactions on Multimedia Computing, Communications and Applications 18, 3 (March 2022), Article 68, 17 pages.","journal-title":"ACM Transactions on Multimedia Computing, Communications and Applications"},{"key":"e_1_3_1_63_2","doi-asserted-by":"crossref","first-page":"3303","DOI":"10.1145\/3474085.3475483","volume-title":"Proceedings of the 29th ACM International Conference on Multimedia","author":"Yang Yuchen","year":"2021","unstructured":"Yuchen Yang, Min Wang, Wengang Zhou, and Houqiang Li. 2021. Cross-modal joint prediction and alignment for composed query image retrieval. In Proceedings of the 29th ACM International Conference on Multimedia. 3303\u20133311."},{"key":"e_1_3_1_64_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Yao Lewei","year":"2021","unstructured":"Lewei Yao, Runhui Huang, Lu Hou, Guansong Lu, Minzhe Niu, Hang Xu, Xiaodan Liang, Zhenguo Li, Xin Jiang, and Chunjing Xu. 2021. FILIP: Fine-grained interactive language-image pre-training. In Proceedings of the International Conference on Learning Representations."},{"key":"e_1_3_1_65_2","first-page":"300","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV\u201918)","author":"Yelamarthi Sasi Kiran","year":"2018","unstructured":"Sasi Kiran Yelamarthi, Shiva Krishna Reddy, Ashish Mishra, and Anurag Mittal. 2018. A zero-shot framework for sketch based image retrieval. In Proceedings of the European Conference on Computer Vision (ECCV\u201918). 300\u2013317."},{"key":"e_1_3_1_66_2","first-page":"1276","article-title":"Adaptive semi-supervised feature selection for cross-modal retrieval","volume":"21","author":"Yu En","year":"2018","unstructured":"En Yu, Jiande Sun, Jing Li, Xiaojun Chang, Xian-Hua Han, and Alexander G. Hauptmann. 2018. Adaptive semi-supervised feature selection for cross-modal retrieval. IEEE Transactions on Multimedia 21 (2018), 1276\u20131288.","journal-title":"IEEE Transactions on Multimedia"},{"key":"e_1_3_1_67_2","first-page":"799","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Yu Qian","year":"2016","unstructured":"Qian Yu, Feng Liu, Yi-Zhe Song, Tao Xiang, Timothy M. Hospedales, and Chen-Change Loy. 2016. Sketch me that shoe. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 799\u2013807."},{"key":"e_1_3_1_68_2","article-title":"CurlingNet: Compositional learning between images and text for Fashion IQ data","author":"Yu Youngjae","year":"2020","unstructured":"Youngjae Yu, Seunghwan Lee, Yuncheol Choi, and Gunhee Kim. 2020. CurlingNet: Compositional learning between images and text for Fashion IQ data. arXiv e-printsarXiv:2003.12299 (2020).","journal-title":"arXiv e-prints"},{"key":"e_1_3_1_69_2","doi-asserted-by":"crossref","first-page":"1000","DOI":"10.1109\/TIP.2021.3138302","article-title":"Geometry sensitive cross-modal reasoning for composed query based image retrieval","volume":"31","author":"Zhang Feifei","year":"2021","unstructured":"Feifei Zhang, Mingliang Xu, and Changsheng Xu. 2021. Geometry sensitive cross-modal reasoning for composed query based image retrieval. IEEE Transactions on Image Processing 31 (2021), 1000\u20131011.","journal-title":"IEEE Transactions on Image Processing"},{"key":"e_1_3_1_70_2","first-page":"23","article-title":"Tell, imagine, and search: End-to-end learning for composing text and image to image retrieval","volume":"18","author":"Zhang Feifei","year":"2022","unstructured":"Feifei Zhang, Mingliang Xu, and Changsheng Xu. 2022. Tell, imagine, and search: End-to-end learning for composing text and image to image retrieval. ACM Transactions on Multimedia Computing, Communications and Applications 18, 2 (March 2022), Article 59, 23 pages.","journal-title":"ACM Transactions on Multimedia Computing, Communications and Applications"},{"key":"e_1_3_1_71_2","doi-asserted-by":"crossref","first-page":"5353","DOI":"10.1145\/3474085.3475659","volume-title":"Proceedings of the 29th ACM International Conference on Multimedia","author":"Zhang Gangjian","year":"2021","unstructured":"Gangjian Zhang, Shikui Wei, Huaxin Pang, and Yao Zhao. 2021. Heterogeneous feature fusion and cross-modal alignment for composed image retrieval. In Proceedings of the 29th ACM International Conference on Multimedia. 5353\u20135362."},{"key":"e_1_3_1_72_2","first-page":"21","article-title":"Attribute-augmented semantic hierarchy: Towards a unified framework for content-based image retrieval","volume":"11","author":"Zhang Hanwang","year":"2014","unstructured":"Hanwang Zhang, Zheng-Jun Zha, Yang Yang, Shuicheng Yan, Yue Gao, and Tat-Seng Chua. 2014. Attribute-augmented semantic hierarchy: Towards a unified framework for content-based image retrieval. ACM Transactions on Multimedia Computing, Communications and Applications 11, 1s (Oct. 2014), Article 21, 21 pages.","journal-title":"ACM Transactions on Multimedia Computing, Communications and Applications"},{"key":"e_1_3_1_73_2","first-page":"1520","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Zhao Bo","year":"2017","unstructured":"Bo Zhao, Jiashi Feng, Xiao Wu, and Shuicheng Yan. 2017. Memory-augmented attribute manipulation networks for interactive fashion search. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 1520\u20131528."},{"key":"e_1_3_1_74_2","article-title":"Uni-perceiver-MoE: Learning sparse generalist models with conditional MoEs","author":"Zhu Jinguo","year":"2022","unstructured":"Jinguo Zhu, Xizhou Zhu, Wenhai Wang, Xiaohua Wang, Hongsheng Li, Xiaogang Wang, and Jifeng Dai. 2022. Uni-perceiver-MoE: Learning sparse generalist models with conditional MoEs. arXiv preprint arXiv:2206.04674 (2022).","journal-title":"arXiv preprint arXiv:2206.04674"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3584703","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3584703","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T17:51:37Z","timestamp":1750182697000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3584703"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,5,30]]},"references-count":73,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2023,11,30]]}},"alternative-id":["10.1145\/3584703"],"URL":"https:\/\/doi.org\/10.1145\/3584703","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,5,30]]},"assertion":[{"value":"2022-08-17","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-02-12","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-05-30","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}