{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,8,22]],"date-time":"2025-08-22T01:40:48Z","timestamp":1755826848361,"version":"3.44.0"},"publisher-location":"New York, NY, USA","reference-count":42,"publisher":"ACM","license":[{"start":{"date-parts":[[2024,5,30]],"date-time":"2024-05-30T00:00:00Z","timestamp":1717027200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2024,5,30]]},"DOI":"10.1145\/3652583.3657611","type":"proceedings-article","created":{"date-parts":[[2024,6,7]],"date-time":"2024-06-07T06:30:40Z","timestamp":1717741840000},"page":"1104-1109","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["CLIP-ProbCR:CLIP-based Probability embedding Combination Retrieval"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-5517-3633","authenticated-orcid":false,"given":"Mingyong","family":"Li","sequence":"first","affiliation":[{"name":"College of Computer and Information Science, Chongqing Normal University, Chongqing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0003-0376-329X","authenticated-orcid":false,"given":"Zongwei","family":"Zhao","sequence":"additional","affiliation":[{"name":"College of Computer and Information Science, Chongqing Normal University, Chongqing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-8429-9669","authenticated-orcid":false,"given":"Xiaolong","family":"Jiang","sequence":"additional","affiliation":[{"name":"College of Computer and Information Science, Chongqing Normal University, Chongqing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-2273-6851","authenticated-orcid":false,"given":"Zheng","family":"Jiang","sequence":"additional","affiliation":[{"name":"College of Computer and Information Science, Chongqing Normal University, Chongqing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2024,6,7]]},"reference":[{"key":"e_1_3_2_1_1_1","volume-title":"Probabilistic Compositional Embeddings for Multimodal Image Retrieval. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops,CVPR Workshops 2022","author":"Andrei Neculai","year":"2022","unstructured":"Neculai Andrei, Yanbei Chen, and Zeynep Akata. 2022. Probabilistic Compositional Embeddings for Multimodal Image Retrieval. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops,CVPR Workshops 2022, New Orleans, LA, USA, June 19--20, 2022. 4546--4556."},{"key":"e_1_3_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW56347.2022.00543"},{"key":"e_1_3_2_1_3_1","volume-title":"BLOCK: Bilinear Superdiagonal Fusion for Visual Question Answering and Visual Relationship Detection. In The Thirty-Third AAAI Conference on Artificial Intelligence, AAAI2019","author":"di Ben-Younes H\u00e9","year":"2019","unstructured":"H\u00e9 di Ben-Younes, R\u00e9 mi Cad\u00e8 ne, Nicolas Thome, and Matthieu Cord. 2019. BLOCK: Bilinear Superdiagonal Fusion for Visual Question Answering and Visual Relationship Detection. In The Thirty-Third AAAI Conference on Artificial Intelligence, AAAI2019. AAAI Press, 8102--8109."},{"key":"e_1_3_2_1_4_1","volume-title":"DAtRNet: Disentangling Fashion Attribute Embedding for Substitute Item Retrieval. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops, CVPR Workshops 2022","author":"Bhattacharya Gaurab","year":"2022","unstructured":"Gaurab Bhattacharya, Nikhil Kilari, Jayavardhana Gubbi, Bagya Lakshmi V, Arpan Pal, and Balamuralidhar P. 2022. DAtRNet: Disentangling Fashion Attribute Embedding for Substitute Item Retrieval. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops, CVPR Workshops 2022, New Orleans, LA, USA, June 19--20, 2022. IEEE, 2282--2286."},{"key":"e_1_3_2_1_5_1","first-page":"1","article-title":"Products and convolutions of Gaussian probability density functions","volume":"3","author":"Bromiley Paul","year":"2003","unstructured":"Paul Bromiley. 2003. Products and convolutions of Gaussian probability density functions. Tina-Vision Memo, Vol. 3, 4 (2003), 1.","journal-title":"Tina-Vision Memo"},{"key":"e_1_3_2_1_6_1","volume-title":"Kyung Woo Park, and Sung Ho Kim.","author":"Bu Hee Hyung","year":"2019","unstructured":"Hee Hyung Bu, Nam Chul Kim, Kyung Woo Park, and Sung Ho Kim. 2019. Content-based image retrieval using combined texture and color features based on multi-resolution multi-direction filtering and color autocorrelogram. Journal of Ambient Intelligence and Humanized Computing 3 (2019)."},{"key":"e_1_3_2_1_7_1","volume-title":"Proceedings of the 37th International Conference on Machine Learning, ICML 2020, 13--18","volume":"119","author":"Chen Ting","year":"2020","unstructured":"Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey E. Hinton. 2020b. A Simple Framework for Contrastive Learning of Visual Representations. In Proceedings of the 37th International Conference on Machine Learning, ICML 2020, 13--18 July 2020, Virtual Event, Vol. 119. PMLR, 1597--1607."},{"key":"e_1_3_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00307"},{"key":"e_1_3_2_1_9_1","volume-title":"Probabilistic Embeddings for Cross-Modal Retrieval. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR2021","author":"Chun Sanghyuk","year":"2021","unstructured":"Sanghyuk Chun, Seong Joon Oh, Rafael Sampaio de Rezende, Yannis Kalantidis, and Diane Larlus. 2021. Probabilistic Embeddings for Cross-Modal Retrieval. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR2021, virtual, June 19--25, 2021. Computer Vision Foundation \/ IEEE, 8415--8424."},{"key":"e_1_3_2_1_10_1","volume-title":"ARTEMIS: Attention-based Retrieval with Text-Explicit Matching and Implicit Similarity. In The Tenth International Conference on Learning Representations, ICLR2022","author":"Delmas Ginger","year":"2022","unstructured":"Ginger Delmas, Rafael Sampaio de Rezende, Gabriela Csurka, and Diane Larlus. 2022. ARTEMIS: Attention-based Retrieval with Text-Explicit Matching and Implicit Similarity. In The Tenth International Conference on Learning Representations, ICLR2022, Virtual Event, April 25--29, 2022. OpenReview.net."},{"key":"e_1_3_2_1_11_1","volume-title":"Modality-Agnostic Attention Fusion for visual search with text feedback. CoRR","author":"Dodds Eric","year":"2020","unstructured":"Eric Dodds, Jack Culpepper, Simao Herdade, Yang Zhang, and Kofi Boakye. 2020. Modality-Agnostic Attention Fusion for visual search with text feedback. CoRR , Vol. abs\/2007.00145 (2020)."},{"key":"e_1_3_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.ins.2024.120575"},{"key":"e_1_3_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1145\/3591106.3592279"},{"key":"e_1_3_2_1_14_1","doi-asserted-by":"crossref","unstructured":"Mingyuan Ge Yewen Li Honghao Wu and Mingyong Li. 2024. JM-CLIP: A Joint Modal Similarity Contrastive Learning Model for Video-Text Retrieval. In ICASSP 2024 - 2024 IEEE International Conference on Acoustics Speech and Signal Processing (ICASSP). 3010--3014.","DOI":"10.1109\/ICASSP48485.2024.10446490"},{"key":"e_1_3_2_1_15_1","volume-title":"Interactive Sketch & Fill: Multiclass Sketch-to-Image Translation. In 2019 IEEE\/CVF International Conference on Computer Vision, ICCV2019","author":"Ghosh Arnab","year":"2019","unstructured":"Arnab Ghosh, Richard Zhang, Puneet K. Dokania, Oliver Wang, Alexei A. Efros, Philip H. S. Torr, and Eli Shechtman. 2019. Interactive Sketch & Fill: Multiclass Sketch-to-Image Translation. In 2019 IEEE\/CVF International Conference on Computer Vision, ICCV2019, Seoul, Korea (South), October 27 - November 2, 2019. IEEE, 1171--1180."},{"key":"e_1_3_2_1_16_1","first-page":"241","article-title":"Deep Image Retrieval: Learning Global Representations for Image Search. In Computer Vision - ECCV 2016 - 14th European Conference, Amsterdam, The Netherlands, October 11--14, 2016","volume":"9910","author":"Gordo Albert","year":"2016","unstructured":"Albert Gordo, Jon Almaz\u00e1 n, J\u00e9 r\u00f4 me Revaud, and Diane Larlus. 2016. Deep Image Retrieval: Learning Global Representations for Image Search. In Computer Vision - ECCV 2016 - 14th European Conference, Amsterdam, The Netherlands, October 11--14, 2016, Proceedings, Part VI, Vol. 9910. 241--257.","journal-title":"Proceedings, Part VI"},{"key":"e_1_3_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.inffus.2024.102417"},{"key":"e_1_3_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00365"},{"key":"e_1_3_2_1_19_1","first-page":"101157","article-title":"Deep learning approaches for multi-modal sensor data analysis and abnormality detection. Measurement","volume":"33","author":"Jadhav Santosh Pandurang","year":"2024","unstructured":"Santosh Pandurang Jadhav, Angalkuditi Srinivas, Patil Dipak Raghunath, M. Ramkumar Prabhu, Jaya Suryawanshi, and Anandakumar Haldorai. 2024. Deep learning approaches for multi-modal sensor data analysis and abnormality detection. Measurement: Sensors , Vol. 33 (2024), 101157--.","journal-title":"Sensors"},{"key":"e_1_3_2_1_20_1","volume-title":"TRACE: Transform Aggregate and Compose Visiolinguistic Representations for Image Search with Text Feedback. CoRR","author":"Jandial Surgan","year":"2020","unstructured":"Surgan Jandial, Ayush Chopra, Pinkesh Badjatiya, Pranit Chawla, Mausoom Sarkar, and Balaji Krishnamurthy. 2020. TRACE: Transform Aggregate and Compose Visiolinguistic Representations for Image Search with Text Feedback. CoRR , Vol. abs\/2009.01485 (2020)."},{"key":"e_1_3_2_1_21_1","volume-title":"Multimodal Residual Learning for Visual QA. In Advances in Neural Information Processing Systems 29: Annual Conference on Neural Information Processing Systems 2016","author":"Kim Jin-Hwa","year":"2016","unstructured":"Jin-Hwa Kim, Sang-Woo Lee, Dong-Hyun Kwak, Min-Oh Heo, Jeonghee Kim, JungWoo Ha, and Byoung-Tak Zhang. 2016. Multimodal Residual Learning for Visual QA. In Advances in Neural Information Processing Systems 29: Annual Conference on Neural Information Processing Systems 2016, December 5--10, 2016, Barcelona, Spain. 361--369."},{"key":"e_1_3_2_1_22_1","volume-title":"Woosang Lim, Jeonghee Kim, Jung-Woo Ha, and Byoung-Tak Zhang.","author":"Kim Jin-Hwa","year":"2017","unstructured":"Jin-Hwa Kim, Kyoung Woon On, Woosang Lim, Jeonghee Kim, Jung-Woo Ha, and Byoung-Tak Zhang. 2017. Hadamard Product for Low-rank Bilinear Pooling. In 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24--26, 2017, Conference Track Proceedings. OpenReview.net."},{"key":"e_1_3_2_1_23_1","volume-title":"Dual Compositional Learning in Interactive Image Retrieval. In Thirty-Fifth AAAI Conference on Artificial Intelligence, AAAI2021","author":"Kim Jongseok","year":"2021","unstructured":"Jongseok Kim, Youngjae Yu, Hoeseong Kim, and Gunhee Kim. 2021. Dual Compositional Learning in Interactive Image Retrieval. In Thirty-Fifth AAAI Conference on Artificial Intelligence, AAAI2021. AAAI Press, 1771--1779."},{"volume-title":"19th European Symposium on Artificial Neural Networks, ESANN 2011, Bruges, Belgium, April 27--29, 2011, Proceedings.","author":"Krizhevsky Alex","key":"e_1_3_2_1_24_1","unstructured":"Alex Krizhevsky and Geoffrey E. Hinton. 2011. Using very deep autoencoders for content-based image retrieval. In 19th European Symposium on Artificial Neural Networks, ESANN 2011, Bruges, Belgium, April 27--29, 2011, Proceedings."},{"key":"e_1_3_2_1_25_1","volume-title":"CoSMo: Content-Style Modulation for Image Retrieval With Text Feedback. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR2021","author":"Lee Seungmin","year":"2021","unstructured":"Seungmin Lee, Dongwan Kim, and Bohyung Han. 2021. CoSMo: Content-Style Modulation for Image Retrieval With Text Feedback. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR2021, virtual, June 19--25, 2021. Computer Vision Foundation \/ IEEE, 802--812."},{"key":"e_1_3_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1007\/S13735-023-00268--7"},{"key":"e_1_3_2_1_27_1","volume-title":"Image Retrieval on Real-life Images with Pre-trained Vision-and-Language Models. In 2021 IEEE\/CVF International Conference on Computer Vision, ICCV2021","author":"Liu Zheyuan","year":"2021","unstructured":"Zheyuan Liu, Cristian Rodriguez Opazo, Damien Teney, and Stephen Gould. 2021. Image Retrieval on Real-life Images with Pre-trained Vision-and-Language Models. In 2021 IEEE\/CVF International Conference on Computer Vision, ICCV2021, Montreal, QC, Canada, October 10--17, 2021. IEEE, 2105--2114."},{"key":"e_1_3_2_1_28_1","volume-title":"Advanced confidence methods in deep learning. Physica A: Statistical Mechanics and its Applications","author":"Meir Yuval","year":"2024","unstructured":"Yuval Meir, Ofek Tevet, Ella Koresh, Yarden Tzach, and Ido Kanter. 2024. Advanced confidence methods in deep learning. Physica A: Statistical Mechanics and its Applications , Vol. 641 (2024), 129758--."},{"key":"e_1_3_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.aej.2020.02.034"},{"key":"e_1_3_2_1_30_1","volume-title":"Modeling uncertainty with hedged instance embedding. arXiv preprint arXiv:1810.00319","author":"Oh Seong Joon","year":"2018","unstructured":"Seong Joon Oh, Kevin Murphy, Jiyan Pan, Joseph Roth, Florian Schroff, and Andrew Gallagher. 2018. Modeling uncertainty with hedged instance embedding. arXiv preprint arXiv:1810.00319 (2018)."},{"key":"e_1_3_2_1_31_1","volume-title":"Proceedings of the 38th International Conference on Machine Learning, ICML 2021, 18--24","volume":"8763","author":"Radford Alec","year":"2021","unstructured":"Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021. Learning Transferable Visual Models From Natural Language Supervision. In Proceedings of the 38th International Conference on Machine Learning, ICML 2021, 18--24 July 2021, Virtual Event (Proceedings of Machine Learning Research, Vol. 139). PMLR, 8748--8763."},{"key":"e_1_3_2_1_32_1","volume-title":"RTIC: Residual Learning for Text and Image Composition using Graph Convolutional Network. CoRR","author":"Shin Minchul","year":"2021","unstructured":"Minchul Shin, Yoonjae Cho, ByungSoo Ko, and Geonmo Gu. 2021. RTIC: Residual Learning for Text and Image Composition using Graph Convolutional Network. CoRR , Vol. abs\/2104.03015 (2021)."},{"key":"e_1_3_2_1_33_1","volume-title":"Image Search with Text Feedback by Additive Attention Compositional Learning. CoRR","author":"Tian Yuxin","year":"2022","unstructured":"Yuxin Tian, Shawn D. Newsam, and Kofi Boakye. 2022. Image Search with Text Feedback by Additive Attention Compositional Learning. CoRR , Vol. abs\/2203.03809 (2022)."},{"key":"e_1_3_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00660"},{"key":"e_1_3_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.01115"},{"key":"e_1_3_2_1_36_1","volume-title":"Ashish Mishra, and Anurag Mittal.","author":"Yelamarthi Sasi Kiran","year":"2018","unstructured":"Sasi Kiran Yelamarthi, M. Shiva Krishna Reddy, Ashish Mishra, and Anurag Mittal. 2018. A Zero-Shot Framework for Sketch Based Image Retrieval. In Computer Vision - ECCV 2018 - 15th European Conference, Munich, Germany, September 8--14, 2018, Proceedings, Part IV, Vol. 11208. Springer, 316--333."},{"key":"e_1_3_2_1_37_1","volume-title":"Thinking Outside the Pool: Active Training Image Creation for Relative Attributes. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2019","author":"Yu Aron","year":"2019","unstructured":"Aron Yu and Kristen Grauman. 2019. Thinking Outside the Pool: Active Training Image Creation for Relative Attributes. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2019, Long Beach, CA, USA, June 16--20, 2019. Computer Vision Foundation \/ IEEE, 708--718."},{"key":"e_1_3_2_1_38_1","volume-title":"CurlingNet: Compositional Learning between Images and Text for Fashion IQ Data. CoRR","author":"Yu Youngjae","year":"2020","unstructured":"Youngjae Yu, Seunghwan Lee, Yuncheol Choi, and Gunhee Kim. 2020. CurlingNet: Compositional Learning between Images and Text for Fashion IQ Data. CoRR , Vol. abs\/2003.12299 (2020)."},{"key":"e_1_3_2_1_39_1","volume-title":"Multi-modal Factorized Bilinear Pooling with Co-attention Learning for Visual Question Answering. In IEEE International Conference on Computer Vision, ICCV 2017","author":"Yu Zhou","year":"2017","unstructured":"Zhou Yu, Jun Yu, Jianping Fan, and Dacheng Tao. 2017a. Multi-modal Factorized Bilinear Pooling with Co-attention Learning for Visual Question Answering. In IEEE International Conference on Computer Vision, ICCV 2017, Venice, Italy, October 22--29, 2017. 1839--1848."},{"key":"e_1_3_2_1_40_1","first-page":"1","article-title":"Beyond Bilinear: Generalized Multi-modal Factorized High-order Pooling for Visual Question Answering","volume":"99","author":"Yu Zhou","year":"2017","unstructured":"Zhou Yu, Jun Yu, Chenchao Xiang, Jianping Fan, and Dacheng Tao. 2017b. Beyond Bilinear: Generalized Multi-modal Factorized High-order Pooling for Visual Question Answering. IEEE Transactions on Neural Networks & Learning Systems, Vol. PP, 99 (2017), 1--13.","journal-title":"PP"},{"key":"e_1_3_2_1_41_1","volume-title":"Context-Aware Attention Network for Image-Text Retrieval. In 2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2020","author":"Zhang Qi","year":"2020","unstructured":"Qi Zhang, Zhen Lei, Zhaoxiang Zhang, and Stan Z. Li. 2020. Context-Aware Attention Network for Image-Text Retrieval. In 2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2020, Seattle, WA, USA, June 13--19, 2020. Computer Vision Foundation \/ IEEE, 3533--3542."},{"key":"e_1_3_2_1_42_1","volume-title":"Progressive Learning for Image Retrieval with Hybrid-Modality Queries. In SIGIR '22: The 45th International ACM SIGIR Conference on Research and Development in Information Retrieval","author":"Zhao Yida","year":"2022","unstructured":"Yida Zhao, Yuqing Song, and Qin Jin. 2022. Progressive Learning for Image Retrieval with Hybrid-Modality Queries. In SIGIR '22: The 45th International ACM SIGIR Conference on Research and Development in Information Retrieval, Madrid, Spain, July 11 - 15, 2022, Enrique Amig\u00f3, Pablo Castells, Julio Gonzalo, Ben Carterette, J. Shane Culpepper, and Gabriella Kazai (Eds.). ACM, 1012--1021."}],"event":{"name":"ICMR '24: International Conference on Multimedia Retrieval","sponsor":["SIGMM ACM Special Interest Group on Multimedia","SIGSOFT ACM Special Interest Group on Software Engineering"],"location":"Phuket Thailand","acronym":"ICMR '24"},"container-title":["Proceedings of the 2024 International Conference on Multimedia Retrieval"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3652583.3657611","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3652583.3657611","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,8,21]],"date-time":"2025-08-21T08:48:28Z","timestamp":1755766108000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3652583.3657611"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,5,30]]},"references-count":42,"alternative-id":["10.1145\/3652583.3657611","10.1145\/3652583"],"URL":"https:\/\/doi.org\/10.1145\/3652583.3657611","relation":{},"subject":[],"published":{"date-parts":[[2024,5,30]]},"assertion":[{"value":"2024-06-07","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}