{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T22:10:03Z","timestamp":1750198203489,"version":"3.41.0"},"reference-count":59,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2021,4,21]],"date-time":"2021-04-21T00:00:00Z","timestamp":1618963200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"JSPS KAKENHI","award":["19H04172"],"award-info":[{"award-number":["19H04172"]}]},{"name":"MIC\/SCOPE","award":["#172107101"],"award-info":[{"award-number":["#172107101"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2021,5,31]]},"abstract":"<jats:p>With the significant development of black-box machine learning algorithms, particularly deep neural networks, the practical demand for reliability assessment is rapidly increasing. On the basis of the concept that \u201cBayesian deep learning knows what it does not know,\u201d the uncertainty of deep neural network outputs has been investigated as a reliability measure for classification and regression tasks. By considering an embedding task as a regression task, several existing studies have quantified the uncertainty of embedded features and improved the retrieval performance of cutting-edge models by model averaging. However, in image-caption embedding-and-retrieval tasks, well-known samples are not always easy to retrieve. This study shows that the existing method has poor performance in reliability assessment and investigates another aspect of image-caption embedding-and-retrieval tasks. We propose posterior uncertainty by considering the retrieval task as a classification task, which can accurately assess the reliability of retrieval results. The consistent performance of the two uncertainty measures is observed with different datasets (MS-COCO and Flickr30k), different deep-learning architectures (dropout and batch normalization), and different similarity functions. To the best of our knowledge, this is the first study to perform a reliability assessment on image-caption embedding-and-retrieval tasks.<\/jats:p>","DOI":"10.1145\/3425663","type":"journal-article","created":{"date-parts":[[2021,4,22]],"date-time":"2021-04-22T06:10:03Z","timestamp":1619071803000},"page":"1-19","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["Exploring Uncertainty Measures for Image-caption Embedding-and-retrieval Task"],"prefix":"10.1145","volume":"17","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-2338-9229","authenticated-orcid":false,"given":"Kenta","family":"Hama","sequence":"first","affiliation":[{"name":"Osaka University, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0642-4800","authenticated-orcid":false,"given":"Takashi","family":"Matsubara","sequence":"additional","affiliation":[{"name":"Osaka University, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Kuniaki","family":"Uehara","sequence":"additional","affiliation":[{"name":"Osaka Gakuin University, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jianfei","family":"Cai","sequence":"additional","affiliation":[{"name":"Monash University, Australia"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2021,4,21]]},"reference":[{"volume-title":"International Conference on Learning Representations Workhosp (ICLRW\u201918)","year":"2018","author":"Atanov Andrei","key":"e_1_2_1_1_1"},{"volume-title":"Bishopt","year":"1997","author":"Barber David","key":"e_1_2_1_2_1"},{"volume-title":"ICML Workshop.","year":"2012","author":"Bengio Yoshua","key":"e_1_2_1_3_1"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2013.50"},{"volume-title":"Pattern Recognition and Machine Learning","author":"Bishop Christopher M.","key":"e_1_2_1_5_1"},{"volume-title":"International Conference on Learning Representations (ICLR\u201919)","year":"2019","author":"Chen Wei-Yu","key":"e_1_2_1_6_1"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/D14-1179"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2018.2832602"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00957"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00419"},{"volume-title":"British Machine Vision Conference (BMVC\u201918)","year":"2018","author":"Faghri Fartash","key":"e_1_2_1_12_1"},{"key":"e_1_2_1_13_1","unstructured":"Andrea Frome Greg S. Corrado Jon Shlens Samy Bengio Jeff Dean Marc\u2019Aurelio Ranzato and Tomas Mikolov. 2013. DeViSE: A deep visual-semantic embedding model. Advances in Neural Information Processing Systems (NIPS\u201913).  Andrea Frome Greg S. Corrado Jon Shlens Samy Bengio Jeff Dean Marc\u2019Aurelio Ranzato and Tomas Mikolov. 2013. DeViSE: A deep visual-semantic embedding model. Advances in Neural Information Processing Systems (NIPS\u201913)."},{"volume-title":"International Conference on Machine Learning (ICML\u201916)","year":"2016","author":"Gal Yarin","key":"e_1_2_1_15_1"},{"volume-title":"International Conference on Machine Learning (ICML\u201915)","year":"2015","author":"Ganin Yaroslav","key":"e_1_2_1_16_1"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00750"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00013"},{"volume-title":"International Conference on Learning Representations (ICLR\u201917)","year":"2017","author":"Higgins Irina","key":"e_1_2_1_20_1"},{"volume-title":"Annual Conference on Computational Learning Theory (COLT\u201993)","author":"Geoffrey","key":"e_1_2_1_21_1"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01246-5_32"},{"volume-title":"International Conference on Machine Learning (ICML\u201915)","year":"2015","author":"Ioffe Sergey","key":"e_1_2_1_23_1"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1145\/2939672.2939756"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298932"},{"key":"e_1_2_1_26_1","unstructured":"Alex Kendall and Yarin Gal. 2017. What uncertainties do we need in Bayesian deep learning for computer vision? In Advances in Neural Information Processing Systems (NIPS\u201917).  Alex Kendall and Yarin Gal. 2017. What uncertainties do we need in Bayesian deep learning for computer vision? In Advances in Neural Information Processing Systems (NIPS\u201917)."},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.5244\/C.31.57"},{"volume-title":"International Conference on Learning Representations (ICLR\u201915)","author":"Diederik","key":"e_1_2_1_28_1"},{"key":"e_1_2_1_29_1","unstructured":"Diederik P. Kingma Tim Salimans Max Welling and Machine Learning Group. 2015. Variational dropout and the local reparameterization trick. In Advances in Neural Information Processing Systems (NIPS\u201915).  Diederik P. Kingma Tim Salimans Max Welling and Machine Learning Group. 2015. Variational dropout and the local reparameterization trick. In Advances in Neural Information Processing Systems (NIPS\u201915)."},{"volume-title":"International Conference on Learning Representations (ICLR\u201914)","author":"Diederik","key":"e_1_2_1_30_1"},{"volume-title":"NIPS Workshop.","author":"Kiros Ryan","key":"e_1_2_1_31_1"},{"volume-title":"Aleatory or epistemic? Does it matter? Struct. Safety 31, 2","year":"2009","author":"Kiureghian Armen Der","key":"e_1_2_1_32_1"},{"volume-title":"NIPS Workshop.","year":"2016","author":"Leibig Christian","key":"e_1_2_1_33_1"},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00475"},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.199"},{"volume-title":"European Conference on Computer Vision (ECCV\u201914)","author":"Lin Tsung Yi","key":"e_1_2_1_36_1"},{"volume-title":"A practical Bayesian framework for backpropagation networks. Neural Comput. 4, 3","year":"1992","author":"MacKay David J. C.","key":"e_1_2_1_37_1"},{"key":"e_1_2_1_38_1","unstructured":"Andrey Malinin and Mark Gales. 2018. Predictive uncertainty estimation via prior networks. In Advances in Neural Information Processing Systems (NIPS\u201918).  Andrey Malinin and Mark Gales. 2018. Predictive uncertainty estimation via prior networks. In Advances in Neural Information Processing Systems (NIPS\u201918)."},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1109\/IJCNN.2018.8489169"},{"key":"e_1_2_1_40_1","unstructured":"Tomas Mikolov Kai Chen Greg Corrado and Jeffrey Dean. 2013. Distributed representations of words and phrases and their compositionality. In Advances in Neural Information Processing Systems (NIPS\u201913).  Tomas Mikolov Kai Chen Greg Corrado and Jeffrey Dean. 2013. Distributed representations of words and phrases and their compositionality. In Advances in Neural Information Processing Systems (NIPS\u201913)."},{"volume-title":"International Conference on Learning Representations (ICLR\u201915)","year":"2015","author":"Miyato Takeru","key":"e_1_2_1_41_1"},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D18-1437"},{"volume-title":"Andrea Caponnetto, Michele Piana, and Alessandro Verri.","year":"2004","author":"Rosasco Lorenzo","key":"e_1_2_1_43_1"},{"key":"e_1_2_1_44_1","unstructured":"Tim Salimans Ian Goodfellow Wojciech Zaremba Vicki Cheung Alec Radford Xi Chen Ishaan Gulrajani Faruk Ahmed Martin Arjovsky Vincent Dumoulin and Aaron Courville. 2017. Improved training of wasserstein GANs. In Advances in Neural Information Processing Systems (NIPS\u201917).  Tim Salimans Ian Goodfellow Wojciech Zaremba Vicki Cheung Alec Radford Xi Chen Ishaan Gulrajani Faruk Ahmed Martin Arjovsky Vincent Dumoulin and Aaron Courville. 2017. Improved training of wasserstein GANs. In Advances in Neural Information Processing Systems (NIPS\u201917)."},{"volume-title":"Deep learning in neural networks: An overview. Neural Netw. 61","year":"2015","author":"Schmidhuber J\u00fcrgen","key":"e_1_2_1_45_1"},{"volume-title":"International Conference on Learning Representations (ICLR\u201915)","year":"2015","author":"Simonyan Karen","key":"e_1_2_1_46_1"},{"key":"e_1_2_1_47_1","unstructured":"Lewis Smith and Yarin Gal. 2018. Understanding measures of uncertainty for adversarial example detection. In Uncertainty in Artificial Intelligence (UAI\u201918).  Lewis Smith and Yarin Gal. 2018. Understanding measures of uncertainty for adversarial example detection. In Uncertainty in Artificial Intelligence (UAI\u201918)."},{"volume-title":"Dropout: A simple way to prevent neural networks from overfitting. J. Mach. Learn. Res. 15","year":"2014","author":"Srivastava Nitish","key":"e_1_2_1_48_1"},{"key":"e_1_2_1_49_1","unstructured":"Ahmed Taha Yi-ting Chen Teruhisa Misu Abhinav Shrivastava and Larry Davis. 2019. Unsupervised data uncertainty learning in visual retrieval systems. CoRR abs\/1902.02586.  Ahmed Taha Yi-ting Chen Teruhisa Misu Abhinav Shrivastava and Larry Davis. 2019. Unsupervised data uncertainty learning in visual retrieval systems. CoRR abs\/1902.02586."},{"key":"e_1_2_1_50_1","unstructured":"Ahmed Taha Yi-Ting Chen Xitong Yang Teruhisa Misu and Larry Davis. 2019. Exploring uncertainty in conditional multi-modal retrieval systems. CoRR abs\/1901.07702.  Ahmed Taha Yi-Ting Chen Xitong Yang Teruhisa Misu and Larry Davis. 2019. Exploring uncertainty in conditional multi-modal retrieval systems. CoRR abs\/1901.07702."},{"volume-title":"Asian Conference on Machine Learning (ACML\u201918)","year":"2018","author":"Takahashi Ryo","key":"e_1_2_1_51_1"},{"volume-title":"Equations of states in singular statistical estimation. Neural Netw. 23, 1","year":"2010","author":"Watanabe Sumio","key":"e_1_2_1_52_1"},{"key":"e_1_2_1_53_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10994-010-5198-3"},{"key":"e_1_2_1_54_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v33i01.33017322"},{"key":"e_1_2_1_55_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D18-1166"},{"key":"e_1_2_1_56_1","doi-asserted-by":"crossref","unstructured":"M. H. Peter Young Alice Lai and J. Hockenmaier. 2014. From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions. Trans. Assoc. Comput. Ling. 2 (2014).  M. H. Peter Young Alice Lai and J. Hockenmaier. 2014. From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions. Trans. Assoc. Comput. Ling. 2 (2014).","DOI":"10.1162\/tacl_a_00166"},{"volume-title":"International Conference on Learning Representations (ICLR\u201918)","year":"2018","author":"Zhang Hongyi","key":"e_1_2_1_57_1"},{"key":"e_1_2_1_58_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.321"},{"volume-title":"AAAI Conference on Artificial Intelligence (AAAI\u201918)","year":"2018","author":"Zhang Quanshi","key":"e_1_2_1_59_1"},{"key":"e_1_2_1_60_1","article-title":"Dual-path convolutional image-text embedding with instance loss","volume":"16","author":"Zheng Zhedong","year":"2021","journal-title":"ACM Transactions on Multimedia Computing, Communications, and Applications"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3425663","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3425663","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T21:31:55Z","timestamp":1750195915000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3425663"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,4,21]]},"references-count":59,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2021,5,31]]}},"alternative-id":["10.1145\/3425663"],"URL":"https:\/\/doi.org\/10.1145\/3425663","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"type":"print","value":"1551-6857"},{"type":"electronic","value":"1551-6865"}],"subject":[],"published":{"date-parts":[[2021,4,21]]},"assertion":[{"value":"2019-12-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2020-09-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2021-04-21","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}