{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,28]],"date-time":"2026-04-28T02:00:16Z","timestamp":1777341616066,"version":"3.51.4"},"reference-count":62,"publisher":"Association for Computing Machinery (ACM)","issue":"1s","license":[{"start":{"date-parts":[[2020,1,31]],"date-time":"2020-01-31T00:00:00Z","timestamp":1580428800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100012166","name":"National Key R&D Program of China","doi-asserted-by":"crossref","award":["2018YFC0806900"],"award-info":[{"award-number":["2018YFC0806900"]}],"id":[{"id":"10.13039\/501100012166","id-type":"DOI","asserted-by":"crossref"}]},{"name":"National Engineering Laboratory for Public Safety Risk Perception and Control by Big Data"},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["61571269, 61701273"],"award-info":[{"award-number":["61571269, 61701273"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2020,1,31]]},"abstract":"<jats:p>Cross-modality human behavior analysis has attracted much attention from both academia and industry. In this article, we focus on the cross-modality image-text retrieval problem for human behavior analysis, which can learn a common latent space for cross-modality data and thus benefit the understanding of human behavior with data from different modalities. Existing state-of-the-art cross-modality image-text retrieval models tend to be fine-grained region-word matching approaches, where they begin with measuring similarities for each image region or text word followed by aggregating them to estimate the global image-text similarity. However, it is observed that such fine-grained approaches often encounter the similarity bias problem, because they only consider matched text words for an image region or matched image regions for a text word for similarity calculation, but they totally ignore unmatched words\/regions, which might still be salient enough to affect the global image-text similarity. In this article, we propose an Adaptive Confidence Matching Network (ACMNet), which is also a fine-grained matching approach, to effectively deal with such a similarity bias. Apart from calculating the local similarity for each region(\/word) with its matched words(\/regions), ACMNet also introduces a confidence score for the local similarity by leveraging the global text(\/image) information, which is expected to help measure the semantic relatedness of the region(\/word) to the whole text(\/image). Moreover, ACMNet also incorporates the confidence scores together with the local similarities in estimating the global image-text similarity. To verify the effectiveness of ACMNet, we conduct extensive experiments and make comparisons with state-of-the-art methods on two benchmark datasets, i.e., Flickr30k and MS COCO. Experimental results show that the proposed ACMNet can outperform the state-of-the-art methods by a clear margin, which well demonstrates the effectiveness of the proposed ACMNet in human behavior analysis and the reasonableness of tackling the mentioned similarity bias issue.<\/jats:p>","DOI":"10.1145\/3362065","type":"journal-article","created":{"date-parts":[[2020,5,4]],"date-time":"2020-05-04T07:01:36Z","timestamp":1588575696000},"page":"1-21","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":9,"title":["ACMNet"],"prefix":"10.1145","volume":"16","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-4180-5801","authenticated-orcid":false,"given":"Hui","family":"Chen","sequence":"first","affiliation":[{"name":"Beijing National Research Center for Information Science and Technology (BNRist); School of Software, Tsinghua University, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Guiguang","family":"Ding","sequence":"additional","affiliation":[{"name":"Beijing National Research Center for Information Science and Technology (BNRist); School of Software, Tsinghua University, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zijia","family":"Lin","sequence":"additional","affiliation":[{"name":"Microsoft Research, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5843-6411","authenticated-orcid":false,"given":"Sicheng","family":"Zhao","sequence":"additional","affiliation":[{"name":"University of California, Berkeley, Berkeley, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xiaopeng","family":"Gu","sequence":"additional","affiliation":[{"name":"National Engineering Laboratory for Public Safety Risk Perception and Control by Big Data(PSRPC); China Academy of Electronics and Information Technology, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Wenyuan","family":"Xu","sequence":"additional","affiliation":[{"name":"Ubiquitous System Security Lab (USSLab), Zhejiang University, Zhejiang, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jungong","family":"Han","sequence":"additional","affiliation":[{"name":"University of Warwick, Coventry, UK"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2020,4,17]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00636"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.279"},{"key":"e_1_2_1_3_1","volume-title":"Neural machine translation by jointly learning to align and translate. Arxiv Preprint Arxiv:1409.0473","author":"Bahdanau Dzmitry","year":"2014","unstructured":"Dzmitry Bahdanau , Kyunghyun Cho , and Yoshua Bengio . 2014. Neural machine translation by jointly learning to align and translate. Arxiv Preprint Arxiv:1409.0473 ( 2014 ). Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2014. Neural machine translation by jointly learning to align and translate. Arxiv Preprint Arxiv:1409.0473 (2014)."},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00051"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.24963\/ijcai.2018\/84"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/D14-1179"},{"key":"e_1_2_1_7_1","volume-title":"Advances in Neural Information Processing Systems","author":"Chorowski Jan K.","unstructured":"Jan K. Chorowski , Dzmitry Bahdanau , Dmitriy Serdyuk , Kyunghyun Cho , and Yoshua Bengio . 2015. Attention-based models for speech recognition . In Advances in Neural Information Processing Systems . MIT Press , 577--585. Jan K. Chorowski, Dzmitry Bahdanau, Dmitriy Serdyuk, Kyunghyun Cho, and Yoshua Bengio. 2015. Attention-based models for speech recognition. In Advances in Neural Information Processing Systems. MIT Press, 577--585."},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"e_1_2_1_9_1","volume-title":"BERT: Pre-training of deep bidirectional transformers for language understanding. Arxiv Preprint Arxiv:1810.04805","author":"Devlin Jacob","year":"2018","unstructured":"Jacob Devlin , Ming-Wei Chang , Kenton Lee , and Kristina Toutanova . 2018 . BERT: Pre-training of deep bidirectional transformers for language understanding. Arxiv Preprint Arxiv:1810.04805 (2018). Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. BERT: Pre-training of deep bidirectional transformers for language understanding. Arxiv Preprint Arxiv:1810.04805 (2018)."},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/TITS.2017.2774778"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2019.2902115"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00957"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.201"},{"key":"e_1_2_1_14_1","volume-title":"Jamie Ryan Kiros, and Sanja Fidler","author":"Faghri Fartash","year":"2017","unstructured":"Fartash Faghri , David J. Fleet , Jamie Ryan Kiros, and Sanja Fidler . 2017 . Vse++: Improving visual-semantic embeddings with hard negatives. Arxiv Preprint Arxiv :1707.05612 (2017). Fartash Faghri, David J. Fleet, Jamie Ryan Kiros, and Sanja Fidler. 2017. Vse++: Improving visual-semantic embeddings with hard negatives. Arxiv Preprint Arxiv:1707.05612 (2017)."},{"key":"e_1_2_1_15_1","volume-title":"Devise: A deep visual-semantic embedding model. In Advances in Neural Information Processing Systems","author":"Frome Andrea","year":"2013","unstructured":"Andrea Frome , Greg S. Corrado , Jon Shlens , Samy Bengio , Jeff Dean , Tomas Mikolov , 2013 . Devise: A deep visual-semantic embedding model. In Advances in Neural Information Processing Systems . MIT Press , 2121--2129. Andrea Frome, Greg S. Corrado, Jon Shlens, Samy Bengio, Jeff Dean, Tomas Mikolov, et al. 2013. Devise: A deep visual-semantic embedding model. In Advances in Neural Information Processing Systems. MIT Press, 2121--2129."},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2014.81"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00750"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.767"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00645"},{"key":"e_1_2_1_21_1","volume-title":"Proceedings of the Association for the Advancement of Artificial Intelligence (AAAI\u201919)","author":"Ding Guiguang","year":"2019","unstructured":"Guiguang Ding , Jianguang Lou , Yusen Zhang , Hui Chen , Zijia Lin , and Borje Karlsson . 2019 . GRN: Gated relation network to enhance convolutional neural network for named entity recognition . In Proceedings of the Association for the Advancement of Artificial Intelligence (AAAI\u201919) . Guiguang Ding, Jianguang Lou, Yusen Zhang, Hui Chen, Zijia Lin, and Borje Karlsson. 2019. GRN: Gated relation network to enhance convolutional neural network for named entity recognition. In Proceedings of the Association for the Advancement of Artificial Intelligence (AAAI\u201919)."},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.5555\/2789272.2789280"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/E17-2068"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298932"},{"key":"e_1_2_1_25_1","volume-title":"Advances in Neural Information Processing Systems","author":"Kim Jin-Hwa","unstructured":"Jin-Hwa Kim , Jaehyun Jun , and Byoung-Tak Zhang . 2018. Bilinear attention networks . In Advances in Neural Information Processing Systems . MIT Press , 1571--1581. Jin-Hwa Kim, Jaehyun Jun, and Byoung-Tak Zhang. 2018. Bilinear attention networks. In Advances in Neural Information Processing Systems. MIT Press, 1571--1581."},{"key":"e_1_2_1_26_1","volume-title":"Zemel","author":"Kiros Ryan","year":"2014","unstructured":"Ryan Kiros , Ruslan Salakhutdinov , and Richard S . Zemel . 2014 . Unifying visual-semantic embeddings with multimodal neural language models. Arxiv Preprint Arxiv :1411.2539 (2014). Ryan Kiros, Ruslan Salakhutdinov, and Richard S. Zemel. 2014. Unifying visual-semantic embeddings with multimodal neural language models. Arxiv Preprint Arxiv:1411.2539 (2014)."},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-016-0981-7"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2015.2481325"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIE.2019.2898618"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2017.2777183"},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01225-0_13"},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1145\/3209978.3210003"},{"key":"e_1_2_1_33_1","volume-title":"Proceedings of the 1st International Conference on Learning Representations (ICLR\u201913)","author":"Mikolov Tomas","year":"2013","unstructured":"Tomas Mikolov , Kai Chen , Greg Corrado , and Jeffrey Dean . 2013 . Efficient estimation of word representations in vector space . In Proceedings of the 1st International Conference on Learning Representations (ICLR\u201913) . Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013. Efficient estimation of word representations in vector space. In Proceedings of the 1st International Conference on Learning Representations (ICLR\u201913)."},{"key":"e_1_2_1_34_1","volume-title":"Advances in Neural Information Processing Systems","author":"Mikolov Tomas","unstructured":"Tomas Mikolov , Ilya Sutskever , Kai Chen , Greg S. Corrado , and Jeff Dean . 2013. Distributed representations of words and phrases and their compositionality . In Advances in Neural Information Processing Systems . MIT Press , 3111--3119. Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S. Corrado, and Jeff Dean. 2013. Distributed representations of words and phrases and their compositionality. In Advances in Neural Information Processing Systems. MIT Press, 3111--3119."},{"key":"e_1_2_1_35_1","volume-title":"Proceedings of the 26th ACM International Conference on Multimedia (MM\u201918)","author":"Mithun Niluthpol Chowdhury","year":"1856","unstructured":"Niluthpol Chowdhury Mithun , Rameswar Panda , Evangelos E. Papalexakis , and Amit K . Roy-Chowdhury. 2018. Webly supervised joint embedding for cross-modal image-text retrieval . In Proceedings of the 26th ACM International Conference on Multimedia (MM\u201918) . ACM, New York, NY , 1856 --1864. DOI:https:\/\/doi.org\/10.1145\/3240508.3240712 10.1145\/3240508.3240712 Niluthpol Chowdhury Mithun, Rameswar Panda, Evangelos E. Papalexakis, and Amit K. Roy-Chowdhury. 2018. Webly supervised joint embedding for cross-modal image-text retrieval. In Proceedings of the 26th ACM International Conference on Multimedia (MM\u201918). ACM, New York, NY, 1856--1864. DOI:https:\/\/doi.org\/10.1145\/3240508.3240712"},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.232"},{"key":"e_1_2_1_37_1","volume-title":"Proceedings of the IEEE International Conference on Computer Vision (ICCV\u201917). 1899--1907","author":"Niu Z.","year":"2017","unstructured":"Z. Niu , M. Zhou , L. Wang , X. Gao , and G. Hua . 2017. Hierarchical multimodal LSTM for dense visual-semantic embedding . In Proceedings of the IEEE International Conference on Computer Vision (ICCV\u201917). 1899--1907 . DOI:https:\/\/doi.org\/10.1109\/ICCV. 2017 .208 10.1109\/ICCV.2017.208 Z. Niu, M. Zhou, L. Wang, X. Gao, and G. Hua. 2017. Hierarchical multimodal LSTM for dense visual-semantic embedding. In Proceedings of the IEEE International Conference on Computer Vision (ICCV\u201917). 1899--1907. DOI:https:\/\/doi.org\/10.1109\/ICCV.2017.208"},{"key":"e_1_2_1_38_1","volume-title":"Proceedings of Neural Information Processing Systems.","author":"Paszke Adam","year":"2017","unstructured":"Adam Paszke , Sam Gross , Soumith Chintala , Gregory Chanan , Edward Yang , Zachary DeVito , Zeming Lin , Alban Desmaison , Luca Antiga , and Adam Lerer . 2017 . Automatic differentiation in PyTorch . In Proceedings of Neural Information Processing Systems. Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer. 2017. Automatic differentiation in PyTorch. In Proceedings of Neural Information Processing Systems."},{"key":"e_1_2_1_39_1","volume-title":"CM-GANs: Cross-modal generative adversarial networks for common representation learning. Arxiv Preprint Arxiv:1710.05106","author":"Peng Yuxin","year":"2017","unstructured":"Yuxin Peng , Jinwei Qi , and Yuxin Yuan . 2017. CM-GANs: Cross-modal generative adversarial networks for common representation learning. Arxiv Preprint Arxiv:1710.05106 ( 2017 ). Yuxin Peng, Jinwei Qi, and Yuxin Yuan. 2017. CM-GANs: Cross-modal generative adversarial networks for common representation learning. Arxiv Preprint Arxiv:1710.05106 (2017)."},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/D14-1162"},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/N18-1202"},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.303"},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2018.2828817"},{"key":"e_1_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.91"},{"key":"e_1_2_1_45_1","volume-title":"Advances in Neural Information Processing Systems","author":"Ren Shaoqing","unstructured":"Shaoqing Ren , Kaiming He , Ross Girshick , and Jian Sun . 2015. Faster r-cnn: Towards real-time object detection with region proposal networks . In Advances in Neural Information Processing Systems . MIT Press , 91--99. Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. 2015. Faster r-cnn: Towards real-time object detection with region proposal networks. In Advances in Neural Information Processing Systems. MIT Press, 91--99."},{"key":"e_1_2_1_46_1","volume-title":"Very deep convolutional networks for large-scale image recognition. Arxiv Preprint Arxiv:1409.1556","author":"Simonyan Karen","year":"2014","unstructured":"Karen Simonyan and Andrew Zisserman . 2014. Very deep convolutional networks for large-scale image recognition. Arxiv Preprint Arxiv:1409.1556 ( 2014 ). Karen Simonyan and Andrew Zisserman. 2014. Very deep convolutional networks for large-scale image recognition. Arxiv Preprint Arxiv:1409.1556 (2014)."},{"key":"e_1_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298594"},{"key":"e_1_2_1_48_1","volume-title":"Advances in Neural Information Processing Systems","author":"Vaswani Ashish","unstructured":"Ashish Vaswani , Noam Shazeer , Niki Parmar , Jakob Uszkoreit , Llion Jones , Aidan N. Gomez , \u0141ukasz Kaiser , and Illia Polosukhin . 2017. Attention is all you need . In Advances in Neural Information Processing Systems . MIT Press , 5998--6008. Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems. MIT Press, 5998--6008."},{"key":"e_1_2_1_49_1","volume-title":"Order-embeddings of images and language. Arxiv Preprint Arxiv:1511.06361","author":"Vendrov Ivan","year":"2015","unstructured":"Ivan Vendrov , Ryan Kiros , Sanja Fidler , and Raquel Urtasun . 2015. Order-embeddings of images and language. Arxiv Preprint Arxiv:1511.06361 ( 2015 ). Ivan Vendrov, Ryan Kiros, Sanja Fidler, and Raquel Urtasun. 2015. Order-embeddings of images and language. Arxiv Preprint Arxiv:1511.06361 (2015)."},{"key":"e_1_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.541"},{"key":"e_1_2_1_51_1","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201918)","author":"Wehrmann Janatas","unstructured":"Janatas Wehrmann and Rodrigo C. Barros . 2018. Bidirectional retrieval made simple . In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201918) . Janatas Wehrmann and Rodrigo C. Barros. 2018. Bidirectional retrieval made simple. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201918)."},{"key":"e_1_2_1_52_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2018.2882155"},{"key":"e_1_2_1_53_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIE.2018.2873547"},{"key":"e_1_2_1_54_1","volume-title":"Proceedings of the International Conference on Machine Learning. 2048--2057","author":"Xu Kelvin","year":"2015","unstructured":"Kelvin Xu , Jimmy Ba , Ryan Kiros , Kyunghyun Cho , Aaron Courville , Ruslan Salakhudinov , Rich Zemel , and Yoshua Bengio . 2015 . Show, attend and tell: Neural image caption generation with visual attention . In Proceedings of the International Conference on Machine Learning. 2048--2057 . Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhudinov, Rich Zemel, and Yoshua Bengio. 2015. Show, attend and tell: Neural image caption generation with visual attention. In Proceedings of the International Conference on Machine Learning. 2048--2057."},{"key":"e_1_2_1_55_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2019.2903448"},{"key":"e_1_2_1_56_1","first-page":"1","article-title":"STAT: spatial-temporal attention mechanism for video captioning","volume":"22","author":"Yan C","year":"2020","unstructured":"C Yan , Y. Tu , X Wang , Y. Zhang , X. Hao , Y. Zhang , and Q. Dai . 2020 . STAT: spatial-temporal attention mechanism for video captioning . IEEE Trans. Multimed. 22 , 1 (Jan 2020), 229--241. https:\/\/doi.org\/10.1109\/TMM.2019.2924576 10.1109\/TMM.2019.2924576 C Yan, Y. Tu, X Wang, Y. Zhang, X. Hao, Y. Zhang, and Q. Dai. 2020. STAT: spatial-temporal attention mechanism for video captioning. IEEE Trans. Multimed. 22, 1 (Jan 2020), 229--241. https:\/\/doi.org\/10.1109\/TMM.2019.2924576","journal-title":"IEEE Trans. Multimed."},{"key":"e_1_2_1_57_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298966"},{"key":"e_1_2_1_58_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2016.2539860"},{"key":"e_1_2_1_59_1","doi-asserted-by":"publisher","DOI":"10.1109\/TNNLS.2016.2586194"},{"key":"e_1_2_1_60_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2015.2406194"},{"key":"e_1_2_1_61_1","volume-title":"Dual-path convolutional image-text embedding with instance loss. Arxiv Preprint Arxiv:1711.05535","author":"Zheng Zhedong","year":"2017","unstructured":"Zhedong Zheng , Liang Zheng , Michael Garrett , Yi Yang , and Yi-Dong Shen . 2017. Dual-path convolutional image-text embedding with instance loss. Arxiv Preprint Arxiv:1711.05535 ( 2017 ). Zhedong Zheng, Liang Zheng, Michael Garrett, Yi Yang, and Yi-Dong Shen. 2017. Dual-path convolutional image-text embedding with instance loss. Arxiv Preprint Arxiv:1711.05535 (2017)."},{"key":"e_1_2_1_62_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICTAI.2017.00104"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3362065","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3362065","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T23:44:54Z","timestamp":1750203894000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3362065"}},"subtitle":["Adaptive Confidence Matching Network for Human Behavior Analysis via Cross-modal Retrieval"],"short-title":[],"issued":{"date-parts":[[2020,1,31]]},"references-count":62,"journal-issue":{"issue":"1s","published-print":{"date-parts":[[2020,1,31]]}},"alternative-id":["10.1145\/3362065"],"URL":"https:\/\/doi.org\/10.1145\/3362065","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,1,31]]},"assertion":[{"value":"2019-04-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2019-09-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2020-04-17","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}