{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,11]],"date-time":"2026-07-11T21:39:36Z","timestamp":1783805976210,"version":"3.55.0"},"reference-count":37,"publisher":"Association for Computing Machinery (ACM)","issue":"1","license":[{"start":{"date-parts":[[2020,2,17]],"date-time":"2020-02-17T00:00:00Z","timestamp":1581897600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"Research Grants Council of the Hong Kong Special Administrative Region, China"},{"name":"Collaborative Research Fund","award":["C1031-18G"],"award-info":[{"award-number":["C1031-18G"]}]},{"name":"Natural Science Foundation of Guangdong Province, China","award":["2018A0303130022"],"award-info":[{"award-number":["2018A0303130022"]}]},{"name":"Science and Technology Program of Guangzhou, China","award":["201904010200"],"award-info":[{"award-number":["201904010200"]}]},{"name":"Science and Technology Planning Project of Guangdong Province, China","award":["2016A010101012"],"award-info":[{"award-number":["2016A010101012"]}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["61703109, 91748107, 61902077, 61876065, and U1611461"],"award-info":[{"award-number":["61703109, 91748107, 61902077, 61876065, and U1611461"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/100012540","name":"Guangdong Innovative Research Team Program","doi-asserted-by":"crossref","award":["2014ZT05G157"],"award-info":[{"award-number":["2014ZT05G157"]}],"id":[{"id":"10.13039\/100012540","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2020,2,29]]},"abstract":"<jats:p>\n            In this article, we propose to learn shared semantic space with correlation alignment (\n            <jats:italic>S<\/jats:italic>\n            <jats:sup>3<\/jats:sup>\n            <jats:italic>CA<\/jats:italic>\n            ) for multimodal data representations, which aligns nonlinear correlations of multimodal data distributions in deep neural networks designed for heterogeneous data. In the context of cross-modal (event) retrieval, we design a neural network with convolutional layers and fully connected layers to extract features for images, including images on Flickr-like social media. Simultaneously, we exploit a fully connected neural network to extract semantic features for text documents, including news articles from news media. In particular, nonlinear correlations of layer activations in the two neural networks are aligned with correlation alignment during the joint training of the networks. Furthermore, we project the multimodal data into a shared semantic space for cross-modal (event) retrieval, where the distances between heterogeneous data samples can be measured directly. In addition, we contribute a Wiki-Flickr Event dataset, where the multimodal data samples are not describing each other in pairs like the existing paired datasets, but all of them are describing semantic events. Extensive experiments conducted on both paired and unpaired datasets manifest the effectiveness of\n            <jats:italic>S<\/jats:italic>\n            <jats:sup>3<\/jats:sup>\n            <jats:italic>CA<\/jats:italic>\n            , outperforming the state-of-the-art methods.\n          <\/jats:p>","DOI":"10.1145\/3374754","type":"journal-article","created":{"date-parts":[[2020,3,4]],"date-time":"2020-03-04T10:23:32Z","timestamp":1583317412000},"page":"1-22","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":23,"title":["Learning Shared Semantic Space with Correlation Alignment for Cross-Modal Event Retrieval"],"prefix":"10.1145","volume":"16","author":[{"given":"Zhenguo","family":"Yang","sequence":"first","affiliation":[{"name":"Guangdong University of Technology and City University of Hong Kong"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Zehang","family":"Lin","sequence":"additional","affiliation":[{"name":"Guangdong University of Technology"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Peipei","family":"Kang","sequence":"additional","affiliation":[{"name":"Guangdong University of Technology"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jianming","family":"LV","sequence":"additional","affiliation":[{"name":"South China University of Technology"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Qing","family":"Li","sequence":"additional","affiliation":[{"name":"Hong Kong Polytechnic University"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Wenyin","family":"Liu","sequence":"additional","affiliation":[{"name":"Guangdong University of Technology"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2020,2,17]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-00767-6_14"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.348"},{"key":"e_1_2_1_3_1","first-page":"449","article-title":"Cross-modal retrieval with CNN visual features: A new baseline","volume":"47","author":"Wei Yunchao","year":"2017","unstructured":"Yunchao Wei , Yao Zhao , Canyi Lu , Shikui Wei , Luoqi Liu , Zhenfeng Zhu , and Shuicheng Yan . 2017 . Cross-modal retrieval with CNN visual features: A new baseline . IEEE Transactions on Cybernetics 47 , 2 (2017), 449 -- 460 . Yunchao Wei, Yao Zhao, Canyi Lu, Shikui Wei, Luoqi Liu, Zhenfeng Zhu, and Shuicheng Yan. 2017. Cross-modal retrieval with CNN visual features: A new baseline. IEEE Transactions on Cybernetics 47, 2 (2017), 449--460.","journal-title":"IEEE Transactions on Cybernetics"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00446"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1162\/0899766042321814"},{"key":"e_1_2_1_6_1","volume-title":"Proceedings of the ACM International Conference on Multimedia. ACM","author":"Li Dongge","unstructured":"Dongge Li , Nevenka Dimitrova , Mingkun Li , and Ishwar K. Sethi . 2003. Multimedia content processing through cross-modal association . In Proceedings of the ACM International Conference on Multimedia. ACM , New York, NY, 604--611. Dongge Li, Nevenka Dimitrova, Mingkun Li, and Ishwar K. Sethi. 2003. Multimedia content processing through cross-modal association. In Proceedings of the ACM International Conference on Multimedia. ACM, New York, NY, 604--611."},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/1873951.1873987"},{"key":"e_1_2_1_8_1","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. IEEE","author":"Sharma Abhishek","unstructured":"Abhishek Sharma , Abhishek Kumar , Hal Daume , and David W. Jacobs . 2012. Generalized multiview analysis: A discriminative latent space . In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. IEEE , Los Alamitos, CA, 2160--2167. Abhishek Sharma, Abhishek Kumar, Hal Daume, and David W. Jacobs. 2012. Generalized multiview analysis: A discriminative latent space. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. IEEE, Los Alamitos, CA, 2160--2167."},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-013-0658-4"},{"key":"e_1_2_1_10_1","volume-title":"Proceedings of the IEEE International Conference on Computer Vision. 4094--4102","author":"Ranjan Viresh","unstructured":"Viresh Ranjan , Nikhil Rasiwasia , and C. V. Jawahar . 2015. Multi-label cross-modal retrieval . In Proceedings of the IEEE International Conference on Computer Vision. 4094--4102 . Viresh Ranjan, Nikhil Rasiwasia, and C. V. Jawahar. 2015. Multi-label cross-modal retrieval. In Proceedings of the IEEE International Conference on Computer Vision. 4094--4102."},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/3281746"},{"key":"e_1_2_1_12_1","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2046--2054","author":"Nhi Tran Thi Quynh","year":"2016","unstructured":"Thi Quynh Nhi Tran , Herv\u00e9 Le Borgne , and Michel Crucianu . 2016 . Aggregating image and text quantized correlated components . In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2046--2054 . Thi Quynh Nhi Tran, Herv\u00e9 Le Borgne, and Michel Crucianu. 2016. Aggregating image and text quantized correlated components. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2046--2054."},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10791-009-9117-9"},{"key":"e_1_2_1_14_1","first-page":"1","article-title":"Learning two-branch neural networks for image-text matching tasks","volume":"99","author":"Wang Liwei","year":"2018","unstructured":"Liwei Wang , Yin Li , Jing Huang , and Svetlana Lazebnik . 2018 . Learning two-branch neural networks for image-text matching tasks . IEEE Transactions on Pattern Analysis and Machine Intelligence PP , 99 (2018), 1 -- 14 . Liwei Wang, Yin Li, Jing Huang, and Svetlana Lazebnik. 2018. Learning two-branch neural networks for image-text matching tasks. IEEE Transactions on Pattern Analysis and Machine Intelligence PP, 99 (2018), 1--14.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence PP"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.327"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/3123266.3123317"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298966"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/2647868.2654902"},{"key":"e_1_2_1_19_1","volume-title":"Proceedings of the International Joint Conferences on Artificial Intelligence. 3846--3853","author":"Peng Yuxin","year":"2016","unstructured":"Yuxin Peng , Xin Huang , and Jinwei Qi . 2016 . Cross-media shared representation by hierarchical learning with multiple deep networks . In Proceedings of the International Joint Conferences on Artificial Intelligence. 3846--3853 . Yuxin Peng, Xin Huang, and Jinwei Qi. 2016. Cross-media shared representation by hierarchical learning with multiple deep networks. In Proceedings of the International Joint Conferences on Artificial Intelligence. 3846--3853."},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2017.2742704"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2016.2558463"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/3123266.3123369"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/3123266.3123326"},{"key":"e_1_2_1_24_1","unstructured":"Xi Zhang Siyu Zhou Jiashi Feng Hanjiang Lai Bo Li Yan Pan Jian Yin and Shuicheng Yan. 2017. HashGAN: Attention-aware deep adversarial hashing for cross modal retrieval. arXiv:1711.09347.  Xi Zhang Siyu Zhou Jiashi Feng Hanjiang Lai Bo Li Yan Pan Jian Yin and Shuicheng Yan. 2017. HashGAN: Attention-aware deep adversarial hashing for cross modal retrieval. arXiv:1711.09347."},{"key":"e_1_2_1_25_1","unstructured":"Karen Simonyan and Andrew Zisserman. 2014. Very deep convolutional networks for large-scale image recognition. arXiv:1409.1556.  Karen Simonyan and Andrew Zisserman. 2014. Very deep convolutional networks for large-scale image recognition. arXiv:1409.1556."},{"key":"e_1_2_1_26_1","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","volume":"6","author":"Sun Baochen","year":"2016","unstructured":"Baochen Sun , Jiashi Feng , and Kate Saenko . 2016 . Return of frustratingly easy domain adaptation . In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 6 . 8. Baochen Sun, Jiashi Feng, and Kate Saenko. 2016. Return of frustratingly easy domain adaptation. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 6. 8."},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1093\/bioinformatics\/btl242"},{"key":"e_1_2_1_28_1","unstructured":"Ian Goodfellow Jean Pouget-Abadie Mehdi Mirza Bing Xu David Warde-Farley Sherjil Ozair Aaron Courville and Yoshua Bengio. 2014. Generative adversarial nets. In Advances in Neural Information Processing Systems. 2672--2680.  Ian Goodfellow Jean Pouget-Abadie Mehdi Mirza Bing Xu David Warde-Farley Sherjil Ozair Aaron Courville and Yoshua Bengio. 2014. Generative adversarial nets. In Advances in Neural Information Processing Systems. 2672--2680."},{"key":"e_1_2_1_29_1","volume-title":"Yoshua Bengio, and Wenjie Li.","author":"Che Tong","year":"2016","unstructured":"Tong Che , Yanran Li , Athul Paul Jacob , Yoshua Bengio, and Wenjie Li. 2016 . Mode regularized generative adversarial networks. arXiv:1612.02136. Tong Che, Yanran Li, Athul Paul Jacob, Yoshua Bengio, and Wenjie Li. 2016. Mode regularized generative adversarial networks. arXiv:1612.02136."},{"key":"e_1_2_1_30_1","unstructured":"Sanjeev Arora Rong Ge Yingyu Liang Tengyu Ma and Yi Zhang. 2017. Generalization and equilibrium in generative adversarial nets (GANs). arXiv:1703.00573.  Sanjeev Arora Rong Ge Yingyu Liang Tengyu Ma and Yi Zhang. 2017. Generalization and equilibrium in generative adversarial nets (GANs). arXiv:1703.00573."},{"key":"e_1_2_1_31_1","unstructured":"Sanjeev Arora and Yi Zhang. 2017. Do GANs actually learn the distribution? arXiv:1706.08224. An empirical study.  Sanjeev Arora and Yi Zhang. 2017. Do GANs actually learn the distribution? arXiv:1706.08224. An empirical study."},{"key":"e_1_2_1_32_1","volume-title":"Jamie Ryan Kiros, and Sanja Fidler","author":"Faghri Fartash","year":"2017","unstructured":"Fartash Faghri , David J. Fleet , Jamie Ryan Kiros, and Sanja Fidler . 2017 . Vse++: Improved visual-semantic embeddings. arXiv:1707.05612. Fartash Faghri, David J. Fleet, Jamie Ryan Kiros, and Sanja Fidler. 2017. Vse++: Improved visual-semantic embeddings. arXiv:1707.05612."},{"key":"e_1_2_1_33_1","volume-title":"Mechanical Turk. In Proceedings of the North American Chapter of the Association for Computational Linguistics Workshop. 139--147","author":"Rashtchian Cyrus","year":"2010","unstructured":"Cyrus Rashtchian , Peter Young , Micah Hodosh , and Julia Hockenmaier . 2010 . Collecting image annotations using Amazon\u2019s Mechanical Turk. In Proceedings of the North American Chapter of the Association for Computational Linguistics Workshop. 139--147 . Cyrus Rashtchian, Peter Young, Micah Hodosh, and Julia Hockenmaier. 2010. Collecting image annotations using Amazon\u2019s Mechanical Turk. In Proceedings of the North American Chapter of the Association for Computational Linguistics Workshop. 139--147."},{"key":"e_1_2_1_34_1","volume-title":"Shared multi-view data representation for multi-domain event detection","author":"Yang Zhenguo","unstructured":"Zhenguo Yang , Qing Li , Wenyin Liu , and Jianming Lv. 2019. Shared multi-view data representation for multi-domain event detection . IEEE Transactions on Pattern Analysis and Machine Intelligence. Epub ahead of print. DOI:10.1109\/TPAMI.2019.2893953 10.1109\/TPAMI.2019.2893953 Zhenguo Yang, Qing Li, Wenyin Liu, and Jianming Lv. 2019. Shared multi-view data representation for multi-domain event detection. IEEE Transactions on Pattern Analysis and Machine Intelligence. Epub ahead of print. DOI:10.1109\/TPAMI.2019.2893953"},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1093\/biomet\/28.3-4.321"},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1162\/0899766042321814"},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298966"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3374754","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3374754","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T22:33:08Z","timestamp":1750199588000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3374754"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,2,17]]},"references-count":37,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2020,2,29]]}},"alternative-id":["10.1145\/3374754"],"URL":"https:\/\/doi.org\/10.1145\/3374754","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,2,17]]},"assertion":[{"value":"2019-05-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2019-12-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2020-02-17","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}