{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,13]],"date-time":"2026-07-13T19:54:09Z","timestamp":1783972449405,"version":"3.55.0"},"reference-count":76,"publisher":"MDPI AG","issue":"2","license":[{"start":{"date-parts":[[2022,1,27]],"date-time":"2022-01-27T00:00:00Z","timestamp":1643241600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Future Internet"],"abstract":"<jats:p>Cross-modal retrieval aims to search samples of one modality via queries of other modalities, which is a hot issue in the community of multimedia. However, two main challenges, i.e., heterogeneity gap and semantic interaction across different modalities, have not been solved efficaciously. Reducing the heterogeneous gap can improve the cross-modal similarity measurement. Meanwhile, modeling cross-modal semantic interaction can capture the semantic correlations more accurately. To this end, this paper presents a novel end-to-end framework, called Dual Attention Generative Adversarial Network (DA-GAN). This technique is an adversarial semantic representation model with a dual attention mechanism, i.e., intra-modal attention and inter-modal attention. Intra-modal attention is used to focus on the important semantic feature within a modality, while inter-modal attention is to explore the semantic interaction between different modalities and then represent the high-level semantic correlation more precisely. A dual adversarial learning strategy is designed to generate modality-invariant representations, which can reduce the cross-modal heterogeneity efficiently. The experiments on three commonly used benchmarks show the better performance of DA-GAN than these competitors.<\/jats:p>","DOI":"10.3390\/fi14020043","type":"journal-article","created":{"date-parts":[[2022,1,27]],"date-time":"2022-01-27T21:59:55Z","timestamp":1643320795000},"page":"43","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":13,"title":["DA-GAN: Dual Attention Generative Adversarial Network for Cross-Modal Retrieval"],"prefix":"10.3390","volume":"14","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-0928-4866","authenticated-orcid":false,"given":"Liewu","family":"Cai","sequence":"first","affiliation":[{"name":"College of Information and Intelligence, Hunan Agricultural University, Changsha 410128, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4569-1429","authenticated-orcid":false,"given":"Lei","family":"Zhu","sequence":"additional","affiliation":[{"name":"College of Information and Intelligence, Hunan Agricultural University, Changsha 410128, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Hongyan","family":"Zhang","sequence":"additional","affiliation":[{"name":"College of Information and Intelligence, Hunan Agricultural University, Changsha 410128, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xinghui","family":"Zhu","sequence":"additional","affiliation":[{"name":"College of Information and Intelligence, Hunan Agricultural University, Changsha 410128, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2022,1,27]]},"reference":[{"key":"ref_1","first-page":"1","article-title":"Survey on deep multi-modal data analytics: Collaboration, rivalry, and fusion","volume":"17","author":"Wang","year":"2021","journal-title":"ACM Trans. Multimed. Comput. Commun. Appl."},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Ranjan, V., Rasiwasia, N., and Jawahar, C.V. (2015, January 7\u201313). Multi-label cross-modal retrieval. Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.466"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Chen, Y., Ren, P., Wang, Y., and de Rijke, M. (2019, January 21\u201325). Bayesian personalized feature interaction selection for factorization machines. Proceedings of the 42nd International ACM SIGIR Conference on Research and Development in Information Retrieval, Paris, France.","DOI":"10.1145\/3331184.3331196"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Wu, Y., and Yang, Y. (2021, January 20\u201325). Exploring Heterogeneous Clues for Weakly-Supervised Audio-Visual Video Parsing. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.00138"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"103589","DOI":"10.1016\/j.artint.2021.103589","article-title":"Bayesian feature interaction selection for factorization machines","volume":"302","author":"Chen","year":"2022","journal-title":"Artif. Intell."},{"key":"ref_6","first-page":"1","article-title":"Multi-graph heterogeneous interaction fusion for social recommendation","volume":"40","author":"Zhang","year":"2021","journal-title":"ACM Trans. Inf. Syst."},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Gu, C., Bu, J., Zhou, X., Yao, C., Ma, D., Yu, Z., and Yan, X. (2021). Cross-modal Image Retrieval with Deep Mutual Information Maximization. arXiv.","DOI":"10.1016\/j.neucom.2022.01.078"},{"key":"ref_8","first-page":"1","article-title":"Hcmsl: Hybrid cross-modal similarity learning for cross-modal retrieval","volume":"17","author":"Zhang","year":"2021","journal-title":"ACM Trans. Multimed. Comput. Commun. Appl."},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Zhang, C., Zhong, Z., Zhu, L., Zhang, S., Cao, D., and Zhang, J. (2021, January 21\u201324). M2guda: Multi-metrics graph-based unsupervised domain adaptation for cross-modal Hashing. Proceedings of the 2021 International Conference on Multimedia Retrieval, Taipei, Taiwan.","DOI":"10.1145\/3460426.3463670"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Thomas, C., and Kovashka, A. (2020). Preserving semantic neighborhoods for robust cross-modal retrieval. European Conference on Computer Vision, Springer.","DOI":"10.1007\/978-3-030-58523-5_19"},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"2639","DOI":"10.1162\/0899766042321814","article-title":"Canonical correlation analysis: An overview with application to learning methods","volume":"16","author":"Hardoon","year":"2004","journal-title":"Neural Comput."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"521","DOI":"10.1109\/TPAMI.2013.142","article-title":"On the role of correlation and abstraction in cross-modal multimedia retrieval","volume":"36","author":"Pereira","year":"2014","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"210","DOI":"10.1007\/s11263-013-0658-4","article-title":"A multi-view embedding space for modeling internet images, tags, and their semantics","volume":"106","author":"Gong","year":"2014","journal-title":"Int. Comput. Vis."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Sharma, A., Kumar, A., Daume, H., and Jacobs, D.W. (2012, January 16\u201321). Generalized multiview analysis: A discriminative latent space. Proceedings of the 2012 IEEE Conference on Computer Vision and Pattern Recognition, Providence, RI, USA.","DOI":"10.1109\/CVPR.2012.6247923"},{"key":"ref_15","unstructured":"Rasiwasia, N., Mahajan, D., Mahadevan, V., and Aggarwal, G. (2014, January 22\u201325). Cluster canonical correlation analysis. Proceedings of the Seventeenth International Conference on Artificial Intelligence and Statistics, AISTATS 2014, Reykjavik, Iceland."},{"key":"ref_16","unstructured":"Lopez-Paz, D., Sra, S., Smola, A., Ghahramani, Z., and Sch\u00f6lkopf, B. (2014, January 21\u201326). Randomized nonlinear component analysis. Proceedings of the International Conference on Machine Learning, Beijing, China."},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"531","DOI":"10.1016\/j.imavis.2006.04.014","article-title":"Locality preserving cca with applications to data visualization and pose estimation","volume":"25","author":"Sun","year":"2007","journal-title":"Image Vis. Comput."},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"436","DOI":"10.1038\/nature14539","article-title":"Deep learning","volume":"521","author":"LeCun","year":"2015","journal-title":"Nature"},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"85","DOI":"10.1016\/j.neunet.2014.09.003","article-title":"Deep learning in neural networks: An overview","volume":"61","author":"Schmidhuber","year":"2015","journal-title":"Neural Netw."},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"1393","DOI":"10.1109\/TIP.2017.2655449","article-title":"Effective multi-query expansions: Collaborative deep networks for robust landmark retrieval","volume":"26","author":"Wang","year":"2017","journal-title":"IEEE Trans. Image Process."},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"4894","DOI":"10.1109\/TIP.2021.3076275","article-title":"Diversifying inference path selection: Moving-mobile-network for landmark recognition","volume":"30","author":"Qian","year":"2021","journal-title":"IEEE Trans. Image Process."},{"key":"ref_22","unstructured":"Benton, A., Khayrallah, H., Gujral, B., Reisinger, D., Zhang, S., and Arora, R. (2017). Deep generalized canonical correlation analysis. arXiv."},{"key":"ref_23","unstructured":"Elmadany, N.E.D., He, Y., and Guan, L. (2016, January 20\u201325). Multiview learning via deep discriminative canonical correlation analysis. Proceedings of the 2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Shanghai, China."},{"key":"ref_24","unstructured":"Wang, W., Arora, R., Livescu, K., and Bilmes, J.A. (2015, January 6\u201311). On deep multi-view representation learning. Proceedings of the 32nd International Conference on Machine Learning, Lille, France."},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"5585","DOI":"10.1109\/TIP.2018.2852503","article-title":"Modality-specific cross-modal similarity measurement with recurrent attention network","volume":"27","author":"Peng","year":"2018","journal-title":"IEEE Trans. Image Process."},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"5412","DOI":"10.1109\/TNNLS.2020.2967597","article-title":"Cross-modal attention with semantic consistence for image-text matching","volume":"31","author":"Xu","year":"2020","journal-title":"IEEE Trans. Neural Netw. Learn. Syst."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Fang, A., Zhao, X., and Zhang, Y. (2019). Cross-modal image fusion theory guided by subjective visual attention. arXiv.","DOI":"10.1016\/j.neucom.2020.07.014"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Zhu, L., Zhang, C., Song, J., Liu, L., Zhang, S., and Li, Y. (2021, January 5\u20139). Multi-graph based hierarchical semantic fusion for cross-modal representation. Proceedings of the 2021 IEEE International Conference on Multimedia and Expo (ICME), Shenzhen, China.","DOI":"10.1109\/ICME51207.2021.9428194"},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"722","DOI":"10.1109\/TNNLS.2020.2979190","article-title":"Deep coattention-based comparator for relative representation learning in person re-identification","volume":"32","author":"Wu","year":"2020","journal-title":"IEEE Trans. Neural Netw. Learn."},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"2010","DOI":"10.1109\/TPAMI.2015.2505311","article-title":"Joint feature selection and subspace learning for cross-modal retrieval","volume":"38","author":"Wang","year":"2015","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"180571","DOI":"10.1109\/ACCESS.2019.2940055","article-title":"An efficient approach for geo-multimedia cross-modal retrieval","volume":"7","author":"Zhu","year":"2019","journal-title":"IEEE Access"},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"79","DOI":"10.1109\/MMUL.2020.3015764","article-title":"Adversarial learning-based semantic correlation representation for cross-modal retrieval","volume":"27","author":"Zhu","year":"2020","journal-title":"IEEE Multimed."},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Wang, C., Yang, H., and Meinel, C. (2015, January 9\u201311). Deep semantic mapping for cross-modal retrieval. Proceedings of the 2015 IEEE 27th International Conference on Tools with Artificial Intelligence (ICTAI), Vietri sul Mare, Italy.","DOI":"10.1109\/ICTAI.2015.45"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Rasiwasia, N., Pereira, J.C., Coviello, E., Doyle, G., Lanckriet, G.R., Levy, R., and Vasconcelos, N. (2010, January 25\u201329). A new approach to cross-modal multimedia Retrieval. Proceedings of the 18th ACM International Conference on Multimedia, Firenze, Italy.","DOI":"10.1145\/1873951.1873987"},{"key":"ref_35","unstructured":"Ngiam, J., Khosla, A., Kim, M., Nam, J., Lee, H., and Ng, A.Y. (July, January 28). Multimodal deep learning. Proceedings of the 28th International Conference on Machine Learning, Bellevue, WA, USA."},{"key":"ref_36","doi-asserted-by":"crossref","first-page":"1602","DOI":"10.1109\/TIP.2018.2878970","article-title":"Cycle-consistent deep generative hashing for cross-modal retrieval","volume":"28","author":"Wu","year":"2018","journal-title":"IEEE Trans. Image Process."},{"key":"ref_37","first-page":"1","article-title":"Deep semantic mapping for heterogeneous multimedia transfer learning using co-occurrence data","volume":"15","author":"Zhao","year":"2019","journal-title":"Acm Trans. Multimed. Comput. Commun. Appl."},{"key":"ref_38","unstructured":"Wang, Y., Zhang, W., Wu, L., Lin, X., Fang, M., and Pan, S. (2016, January 9\u201315). Iterative views agreement: An iterative low-rank based structured optimization method to multi-view spectral clustering. Proceedings of the Twenty-Fifth International Joint Conference on Artificial Intelligence, New York, NY, USA."},{"key":"ref_39","first-page":"1","article-title":"Deep learning\u2013based multimedia analytics: A review","volume":"15","author":"Zhang","year":"2019","journal-title":"ACM Trans. Multimed. Comput. Appl."},{"key":"ref_40","first-page":"449","article-title":"Cross-modal retrieval with cnn visual features: A new baseline","volume":"47","author":"Wei","year":"2016","journal-title":"IEEE Trans. Cybern."},{"key":"ref_41","first-page":"1","article-title":"Caesar: Concept augmentation based semantic representation for cross-modal retrieval","volume":"1","author":"Zhu","year":"2020","journal-title":"Multimed. Tools Appl."},{"key":"ref_42","unstructured":"Andrew, G., Arora, R., Bilmes, J.A., and Livescu, K. (2013, January 16\u201321). Deep canonical correlation Analysis. Proceedings of the 30th International Conference on Machine Learning, Atlanta, GA, USA."},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Gu, J., Cai, J., Joty, S.R., Niu, L., and Wang, G. (2018, January 18\u201323). Look, imagine and match: Improving textual-visual cross-modal retrieval with generative models. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00750"},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"Zhen, L., Hu, P., Wang, X., and Peng, D. (2019, January 16\u201320). Deep supervised cross-modal Retrieval. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.01064"},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Gao, P., Jiang, Z., You, H., Lu, P., Hoi, S.C.H., Wang, X., and Li, H. (2019, January 16\u201320). Dynamic fusion with intra- and inter-modality attention flow for visual question answering. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00680"},{"key":"ref_46","unstructured":"Xu, K., Ba, J., Kiros, R., Cho, K., Courville, A., Salakhudinov, R., Zemel, R., and Bengio, Y. (2015, January 6\u201311). Show, attend and tell: Neural image caption generation with visual attention. Proceedings of the 32nd International Conference on Machine Learning, Lille, France."},{"key":"ref_47","doi-asserted-by":"crossref","unstructured":"Liu, J., Wang, G., Hu, P., Duan, L.-Y., and Kot, A.C. (2017, January 21\u201326). Global context-aware attention lstm networks for 3d action recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.391"},{"key":"ref_48","unstructured":"Xiao, T., Xu, Y., Yang, K., Zhang, J., Peng, Y., and Zhang, Z. (2015, January 7\u201312). The application of two-level attention models in deep convolutional neural network for fine-grained image classification. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA."},{"key":"ref_49","unstructured":"Lu, J., Yang, J., Batra, D., and Parikh, D. Hierarchical question-image co-attention for visual question answering. In Advances in Neural Information Processing Systems; 2016; pp. 289\u2013297."},{"key":"ref_50","doi-asserted-by":"crossref","first-page":"1791","DOI":"10.1109\/TCYB.2018.2813971","article-title":"Deep attention-based spatially recursive networks for fine-grained visual recognition","volume":"49","author":"Wu","year":"2019","journal-title":"IEEE Trans. Cybern."},{"key":"ref_51","doi-asserted-by":"crossref","unstructured":"Sudhakaran, S., Escalera, S., and Lanz, O. (2019, January 15\u201320). Lsta: Long short-term attention for egocentric action recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.01019"},{"key":"ref_52","doi-asserted-by":"crossref","unstructured":"Wang, X., Wang, Y.-F., and Wang, W.Y. (2018, January 1\u20136). Watch, listen, and describe: Globally and locally aligned cross-modal attentions for video captioning. Proceedings of the NAACL-HLT, New Orleans, LA, USA.","DOI":"10.18653\/v1\/N18-2125"},{"key":"ref_53","doi-asserted-by":"crossref","unstructured":"Liu, X., Wang, Z., Shao, J., Wang, X., and Li, H. (2019, January 15\u201320). Improving referring expression grounding with cross-modal attention-guided erasing. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00205"},{"key":"ref_54","doi-asserted-by":"crossref","unstructured":"Huang, P.-Y., Chang, X., and Hauptmann, A.G. (2019, January 10\u201313). Improving what cross-modal retrieval models learn through object-oriented inter-and intra-modal attention networks. Proceedings of the 2019 on International Conference on Multimedia Retrieval, Ottawa, ON, Canada.","DOI":"10.1145\/3323873.3325043"},{"key":"ref_55","unstructured":"Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y. (2014). Generative Adversarial Nets. Advances in Neural Information Processing Systems, Available online: https:\/\/proceedings.neurips.cc\/paper\/2014\/file\/5ca3e9b122f61f8f06494c97b1afccf3-Paper.pdf."},{"key":"ref_56","doi-asserted-by":"crossref","first-page":"657","DOI":"10.1007\/s11280-018-0541-x","article-title":"Deep adversarial metric learning for cross-modal retrieval","volume":"22","author":"Xu","year":"2019","journal-title":"World Wide Web"},{"key":"ref_57","unstructured":"Liu, Q., Lienhart, R., Wang, H., Chen, S.K., Boll, S., Chen, Y.P., Friedland, G., Li, J., and Yan, S. (2017). Adversarial cross-modal Retrieval. Proceedings of the 2017 ACM on Multimedia Conference, Mountain View, CA, USA, 23\u201327 October 2017, ACM."},{"key":"ref_58","first-page":"1","article-title":"Modality-invariant image-text embedding for image-sentence matching","volume":"15","author":"Liu","year":"2019","journal-title":"ACM Trans. Multimed. Comput. Commun. Appl."},{"key":"ref_59","doi-asserted-by":"crossref","first-page":"1047","DOI":"10.1109\/TCYB.2018.2879846","article-title":"MHTN: Modal-adversarial hybrid transfer network for cross-modal retrieval","volume":"50","author":"Huang","year":"2020","journal-title":"IEEE Trans. Cybern."},{"key":"ref_60","doi-asserted-by":"crossref","first-page":"1059","DOI":"10.1109\/TPAMI.2016.2645565","article-title":"Hetero-manifold regularisation for cross-modal hashing","volume":"40","author":"Zheng","year":"2018","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_61","unstructured":"Baeza-Yates, R., Lalmas, M., Moffat, A., and Ribeiro-Neto, B.A. (2015). LBMCH: Learning bridging mapping for cross-modal hashing. Proceedings of the 38th International ACM SIGIR Conference on Research and Development in Information Retrieval, Santiago, Chile, 9\u201313 August 2015, ACM."},{"key":"ref_62","doi-asserted-by":"crossref","first-page":"489","DOI":"10.1109\/TCYB.2018.2868826","article-title":"SCH-GAN: Semi-supervised cross-modal hashing by generative adversarial network","volume":"50","author":"Zhang","year":"2020","journal-title":"IEEE Trans. Cybern."},{"key":"ref_63","doi-asserted-by":"crossref","unstructured":"Graves, A., Mohamed, A., and Hinton, G.E. (2013, January 26\u201331). Speech recognition with deep recurrent neural networks. Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing, Vancouver, BC, Canada.","DOI":"10.1109\/ICASSP.2013.6638947"},{"key":"ref_64","unstructured":"Krizhevsky, A., Sutskever, I., and Hinton, G.E. (2012, January 3\u20136). Imagenet classification with deep convolutional neural networks. Proceedings of the Advances in Neural Information Processing Systems 25: 26th Annual Conference on Neural Information Processing Systems 2012, Lake Tahoe, NV, USA."},{"key":"ref_65","first-page":"2493","article-title":"Natural language processing (almost) from scratch","volume":"12","author":"Collobert","year":"2011","journal-title":"J. Mach. Learn. Res."},{"key":"ref_66","unstructured":"Kingma, D.P., and Ba, J. (2015, January 7\u20139). Adam: A method for stochastic optimization. Proceedings of the 3rd International Conference on Learning Representations, San Diego, CA, USA."},{"key":"ref_67","doi-asserted-by":"crossref","unstructured":"Chua, T.-S., Tang, J., Hong, R., Li, H., Luo, Z., and Zheng, Y. (2009). Nus-wide: A real-world web image database from national university of Singapore. Proceedings of the ACM International Conference on Image and Video Retrieval, Fira, Greece, 8\u201310 July 2009, ACM.","DOI":"10.1145\/1646396.1646452"},{"key":"ref_68","unstructured":"Rashtchian, C., Young, P., Hodosh, M., and Hockenmaier, J. (2010). Collecting image annotations using amazon\u2019s mechanical turk. Proceedings of the NAACL HLT 2010 Workshop on Creating Speech and Language Data with Amazon\u2019s Mechanical Turk, Los Angeles, CA, USA, 6 June 2010, Association for Computational Linguistics."},{"key":"ref_69","doi-asserted-by":"crossref","first-page":"321","DOI":"10.1093\/biomet\/28.3-4.321","article-title":"Relations between two sets of variates","volume":"28","author":"Hotelling","year":"1936","journal-title":"Biometrika"},{"key":"ref_70","unstructured":"Rupnik, J., and Shawe-Taylor, J. (2010, January 12). Multi-view canonical correlation analysis. Proceedings of the Conference on Data Mining and Data Warehouses (SiKDD 2010), Ljubljana, Slovenia."},{"key":"ref_71","first-page":"808","article-title":"Multi-view discriminant Analysis","volume":"7572","author":"Fitzgibbon","year":"2012","journal-title":"Proceedings of the Computer Vision\u2014ECCV 2012\u201412th European Conference on Computer Vision, Florence, Italy, 7\u201313 October 2012"},{"key":"ref_72","doi-asserted-by":"crossref","first-page":"188","DOI":"10.1109\/TPAMI.2015.2435740","article-title":"Multi-view discriminant analysis","volume":"38","author":"Kan","year":"2016","journal-title":"IEEE Trans. Pattern Anal. Machine Intell."},{"key":"ref_73","doi-asserted-by":"crossref","first-page":"965","DOI":"10.1109\/TCSVT.2013.2276704","article-title":"Learning cross-media joint representation with sparse and semisupervised regularization","volume":"24","author":"Zhai","year":"2014","journal-title":"IEEE Trans. Circuits Syst. Video Techn."},{"key":"ref_74","doi-asserted-by":"crossref","first-page":"405","DOI":"10.1109\/TMM.2017.2742704","article-title":"CCL: Cross-modal correlation learning with multigrained fusion by hierarchical network","volume":"20","author":"Peng","year":"2018","journal-title":"IEEE Trans. Multimed."},{"key":"ref_75","unstructured":"Peng, Y., Huang, X., and Qi, J. (2016, January 9\u201315). Cross-media shared representation by hierarchical learning with multiple deep networks. Proceedings of the Twenty-Fifth International Joint Conference on Artificial Intelligence, New York, NY, USA."},{"key":"ref_76","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3284750","article-title":"Cm-gans: Cross-modal generative adversarial networks for common representation learning","volume":"15","author":"Peng","year":"2019","journal-title":"ACM Trans. Multimed. Comput. Commun. Appl."}],"container-title":["Future Internet"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1999-5903\/14\/2\/43\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T22:09:30Z","timestamp":1760134170000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1999-5903\/14\/2\/43"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,1,27]]},"references-count":76,"journal-issue":{"issue":"2","published-online":{"date-parts":[[2022,2]]}},"alternative-id":["fi14020043"],"URL":"https:\/\/doi.org\/10.3390\/fi14020043","relation":{},"ISSN":["1999-5903"],"issn-type":[{"value":"1999-5903","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,1,27]]}}}