{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,12]],"date-time":"2025-10-12T03:17:29Z","timestamp":1760239049644,"version":"build-2065373602"},"reference-count":52,"publisher":"MDPI AG","issue":"9","license":[{"start":{"date-parts":[[2020,9,13]],"date-time":"2020-09-13T00:00:00Z","timestamp":1599955200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"National High-level Personnel for Defense Technology Program","award":["2017-JCJQ-ZQ-013"],"award-info":[{"award-number":["2017-JCJQ-ZQ-013"]}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["61902405"],"award-info":[{"award-number":["61902405"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"name":"HUNAN Province Science Foundation","award":["2017RS3045"],"award-info":[{"award-number":["2017RS3045"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Symmetry"],"abstract":"<jats:p>Multimodal representations play an important role in multimodal learning tasks, including cross-modal retrieval and intra-modal clustering. However, existing multimodal representation learning approaches focus on building one common space by aligning different modalities and ignore the complementary information across the modalities, such as the intra-modal local structures. In other words, they only focus on the object-level alignment and ignore structure-level alignment. To tackle the problem, we propose a novel symmetric multimodal representation learning framework by transferring local structures across different modalities, namely MTLS. A customized soft metric learning strategy and an iterative parameter learning process are designed to symmetrically transfer local structures and enhance the cluster structures in intra-modal representations. The bidirectional retrieval loss based on multi-layer neural networks is utilized to align two modalities. MTLS is instantiated with image and text data and shows its superior performance on image-text retrieval and image clustering. MTLS outperforms the state-of-the-art multimodal learning methods by up to 32% in terms of R@1 on text-image retrieval and 16.4% in terms of AMI onclustering.<\/jats:p>","DOI":"10.3390\/sym12091504","type":"journal-article","created":{"date-parts":[[2020,9,13]],"date-time":"2020-09-13T22:01:01Z","timestamp":1600034461000},"page":"1504","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Learning Multimodal Representations by Symmetrically Transferring Local Structures"],"prefix":"10.3390","volume":"12","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-1869-3784","authenticated-orcid":false,"given":"Bin","family":"Dong","sequence":"first","affiliation":[{"name":"College of Computer, National University of Defense Technology, Changsha 410000, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Songlei","family":"Jian","sequence":"additional","affiliation":[{"name":"College of Computer, National University of Defense Technology, Changsha 410000, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Kai","family":"Lu","sequence":"additional","affiliation":[{"name":"College of Computer, National University of Defense Technology, Changsha 410000, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2020,9,13]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Vinyals, O., Toshev, A., Bengio, S., and Erhan, D. (2015, January 7\u201312). Show and tell: A neural image caption generator. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298935"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Johnson, J., Karpathy, A., and Li, F.F. (2016, January 27\u201330). Densecap: Fully convolutional localization networks for dense captioning. Proceedings of the CVPR, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.494"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Antol, S., Agrawal, A., Lu, J., Mitchell, M., Batra, D., Lawrence, Z., and Parikh, D. (2015, January 7\u201313). Vqa: Visual question answering. In Proceedings of the ICCV, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.279"},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"32","DOI":"10.1007\/s11263-016-0981-7","article-title":"Visual genome: Connecting language and vision using crowdsourced dense image annotations","volume":"123","author":"Krishna","year":"2017","journal-title":"Int. J. Comput. Vis."},{"key":"ref_5","unstructured":"Ngiam, J., Khosla, A., Kim, M., Nam, J., Lee, H., and Ng, A.Y. (2011, January 2). Multimodal deep learning. Proceedings of the ICML-11, New Brunswick, NJ, USA."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"1798","DOI":"10.1109\/TPAMI.2013.50","article-title":"Representation learning: A review and new perspectives","volume":"35","author":"Bengio","year":"2013","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_7","first-page":"423","article-title":"Multimodal machine learning: A survey and taxonomy","volume":"41","author":"Ahuja","year":"2018","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Jian, S., Hu, L., Cao, L., and Lu, K. (2020, January 7\u201312). Representation Learning with Multiple Lipschitz-Constrained Alignments on Partially-Labeled Cross-Domain Data. Proceedings of the AAAI, Hilton New York Midtown, New York, NY, USA.","DOI":"10.1609\/aaai.v34i04.5856"},{"key":"ref_9","unstructured":"Jian, S., Hu, L., Cao, L., Gao, H., and Lu, K. (February, January 27). Evolutionarily learning multi-aspect interactions and influences from network structure and node content. Proceedings of the AAAI Conference on Artificial Intelligence. Hilton Hawaiian Village, Honolulu, HI, USA."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Silberer, C., and Lapata, M. (2014, January 23\u201325). Learning grounded meaning representations with autoencoders. Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Baltimore, MD, USA.","DOI":"10.3115\/v1\/P14-1068"},{"key":"ref_11","unstructured":"Frome, A., Corrado, G.S., Shlens, J., Bengio, S., Dean, J., Ranzato, M.A., and Mikolov, T. (2013, January 5\u20138). Devise: A deep visual-semantic embedding model. Proceedings of the NIPS, Harrahs and Harveys, Lake Tahoe, CA, USA."},{"key":"ref_12","unstructured":"Kiros, R., Salakhutdinov, R., and Zemel, R. (2014). Unifying visual-semantic embeddings with multimodal neural language models. arXiv."},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"394","DOI":"10.1109\/TPAMI.2018.2797921","article-title":"Learning two-branch neural networks for image-text matching tasks","volume":"41","author":"Wang","year":"2018","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Ouyang, W., Chu, X., and Wang, X. (2014, January 23\u201328). Multi-source deep learning for human pose estimation. Proceedings of the CVPR, Columbus, OH, USA.","DOI":"10.1109\/CVPR.2014.299"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Zhang, H., Hu, Z., Deng, Y., Sachan, M., Yan, Z., and Xing, E. (2016, January 7\u201312). Learning Concept Taxonomies from Multi-modal Data. Proceedings of the ACL, Berlin, Germany.","DOI":"10.18653\/v1\/P16-1169"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Zhang, J., Peng, Y., and Yuan, M. (2018, January 2\u20137). Unsupervised Generative Adversarial Cross-modal Hashing. Proceedings of the AAAI, New Orleans, LA, USA.","DOI":"10.1609\/aaai.v32i1.11263"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Zhang, Y., and Lu, H. (2018, January 8\u201314). Deep cross-modal projection learning for image-text matching. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01246-5_42"},{"key":"ref_18","unstructured":"Weston, J., Bengio, S., and Usunier, N. (2011, January 16\u201322). Wsabie: Scaling up to large vocabulary image annotation. Proceedings of the Twenty-Second International Joint Conference on Artificial Intelligence, Catalonia, Spain."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Wang, T., Xu, X., Yang, Y., Hanjalic, A., Shen, H.T., and Song, J. (2019, January 21\u201325). Matching Images and Text with Multi-modal Tensor Fusion and Re-ranking. Proceedings of the 27th ACM International Conference on Multimedia, Nice, France.","DOI":"10.1145\/3343031.3350875"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Huang, Y., Wang, W., and Wang, L. (2017, January 21\u201326). Instance-aware image and sentence matching with selective multimodal lstm. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.767"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Li, S., Xiao, T., Li, H., Yang, W., and Wang, X. (2017, January 22\u201329). Identity-aware textual-visual matching with latent co-attention. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.209"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Fukui, A., Park, D.H., Yang, D., Rohrbach, A., Darrell, T., and Rohrbach, M. (2016). Multimodal compact bilinear pooling for visual question answering and visual grounding. arXiv.","DOI":"10.18653\/v1\/D16-1044"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Jabri, A., Joulin, A., and Van Der Maaten, L. (2016, January 8\u201316). Revisiting visual question answering baselines. Proceedings of the European Conference on Computer Vision, Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46484-8_44"},{"key":"ref_24","unstructured":"Vendrov, I., Kiros, R., Fidler, S., and Urtasun, R. (2015). Order-embeddings of images and language. arXiv."},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"365","DOI":"10.1142\/S012906570000034X","article-title":"Kernel and nonlinear canonical correlation analysis","volume":"10","author":"Lai","year":"2000","journal-title":"Int. J. Neural Syst."},{"key":"ref_26","unstructured":"Andrew, G., Arora, R., Bilmes, J., and Livescu, K. (2013, January 16\u201321). Deep canonical correlation analysis. Proceedings of the International Conference on Machine Learning, Atlanta, GA, USA."},{"key":"ref_27","unstructured":"Klein, B., Lev, G., Sadeh, G., and Wolf, L. (2014). Fisher vectors derived from hybrid gaussian-laplacian mixture models for image annotation. arXiv."},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Rasiwasia, N., Costa Pereira, J., Coviello, E., Doyle, G., Lanckriet, G.R., Levy, R., and Vasconcelos, N. (2010, January 25\u201329). A new approach to cross-modal multimedia retrieval. Proceedings of the 18th ACM international Conference on Multimedia, Firenze, Italy.","DOI":"10.1145\/1873951.1873987"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Lee, K.H., Chen, X., Hua, G., Hu, H., and He, X. (2018, January 8\u201314). Stacked cross attention for image-text matching. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01225-0_13"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Liu, C., Mao, Z., Liu, A.A., Zhang, T., Wang, B., and Zhang, Y. (2019, January 21\u201325). Focus Your Attention: A Bidirectional Focal Attention Network for Image-Text Matching. Proceedings of the 27th ACM International Conference on Multimedia, Nice, France.","DOI":"10.1145\/3343031.3350869"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Xu, R., Li, C., Yan, J., Deng, C., and Liu, X. (2019, January 10\u201316). Graph Convolutional Network Hashing for Cross-Modal Retrieval. Proceedings of the IJCAI, Macao, China.","DOI":"10.24963\/ijcai.2019\/138"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Shi, Y., You, X., Zheng, F., Wang, S., and Peng, Q. (2019, January 10\u201316). Equally-Guided Discriminative Hashing for Cross-modal Retrieval. Proceedings of the IJCAI, Macao, China.","DOI":"10.24963\/ijcai.2019\/662"},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"93","DOI":"10.1016\/j.neucom.2019.04.041","article-title":"Adversarial cross-modal retrieval based on dictionary learning","volume":"355","author":"Shang","year":"2019","journal-title":"Neurocomputing"},{"key":"ref_34","doi-asserted-by":"crossref","first-page":"2675","DOI":"10.1109\/TMM.2019.2903448","article-title":"Cross-modality bridging and knowledge transferring for image understanding","volume":"21","author":"Yan","year":"2019","journal-title":"IEEE Trans. Multimed."},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"84","DOI":"10.1016\/j.neucom.2019.12.058","article-title":"Unsupervised deep cross-modal hashing with virtual label regression","volume":"386","author":"Wang","year":"2020","journal-title":"Neurocomputing"},{"key":"ref_36","unstructured":"Simonyan, K., and Zisserman, A. (2014). Very deep convolutional networks for large-scale image recognition. arXiv."},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep residual learning for image recognition. Proceedings of the CVPR, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Cho, K., Van Merri\u00ebnboer, B., Gulcehre, C., Bahdanau, D., Bougares, F., Schwenk, H., and Bengio, Y. (2014). Learning phrase representations using RNN encoder-decoder for statistical machine translation. arXiv.","DOI":"10.3115\/v1\/D14-1179"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Frome, A., Singer, Y., Sha, F., and Malik, J. (2007, January 20). Learning globally-consistent local distance functions for shape-based image retrieval and classification. Proceedings of the ICCV, Rio de Janeiro, Brazil.","DOI":"10.1109\/ICCV.2007.4408839"},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Jian, S., Hu, L., Cao, L., and Lu, K. (2018, January 2\u20137). Metric-based auto-instructor for learning mixed data representation. Proceedings of the Thirty-Second AAAI Conference on Artificial Intelligence, Hilton New Orleans Riverside, New Orleans, LA, USA.","DOI":"10.1609\/aaai.v32i1.11597"},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"LeCun, Y., Chopra, S., Hadsell, R., Ranzato, M., and Huang, F. (2006). A Tutorial on Energy-Based Learning, MIT Press.","DOI":"10.7551\/mitpress\/7443.003.0014"},{"key":"ref_42","doi-asserted-by":"crossref","first-page":"1627","DOI":"10.1109\/TPAMI.2009.167","article-title":"Object detection with discriminatively trained part-based models","volume":"32","author":"Felzenszwalb","year":"2009","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Yu, C.N.J., and Joachims, T. (2009, January 14\u201318). Learning structural SVMs with latent variables. Proceedings of the ICML, Montreal, QC, Canada.","DOI":"10.1145\/1553374.1553523"},{"key":"ref_44","unstructured":"Faghri, F., Fleet, D.J., Kiros, J.R., and Fidler, S. (2017). Vse++: Improving visual-semantic embeddings with hard negatives. arXiv."},{"key":"ref_45","unstructured":"Kingma, D., and Ba, J. (2014). Adam: A method for stochastic optimization. arXiv."},{"key":"ref_46","doi-asserted-by":"crossref","first-page":"67","DOI":"10.1162\/tacl_a_00166","article-title":"From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions","volume":"2","author":"Young","year":"2014","journal-title":"Trans. Assoc. Comput. Linguist."},{"key":"ref_47","doi-asserted-by":"crossref","unstructured":"Lin, T., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Doll\u00e1r, P., and Zitnick, C. (2014, January 6\u201312). Microsoft coco: Common objects in context. Proceedings of the ECCV, Zurich, Switzerland.","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"ref_48","doi-asserted-by":"crossref","unstructured":"Karpathy, A., and Li, F. (2015, January 7\u201312). Deep visual-semantic alignments for generating image descriptions. Proceedings of the CVPR, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298932"},{"key":"ref_49","doi-asserted-by":"crossref","first-page":"853","DOI":"10.1613\/jair.3994","article-title":"Framing image description as a ranking task: Data, models and evaluation metrics","volume":"47","author":"Hodosh","year":"2013","journal-title":"J. Artif. Intell. Res."},{"key":"ref_50","doi-asserted-by":"crossref","unstructured":"Zhong, Z., Zheng, L., Cao, D., and Li, S. (2017, January 21\u201326). Re-ranking person re-identification with k-reciprocal encoding. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.389"},{"key":"ref_51","doi-asserted-by":"crossref","unstructured":"Ramirez, E.H., Brena, R., Magatti, D., and Stella, F. (September, January 31). Probabilistic metrics for soft-clustering and topic model validation. Proceedings of the 2010 IEEE\/WIC\/ACM International Conference on Web Intelligence and Intelligent Agent Technology, Toronto, ON, Canada.","DOI":"10.1109\/WI-IAT.2010.148"},{"key":"ref_52","unstructured":"Romano, S., Bailey, J., Nguyen, V., and Verspoor, K. (2014, January 21\u201326). Standardized mutual information for clustering comparisons: One step further in adjustment for chance. Proceedings of the International Conference on Machine Learning, Beijing, China."}],"container-title":["Symmetry"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2073-8994\/12\/9\/1504\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T10:09:32Z","timestamp":1760177372000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2073-8994\/12\/9\/1504"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,9,13]]},"references-count":52,"journal-issue":{"issue":"9","published-online":{"date-parts":[[2020,9]]}},"alternative-id":["sym12091504"],"URL":"https:\/\/doi.org\/10.3390\/sym12091504","relation":{},"ISSN":["2073-8994"],"issn-type":[{"type":"electronic","value":"2073-8994"}],"subject":[],"published":{"date-parts":[[2020,9,13]]}}}