{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,30]],"date-time":"2026-01-30T04:27:35Z","timestamp":1769747255080,"version":"3.49.0"},"reference-count":32,"publisher":"Association for Computing Machinery (ACM)","issue":"1s","license":[{"start":{"date-parts":[[2015,10,21]],"date-time":"2015-10-21T00:00:00Z","timestamp":1445385600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"National High Technology Research and Development Program of China","award":["2012AA011103"],"award-info":[{"award-number":["2012AA011103"]}]},{"name":"discipline building plan in 111 base","award":["B08004"],"award-info":[{"award-number":["B08004"]}]},{"DOI":"10.13039\/501100012226","name":"Fundamental Research Funds for the Central Universities","doi-asserted-by":"crossref","award":["2013RC0304"],"award-info":[{"award-number":["2013RC0304"]}],"id":[{"id":"10.13039\/501100012226","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["61273365"],"award-info":[{"award-number":["61273365"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2015,10,21]]},"abstract":"<jats:p>This article considers the problem of cross-modal retrieval, such as using a text query to search for images and vice-versa. Based on different autoencoders, several novel models are proposed here for solving this problem. These models are constructed by correlating hidden representations of a pair of autoencoders. A novel optimal objective, which minimizes a linear combination of the representation learning errors for each modality and the correlation learning error between hidden representations of two modalities, is used to train the model as a whole. Minimizing the correlation learning error forces the model to learn hidden representations with only common information in different modalities, while minimizing the representation learning error makes hidden representations good enough to reconstruct inputs of each modality. To balance the two kind of errors induced by representation learning and correlation learning, we set a specific parameter in our models. Furthermore, according to the modalities the models attempt to reconstruct they are divided into two groups. One group including three models is named multimodal reconstruction correspondence autoencoder since it reconstructs both modalities. The other group including two models is named unimodal reconstruction correspondence autoencoder since it reconstructs a single modality. The proposed models are evaluated on three publicly available datasets. And our experiments demonstrate that our proposed correspondence autoencoders perform significantly better than three canonical correlation analysis based models and two popular multimodal deep models on cross-modal retrieval tasks.<\/jats:p>","DOI":"10.1145\/2808205","type":"journal-article","created":{"date-parts":[[2015,10,24]],"date-time":"2015-10-24T18:27:12Z","timestamp":1445711232000},"page":"1-22","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":23,"title":["Correspondence Autoencoders for Cross-Modal Retrieval"],"prefix":"10.1145","volume":"12","author":[{"given":"Fangxiang","family":"Feng","sequence":"first","affiliation":[{"name":"Beijing University of Posts and Telecommunications, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xiaojie","family":"Wang","sequence":"additional","affiliation":[{"name":"Beijing University of Posts and Telecommunications, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ruifan","family":"Li","sequence":"additional","affiliation":[{"name":"Beijing University of Posts and Telecommunications, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ibrar","family":"Ahmad","sequence":"additional","affiliation":[{"name":"Beijing University of Posts and Telecommunications and University of Peshawar, Pakistan"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2015,10,21]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1109\/MMUL.2009.74"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1561\/2200000006"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.3115\/1225403.1225421"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1145\/860435.860460"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.5555\/944919.944937"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2007.4409066"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/1646396.1646452"},{"key":"e_1_2_1_8_1","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV'10)","author":"Farhadi Ali"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/2647868.2654902"},{"key":"e_1_2_1_10_1","unstructured":"Andrea Frome Greg Corrado Jon Shlens Samy Bengio Jeffrey Dean Marc'Aurelio Ranzato and Tomas Mikolov. 2013. DeViSE: A deep visual-semantic embedding model. In Neural Information Processing Systems (NIPS'13) 2121--2129.  Andrea Frome Greg Corrado Jon Shlens Samy Bengio Jeffrey Dean Marc'Aurelio Ranzato and Tomas Mikolov. 2013. DeViSE: A deep visual-semantic embedding model. In Neural Information Processing Systems (NIPS'13) 2121--2129."},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1162\/0899766042321814"},{"key":"e_1_2_1_12_1","doi-asserted-by":"crossref","unstructured":"G. Hinton and R. Salakhutdinov. 2006. Reducing the dimensionality of data with neural networks. Science 313 5786 504--507.  G. Hinton and R. Salakhutdinov. 2006. Reducing the dimensionality of data with neural networks. Science 313 5786 504--507.","DOI":"10.1126\/science.1127647"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1162\/089976602760128018"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1162\/neco.2006.18.7.1527"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2011.6126524"},{"key":"e_1_2_1_16_1","volume-title":"Proceedings of the 25th International Conference on Computational Linguistics (COLING'12)","author":"Kim Jungi","year":"2012"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1109\/76.927424"},{"key":"e_1_2_1_18_1","volume-title":"Proceedings of the 28th International Conference on Machine Learning (ICML'11)","author":"Ngiam J."},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1023\/A:1011139631724"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/1873951.1873987"},{"key":"e_1_2_1_21_1","unstructured":"R. Salakhutdinov and G. Hinton. 2009. Replicated Softmax: an Undirected Topic Model. In Neural Information Processing Systems (NIPS'09) 1607--1614.  R. Salakhutdinov and G. Hinton. 2009. Replicated Softmax: an Undirected Topic Model. In Neural Information Processing Systems (NIPS'09) 1607--1614."},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1162\/NECO_a_00311"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/P14-1068"},{"key":"e_1_2_1_24_1","unstructured":"P. Smolensky. 1986. Parallel Distributed Processing: Explorations in the Microstructure of Cognition Vol. 1. MIT Press Cambridge MA Chapter Information processing in dynamical systems: foundations of harmony theory 194--281.   P. Smolensky. 1986. Parallel Distributed Processing: Explorations in the Microstructure of Cognition Vol. 1. MIT Press Cambridge MA Chapter Information processing in dynamical systems: foundations of harmony theory 194--281."},{"key":"e_1_2_1_25_1","unstructured":"Richard Socher Milind Ganjoo Christopher D. Manning and Andrew Ng. 2013. Zero-shot learning through cross-modal transfer. In Neural Information Processing Systems (NIPS'13) 935--943.  Richard Socher Milind Ganjoo Christopher D. Manning and Andrew Ng. 2013. Zero-shot learning through cross-modal transfer. In Neural Information Processing Systems (NIPS'13) 935--943."},{"key":"e_1_2_1_26_1","volume-title":"Proceedings of the International Conference on Machine Learning Representation Learning Workshop.","author":"Srivastava N."},{"key":"e_1_2_1_27_1","unstructured":"N. Srivastava and R. Salakhutdinov. 2012b. Multimodal learning with deep Boltzmann machines. In Neural Information Processing Systems (NIPS'12) 2231--2239.  N. Srivastava and R. Salakhutdinov. 2012b. Multimodal learning with deep Boltzmann machines. In Neural Information Processing Systems (NIPS'12) 2231--2239."},{"key":"e_1_2_1_28_1","first-page":"2579","article-title":"Visualizing High-Dimensional Data Using t-SNE","volume":"9","author":"van der Maaten L. J. P.","year":"2008","journal-title":"J. Machine Learn. Res."},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/1873951.1874249"},{"key":"e_1_2_1_30_1","unstructured":"M. Welling M. Rosen-Zvi and G. Hinton. 2004. Exponential family harmoniums with an application to information retrieval. In Neural Information Processing Systems (NIPS'04) 501--508.  M. Welling M. Rosen-Zvi and G. Hinton. 2004. Exponential family harmoniums with an application to information retrieval. In Neural Information Processing Systems (NIPS'04) 501--508."},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10994-010-5198-3"},{"key":"e_1_2_1_32_1","volume-title":"Proceedings of the 27th AAAI Conference on Artificial Intelligence (AAAI'13)","author":"Zhuang Yueting","year":"2013"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2808205","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/2808205","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T06:12:40Z","timestamp":1750227160000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2808205"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2015,10,21]]},"references-count":32,"journal-issue":{"issue":"1s","published-print":{"date-parts":[[2015,10,21]]}},"alternative-id":["10.1145\/2808205"],"URL":"https:\/\/doi.org\/10.1145\/2808205","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2015,10,21]]},"assertion":[{"value":"2015-01-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2015-07-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2015-10-21","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}