{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,6]],"date-time":"2026-05-06T15:38:40Z","timestamp":1778081920959,"version":"3.51.4"},"publisher-location":"New York, NY, USA","reference-count":45,"publisher":"ACM","license":[{"start":{"date-parts":[[2021,10,17]],"date-time":"2021-10-17T00:00:00Z","timestamp":1634428800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2021,10,17]]},"DOI":"10.1145\/3474085.3475448","type":"proceedings-article","created":{"date-parts":[[2021,10,18]],"date-time":"2021-10-18T05:04:15Z","timestamp":1634533455000},"page":"2862-2870","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":3,"title":["Learning Disentangled Factors from Paired Data in Cross-Modal Retrieval: An Implicit Identifiable VAE Approach"],"prefix":"10.1145","author":[{"given":"Minyoung","family":"Kim","sequence":"first","affiliation":[{"name":"Samsung AI Center Cambridge, Cambridge, United Kingdom"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ricardo","family":"Guerrero","sequence":"additional","affiliation":[{"name":"Samsung AI Center Cambridge, Cambridge, United Kingdom"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Vladimir","family":"Pavlovic","sequence":"additional","affiliation":[{"name":"Samsung AI Center Cambridge, Cambridge, United Kingdom"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2021,10,17]]},"reference":[{"key":"e_1_3_2_2_1_1","doi-asserted-by":"publisher","DOI":"10.5555\/3042817.3043076"},{"key":"e_1_3_2_2_2_1","unstructured":"Abdul Fatir Ansari and Harold Soh. 2018. Hyperprior induced unsupervised disentanglement of latent representations. arXiv:1809.04497.  Abdul Fatir Ansari and Harold Soh. 2018. Hyperprior induced unsupervised disentanglement of latent representations. arXiv:1809.04497."},{"key":"e_1_3_2_2_3_1","volume-title":"Technical Report 688, Department of Statistics","author":"Bach F. R.","year":"2005","unstructured":"F. R. Bach and M. I. Jordan . 2005 . A probabilistic interpretation of canonical correlation analysis. Technical Report 688, Department of Statistics , University of California, Berkeley . F. R. Bach and M. I. Jordan. 2005. A probabilistic interpretation of canonical correlation analysis. Technical Report 688, Department of Statistics, University of California, Berkeley."},{"key":"e_1_3_2_2_4_1","doi-asserted-by":"publisher","DOI":"10.1137\/S0895479895296896"},{"key":"e_1_3_2_2_5_1","unstructured":"Philemon Brakel and Yoshua Bengio. 2017. Learning independent features with adversarial nets for non-linear ICA. In arXiv preprint. https:\/\/arxiv.org\/abs\/1710.05050  Philemon Brakel and Yoshua Bengio. 2017. Learning independent features with adversarial nets for non-linear ICA. In arXiv preprint. https:\/\/arxiv.org\/abs\/1710.05050"},{"key":"e_1_3_2_2_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/2939672.2939812"},{"key":"e_1_3_2_2_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/3123266.3123428"},{"key":"e_1_3_2_2_8_1","volume-title":"International Conference on Multimedia Modeling.","author":"Chen J.","unstructured":"J. Chen , L. Pang , and C. W. Ngo . 2017b. Cross-modal recipe retrieval: How to cook this dish? International Conference on Multimedia Modeling. J. Chen, L. Pang, and C. W. Ngo. 2017b. Cross-modal recipe retrieval: How to cook this dish? International Conference on Multimedia Modeling."},{"key":"e_1_3_2_2_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/3240508.3240627"},{"key":"e_1_3_2_2_10_1","doi-asserted-by":"publisher","DOI":"10.5555\/3327144.3327186"},{"key":"e_1_3_2_2_11_1","doi-asserted-by":"publisher","DOI":"10.5555\/3157096.3157340"},{"key":"e_1_3_2_2_12_1","volume-title":"Thomas Kipf, and Jakub M. Tomczak.","author":"Davidson Tim R.","year":"2018","unstructured":"Tim R. Davidson , Luca Falorsi , Nicola De Cao , Thomas Kipf, and Jakub M. Tomczak. 2018 . Hyperspherical Variational Auto-Encoders. Uncertainty in AI. Tim R. Davidson, Luca Falorsi, Nicola De Cao, Thomas Kipf, and Jakub M. Tomczak. 2018. Hyperspherical Variational Auto-Encoders. Uncertainty in AI."},{"key":"e_1_3_2_2_13_1","volume-title":"Proceedings of the Second International Conference on Learning Representations, ICLR.","author":"Eastwood Cian","unstructured":"Cian Eastwood and Christopher K. I. Williams . 2018. A Framework for the Quantitative Evaluation of Disentangled Representations . In Proceedings of the Second International Conference on Learning Representations, ICLR. Cian Eastwood and Christopher K. I. Williams. 2018. A Framework for the Quantitative Evaluation of Disentangled Representations. In Proceedings of the Second International Conference on Learning Representations, ICLR."},{"key":"e_1_3_2_2_14_1","unstructured":"Babak Esmaeili Hao Wu Sarthak Jain Alican Bozkurt N. Siddharth Brooks Paige Dana H. Brooks Jennifer Dy and Jan-Willem van de Meent. 2018. Structured disentangled representations. arXiv:1804.02086v4.  Babak Esmaeili Hao Wu Sarthak Jain Alican Bozkurt N. Siddharth Brooks Paige Dana H. Brooks Jennifer Dy and Jan-Willem van de Meent. 2018. Structured disentangled representations. arXiv:1804.02086v4."},{"key":"e_1_3_2_2_15_1","volume-title":"International Conference on Learning Representations.","author":"Higgins Irina","year":"2017","unstructured":"Irina Higgins , Loic Matthey , Arka Pal , Christopher Burgess , Xavier Glorot , Matthew Botvinick , Shakir Mohamed , and Alexander Lerchner . 2017 a. Beta-VAE: Learning basic visual concepts with a constrained variational framework .. In International Conference on Learning Representations. Irina Higgins, Loic Matthey, Arka Pal, Christopher Burgess, Xavier Glorot, Matthew Botvinick, Shakir Mohamed, and Alexander Lerchner. 2017a. Beta-VAE: Learning basic visual concepts with a constrained variational framework.. In International Conference on Learning Representations."},{"key":"e_1_3_2_2_16_1","volume-title":"SCAN: Learning Hierarchical Compositional Visual Concepts. arXiv:1707.03389.","author":"Higgins Irina","year":"2017","unstructured":"Irina Higgins , Nicolas Sonnerat , Loic Matthey , Arka Pal , Christopher P Burgess , Matko Bosnjak , Murray Shanahan , Matthew Botvinick , Demis Hassabis , and Alexander Lerchner . 2017 b. SCAN: Learning Hierarchical Compositional Visual Concepts. arXiv:1707.03389. Irina Higgins, Nicolas Sonnerat, Loic Matthey, Arka Pal, Christopher P Burgess, Matko Bosnjak, Murray Shanahan, Matthew Botvinick, Demis Hassabis, and Alexander Lerchner. 2017b. SCAN: Learning Hierarchical Compositional Visual Concepts. arXiv:1707.03389."},{"key":"e_1_3_2_2_17_1","unstructured":"Wei-Ning Hsu and James Glass. 2018. Disentangling by partitioning: A representation learning framework for multimodal sensory data. arXiv:1805.11264.  Wei-Ning Hsu and James Glass. 2018. Disentangling by partitioning: A representation learning framework for multimodal sensory data. arXiv:1805.11264."},{"key":"e_1_3_2_2_18_1","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition.","author":"Jiang Q.","unstructured":"Q. Jiang and W. Li . 2017. Deep cross-modal hashing . IEEE Conference on Computer Vision and Pattern Recognition. Q. Jiang and W. Li. 2017. Deep cross-modal hashing. IEEE Conference on Computer Vision and Pattern Recognition."},{"key":"e_1_3_2_2_19_1","volume-title":"Ricardo Pio Monti, and Aapo Hyv\"arinen","author":"Khemakhem Ilyes","year":"2020","unstructured":"Ilyes Khemakhem , Diederik P. Kingma , Ricardo Pio Monti, and Aapo Hyv\"arinen . 2020 . Variational Autoencoders and Nonlinear ICA: A Unifying Framework. Artificial Intelligence and Statistics . Ilyes Khemakhem, Diederik P. Kingma, Ricardo Pio Monti, and Aapo Hyv\"arinen. 2020. Variational Autoencoders and Nonlinear ICA: A Unifying Framework. Artificial Intelligence and Statistics."},{"key":"e_1_3_2_2_20_1","volume-title":"International Conference on Machine Learning.","author":"Kim Hyunjik","year":"2018","unstructured":"Hyunjik Kim and Andriy Mnih . 2018 . Disentangling by Factorising. (2018) . International Conference on Machine Learning. Hyunjik Kim and Andriy Mnih. 2018. Disentangling by Factorising. (2018). International Conference on Machine Learning."},{"key":"e_1_3_2_2_21_1","volume-title":"Bayes-Factor-VAE: Hierarchical Bayesian Deep Auto-Encoder Models for Factor Disentanglement. In IEEE International Conference on Computer Vision, ICCV.","author":"Kim Minyoung","year":"2019","unstructured":"Minyoung Kim , Yuting Wang , Pritish Sahu , and Vladimir Pavlovic . 2019 a . Bayes-Factor-VAE: Hierarchical Bayesian Deep Auto-Encoder Models for Factor Disentanglement. In IEEE International Conference on Computer Vision, ICCV. Minyoung Kim, Yuting Wang, Pritish Sahu, and Vladimir Pavlovic. 2019 a. Bayes-Factor-VAE: Hierarchical Bayesian Deep Auto-Encoder Models for Factor Disentanglement. In IEEE International Conference on Computer Vision, ICCV."},{"key":"e_1_3_2_2_22_1","unstructured":"Minyoung Kim Yuting Wang Pritish Sahu and Vladimir Pavlovic. 2019 b. Relevance Factor VAE: Learning and Identifying Disentangled Factors. https:\/\/arxiv.org\/abs\/1902.01568 arXiv:1902.01568.  Minyoung Kim Yuting Wang Pritish Sahu and Vladimir Pavlovic. 2019 b. Relevance Factor VAE: Learning and Identifying Disentangled Factors. https:\/\/arxiv.org\/abs\/1902.01568 arXiv:1902.01568."},{"key":"e_1_3_2_2_23_1","volume-title":"Proceedings of the Second International Conference on Learning Representations, ICLR.","author":"Diederik","unstructured":"Diederik P. Kingma and Max Welling. 2014. Auto-Encoding Variational Bayes . In Proceedings of the Second International Conference on Learning Representations, ICLR. Diederik P. Kingma and Max Welling. 2014. Auto-Encoding Variational Bayes. In Proceedings of the Second International Conference on Learning Representations, ICLR."},{"key":"e_1_3_2_2_24_1","volume-title":"International Conference on Learning Representations.","author":"Kumar Abhishek","year":"2018","unstructured":"Abhishek Kumar , Prasanna Sattigeri , and Avinash Balakrishnan . 2018 . Variation inference of disentangled latent concepts from unlabeled observations . In International Conference on Learning Representations. Abhishek Kumar, Prasanna Sattigeri, and Avinash Balakrishnan. 2018. Variation inference of disentangled latent concepts from unlabeled observations. In International Conference on Learning Representations."},{"key":"e_1_3_2_2_25_1","doi-asserted-by":"publisher","DOI":"10.1109\/5.726791"},{"key":"e_1_3_2_2_26_1","volume-title":"Challenging Common Assumptions in the Unsupervised Learning of Disentangled Representations. International Conference on Machine Learning.","author":"Locatello Francesco","year":"2019","unstructured":"Francesco Locatello , Stefan Bauer , Mario Lucic , Gunnar R\"atsch, Sylvain Gelly , Bernhard Sch\u00f6lkopf , and Olivier Bachem . 2019 . Challenging Common Assumptions in the Unsupervised Learning of Disentangled Representations. International Conference on Machine Learning. Francesco Locatello, Stefan Bauer, Mario Lucic, Gunnar R\"atsch, Sylvain Gelly, Bernhard Sch\u00f6lkopf, and Olivier Bachem. 2019. Challenging Common Assumptions in the Unsupervised Learning of Disentangled Representations. International Conference on Machine Learning."},{"key":"e_1_3_2_2_27_1","volume-title":"Adversarial Autoencoders. In International Conference on Learning Representations. http:\/\/arxiv.org\/abs\/1511","author":"Makhzani Alireza","year":"2016","unstructured":"Alireza Makhzani , Jonathon Shlens , Navdeep Jaitly , and Ian Goodfellow . 2016 . Adversarial Autoencoders. In International Conference on Learning Representations. http:\/\/arxiv.org\/abs\/1511 .05644 Alireza Makhzani, Jonathon Shlens, Navdeep Jaitly, and Ian Goodfellow. 2016. Adversarial Autoencoders. In International Conference on Learning Representations. http:\/\/arxiv.org\/abs\/1511.05644"},{"key":"e_1_3_2_2_28_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2019.2927476"},{"key":"e_1_3_2_2_29_1","volume-title":"International Conference on Machine Learning.","author":"Mathieu E.","unstructured":"E. Mathieu , T. Rainforth , N. Siddharth , and Y. W. Teh . 2019. Disentangling disentanglement in variational autoencoders . International Conference on Machine Learning. E. Mathieu, T. Rainforth, N. Siddharth, and Y. W. Teh. 2019. Disentangling disentanglement in variational autoencoders. International Conference on Machine Learning."},{"key":"e_1_3_2_2_30_1","doi-asserted-by":"publisher","DOI":"10.5555\/3157382.3157661"},{"key":"e_1_3_2_2_31_1","unstructured":"Loic Matthey Irina Higgins Demis Hassabis and Alexander Lerchner. 2017. dSprites: Disentanglement testing Sprites dataset. https:\/\/github.com\/deepmind\/dsprites-dataset\/  Loic Matthey Irina Higgins Demis Hassabis and Alexander Lerchner. 2017. dSprites: Disentanglement testing Sprites dataset. https:\/\/github.com\/deepmind\/dsprites-dataset\/"},{"key":"e_1_3_2_2_32_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2017.2742704"},{"key":"e_1_3_2_2_33_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.327"},{"key":"e_1_3_2_2_34_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICRA.2017.7989160"},{"key":"e_1_3_2_2_35_1","doi-asserted-by":"publisher","DOI":"10.5555\/2969442.2969628"},{"key":"e_1_3_2_2_36_1","volume-title":"Deep Joint-Semantics Reconstructing Hashing for Large-Scale Unsupervised Cross-Modal Retrieval. In The IEEE International Conference on Computer Vision (ICCV).","author":"Su Shupeng","year":"2019","unstructured":"Shupeng Su , Zhisheng Zhong , and Chao Zhang . 2019 . Deep Joint-Semantics Reconstructing Hashing for Large-Scale Unsupervised Cross-Modal Retrieval. In The IEEE International Conference on Computer Vision (ICCV). Shupeng Su, Zhisheng Zhong, and Chao Zhang. 2019. Deep Joint-Semantics Reconstructing Hashing for Large-Scale Unsupervised Cross-Modal Retrieval. In The IEEE International Conference on Computer Vision (ICCV)."},{"key":"e_1_3_2_2_37_1","volume-title":"Tomczak and Max Welling","author":"Jakub","year":"2018","unstructured":"Jakub M. Tomczak and Max Welling . 2018 . VAE with a VampPrior. Artificial Intelligence and Statistics . Jakub M. Tomczak and Max Welling. 2018. VAE with a VampPrior. Artificial Intelligence and Statistics."},{"key":"e_1_3_2_2_38_1","volume-title":"Amir Zadeh, Louis-Philippe Morency, and Ruslan Salakhutdinov.","author":"Hubert Tsai Yao-Hung","year":"2018","unstructured":"Yao-Hung Hubert Tsai , Paul Pu Liang , Amir Zadeh, Louis-Philippe Morency, and Ruslan Salakhutdinov. 2018 . Learning Factorized Multimodal Representations . arXiv:1806.06176. Yao-Hung Hubert Tsai, Paul Pu Liang, Amir Zadeh, Louis-Philippe Morency, and Ruslan Salakhutdinov. 2018. Learning Factorized Multimodal Representations. arXiv:1806.06176."},{"key":"e_1_3_2_2_39_1","unstructured":"Ramakrishna Vedantam Ian Fischer Jonathan Huang and Kevin Murphy. 2017. Generative models of visually grounded imagination. arXiv:1705.10762.  Ramakrishna Vedantam Ian Fischer Jonathan Huang and Kevin Murphy. 2017. Generative models of visually grounded imagination. arXiv:1705.10762."},{"key":"e_1_3_2_2_40_1","doi-asserted-by":"publisher","DOI":"10.1145\/3123266.3123326"},{"key":"e_1_3_2_2_41_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.01184"},{"key":"e_1_3_2_2_42_1","doi-asserted-by":"crossref","unstructured":"Weiran Wang Xinchen Yan Honglak Lee and Karen Livescu. 2016. Deep Variational Canonical Correlation Analysis. arXiv:1610.03454.  Weiran Wang Xinchen Yan Honglak Lee and Karen Livescu. 2016. Deep Variational Canonical Correlation Analysis. arXiv:1610.03454.","DOI":"10.21437\/Interspeech.2017-1581"},{"key":"e_1_3_2_2_43_1","doi-asserted-by":"publisher","DOI":"10.5555\/3327345.3327461"},{"key":"e_1_3_2_2_44_1","doi-asserted-by":"crossref","unstructured":"F. Zheng Y. Tang and L. Shao. 2018. Hetero-manifold regularisation for cross-modal hashing. IEEE transactions on pattern analysis and machine intelligence Vol. 40 5 (2018) 1059--1071.  F. Zheng Y. Tang and L. Shao. 2018. Hetero-manifold regularisation for cross-modal hashing. IEEE transactions on pattern analysis and machine intelligence Vol. 40 5 (2018) 1059--1071.","DOI":"10.1109\/TPAMI.2016.2645565"},{"key":"e_1_3_2_2_45_1","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition.","author":"Zhu B.","unstructured":"B. Zhu , C. W. Ngo , J. Chen , and Y. Hao . 2019. R2GAN: Cross-modal recipe retrieval with generative adversarial network . IEEE Conference on Computer Vision and Pattern Recognition. B. Zhu, C. W. Ngo, J. Chen, and Y. Hao. 2019. R2GAN: Cross-modal recipe retrieval with generative adversarial network. IEEE Conference on Computer Vision and Pattern Recognition."}],"event":{"name":"MM '21: ACM Multimedia Conference","location":"Virtual Event China","acronym":"MM '21","sponsor":["SIGMM ACM Special Interest Group on Multimedia"]},"container-title":["Proceedings of the 29th ACM International Conference on Multimedia"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3474085.3475448","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3474085.3475448","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T20:48:33Z","timestamp":1750193313000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3474085.3475448"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,10,17]]},"references-count":45,"alternative-id":["10.1145\/3474085.3475448","10.1145\/3474085"],"URL":"https:\/\/doi.org\/10.1145\/3474085.3475448","relation":{},"subject":[],"published":{"date-parts":[[2021,10,17]]},"assertion":[{"value":"2021-10-17","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}