{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,16]],"date-time":"2026-07-16T02:33:01Z","timestamp":1784169181937,"version":"3.55.0"},"reference-count":52,"publisher":"Association for Computing Machinery (ACM)","issue":"1","license":[{"start":{"date-parts":[[2023,1,5]],"date-time":"2023-01-05T00:00:00Z","timestamp":1672876800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/100012542","name":"Science and Technology Project of Sichuan","doi-asserted-by":"crossref","award":["2022ZHCG0033, 2021YFG0314, and 2020YFG0459"],"award-info":[{"award-number":["2022ZHCG0033, 2021YFG0314, and 2020YFG0459"]}],"id":[{"id":"10.13039\/100012542","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["U19A2078"],"award-info":[{"award-number":["U19A2078"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2023,1,31]]},"abstract":"<jats:p>Zero-shot learning (ZSL) aims to recognize image instances of unseen classes solely based on the semantic descriptions of the unseen classes. In this field, Generalized Zero-Shot Learning (GZSL) is a challenging problem in which the images of both seen and unseen classes are mixed in the testing phase of learning. Existing methods formulate GZSL as a semantic-visual correspondence problem and apply generative models such as Generative Adversarial Networks and Variational Autoencoders to solve the problem. However, these methods suffer from the bias problem since the images of unseen classes are often misclassified into seen classes. In this work, a novel model named the Dual Projective model for Zero-Shot Learning (DPZSL) is proposed using text descriptions. In order to alleviate the bias problem, we leverage two autoencoders to project the visual and semantic features into a latent space and evaluate the embeddings by a visual-semantic correspondence loss function. An additional novel classifier is also introduced to ensure the discriminability of the embedded features. Our method focuses on a more challenging inductive ZSL setting in which only the labeled data from seen classes are used in the training phase. The experimental results, obtained from two popular datasets\u2014Caltech-UCSD Birds-200-2011 (CUB) and North America Birds (NAB)\u2014show that the proposed DPZSL model significantly outperforms both the inductive ZSL and GZSL settings. Particularly in the GZSL setting, our model yields an improvement up to 15.2% in comparison with state-of-the-art CANZSL on datasets CUB and NAB with two splittings.<\/jats:p>","DOI":"10.1145\/3514247","type":"journal-article","created":{"date-parts":[[2022,7,29]],"date-time":"2022-07-29T11:49:08Z","timestamp":1659095348000},"page":"1-17","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":9,"title":["Dual Projective Zero-Shot Learning Using Text Descriptions"],"prefix":"10.1145","volume":"19","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-5433-7379","authenticated-orcid":false,"given":"Yunbo","family":"Rao","sequence":"first","affiliation":[{"name":"University of Electronic Science and Technology of China, Chengdu, Sichuan, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9761-2705","authenticated-orcid":false,"given":"Ziqiang","family":"Yang","sequence":"additional","affiliation":[{"name":"University of Electronic Science and Technology of China, Chengdu, Sichuan, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4384-8787","authenticated-orcid":false,"given":"Shaoning","family":"Zeng","sequence":"additional","affiliation":[{"name":"University of Electronic Science and Technology of China, Chengdu, Sichuan, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7476-0190","authenticated-orcid":false,"given":"Qifeng","family":"Wang","sequence":"additional","affiliation":[{"name":"Google Berkeley, Berkeley, California, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4284-6958","authenticated-orcid":false,"given":"Jiansu","family":"Pu","sequence":"additional","affiliation":[{"name":"University of Electronic Science and Technology of China, Chengdu, Sichuan, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2023,1,5]]},"reference":[{"key":"e_1_3_1_2_2","first-page":"59","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Akata Zeynep","year":"2016","unstructured":"Zeynep Akata, Mateusz Malinowski, Mario Fritz, and Bernt Schiele. 2016. Multi-cue zero-shot learning with strong supervision. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. IEEE, Las Vegas, Nevada, USA, 59\u201368."},{"issue":"7","key":"e_1_3_1_3_2","doi-asserted-by":"crossref","first-page":"1425","DOI":"10.1109\/TPAMI.2015.2487986","article-title":"Label-embedding for image classification","volume":"38","author":"Akata Zeynep","year":"2015","unstructured":"Zeynep Akata, Florent Perronnin, Zaid Harchaoui, and Cordelia Schmid. 2015. Label-embedding for image classification. IEEE Transactions on Pattern Analysis and Machine Intelligence 38, 7 (2015), 1425\u20131438.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298911"},{"key":"e_1_3_1_5_2","first-page":"7603","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Annadani Yashas","year":"2018","unstructured":"Yashas Annadani and Soma Biswas. 2018. Preserving semantic relations for zero-shot learning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. IEEE, Salt Lake City, Utah, USA, 7603\u20137612."},{"key":"e_1_3_1_6_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2016.2644615"},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.575"},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-019-01193-1"},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46475-6_4"},{"key":"e_1_3_1_10_2","first-page":"874","volume-title":"Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision","author":"Chen Zhi","year":"2020","unstructured":"Zhi Chen, Jingjing Li, Yadan Luo, Zi Huang, and Yang Yang. 2020. CANZSL: Cycle-consistent adversarial networks for zero-shot learning from natural language. In Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision. IEEE, Snowmass village, Colorado, 874\u2013883."},{"key":"e_1_3_1_11_2","volume-title":"International Conference on Learning Representations","author":"Chou Yu-Ying","year":"2020","unstructured":"Yu-Ying Chou, Hsuan-Tien Lin, and Tyng-Luh Liu. 2020. Adaptive and generative zero-shot learning. In International Conference on Learning Representations. IEEE, Vienna, Austria."},{"issue":"1","key":"e_1_3_1_12_2","first-page":"198","article-title":"General knowledge embedded image representation learning","volume":"20","author":"Cui Peng","year":"2017","unstructured":"Peng Cui, Shaowei Liu, and Wenwu Zhu. 2017. General knowledge embedded image representation learning. IEEE Transactions on Multimedia 20, 1 (2017), 198\u2013207.","journal-title":"IEEE Transactions on Multimedia"},{"key":"e_1_3_1_13_2","first-page":"5784","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Elhoseiny Mohamed","year":"2019","unstructured":"Mohamed Elhoseiny and Mohamed Elfeki. 2019. Creativity inspired zero-shot learning. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. IEEE, Seoul, Korea (South), 5784\u20135793."},{"issue":"12","key":"e_1_3_1_14_2","doi-asserted-by":"crossref","first-page":"2539","DOI":"10.1109\/TPAMI.2016.2643667","article-title":"Write a classifier: Predicting visual classifiers from unstructured text","volume":"39","author":"Elhoseiny Mohamed","year":"2016","unstructured":"Mohamed Elhoseiny, Ahmed Elgammal, and Babak Saleh. 2016. Write a classifier: Predicting visual classifiers from unstructured text. IEEE Transactions on Pattern Analysis and Machine Intelligence 39, 12 (2016), 2539\u20132553.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"e_1_3_1_15_2","first-page":"5640","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Elhoseiny Mohamed","year":"2017","unstructured":"Mohamed Elhoseiny, Yizhe Zhu, Han Zhang, and Ahmed Elgammal. 2017. Link the head to the \u201cbeak\u201d: Zero shot learning from noisy text description at part precision. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. IEEE, Honolulu, Hawaii, USA, 5640\u20135649."},{"key":"e_1_3_1_16_2","article-title":"DeVISE: A deep visual-semantic embedding model","volume":"26","author":"Frome Andrea","year":"2013","unstructured":"Andrea Frome, Greg S. Corrado, Jon Shlens, Samy Bengio, Jeff Dean, Marc\u2019Aurelio Ranzato, and Tomas Mikolov. 2013. DeVISE: A deep visual-semantic embedding model. Advances in Neural Information Processing Systems 26 (2013).","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.169"},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1145\/3394171.3413593"},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2017.2671414"},{"key":"e_1_3_1_20_2","first-page":"801","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Huang He","year":"2019","unstructured":"He Huang, Changhu Wang, Philip S. Yu, and Chang-Dong Wang. 2019. Generative dual adversarial network for generalized zero-shot learning. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. IEEE, Long Beach, California, 801\u2013810."},{"key":"e_1_3_1_21_2","unstructured":"Diederik P. Kingma and Jimmy Ba. 2015. Adam: A method for stochastic optimization. International Conference on Learning Representations (ICLR\u201915) . San Diego CA arXiv preprint arXiv:1412.6980 9 (2015)."},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.282"},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.473"},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2013.140"},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1038\/nature14539"},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.1145\/3343031.3350901"},{"key":"e_1_3_1_27_2","first-page":"3583","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Li Kai","year":"2019","unstructured":"Kai Li, Martin Renqiang Min, and Yun Fu. 2019. Rethinking zero-shot learning: A conditional visual classification perspective. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. IEEE, Seoul, South Korea, 3583\u20133592."},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00518"},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW.2018.00294"},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.220"},{"key":"e_1_3_1_31_2","first-page":"2189","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Pal Arghya","year":"2019","unstructured":"Arghya Pal and Vineeth N. Balasubramanian. 2019. Zero-shot task transfer. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. IEEE, Long Beach, California, USA, 2189\u20132198."},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.1145\/3284750"},{"key":"e_1_3_1_33_2","first-page":"2249","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Qiao Ruizhi","year":"2016","unstructured":"Ruizhi Qiao, Lingqiao Liu, Chunhua Shen, and Anton Van Den Hengel. 2016. Less is more: Zero-shot learning from online textual documents with noise suppression. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. IEEE, Las Vegas, Nevada, USA, 2249\u20132257."},{"issue":"1","key":"e_1_3_1_34_2","doi-asserted-by":"crossref","first-page":"242","DOI":"10.1109\/TMM.2019.2924511","article-title":"Deep0tag: Deep multiple instance learning for zero-shot image tagging","volume":"22","author":"Rahman Shafin","year":"2019","unstructured":"Shafin Rahman, Salman Khan, and Nick Barnes. 2019. Deep0tag: Deep multiple instance learning for zero-shot image tagging. IEEE Transactions on Multimedia 22, 1 (2019), 242\u2013255.","journal-title":"IEEE Transactions on Multimedia"},{"key":"e_1_3_1_35_2","first-page":"2152","volume-title":"International Conference on Machine Learning","author":"Romera-Paredes Bernardino","year":"2015","unstructured":"Bernardino Romera-Paredes and Philip Torr. 2015. An embarrassingly simple approach to zero-shot learning. In International Conference on Machine Learning. JMLR.org, Lille, France, 2152\u20132161."},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","DOI":"10.1016\/0306-4573(88)90021-0"},{"key":"e_1_3_1_37_2","first-page":"8247","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Schonfeld Edgar","year":"2019","unstructured":"Edgar Schonfeld, Sayna Ebrahimi, Samarth Sinha, Trevor Darrell, and Zeynep Akata. 2019. Generalized zero- and few-shot learning via aligned variational autoencoders. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. IEEE, Long Beach, CA, USA, 8247\u20138255."},{"key":"e_1_3_1_38_2","article-title":"Very deep convolutional networks for large-scale image recognition","author":"Simonyan Karen","year":"2014","unstructured":"Karen Simonyan and Andrew Zisserman. 2014. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556 (2014).","journal-title":"arXiv preprint arXiv:1409.1556"},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00113"},{"issue":"11","key":"e_1_3_1_40_2","article-title":"Visualizing data using t-SNE.","volume":"9","author":"Maaten Laurens Van der","year":"2008","unstructured":"Laurens Van der Maaten and Geoffrey Hinton. 2008. Visualizing data using t-SNE.Journal of Machine Learning Research 9, 11 (2008).","journal-title":"Journal of Machine Learning Research"},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298658"},{"key":"e_1_3_1_42_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i04.6069"},{"key":"e_1_3_1_43_2","unstructured":"Catherine Wah Steve Branson Peter Welinder Pietro Perona and Serge Belongie. 2011. The Caltech-UCSD birds-200-2011 dataset. (2011)."},{"issue":"2","key":"e_1_3_1_44_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3293318","article-title":"A survey of zero-shot learning: Settings, methods, and applications","volume":"10","author":"Wang Wei","year":"2019","unstructured":"Wei Wang, Vincent W. Zheng, Han Yu, and Chunyan Miao. 2019. A survey of zero-shot learning: Settings, methods, and applications. ACM Transactions on Intelligent Systems and Technology 10, 2 (2019), 1\u201337.","journal-title":"ACM Transactions on Intelligent Systems and Technology"},{"issue":"2","key":"e_1_3_1_45_2","doi-asserted-by":"crossref","first-page":"260","DOI":"10.1109\/TMM.2015.2505083","article-title":"Zero-shot person re-identification via cross-view consistency","volume":"18","author":"Wang Zheng","year":"2015","unstructured":"Zheng Wang, Ruimin Hu, Chao Liang, Yi Yu, Junjun Jiang, Mang Ye, Jun Chen, and Qingming Leng. 2015. Zero-shot person re-identification via cross-view consistency. IEEE Transactions on Multimedia 18, 2 (2015), 260\u2013272.","journal-title":"IEEE Transactions on Multimedia"},{"key":"e_1_3_1_46_2","first-page":"478","volume-title":"International Conference on Machine Learning","author":"Xie Junyuan","year":"2016","unstructured":"Junyuan Xie, Ross Girshick, and Ali Farhadi. 2016. Unsupervised deep embedding for clustering analysis. In International Conference on Machine Learning. PMLR, New York, USA, 478\u2013487."},{"issue":"9","key":"e_1_3_1_47_2","doi-asserted-by":"crossref","first-page":"2387","DOI":"10.1109\/TMM.2019.2898777","article-title":"Adversarially approximated autoencoder for image generation and manipulation","volume":"21","author":"Xu Wenju","year":"2019","unstructured":"Wenju Xu, Shawn Keshmiri, and Guanghui Wang. 2019. Adversarially approximated autoencoder for image generation and manipulation. IEEE Transactions on Multimedia 21, 9 (2019), 2387\u20132396.","journal-title":"IEEE Transactions on Multimedia"},{"key":"e_1_3_1_48_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46493-0_47"},{"key":"e_1_3_1_49_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2020.2984666"},{"key":"e_1_3_1_50_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.321"},{"key":"e_1_3_1_51_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.474"},{"key":"e_1_3_1_52_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.649"},{"key":"e_1_3_1_53_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00111"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3514247","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3514247","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T18:10:14Z","timestamp":1750183814000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3514247"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,1,5]]},"references-count":52,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2023,1,31]]}},"alternative-id":["10.1145\/3514247"],"URL":"https:\/\/doi.org\/10.1145\/3514247","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,1,5]]},"assertion":[{"value":"2021-05-30","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2022-01-25","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-01-05","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}