{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,16]],"date-time":"2026-06-16T04:34:56Z","timestamp":1781584496082,"version":"3.54.5"},"reference-count":63,"publisher":"Association for Computing Machinery (ACM)","issue":"1","funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62066033"],"award-info":[{"award-number":["62066033"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Distinguished Young Foundation of Inner Mongolia Autonomous Region","award":["2022JQ05"],"award-info":[{"award-number":["2022JQ05"]}]},{"name":"Applied Technology Research and Development Foundation of Inner Mongolia Autonomous Region","award":["2019GG372, 2020GG0046, 2021GG0158, 2021GG0165, 2020PT0002"],"award-info":[{"award-number":["2019GG372, 2020GG0046, 2021GG0158, 2021GG0165, 2020PT0002"]}]},{"name":"Achievements Transformation Project of Inner Mongolia Autonomous Region","award":["2019CG028"],"award-info":[{"award-number":["2019CG028"]}]},{"name":"Key R&D and Achievement Transformation Program of Inner Mongolia Autonomous Region","award":["2022YFHH0077"],"award-info":[{"award-number":["2022YFHH0077"]}]},{"name":"Central Government Fund for Promoting Local Scientific and Technological Development","award":["2022ZY0198"],"award-info":[{"award-number":["2022ZY0198"]}]},{"name":"Research Foundation for Young Scholars of Inner Mongolia University","award":["10000-A25206024"],"award-info":[{"award-number":["10000-A25206024"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2026,1,31]]},"abstract":"<jats:p>\n                    The\n                    <jats:bold>Zero-Shot Sketch-Based Image Retrieval<\/jats:bold>\n                    (ZS-SBIR) task aims to retrieve images associated with sketches from unseen classes, bringing great convenience to the engineering field. To address the modality gap, most existing works project images and sketches into a shared Euclidean space. However, the hierarchical structure of image data makes the Euclidean space not the optimal choice as an embedding space for representing complex structured image data. Meanwhile, existing text and hierarchical models are not effective enough for addressing the problem of knowledge transfer. To address these issues, this article proposes an original\n                    <jats:bold>Hyperbolic-Based Cross-Modal Semantic Remodeling Network<\/jats:bold>\n                    (called HCMSN) for ZS-SBIR. Specifically, this article proposes to extract category-level word embeddings based on BERT model, then align image features and sketches with the word embeddings using adversarial methods. Meanwhile, this article further proposes a cross-modal retrieval feature reconstruction network for improving the informativeness and robustness of retrieval features. Moreover, this article presents a feature projection network that maps the retrieval features to the hyperbolic space to generate the hyperbolic retrieval features, thus effectively representing the data with hierarchical structure. Extensive experiments demonstrate that the mAP@all of our HCMSN model surpasses CNN-based models by 20.9% on the Sketchy dataset, 1.2% on the more difficult TU-Berlin dataset, and 13.6% on the more challenging QuickDraw dataset.\n                  <\/jats:p>","DOI":"10.1145\/3777371","type":"journal-article","created":{"date-parts":[[2025,11,17]],"date-time":"2025-11-17T14:09:46Z","timestamp":1763388586000},"page":"1-23","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["Hyperbolic-Based Cross-Modal Semantic Remodeling Network for Zero-Shot Sketch-Based Image Retrieval"],"prefix":"10.1145","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0009-0005-4472-630X","authenticated-orcid":false,"given":"Qing","family":"Zhang","sequence":"first","affiliation":[{"name":"School of Computer Science, Inner Mongolia University, Hohhot, China, Inner Mongolia Key Laboratory of Mongolian Information Processing Technology, Inner Mongolia University, Hohhot, China, and National and Local Joint Engineering Research Center for Mongolian Intelligent Information Processing Technology, Inner Mongolia University, Hohhot, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0005-7206-6441","authenticated-orcid":false,"given":"Jing","family":"Zhang","sequence":"additional","affiliation":[{"name":"School of Computer Science, Inner Mongolia University, Hohhot, China, Inner Mongolia Key Laboratory of Mongolian Information Processing Technology, Inner Mongolia University, Hohhot, China, and National and Local Joint Engineering Research Center for Mongolian Intelligent Information Processing Technology, Inner Mongolia University, Hohhot, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8061-1474","authenticated-orcid":false,"given":"Xiangdong","family":"Su","sequence":"additional","affiliation":[{"name":"School of Computer Science, Inner Mongolia University, Hohhot, China, Inner Mongolia Key Laboratory of Mongolian Information Processing Technology, Inner Mongolia University, Hohhot, China, and National and Local Joint Engineering Research Center for Mongolian Intelligent Information Processing Technology, Inner Mongolia University, Hohhot, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7312-1629","authenticated-orcid":false,"given":"Feilong","family":"Bao","sequence":"additional","affiliation":[{"name":"School of Computer Science, Inner Mongolia University, Hohhot, China, Inner Mongolia Key Laboratory of Mongolian Information Processing Technology, Inner Mongolia University, Hohhot, China, and National and Local Joint Engineering Research Center for Mongolian Intelligent Information Processing Technology, Inner Mongolia University, Hohhot, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0005-5513-1192","authenticated-orcid":false,"given":"Guanglai","family":"Gao","sequence":"additional","affiliation":[{"name":"School of Computer Science, Inner Mongolia University, Hohhot, China, Inner Mongolia Key Laboratory of Mongolian Information Processing Technology, Inner Mongolia University, Hohhot, China, and National and Local Joint Engineering Research Center for Mongolian Intelligent Information Processing Technology, Inner Mongolia University, Hohhot, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,1,12]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.1145\/2637291"},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2021.108291"},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.1145\/3565368"},{"key":"e_1_3_1_5_2","first-page":"300","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV),","author":"Yelamarthi S. K.","year":"2018","unstructured":"S. K. Yelamarthi, S. K. Reddy, A. Mishra, and A. Mittal. 2018. A zero-shot framework for sketch based image retrieval. In Proceedings of the European Conference on Computer Vision (ECCV), 300\u2013317."},{"key":"e_1_3_1_6_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00228"},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i07.6817"},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00376"},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00379"},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2021.3123315"},{"key":"e_1_3_1_11_2","doi-asserted-by":"crossref","unstructured":"Z. Wang H. Wang J. Yan A. Wu and C. Deng. 2021. Domain-smoothing network for zero-shot sketch-based image retrieval. arXiv:2106.11841. Retrieved from https:\/\/arxiv.org\/abs\/2106.11841","DOI":"10.24963\/ijcai.2021\/158"},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2023.3248646"},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2022.09.104"},{"key":"e_1_3_1_14_2","doi-asserted-by":"crossref","first-page":"8892","DOI":"10.1109\/TIP.2020.3020383","article-title":"Progressive cross-modal semantic network for zero-shot sketch-based image retrieval","volume":"29","author":"Deng C.","year":"2020","unstructured":"C. Deng, X. Xu, H. Wang, M. Yang, and D. Tao. 2020. Progressive cross-modal semantic network for zero-shot sketch-based image retrieval. IEEE Transactions on Image Processing 29 (2020), 8892\u20138902.","journal-title":"IEEE Transactions on Image Processing"},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-020-01350-x"},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i07.6993"},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v36i2.20136"},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1145\/3503161.3548382"},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","DOI":"10.1145\/3591106.3592287"},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2025.3530248"},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.1145\/3474085.3475676"},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2024.110452"},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.asoc.2024.112474"},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2023.3265697"},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.00130"},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00929"},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2018.2857768"},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.01288"},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.1109\/JAS.2023.123207"},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2013.140"},{"key":"e_1_3_1_31_2","first-page":"3337","volume-title":"Proceedings of the 2011 IEEE Conference on Computer Vision and Pattern Recognition (CVPR \u201911)","author":"Liu J.","year":"2011","unstructured":"J. Liu, B. Kuipers, and S. Savarese. 2011. Recognizing human actions by attributes. In Proceedings of the 2011 IEEE Conference on Computer Vision and Pattern Recognition (CVPR \u201911). IEEE, 3337\u20133344."},{"key":"e_1_3_1_32_2","doi-asserted-by":"crossref","first-page":"730","DOI":"10.1007\/978-3-319-46454-1_44","volume-title":"Proceedings of the Computer Vision\u2013ECCV 2016: 14th European Conference","author":"Bucher M.","year":"2016","unstructured":"M. Bucher, S. Herbin, and F. Jurie. 2016. Improving semantic embedding consistency by metric learning for zero-shot classification. In Proceedings of the Computer Vision\u2013ECCV 2016: 14th European Conference, Part V 14. Springer, 730\u2013746."},{"key":"e_1_3_1_33_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.376"},{"key":"e_1_3_1_34_2","doi-asserted-by":"crossref","first-page":"1532","DOI":"10.3115\/v1\/D14-1162","volume-title":"Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP),","author":"Pennington J.","year":"2014","unstructured":"J. Pennington, R. Socher, and C. D. Manning. 2014. GloVe: Global vectors for word representation. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), 1532\u20131543."},{"key":"e_1_3_1_35_2","unstructured":"J. J. Jiang and D. W. Conrath. 1997. Semantic similarity based on corpus statistics and lexical taxonomy. arXiv:cmp-lg\/9709008. Retrieved from https:\/\/arxiv.org\/abs\/cmp-lg\/9709008"},{"key":"e_1_3_1_36_2","unstructured":"J. Devlin M.-W. Chang K. Lee and K. Toutanova. 2018. BERT: Pre-training of deep bidirectional transformers for language understanding. arXiv:1810.04805. Retrieved from https:\/\/arxiv.org\/abs\/1810.04805"},{"key":"e_1_3_1_37_2","first-page":"6338","article-title":"Poincar\u00e9 embeddings for learning hierarchical representations","volume":"30","author":"Nickel M.","year":"2017","unstructured":"M. Nickel and D. Kiela. 2017. Poincar\u00e9 embeddings for learning hierarchical representations. In Advances in Neural Information Processing Systems 30, 6338\u20136347.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00645"},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.00441"},{"key":"e_1_3_1_40_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00122"},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2022.108528"},{"key":"e_1_3_1_42_2","doi-asserted-by":"crossref","first-page":"102398","DOI":"10.1016\/j.aei.2024.102398","article-title":"Indicative vision transformer for end-to-end zero-shot sketch-based image retrieval","volume":"60","author":"Zhang H.","year":"2024","unstructured":"H. Zhang, D. Cheng, Q. Kou, M. Asad, and H. Jiang. 2024. Indicative vision transformer for end-to-end zero-shot sketch-based image retrieval. Advanced Engineering Informatics 60 (2024), 102398.","journal-title":"Advanced Engineering Informatics"},{"key":"e_1_3_1_43_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.02236"},{"key":"e_1_3_1_44_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.00731"},{"key":"e_1_3_1_45_2","doi-asserted-by":"publisher","DOI":"10.1145\/3503161.3548237"},{"key":"e_1_3_1_46_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.00277"},{"key":"e_1_3_1_47_2","doi-asserted-by":"crossref","first-page":"10342","DOI":"10.1109\/TMM.2024.3407664","article-title":"Estimating the semantics via sector embedding for image-text retrieval","volume":"26","author":"Wang Z.","year":"2024","unstructured":"Z. Wang, Z. Gao, M. Han, Y. Yang, and H. T. Shen. 2024. Estimating the semantics via sector embedding for image-text retrieval. IEEE Transactions on Multimedia 26 (2024), 10342\u201310353.","journal-title":"IEEE Transactions on Multimedia"},{"key":"e_1_3_1_48_2","first-page":"8748","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Radford A.","year":"2021","unstructured":"A. Radford, et al. 2021. Learning transferable visual models from natural language supervision. In Proceedings of the International Conference on Machine Learning, 8748\u20138763."},{"key":"e_1_3_1_49_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2024.3374111"},{"key":"e_1_3_1_50_2","doi-asserted-by":"publisher","DOI":"10.1109\/TNNLS.2024.3381347"},{"key":"e_1_3_1_51_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.00271"},{"key":"e_1_3_1_52_2","doi-asserted-by":"publisher","DOI":"10.1109\/WACV57701.2024.00555"},{"key":"e_1_3_1_53_2","first-page":"1780","volume-title":"Proceedings of the 33rd International Joint Conference on Artificial Intelligence","author":"Zhou Y.","year":"2024","unstructured":"Y. Zhou, D. Liu, and P. Mok. 2024. Zero-shot sketch based image retrieval via modality capacity guidance. In Proceedings of the 33rd International Joint Conference on Artificial Intelligence, 1780\u20131787."},{"issue":"2","key":"e_1_3_1_54_2","doi-asserted-by":"crossref","first-page":"2014","DOI":"10.1109\/TETCI.2024.3502430","article-title":"Triplet bridge for zero-shot sketch-based image retrieval","volume":"9","author":"Zheng J.","year":"2024","unstructured":"J. Zheng, Y. Tang, and D. Wu. 2024. Triplet bridge for zero-shot sketch-based image retrieval. IEEE Transactions on Emerging Topics in Computational Intelligence 9, 2 (2024), 2014\u20132025.","journal-title":"IEEE Transactions on Emerging Topics in Computational Intelligence"},{"key":"e_1_3_1_55_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2025.3543035"},{"key":"e_1_3_1_56_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2025.111529"},{"key":"e_1_3_1_57_2","first-page":"59","volume-title":"Flavors of Geometry","author":"Cannon J. W.","unstructured":"J. W. Cannon, W. J. Floyd, R. Kenyon, and W. R. Parry. 1997. Hyperbolic geometry. In Flavors of Geometry, Vol. 31. Silvio Levy (Ed.), Cambridge University Press, 59\u2013115."},{"key":"e_1_3_1_58_2","doi-asserted-by":"crossref","first-page":"355","DOI":"10.1007\/978-3-642-25878-7_34","volume-title":"Proceedings of the Graph Drawing: 19th International Symposium, Revised Selected Papers 19","author":"Sarkar R.","year":"2012","unstructured":"R. Sarkar. 2012. Low distortion Delaunay embedding of trees in hyperbolic plane. In Proceedings of the Graph Drawing: 19th International Symposium, Revised Selected Papers 19. Springer, 355\u2013366."},{"key":"e_1_3_1_59_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.247"},{"key":"e_1_3_1_60_2","doi-asserted-by":"publisher","DOI":"10.1145\/2185520.2185540"},{"key":"e_1_3_1_61_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2012.193"},{"key":"e_1_3_1_62_2","first-page":"2579","article-title":"Visualizing data using t-SNE","volume":"9","author":"Van der Maaten L.","year":"2008","unstructured":"L. Van der Maaten and G. Hinton. 2008. Visualizing data using t-SNE. Journal of Machine Learning Research 9 (Nov. 2008), 2579\u20132605.","journal-title":"Journal of Machine Learning Research"},{"key":"e_1_3_1_63_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.319"},{"key":"e_1_3_1_64_2","first-page":"1","article-title":"Hyperbolic deep learning in computer vision: A survey","author":"Mettes P.","unstructured":"P. Mettes, M. Ghadimi Atigh, M. Keller-Ressel, J. Gu, and S. Yeung. 2024. Hyperbolic deep learning in computer vision: A survey. International Journal of Computer Vision, 1\u201325.","journal-title":"International Journal of Computer Vision"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3777371","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,1,12]],"date-time":"2026-01-12T14:29:50Z","timestamp":1768228190000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3777371"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,1,12]]},"references-count":63,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2026,1,31]]}},"alternative-id":["10.1145\/3777371"],"URL":"https:\/\/doi.org\/10.1145\/3777371","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,1,12]]},"assertion":[{"value":"2024-07-24","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-11-10","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-01-12","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}