{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,20]],"date-time":"2026-03-20T14:15:10Z","timestamp":1774016110605,"version":"3.50.1"},"reference-count":44,"publisher":"World Scientific Pub Co Pte Ltd","issue":"07","funder":[{"name":"the National Science Foundation of China","award":["62088102"],"award-info":[{"award-number":["62088102"]}]},{"name":"STI2030-Major Projects","award":["2021ZD0113604"],"award-info":[{"award-number":["2021ZD0113604"]}]},{"name":"China National Postdoctoral Program for Innovative Talents from China Postdoctoral Science Foundation","award":["BX2021239"],"award-info":[{"award-number":["BX2021239"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Int. J. Neur. Syst."],"published-print":{"date-parts":[[2023,7]]},"abstract":"<jats:p>Zero-shot detection (ZSD) aims to locate and classify unseen objects in pictures or videos by semantic auxiliary information without additional training examples. Most of the existing ZSD methods are based on two-stage models, which achieve the detection of unseen classes by aligning object region proposals with semantic embeddings. However, these methods have several limitations, including poor region proposals for unseen classes, lack of consideration of semantic representations of unseen classes or their inter-class correlations, and domain bias towards seen classes, which can degrade overall performance. To address these issues, the Trans-ZSD framework is proposed, which is a transformer-based multi-scale contextual detection framework that explicitly exploits inter-class correlations between seen and unseen classes and optimizes feature distribution to learn discriminative features. Trans-ZSD is a single-stage approach that skips proposal generation and performs detection directly, allowing the encoding of long-term dependencies at multiple scales to learn contextual features while requiring fewer inductive biases. Trans-ZSD also introduces a foreground\u2013background separation branch to alleviate the confusion of unseen classes and backgrounds, contrastive learning to learn inter-class uniqueness and reduce misclassification between similar classes, and explicit inter-class commonality learning to facilitate generalization between related classes. Trans-ZSD addresses the domain bias problem in end-to-end generalized zero-shot detection (GZSD) models by using balance loss to maximize response consistency between seen and unseen predictions, ensuring that the model does not bias towards seen classes. The Trans-ZSD framework is evaluated on the PASCAL VOC and MS COCO datasets, demonstrating significant improvements over existing ZSD models.<\/jats:p>","DOI":"10.1142\/s0129065723500351","type":"journal-article","created":{"date-parts":[[2023,4,23]],"date-time":"2023-04-23T08:58:24Z","timestamp":1682240304000},"source":"Crossref","is-referenced-by-count":13,"title":["Transformer-Based Approach Via Contrastive Learning for Zero-Shot Detection"],"prefix":"10.1142","volume":"33","author":[{"given":"Wei","family":"Liu","sequence":"first","affiliation":[{"name":"Institute of Artificial Intelligence and Robotics, Xian Jiaotong University, Xian, Shaanxi 710049, P.\u00a0R.\u00a0China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hui","family":"Chen","sequence":"additional","affiliation":[{"name":"Institute of Artificial Intelligence and Robotics, Xian Jiaotong University, Xian, Shaanxi 710049, P.\u00a0R.\u00a0China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yongqiang","family":"Ma","sequence":"additional","affiliation":[{"name":"Institute of Artificial Intelligence and Robotics, Xian Jiaotong University, Xian, Shaanxi 710049, P.\u00a0R.\u00a0China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jianji","family":"Wang","sequence":"additional","affiliation":[{"name":"Institute of Artificial Intelligence and Robotics, Xian Jiaotong University, Xian, Shaanxi 710049, P.\u00a0R.\u00a0China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Nanning","family":"Zheng","sequence":"additional","affiliation":[{"name":"Institute of Artificial Intelligence and Robotics, Xian Jiaotong University, Xian, Shaanxi 710049, P.\u00a0R.\u00a0China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"219","published-online":{"date-parts":[[2023,6,14]]},"reference":[{"key":"S0129065723500351BIB001","first-page":"6154","volume-title":"Proc. IEEE Conf. Computer Vision and Pattern Recognition","author":"Cai Z.","year":"2018"},{"key":"S0129065723500351BIB002","first-page":"2961","volume-title":"Proc. IEEE Int. Conf. Computer Vision","author":"He K.","year":"2017"},{"key":"S0129065723500351BIB003","first-page":"7263","volume-title":"Proc. IEEE Conf. Computer Vision and Pattern Recognition","author":"Redmon J.","year":"2017"},{"key":"S0129065723500351BIB004","volume-title":"Advances in Neural Information Processing Systems","author":"Ren S.","year":"2015"},{"key":"S0129065723500351BIB005","doi-asserted-by":"crossref","first-page":"2250052","DOI":"10.1142\/S0129065722500526","volume":"32","author":"K\u00fc\u00e7\u00fcko\u011flu B.","year":"2022","journal-title":"Int. J. Neural Syst."},{"key":"S0129065723500351BIB006","first-page":"384","volume-title":"Proc. Eur. Conf. Computer Vision (ECCV)","author":"Bansal A.","year":"2018"},{"issue":"1","key":"S0129065723500351BIB008","doi-asserted-by":"crossref","first-page":"32","DOI":"10.1007\/s11263-016-0981-7","volume":"123","author":"Krishna R.","year":"2017","journal-title":"Int. J. Comput. Vision"},{"issue":"12","key":"S0129065723500351BIB009","doi-asserted-by":"crossref","first-page":"2979","DOI":"10.1007\/s11263-020-01355-6","volume":"128","author":"Rahman S.","year":"2020","journal-title":"Int. J. Comput. Vision"},{"key":"S0129065723500351BIB010","first-page":"11932","volume-title":"Proc. AAAI Conf. Artificial Intelligence","volume":"34","author":"Rahman S.","year":"2020"},{"key":"S0129065723500351BIB011","first-page":"8690","volume-title":"Proc. AAAI Conf. on Artificial Intelligence","volume":"33","author":"Li Z.","year":"2019"},{"key":"S0129065723500351BIB012","first-page":"3111","volume-title":"NIPS\u201913: Proc. 26th Int. Conf. Neural Information Processing Systems","author":"Mikolov T.","year":"2013"},{"key":"S0129065723500351BIB013","first-page":"1209","volume-title":"Proc. IEEE\/CVF Winter Conf. Applications of Computer Vision","author":"Gupta D.","year":"2020"},{"key":"S0129065723500351BIB014","first-page":"230","volume-title":"2020 IEEE 32nd Int. Conf. Tools with Artificial Intelligence (ICTAI)","author":"Wang K.","year":"2020"},{"key":"S0129065723500351BIB016","first-page":"107","volume-title":"Proc. Asian Conf. Computer Vision","author":"Zheng Y.","year":"2020"},{"key":"S0129065723500351BIB018","first-page":"7363","volume-title":"Proc. IEEE\/CVF Conf. Computer Vision and Pattern Recognition","author":"Lu X.","year":"2019"},{"key":"S0129065723500351BIB019","first-page":"1440","volume-title":"Proc. IEEE Int. Conf. Computer Vision","author":"Girshick R.","year":"2015"},{"key":"S0129065723500351BIB020","first-page":"379","volume-title":"NIPS\u201916: Proc. 30th Int. Conf. Neural Information Processing Systems","author":"Dai J.","year":"2016"},{"key":"S0129065723500351BIB021","first-page":"213","volume-title":"Eur. Conf. Computer Vision","author":"Carion N.","year":"2020"},{"key":"S0129065723500351BIB022","first-page":"764","volume-title":"Proc. IEEE Int. Conf Computer Vision","author":"Dai J.","year":"2017"},{"key":"S0129065723500351BIB023","first-page":"935","volume-title":"NIPS\u201913: Proc. 26th Int. Conf. on Neural Information Processing Systems","author":"Socher R.","year":"2013"},{"key":"S0129065723500351BIB024","first-page":"2121","volume-title":"NIPS\u201913: Proc. 26th Int. Conf. Neural Information Processing Systems","author":"Frome A.","year":"2013"},{"issue":"4","key":"S0129065723500351BIB025","doi-asserted-by":"crossref","first-page":"500","DOI":"10.1111\/mice.12755","volume":"37","author":"\u017barski M.","year":"2022","journal-title":"Comput.-Aided Civ. Infrastruct. Eng."},{"issue":"7","key":"S0129065723500351BIB026","doi-asserted-by":"crossref","first-page":"2250032","DOI":"10.1142\/S0129065722500320","volume":"32","author":"Yu Z.","year":"2022","journal-title":"Int. J. Neural Syst."},{"key":"S0129065723500351BIB027","doi-asserted-by":"crossref","first-page":"227","DOI":"10.3233\/ICA-220680","volume":"29","author":"Wolyn S.","year":"2022","journal-title":"Integr. Comput.-Aided Eng."},{"key":"S0129065723500351BIB028","first-page":"5542","volume-title":"Proc. IEEE Conf. Computer Vision and Pattern Recognition","author":"Xian Y.","year":"2018"},{"key":"S0129065723500351BIB029","first-page":"21","volume-title":"Proc. Eur. Conf. Computer Vision (ECCV)","author":"Felix R.","year":"2018"},{"key":"S0129065723500351BIB030","first-page":"1597","volume-title":"Int. Conf. Machine Learning","author":"Chen T.","year":"2020"},{"key":"S0129065723500351BIB031","first-page":"8392","volume-title":"Proc. IEEE\/CVF Int. Conf. Computer Vision","author":"Xie E.","year":"2021"},{"key":"S0129065723500351BIB032","first-page":"319","volume-title":"Eur. Conf. Computer Vision","author":"Park T.","year":"2020"},{"issue":"11","key":"S0129065723500351BIB033","doi-asserted-by":"crossref","first-page":"1382","DOI":"10.1111\/mice.12640","volume":"36","author":"Hsieh Y.-A.","year":"2021","journal-title":"Comput.-Aided Civ. Infrastruct. Eng."},{"key":"S0129065723500351BIB034","first-page":"18661","volume-title":"NIPS\u201920: Proc. 34th Int. Conf. Neural Information Processing Systems","author":"Khosla P.","year":"2020"},{"key":"S0129065723500351BIB036","doi-asserted-by":"publisher","DOI":"10.1142\/S0129065722500162"},{"key":"S0129065723500351BIB037","first-page":"658","volume-title":"Proc. IEEE\/CVF Conf. Computer Vision and Pattern Recognition","author":"Rezatofighi H.","year":"2019"},{"key":"S0129065723500351BIB038","first-page":"6827","volume-title":"NIPS\u201920: Proc. 34th Int. Conf. Neural Information Processing Systems","author":"Tian Y.","year":"2020"},{"key":"S0129065723500351BIB039","doi-asserted-by":"crossref","first-page":"211","DOI":"10.1007\/s11263-015-0816-y","volume":"115","author":"Russakovsky O.","year":"2015","journal-title":"Int. J. Comput. Vision"},{"key":"S0129065723500351BIB040","first-page":"944","volume-title":"Proc. IEEE\/CVF Conf. Computer Vision and Pattern Recognition Workshops","author":"Li Y.","year":"2020"},{"key":"S0129065723500351BIB041","first-page":"155","volume-title":"Proc. Asian Conf. Computer Vision","author":"Hayat N.","year":"2020"},{"key":"S0129065723500351BIB042","first-page":"1993","volume-title":"Proc. AAAI Conf. Artificial Intelligence","volume":"35","author":"Li Y.","year":"2021"},{"key":"S0129065723500351BIB043","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1109\/TPAMI.2022.3226498","author":"Yan C.","year":"2022","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"S0129065723500351BIB044","first-page":"547","volume-title":"Asian Conf. Computer Vision","author":"Rahman S.","year":"2018"},{"key":"S0129065723500351BIB045","first-page":"6082","volume-title":"Proc. IEEE\/CVF Int. Conf. Computer Vision","author":"Rahman S.","year":"2019"},{"key":"S0129065723500351BIB046","first-page":"2980","volume-title":"Proc. IEEE Int. Conf. Computer Vision","author":"Lin T.-Y.","year":"2017"},{"issue":"10","key":"S0129065723500351BIB048","doi-asserted-by":"crossref","first-page":"2941","DOI":"10.1109\/TCSVT.2018.2870832","volume":"29","author":"Cong R.","year":"2018","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"issue":"8","key":"S0129065723500351BIB049","doi-asserted-by":"crossref","first-page":"739","DOI":"10.1109\/LSP.2010.2053200","volume":"17","author":"Yan J.","year":"2010","journal-title":"IEEE Signal Process Lett."}],"container-title":["International Journal of Neural Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.worldscientific.com\/doi\/pdf\/10.1142\/S0129065723500351","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,10,19]],"date-time":"2024-10-19T01:23:20Z","timestamp":1729301000000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.worldscientific.com\/doi\/10.1142\/S0129065723500351"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,6,14]]},"references-count":44,"journal-issue":{"issue":"07","published-print":{"date-parts":[[2023,7]]}},"alternative-id":["10.1142\/S0129065723500351"],"URL":"https:\/\/doi.org\/10.1142\/s0129065723500351","relation":{},"ISSN":["0129-0657","1793-6462"],"issn-type":[{"value":"0129-0657","type":"print"},{"value":"1793-6462","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,6,14]]},"article-number":"2350035"}}