{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,18]],"date-time":"2026-07-18T17:54:06Z","timestamp":1784397246597,"version":"3.55.0"},"publisher-location":"New York, NY, USA","reference-count":75,"publisher":"ACM","license":[{"start":{"date-parts":[[2022,10,10]],"date-time":"2022-10-10T00:00:00Z","timestamp":1665360000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["62176188"],"award-info":[{"award-number":["62176188"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Key Research and Development Program of Hubei Province","award":["2021BAA187"],"award-info":[{"award-number":["2021BAA187"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2022,10,10]]},"DOI":"10.1145\/3503161.3548770","type":"proceedings-article","created":{"date-parts":[[2022,10,10]],"date-time":"2022-10-10T15:42:35Z","timestamp":1665416555000},"page":"7317-7326","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":20,"title":["Pyramidal Transformer with Conv-Patchify for Person Re-identification"],"prefix":"10.1145","author":[{"given":"He","family":"Li","sequence":"first","affiliation":[{"name":"Wuhan University, Wuhan, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Mang","family":"Ye","sequence":"additional","affiliation":[{"name":"Wuhan University &amp; Hubei Luojia Laboratory, Wuhan, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Cong","family":"Wang","sequence":"additional","affiliation":[{"name":"Huawei Technologies Ltd., Zhejiang, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Bo","family":"Du","sequence":"additional","affiliation":[{"name":"Wuhan University &amp; Hubei Luojia Laboratory, Wuhan, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2022,10,10]]},"reference":[{"key":"e_1_3_2_2_1_1","unstructured":"2020. MindSpore. https:\/\/www.mindspore.cn\/  2020. MindSpore. https:\/\/www.mindspore.cn\/"},{"key":"e_1_3_2_2_2_1","volume-title":"Why do deep convolutional networks generalize so poorly to small image transformations Journal of Machine Learning Research 20","author":"Azulay Aharon","year":"2019","unstructured":"Aharon Azulay and Yair Weiss . 2019. Why do deep convolutional networks generalize so poorly to small image transformations Journal of Machine Learning Research 20 ( 2019 ), 1--25. Aharon Azulay and Yair Weiss. 2019. Why do deep convolutional networks generalize so poorly to small image transformations Journal of Machine Learning Research 20 (2019), 1--25."},{"key":"e_1_3_2_2_3_1","volume-title":"Int. Conf. Comput. Vis.","author":"Chen Binghui","year":"2019","unstructured":"Binghui Chen , Weihong Deng , Jiani Hu , Jiani Hu , and Jiani Hu . 2019 . Mixed high-order attention network for person reidentification . Int. Conf. Comput. Vis. (2019), 371--381. Binghui Chen, Weihong Deng, Jiani Hu, Jiani Hu, and Jiani Hu. 2019. Mixed high-order attention network for person reidentification. Int. Conf. Comput. Vis. (2019), 371--381."},{"key":"e_1_3_2_2_4_1","volume-title":"Int. Conf. Comput. Vis.","author":"Chen Tianlong","year":"2019","unstructured":"Tianlong Chen , Shaojin Ding , Jingyi Xie , Ye Yuan , Wuyang Chen , Yang Yang , Zhou Ren , and Zhangyang Wang . 2019 . Abdnet: Attentive but diverse person re-identification . Int. Conf. Comput. Vis. (2019), 8351--8361. Tianlong Chen, Shaojin Ding, Jingyi Xie, Ye Yuan, Wuyang Chen, Yang Yang, Zhou Ren, and Zhangyang Wang. 2019. Abdnet: Attentive but diverse person re-identification. Int. Conf. Comput. Vis. (2019), 8351--8361."},{"key":"e_1_3_2_2_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00336"},{"key":"e_1_3_2_2_6_1","volume-title":"OH-Former: Omni- Relational High-Order Transformer for Person Re-Identification. arXiv preprint arXiv:2109.11159","author":"Chen Xianing","year":"2021","unstructured":"Xianing Chen , Jialang Xu , Jiale Xu , and Shenghua Gao . 2021. OH-Former: Omni- Relational High-Order Transformer for Person Re-Identification. arXiv preprint arXiv:2109.11159 ( 2021 ). Xianing Chen, Jialang Xu, Jiale Xu, and Shenghua Gao. 2021. OH-Former: Omni- Relational High-Order Transformer for Person Re-Identification. arXiv preprint arXiv:2109.11159 (2021)."},{"key":"e_1_3_2_2_7_1","volume-title":"Xception: Deep Learning With Depthwise Separable Convolutions. IEEE Conf. Comput. Vis. Pattern Recog.","author":"Chollet Francois","year":"2017","unstructured":"Francois Chollet . 2017 . Xception: Deep Learning With Depthwise Separable Convolutions. IEEE Conf. Comput. Vis. Pattern Recog. (2017), 1251--1258. Francois Chollet. 2017. Xception: Deep Learning With Depthwise Separable Convolutions. IEEE Conf. Comput. Vis. Pattern Recog. (2017), 1251--1258."},{"key":"e_1_3_2_2_8_1","volume-title":"Conditional positional encodings for vision transformers. arXiv preprint arXiv:2102.10882","author":"Chu Xiangxiang","year":"2021","unstructured":"Xiangxiang Chu , Bo Zhang , Zhi Tian , Xiaolin Wei , and Huaxia Xia . 2021. Conditional positional encodings for vision transformers. arXiv preprint arXiv:2102.10882 ( 2021 ). Xiangxiang Chu, Bo Zhang, Zhi Tian, Xiaolin Wei, and Huaxia Xia. 2021. Conditional positional encodings for vision transformers. arXiv preprint arXiv:2102.10882 (2021)."},{"key":"e_1_3_2_2_9_1","volume-title":"Int. Conf. Learn. Represent.","author":"Dosovitskiy Alexey","year":"2021","unstructured":"Alexey Dosovitskiy , Lucas Beyer , Alexander Kolesnikov , Dirk Weissenborn , and Xiaohua et al. Zhai. 2021. An image is worth 16x16 words: Transformers for image recognition at scale . Int. Conf. Learn. Represent. ( 2021 ). Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, and Xiaohua et al. Zhai. 2021. An image is worth 16x16 words: Transformers for image recognition at scale. Int. Conf. Learn. Represent. (2021)."},{"key":"e_1_3_2_2_10_1","volume-title":"CMT: Convolutional Neural Networks Meet Vision Transformers. arXiv preprint arXiv:2107","author":"Guo Jianyuan","year":"2021","unstructured":"Jianyuan Guo , Kai Han , HanWu, Chang Xu , Yehui Tang , Chunjing Xu , and Yunhe Wang . 2021 . CMT: Convolutional Neural Networks Meet Vision Transformers. arXiv preprint arXiv:2107 .06263 (2021). Jianyuan Guo, Kai Han, HanWu, Chang Xu, Yehui Tang, Chunjing Xu, and Yunhe Wang. 2021. CMT: Convolutional Neural Networks Meet Vision Transformers. arXiv preprint arXiv:2107.06263 (2021)."},{"key":"e_1_3_2_2_11_1","volume-title":"Transformer in transformer. arXiv preprint arXiv:2103.00112","author":"Han Kai","year":"2021","unstructured":"Kai Han , An Xiao , Enhua Wu , Jianyuan Guo , Chunjing Xu , and YunheWang. 2021. Transformer in transformer. arXiv preprint arXiv:2103.00112 ( 2021 ). Kai Han, An Xiao, Enhua Wu, Jianyuan Guo, Chunjing Xu, and YunheWang. 2021. Transformer in transformer. arXiv preprint arXiv:2103.00112 (2021)."},{"key":"e_1_3_2_2_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_2_2_13_1","volume-title":"TransReID: Transformer-Based Object Re-Identification. Int. Conf. Comput. Vis.","author":"He Shuting","year":"2021","unstructured":"Shuting He , Hao Luo , Pichao Wang , Fan Wang , Hao Li , and Wei Jiang . 2021 . TransReID: Transformer-Based Object Re-Identification. Int. Conf. Comput. Vis. (2021), 15013--15022. Shuting He, Hao Luo, Pichao Wang, Fan Wang, Hao Li, and Wei Jiang. 2021. TransReID: Transformer-Based Object Re-Identification. Int. Conf. Comput. Vis. (2021), 15013--15022."},{"key":"e_1_3_2_2_14_1","volume-title":"Int. Conf. Comput. Vis.","author":"Howard Andrew","year":"2019","unstructured":"Andrew Howard , Mark Sandler , Grace Chu , Liang-Chieh Chen , Bo Chen , Mingxing Tan , Weijun Wang , Yukun Zhu , Ruoming Pang , Vijay Vasudevan , and et al. 2019. Searching for mobilenetv3 . Int. Conf. Comput. Vis. ( 2019 ), 1314--1324. Andrew Howard, Mark Sandler, Grace Chu, Liang-Chieh Chen, Bo Chen, Mingxing Tan, Weijun Wang, Yukun Zhu, Ruoming Pang, Vijay Vasudevan, and et al. 2019. Searching for mobilenetv3. Int. Conf. Comput. Vis. (2019), 1314--1324."},{"key":"e_1_3_2_2_15_1","volume-title":"Gatherexcite: Exploiting feature context in convolutional neural networks. arXiv preprint arXiv:1810.12348","author":"Hu Jie","year":"2018","unstructured":"Jie Hu , Li Shen , Samuel Albanie , Gang Sun , and Andrea Vedaldi . 2018 . Gatherexcite: Exploiting feature context in convolutional neural networks. arXiv preprint arXiv:1810.12348 (2018). Jie Hu, Li Shen, Samuel Albanie, Gang Sun, and Andrea Vedaldi. 2018. Gatherexcite: Exploiting feature context in convolutional neural networks. arXiv preprint arXiv:1810.12348 (2018)."},{"key":"e_1_3_2_2_16_1","volume-title":"Int. Conf. Mach. Learn. 37","author":"Ioffe Sergey","year":"2015","unstructured":"Sergey Ioffe and Christian Szegedy . 2015 . Batch normalization: Accelerating deep network training by reducing internal covariate shift . Int. Conf. Mach. Learn. 37 (2015), 448--456. Sergey Ioffe and Christian Szegedy. 2015. Batch normalization: Accelerating deep network training by reducing internal covariate shift. Int. Conf. Mach. Learn. 37 (2015), 448--456."},{"key":"e_1_3_2_2_17_1","volume-title":"How Much Position Information Do Convolutional Neural Networks Encode Int. Conf. Learn. Represent.","author":"Islam Md Amirul","year":"2020","unstructured":"Md Amirul Islam , Sen Jia , and Neil D. B. Bruce . 2020 . How Much Position Information Do Convolutional Neural Networks Encode Int. Conf. Learn. Represent. ( 2020 ). Md Amirul Islam, Sen Jia, and Neil D. B. Bruce. 2020. How Much Position Information Do Convolutional Neural Networks Encode Int. Conf. Learn. Represent. (2020)."},{"key":"e_1_3_2_2_18_1","volume-title":"All Tokens Matter: Token Labeling for Training Better Vision Transformers. arXiv preprint arXiv:2104.10858","author":"Jiang Zihang","year":"2021","unstructured":"Zihang Jiang , Qibin Hou , Li Yuan , Daquan Zhou , Yujun Shi , Xiaojie Jin , Anran Wang , and Jiashi Feng . 2021. All Tokens Matter: Token Labeling for Training Better Vision Transformers. arXiv preprint arXiv:2104.10858 ( 2021 ). Zihang Jiang, Qibin Hou, Li Yuan, Daquan Zhou, Yujun Shi, Xiaojie Jin, Anran Wang, and Jiashi Feng. 2021. All Tokens Matter: Token Labeling for Training Better Vision Transformers. arXiv preprint arXiv:2104.10858 (2021)."},{"key":"e_1_3_2_2_19_1","volume-title":"On Translation Invariance in CNNs: Convolutional Layers Can Exploit Absolute Spatial Location. IEEE Conf. Comput. Vis. Pattern Recog.","author":"Kayhan Osman Semih","year":"2020","unstructured":"Osman Semih Kayhan and Jan C . van Gemert. 2020 . On Translation Invariance in CNNs: Convolutional Layers Can Exploit Absolute Spatial Location. IEEE Conf. Comput. Vis. Pattern Recog. ( 2020 ), 14274--14285. Osman Semih Kayhan and Jan C. van Gemert. 2020. On Translation Invariance in CNNs: Convolutional Layers Can Exploit Absolute Spatial Location. IEEE Conf. Comput. Vis. Pattern Recog. (2020), 14274--14285."},{"key":"e_1_3_2_2_20_1","volume-title":"Imagenet classification with deep convolutional neural networks. Advances in neural information processing systems","author":"Krizhevsky Alex","year":"2012","unstructured":"Alex Krizhevsky , Ilya Sutskever , and Geoffrey E Hinton . 2012. Imagenet classification with deep convolutional neural networks. Advances in neural information processing systems ( 2012 ). Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. 2012. Imagenet classification with deep convolutional neural networks. Advances in neural information processing systems (2012)."},{"key":"e_1_3_2_2_21_1","volume-title":"Combined Depth Space Based Architecture Search for Person Re-Identification. IEEE Conf. Comput. Vis. Pattern Recog.","author":"Li Hanjun","year":"2021","unstructured":"Hanjun Li , Gaojie Wu , and Wei-Shi Zheng . 2021 . Combined Depth Space Based Architecture Search for Person Re-Identification. IEEE Conf. Comput. Vis. Pattern Recog. (2021), 6729--6738. Hanjun Li, Gaojie Wu, and Wei-Shi Zheng. 2021. Combined Depth Space Based Architecture Search for Person Re-Identification. IEEE Conf. Comput. Vis. Pattern Recog. (2021), 6729--6738."},{"key":"e_1_3_2_2_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/3474085.3475455"},{"key":"e_1_3_2_2_23_1","volume-title":"Contextual Transformer Networks for Visual Recognition. arXiv preprint arXiv:2107.12292","author":"Li Yehao","year":"2021","unstructured":"Yehao Li , Ting Yao , Yingwei Pan , and Tao Mei . 2021. Contextual Transformer Networks for Visual Recognition. arXiv preprint arXiv:2107.12292 ( 2021 ). Yehao Li, Ting Yao, Yingwei Pan, and Tao Mei. 2021. Contextual Transformer Networks for Visual Recognition. arXiv preprint arXiv:2107.12292 (2021)."},{"key":"e_1_3_2_2_24_1","volume-title":"Interpretable and Generalizable Person Re- Identification with Query-Adaptive Convolution and Temporal Lifting. Eur. Conf. Comput. Vis.","author":"Liao Shengcai","year":"2020","unstructured":"Shengcai Liao and Ling Shao . 2020 . Interpretable and Generalizable Person Re- Identification with Query-Adaptive Convolution and Temporal Lifting. Eur. Conf. Comput. Vis. (2020), 456--474. Shengcai Liao and Ling Shao. 2020. Interpretable and Generalizable Person Re- Identification with Query-Adaptive Convolution and Temporal Lifting. Eur. Conf. Comput. Vis. (2020), 456--474."},{"key":"e_1_3_2_2_25_1","doi-asserted-by":"crossref","unstructured":"Shengcai Liao and Ling Shao. 2021. Transformer-Based Deep Image Matching for Generalizable Person Re-identification. Adv. Neural Inform. Process. Syst. (2021).  Shengcai Liao and Ling Shao. 2021. Transformer-Based Deep Image Matching for Generalizable Person Re-identification. Adv. Neural Inform. Process. Syst. (2021).","DOI":"10.1109\/CVPR52688.2022.00721"},{"key":"e_1_3_2_2_26_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2019.06.006"},{"key":"e_1_3_2_2_27_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2017.2700762"},{"key":"e_1_3_2_2_28_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00986"},{"key":"e_1_3_2_2_29_1","volume-title":"SGDR: Stochastic Gradient Descent with Warm Restarts. Int. Conf. Learn. Represent.","author":"Loshchilov Ilya","year":"2017","unstructured":"Ilya Loshchilov and Frank Hutter . 2017 . SGDR: Stochastic Gradient Descent with Warm Restarts. Int. Conf. Learn. Represent. (2017). Ilya Loshchilov and Frank Hutter. 2017. SGDR: Stochastic Gradient Descent with Warm Restarts. Int. Conf. Learn. Represent. (2017)."},{"key":"e_1_3_2_2_30_1","volume-title":"DecoupledWeight Decay Regularization. Int. Conf. Learn. Represent.","author":"Loshchilov Ilya","year":"2019","unstructured":"Ilya Loshchilov and Frank Hutter . 2019 . DecoupledWeight Decay Regularization. Int. Conf. Learn. Represent. (2019). Ilya Loshchilov and Frank Hutter. 2019. DecoupledWeight Decay Regularization. Int. Conf. Learn. Represent. (2019)."},{"key":"e_1_3_2_2_31_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW.2019.00190"},{"key":"e_1_3_2_2_32_1","unstructured":"Wenjie Luo Yujia Li Raquel Urtasun and Richard Zemel. 2016. Understanding the effective receptive field in deep convolutional neural networks. Adv. Neural Inform. Process. Syst. (2016) 4905--4913.  Wenjie Luo Yujia Li Raquel Urtasun and Richard Zemel. 2016. Understanding the effective receptive field in deep convolutional neural networks. Adv. Neural Inform. Process. Syst. (2016) 4905--4913."},{"key":"e_1_3_2_2_33_1","first-page":"3221","article-title":"Accelerating t-SNE using tree-based algorithms","volume":"15","author":"Der Maaten Laurens Van","year":"2014","unstructured":"Laurens Van Der Maaten . 2014 . Accelerating t-SNE using tree-based algorithms . Journal of Machine Learning Research 15 (2014), 3221 -- 3245 . Laurens Van Der Maaten. 2014. Accelerating t-SNE using tree-based algorithms. Journal of Machine Learning Research 15 (2014), 3221--3245.","journal-title":"Journal of Machine Learning Research"},{"key":"e_1_3_2_2_34_1","volume-title":"Bam: Bottleneck attention module. arXiv preprint arXiv:1807.06514","author":"Park Jongchan","year":"2018","unstructured":"Jongchan Park , Sanghyun Woo , Joon-Young Lee , and In So Kweon . 2018 . Bam: Bottleneck attention module. arXiv preprint arXiv:1807.06514 (2018). Jongchan Park, Sanghyun Woo, Joon-Young Lee, and In So Kweon. 2018. Bam: Bottleneck attention module. arXiv preprint arXiv:1807.06514 (2018)."},{"key":"e_1_3_2_2_35_1","volume-title":"Stripe-based and attribute-aware network: A two-branch deep model for vehicle re-identification. Measurement Science and Technology 31, 9","author":"Qian Jingjing","year":"2020","unstructured":"Jingjing Qian , Wei Jiang , Hao Luo , and Hongyan Yu. 2020. Stripe-based and attribute-aware network: A two-branch deep model for vehicle re-identification. Measurement Science and Technology 31, 9 ( 2020 ). Jingjing Qian, Wei Jiang, Hao Luo, and Hongyan Yu. 2020. Stripe-based and attribute-aware network: A two-branch deep model for vehicle re-identification. Measurement Science and Technology 31, 9 (2020)."},{"key":"e_1_3_2_2_36_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2018.2846566"},{"key":"e_1_3_2_2_37_1","volume-title":"Learning Instance-Level Spatial-Temporal Patterns for Person Re-Identification. Int. Conf. Comput. Vis.","author":"Ren Min","year":"2021","unstructured":"Min Ren , Lingxiao He , Xingyu Liao , Wu Liu , YunlongWang, and Tieniu Tan . 2021 . Learning Instance-Level Spatial-Temporal Patterns for Person Re-Identification. Int. Conf. Comput. Vis. (2021), 14930--14939. Min Ren, Lingxiao He, Xingyu Liao,Wu Liu, YunlongWang, and Tieniu Tan. 2021. Learning Instance-Level Spatial-Temporal Patterns for Person Re-Identification. Int. Conf. Comput. Vis. (2021), 14930--14939."},{"key":"e_1_3_2_2_38_1","unstructured":"Tal Ridnik Emanuel Ben-Baruch Asaf Noy and Lihi Zelnik-Manor. 2021. ImageNet-21K Pretraining for the Masses. Adv. Neural Inform. Process. Syst. (2021).  Tal Ridnik Emanuel Ben-Baruch Asaf Noy and Lihi Zelnik-Manor. 2021. ImageNet-21K Pretraining for the Masses. Adv. Neural Inform. Process. Syst. (2021)."},{"key":"e_1_3_2_2_39_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00474"},{"key":"e_1_3_2_2_40_1","volume-title":"Dual Attention Matching Network for Context-Aware Feature Sequence Based Person Re-Identification. IEEE Conf. Comput. Vis. Pattern Recog.","author":"Si Jianlou","year":"2018","unstructured":"Jianlou Si , Honggang Zhang , Chun-Guang Li , Jason Kuen , Xiangfei Kong , Alex C. Kot , and GangWang. 2018 . Dual Attention Matching Network for Context-Aware Feature Sequence Based Person Re-Identification. IEEE Conf. Comput. Vis. Pattern Recog. (2018), 5363--5372. Jianlou Si, Honggang Zhang, Chun-Guang Li, Jason Kuen, Xiangfei Kong, Alex C. Kot, and GangWang. 2018. Dual Attention Matching Network for Context-Aware Feature Sequence Based Person Re-Identification. IEEE Conf. Comput. Vis. Pattern Recog. (2018), 5363--5372."},{"key":"e_1_3_2_2_41_1","volume-title":"Very Deep Convolutional Networks for Large-Scale Image Recognition. arXiv preprint arXiv:1409.1556","author":"Simonyan Karen","year":"2014","unstructured":"Karen Simonyan and Andrew Zisserman . 2014. Very Deep Convolutional Networks for Large-Scale Image Recognition. arXiv preprint arXiv:1409.1556 ( 2014 ). Karen Simonyan and Andrew Zisserman. 2014. Very Deep Convolutional Networks for Large-Scale Image Recognition. arXiv preprint arXiv:1409.1556 (2014)."},{"key":"e_1_3_2_2_42_1","volume-title":"Mask-Guided Contrastive Attention Model for Person Re-Identification. IEEE Conf. Comput. Vis. Pattern Recog.","author":"Song Chunfeng","year":"2018","unstructured":"Chunfeng Song , Yan Huang , Wanli Ouyang , and LiangWang. 2018 . Mask-Guided Contrastive Attention Model for Person Re-Identification. IEEE Conf. Comput. Vis. Pattern Recog. (2018), 1179--1188. Chunfeng Song, Yan Huang,Wanli Ouyang, and LiangWang. 2018. Mask-Guided Contrastive Attention Model for Person Re-Identification. IEEE Conf. Comput. Vis. Pattern Recog. (2018), 1179--1188."},{"key":"e_1_3_2_2_43_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.97"},{"key":"e_1_3_2_2_44_1","volume-title":"Dissecting Person Re-Identification From the Viewpoint of Viewpoint. IEEE Conf. Comput. Vis. Pattern Recog.","author":"Sun Xiaoxiao","year":"2019","unstructured":"Xiaoxiao Sun and Liang Zheng . 2019 . Dissecting Person Re-Identification From the Viewpoint of Viewpoint. IEEE Conf. Comput. Vis. Pattern Recog. (2019), 608--617. Xiaoxiao Sun and Liang Zheng. 2019. Dissecting Person Re-Identification From the Viewpoint of Viewpoint. IEEE Conf. Comput. Vis. Pattern Recog. (2019), 608--617."},{"key":"e_1_3_2_2_45_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00643"},{"key":"e_1_3_2_2_46_1","volume-title":"Eur. Conf. Comput. Vis.","author":"Sun Yifan","year":"2018","unstructured":"Yifan Sun , Liang Zheng , Yi Yang , Qi Tian , and Shengjin Wang . 2018 . Circle loss: A unified perspective of pair similarity optimization . Eur. Conf. Comput. Vis. (2018), 480--496. Yifan Sun, Liang Zheng, Yi Yang, Qi Tian, and Shengjin Wang. 2018. Circle loss: A unified perspective of pair similarity optimization. Eur. Conf. Comput. Vis. (2018), 480--496."},{"key":"e_1_3_2_2_47_1","volume-title":"inception-resnet and the impact of residual connections on learning. AAAI","author":"Szegedy Christian","year":"2017","unstructured":"Christian Szegedy , Sergey Ioffe , Vincent Vanhoucke , and Alexander Alemi . 2017. Inception-v4 , inception-resnet and the impact of residual connections on learning. AAAI ( 2017 ). Christian Szegedy, Sergey Ioffe, Vincent Vanhoucke, and Alexander Alemi. 2017. Inception-v4, inception-resnet and the impact of residual connections on learning. AAAI (2017)."},{"key":"e_1_3_2_2_48_1","volume-title":"Int. Conf. Mach. Learn. 97","author":"Tan Mingxing","year":"2017","unstructured":"Mingxing Tan and Quoc Le . 2017 . Efficientnet: Rethinking model scaling for convolutional neural networks . Int. Conf. Mach. Learn. 97 (2017), 6105--6114. Mingxing Tan and Quoc Le. 2017. Efficientnet: Rethinking model scaling for convolutional neural networks. Int. Conf. Mach. Learn. 97 (2017), 6105--6114."},{"key":"e_1_3_2_2_49_1","volume-title":"AANet: Attribute Attention Network for Person Re-Identifications. IEEE Conf. Comput. Vis. Pattern Recog.","author":"Tay Chiat-Pin","year":"2019","unstructured":"Chiat-Pin Tay , Sharmili Roy , and Kim-Hui Yap . 2019 . AANet: Attribute Attention Network for Person Re-Identifications. IEEE Conf. Comput. Vis. Pattern Recog. (2019), 7134--7143. Chiat-Pin Tay, Sharmili Roy, and Kim-Hui Yap. 2019. AANet: Attribute Attention Network for Person Re-Identifications. IEEE Conf. Comput. Vis. Pattern Recog. (2019), 7134--7143."},{"key":"e_1_3_2_2_50_1","volume-title":"Training data-efficient image transformers & distillation through attention. arXiv preprint arXiv:2012.12877","author":"Touvron Hugo","year":"2020","unstructured":"Hugo Touvron , Matthieu Cord , Matthijs Douze , Francisco Massa , Alexandre Sablayrolles , and Herv\u00b4e J\u00b4egou . 2020. Training data-efficient image transformers & distillation through attention. arXiv preprint arXiv:2012.12877 ( 2020 ). Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Herv\u00b4e J\u00b4egou. 2020. Training data-efficient image transformers & distillation through attention. arXiv preprint arXiv:2012.12877 (2020)."},{"key":"e_1_3_2_2_51_1","volume-title":"Attention is all you need. arXiv preprint arXiv:1706.03762","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani , Noam Shazeer , Niki Parmar , Jakob Uszkoreit , Llion Jones , Aidan N Gomez , Lukasz Kaiser , and Illia Polosukhin . 2017. Attention is all you need. arXiv preprint arXiv:1706.03762 ( 2017 ). Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. arXiv preprint arXiv:1706.03762 (2017)."},{"key":"e_1_3_2_2_52_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.683"},{"key":"e_1_3_2_2_53_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v33i01.33018933"},{"key":"e_1_3_2_2_54_1","volume-title":"ACM Int. Conf. Multimedia","author":"Yuan Yufeng","year":"2018","unstructured":"GuanshuoWang, Yufeng Yuan , Xiong Chen , Jiwei Li , and Xi Zhou . 2018 . Learning discriminative features with multiple granularities for person re-identification . ACM Int. Conf. Multimedia (2018), 274--282. GuanshuoWang, Yufeng Yuan, Xiong Chen, Jiwei Li, and Xi Zhou. 2018. Learning discriminative features with multiple granularities for person re-identification. ACM Int. Conf. Multimedia (2018), 274--282."},{"key":"e_1_3_2_2_55_1","volume-title":"PVTv2: Improved Baselines with Pyramid Vision Transformer. arXiv preprint arXiv:2106.13797","author":"Wang Wenhai","year":"2021","unstructured":"Wenhai Wang , Enze Xie , Xiang Li , Deng-Ping Fan , Kaitao Song , Ding Liang , Tong Lu , Ping Luo , and Ling Shao . 2021. PVTv2: Improved Baselines with Pyramid Vision Transformer. arXiv preprint arXiv:2106.13797 ( 2021 ). Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, and Ling Shao. 2021. PVTv2: Improved Baselines with Pyramid Vision Transformer. arXiv preprint arXiv:2106.13797 (2021)."},{"key":"e_1_3_2_2_56_1","volume-title":"Int. Conf. Comput. Vis.","author":"Xie Enze","year":"2021","unstructured":"WenhaiWang, Enze Xie , Xiang Li , Deng-Ping Fan , Kaitao Song , Ding Liang , Tong Lu , Ping Luo , and Ling Shao . 2021 . Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions . Int. Conf. Comput. Vis. (2021). WenhaiWang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, and Ling Shao. 2021. Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions. Int. Conf. Comput. Vis. (2021)."},{"key":"e_1_3_2_2_57_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00016"},{"key":"e_1_3_2_2_58_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01234-2_1"},{"key":"e_1_3_2_2_59_1","volume-title":"Cvt: Introducing convolutions to vision transformers. arXiv preprint arXiv:2103.15808","author":"Wu Haiping","year":"2021","unstructured":"Haiping Wu , Bin Xiao , Noel Codella , Mengchen Liu , Xiyang Dai , Lu Yuan , and Lei Zhang . 2021 . Cvt: Introducing convolutions to vision transformers. arXiv preprint arXiv:2103.15808 (2021). Haiping Wu, Bin Xiao, Noel Codella, Mengchen Liu, Xiyang Dai, Lu Yuan, and Lei Zhang. 2021. Cvt: Introducing convolutions to vision transformers. arXiv preprint arXiv:2103.15808 (2021)."},{"key":"e_1_3_2_2_60_1","volume-title":"Early Convolutions Help Transformers See Better. arXiv preprint arXiv:2106.14881","author":"Xiao Tete","year":"2021","unstructured":"Tete Xiao , Mannat Singh , Eric Mintun , Trevor Darrell , Piotr Doll\u00e1r , and Ross Girshick . 2021. Early Convolutions Help Transformers See Better. arXiv preprint arXiv:2106.14881 ( 2021 ). Tete Xiao, Mannat Singh, Eric Mintun, Trevor Darrell, Piotr Doll\u00e1r, and Ross Girshick. 2021. Early Convolutions Help Transformers See Better. arXiv preprint arXiv:2106.14881 (2021)."},{"key":"e_1_3_2_2_61_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.634"},{"key":"e_1_3_2_2_62_1","volume-title":"Hoi","author":"Ye Mang","year":"2021","unstructured":"Mang Ye , Jianbing Shen , Gaojie Lin , Tao Xiang , Ling Shao , and Steven C. H . Hoi . 2021 . Deep learning for person re-identification: A survey and outlook. IEEE Trans. Pattern Anal. Mach. Intell . (2021), 1--1. Mang Ye, Jianbing Shen, Gaojie Lin, Tao Xiang, Ling Shao, and Steven C. H. Hoi. 2021. Deep learning for person re-identification: A survey and outlook. IEEE Trans. Pattern Anal. Mach. Intell. (2021), 1--1."},{"key":"e_1_3_2_2_63_1","volume-title":"Yujun Shi Weihao Yu, Francis EH Tay, Jiashi Feng, and Shuicheng Yan.","author":"Yuan Li","year":"2021","unstructured":"Li Yuan , Tao Wang Yunpeng Chen , Yujun Shi Weihao Yu, Francis EH Tay, Jiashi Feng, and Shuicheng Yan. 2021 . Tokens-to-token vit: Training vision transformers from scratch on imagenet. arXiv preprint arXiv:2101.11986 (2021). Li Yuan, Tao Wang Yunpeng Chen, Yujun Shi Weihao Yu, Francis EH Tay, Jiashi Feng, and Shuicheng Yan. 2021. Tokens-to-token vit: Training vision transformers from scratch on imagenet. arXiv preprint arXiv:2101.11986 (2021)."},{"key":"e_1_3_2_2_64_1","volume-title":"HAT: Hierarchical Aggregation Transformers for Person Re-identification. ACM Int. Conf. Multimedia","author":"Zhang Guowen","year":"2021","unstructured":"Guowen Zhang , Pingping Zhang , Jinqing Qi , and Huchuan Lu . 2021 . HAT: Hierarchical Aggregation Transformers for Person Re-identification. ACM Int. Conf. Multimedia (2021), 516--525. Guowen Zhang, Pingping Zhang, Jinqing Qi, and Huchuan Lu. 2021. HAT: Hierarchical Aggregation Transformers for Person Re-identification. ACM Int. Conf. Multimedia (2021), 516--525."},{"key":"e_1_3_2_2_65_1","volume-title":"Making Convolutional Networks Shift-Invariant Again. Int. Conf. Mach. Learn.","author":"Zhang Richard","year":"2019","unstructured":"Richard Zhang . 2019 . Making Convolutional Networks Shift-Invariant Again. Int. Conf. Mach. Learn. (2019), 7324--7334. Richard Zhang. 2019. Making Convolutional Networks Shift-Invariant Again. Int. Conf. Mach. Learn. (2019), 7324--7334."},{"key":"e_1_3_2_2_66_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.133"},{"key":"e_1_3_2_2_67_1","doi-asserted-by":"publisher","DOI":"10.1145\/3159171"},{"key":"e_1_3_2_2_68_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.389"},{"key":"e_1_3_2_2_69_1","volume-title":"Domain Generalization: A Survey. arXiv preprint arXiv:2103.02503","author":"Zhou Kaiyang","year":"2021","unstructured":"Kaiyang Zhou , Ziwei Liu , Yu Qiao , Tao Xiang , and Chen Change Loy . 2021 . Domain Generalization: A Survey. arXiv preprint arXiv:2103.02503 (2021). Kaiyang Zhou, Ziwei Liu, Yu Qiao, Tao Xiang, and Chen Change Loy. 2021. Domain Generalization: A Survey. arXiv preprint arXiv:2103.02503 (2021)."},{"key":"e_1_3_2_2_70_1","volume-title":"Omni- Scale Feature Learning for Person Re-Identification. Int. Conf. Comput. Vis.","author":"Zhou Kaiyang","year":"2019","unstructured":"Kaiyang Zhou , Yongxin Yang , Andrea Cavallaro , and Tao Xiang . 2019 . Omni- Scale Feature Learning for Person Re-Identification. Int. Conf. Comput. Vis. (2019), 3702--3712. Kaiyang Zhou, Yongxin Yang, Andrea Cavallaro, and Tao Xiang. 2019. Omni- Scale Feature Learning for Person Re-Identification. Int. Conf. Comput. Vis. (2019), 3702--3712."},{"key":"e_1_3_2_2_71_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2021.3069237"},{"key":"e_1_3_2_2_72_1","volume-title":"Domain Generalization with MixStyle. Int. Conf. Mach. Learn.","author":"Zhou Kaiyang","year":"2021","unstructured":"Kaiyang Zhou , Yongxin Yang , Yu Qiao , and Tao Xiang . 2021 . Domain Generalization with MixStyle. Int. Conf. Mach. Learn. (2021). Kaiyang Zhou, Yongxin Yang, Yu Qiao, and Tao Xiang. 2021. Domain Generalization with MixStyle. Int. Conf. Mach. Learn. (2021)."},{"key":"e_1_3_2_2_73_1","volume-title":"AAformer: Auto-Aligned Transformer for Person Re-Identification. arXiv preprint arXiv:2104.00921","author":"Zhu Kuan","year":"2021","unstructured":"Kuan Zhu , Haiyun Guo , Shiliang Zhang , Yaowei Wang , Gaopan Huang , Honglin Qiao , Jing Liu , Jinqiao Wang , and Ming Tang . 2021. AAformer: Auto-Aligned Transformer for Person Re-Identification. arXiv preprint arXiv:2104.00921 ( 2021 ). Kuan Zhu, Haiyun Guo, Shiliang Zhang, Yaowei Wang, Gaopan Huang, Honglin Qiao, Jing Liu, Jinqiao Wang, and Ming Tang. 2021. AAformer: Auto-Aligned Transformer for Person Re-Identification. arXiv preprint arXiv:2104.00921 (2021)."},{"key":"e_1_3_2_2_74_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i07.7014"},{"key":"e_1_3_2_2_75_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58610-2_9"}],"event":{"name":"MM '22: The 30th ACM International Conference on Multimedia","location":"Lisboa Portugal","acronym":"MM '22","sponsor":["SIGMM ACM Special Interest Group on Multimedia"]},"container-title":["Proceedings of the 30th ACM International Conference on Multimedia"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3503161.3548770","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3503161.3548770","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T17:49:18Z","timestamp":1750182558000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3503161.3548770"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,10,10]]},"references-count":75,"alternative-id":["10.1145\/3503161.3548770","10.1145\/3503161"],"URL":"https:\/\/doi.org\/10.1145\/3503161.3548770","relation":{},"subject":[],"published":{"date-parts":[[2022,10,10]]},"assertion":[{"value":"2022-10-10","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}