{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T05:04:01Z","timestamp":1750309441137,"version":"3.41.0"},"reference-count":60,"publisher":"Association for Computing Machinery (ACM)","issue":"11","license":[{"start":{"date-parts":[[2024,11,14]],"date-time":"2024-11-14T00:00:00Z","timestamp":1731542400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62272144, U22A2094, 72188101, 62020106007, 62076086, and U20A20183"],"award-info":[{"award-number":["62272144, U22A2094, 72188101, 62020106007, 62076086, and U20A20183"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Major Project of Anhui Province","award":["202203a05020011"],"award-info":[{"award-number":["202203a05020011"]}]},{"DOI":"10.13039\/501100012226","name":"Fundamental Research Funds for the Central Universities","doi-asserted-by":"crossref","award":["JZ2024HGTG0309, JZ2024AHST0337, and JZ2023YQTD0072"],"award-info":[{"award-number":["JZ2024HGTG0309, JZ2024AHST0337, and JZ2023YQTD0072"]}],"id":[{"id":"10.13039\/501100012226","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2024,11,30]]},"abstract":"<jats:p>\n            Gaze following aims to predict where a person is looking in a scene. Existing methods tend to prioritize traditional 2D RGB visual cues or require burdensome prior knowledge and extra expensive datasets annotated in 3D coordinate systems to train specialized modules to enhance scene modeling. In this work, we introduce a novel framework deployed on a simple ResNet backbone, which exclusively uses image and depth maps to mimic human visual preferences and realize 3D-like depth perception. We first leverage depth maps to formulate spatial-based proximity information regarding the objects with the target person. This process sharpens the focus of the gaze cone on the specific region of interest pertaining to the target while diminishing the impact of surrounding distractions. To capture the diverse dependence of scene context on the saliency gaze cone, we then introduce a learnable grid-level regularized attention that anticipates coarse-grained regions of interest, thereby refining the mapping of the saliency feature to pixel-level heatmaps. This allows our model to better account for individual differences when predicting others\u2019 gaze locations. Finally, we employ the KL-divergence loss to super the grid-level regularized attention, which combines the gaze direction, heatmap regression, and in\/out classification losses, providing comprehensive supervision for model optimization. Experimental results on two publicly available datasets demonstrate the comparable performance of our model with less help of modal information. Quantitative visualization results further validate the interpretability of our method. The source code will be available at\n            <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" ext-link-type=\"url\" xlink:href=\"https:\/\/github.com\/VUT-HFUT\/DepthMatters\">https:\/\/github.com\/VUT-HFUT\/DepthMatters<\/jats:ext-link>\n            .\n          <\/jats:p>","DOI":"10.1145\/3689643","type":"journal-article","created":{"date-parts":[[2024,8,26]],"date-time":"2024-08-26T15:19:40Z","timestamp":1724685580000},"page":"1-24","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Depth Matters: Spatial Proximity-Based Gaze Cone Generation for Gaze Following in Wild"],"prefix":"10.1145","volume":"20","author":[{"ORCID":"https:\/\/orcid.org\/0009-0007-7057-8739","authenticated-orcid":false,"given":"Feiyang","family":"Liu","sequence":"first","affiliation":[{"name":"Hefei University of Technology, Hefei, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5083-2145","authenticated-orcid":false,"given":"Kun","family":"Li","sequence":"additional","affiliation":[{"name":"Hefei University of Technology, Hefei, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8202-0544","authenticated-orcid":false,"given":"Zhun","family":"Zhong","sequence":"additional","affiliation":[{"name":"University of Nottingham, Nottingham, United Kingdom"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5628-6237","authenticated-orcid":false,"given":"Wei","family":"Jia","sequence":"additional","affiliation":[{"name":"Hefei University of Technology, Hefei, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3514-5413","authenticated-orcid":false,"given":"Bin","family":"Hu","sequence":"additional","affiliation":[{"name":"Lanzhou University, Lanzhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0201-1638","authenticated-orcid":false,"given":"Xun","family":"Yang","sequence":"additional","affiliation":[{"name":"University of Science and Technology of China, Hefei, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3094-7735","authenticated-orcid":false,"given":"Meng","family":"Wang","sequence":"additional","affiliation":[{"name":"Hefei University of Technology, Hefei, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2594-254X","authenticated-orcid":false,"given":"Dan","family":"Guo","sequence":"additional","affiliation":[{"name":"Hefei University of Technology, Hefei, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2024,11,14]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.1145\/1314303.1314307"},{"key":"e_1_3_1_3_2","first-page":"14126","article-title":"ESCNET: Gaze target detection with the understanding of 3d scenes","author":"Bao Jun","year":"2022","unstructured":"Jun Bao, Buyu Liu, and Jun Yu. 2022. ESCNET: Gaze target detection with the understanding of 3d scenes. In CVPR, 14126\u201314135.","journal-title":"CVPR"},{"key":"e_1_3_1_4_2","first-page":"612","article-title":"Multiple-gaze geometry: Inferring novel 3d locations from gazes observed in monocular video","author":"Brau Ernesto","year":"2018","unstructured":"Ernesto Brau, Jinyan Guan, Tanya Jeffries, and Kobus Barnard. 2018. Multiple-gaze geometry: Inferring novel 3d locations from gazes observed in monocular video. In ECCV, 612\u2013630.","journal-title":"ECCV"},{"key":"e_1_3_1_5_2","first-page":"213","article-title":"End-to-end object detection with transformers","author":"Carion Nicolas","year":"2020","unstructured":"Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. 2020. End-to-end object detection with transformers. In ECCV, 213\u2013229.","journal-title":"ECCV"},{"key":"e_1_3_1_6_2","first-page":"5410","article-title":"Pyramid stereo matching network","author":"Chang Jia-Ren","year":"2018","unstructured":"Jia-Ren Chang and Yong-Sheng Chen. 2018. Pyramid stereo matching network. In CVPR, 5410\u20135418.","journal-title":"CVPR"},{"issue":"1","key":"e_1_3_1_7_2","first-page":"1","article-title":"Dynamic message propagation network for RGB-D and video salient object detection","volume":"20","author":"Chen Baian","year":"2023","unstructured":"Baian Chen, Zhilei Chen, Xiaowei Hu, Jun Xu, Haoran Xie, Jing Qin, and Mingqiang Wei. 2023. Dynamic message propagation network for RGB-D and video salient object detection. ACM TOMM 20, 1 (2023), 1\u201321.","journal-title":"ACM TOMM"},{"key":"e_1_3_1_8_2","first-page":"5259","article-title":"Gaze estimation by exploring two-eye asymmetry","volume":"29","author":"Cheng Yihua","year":"2020","unstructured":"Yihua Cheng, Xucong Zhang, Feng Lu, and Yoichi Sato. 2020. Gaze estimation by exploring two-eye asymmetry. IEEE TIP 29 (2020), 5259\u20135272.","journal-title":"IEEE TIP"},{"key":"e_1_3_1_9_2","first-page":"383","article-title":"Connecting gaze, scene, and attention: Generalized attention estimation via joint modeling of gaze and scene saliency","author":"Chong Eunji","year":"2018","unstructured":"Eunji Chong, Nataniel Ruiz, Yongxin Wang, Yun Zhang, Agata Rozga, and James M. Rehg. 2018. Connecting gaze, scene, and attention: Generalized attention estimation via joint modeling of gaze and scene saliency. In ECCV, 383\u2013398.","journal-title":"ECCV"},{"key":"e_1_3_1_10_2","first-page":"5396","article-title":"Detecting attended visual targets in video","author":"Chong Eunji","year":"2020","unstructured":"Eunji Chong, Yongxin Wang, Nataniel Ruiz, and James M. Rehg. 2020. Detecting attended visual targets in video. In CVPR, 5396\u20135406.","journal-title":"CVPR"},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2012.49"},{"key":"e_1_3_1_12_2","first-page":"6460","article-title":"Inferring shared attention in social scene videos","author":"Fan Lifeng","year":"2018","unstructured":"Lifeng Fan, Yixin Chen, Ping Wei, Wenguan Wang, and Song-Chun Zhu. 2018. Inferring shared attention in social scene videos. In CVPR, 6460\u20136468.","journal-title":"CVPR"},{"key":"e_1_3_1_13_2","first-page":"11390","article-title":"Dual attention guided gaze target detection in the wild","author":"Fang Yi","year":"2021","unstructured":"Yi Fang, Jiapeng Tang, Wang Shen, Wei Shen, Xiao Gu, Li Song, and Guangtao Zhai. 2021. Dual attention guided gaze target detection in the wild. In CVPR, 11390\u201311399.","journal-title":"CVPR"},{"key":"e_1_3_1_14_2","first-page":"314","article-title":"Learning to recognize daily actions using gaze","author":"Fathi Alireza","year":"2012","unstructured":"Alireza Fathi, Yin Li, and James M. Rehg. 2012. Learning to recognize daily actions using gaze. In ECCV, 314\u2013327.","journal-title":"ECCV"},{"key":"e_1_3_1_15_2","first-page":"15","article-title":"SVO: Fast semi-direct monocular visual odometry","author":"Forster Christian","year":"2014","unstructured":"Christian Forster, Matia Pizzoli, and Davide Scaramuzza. 2014. SVO: Fast semi-direct monocular visual odometry. In ICRA, 15\u201322.","journal-title":"ICRA"},{"key":"e_1_3_1_16_2","article-title":"Learning with a Wasserstein loss","volume":"28","author":"Frogner Charlie","year":"2015","unstructured":"Charlie Frogner, Chiyuan Zhang, Hossein Mobahi, Mauricio Araya, and Tomaso A. Poggio. 2015. Learning with a Wasserstein loss. In NeurIPS, Vol. 28.","journal-title":"NeurIPS,"},{"key":"e_1_3_1_17_2","first-page":"255","article-title":"Eyediap: A database for the development and evaluation of gaze estimation algorithms from RGB and RGB-D cameras","author":"Mora Kenneth Alberto Funes","year":"2014","unstructured":"Kenneth Alberto Funes Mora, Florent Monay, and Jean-Marc Odobez. 2014. Eyediap: A database for the development and evaluation of gaze estimation algorithms from RGB and RGB-D cameras. In ETRA, 255\u2013258.","journal-title":"ETRA"},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.2307\/1574154"},{"key":"e_1_3_1_19_2","first-page":"7297","article-title":"Densepose: Dense human pose estimation in the wild","author":"G\u00fcler Ri\u0307za Alp","year":"2018","unstructured":"Ri\u0307za Alp G\u00fcler, Natalia Neverova, and Iasonas Kokkinos. 2018. Densepose: Dense human pose estimation in the wild. In CVPR, 7297\u20137306.","journal-title":"CVPR"},{"issue":"7","key":"e_1_3_1_20_2","first-page":"6238","article-title":"Benchmarking micro-action recognition: Dataset, methods, and applications","volume":"34","author":"Guo Dan","year":"2024","unstructured":"Dan Guo, Kun Li, Bin Hu, Yan Zhang, and Meng Wang. 2024. Benchmarking micro-action recognition: Dataset, methods, and applications. IEEE TCSVT 34, 7 (2024), 6238\u20136252.","journal-title":"IEEE TCSVT"},{"key":"e_1_3_1_21_2","first-page":"770","article-title":"Deep residual learning for image recognition","author":"He Kaiming","year":"2016","unstructured":"Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep residual learning for image recognition. In CVPR, 770\u2013778.","journal-title":"CVPR"},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.1145\/3311784"},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.1016\/S0042-6989(99)00163-7"},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.1017\/S0140525X18001826"},{"key":"e_1_3_1_25_2","first-page":"01","article-title":"Multi-person gaze-following with numerical coordinate regression","author":"Jin Tianlei","year":"2021","unstructured":"Tianlei Jin, Zheyuan Lin, Shiqiang Zhu, Wen Wang, and Shunda Hu. 2021. Multi-person gaze-following with numerical coordinate regression. In IEEE FG, 01\u201308.","journal-title":"IEEE FG"},{"key":"e_1_3_1_26_2","doi-asserted-by":"crossref","unstructured":"Tianlei Jin Qizhi Yu Shiqiang Zhu Zheyuan Lin Jie Ren Yuanhai Zhou and Wei Song. 2022. Depth-aware gaze-following via auxiliary networks for robotics. EAAI 113 (2022) 104924.","DOI":"10.1016\/j.engappai.2022.104924"},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2009.5459462"},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00701"},{"key":"e_1_3_1_29_2","first-page":"2100","article-title":"Dense visual SLAM for RGB-D cameras","author":"Kerl Christian","year":"2013","unstructured":"Christian Kerl, J\u00fcrgen Sturm, and Daniel Cremers. 2013. Dense visual SLAM for RGB-D cameras. In IROS, 2100\u20132106.","journal-title":"IROS"},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v35i3.16285"},{"issue":"10","key":"e_1_3_1_31_2","first-page":"202102","article-title":"ViGT: Proposal-free video grounding with a learnable token in the transformer","volume":"66","author":"Li Kun","year":"2023","unstructured":"Kun Li, Dan Guo, and Meng Wang. 2023. ViGT: Proposal-free video grounding with a learnable token in the transformer. SCIS 66, 10 (2023), 202102.","journal-title":"SCIS"},{"issue":"6","key":"e_1_3_1_32_2","first-page":"1","article-title":"Transformer-based visual grounding with cross-modality interaction","volume":"19","author":"Li Kun","year":"2023","unstructured":"Kun Li, Jiaxiu Li, Dan Guo, Xun Yang, and Meng Wang. 2023. Transformer-based visual grounding with cross-modality interaction. ACM TOMM 19, 6 (2023), 1\u201319.","journal-title":"ACM TOMM"},{"key":"e_1_3_1_33_2","first-page":"4521","article-title":"Learning the depths of moving people by watching frozen people","author":"Li Zhengqi","year":"2019","unstructured":"Zhengqi Li, Tali Dekel, Forrester Cole, Richard Tucker, Noah Snavely, Ce Liu, and William T. Freeman. 2019. Learning the depths of moving people by watching frozen people. In CVPR, 4521\u20134530.","journal-title":"CVPR"},{"key":"e_1_3_1_34_2","first-page":"35","article-title":"Believe it or not, we know what you are looking at!","author":"Lian Dongze","year":"2018","unstructured":"Dongze Lian, Zehao Yu, and Shenghua Gao. 2018. Believe it or not, we know what you are looking at!. In ACCV, 35\u201350.","journal-title":"ACCV"},{"key":"e_1_3_1_35_2","first-page":"740","article-title":"Microsoft COCO: Common objects in context","author":"Lin Tsung-Yi","year":"2014","unstructured":"Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll\u00e1r, and C. Lawrence Zitnick. 2014. Microsoft COCO: Common objects in context. In ECCV, 740\u2013755.","journal-title":"ECCV"},{"issue":"5","key":"e_1_3_1_36_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3638559","article-title":"DPDFormer: A coarse-to-fine model for monocular depth estimation","volume":"20","author":"Liu Chunpu","year":"2024","unstructured":"Chunpu Liu, Guanglei Yang, Wangmeng Zuo, and Tianyi Zang. 2024. DPDFormer: A coarse-to-fine model for monocular depth estimation. ACM TOMM 20, 5 (2024), 1\u201321.","journal-title":"ACM TOMM"},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.1016\/B978-0-323-99864-2.00016-0"},{"key":"e_1_3_1_38_2","doi-asserted-by":"crossref","first-page":"971","DOI":"10.1007\/978-3-030-20085-5_23","article-title":"Eye movements and human-computer interaction","author":"Majaranta P\u00e4ivi","year":"2019","unstructured":"P\u00e4ivi Majaranta, Kari-Jouko R\u00e4ih\u00e4, Aulikki Hyrskykari, and Oleg \u0160pakov. 2019. Eye movements and human-computer interaction. Eye movement research: An introduction to its scientific foundations and applications (2019), 971\u20131015.","journal-title":"Eye movement research: An introduction to its scientific foundations and applications"},{"key":"e_1_3_1_39_2","first-page":"3477","article-title":"Laeo-net: revisiting people looking at each other in videos","author":"Marin-Jimenez Manuel J.","year":"2019","unstructured":"Manuel J. Marin-Jimenez, Vicky Kalogeiton, Pablo Medina-Suarez, and Andrew Zisserman. 2019. Laeo-net: revisiting people looking at each other in videos. In CVPR, 3477\u20133485.","journal-title":"CVPR"},{"key":"e_1_3_1_40_2","first-page":"11","article-title":"Tracking gaze and visual focus of attention of people involved in social interaction","volume":"40","author":"Mass\u00e9 Beno\u00eet","year":"2017","unstructured":"Beno\u00eet Mass\u00e9, Sil\u00e8ye Ba, and Radu Horaud. 2017. Tracking gaze and visual focus of attention of people involved in social interaction. IEEE TPAMI 40, 11 (2017), 2711\u20132724.","journal-title":"IEEE TPAMI"},{"key":"e_1_3_1_41_2","first-page":"1","article-title":"Extended gaze following: Detecting objects in videos beyond the camera field of view","author":"Mass\u00e9 Benoit","year":"2019","unstructured":"Benoit Mass\u00e9, St\u00e9phane Lathuili\u00e8re, Pablo Mesejo, and Radu Horaud. 2019. Extended gaze following: Detecting objects in videos beyond the camera field of view. In IEEE FG, 1\u20138.","journal-title":"IEEE FG"},{"key":"e_1_3_1_42_2","first-page":"120","article-title":"Single-shot multi-person 3d pose estimation from monocular RGB","author":"Mehta Dushyant","year":"2018","unstructured":"Dushyant Mehta, Oleksandr Sotnychenko, Franziska Mueller, Weipeng Xu, Srinath Sridhar, Gerard Pons-Moll, and Christian Theobalt. 2018. Single-shot multi-person 3d pose estimation from monocular RGB. In 3DV, 120\u2013130.","journal-title":"3DV"},{"key":"e_1_3_1_43_2","doi-asserted-by":"publisher","DOI":"10.1016\/S0016-0032(96)00063-4"},{"key":"e_1_3_1_44_2","first-page":"880","article-title":"Patch-level gaze distribution prediction for gaze following","author":"Miao Qiaomu","year":"2023","unstructured":"Qiaomu Miao, Minh Hoai, and Dimitris Samaras. 2023. Patch-level gaze distribution prediction for gaze following. In WACV, 880\u2013889.","journal-title":"WACV"},{"key":"e_1_3_1_45_2","first-page":"598","article-title":"Shallow and deep convolutional networks for saliency prediction","author":"Pan Junting","year":"2016","unstructured":"Junting Pan, Elisa Sayrol, Xavier Giro-i Nieto, Kevin McGuinness, and Noel E. O\u2019Connor. 2016. Shallow and deep convolutional networks for saliency prediction. In CVPR, 598\u2013606.","journal-title":"CVPR"},{"key":"e_1_3_1_46_2","article-title":"Stand-alone self-attention in vision models","volume":"32","author":"Ramachandran Prajit","year":"2019","unstructured":"Prajit Ramachandran, Niki Parmar, Ashish Vaswani, Irwan Bello, Anselm Levskaya, and Jon Shlens. 2019. Stand-alone self-attention in vision models. In NeurIPS, Vol. 32.","journal-title":"NeurIPS,"},{"key":"e_1_3_1_47_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2020.3019967"},{"key":"e_1_3_1_48_2","first-page":"1","article-title":"Where are they looking?","author":"Recasens Adria","year":"2015","unstructured":"Adria Recasens, Aditya Khosla, Carl Vondrick, and Antonio Torralba. 2015. Where are they looking? In NeurIPS, 1\u20139.","journal-title":"NeurIPS"},{"key":"e_1_3_1_49_2","first-page":"1435","article-title":"Following gaze in video","author":"Recasens Adria","year":"2017","unstructured":"Adria Recasens, Carl Vondrick, Aditya Khosla, and Antonio Torralba. 2017. Following gaze in video. In ICCV, 1435\u20131443.","journal-title":"ICCV"},{"key":"e_1_3_1_50_2","first-page":"3327","article-title":"Attention flow: End-to-end joint attention estimation","author":"Sumer Omer","year":"2020","unstructured":"Omer Sumer, Peter Gerjets, Ulrich Trautwein, and Enkelejda Kasneci. 2020. Attention flow: End-to-end joint attention estimation. In WACV, 3327\u20133336.","journal-title":"WACV"},{"key":"e_1_3_1_51_2","first-page":"6243","article-title":"CNN-SLAM: Real-time dense monocular slam with learned depth prediction","author":"Tateno Keisuke","year":"2017","unstructured":"Keisuke Tateno, Federico Tombari, Iro Laina, and Nassir Navab. 2017. CNN-SLAM: Real-time dense monocular slam with learned depth prediction. In CVPR, 6243\u20136252.","journal-title":"CVPR"},{"key":"e_1_3_1_52_2","first-page":"21860","article-title":"Object-aware gaze target detection","author":"Tonini Francesco","year":"2023","unstructured":"Francesco Tonini, Nicola Dall\u2019Asen, Cigdem Beyan, and Elisa Ricci. 2023. Object-aware gaze target detection. In ICCV, 21860\u201321869.","journal-title":"ICCV"},{"key":"e_1_3_1_53_2","first-page":"2192","article-title":"End-to-end human-gaze-target detection with transformers","author":"Tu Danyang","year":"2022","unstructured":"Danyang Tu, Xiongkuo Min, Huiyu Duan, Guodong Guo, Guangtao Zhai, and Wei Shen. 2022. End-to-end human-gaze-target detection with transformers. In CVPR, 2192\u20132200.","journal-title":"CVPR"},{"key":"e_1_3_1_54_2","first-page":"5998","article-title":"Attention is all you need","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In NeurIPS, 5998\u20136008.","journal-title":"NeurIPS"},{"key":"e_1_3_1_55_2","first-page":"7364","article-title":"Dada: Depth-aware domain adaptation in semantic segmentation","author":"Vu Tuan-Hung","year":"2019","unstructured":"Tuan-Hung Vu, Himalaya Jain, Maxime Bucher, Matthieu Cord, and Patrick P\u00e9rez. 2019. Dada: Depth-aware domain adaptation in semantic segmentation. In ICCV, 7364\u20137373.","journal-title":"ICCV"},{"key":"e_1_3_1_56_2","first-page":"454","article-title":"Depth-conditioned dynamic message propagation for monocular 3d object detection","author":"Wang Li","year":"2021","unstructured":"Li Wang, Liang Du, Xiaoqing Ye, Yanwei Fu, Guodong Guo, Xiangyang Xue, Jianfeng Feng, and Li Zhang. 2021. Depth-conditioned dynamic message propagation for monocular 3d object detection. In CVPR, 454\u2013463.","journal-title":"CVPR"},{"key":"e_1_3_1_57_2","first-page":"6801","article-title":"Where and why are they looking? Jointly inferring human attention and intentions in complex tasks","author":"Wei Ping","year":"2018","unstructured":"Ping Wei, Yang Liu, Tianmin Shu, Nanning Zheng, and Song-Chun Zhu. 2018. Where and why are they looking? Jointly inferring human attention and intentions in complex tasks. In CVPR, 6801\u20136809.","journal-title":"CVPR"},{"key":"e_1_3_1_58_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-019-01263-4"},{"key":"e_1_3_1_59_2","first-page":"550","article-title":"Smap: Single-shot multi-person absolute 3d pose estimation","author":"Zhen Jianan","year":"2020","unstructured":"Jianan Zhen, Qi Fang, Jiaming Sun, Wentao Liu, Wei Jiang, Hujun Bao, and Xiaowei Zhou. 2020. Smap: Single-shot multi-person absolute 3d pose estimation. In ECCV, 550\u2013566.","journal-title":"ECCV"},{"key":"e_1_3_1_60_2","article-title":"Learning deep features for scene recognition using places database","volume":"27","author":"Zhou Bolei","year":"2014","unstructured":"Bolei Zhou, Agata Lapedriza, Jianxiong Xiao, Antonio Torralba, and Aude Oliva. 2014. Learning deep features for scene recognition using places database. In NeurIPS, Vol. 27.","journal-title":"NeurIPS,"},{"issue":"10","key":"e_1_3_1_61_2","first-page":"3637","article-title":"MUGGLE: MUlti-stream group gaze learning and estimation","volume":"30","author":"Zhuang Ning","year":"2019","unstructured":"Ning Zhuang, Bingbing Ni, Yi Xu, Xiaokang Yang, Wenjun Zhang, Zefan Li, and Wen Gao. 2019. MUGGLE: MUlti-stream group gaze learning and estimation. IEEE TCSVT 30, 10 (2019), 3637\u20133650.","journal-title":"IEEE TCSVT"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3689643","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3689643","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T01:09:47Z","timestamp":1750295387000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3689643"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,11,14]]},"references-count":60,"journal-issue":{"issue":"11","published-print":{"date-parts":[[2024,11,30]]}},"alternative-id":["10.1145\/3689643"],"URL":"https:\/\/doi.org\/10.1145\/3689643","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"type":"print","value":"1551-6857"},{"type":"electronic","value":"1551-6865"}],"subject":[],"published":{"date-parts":[[2024,11,14]]},"assertion":[{"value":"2024-03-06","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-08-13","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-11-14","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}