{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,19]],"date-time":"2026-06-19T04:10:11Z","timestamp":1781842211718,"version":"3.54.5"},"reference-count":36,"publisher":"MDPI AG","issue":"9","license":[{"start":{"date-parts":[[2025,8,24]],"date-time":"2025-08-24T00:00:00Z","timestamp":1755993600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Natural Science Foundation of Fujian Province, China","award":["2023J01803"],"award-info":[{"award-number":["2023J01803"]}]},{"name":"Natural Science Foundation of Fujian Province, China","award":["2025J01345"],"award-info":[{"award-number":["2025J01345"]}]},{"name":"Natural Science Foundation of Fujian Province, China","award":["3502Z202371019"],"award-info":[{"award-number":["3502Z202371019"]}]},{"name":"Natural Science Foundation of Fujian Province, China","award":["42371457"],"award-info":[{"award-number":["42371457"]}]},{"name":"Natural Science Foundation of Xiamen, China","award":["2023J01803"],"award-info":[{"award-number":["2023J01803"]}]},{"name":"Natural Science Foundation of Xiamen, China","award":["2025J01345"],"award-info":[{"award-number":["2025J01345"]}]},{"name":"Natural Science Foundation of Xiamen, China","award":["3502Z202371019"],"award-info":[{"award-number":["3502Z202371019"]}]},{"name":"Natural Science Foundation of Xiamen, China","award":["42371457"],"award-info":[{"award-number":["42371457"]}]},{"name":"National Natural Science Foundation of China","award":["2023J01803"],"award-info":[{"award-number":["2023J01803"]}]},{"name":"National Natural Science Foundation of China","award":["2025J01345"],"award-info":[{"award-number":["2025J01345"]}]},{"name":"National Natural Science Foundation of China","award":["3502Z202371019"],"award-info":[{"award-number":["3502Z202371019"]}]},{"name":"National Natural Science Foundation of China","award":["42371457"],"award-info":[{"award-number":["42371457"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Algorithms"],"abstract":"<jats:p>Gaze estimation is a cornerstone of applications such as human\u2013computer interaction and behavioral analysis, e.g., for intelligent transport systems. Nevertheless, existing methods predominantly rely on coarse-grained features from deep layers of visual encoders, overlooking the critical role that fine-grained details from shallow layers play in gaze estimation. To address this gap, we propose a novel Hierarchical Fine-Grained Attention Decoder (HFGAD), a lightweight fine-grained decoder that emphasizes the importance of shallow-layer information in gaze estimation. Specifically, HFGAD integrates a fine-grained amplifier MSCSA that employs multi-scale spatial-channel attention to direct focus toward gaze-relevant regions, and also incorporates a shallow-to-deep fusion module SFM to facilitate interaction between coarse-grained and fine-grained information. Extensive experiments on three benchmark datasets demonstrate the superiority of HFGAD over existing methods, achieving a remarkable 1.13\u00b0 improvement in gaze estimation accuracy for in-car scenarios.<\/jats:p>","DOI":"10.3390\/a18090538","type":"journal-article","created":{"date-parts":[[2025,8,25]],"date-time":"2025-08-25T00:09:53Z","timestamp":1756080593000},"page":"538","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":2,"title":["HFGAD: Hierarchical Fine-Grained Attention Decoder for Gaze Estimation"],"prefix":"10.3390","volume":"18","author":[{"ORCID":"https:\/\/orcid.org\/0009-0005-8912-0627","authenticated-orcid":false,"given":"Shaojie","family":"Huang","sequence":"first","affiliation":[{"name":"School of Computer Engineering, Jimei University, Xiamen 361021, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Tianzhong","family":"Wang","sequence":"additional","affiliation":[{"name":"School of Computer Engineering, Jimei University, Xiamen 361021, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5934-1139","authenticated-orcid":false,"given":"Weiquan","family":"Liu","sequence":"additional","affiliation":[{"name":"School of Computer Engineering, Jimei University, Xiamen 361021, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yingchao","family":"Piao","sequence":"additional","affiliation":[{"name":"Computer Network Information Center, Chinese Academy of Sciences, Beijing 100000, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1707-5685","authenticated-orcid":false,"given":"Jinhe","family":"Su","sequence":"additional","affiliation":[{"name":"School of Computer Engineering, Jimei University, Xiamen 361021, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Guorong","family":"Cai","sequence":"additional","affiliation":[{"name":"School of Computer Engineering, Jimei University, Xiamen 361021, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Huilin","family":"Xu","sequence":"additional","affiliation":[{"name":"School of Computer Engineering, Jimei University, Xiamen 361021, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2025,8,24]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"7509","DOI":"10.1109\/TPAMI.2024.3393571","article-title":"Appearance-based gaze estimation with deep learning: A review and benchmark","volume":"46","author":"Cheng","year":"2024","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Steil, J., Huang, M.X., and Bulling, A. (2018, January 14\u201317). Fixation detection for head-mounted eye tracking based on visual similarity of gaze targets. Proceedings of the 2018 ACM Symposium on Eye Tracking Research & Applications, Warsaw, Poland.","DOI":"10.1145\/3204493.3204538"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"1720","DOI":"10.1109\/TPAMI.2018.2845370","article-title":"Predicting the driver\u2019s focus of attention: The dr (eye) ve project","volume":"41","author":"Palazzi","year":"2018","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"141","DOI":"10.1109\/TIV.2018.2804160","article-title":"Dynamics of Driver\u2019s Gaze: Explorations in Behavior Modeling and Maneuver Prediction","volume":"3","author":"Martin","year":"2018","journal-title":"IEEE Trans. Intell. Veh."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"8000","DOI":"10.1109\/TIV.2024.3405990","article-title":"Driver Distraction Behavior Recognition for Autonomous Driving: Approaches, Datasets and Challenges","volume":"9","author":"Tan","year":"2024","journal-title":"IEEE Trans. Intell. Veh."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Cheng, Y., Huang, S., Wang, F., Qian, C., and Lu, F. (2020, January 7\u201312). A coarse-to-fine adaptive network for appearance-based gaze estimation. Proceedings of the AAAI Conference on Artificial Intelligence, New York, NY, USA.","DOI":"10.1609\/aaai.v34i07.6636"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Zhang, X., Sugano, Y., Fritz, M., and Bulling, A. (2017, January 21\u201326). It\u2019s written all over your face: Full-face appearance-based gaze estimation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, Honolulu, HI, USA.","DOI":"10.1109\/CVPRW.2017.284"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Fischer, T., Chang, H.J., and Demiris, Y. (2018, January 8\u201314). Rt-gene: Real-time eye gaze estimation in natural environments. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01249-6_21"},{"key":"ref_9","unstructured":"Kellnhofer, P., Recasens, A., Stent, S., Matusik, W., and Torralba, A. (November, January 27). Gaze360: Physically unconstrained gaze estimation in the wild. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Seoul, Republic of Korea."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Cheng, Y., Lu, F., and Zhang, X. (2018, January 8\u201314). Appearance-based gaze estimation via evaluation-guided asymmetric regression. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01264-9_7"},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"5259","DOI":"10.1109\/TIP.2020.2982828","article-title":"Gaze estimation by exploring two-eye asymmetry","volume":"29","author":"Cheng","year":"2020","journal-title":"IEEE Trans. Image Process."},{"key":"ref_12","unstructured":"Biswas, P. (2021, January 19\u201325). Appearance-based gaze estimation using attention and difference mechanism. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Cheng, Y., and Lu, F. (2022, January 21\u201325). Gaze estimation using transformer. Proceedings of the 2022 26th International Conference on Pattern Recognition (ICPR), Montr\u00e9al, QC, Canada.","DOI":"10.1109\/ICPR56361.2022.9956687"},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"38410","DOI":"10.1109\/JIOT.2024.3449409","article-title":"Appearance-Based Driver 3D Gaze Estimation Using GRM and Mixed Loss Strategies","volume":"11","author":"Li","year":"2024","journal-title":"IEEE Internet Things J."},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"445","DOI":"10.1007\/s00138-017-0852-4","article-title":"Tabletgaze: Dataset and analysis for unconstrained appearance-based gaze estimation in mobile tablets","volume":"28","author":"Huang","year":"2017","journal-title":"Mach. Vis. Appl."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Cheng, Y., Zhu, Y., Wang, Z., Hao, H., Liu, Y., Cheng, S., Wang, X., and Chang, H.J. (2024, January 17\u201321). What Do You See in Vehicle? Comprehensive Vision Solution for In-Vehicle Gaze Estimation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR52733.2024.00154"},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"4","DOI":"10.1016\/j.cviu.2004.07.010","article-title":"Eye gaze tracking techniques for interactive applications","volume":"98","author":"Morimoto","year":"2005","journal-title":"Comput. Vis. Image Underst."},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"137","DOI":"10.3758\/BF03204486","article-title":"Heuristic filtering and reliable calibration methods for video-based pupil-tracking systems","volume":"25","author":"Stampe","year":"1993","journal-title":"Behav. Res. Methods Instrum. Comput."},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"357","DOI":"10.1006\/rtim.2002.0279","article-title":"Real-time eye, gaze, and face pose tracking for monitoring driver vigilance","volume":"8","author":"Ji","year":"2002","journal-title":"Real-Time Imaging"},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"1124","DOI":"10.1109\/TBME.2005.863952","article-title":"General theory of remote gaze estimation using the pupil center and corneal reflections","volume":"53","author":"Guestrin","year":"2006","journal-title":"IEEE Trans. Biomed. Eng."},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"2246","DOI":"10.1109\/TBME.2007.895750","article-title":"Novel eye gaze tracking techniques under natural head movement","volume":"54","author":"Zhu","year":"2007","journal-title":"IEEE Trans. Biomed. Eng."},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"802","DOI":"10.1109\/TIP.2011.2162740","article-title":"Combining head pose and eye location information for gaze estimation","volume":"21","author":"Valenti","year":"2011","journal-title":"IEEE Trans. Image Process."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Alberto Funes Mora, K., and Odobez, J.M. (2014, January 23\u201328). Geometric generative gaze estimation (g3e) for remote rgb-d cameras. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA.","DOI":"10.1109\/CVPR.2014.229"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Park, S., Spurr, A., and Hilliges, O. (2018, January 8\u201314). Deep pictorial gaze estimation. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01261-8_44"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Cai, X., Zeng, J., Shan, S., and Chen, X. (2023, January 18\u201322). Source-free adaptive gaze estimation by uncertainty reduction. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada.","DOI":"10.1109\/CVPR52729.2023.02110"},{"key":"ref_26","unstructured":"Cheng, Y., Bao, Y., and Lu, F. (March, January 22). Puregaze: Purifying gaze feature for generalizable gaze estimation. Proceedings of the AAAI Conference on Artificial Intelligence, Vancouver, BC, Canada."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Cheng, Y., and Lu, F. (2023, January 2\u20136). Dvgaze: Dual-view gaze estimation. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Paris, France.","DOI":"10.1109\/ICCV51070.2023.01886"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Li, F., Yan, H., and Shi, L. (2024). Multi-scale coupled attention for visual object detection. Sci. Rep., 14.","DOI":"10.1038\/s41598-024-60897-8"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Shang, C., Wang, Z., Wang, H., and Meng, X. (2025, January 11\u201315). SCSA: A Plug-and-Play Semantic Continuous-Sparse Attention for Arbitrary Semantic Style Transfer. Proceedings of the Computer Vision and Pattern Recognition Conference, Nashville, TN, USA.","DOI":"10.1109\/CVPR52734.2025.01218"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Ouyang, D., He, S., Zhang, G., Luo, M., Guo, H., Zhan, J., and Huang, Z. (2023, January 4\u201310). Efficient multi-scale attention module with cross-spatial learning. Proceedings of the ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Rhodes, Greece.","DOI":"10.1109\/ICASSP49357.2023.10096516"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Liu, S., Qi, L., Qin, H., Shi, J., and Jia, J. (2018, January 18\u201322). Path aggregation network for instance segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00913"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Xu, W., and Wan, Y. (2024). ELA: Efficient Local Attention for Deep Convolutional Neural Networks. arXiv.","DOI":"10.1007\/s11554-025-01719-6"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Wang, Q., Wu, B., Zhu, P., Li, P., Zuo, W., and Hu, Q. (2020, January 14\u201319). ECA-Net: Efficient channel attention for deep convolutional neural networks. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.01155"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Selvaraju, R.R., Cogswell, M., Das, A., Vedantam, R., Parikh, D., and Batra, D. (2017, January 22\u201329). Grad-cam: Visual explanations from deep networks via gradient-based localization. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.74"},{"key":"ref_35","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (July, January 26). Deep residual learning for image recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA."},{"key":"ref_36","doi-asserted-by":"crossref","first-page":"415","DOI":"10.1007\/s41095-022-0274-8","article-title":"Pvt v2: Improved baselines with pyramid vision transformer","volume":"8","author":"Wang","year":"2022","journal-title":"Comput. Vis. Media"}],"container-title":["Algorithms"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1999-4893\/18\/9\/538\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,9]],"date-time":"2025-10-09T18:35:22Z","timestamp":1760034922000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1999-4893\/18\/9\/538"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,8,24]]},"references-count":36,"journal-issue":{"issue":"9","published-online":{"date-parts":[[2025,9]]}},"alternative-id":["a18090538"],"URL":"https:\/\/doi.org\/10.3390\/a18090538","relation":{},"ISSN":["1999-4893"],"issn-type":[{"value":"1999-4893","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,8,24]]}}}