{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,6]],"date-time":"2026-04-06T19:15:11Z","timestamp":1775502911241,"version":"3.50.1"},"reference-count":39,"publisher":"MDPI AG","issue":"24","license":[{"start":{"date-parts":[[2023,12,18]],"date-time":"2023-12-18T00:00:00Z","timestamp":1702857600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Open Fund of Key Laboratory of Geospatial Technology for the Middle and Lower Yellow River Regions (Henan University), Ministry of Education","award":["GTYR202209"],"award-info":[{"award-number":["GTYR202209"]}]},{"name":"Open Fund of Key Laboratory of Geospatial Technology for the Middle and Lower Yellow River Regions (Henan University), Ministry of Education","award":["W2023JSFW0173"],"award-info":[{"award-number":["W2023JSFW0173"]}]},{"name":"University-Enterprise Collaboration Project(Hefei University of Technology)","award":["GTYR202209"],"award-info":[{"award-number":["GTYR202209"]}]},{"name":"University-Enterprise Collaboration Project(Hefei University of Technology)","award":["W2023JSFW0173"],"award-info":[{"award-number":["W2023JSFW0173"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Head pose estimation serves various applications, such as gaze estimation, fatigue-driven detection, and virtual reality. Nonetheless, achieving precise and efficient predictions remains challenging owing to the reliance on singular data sources. Therefore, this study introduces a technique involving multimodal feature fusion to elevate head pose estimation accuracy. The proposed method amalgamates data derived from diverse sources, including RGB and depth images, to construct a comprehensive three-dimensional representation of the head, commonly referred to as a point cloud. The noteworthy innovations of this method encompass a residual multilayer perceptron structure within PointNet, designed to tackle gradient-related challenges, along with spatial self-attention mechanisms aimed at noise reduction. The enhanced PointNet and ResNet networks are utilized to extract features from both point clouds and images. These extracted features undergo fusion. Furthermore, the incorporation of a scoring module strengthens robustness, particularly in scenarios involving facial occlusion. This is achieved by preserving features from the highest-scoring point cloud. Additionally, a prediction module is employed, combining classification and regression methodologies to accurately estimate head poses. The proposed method improves the accuracy and robustness of head pose estimation, especially in cases involving facial obstructions. These advancements are substantiated by experiments conducted using the BIWI dataset, demonstrating the superiority of this method over existing techniques.<\/jats:p>","DOI":"10.3390\/s23249894","type":"journal-article","created":{"date-parts":[[2023,12,18]],"date-time":"2023-12-18T12:57:56Z","timestamp":1702904276000},"page":"9894","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":5,"title":["Self-Attention Mechanism-Based Head Pose Estimation Network with Fusion of Point Cloud and Image Features"],"prefix":"10.3390","volume":"23","author":[{"ORCID":"https:\/\/orcid.org\/0009-0002-4825-984X","authenticated-orcid":false,"given":"Kui","family":"Chen","sequence":"first","affiliation":[{"name":"College of Civil Engineering, Hefei University of Technology, Hefei 230009, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zhaofu","family":"Wu","sequence":"additional","affiliation":[{"name":"College of Civil Engineering, Hefei University of Technology, Hefei 230009, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0207-8800","authenticated-orcid":false,"given":"Jianwei","family":"Huang","sequence":"additional","affiliation":[{"name":"College of Civil Engineering, Hefei University of Technology, Hefei 230009, China"},{"name":"Key Laboratory of Geospatial Technology for the Middle and Lower Yellow River Regions, Henan University, Ministry of Education, Kaifeng 475004, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yiming","family":"Su","sequence":"additional","affiliation":[{"name":"College of Civil Engineering, Hefei University of Technology, Hefei 230009, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2023,12,18]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Rossi, S., Leone, E., and Staffa, M. (December, January 29). Using random forests for the estimation of multiple users\u2019 visual focus of attention from head pose. Proceedings of the XV of AI* IA 2016 Advances in Artificial Intelligence: XVth International Conference of the Italian Association for Artificial Intelligence, Genova, Italy.","DOI":"10.1007\/978-3-319-49130-1_8"},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"103402","DOI":"10.1016\/j.jvcir.2021.103402","article-title":"A new head pose tracking method based on stereo visual SLAM","volume":"82","author":"Huang","year":"2022","journal-title":"J. Vis. Commun. Image Represent."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"7107","DOI":"10.1109\/TII.2022.3143605","article-title":"ARHPE: Asymmetric Relation-Aware Representation Learning for Head Pose Estimation in Industrial Human-Computer Interaction","volume":"18","author":"Liu","year":"2022","journal-title":"IEEE Trans. Ind. Inf."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"13533","DOI":"10.1007\/s11042-019-08590-1","article-title":"MIFTel: A multimodal interactive framework based on temporal logic rules","volume":"79","author":"Avola","year":"2020","journal-title":"Multimed. Tools Appl."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"310","DOI":"10.1016\/j.neucom.2020.09.068","article-title":"Anisotropic angle distribution learning for head pose estimation and attention understanding in human-computer interaction","volume":"433","author":"Liu","year":"2021","journal-title":"Neurocomputing"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Wongphanngam, J., and Pumrin, S. (July, January 28). Fatigue warning system for driver nodding off using depth image from Kinect. Proceedings of the 2016 13th International Conference on Electrical Engineering\/Electronics, Computer, Telecommunications and Information Technology, Chiang Mai, Thailand.","DOI":"10.1109\/ECTICon.2016.7561274"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Baltru\u0161aitis, T., Robinson, P., and Morency, L.P. (2016, January 7\u201310). OpenFace: An open source facial behavior analysis toolkit. Proceedings of the 2016 IEEE Winter Conference on Applications of Computer Vision, Lake Placid, NY, USA.","DOI":"10.1109\/WACV.2016.7477553"},{"key":"ref_8","first-page":"310","article-title":"Head attitude estimation method of eye tracker based on binocular camera","volume":"58","author":"Han","year":"2019","journal-title":"Adv. Laser Optoelectron."},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Zhao, G., Chen, L., Song, J., and Chen, G. (2007, January 25\u201329). Large head movement tracking using sift-based registration. Proceedings of the 15th ACM International Conference on Multimedia, Augsburg, Germany.","DOI":"10.1145\/1291233.1291416"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Liu, L., Ke, Z., Huo, J., and Chen, J. (2021). Head pose estimation through keypoints matching between reconstructed 3D face model and 2D image. Sensors, 21.","DOI":"10.3390\/s21051841"},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"6289","DOI":"10.1109\/TIP.2023.3331309","article-title":"Orientation Cues-Aware Facial Relationship Representation for Head Pose Estimation via Transformer","volume":"32","author":"Liu","year":"2023","journal-title":"IEEE Trans. Image Process."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"1974","DOI":"10.1109\/TPAMI.2020.3029585","article-title":"Head pose estimation based on multivariate label distribution","volume":"44","author":"Geng","year":"2020","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Zhang, C., Liu, H., Deng, Y., Xie, B., and Li, Y. (2023, January 18\u201322). TokenHPE: Learning Orientation Tokens for Efficient Head Pose Estimation via Transformers. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, BC, Canada.","DOI":"10.1109\/CVPR52729.2023.00859"},{"key":"ref_14","first-page":"2449","article-title":"MFDNet: Collaborative Poses Perception and Matrix Fisher Distribution for Head Pose Estimation. IEEE Trans","volume":"24","author":"Liu","year":"2022","journal-title":"Multimedia"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Ruiz, N., Chong, E., and Rehg, J.M. (2018, January 18\u201322). Fine-grained head pose estimation without keypoints. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPRW.2018.00281"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Yang, T.-Y., Chen, Y.-T., Lin, Y.-Y., and Chuang, Y.-Y. (2019, January 15\u201320). FSA-Net: Learning fine-grained structure aggregation for head pose estimation from a single image. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00118"},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"153318","DOI":"10.1007\/s11704-020-8272-4","article-title":"Practical age estimation using deep label distribution learning","volume":"15","author":"Zhang","year":"2021","journal-title":"Front. Comput. Sci."},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"210","DOI":"10.1016\/j.neucom.2020.12.090","article-title":"NGDNet: Nonuniform Gaussian-label distribution learning for infrared head pose estimation and on-task behavior understanding in the classroom","volume":"436","author":"Liu","year":"2021","journal-title":"Neurocomputing"},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"339","DOI":"10.1016\/j.neucom.2018.12.074","article-title":"Head pose estimation with soft labels using regularized convolutional neural network","volume":"337","author":"Xu","year":"2017","journal-title":"Neurocomputing"},{"key":"ref_20","first-page":"2309","article-title":"Real-time head attitude estimation based on Kalman filter and random regression forest","volume":"29","author":"Chenglong","year":"2017","journal-title":"J. Comput. Aid. Des. Graph."},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Wang, Y., Yuan, G., and Fu, X. (2022). Driver\u2019s head pose and gaze zone estimation based on multi-zone templates registration and multi-frame point cloud fusion. Sensors, 22.","DOI":"10.3390\/s22093154"},{"key":"ref_22","first-page":"996","article-title":"3D point cloud head attitude estimation based on Deep learning","volume":"40","author":"Shihua","year":"2020","journal-title":"J. Comput. Appl."},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"108210","DOI":"10.1016\/j.patcog.2021.108210","article-title":"Head pose estimation using deep neural networks and 3D point clouds","volume":"121","author":"Xu","year":"2022","journal-title":"Pattern Recog."},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"259","DOI":"10.1016\/j.neucom.2020.05.010","article-title":"Learning from discrete Gaussian label distribution and spatial channel-aware residual attention for head pose estimation","volume":"407","author":"Zhang","year":"2020","journal-title":"Neurocomputing"},{"key":"ref_25","first-page":"115","article-title":"Les valeurs extr\u00eames des distributions statistiques","volume":"5","author":"Gumbel","year":"1935","journal-title":"Ann. De L\u2019Institut Henri Poincar\u00e9"},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"359","DOI":"10.1016\/0893-6080(89)90020-8","article-title":"Multilayer feedforward networks are universal approximators","volume":"2","author":"Hornik","year":"1989","journal-title":"Neural Netw."},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"99","DOI":"10.1145\/3503250","article-title":"NeRF: Representing scenes as neural radiance fields for view synthesis","volume":"65","author":"Mildenhall","year":"2021","journal-title":"Commun. ACM"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Charles, R.Q., Su, H., Mo, K., and Guibas, L.J. (2017, January 21\u201326). PointNet: Deep learning on point sets for 3D classification and segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.16"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep residual learning for image recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"748","DOI":"10.1016\/j.asoc.2018.09.010","article-title":"A convolutional neural network with feature fusion for real-time hand posture recognition","volume":"73","author":"Chevtchenko","year":"2018","journal-title":"Appl. Soft Comput."},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"48","DOI":"10.1109\/TIV.2022.3164899","article-title":"MTANet: Multitask-aware network with hierarchical multimodal fusion for RGB-T urban scene understanding","volume":"8","author":"Zhou","year":"2022","journal-title":"IEEE Trans. Intell. Veh."},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Xu, D., Anguelov, D., and Jain, A. (2018, January 18\u201323). PointFusion: Deep sensor fusion for 3D bounding box estimation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00033"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Wang, C., Xu, D., Zhu, Y., Mart\u00edn-Mart\u00edn, R., Lu, C., Fei-Fei, L., and Savarese, S. (2019, January 15\u201320). DenseFusion: 6D object pose estimation by iterative dense fusion. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00346"},{"key":"ref_34","doi-asserted-by":"crossref","first-page":"510","DOI":"10.1016\/j.neucom.2020.06.066","article-title":"Infrared head pose estimation with multi-scales feature fusion on the IRHP database for human attention recognition","volume":"411","author":"Liu","year":"2020","journal-title":"Neurocomputing"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Fanelli, G., Gall, J., and Gool, L.V. (2011, January 20\u201325). Real time head pose estimation with random regression forests. Proceedings of the Conference on Computer Vision and Pattern Recognition 2011, Colorado Springs, CO, USA.","DOI":"10.1109\/CVPR.2011.5995458"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Xu, X., and Kakadiaris, I.A. (June, January 30). Joint head pose estimation and face alignment framework using global and local CNN features. Proceedings of the 2017 12th IEEE International Conference on Automatic Face & Gesture Recognition (FG 2017), Washington, DC, USA.","DOI":"10.1109\/FG.2017.81"},{"key":"ref_37","doi-asserted-by":"crossref","first-page":"196","DOI":"10.1016\/j.patcog.2019.05.026","article-title":"A deep coarse-to-fine network for head pose estimation from synthetic data","volume":"94","author":"Wang","year":"2019","journal-title":"Pattern Recog."},{"key":"ref_38","doi-asserted-by":"crossref","first-page":"596","DOI":"10.1109\/TPAMI.2018.2885472","article-title":"Face-from-depth for head pose estimation on depth images","volume":"42","author":"Borghi","year":"2018","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Meyer, G.P., Gupta, S., Frosio, I., Reddy, D., and Kautz, J. (2015, January 7\u201313). Robust model-based 3D head pose estimation. Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.416"}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/24\/9894\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T21:40:37Z","timestamp":1760132437000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/24\/9894"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,12,18]]},"references-count":39,"journal-issue":{"issue":"24","published-online":{"date-parts":[[2023,12]]}},"alternative-id":["s23249894"],"URL":"https:\/\/doi.org\/10.3390\/s23249894","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,12,18]]}}}