{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,5]],"date-time":"2026-06-05T16:11:30Z","timestamp":1780675890690,"version":"3.54.1"},"reference-count":44,"publisher":"MDPI AG","issue":"15","license":[{"start":{"date-parts":[[2022,8,5]],"date-time":"2022-08-05T00:00:00Z","timestamp":1659657600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Estimating the driver\u2019s gaze in a natural real-world setting can be problematic for different challenging scenario conditions. For example, faces will undergo facial occlusions, illumination, or various face positions while driving. In this effort, we aim to reduce misclassifications in driving situations when the driver has different face distances regarding the camera. Three-dimensional Convolutional Neural Networks (CNN) models can make a spatio-temporal driver\u2019s representation that extracts features encoded in multiple adjacent frames that can describe motions. This characteristic may help ease the deficiencies of a per-frame recognition system due to the lack of context information. For example, the front, navigator, right window, left window, back mirror, and speed meter are part of the known common areas to be checked by drivers. Based on this, we implement and evaluate a model that is able to detect the head direction toward these regions having various distances from the camera. In our evaluation, the 2D CNN model had a mean average recall of 74.96% across the three models, whereas the 3D CNN model had a mean average recall of 87.02%. This result show that our proposed 3D CNN-based approach outperforms a 2D CNN per-frame recognition approach in driving situations when the driver\u2019s face has different distances from the camera.<\/jats:p>","DOI":"10.3390\/s22155857","type":"journal-article","created":{"date-parts":[[2022,8,9]],"date-time":"2022-08-09T04:16:55Z","timestamp":1660018615000},"page":"5857","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":7,"title":["Single Camera Face Position-Invariant Driver\u2019s Gaze Zone Classifier Based on Frame-Sequence Recognition Using 3D Convolutional Neural Networks"],"prefix":"10.3390","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-7478-423X","authenticated-orcid":false,"given":"Catherine","family":"Lollett","sequence":"first","affiliation":[{"name":"Graduate School of Creative Science and Engineering, Waseda University, Tokyo 169-8555, Japan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Mitsuhiro","family":"Kamezaki","sequence":"additional","affiliation":[{"name":"Research Institute for Science and Engineering (RISE), Waseda University, Tokyo 162-0044, Japan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9331-2446","authenticated-orcid":false,"given":"Shigeki","family":"Sugano","sequence":"additional","affiliation":[{"name":"Graduate School of Creative Science and Engineering, Waseda University, Tokyo 169-8555, Japan"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2022,8,5]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Ashraf, I., Hur, S., Shafiq, M., and Park, Y. (2019). Catastrophic factors involved in road accidents: Underlying causes and descriptive analysis. PLoS ONE, 14.","DOI":"10.1371\/journal.pone.0223473"},{"key":"ref_2","first-page":"16663","article-title":"Prediction of road accidents severity using various algorithms","volume":"119","author":"Ramachandiran","year":"2018","journal-title":"Int. J. Pure Appl. Math."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"144","DOI":"10.5958\/0973-9130.2019.00102.6","article-title":"Road traffic accidents in India: Need for urgent attention and solutions to ensure road safety","volume":"13","author":"Kini","year":"2019","journal-title":"Indian J. Forensic Med. Toxicol."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Hayashi, H., Kamezaki, M., Manawadu, U.E., Kawano, T., Ema, T., Tomita, T., Catherine, L., and Sugano, S. (2019, January 9\u201312). A Driver Situational Awareness Estimation System Based on Standard Glance Model for Unscheduled Takeover Situations. Proceedings of the IEEE Intelligent Vehicles Symposium, Paris, France.","DOI":"10.1109\/IVS.2019.8814067"},{"key":"ref_5","first-page":"167","article-title":"Development of a Situational Awareness Estimation Model Considering Traffic Environment for Unscheduled Takeover Situations","volume":"19","author":"Hayashi","year":"2020","journal-title":"Int. J. Intell. Transp. Res."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Manawadu, U.E., Kawano, T., Murata, S., Kamezaki, M., Muramatsu, J., and Sugano, S. (2018, January 26\u201330). Multiclass Classification of Driver Perceived Workload Using Long Short-Term Memory based Recurrent Neural Network. Proceedings of the IEEE Intelligent Vehicles Symposium, Changshu, China.","DOI":"10.1109\/IVS.2018.8500410"},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"240","DOI":"10.1109\/OJITS.2021.3102125","article-title":"Toward Health-Related Accident Prevention: Symptom Detection and Intervention based on Driver Monitoring and Verbal Interaction","volume":"2","author":"Hayashi","year":"2021","journal-title":"IEEE Open Intell. Transp. Syst."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Lollett, C., Hayashi, H., Kamezaki, M., and Sugano, S. (2020, January 11\u201314). A Robust Driver\u2019s Gaze Zone Classification using a Single Camera for Self-occlusions and Non-aligned Head and Eyes Direction Driving Situations. Proceedings of the 2020 IEEE International Conference on Systems, Man, and Cybernetics (SMC), Toronto, ON, Canada.","DOI":"10.1109\/SMC42975.2020.9283470"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Lollett, C., Kamezaki, M., and Sugano, S. (2021, January 11\u201317). Towards a Driver\u2019s Gaze Zone Classifier using a Single Camera Robust to Temporal and Permanent Face Occlusions. Proceedings of the 2021 IEEE Intelligent Vehicles Symposium (IV), Nagoya, Japan.","DOI":"10.1109\/IV48863.2021.9575367"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Lollett, C., Kamezaki, M., and Sugano, S. (2022, January 4\u20139). Driver\u2019s Drowsiness Classifier using a Single Camera Robust to Mask-wearing Situations using an Eyelid, Face Contour and Chest Movement Feature Vector GRU-based Model. Proceedings of the 2022 IEEE Intelligent Vehicles Symposium (IV), Aachen, Germany.","DOI":"10.1109\/IV51971.2022.9827229"},{"key":"ref_11","unstructured":"Chen, M., and Hauptmann, A. (2022, June 29). Mosift: Recognizing Human Actions in Surveillance Videos. Available online: https:\/\/kilthub.cmu.edu\/articles\/journal_contribution\/MoSIFT_Recognizing_Human_Actions_in_Surveillance_Videos\/6607523."},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Li, B. (2017, January 24\u201328). 3D Fully convolutional network for vehicle detection in point cloud. Proceedings of the 2017 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS), Vancouver, BC, Canada.","DOI":"10.1109\/IROS.2017.8205955"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Tawari, A., Chen, K.H., and Trivedi, M.M. (2014, January 8\u201311). Where is the driver looking: Analysis of head, eye and iris for robust gaze zone estimation. Proceedings of the 17th International IEEE Conference on Intelligent Transportation Systems (ITSC), Qingdao, China.","DOI":"10.1109\/ITSC.2014.6957817"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Fridman, L., Toyoda, H., Seaman, S., Seppelt, B., Angell, L., Lee, J., Mehler, B., and Reimer, B. (2017, January 6\u201311). What can be predicted from six seconds of driver glances?. Proceedings of the Human Factors in Computing Systems, Denver, CO, USA.","DOI":"10.1145\/3025453.3025929"},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"308","DOI":"10.1049\/iet-cvi.2015.0296","article-title":"\u2018Owl\u2019and \u2018Lizard\u2019: Patterns of head pose and eye pose in driver gaze classification","volume":"10","author":"Fridman","year":"2016","journal-title":"IET Comput. Vis."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Chuang, M.C., Bala, R., Bernal, E.A., Paul, P., and Burry, A. (2014, January 23\u201328). Estimating gaze direction of vehicle drivers using a smartphone camera. Proceedings of the 2014 IEEE Conference on Computer Vision and Pattern Recognition Workshops, Columbus, OH, USA.","DOI":"10.1109\/CVPRW.2014.30"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Naqvi, A., Arsalan, M., Batchuluun, G., Yoon, S., and Park, R. (2018). Deep Learning-Based Gaze Detection System for Automobile Drivers Using a NIR Camera Sensor. Sensors, 18.","DOI":"10.3390\/s18020456"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Liu, Z., Luo, P., Wang, X., and Tang, X. (2015, January 7\u201313). Deep learning face attributes in the wild. Proceedings of the 2015 IEEE International Conference on Computer Vision (ICCV), Santiago, Chile.","DOI":"10.1109\/ICCV.2015.425"},{"key":"ref_19","unstructured":"Martin, S.C. (2016). Vision based, Multi-cue Driver Models for Intelligent Vehicles. [Ph.D. Dissertation, University of California]."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Burgos-Artizzu, X., Perona, P., and Doll\u2019ar, P. (2013, January 1\u20138). Robust face landmark estimation under occlusion. Proceedings of the 2013 IEEE International Conference on Computer Vision, Sydney, Australia.","DOI":"10.1109\/ICCV.2013.191"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Tayibnapis, I.R., Choi, M.K., and Kwon, S. (2018, January 12\u201314). Driver\u2019s gaze zone estimation by transfer learning. Proceedings of the 2018 IEEE International Conference on Consumer Electronics (ICCE), Las Vegas, NV, USA.","DOI":"10.1109\/ICCE.2018.8326308"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Shan, X., Wang, Z., Liu, X., Lin, M., Zhao, L., Wang, J., and Wang, G. (2020, January 28\u201329). Driver Gaze Region Estimation Based on Computer Vision. Proceedings of the Measuring Technology and Mechatronics Automation (ICMTMA), Phuket, Thailand.","DOI":"10.1109\/ICMTMA50254.2020.00085"},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"254","DOI":"10.1109\/TIV.2018.2843120","article-title":"Driver gaze zone estimation using convolutional neural networks: A general framework and ablative analysis","volume":"3","author":"Vora","year":"2018","journal-title":"IEEE Trans. Intell. Transp."},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Schwehr, J., and Willert, V. (2017, January 16\u201319). Driver\u2019s gaze prediction in dynamic automotive scenes. Proceedings of the 2017 IEEE 20th International Conference on Intelligent Transportation Systems (ITSC), Yokohama, Japan.","DOI":"10.1109\/ITSC.2017.8317586"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Guasconi, S., Porta, M., Resta, C., and Rottenbacher, C. (2017, January 17\u201319). A low-cost implementation of an eye tracking system for driver\u2019s gaze analysis. Proceedings of the 10th International Conference on Human System Interactions (HSI), Ulsan, Korea.","DOI":"10.1109\/HSI.2017.8005043"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Wang, Y., Yuan, G., Mi, Z., Peng, J., Ding, X., Liang, Z., and Fu, X. (2019). Continuous driver\u2019s gaze zone estimation using rgb-d camera. Sensors, 19.","DOI":"10.3390\/s19061287"},{"key":"ref_27","unstructured":"Wang, Y., Zhao, T., Ding, X., Bian, J., and Fu, X. (2017, January 13\u201316). Head pose-free eye gaze prediction for driver attention study. Proceedings of the 2017 IEEE International Conference on Big Data and Smart Computing (BigComp), Jeju, Korea."},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Jha, S., and Busso, C. (2018, January 4\u20137). Probabilistic Estimation of the Gaze Region of the Driver using Dense Classification. Proceedings of the 21st International Conference on Intelligent Transportation Systems (ITSC), Maui, HI, USA.","DOI":"10.1109\/ITSC.2018.8569709"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Yuen, K., Martin, S., and Trivedi, M.M. (2016, January 1\u20134). Looking at faces in a vehicle: A deep CNN based approach and evaluation. Proceedings of the IEEE 19th International Conference on Intelligent Transportation Systems (ITSC), Rio de Janeiro, Brazil.","DOI":"10.1109\/ITSC.2016.7795622"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Hu, T., Jha, S., and Busso, C. (November, January 19). Robust Driver Head Pose Estimation in Naturalistic Conditions from Point-Cloud Data. Proceedings of the 2020 IEEE Intelligent Vehicles Symposium (IV), Las Vegas, NV, USA.","DOI":"10.1109\/IV47402.2020.9304592"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Rangesh, A., Zhang, B., and Trivedi, M. (November, January 19). Driver Gaze Estimation in the Real World: Overcoming the Eyeglass Challenge. Proceedings of the 2020 IEEE Intelligent Vehicles Symposium (IV), Las Vegas, NV, USA.","DOI":"10.1109\/IV47402.2020.9304573"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Dari, S., Kadrileev, N., and H\u00fcllermeier, E. (2020, January 19\u201324). A Neural Network-Based Driver Gaze Classification System with Vehicle Signals. Proceedings of the 2020 International Joint Conference on Neural Networks (IJCNN), Glasgow, UK.","DOI":"10.1109\/IJCNN48605.2020.9207709"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Ruiz, N., Chong, E., and Rehg, J.M. (2018, January 18\u201322). Fine-grained head pose estimation without keypoints. Proceedings of the 2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPRW.2018.00281"},{"key":"ref_34","doi-asserted-by":"crossref","first-page":"4206","DOI":"10.1109\/TITS.2018.2883823","article-title":"Driver drowsiness detection using condition-adaptive representation learning framework","volume":"20","author":"Yu","year":"2018","journal-title":"IEEE Trans. Intell. Transp. Syst."},{"key":"ref_35","unstructured":"Huynh, X.P., Park, S.M., and Kim, Y.G. (2016). Detection of driver drowsiness using 3D deep neural network and semi-supervised gradient boosting machine. Asian Conference on Computer Vision, Springer."},{"key":"ref_36","first-page":"127","article-title":"Facial feature detection using Haar classifiers","volume":"21","author":"Wilson","year":"2006","journal-title":"J. Comput. Sci. Coll."},{"key":"ref_37","doi-asserted-by":"crossref","first-page":"1598","DOI":"10.1016\/j.patrec.2011.01.004","article-title":"Face recognition using histograms of oriented gradients","volume":"32","author":"Bueno","year":"2011","journal-title":"Pattern Recognit. Lett."},{"key":"ref_38","unstructured":"Minaee, S., Luo, P., Lin, Z., and Bowyer, K. (2021). Going Deeper Into Face Detection: A Survey. arXiv."},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Zhang, S., Zhu, X., Lei, Z., Shi, H., Wang, X., and Li, S.Z. (2017, January 22\u201329). S3fd: Single shot scale-invariant face detector. Proceedings of the 2017 IEEE International Conference on Computer Vision (ICCV), Venice, Italy.","DOI":"10.1109\/ICCV.2017.30"},{"key":"ref_40","unstructured":"Witten, I.H., Frank, E., Hall, M.A., and Pal, C.J. (2016). Data Mining: Practical Machine Learning Tools and Techniques, Morgan Kaufmann."},{"key":"ref_41","doi-asserted-by":"crossref","first-page":"504","DOI":"10.1126\/science.1127647","article-title":"Reducing the dimensionality of data with neural networks","volume":"313","author":"Hinton","year":"2006","journal-title":"Science"},{"key":"ref_42","doi-asserted-by":"crossref","first-page":"436","DOI":"10.1038\/nature14539","article-title":"Deep learning","volume":"521","author":"LeCun","year":"2015","journal-title":"Nature"},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Tran, D., Bourdev, L., Fergus, R., Torresani, L., and Paluri, M. (2015, January 7\u201313). Learningspatiotemporal features with 3d convolutional networks. Proceedings of the 2015 IEEE International Conference on Computer Vision (ICCV), Santiago, Chile.","DOI":"10.1109\/ICCV.2015.510"},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"Khan, M.Q., and Lee, S. (2019). A comprehensive survey of driving monitoring and assistance systems. Sensors, 19.","DOI":"10.3390\/s19112574"}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/22\/15\/5857\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T00:04:36Z","timestamp":1760141076000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/22\/15\/5857"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,8,5]]},"references-count":44,"journal-issue":{"issue":"15","published-online":{"date-parts":[[2022,8]]}},"alternative-id":["s22155857"],"URL":"https:\/\/doi.org\/10.3390\/s22155857","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,8,5]]}}}