{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,31]],"date-time":"2026-01-31T08:44:30Z","timestamp":1769849070164,"version":"3.49.0"},"reference-count":36,"publisher":"MDPI AG","issue":"10","license":[{"start":{"date-parts":[[2021,10,9]],"date-time":"2021-10-09T00:00:00Z","timestamp":1633737600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"the National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["61772152"],"award-info":[{"award-number":["61772152"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Entropy"],"abstract":"<jats:p>Vigilance estimation of drivers is a hot research field of current traffic safety. Wearable devices can monitor information regarding the driver\u2019s state in real time, which is then analyzed by a data analysis model to provide an estimation of vigilance. The accuracy of the data analysis model directly affects the effect of vigilance estimation. In this paper, we propose a deep coupling recurrent auto-encoder (DCRA) that combines electroencephalography (EEG) and electrooculography (EOG). This model uses a coupling layer to connect two single-modal auto-encoders to construct a joint objective loss function optimization model, which consists of single-modal loss and multi-modal loss. The single-modal loss is measured by Euclidean distance, and the multi-modal loss is measured by a Mahalanobis distance of metric learning, which can effectively reflect the distance between different modal data so that the distance between different modes can be described more accurately in the new feature space based on the metric matrix. In order to ensure gradient stability in the long sequence learning process, a multi-layer gated recurrent unit (GRU) auto-encoder model was adopted. The DCRA integrates data feature extraction and feature fusion. Relevant comparative experiments show that the DCRA is better than the single-modal method and the latest multi-modal fusion. The DCRA has a lower root mean square error (RMSE) and a higher Pearson correlation coefficient (PCC).<\/jats:p>","DOI":"10.3390\/e23101316","type":"journal-article","created":{"date-parts":[[2021,10,10]],"date-time":"2021-10-10T21:19:32Z","timestamp":1633900772000},"page":"1316","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":16,"title":["Deep Coupling Recurrent Auto-Encoder with Multi-Modal EEG and EOG for Vigilance Estimation"],"prefix":"10.3390","volume":"23","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-4550-3173","authenticated-orcid":false,"given":"Kuiyong","family":"Song","sequence":"first","affiliation":[{"name":"College of Computer Science and Technology, Harbin Engineering University, Harbin 150001, China"},{"name":"Department of Information Engineering, Hulunbuir Vocational Technical College, Hulunbuir 021000, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Lianke","family":"Zhou","sequence":"additional","affiliation":[{"name":"College of Computer Science and Technology, Harbin Engineering University, Harbin 150001, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hongbin","family":"Wang","sequence":"additional","affiliation":[{"name":"College of Computer Science and Technology, Harbin Engineering University, Harbin 150001, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2021,10,9]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"1052","DOI":"10.1109\/TVT.2004.830974","article-title":"Real-Time Nonintrusive Monitoring and Prediction of Driver Fatigue","volume":"53","author":"Ji","year":"2004","journal-title":"IEEE Trans. Veh. Technol."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"026017","DOI":"10.1088\/1741-2552\/aa5a98","article-title":"A Multimodal Approach to Estimating Vigilance Using EEG and Forehead EOG","volume":"14","author":"Zheng","year":"2017","journal-title":"J. Neural Eng."},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Du, L.-H., Liu, W., Zheng, W.-L., and Lu, B.-L. (2017, January 25\u201328). Detecting driving fatigue with multimodal deep learning. Proceedings of the 2017 8th International IEEE\/EMBS Conference on Neural Engineering (NER), Shanghai, China.","DOI":"10.1109\/NER.2017.8008295"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Li, H., Zheng, W.-L., and Lu, B.-L. (2018, January 8\u201313). Multimodal vigilance estimation with adversarial domain adaptation networks. Proceedings of the 2018 International Joint Conference on Neural Networks (IJCNN), Rio de Janeiro, Brazil.","DOI":"10.1109\/IJCNN.2018.8489212"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"193","DOI":"10.1016\/j.neunet.2020.11.002","article-title":"An EEG channel selection method for motor imagery based brain\u2013computer interface and neurofeedback using Granger causality","volume":"133","author":"Varsehi","year":"2021","journal-title":"Neural Netw."},{"key":"ref_6","first-page":"423","article-title":"Multimodal machine learning: A survey and taxonomy","volume":"41","author":"Ahuja","year":"2018","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"b345","DOI":"10.1007\/s00530-010-0182-0","article-title":"Multimodal Fusion for Multimedia Analysis: A survey","volume":"16","author":"Atrey","year":"2010","journal-title":"Multimed. Syst."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"478","DOI":"10.1109\/JSTSP.2020.2987728","article-title":"Multimodal intelligence: Representation learning, information fusion, and applications","volume":"14","author":"Zhang","year":"2020","journal-title":"IEEE J. Sel. Top. Signal Process."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"1693","DOI":"10.1109\/TIM.2017.2669947","article-title":"Multisensor feature fusion for bearing fault diagnosis using sparse autoencoder and deep belief network","volume":"66","author":"Chen","year":"2017","journal-title":"IEEE Trans. Instrum. Meas."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"1137","DOI":"10.1109\/TII.2018.2793246","article-title":"Deep coupling autoencoder for fault diagnosis with multimodal sensory data","volume":"14","author":"Ma","year":"2018","journal-title":"IEEE Trans. Ind. Inform."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Feng, F., Wang, X., and Li, R. (2014, January 21\u201325). Cross-modal retrieval with correspondence autoencoder. Proceedings of the 22nd ACM International Conference on Multimedia, Nice, France.","DOI":"10.1145\/2647868.2654902"},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"165","DOI":"10.1016\/j.neucom.2018.11.042","article-title":"Multi-modal semantic autoencoder for cross-modal retrieval","volume":"331","author":"Wu","year":"2019","journal-title":"Neurocomputing"},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"63373","DOI":"10.1109\/ACCESS.2019.2916887","article-title":"Deep multimodal representation learning: A survey","volume":"7","author":"Guo","year":"2019","journal-title":"IEEE Access"},{"key":"ref_14","unstructured":"Zhang, G., and Etemad, A. (2019). Capsule attention for multimodal eeg and eog spatiotemporal representation learning with application to driver vigilance estimation. arXiv."},{"key":"ref_15","unstructured":"Huo, X.-Q., Zheng, W.-L., and Lu, B.-L. (2016, January 24\u201329). Driving fatigue detection with fusion of EEG and forehead EOG. Proceedings of the 2016 International Joint Conference on Neural Networks (IJCNN), Vancouver, BC, Canada."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Zhang, N., Zheng, W.-L., Liu, W., and Lu, B.-L. (2016, January 16\u201321). Continuous vigilance estimation using LSTM neural networks. Proceedings of the International Conference on Neural Information Processing, Kyoto, Japan.","DOI":"10.1007\/978-3-319-46672-9_59"},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"209","DOI":"10.1109\/TCDS.2018.2889223","article-title":"A regression method with subnetwork neurons for vigilance estimation using EOG and EEG","volume":"13","author":"Wu","year":"2018","journal-title":"IEEE Trans. Cogn. Dev. Syst."},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"10787","DOI":"10.1007\/s00521-020-05046-8","article-title":"Neural image reconstruction using a heuristic validation mechanism","volume":"33","author":"Srivastava","year":"2021","journal-title":"Neural Comput. Appl."},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"1110","DOI":"10.1109\/TCYB.2018.2797176","article-title":"Emotionmeter: A multimodal framework for recognizing human emotions","volume":"49","author":"Zheng","year":"2018","journal-title":"IEEE Trans. Cybern."},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"170","DOI":"10.1109\/TITS.2018.2889962","article-title":"Vigilance estimation using a wearable EOG device in real driving environment","volume":"21","author":"Zheng","year":"2019","journal-title":"IEEE Trans. Intell. Transp. Syst."},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Lan, Y.-T., Liu, W., and Lu, B.-L. (2020, January 19\u201324). Multimodal emotion recognition using deep generalized canonical correlation analysis with an attention mechanism. Proceedings of the 2020 International Joint Conference on Neural Networks (IJCNN), Glasgow, UK.","DOI":"10.1109\/IJCNN48605.2020.9207625"},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1016\/S0169-7439(99)00047-7","article-title":"The mahalanobis distance","volume":"50","author":"Massart","year":"2000","journal-title":"Chemom. Intell. Lab. Syst."},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"107424","DOI":"10.1016\/j.patcog.2020.107424","article-title":"Deep features for person re-identification on metric learning","volume":"110","author":"Wu","year":"2021","journal-title":"Pattern Recognit."},{"key":"ref_24","first-page":"3","article-title":"Autoencoders, minimum description length, and Helmholtz free energy","volume":"6","author":"Hinton","year":"1994","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"1363","DOI":"10.1109\/TCYB.2015.2426723","article-title":"Learning a mahalanobis distance-based dynamic time warping measure for multivariate time series classification","volume":"46","author":"Mei","year":"2015","journal-title":"IEEE Trans. Cybern."},{"key":"ref_26","first-page":"521","article-title":"Distance metric learning with application to clustering with side-information","volume":"15","author":"Xing","year":"2002","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_27","unstructured":"Yang, L., and Jin, R. (2006). Distance Metric Learning: A Comprehensive Survey, Michigan State Universiy."},{"key":"ref_28","first-page":"207","article-title":"Distance metric learning for large margin nearest neighbor classification","volume":"10","author":"Weinberger","year":"2009","journal-title":"J. Mach. Learn. Res."},{"key":"ref_29","unstructured":"Wang, S., and Jin, R. (2021, September 10). An information geometry approach for distance metric learning. Artificial Intelligence and Statistics, Available online: http:\/\/proceedings.mlr.press\/v5\/wang09c.html."},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"1735","DOI":"10.1162\/neco.1997.9.8.1735","article-title":"Long short-term memory","volume":"9","author":"Hochreiter","year":"1997","journal-title":"Neural Comput."},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Sak, H., Senior, A., and Beaufays, F. (2014). Long short-term memory based recurrent neural network architectures for large vocabulary speech recognition. arXiv.","DOI":"10.21437\/Interspeech.2014-80"},{"key":"ref_32","unstructured":"Zaremba, W., Sutskever, I., and Vinyals, O. (2014). Recurrent neural network regularization. arXiv."},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Cho, K., Van Merri\u00ebnboer, B., Gulcehre, C., Bahdanau, D., Bougares, F., Schwenk, H., and Bengio, Y. (2014). Learning phrase representations using RNN encoder-decoder for statistical machine translation. arXiv.","DOI":"10.3115\/v1\/D14-1179"},{"key":"ref_34","unstructured":"Dinges, D.F., and Grace, R. (1998). PERCLOS: A Valid Psychophysiological Measure of Alertness as Assessed by Psychomotor Vigilance, Publication Number FHWA-MCRT-98-006."},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"92","DOI":"10.1109\/T-AFFC.2011.9","article-title":"Continuous prediction of spontaneous affect from multiple cues and modalities in valence-arousal space","volume":"2","author":"Nicolaou","year":"2011","journal-title":"IEEE Trans. Affect. Comput."},{"key":"ref_36","unstructured":"Ioffe, S., and Szegedy, C. (2015, January 7\u20139). Batch normalization: Accelerating deep network training by reducing internal covariate shift. Proceedings of the International conference on machine learning, Lille, France."}],"container-title":["Entropy"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1099-4300\/23\/10\/1316\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T07:10:52Z","timestamp":1760166652000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1099-4300\/23\/10\/1316"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,10,9]]},"references-count":36,"journal-issue":{"issue":"10","published-online":{"date-parts":[[2021,10]]}},"alternative-id":["e23101316"],"URL":"https:\/\/doi.org\/10.3390\/e23101316","relation":{},"ISSN":["1099-4300"],"issn-type":[{"value":"1099-4300","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,10,9]]}}}