{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,21]],"date-time":"2026-07-21T14:58:21Z","timestamp":1784645901405,"version":"3.55.0"},"reference-count":115,"publisher":"Institution of Engineering and Technology (IET)","issue":"1","license":[{"start":{"date-parts":[[2025,7,14]],"date-time":"2025-07-14T00:00:00Z","timestamp":1752451200000},"content-version":"vor","delay-in-days":194,"URL":"http:\/\/creativecommons.org\/licenses\/by-nc\/4.0\/"},{"start":{"date-parts":[[2025,1,1]],"date-time":"2025-01-01T00:00:00Z","timestamp":1735689600000},"content-version":"tdm","delay-in-days":0,"URL":"http:\/\/doi.wiley.com\/10.1002\/tdm_license_1.1"}],"content-domain":{"domain":["ietresearch.onlinelibrary.wiley.com"],"crossmark-restriction":true},"short-container-title":["IET Image Processing"],"published-print":{"date-parts":[[2025,1]]},"abstract":"<jats:title>ABSTRACT<\/jats:title>\n                  <jats:p>Lip reading models improve information processing and decision\u2010making by quickly and accurately comprehending enormous amounts of text. This study dives into the important role that lip reading plays in making communication more inclusive, especially for individuals with hearing impairments. From 2020 to 2024, the researchers carefully examine the progress made in lip\u2010reading algorithms. They take a close look at the methods, innovations and principles used to decode spoken content from videos, specifically using visual speech recognition techniques. The study also emphasises the use of datasets like LRW, LRS2 and LRS3, which are crucial for this exploration. This paper offers valuable insights into recent advancements and highlights the importance of diverse datasets in improving lip\u2010reading models. Its findings aim to guide future research efforts in making communication more accessible for people with hearing impairments.<\/jats:p>","DOI":"10.1049\/ipr2.70095","type":"journal-article","created":{"date-parts":[[2025,7,14]],"date-time":"2025-07-14T07:40:27Z","timestamp":1752478827000},"update-policy":"https:\/\/doi.org\/10.1002\/crossmark_policy","source":"Crossref","is-referenced-by-count":5,"title":["A Comprehensive Survey of Advancement in Lip Reading Models: Techniques and Future Directions"],"prefix":"10.1049","volume":"19","author":[{"ORCID":"https:\/\/orcid.org\/0009-0005-0841-3234","authenticated-orcid":false,"given":"Sampada","family":"Deshpande","sequence":"first","affiliation":[{"name":"Department of Computer Engineering and Information Technology Veermata Jijabai Technological Institute  Mumbai India"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Kalyani","family":"Shirsath","sequence":"additional","affiliation":[{"name":"Department of Computer Engineering and Information Technology Veermata Jijabai Technological Institute  Mumbai India"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Amey","family":"Pashte","sequence":"additional","affiliation":[{"name":"Department of Computer Engineering and Information Technology Veermata Jijabai Technological Institute  Mumbai India"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Pratham","family":"Loya","sequence":"additional","affiliation":[{"name":"Department of Computer Engineering and Information Technology Veermata Jijabai Technological Institute  Mumbai India"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Sandip","family":"Shingade","sequence":"additional","affiliation":[{"name":"Department of Computer Engineering and Information Technology Veermata Jijabai Technological Institute  Mumbai India"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Vijay","family":"Sambhe","sequence":"additional","affiliation":[{"name":"Department of Computer Engineering and Information Technology Veermata Jijabai Technological Institute  Mumbai India"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"265","published-online":{"date-parts":[[2025,7,14]]},"reference":[{"key":"e_1_2_17_2_1","doi-asserted-by":"publisher","DOI":"10.1109\/JPROC.2003.817150"},{"key":"e_1_2_17_3_1","doi-asserted-by":"publisher","DOI":"10.1080\/01690965.2013.834370"},{"key":"e_1_2_17_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2021.3107946"},{"key":"e_1_2_17_5_1","doi-asserted-by":"publisher","DOI":"10.1186\/s40537-019-0197-0"},{"key":"e_1_2_17_6_1","doi-asserted-by":"crossref","unstructured":"X.Zhao S.Yang S.Shan andXChen \u201cMutual Information Maximization for Effective Lip Reading \u201d in2020 15th IEEE International Conference on Automatic Face and Gesture Recognition (FG 2020)(IEEE 2020) 420\u2013427.","DOI":"10.1109\/FG47880.2020.00133"},{"key":"e_1_2_17_7_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.tins.2010.11.002"},{"key":"e_1_2_17_8_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-540-35488-8_1"},{"key":"e_1_2_17_9_1","doi-asserted-by":"publisher","DOI":"10.1006\/csla.2001.0185"},{"key":"e_1_2_17_10_1","doi-asserted-by":"crossref","unstructured":"J. S.ChungandA.Zisserman \u201cLip Reading in the Wild \u201d in13th Asian Conference on Computer Vision(Springer 2017) 87\u2013103.","DOI":"10.1007\/978-3-319-54184-6_6"},{"key":"e_1_2_17_11_1","unstructured":"J.Son Chung A.Senior O.Vinyals andAZisserman \u201cLip Reading Sentences in the Wild \u201d inProceedings of the IEEE Conference on Computer Vision and Pattern Recognition(IEEE 2017) 6447\u20136456 accessed 11 May 2023."},{"key":"e_1_2_17_12_1","unstructured":"T.Afouras J. S.Chung andA.Zisserman \u201cLRS3\u2010TED: A Large\u2010Scale Dataset for Visual Speech Recognition \u201d arXiv preprintarXiv:1809.00496 \u00a0 accessed 11 May 2023."},{"key":"e_1_2_17_13_1","unstructured":"D.Feng S.Yang S.Shan andXChen \u201cLearn an Effective Lip Reading Model Without Pains \u201d arXiv preprint arXiv:2011.07557 accessed 2020."},{"key":"e_1_2_17_14_1","doi-asserted-by":"crossref","unstructured":"B.Martinez P.Ma S.Petridis andMPantic \u201cLipreading Using Temporal Convolutional Networks \u201d in2020 IEEE International Conference on Acoustics Speech and Signal Processing (ICASSP)(IEEE 2020) 6319\u20136323.","DOI":"10.1109\/ICASSP40776.2020.9053841"},{"key":"e_1_2_17_15_1","doi-asserted-by":"crossref","unstructured":"P.Ma B.Martinez S.Petridis andMPantic \u201cTowards Practical Lipreading With Distilled and Efficient Models \u201d in2021 IEEE International Conference on Acoustics Speech and Signal Processing (ICASSP)(IEEE 2021) 7608\u20137612.","DOI":"10.1109\/ICASSP39728.2021.9415063"},{"key":"e_1_2_17_16_1","doi-asserted-by":"crossref","unstructured":"M.Kim J.Hong S. J.Park andY. M.Ro \u201cMulti\u2010Modality Associative Bridging Through Memory: Speech Sound Recollected From Face Video \u201d inProceedings of the IEEE\/CVF International Conference on Computer Vision(IEEE 2021) 296\u2013306.","DOI":"10.1109\/ICCV48922.2021.00036"},{"key":"e_1_2_17_17_1","doi-asserted-by":"crossref","unstructured":"P.Ma Y.Wang S.Petridis J.Shen andMPantic \u201cTraining Strategies for Improved Lip\u2010Reading \u201d in2022 IEEE International Conference on Acoustics Speech and Signal Processing (ICASSP)(IEEE 2022) 8472\u20138476.","DOI":"10.1109\/ICASSP43922.2022.9746706"},{"key":"e_1_2_17_18_1","doi-asserted-by":"crossref","unstructured":"D.Ivanko D.Ryumin A.Kashevnik A.Axyonov andA.Karnov \u201cVisual Speech Recognition in a Driver Assistance System \u201d in2022 30th European Signal Processing Conference (EUSIPCO)(IEEE 2022) 1131\u20131135.","DOI":"10.23919\/EUSIPCO55093.2022.9909819"},{"key":"e_1_2_17_19_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v36i1.20003"},{"key":"e_1_2_17_20_1","doi-asserted-by":"crossref","unstructured":"P.Ma S.Petridis andM.Pantic \u201cEnd\u2010to\u2010End Audio\u2010Visual Speech Recognition With Conformers \u201d in2021 IEEE International Conference on Acoustics Speech and Signal Processing (ICASSP)(IEEE 2021) 7613\u20137617.","DOI":"10.1109\/ICASSP39728.2021.9414567"},{"key":"e_1_2_17_21_1","doi-asserted-by":"crossref","unstructured":"K. R.Prajwal T.Afouras andA.Zisserman \u201cSub\u2010word Level Lip Reading With Visual Attention \u201d inProceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(IEEE 2022) 5162\u20135172.","DOI":"10.1109\/CVPR52688.2022.00510"},{"key":"e_1_2_17_22_1","doi-asserted-by":"publisher","DOI":"10.1038\/s42256-022-00550-z"},{"key":"e_1_2_17_23_1","doi-asserted-by":"crossref","unstructured":"P.Ma A.Haliassos A.Fernandez\u2010Lopez H.Chen S.Petridis andMPantic \u201cAuto\u2010AVSR: Audio\u2010Visual Speech Recognition With Automatic Labels \u201d in:2023 IEEE International Conference on Acoustics Speech and Signal Processing (ICASSP)(IEEE 2023) 1\u20135.","DOI":"10.1109\/ICASSP49357.2023.10096889"},{"key":"e_1_2_17_24_1","unstructured":"A.Haliassos P.Ma R.Mira S.Petridis andM.Pantic \u201cJointly Learning Visual and Auditory Speech Representations From Raw Data \u201d arXiv preprint arXiv:2212.06246 accessed 12 May 2023."},{"key":"e_1_2_17_25_1","doi-asserted-by":"crossref","unstructured":"B.Xu C.Lu Y.Guo andJWang \u201cDiscriminative Multi\u2010Modality Speech Recognition \u201d inProceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(IEEE 2020) 14433\u201314442.","DOI":"10.1109\/CVPR42600.2020.01444"},{"key":"e_1_2_17_26_1","unstructured":"S. L.Phung A.Bouzerdoum D.Chai andAWatson \u201cNaive Bayes Face\u2010Nonface Classifier: A Study of Preprocessing and Feature Extraction Techniques \u201d in2004 International Conference on Image Processing (ICIP)(IEEE 2004) 1385\u20131388."},{"key":"e_1_2_17_27_1","doi-asserted-by":"crossref","unstructured":"A.Chaari S.Lelandais V.Vigneron andM.Bedda \u201cFace Localization by Neural Networks Trained With Zernike Moments and Eigenfaces Feature Vectors. A Comparison \u201d in2007 IEEE Conference on Advanced Video and Signal Based Surveillance(IEEE 2007) 377\u2013382.","DOI":"10.1109\/AVSS.2007.4425340"},{"issue":"02","key":"e_1_2_17_28_1","article-title":"A Review on Face Detection Methods","volume":"11","author":"Rizvi Q. M.","year":"2011","journal-title":"Journal of Management Development and Information Technology"},{"key":"e_1_2_17_29_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10921-021-00768-8"},{"key":"e_1_2_17_30_1","doi-asserted-by":"publisher","DOI":"10.1109\/6046.865479"},{"key":"e_1_2_17_31_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.imavis.2014.06.004"},{"key":"e_1_2_17_32_1","doi-asserted-by":"crossref","unstructured":"N.ShrivastavaandV.Tyagi \u201cA Review of ROI Image Retrieval Techniques \u201d inProceedings of the 3rd International Conference on Frontiers of Intelligent Computing: Theory and Applications (FICTA) 2014(Springer 2015) 509\u2013520.","DOI":"10.1007\/978-3-319-12012-6_56"},{"key":"e_1_2_17_33_1","doi-asserted-by":"crossref","unstructured":"S.S\u00fcsstrunk R.Buckley andS.Swen \u201cStandard RGB Color Spaces \u201d inColor and Imaging Conference(Society of Imaging Science and Technology 1999) 127\u2013134.","DOI":"10.2352\/CIC.1999.7.1.art00024"},{"key":"e_1_2_17_34_1","doi-asserted-by":"publisher","DOI":"10.14257\/ijseia.2016.10.1.02"},{"key":"e_1_2_17_35_1","unstructured":"S. G. K.PatroandK. K.Sahu \u201cNormalization: A Preprocessing Stage \u201d arXiv preprint arXiv:1503.06462 accessed 30 May 2023."},{"key":"e_1_2_17_36_1","doi-asserted-by":"crossref","unstructured":"N.Fei Y.Gao Z.Lu andTXiang \u201cZ\u2010Score Normalization Hubness and Few\u2010Shot Learning \u201d inProceedings of the IEEE\/CVF International Conference on Computer Vision(IEEE 2021) 142\u2013151.","DOI":"10.1109\/ICCV48922.2021.00021"},{"key":"e_1_2_17_37_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.gltp.2022.04.020"},{"key":"e_1_2_17_38_1","doi-asserted-by":"crossref","unstructured":"J.Hu L.Shen andG.Sun \u201cSqueeze\u2010and Excitation Networks \u201d inProceedings of the IEEE Conference on Computer Vision and Pattern Recognition(IEEE 2018) 7132\u20137141.","DOI":"10.1109\/CVPR.2018.00745"},{"key":"e_1_2_17_39_1","doi-asserted-by":"publisher","DOI":"10.3390\/s22010072"},{"key":"e_1_2_17_40_1","doi-asserted-by":"publisher","DOI":"10.1109\/34.927467"},{"key":"e_1_2_17_41_1","first-page":"223","article-title":"An Introduction to Active Shape Models","volume":"328","author":"Cootes T.","year":"2000","journal-title":"Image Processing and Analysis"},{"key":"e_1_2_17_42_1","doi-asserted-by":"publisher","DOI":"10.1109\/T-C.1974.223784"},{"key":"e_1_2_17_43_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-1-4419-9878-1_4"},{"key":"e_1_2_17_44_1","doi-asserted-by":"publisher","DOI":"10.1002\/wics.101"},{"key":"e_1_2_17_45_1","doi-asserted-by":"crossref","unstructured":"Y.Fu X.Zhou M.Liu M.Hasegawa\u2010Johnson andT. SHuang \u201cLipreading by Locality Discriminant Graph \u201d in2007 IEEE International Conference on Image Processing(IEEE 2007) 325\u2013328.","DOI":"10.1109\/ICIP.2007.4379312"},{"key":"e_1_2_17_46_1","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2022.3175867"},{"key":"e_1_2_17_47_1","doi-asserted-by":"crossref","unstructured":"J.Peymanfard M. R.Mohammadi H.Zeinali andNMozayani \u201cLip Reading Using External Viseme Decoding \u201d in2022 International Conference on Machine Vision and Image Processing (MVIP)(IEEE 2022) 1\u20135.","DOI":"10.1109\/MVIP53647.2022.9738749"},{"key":"e_1_2_17_48_1","doi-asserted-by":"publisher","DOI":"10.1002\/9781119241485.ch25"},{"key":"e_1_2_17_49_1","doi-asserted-by":"crossref","unstructured":"T.Nunnally P.Chi K.Abdullah A. S.Uluagac J. A.Copeland andR.Beyah \u201cP3D: A Parallel 3D Coordinate Visualization for Advanced Network Scans \u201d in2013 IEEE International Conference on Communications (ICC)(IEEE 2013) 2052\u20132057.","DOI":"10.1109\/ICC.2013.6654828"},{"key":"e_1_2_17_50_1","doi-asserted-by":"publisher","DOI":"10.3390\/app10207208"},{"key":"e_1_2_17_51_1","doi-asserted-by":"crossref","unstructured":"J.Wu Y.Xu S. X.Zhang et\u00a0al. \u201cTime Domain Audio Visual Speech Separation \u201d in2019 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU)(IEEE 2019) 667\u2013673.","DOI":"10.1109\/ASRU46091.2019.9003983"},{"key":"e_1_2_17_52_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.engappai.2024.107862"},{"key":"e_1_2_17_53_1","doi-asserted-by":"crossref","unstructured":"P.Ma Y.Wang J.Shen S.Petridis andMPantic \u201cLip\u2010Reading With Densely Connected Temporal Convolutional Networks \u201d inProceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision(IEEE 2021) 2857\u20132866.","DOI":"10.1109\/WACV48630.2021.00290"},{"key":"e_1_2_17_54_1","doi-asserted-by":"crossref","unstructured":"K.He X.Zhang S.Ren andJ.Sun \u201cDeep Residual Learning for Image Recognition \u201d inProceedings of the IEEE Conference on Computer Vision and Pattern Recognition(IEEE 2016) 770\u2013778.","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_2_17_55_1","doi-asserted-by":"crossref","unstructured":"J.Huang L.Teng Y.Xiao A.Zhu andX.Liu \u201cLip Reading Using Temporal Adaptive Module \u201dinInternational Conference on Neural Information Processing(Springer 2023) 347\u2013356.","DOI":"10.1007\/978-981-99-8141-0_26"},{"key":"e_1_2_17_56_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11042-018-6912-6"},{"key":"e_1_2_17_57_1","doi-asserted-by":"publisher","DOI":"10.3390\/biomedinformatics4010023"},{"key":"e_1_2_17_58_1","unstructured":"M.Oghbaie A.Sabaghi K.Hashemifard andM.Akbari \u201cAdvances and Challenges in Deep Lip Reading \u201d arXiv preprint arXiv:2110.07879 accessed 1 June 2023."},{"key":"e_1_2_17_59_1","doi-asserted-by":"publisher","DOI":"10.28991\/HIJ-2023-04-02-010"},{"key":"e_1_2_17_60_1","unstructured":"A.Garg J.Noyola andS.Bagadia \u201cLip Reading Using CNN and LSTM \u201d2016 accessed 10 June 2023 https:\/\/api.semanticscholar.org\/CorpusID:22889293."},{"key":"e_1_2_17_61_1","doi-asserted-by":"crossref","unstructured":"S.Petridis Z.Li andM.Pantic \u201cEndto\u2010End Visual Speech Recognition With LSTMs \u201d in2017 IEEE International Conference on Acoustics Speech and Signal Processing (ICASSP)(IEEE 2017) 2592\u20132596.","DOI":"10.1109\/ICASSP.2017.7952625"},{"key":"e_1_2_17_62_1","doi-asserted-by":"crossref","unstructured":"A.David Rasamoelina F.Adjailia andP.Sin\u010d\u00e1k \u201cA Review of Activation Function for Artificial Neural Network \u201d in2020 IEEE 18th World Symposium on Applied Machine Intelligence and Informatics (SAMI)(IEEE 2020) 281\u2013286.","DOI":"10.1109\/SAMI48414.2020.9108717"},{"key":"e_1_2_17_63_1","doi-asserted-by":"crossref","unstructured":"B.Hampiholi C.Jarvers W.Mader andH.Neumann \u201cDepthwise Separable Temporal Convolutional Network for Action Segmentation \u201d in2020 International Conference on 3D Vision (3DV)(IEEE 2020) 633\u2013641.","DOI":"10.1109\/3DV50981.2020.00073"},{"key":"e_1_2_17_64_1","doi-asserted-by":"publisher","DOI":"10.3390\/app132011127"},{"key":"e_1_2_17_65_1","doi-asserted-by":"publisher","DOI":"10.1162\/neco.1997.9.8.1735"},{"key":"e_1_2_17_66_1","doi-asserted-by":"publisher","DOI":"10.3758\/BF03196492"},{"key":"e_1_2_17_67_1","doi-asserted-by":"crossref","unstructured":"A.PandeyandD.eL.Wang \u201cTCNN: Temporal Convolutional Neural Network for Real\u2010Time Speech Enhancement in the Time Domain \u201d in2019 IEEE International Conference on Acoustics Speech and Signal Processing (ICASSP)(IEEE 2019) 6875\u20136879.","DOI":"10.1109\/ICASSP.2019.8683634"},{"key":"e_1_2_17_68_1","doi-asserted-by":"publisher","DOI":"10.3390\/app11156975"},{"issue":"9","key":"e_1_2_17_69_1","first-page":"2577","article-title":"Survey on Text Transformation Using Bi\u2010LSTM in Natural Language Processing With Text Data","volume":"12","author":"Preethi V.","year":"2021","journal-title":"Turkish Journal of Computer and Mathematics Education (TURCOMAT)"},{"key":"e_1_2_17_70_1","doi-asserted-by":"crossref","unstructured":"R. C.Joshi A.Juyal V.Jain andS.Chaturvedi \u201cAn Efficient Approach to Lip\u2010Reading With 3D CNN and Bi\u2010LSTM Fusion Model \u201d inInternational Conference on Cognitive Computing and Cyber Physical Systems(Springer 2025) 15\u201328.","DOI":"10.1007\/978-981-97-7371-8_2"},{"key":"e_1_2_17_71_1","unstructured":"F. M.Shiri T.Perumal N.Mustapha andR.Mohamed \u201cA Comprehensive Overview and Comparative Analysis on Deep Learning Models: CNN RNN LSTM GRU \u201d accessed May 2023 https:\/\/doi.org\/10.48550\/arXiv.2305.17473."},{"key":"e_1_2_17_72_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2023.121648"},{"key":"e_1_2_17_73_1","unstructured":"D. P.KingmaandJ.Ba \u201cAdam: A Method for Stochastic Optimization \u201d arXiv preprint arXiv:1412.6980 accessed 12 June 2023."},{"key":"e_1_2_17_74_1","doi-asserted-by":"crossref","unstructured":"B.Alabassy M.Safar andM. W.El\u2010Kharashi \u201cA High\u2010Accuracy Implementation for Softmax Layer in Deep Neural Networks \u201d in2020 15th Design & Technology of Integrated Systems in Nanoscale Era (DTIS)(IEEE 2020) 1\u20136.","DOI":"10.1109\/DTIS48698.2020.9081313"},{"key":"e_1_2_17_75_1","doi-asserted-by":"crossref","unstructured":"J.XiaoandZ.Zhou \u201cResearch Progress of RNN Language Model \u201d in2020 IEEE International Conference on Artificial Intelligence and Computer Applications (ICAICA)(IEEE 2020) 1285\u20131288.","DOI":"10.1109\/ICAICA50127.2020.9182390"},{"key":"e_1_2_17_76_1","unstructured":"K.Bae H.Ryu andH.Shin \u201cDoes Adam Optimizer Keep Close to the Optimal Point? \u201d arXiv preprintarXiv:1911.00289 accessed 21 June 2023."},{"key":"e_1_2_17_77_1","doi-asserted-by":"crossref","unstructured":"B.\u2010Y.Hsueh W.Li andI.\u2010C.Wu \u201cStochastic Gradient Descent With Hyperbolic\u2010tangent Decay on Classification \u201d in2019 IEEE Winter Conference on Applications of Computer Vision (WACV)(IEEE 2019) 435\u2013442.","DOI":"10.1109\/WACV.2019.00052"},{"key":"e_1_2_17_78_1","doi-asserted-by":"crossref","unstructured":"M.Wand J.Koutn\u00edk andJ.Schmidhuber \u201cLipreading With Long Short\u2010Term Memory \u201d in2016 IEEE International Conference on Acoustics Speech and Signal Processing (ICASSP)(IEEE 2016) 6115\u20136119.","DOI":"10.1109\/ICASSP.2016.7472852"},{"key":"e_1_2_17_79_1","unstructured":"T.StafylakisandG.Tzimiropoulos \u201cCombining Residual Networks With LSTMs for Lipreading \u201d arXiv preprint arXiv:1703.04105 accessed 29 June 2023."},{"key":"e_1_2_17_80_1","doi-asserted-by":"publisher","DOI":"10.1109\/TASLP.2023.3282109"},{"key":"e_1_2_17_81_1","doi-asserted-by":"crossref","unstructured":"A.Al\u2010Sabaawi H. M.Ibrahim Z. M.Arkah M.Al\u2010Amidie andLAlzubaidi \u201cAmended Convolutional Neural Network With Global Average Pooling for Image Classification \u201d inInternational Conference on Intelligent Systems Design and Applications(Springer 2020) 171\u2013180.","DOI":"10.1007\/978-3-030-71187-0_16"},{"key":"e_1_2_17_82_1","doi-asserted-by":"crossref","unstructured":"R.Liao D.Moyer M.Cha et\u00a0al. \u201cMultimodal Representation Learning via Maximization of Local Mutual Information \u201d inMedical Image Computing and Computer Assisted Intervention\u2013MICCAI 2021 Part II(Springer 2021) 273\u2013283.","DOI":"10.1007\/978-3-030-87196-3_26"},{"key":"e_1_2_17_83_1","doi-asserted-by":"crossref","unstructured":"B.Xu J.Wang C.Lu andYGuo \u201cWatch to Listen Clearly: Visual Speech Enhancement Driven Multi\u2010Modality Speech Recognition \u201d inProceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision(IEEE 2020) 1637\u20131646.","DOI":"10.1109\/WACV45572.2020.9093314"},{"key":"e_1_2_17_84_1","doi-asserted-by":"publisher","DOI":"10.1117\/1.JEI.32.2.023001"},{"key":"e_1_2_17_85_1","doi-asserted-by":"crossref","unstructured":"K.Suresh G.Gopakumar andS.Duttagupta \u201cGenerating Audio From Lip Movements Visual Input: A Survey \u201d inIntelligent Systems Technologies and Applications: Proceedings of Sixth ISTA2020(Springer 2021) 315\u2013326.","DOI":"10.1007\/978-981-16-0730-1_21"},{"key":"e_1_2_17_86_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2020.3021756"},{"key":"e_1_2_17_87_1","doi-asserted-by":"crossref","unstructured":"N.Ma X.Zhang H. T.Zheng andJ.Sun \u201cShufflenet V2: Practical Guidelines for Efficient CNN Architecture Design \u201d inProceedings of the European Conference on Computer Vision (ECCV)(Springer 2018) 116\u2013131.","DOI":"10.1007\/978-3-030-01264-9_8"},{"key":"e_1_2_17_88_1","doi-asserted-by":"crossref","unstructured":"F.Chollet \u201cXception: Deep Learning With Depthwise Separable Convolutions \u201d inProceedings of the IEEE Conference on Computer Vision and Pattern Recognition(IEEE 2017) 1251\u20131258.","DOI":"10.1109\/CVPR.2017.195"},{"key":"e_1_2_17_89_1","unstructured":"T.Furlanello Z.Lipton M.Tschannen L.Itti andAAnandkumar \u201cBorn Again Neural Net\u2010Works \u201d inInternational Conference on Machine Learning(PMLR 2018) 1607\u20131616."},{"key":"e_1_2_17_90_1","unstructured":"J.Weston S.Chopra andA.Bordes \u201cMemory Networks \u201d arXiv preprint arXiv:1410.3916 accessed 2 July 2023."},{"key":"e_1_2_17_91_1","doi-asserted-by":"crossref","unstructured":"Z.Wang S.Chang J.Zhou M.Wang andT. SHuang. \u201cLearning a Task\u2010Specific Deep Architecture for Clustering \u201d inProceedings of the 2016 SIAM International Conference on Data Mining(SIAM 2016) 369\u2013377.","DOI":"10.1137\/1.9781611974348.42"},{"key":"e_1_2_17_92_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-1-4842-4470-8_31"},{"key":"e_1_2_17_93_1","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2022.3204677"},{"key":"e_1_2_17_94_1","doi-asserted-by":"publisher","DOI":"10.1145\/3530811"},{"key":"e_1_2_17_95_1","unstructured":"M. V.Koroteev \u201cBERT: A Review of Applications in Natural Language Processing and Understanding \u201d arXiv preprint arXiv:2103.11943 accessed 10 July 2023."},{"key":"e_1_2_17_96_1","unstructured":"X.Song A.Salcianu Y.Song D.Dopson andDZhou \u201cFast Wordpiece Tokenization \u201d arXiv preprint arXiv:2012.15524 accessed 21 July 2023."},{"key":"e_1_2_17_97_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-33808-3_14"},{"key":"e_1_2_17_98_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.array.2022.100258"},{"key":"e_1_2_17_99_1","unstructured":"Y.KhaireddinandZ.Chen \u201cFacial Emotion Recognition: State of the Art Performance on FER2013 \u201d arXiv preprint arXiv:2105.03588 accessed 25 July 2023."},{"key":"e_1_2_17_100_1","doi-asserted-by":"crossref","unstructured":"A. A.Bengeri S.Jain R.Devaranavadagi et\u00a0al. \u201cFace Counting Based on Pretrained Machine Learning Models: A Brief Systematic Review \u201d inInternational Conference on Data Science and Applications(Springer 2023) 353\u2013364.","DOI":"10.1007\/978-981-99-7820-5_29"},{"key":"e_1_2_17_101_1","unstructured":"Y.Kartynnik A.Ablavatski I.Grishchenko andM.Grundmann \u201cReal\u2010Time Facial Surface Geometry From Monocular Video on Mobile GPUs \u201d arXiv preprint arXiv:1907.06724 accessed 16 July 2023."},{"key":"e_1_2_17_102_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10479-005-5724-z"},{"key":"e_1_2_17_103_1","doi-asserted-by":"publisher","DOI":"10.54187\/jnrs.1011739"},{"key":"e_1_2_17_104_1","doi-asserted-by":"crossref","unstructured":"A.Ephrat I.Mosseri O.Lang T.Dekel K.Wilson andA.Hassidim \u201cLooking to Listen at the Cocktail Party: A Speaker\u2010Independent Audio\u2010Visual Model for Speech Separation\u201d (2018) https:\/\/doi.org\/10.48550\/arXiv.1804.03619.","DOI":"10.1145\/3197517.3201357"},{"key":"e_1_2_17_105_1","doi-asserted-by":"crossref","unstructured":"J. S.Chung A.Nagrani andA.Zisserman \u201cVoxceleb2: Deep Speaker Recognition \u201d (2018) https:\/\/doi.org\/10.48550\/arXiv.1806.05622.","DOI":"10.21437\/Interspeech.2018-1929"},{"key":"e_1_2_17_106_1","unstructured":"C.Yi J.Wang N.Cheng S.Zhou andB.Xu \u201cApplying wav2vec2. 0 to Speech Recognition in Various Low\u2010Resource Languages \u201d (2020) https:\/\/doi.org\/10.48550\/arXiv.2012.12121."},{"key":"e_1_2_17_107_1","doi-asserted-by":"publisher","DOI":"10.1109\/TASLP.2021.3122291"},{"key":"e_1_2_17_108_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.aiopen.2022.10.001"},{"key":"e_1_2_17_109_1","doi-asserted-by":"publisher","DOI":"10.1561\/116.00000050"},{"key":"e_1_2_17_110_1","doi-asserted-by":"publisher","DOI":"10.3390\/s23042284"},{"key":"e_1_2_17_111_1","first-page":"1","article-title":"Lip Reading Using Various Deep Learning Models With Visual Turkish Data","volume":"1","author":"Berkol A.","year":"2024","journal-title":"Gazi University Journal of Science"},{"key":"e_1_2_17_112_1","doi-asserted-by":"crossref","unstructured":"M.BurchiandV.Vielzeuf \u201cEfficient Conformer: Progressive Down Sampling and Grouped Attention for Automatic Speech Recognition \u201d in2021 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU)(IEEE 2021) 8\u201315.","DOI":"10.1109\/ASRU51503.2021.9687874"},{"key":"e_1_2_17_113_1","doi-asserted-by":"crossref","unstructured":"Y.Fan X.Lu D.Li andYLiu \u201cVideo\u2010Based Emotion Recognition Using CNN\u2010RNN and C3D Hybrid Networks \u201d inProceedings of the 18th ACM International Conference on Multimodal Interaction(2016) 445\u2013450.","DOI":"10.1145\/2993148.2997632"},{"key":"e_1_2_17_114_1","unstructured":"S. F.Chen D.Beeferman andR.Rosenfeld \u201cEvaluation Metrics for Language Models \u201d February1998 https:\/\/www.cs.cmu.edu\/~roni\/papers\/eval-metrics-bntuw-9802.pdf."},{"key":"e_1_2_17_115_1","doi-asserted-by":"crossref","unstructured":"A.Das J.Li R.Zhao andY.Gong \u201cAdvancing Connectionist Temporal Classification With Attention Modeling \u201d in2018 IEEE International Conference on Acoustics Speech and Signal Processing (ICASSP)(IEEE 2018) 4769\u20134773.","DOI":"10.1109\/ICASSP.2018.8461558"},{"key":"e_1_2_17_116_1","doi-asserted-by":"crossref","unstructured":"M.Rane V.Chaudhari M.Vasant B.Nambiar S.Umak andAKulkarni \u201cSeeing the Unheard: Lip Reading With Deep Learning \u201d in2023 International Conference on Evolutionary Algorithms and Soft Computing Techniques (EASCT)(IEEE 2023) 1\u20136.","DOI":"10.1109\/EASCT59475.2023.10393197"}],"container-title":["IET Image Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/ietresearch.onlinelibrary.wiley.com\/doi\/pdf\/10.1049\/ipr2.70095","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/ietresearch.onlinelibrary.wiley.com\/doi\/full-xml\/10.1049\/ipr2.70095","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/ietresearch.onlinelibrary.wiley.com\/doi\/pdf\/10.1049\/ipr2.70095","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,5,8]],"date-time":"2026-05-08T23:37:15Z","timestamp":1778283435000},"score":1,"resource":{"primary":{"URL":"https:\/\/ietresearch.onlinelibrary.wiley.com\/doi\/10.1049\/ipr2.70095"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,1]]},"references-count":115,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2025,1]]}},"alternative-id":["10.1049\/ipr2.70095"],"URL":"https:\/\/doi.org\/10.1049\/ipr2.70095","archive":["Portico"],"relation":{},"ISSN":["1751-9659","1751-9667"],"issn-type":[{"value":"1751-9659","type":"print"},{"value":"1751-9667","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,1]]},"assertion":[{"value":"2024-04-18","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-04-21","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-07-14","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}],"article-number":"e70095"}}