{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,1]],"date-time":"2026-04-01T02:30:50Z","timestamp":1775010650794,"version":"3.50.1"},"publisher-location":"New York, NY, USA","reference-count":30,"publisher":"ACM","license":[{"start":{"date-parts":[[2022,11,7]],"date-time":"2022-11-07T00:00:00Z","timestamp":1667779200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2022,11,7]]},"DOI":"10.1145\/3536221.3556571","type":"proceedings-article","created":{"date-parts":[[2022,11,4]],"date-time":"2022-11-04T15:54:14Z","timestamp":1667577254000},"page":"368-372","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":6,"title":["Is Lip Region-of-Interest Sufficient for Lipreading?"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-4341-3174","authenticated-orcid":false,"given":"Jing-Xuan","family":"Zhang","sequence":"first","affiliation":[{"name":"iFLYTEK Research, iFLYTEK Co., Ltd., China and University of Science and Technology of China, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Genshun","family":"Wan","sequence":"additional","affiliation":[{"name":"iFLYTEK Research, iFLYTEK Co., Ltd, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jia","family":"Pan","sequence":"additional","affiliation":[{"name":"iFLYTEK Research, iFLYTEK Co., Ltd., China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2022,11,7]]},"reference":[{"key":"e_1_3_2_1_1_1","volume-title":"Deep audio-visual speech recognition","author":"Afouras Triantafyllos","year":"2018","unstructured":"Triantafyllos Afouras , Joon\u00a0Son Chung , Andrew Senior , Oriol Vinyals , and Andrew Zisserman . 2018. Deep audio-visual speech recognition . IEEE transactions on pattern analysis and machine intelligence ( 2018 ). Triantafyllos Afouras, Joon\u00a0Son Chung, Andrew Senior, Oriol Vinyals, and Andrew Zisserman. 2018. Deep audio-visual speech recognition. IEEE transactions on pattern analysis and machine intelligence (2018)."},{"key":"e_1_3_2_1_2_1","unstructured":"T. Afouras J.\u00a0S. Chung and A. Zisserman. 2018. LRS3-TED: a large-scale dataset for visual speech recognition. In arXiv preprint arXiv:1809.00496.  T. Afouras J.\u00a0S. Chung and A. Zisserman. 2018. LRS3-TED: a large-scale dataset for visual speech recognition. In arXiv preprint arXiv:1809.00496."},{"key":"e_1_3_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP40776.2020.9054253"},{"key":"e_1_3_2_1_5_1","first-page":"12449","article-title":"Wav2vec 2.0: A framework for self-supervised learning of speech representations","volume":"33","author":"Baevski Alexei","year":"2020","unstructured":"Alexei Baevski , Yuhao Zhou , Abdelrahman Mohamed , and Michael Auli . 2020 . Wav2vec 2.0: A framework for self-supervised learning of speech representations . Advances in Neural Information Processing Systems 33 (2020), 12449 \u2013 12460 . Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli. 2020. Wav2vec 2.0: A framework for self-supervised learning of speech representations. Advances in Neural Information Processing Systems 33 (2020), 12449\u201312460.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_1_6_1","volume-title":"BEiT: BERT Pre-Training of Image Transformers. In ICLR","author":"Bao Hangbo","year":"2022","unstructured":"Hangbo Bao , Li Dong , Songhao Piao , and Furu Wei . 2022 . BEiT: BERT Pre-Training of Image Transformers. In ICLR 2022. Hangbo Bao, Li Dong, Songhao Piao, and Furu Wei. 2022. BEiT: BERT Pre-Training of Image Transformers. In ICLR 2022."},{"key":"e_1_3_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2017.7953127"},{"key":"e_1_3_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2020.3036865"},{"key":"e_1_3_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/TASLP.2021.3122291"},{"key":"e_1_3_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/ASRU.2011.6163922"},{"key":"e_1_3_2_1_11_1","first-page":"1597","article-title":"Design and implementation of a real-time lipreading system using PCA and HMM","volume":"7","author":"Lee CG","year":"2004","unstructured":"CG Lee , ES Lee , ST Jung , and SS Lee . 2004 . Design and implementation of a real-time lipreading system using PCA and HMM . Journal of Korea Multimedia Society 7 , 11 (2004), 1597 \u2013 1609 . CG Lee, ES Lee, ST Jung, and SS Lee. 2004. Design and implementation of a real-time lipreading system using PCA and HMM. Journal of Korea Multimedia Society 7, 11 (2004), 1597\u20131609.","journal-title":"Journal of Korea Multimedia Society"},{"key":"e_1_3_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICSLP.1996.607030"},{"key":"e_1_3_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP39728.2021.9414567"},{"key":"e_1_3_2_1_14_1","volume-title":"Recurrent neural network transducer for audio-visual speech recognition. In 2019 IEEE automatic speech recognition and understanding workshop (ASRU)","author":"Makino Takaki","unstructured":"Takaki Makino , Hank Liao , Yannis Assael , Brendan Shillingford , Basilio Garcia , Otavio Braga , and Olivier Siohan . 2019. Recurrent neural network transducer for audio-visual speech recognition. In 2019 IEEE automatic speech recognition and understanding workshop (ASRU) . IEEE , 905\u2013912. Takaki Makino, Hank Liao, Yannis Assael, Brendan Shillingford, Basilio Garcia, Otavio Braga, and Olivier Siohan. 2019. Recurrent neural network transducer for audio-visual speech recognition. In 2019 IEEE automatic speech recognition and understanding workshop (ASRU). IEEE, 905\u2013912."},{"key":"e_1_3_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP40776.2020.9053841"},{"key":"e_1_3_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2015.7178347"},{"key":"e_1_3_2_1_17_1","doi-asserted-by":"crossref","unstructured":"Kuniaki Noda Yuki Yamaguchi Kazuhiro Nakadai Hiroshi\u00a0G Okuno and Tetsuya Ogata. 2014. Lipreading using convolutional neural network. In fifteenth annual conference of the international speech communication association.  Kuniaki Noda Yuki Yamaguchi Kazuhiro Nakadai Hiroshi\u00a0G Okuno and Tetsuya Ogata. 2014. Lipreading using convolutional neural network. In fifteenth annual conference of the international speech communication association.","DOI":"10.21437\/Interspeech.2014-293"},{"key":"e_1_3_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2018.8461326"},{"key":"e_1_3_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/SLT.2018.8639643"},{"key":"e_1_3_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-25903-1_33"},{"key":"e_1_3_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2014-503"},{"key":"e_1_3_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP40776.2020.9054249"},{"key":"e_1_3_2_1_23_1","unstructured":"Bowen Shi Wei-Ning Hsu Kushal Lakhotia and Abdelrahman Mohamed. 2022. Learning audio-visual speech representation by masked multimodal cluster prediction. arXiv preprint arXiv:2201.02184(2022).  Bowen Shi Wei-Ning Hsu Kushal Lakhotia and Abdelrahman Mohamed. 2022. Learning audio-visual speech representation by masked multimodal cluster prediction. arXiv preprint arXiv:2201.02184(2022)."},{"key":"e_1_3_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.21437\/AVSP.2017-14"},{"key":"e_1_3_2_1_25_1","unstructured":"Georgios\u00a0Tzimiropoulos Themos\u00a0Stafylakis. 2017. Combining Residual Networks with LSTMs for Lipreading. In INTERSPEECH. 3652\u20133656.  Georgios\u00a0Tzimiropoulos Themos\u00a0Stafylakis. 2017. Combining Residual Networks with LSTMs for Lipreading. In INTERSPEECH. 3652\u20133656."},{"key":"e_1_3_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2014.6854363"},{"key":"e_1_3_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2016.7472852"},{"key":"e_1_3_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.01444"},{"key":"e_1_3_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v35i16.17693"},{"key":"e_1_3_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1109\/FG47880.2020.00134"},{"key":"e_1_3_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2011.5995345"}],"event":{"name":"ICMI '22: INTERNATIONAL CONFERENCE ON MULTIMODAL INTERACTION","location":"Bengaluru India","acronym":"ICMI '22","sponsor":["SIGCHI ACM Special Interest Group on Computer-Human Interaction"]},"container-title":["Proceedings of the 2022 International Conference on Multimodal Interaction"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3536221.3556571","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3536221.3556571","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T17:48:52Z","timestamp":1750182532000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3536221.3556571"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,11,7]]},"references-count":30,"alternative-id":["10.1145\/3536221.3556571","10.1145\/3536221"],"URL":"https:\/\/doi.org\/10.1145\/3536221.3556571","relation":{},"subject":[],"published":{"date-parts":[[2022,11,7]]},"assertion":[{"value":"2022-11-07","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}