{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,3]],"date-time":"2026-07-03T07:04:57Z","timestamp":1783062297511,"version":"3.54.6"},"reference-count":27,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2015,1,7]],"date-time":"2015-01-07T00:00:00Z","timestamp":1420588800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2015,1,7]]},"abstract":"<jats:p>We propose a new method for improving the presentation of subtitles in video (e.g., TV and movies). With conventional subtitles, the viewer has to constantly look away from the main viewing area to read the subtitles at the bottom of the screen, which disrupts the viewing experience and causes unnecessary eyestrain. Our method places on-screen subtitles next to the respective speakers to allow the viewer to follow the visual content while simultaneously reading the subtitles. We use novel identification algorithms to detect the speakers based on audio and visual information. Then the placement of the subtitles is determined using global optimization. A comprehensive usability study indicated that our subtitle placement method outperformed both conventional fixed-position subtitling and another previous dynamic subtitling method in terms of enhancing the overall viewing experience and reducing eyestrain.<\/jats:p>","DOI":"10.1145\/2632111","type":"journal-article","created":{"date-parts":[[2015,1,12]],"date-time":"2015-01-12T20:02:10Z","timestamp":1421092930000},"page":"1-17","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":42,"title":["Speaker-Following Video Subtitles"],"prefix":"10.1145","volume":"11","author":[{"given":"Yongtao","family":"Hu","sequence":"first","affiliation":[{"name":"The University of Hong Kong"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jan","family":"Kautz","sequence":"additional","affiliation":[{"name":"University College London"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yizhou","family":"Yu","sequence":"additional","affiliation":[{"name":"The University of Hong Kong"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Wenping","family":"Wang","sequence":"additional","affiliation":[{"name":"The University of Hong Kong"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2015,1,7]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1109\/TASL.2011.2125954"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1007\/11919629_58"},{"key":"e_1_2_1_3_1","volume-title":"Proceedings of the 5th EURASIP Conference on Speech and Image Processing, Multimedia Communications and Service.","author":"Dimou A.","year":"2005","unstructured":"A. Dimou , O. Nemethova , and M. Rupp 2005 . Scene change detection for h. 264 using dynamic threshold techniques . In Proceedings of the 5th EURASIP Conference on Speech and Image Processing, Multimedia Communications and Service. A. Dimou, O. Nemethova, and M. Rupp 2005. Scene change detection for h. 264 using dynamic threshold techniques. In Proceedings of the 5th EURASIP Conference on Speech and Image Processing, Multimedia Communications and Service."},{"key":"e_1_2_1_4_1","doi-asserted-by":"crossref","unstructured":"J. Driver. 1996. Enhancement of selective listening by illusory mislocation of speech sounds due to lip-reading. Nature 381 6577 66--8.  J. Driver. 1996. Enhancement of selective listening by illusory mislocation of speech sounds due to lip-reading. Nature 381 6577 66--8.","DOI":"10.1038\/381066a0"},{"key":"e_1_2_1_5_1","volume-title":"Proceedings of the 17th British Machine Vision Conference (BMVC'06)","author":"Everingham M.","unstructured":"M. Everingham , J. Sivic , and A. Zisserman . 2006. Hello! my name is... buffy--automatic naming of characters in TV video . In Proceedings of the 17th British Machine Vision Conference (BMVC'06) . M. Everingham, J. Sivic, and A. Zisserman. 2006. Hello! my name is... buffy--automatic naming of characters in TV video. In Proceedings of the 17th British Machine Vision Conference (BMVC'06)."},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1049\/ecej:20010302"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1155\/S1110865702207039"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/1873951.1874013"},{"key":"e_1_2_1_9_1","volume-title":"Proceedings of RIAO: Coupling Approaches, Coupling Media and Coupling Languages for Information Retrieval. 314--325","author":"Jaffr\u00e9 G.","year":"2004","unstructured":"G. Jaffr\u00e9 , P. Joly , 2004 . Costume: A new feature for automatic video content indexing . In Proceedings of RIAO: Coupling Approaches, Coupling Media and Coupling Languages for Information Retrieval. 314--325 . G. Jaffr\u00e9, P. Joly, et al. 2004. Costume: A new feature for automatic video content indexing. In Proceedings of RIAO: Coupling Approaches, Coupling Media and Coupling Languages for Information Retrieval. 314--325."},{"key":"e_1_2_1_10_1","unstructured":"M. A. Just and P. A. Carpenter. 1987. The Psychology of Reading and Language Comprehension. ERIC.  M. A. Just and P. A. Carpenter. 1987. The Psychology of Reading and Language Comprehension. ERIC."},{"key":"e_1_2_1_11_1","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. IEEE, 685--692","author":"Kuo C.","unstructured":"C. Kuo , C. Huang , and R. Nevatia . 2010. Multi-target tracking by on-line learned discriminative appearance models . In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. IEEE, 685--692 . C. Kuo, C. Huang, and R. Nevatia. 2010. Multi-target tracking by on-line learned discriminative appearance models. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. IEEE, 685--692."},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/237170.237260"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.3758\/BF03208086"},{"key":"e_1_2_1_14_1","volume-title":"Proceedings of the 19th European Signal Processing Conference.","author":"Monaci G.","year":"2011","unstructured":"G. Monaci . 2011 . Towards real-time audiovisual speaker localization . In Proceedings of the 19th European Signal Processing Conference. G. Monaci. 2011. Towards real-time audiovisual speaker localization. In Proceedings of the 19th European Signal Processing Conference."},{"key":"e_1_2_1_15_1","doi-asserted-by":"crossref","unstructured":"H. Nock G. Iyengar and C. Neti. 2003. Speaker localisation using audio-visual synchrony: An empirical study. In Image and Video Retrieval 565--570.   H. Nock G. Iyengar and C. Neti. 2003. Speaker localisation using audio-visual synchrony: An empirical study. In Image and Video Retrieval 565--570.","DOI":"10.1007\/3-540-45113-7_48"},{"key":"e_1_2_1_16_1","volume-title":"Proceedings of the International Conference on Games Research and Development.","author":"Park S.-H.","year":"2008","unstructured":"S.-H. Park , S.-H. Ji , D.-S. Ryu , and H.-G. Cho . 2008 a. A smart and realistic chatting interface for gaming agents in 3-d virtual space . In Proceedings of the International Conference on Games Research and Development. S.-H. Park, S.-H. Ji, D.-S. Ryu, and H.-G. Cho. 2008a. A smart and realistic chatting interface for gaming agents in 3-d virtual space. In Proceedings of the International Conference on Games Research and Development."},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICHIT.2008.186"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/JPROC.2003.817150"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1016\/0010-0285(75)90005-5"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/1027933.1027960"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2005.251"},{"key":"e_1_2_1_22_1","volume-title":"Proceedings of IS&T\/SPIE's Symposium on Electronic Imaging: Science & Technology. International Society for Optics and Photonics, 329--338","author":"Sethi I. K.","unstructured":"I. K. Sethi and N. V. Patel . 1995. Statistical approach to scene change detection . In Proceedings of IS&T\/SPIE's Symposium on Electronic Imaging: Science & Technology. International Society for Optics and Photonics, 329--338 . I. K. Sethi and N. V. Patel. 1995. Statistical approach to scene change detection. In Proceedings of IS&T\/SPIE's Symposium on Electronic Imaging: Science & Technology. International Society for Optics and Photonics, 329--338."},{"key":"e_1_2_1_23_1","volume-title":"Proceedings of the 7th International Conference on Computer Vision Theory and Applications.","volume":"1","author":"Uric\u00e1r M.","unstructured":"M. Uric\u00e1r , V. Franc , and V. Hlav\u00e1c . 2012. Detector of facial landmarks learned by the structured output svm . In Proceedings of the 7th International Conference on Computer Vision Theory and Applications. Vol. 1 , 547--556. M. Uric\u00e1r, V. Franc, and V. Hlav\u00e1c. 2012. Detector of facial landmarks learned by the structured output svm. In Proceedings of the 7th International Conference on Computer Vision Theory and Applications. Vol. 1, 547--556."},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1023\/B:VISI.0000013087.49260.fb"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1007\/s00221-004-1899-9"},{"key":"e_1_2_1_26_1","unstructured":"Wikipedia. 2012. Vision span.  Wikipedia. 2012. Vision span."},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/957013.957090"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2632111","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/2632111","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T06:56:12Z","timestamp":1750229772000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2632111"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2015,1,7]]},"references-count":27,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2015,1,7]]}},"alternative-id":["10.1145\/2632111"],"URL":"https:\/\/doi.org\/10.1145\/2632111","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2015,1,7]]},"assertion":[{"value":"2013-10-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2014-04-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2015-01-07","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}