{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2022,4,5]],"date-time":"2022-04-05T22:58:30Z","timestamp":1649199510624},"reference-count":19,"publisher":"World Scientific Pub Co Pte Lt","issue":"02","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Int. J. Semantic Computing"],"published-print":{"date-parts":[[2012,6]]},"abstract":"<jats:p> We propose a method for discriminating between a speech shot and a narrated shot to extract genuine speech shots from a broadcast news video. Speech shots in news videos contain a wealth of multimedia information of the speaker, and could thus be considered valuable as archived material. In order to extract speech shots from news videos, there is an approach that uses the position and size of a face region. However, it is difficult to extract them with only such an approach, since news videos contain non-speech shots where the speaker is not the subject that appears in the screen, namely, narrated shots. To solve this problem, we propose a method to discriminate between a speech shot and a narrated shot in two stages. The first stage of the proposed method directly evaluates the inconsistency between a subject and a speaker based on the co-occurrence between lip motion and voice. The second stage of the proposed method evaluates based on the intra- and inter-shot features that focus on the tendency of speech shots. With the combination of both stages, the proposed method accurately discriminates between a speech shot and a narrated shot. In the experiments, the overall accuracy of speech shots extraction by the proposed method was 0.871. Therefore, we confirmed the effectiveness of the proposed method. <\/jats:p>","DOI":"10.1142\/s1793351x12400077","type":"journal-article","created":{"date-parts":[[2012,10,22]],"date-time":"2012-10-22T09:26:52Z","timestamp":1350898012000},"page":"179-204","source":"Crossref","is-referenced-by-count":5,"title":["SPEECH SHOT EXTRACTION FROM BROADCAST NEWS VIDEOS"],"prefix":"10.1142","volume":"06","author":[{"given":"SHOGO","family":"KUMAGAI","sequence":"first","affiliation":[{"name":"Graduate School of Information Science, Nagoya University, Furo-cho, Chikusa-ku, Nagoya, Aichi, 464-8601, Japan"},{"name":"Currently at Ricoh Company, Ltd., Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"KEISUKE","family":"DOMAN","sequence":"additional","affiliation":[{"name":"Graduate School of Information Science, Nagoya University, Furo-cho, Chikusa-ku, Nagoya, Aichi, 464-8601, Japan"},{"name":"Japan Society for the Promotion of Science (JSPS), Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"TOMOKAZU","family":"TAKAHASHI","sequence":"additional","affiliation":[{"name":"Faculty of Economics and Information, Gifu Shotoku Gakuen University, 1-38 Nakauzura, Gifu, 500-8288, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"DAISUKE","family":"DEGUCHI","sequence":"additional","affiliation":[{"name":"Information and Communications Headquarters, Nagoya University, Furo-cho, Chikusa-ku, Nagoya, Aichi, 464-8601, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"ICHIRO","family":"IDE","sequence":"additional","affiliation":[{"name":"Graduate School of Information Science, Nagoya University, Furo-cho, Chikusa-ku, Nagoya, Aichi, 464-8601, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"HIROSHI","family":"MURASE","sequence":"additional","affiliation":[{"name":"Graduate School of Information Science, Nagoya University, Furo-cho, Chikusa-ku, Nagoya, Aichi, 464-8601, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"219","published-online":{"date-parts":[[2012,10,22]]},"reference":[{"key":"rf1","doi-asserted-by":"publisher","DOI":"10.1109\/93.752960"},{"key":"rf2","unstructured":"D.\u00a0Ozkan and P.\u00a0Duygulu, Image and Video Retrieval, Lecture Notes in Computer Science\u00a04071, eds. H.\u00a0Sundaram (2000)\u00a0pp. 173\u2013182."},{"key":"rf3","doi-asserted-by":"publisher","DOI":"10.1007\/11562382_42"},{"key":"rf4","unstructured":"A. F.\u00a0Smeaton, P.\u00a0Over and W.\u00a0Kraaij, Multimedia Content Analysis, Theory and Applications, Signals and Communication Technology Series, ed. A.\u00a0Divakaran (Springer-Verlag, 2009)\u00a0pp. 151\u2013174."},{"key":"rf5","unstructured":"H. J.\u00a0Nock, G.\u00a0Iyengar and C.\u00a0Neti, Image and Video Retrieval, Lecture Notes in Computer Science\u00a02728, eds. E. M.\u00a0Bakker (2003)\u00a0pp. 565\u2013570."},{"key":"rf6","doi-asserted-by":"publisher","DOI":"10.1109\/TASL.2006.878256"},{"key":"rf8","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2009.2030637"},{"key":"rf9","doi-asserted-by":"publisher","DOI":"10.1002\/vis.4340020404"},{"key":"rf10","first-page":"271","volume":"12","author":"R\u00faa E. A.","journal-title":"Pattern Analysis and Applications"},{"key":"rf13","doi-asserted-by":"publisher","DOI":"10.1023\/A:1008122917811"},{"key":"rf14","doi-asserted-by":"publisher","DOI":"10.1109\/34.927467"},{"key":"rf15","doi-asserted-by":"publisher","DOI":"10.1109\/34.982900"},{"key":"rf16","doi-asserted-by":"publisher","DOI":"10.1016\/S0031-3203(01)00231-X"},{"key":"rf17","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2003.1233896"},{"key":"rf18","first-page":"148","volume":"7","author":"Jang K. S.","journal-title":"Intl. J. of Computer Science and Network Security"},{"key":"rf21","unstructured":"G.\u00a0Potamianos, Visual and Audio-Visual Speech Processing, eds. G.\u00a0Bailly, E.\u00a0Vatikiotis-Bateson and P.\u00a0Perrier (MIT Press, 2004)\u00a0pp. 1\u201330."},{"key":"rf22","doi-asserted-by":"publisher","DOI":"10.1109\/89.848229"},{"key":"rf24","volume-title":"The Nature of Statistical Learning Theory","author":"Vapnik V. N.","year":"1999"},{"key":"rf25","first-page":"821","volume":"25","author":"Aizerman M. A.","journal-title":"Automation and Remote Control"}],"container-title":["International Journal of Semantic Computing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.worldscientific.com\/doi\/pdf\/10.1142\/S1793351X12400077","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2019,8,6]],"date-time":"2019-08-06T18:48:37Z","timestamp":1565117317000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.worldscientific.com\/doi\/abs\/10.1142\/S1793351X12400077"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2012,6]]},"references-count":19,"journal-issue":{"issue":"02","published-online":{"date-parts":[[2012,10,22]]},"published-print":{"date-parts":[[2012,6]]}},"alternative-id":["10.1142\/S1793351X12400077"],"URL":"https:\/\/doi.org\/10.1142\/s1793351x12400077","relation":{},"ISSN":["1793-351X","1793-7108"],"issn-type":[{"value":"1793-351X","type":"print"},{"value":"1793-7108","type":"electronic"}],"subject":[],"published":{"date-parts":[[2012,6]]}}}