{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T04:14:02Z","timestamp":1750306442496,"version":"3.41.0"},"reference-count":31,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2015,11,2]],"date-time":"2015-11-02T00:00:00Z","timestamp":1446422400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"Research Grants Council of the Hong Kong Special Administrative Region, China"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2016,3,3]]},"abstract":"<jats:p>This article considers multimedia question answering beyond factoid and how-to questions. We are interested in searching videos for answering opinion-oriented questions that are controversial and hotly debated. Examples of questions include \u201cShould Edward Snowden be pardoned?\u201d and \u201cObamacare\u2014unconstitutional or not?\u201d. These questions often invoke emotional response, either positively or negatively, hence are likely to be better answered by videos than texts, due to the vivid display of emotional signals visible through facial expression and speaking tone. Nevertheless, a potential answer of duration 60s may be embedded in a video of 10min, resulting in degraded user experience compared to reading the answer in text only. Furthermore, a text-based opinion question may be short and vague, while the video answers could be verbal, less structured grammatically, and noisy because of errors in speech transcription. Direct matching of words or syntactic analysis of sentence structure, such as adopted by factoid and how-to question-answering, is unlikely to find video answers. The first problem, the answer localization, is addressed by audiovisual analysis of the emotional signals in videos for locating video segments likely expressing opinions. The second problem, questions and answers matching, is tackled by a deep architecture that nonlinearly matches text words in questions and speeches in videos. Experiments are conducted on eight controversial topics based on questions crawled from Yahoo! Answers and Internet videos from YouTube.<\/jats:p>","DOI":"10.1145\/2818711","type":"journal-article","created":{"date-parts":[[2015,11,2]],"date-time":"2015-11-02T17:09:35Z","timestamp":1446484175000},"page":"1-19","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":4,"title":["Opinion Question Answering by Sentiment Clip Localization"],"prefix":"10.1145","volume":"12","author":[{"given":"Lei","family":"Pang","sequence":"first","affiliation":[{"name":"City University of Hong Kong, Hong Kong"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Chong-Wah","family":"Ngo","sequence":"additional","affiliation":[{"name":"City University of Hong Kong, Hong Kong"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2015,11,2]]},"reference":[{"volume-title":"Proceedings of the 7th International Conference on Language Resources and Evaluation.","year":"2010","author":"Baccianella Stefano","key":"e_1_2_1_1_1"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1145\/2502081.2502282"},{"volume-title":"Proceedings of the 10th Text REtrieval Conference (TREC).","year":"2001","author":"Brill Eric","key":"e_1_2_1_3_1"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1145\/996350.996398"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/2647868.2654935"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/1631058.1631069"},{"volume-title":"Proceedings of the British Machine Vision Conference.","author":"Everingham M.","key":"e_1_2_1_7_1"},{"volume-title":"Proceedings of TREC","year":"2002","author":"Hermjakob Ulf","key":"e_1_2_1_8_1"},{"key":"e_1_2_1_9_1","unstructured":"Gary Kacmarcik. 2005. Multi-modal question-answering: Questions without keyboards. In Asia Federation of Natural Language Processing.  Gary Kacmarcik. 2005. Multi-modal question-answering: Questions without keyboards. In Asia Federation of Natural Language Processing."},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/2461466.2461484"},{"key":"e_1_2_1_11_1","doi-asserted-by":"crossref","unstructured":"Y. Lecun L. Bottou G. B. Orr and K. R. M\u00fcller. 1998. Efficient backprop. In Neural Networks: Tricks of the Trade.   Y. Lecun L. Bottou G. B. Orr and K. R. M\u00fcller. 1998. Efficient backprop. In Neural Networks: Tricks of the Trade.","DOI":"10.1007\/3-540-49430-8_2"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1002\/asi.v60:3"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/MMUL.2010.47"},{"key":"e_1_2_1_14_1","unstructured":"Zhengdong Lu and Hang Li. 2013. A deep architecture for matching short texts. In Advances in Neural Information Processing Systems.  Zhengdong Lu and Hang Li. 2013. A deep architecture for matching short texts. In Advances in Neural Information Processing Systems."},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/1873951.1873965"},{"key":"e_1_2_1_16_1","doi-asserted-by":"crossref","unstructured":"Christopher D. Manning Prabhakar Raghavan and Hinrich Sch\u00fctze. 2008. Introduction to Information Retrieval. Cambridge University Press New York NY.   Christopher D. Manning Prabhakar Raghavan and Hinrich Sch\u00fctze. 2008. Introduction to Information Retrieval. Cambridge University Press New York NY.","DOI":"10.1017\/CBO9780511809071"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/1273496.1273576"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/2009916.2010010"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/502585.502610"},{"key":"e_1_2_1_20_1","unstructured":"S. E. Robertson S. Walker S. Jones M. M. Hancock-Beaulieu and M. Gatford. 1996. Okapi at TREC-3. 109--126.  S. E. Robertson S. Walker S. Jones M. M. Hancock-Beaulieu and M. Gatford. 1996. Okapi at TREC-3. 109--126."},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1007\/11752790_2"},{"key":"e_1_2_1_22_1","doi-asserted-by":"crossref","unstructured":"Mickael Rouvier Gregor Dupuy Paul Gay Elie Khoury Teva Merlin and Sylvain Meignier. 2013. An open-source state-of-the-art toolbox for broadcast news diarization. In INTERSPEECH.  Mickael Rouvier Gregor Dupuy Paul Gay Elie Khoury Teva Merlin and Sylvain Meignier. 2013. An open-source state-of-the-art toolbox for broadcast news diarization. In INTERSPEECH.","DOI":"10.21437\/Interspeech.2013-383"},{"volume-title":"IEEE Conference on Computer Vision and Pattern Recognition. 593--600","year":"1994","author":"Shi Jianbo","key":"e_1_2_1_23_1"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1023\/B:VISI.0000013087.49260.fb"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/1571941.1571975"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1145\/2433396.2433481"},{"volume-title":"Proceedings of the IEEE 6th International Symposium on Multimedia Software Engineering.","year":"2004","author":"Wu Yu-Chyeh","key":"e_1_2_1_27_1"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2008.2002831"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/860435.860444"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1145\/1459359.1459412"},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1145\/2393347.2393432"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2818711","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/2818711","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T05:42:49Z","timestamp":1750225369000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2818711"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2015,11,2]]},"references-count":31,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2016,3,3]]}},"alternative-id":["10.1145\/2818711"],"URL":"https:\/\/doi.org\/10.1145\/2818711","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"type":"print","value":"1551-6857"},{"type":"electronic","value":"1551-6865"}],"subject":[],"published":{"date-parts":[[2015,11,2]]},"assertion":[{"value":"2014-08-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2015-06-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2015-11-02","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}