{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,26]],"date-time":"2026-08-26T15:11:26Z","timestamp":1787757086403,"version":"build-2784847793"},"reference-count":67,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2016,8,3]],"date-time":"2016-08-03T00:00:00Z","timestamp":1470182400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100012166","name":"National Basic Research Program of China","doi-asserted-by":"crossref","award":["2012CB316304"],"award-info":[{"award-number":["2012CB316304"]}],"id":[{"id":"10.13039\/501100012166","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Importation and Development of High-Caliber Talents Project of Beijing Municipal Institutions","award":["IDHT20140224"],"award-info":[{"award-number":["IDHT20140224"]}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["61225009, 61303173, 61432019, 61572498, 61532009, 61472379 and 61572296"],"award-info":[{"award-number":["61225009, 61303173, 61432019, 61572498, 61532009, 61472379 and 61572296"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2016,8,24]]},"abstract":"<jats:p>Content-based video understanding is extremely difficult due to the semantic gap between low-level vision signals and the various semantic concepts (object, action, and scene) in videos. Though feature extraction from videos has achieved significant progress, most of the previous methods rely only on low-level features, such as the appearance and motion features. Recently, visual-feature extraction has been improved significantly with machine-learning algorithms, especially deep learning. However, there is still not enough work focusing on extracting semantic features from videos directly. The goal of this article is to adopt unlabeled videos with the help of text descriptions to learn an embedding function, which can be used to extract more effective semantic features from videos when only a few labeled samples are available for video recognition. To achieve this goal, we propose a novel embedding convolutional neural network (ECNN). We evaluate our algorithm by comparing its performance on three challenging benchmarks with several popular state-of-the-art methods. Extensive experimental results show that the proposed ECNN consistently and significantly outperforms the existing methods.<\/jats:p>","DOI":"10.1145\/2962719","type":"journal-article","created":{"date-parts":[[2016,8,4]],"date-time":"2016-08-04T13:26:34Z","timestamp":1470317194000},"page":"1-22","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":18,"title":["Semantic Feature Mining for Video Event Understanding"],"prefix":"10.1145","volume":"12","author":[{"given":"Xiaoshan","family":"Yang","sequence":"first","affiliation":[{"name":"Institute of Automation, Chinese Academy of Sciences, Beijing, P. R. China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Tianzhu","family":"Zhang","sequence":"additional","affiliation":[{"name":"Institute of Automation, Chinese Academy of Sciences, Beijing, P. R. China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Changsheng","family":"Xu","sequence":"additional","affiliation":[{"name":"Institute of Automation, Chinese Academy of Sciences, Beijing, P. R. China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2016,8,3]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.5555\/944919.944965"},{"key":"e_1_2_1_2_1","doi-asserted-by":"crossref","unstructured":"Yoshua Bengio Pascal Lamblin Dan Popovici and Hugo Larochelle. 2006. Greedy layer-wise training of deep networks. In NIPS. 153--160. Yoshua Bengio Pascal Lamblin Dan Popovici and Hugo Larochelle. 2006. Greedy layer-wise training of deep networks. In NIPS. 153--160.","DOI":"10.7551\/mitpress\/7503.003.0024"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.507"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298856"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298878"},{"key":"e_1_2_1_6_1","volume-title":"Exploiting web images for event recognition in consumer videos: A multiple source domain adaptation approach","author":"Duan Lixin"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.5555\/1953048.2021068"},{"key":"e_1_2_1_8_1","volume-title":"Proceedings of Computer Vision - 7th European Conference on Computer Vision (ECCV\u201902)","author":"Duygulu Pinar","year":"2002"},{"key":"e_1_2_1_9_1","volume-title":"Proceedings of Computer Vision - 11th European Conference on Computer Vision (ECCV\u201910)","author":"Farhadi Ali","year":"2010"},{"key":"e_1_2_1_10_1","volume-title":"Advances in Neural Information Processing Systems 26: 27th Annual Conference on Neural Information Processing Systems 2013. Proceedings of a meeting held December 5--8","author":"Frome Andrea","year":"2013"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-013-0658-4"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2013.337"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1145\/2647868.2654913"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1162\/neco.2006.18.7.1527"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2014.242"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/1631272.1631371"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2012.59"},{"key":"e_1_2_1_18_1","volume-title":"Caffe: Convolutional architecture for fast feature embedding. arXiv preprint arXiv:1408.5093.","author":"Jia Yangqing","year":"2014"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/2393347.2393412"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/1991996.1992025"},{"key":"e_1_2_1_21_1","doi-asserted-by":"crossref","unstructured":"Nal Kalchbrenner Edward Grefenstette and Phil Blunsom. 2014. A convolutional neural network for modelling sentences. In ACLs. 655--665. Nal Kalchbrenner Edward Grefenstette and Phil Blunsom. 2014. A convolutional neural network for modelling sentences. In ACLs. 655--665.","DOI":"10.3115\/v1\/P14-1062"},{"key":"e_1_2_1_22_1","volume-title":"Advances in Neural Information Processing Systems 27: Annual Conference on Neural Information Processing Systems, December 8--13","author":"Karpathy Andrej","year":"2014"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298932"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2014.223"},{"key":"e_1_2_1_25_1","volume-title":"Zemel","author":"Kiros Ryan","year":"2014"},{"key":"e_1_2_1_26_1","volume-title":"Proceedings of the 27th AAAI Conference on Artificial Intelligence, July 14--18","author":"Krishnamoorthy Niveda","year":"2013"},{"key":"e_1_2_1_27_1","volume-title":"Hinton","author":"Krizhevsky Alex","year":"2012"},{"key":"e_1_2_1_28_1","doi-asserted-by":"crossref","first-page":"351","DOI":"10.1162\/tacl_a_00188","article-title":"TREETALK: Composition and compression of trees for image descriptions","volume":"2","author":"Kuznetsova Polina","year":"2014","journal-title":"TACL"},{"key":"e_1_2_1_29_1","volume-title":"Proceedings of the 32nd International Conference on Machine Learning (ICML\u201915)","author":"Lebret R\u00e9mi","year":"2015"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1109\/5.726791"},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCVW.2015.44"},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1145\/1459359.1459574"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1145\/2502081.2502119"},{"key":"e_1_2_1_34_1","volume-title":"Yuille","author":"Mao Junhua","year":"2014"},{"key":"e_1_2_1_35_1","unstructured":"Tomas Mikolov Ilya Sutskever Kai Chen Gregory S. Corrado and Jeffrey Dean. 2013. Distributed representations of words and phrases and their compositionality. In NIPS. 3111--3119. Tomas Mikolov Ilya Sutskever Kai Chen Gregory S. Corrado and Jeffrey Dean. 2013. Distributed representations of words and phrases and their compositionality. In NIPS. 3111--3119."},{"key":"e_1_2_1_36_1","volume-title":"TRECVID 2013 -- An overview of the goals, tasks, data, evaluation mechanisms and metrics. In Proceedings of TRECVID 2013. NIST.","author":"Over Paul","year":"2013"},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1145\/2733373.2806234"},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1145\/2659521"},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1109\/tmm.2015.2510329"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2013.117"},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298940"},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2013.61"},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCVW.2015.47"},{"key":"e_1_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-015-0816-y"},{"key":"e_1_2_1_45_1","volume-title":"Proceedings of the International Conference on AISTATS. 448--455","author":"Salakhutdinov Ruslan"},{"key":"e_1_2_1_46_1","unstructured":"Karen Simonyan and Andrew Zisserman. 2014. Very deep convolutional networks for large-scale image recognition. CoRR abs\/1409.1556. Karen Simonyan and Andrew Zisserman. 2014. Very deep convolutional networks for large-scale image recognition. CoRR abs\/1409.1556."},{"key":"e_1_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00177"},{"key":"e_1_2_1_48_1","volume-title":"Proceedings of the 8th International Conference on Language Resources and Evaluation. 2573--2577","author":"Strassel Stephanie","year":"2012"},{"key":"e_1_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2013.446"},{"key":"e_1_2_1_50_1","doi-asserted-by":"crossref","unstructured":"Christian Szegedy Wei Liu Yangqing Jia Pierre Sermanet Scott Reed Dragomir Anguelov Dumitru Erhan Vincent Vanhoucke and Andrew Rabinovich. 2014. Going deeper with convolutions. CoRR abs\/1409.4842. Christian Szegedy Wei Liu Yangqing Jia Pierre Sermanet Scott Reed Dragomir Anguelov Dumitru Erhan Vincent Vanhoucke and Andrew Rabinovich. 2014. Going deeper with convolutions. CoRR abs\/1409.4842.","DOI":"10.1109\/CVPR.2015.7298594"},{"key":"e_1_2_1_51_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298792"},{"key":"e_1_2_1_52_1","volume-title":"Proceedings of the 25th International Conference on Computational Linguistics (COLING\u201914)","author":"Thomason Jesse","year":"2014"},{"key":"e_1_2_1_53_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.515"},{"key":"e_1_2_1_54_1","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/N15-1173"},{"key":"e_1_2_1_55_1","doi-asserted-by":"publisher","DOI":"10.1145\/1390156.1390294"},{"key":"e_1_2_1_56_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298935"},{"key":"e_1_2_1_57_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2011.5995407"},{"key":"e_1_2_1_58_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2013.441"},{"key":"e_1_2_1_59_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2014.180"},{"key":"e_1_2_1_60_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW.2015.7301333"},{"key":"e_1_2_1_61_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCVW.2015.46"},{"key":"e_1_2_1_62_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2014.2375793"},{"key":"e_1_2_1_63_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2015.2393635"},{"key":"e_1_2_1_64_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2013.456"},{"key":"e_1_2_1_65_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.512"},{"key":"e_1_2_1_66_1","volume-title":"Zeiler and Rob Fergus","author":"Matthew","year":"2014"},{"key":"e_1_2_1_67_1","doi-asserted-by":"publisher","DOI":"10.1145\/2602633"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2962719","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/2962719","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T04:54:03Z","timestamp":1750222443000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2962719"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2016,8,3]]},"references-count":67,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2016,8,24]]}},"alternative-id":["10.1145\/2962719"],"URL":"https:\/\/doi.org\/10.1145\/2962719","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2016,8,3]]},"assertion":[{"value":"2015-11-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2016-05-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2016-08-03","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}