{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T04:27:32Z","timestamp":1750220852629,"version":"3.41.0"},"publisher-location":"New York, NY, USA","reference-count":14,"publisher":"ACM","license":[{"start":{"date-parts":[[2019,10,29]],"date-time":"2019-10-29T00:00:00Z","timestamp":1572307200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2019,10,29]]},"DOI":"10.1145\/3323503.3345029","type":"proceedings-article","created":{"date-parts":[[2019,10,10]],"date-time":"2019-10-10T13:04:27Z","timestamp":1570712667000},"page":"21-23","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":6,"title":["Deep learning methods for video understanding"],"prefix":"10.1145","author":[{"given":"Gabriel N. P. dos","family":"Santos","sequence":"first","affiliation":[{"name":"PUC-Rio"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Pedro V. A.","family":"de Freitas","sequence":"additional","affiliation":[{"name":"PUC-Rio"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Antonio Jos\u00e9 G.","family":"Busson","sequence":"additional","affiliation":[{"name":"PUC-Rio"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"\u00c1lan L. V.","family":"Guedes","sequence":"additional","affiliation":[{"name":"PUC-Rio"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ruy","family":"Milidi\u00fa","sequence":"additional","affiliation":[{"name":"PUC-Rio"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"S\u00e9rgio","family":"Colcher","sequence":"additional","affiliation":[{"name":"PUC-Rio"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2019,10,29]]},"reference":[{"key":"e_1_3_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00742"},{"key":"e_1_3_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_2_1_3_1","volume-title":"CNN Architectures for Large-Scale Audio Classification. In International Conference on Acoustics, Speech and Signal Processing (ICASSP). https:\/\/arxiv.org\/abs\/1609","author":"Hershey Shawn","year":"2017","unstructured":"Shawn Hershey , Sourish Chaudhuri , Daniel P. W. Ellis , Jort F. Gemmeke , Aren Jansen , Channing Moore , Manoj Plakal , Devin Platt , Rif A. Saurous , Bryan Seybold , Malcolm Slaney , Ron Weiss , and Kevin Wilson . 2017 . CNN Architectures for Large-Scale Audio Classification. In International Conference on Acoustics, Speech and Signal Processing (ICASSP). https:\/\/arxiv.org\/abs\/1609 .09430 Shawn Hershey, Sourish Chaudhuri, Daniel P. W. Ellis, Jort F. Gemmeke, Aren Jansen, Channing Moore, Manoj Plakal, Devin Platt, Rif A. Saurous, Bryan Seybold, Malcolm Slaney, Ron Weiss, and Kevin Wilson. 2017. CNN Architectures for Large-Scale Audio Classification. In International Conference on Acoustics, Speech and Signal Processing (ICASSP). https:\/\/arxiv.org\/abs\/1609.09430"},{"key":"e_1_3_2_1_4_1","unstructured":"Alex Krizhevsky Ilya Sutskever and Geoffrey E Hinton. 2012. Imagenet classification with deep convolutional neural networks. In Advances in neural information processing systems. 1097--1105.  Alex Krizhevsky Ilya Sutskever and Geoffrey E Hinton. 2012. Imagenet classification with deep convolutional neural networks. In Advances in neural information processing systems. 1097--1105."},{"key":"e_1_3_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46484-8_29"},{"key":"e_1_3_2_1_6_1","volume-title":"Wavenet: A generative model for raw audio. arXiv preprint arXiv:1609.03499","author":"van den Oord Aaron","year":"2016","unstructured":"Aaron van den Oord , Sander Dieleman , Heiga Zen , Karen Simonyan , Oriol Vinyals , Alex Graves , Nal Kalchbrenner , Andrew Senior , and Koray Kavukcuoglu . 2016 . Wavenet: A generative model for raw audio. arXiv preprint arXiv:1609.03499 (2016). Aaron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu. 2016. Wavenet: A generative model for raw audio. arXiv preprint arXiv:1609.03499 (2016)."},{"key":"e_1_3_2_1_7_1","volume-title":"Yolov3: An incremental improvement. arXiv preprint arXiv:1804.02767","author":"Redmon Joseph","year":"2018","unstructured":"Joseph Redmon and Ali Farhadi . 2018. Yolov3: An incremental improvement. arXiv preprint arXiv:1804.02767 ( 2018 ). Joseph Redmon and Ali Farhadi. 2018. Yolov3: An incremental improvement. arXiv preprint arXiv:1804.02767 (2018)."},{"key":"e_1_3_2_1_8_1","unstructured":"Shaoqing Ren Kaiming He Ross Girshick and Jian Sun. 2015. Faster r-cnn: Towards real-time object detection with region proposal networks. In Advances in neural information processing systems. 91--99.  Shaoqing Ren Kaiming He Ross Girshick and Jian Sun. 2015. Faster r-cnn: Towards real-time object detection with region proposal networks. In Advances in neural information processing systems. 91--99."},{"key":"e_1_3_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-015-0816-y"},{"key":"e_1_3_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298682"},{"key":"e_1_3_2_1_11_1","first-page":"12","article-title":"Inception-v4, inception-resnet and the impact of residual connections on learning","volume":"4","author":"Szegedy Christian","year":"2017","unstructured":"Christian Szegedy , Sergey Ioffe , Vincent Vanhoucke , and Alexander A Alemi . 2017 . Inception-v4, inception-resnet and the impact of residual connections on learning .. In AAAI , Vol. 4. 12 . Christian Szegedy, Sergey Ioffe, Vincent Vanhoucke, and Alexander A Alemi. 2017. Inception-v4, inception-resnet and the impact of residual connections on learning.. In AAAI, Vol. 4. 12.","journal-title":"AAAI"},{"key":"e_1_3_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298594"},{"key":"e_1_3_2_1_13_1","volume-title":"European Conference on Computer Vision. Springer, 219--228","author":"Tang Yongyi","year":"2018","unstructured":"Yongyi Tang , Xing Zhang , Jingwen Wang , Shaoxiang Chen , Lin Ma , and Yu-Gang Jiang . 2018 . Non-local netVLAD encoding for video classification . In European Conference on Computer Vision. Springer, 219--228 . Yongyi Tang, Xing Zhang, Jingwen Wang, Shaoxiang Chen, Lin Ma, and Yu-Gang Jiang. 2018. Non-local netVLAD encoding for video classification. In European Conference on Computer Vision. Springer, 219--228."},{"key":"e_1_3_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01231-1_29"}],"event":{"name":"WebMedia '19: Brazilian Symposium on Multimedia and the Web","acronym":"WebMedia '19","location":"Rio de Janeiro Brazil"},"container-title":["Proceedings of the 25th Brazillian Symposium on Multimedia and the Web"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3323503.3345029","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3323503.3345029","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T23:23:17Z","timestamp":1750202597000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3323503.3345029"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019,10,29]]},"references-count":14,"alternative-id":["10.1145\/3323503.3345029","10.1145\/3323503"],"URL":"https:\/\/doi.org\/10.1145\/3323503.3345029","relation":{},"subject":[],"published":{"date-parts":[[2019,10,29]]},"assertion":[{"value":"2019-10-29","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}