{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,16]],"date-time":"2026-05-16T00:02:54Z","timestamp":1778889774898,"version":"3.51.4"},"publisher-location":"New York, NY, USA","reference-count":42,"publisher":"ACM","license":[{"start":{"date-parts":[[2016,10,1]],"date-time":"2016-10-01T00:00:00Z","timestamp":1475280000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2016,10]]},"DOI":"10.1145\/2964284.2964288","type":"proceedings-article","created":{"date-parts":[[2016,9,29]],"date-time":"2016-09-29T19:17:32Z","timestamp":1475176652000},"page":"1008-1017","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":85,"title":["Robust Visual-Textual Sentiment Analysis"],"prefix":"10.1145","author":[{"given":"Quanzeng","family":"You","sequence":"first","affiliation":[{"name":"University of Rochester, Rochester, NY, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Liangliang","family":"Cao","sequence":"additional","affiliation":[{"name":"Yahoo Labs, New York, NY, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hailin","family":"Jin","sequence":"additional","affiliation":[{"name":"Adobe, San Jose, CA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jiebo","family":"Luo","sequence":"additional","affiliation":[{"name":"University of Rochester, Rochester, NY, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2016,10]]},"reference":[{"key":"e_1_3_2_1_1_1","volume-title":"Neural machine translation by jointly learning to align and translate. arXiv preprint arXiv:1409.0473","author":"Bahdanau D.","year":"2014","unstructured":"D. Bahdanau , K. Cho , and Y. Bengio . Neural machine translation by jointly learning to align and translate. arXiv preprint arXiv:1409.0473 , 2014 . D. Bahdanau, K. Cho, and Y. Bengio. Neural machine translation by jointly learning to align and translate. arXiv preprint arXiv:1409.0473, 2014."},{"key":"e_1_3_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1145\/2502081.2502268"},{"key":"e_1_3_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/2502081.2502282"},{"key":"e_1_3_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1007\/s00530-014-0407-8"},{"key":"e_1_3_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298856"},{"key":"e_1_3_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298754"},{"key":"e_1_3_2_1_7_1","first-page":"2121","volume-title":"Advances in Neural Information Processing Systems (NIPS)","author":"Frome A.","year":"2013","unstructured":"A. Frome , G. S. Corrado , J. Shlens , S. Bengio , J. Dean , M. Ranzato , and T. Mikolov . Devise: A deep visual-semantic embedding model . In Advances in Neural Information Processing Systems (NIPS) , pages 2121 -- 2129 , 2013 . A. Frome, G. S. Corrado, J. Shlens, S. Bengio, J. Dean, M. Ranzato, and T. Mikolov. Devise: A deep visual-semantic embedding model. In Advances in Neural Information Processing Systems (NIPS), pages 2121--2129, 2013."},{"key":"e_1_3_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1162\/153244303768966139"},{"key":"e_1_3_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/ASRU.2013.6707742"},{"key":"e_1_3_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2013.6638947"},{"key":"e_1_3_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1162\/neco.1997.9.8.1735"},{"key":"e_1_3_2_1_13_1","volume-title":"ICWSM","author":"Hutto C. J.","year":"2014","unstructured":"C. J. Hutto and E. Gilbert . VADER: A parsimonious rule-based model for sentiment analysis of social media text . In ICWSM , 2014 . C. J. Hutto and E. Gilbert. VADER: A parsimonious rule-based model for sentiment analysis of social media text. In ICWSM, 2014."},{"key":"e_1_3_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298932"},{"key":"e_1_3_2_1_15_1","volume-title":"Unifying visual-semantic embeddings with multimodal neural language models. CoRR, abs\/1411.2539","author":"Kiros R.","year":"2014","unstructured":"R. Kiros , R. Salakhutdinov , and R. S. Zemel . Unifying visual-semantic embeddings with multimodal neural language models. CoRR, abs\/1411.2539 , 2014 . R. Kiros, R. Salakhutdinov, and R. S. Zemel. Unifying visual-semantic embeddings with multimodal neural language models. CoRR, abs\/1411.2539, 2014."},{"key":"e_1_3_2_1_16_1","first-page":"1106","volume-title":"Advances in Neural Information Processing Systems (NIPS)","author":"Krizhevsky A.","year":"2012","unstructured":"A. Krizhevsky , I. Sutskever , and G. E. Hinton . Imagenet classification with deep convolutional neural networks . In Advances in Neural Information Processing Systems (NIPS) , pages 1106 -- 1114 , 2012 . A. Krizhevsky, I. Sutskever, and G. E. Hinton. Imagenet classification with deep convolutional neural networks. In Advances in Neural Information Processing Systems (NIPS), pages 1106--1114, 2012."},{"key":"e_1_3_2_1_17_1","first-page":"1188","volume-title":"Proceedings of the 28th International Conference on Machine Learning (ICML)","author":"Le Q. V.","year":"2014","unstructured":"Q. V. Le and T. Mikolov . Distributed representations of sentences and documents . In Proceedings of the 28th International Conference on Machine Learning (ICML) , pages 1188 -- 1196 , 2014 . Q. V. Le and T. Mikolov. Distributed representations of sentences and documents. In Proceedings of the 28th International Conference on Machine Learning (ICML), pages 1188--1196, 2014."},{"key":"e_1_3_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/2647868.2654927"},{"key":"e_1_3_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.301"},{"key":"e_1_3_2_1_20_1","volume-title":"ICLR","author":"Mao J.","year":"2015","unstructured":"J. Mao , W. Xu , Y. Yang , J. Wang , Z. Huang , and A. Yuille . Deep captioning with multimodal recurrent neural networks (m-rnn) . ICLR , 2015 . J. Mao, W. Xu, Y. Yang, J. Wang, Z. Huang, and A. Yuille. Deep captioning with multimodal recurrent neural networks (m-rnn). ICLR, 2015."},{"key":"e_1_3_2_1_21_1","first-page":"3111","volume-title":"Advances in Neural Information Processing Systems 26 (NIPS)","author":"Mikolov T.","year":"2013","unstructured":"T. Mikolov , I. Sutskever , K. Chen , G. S. Corrado , and J. Dean . Distributed representations of words and phrases and their compositionality . In Advances in Neural Information Processing Systems 26 (NIPS) , pages 3111 -- 3119 , 2013 . T. Mikolov, I. Sutskever, K. Chen, G. S. Corrado, and J. Dean. Distributed representations of words and phrases and their compositionality. In Advances in Neural Information Processing Systems 26 (NIPS), pages 3111--3119, 2013."},{"key":"e_1_3_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1561\/0600000033"},{"key":"e_1_3_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1561\/1500000011"},{"key":"e_1_3_2_1_24_1","first-page":"1310","volume-title":"Proceedings of the 28th International Conference on Machine Learning (ICML)","author":"Pascanu R.","year":"2013","unstructured":"R. Pascanu , T. Mikolov , and Y. Bengio . On the difficulty of training recurrent neural networks . In Proceedings of the 28th International Conference on Machine Learning (ICML) , pages 1310 -- 1318 , 2013 . R. Pascanu, T. Mikolov, and Y. Bengio. On the difficulty of training recurrent neural networks. In Proceedings of the 28th International Conference on Machine Learning (ICML), pages 1310--1318, 2013."},{"key":"e_1_3_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/D14-1162"},{"key":"e_1_3_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1145\/1873951.1874060"},{"key":"e_1_3_2_1_27_1","volume-title":"Very deep convolutional networks for large-scale image recognition. CoRR, abs\/1409.1556","author":"Simonyan K.","year":"2014","unstructured":"K. Simonyan and A. Zisserman . Very deep convolutional networks for large-scale image recognition. CoRR, abs\/1409.1556 , 2014 . K. Simonyan and A. Zisserman. Very deep convolutional networks for large-scale image recognition. CoRR, abs\/1409.1556, 2014."},{"key":"e_1_3_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00177"},{"key":"e_1_3_2_1_29_1","first-page":"129","volume-title":"Proceedings of the 28th International Conference on Machine Learning (ICML)","author":"Socher R.","year":"2011","unstructured":"R. Socher , C. C. Lin , A. Y. Ng , and C. D. Manning . Parsing natural scenes and natural language with recursive neural networks . In Proceedings of the 28th International Conference on Machine Learning (ICML) , pages 129 -- 136 , 2011 . R. Socher, C. C. Lin, A. Y. Ng, and C. D. Manning. Parsing natural scenes and natural language with recursive neural networks. In Proceedings of the 28th International Conference on Machine Learning (ICML), pages 129--136, 2011."},{"key":"e_1_3_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.5555\/2627435.2697059"},{"key":"e_1_3_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298594"},{"key":"e_1_3_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/P15-1150"},{"key":"e_1_3_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298935"},{"key":"e_1_3_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1145\/2632856.2632912"},{"key":"e_1_3_2_1_36_1","first-page":"473","volume-title":"ICWSM","author":"Wang Y.","year":"2015","unstructured":"Y. Wang , Y. Hu , S. Kambhampati , and B. Li . Inferring sentiment from web images with joint inference on visual and social cues: A regulated matrix factorization approach . In ICWSM , pages 473 -- 482 , 2015 . Y. Wang, Y. Hu, S. Kambhampati, and B. Li. Inferring sentiment from web images with joint inference on visual and social cues: A regulated matrix factorization approach. In ICWSM, pages 473--482, 2015."},{"key":"e_1_3_2_1_37_1","first-page":"2378","volume-title":"Proceedings of the Twenty-Fourth International Joint Conference on Artificial Intelligence (IJCAI)","author":"Wang Y.","year":"2015","unstructured":"Y. Wang , S. Wang , J. Tang , H. Liu , and B. Li . Unsupervised sentiment analysis for social media images . In Proceedings of the Twenty-Fourth International Joint Conference on Artificial Intelligence (IJCAI) , pages 2378 -- 2379 , 2015 . Y. Wang, S. Wang, J. Tang, H. Liu, and B. Li. Unsupervised sentiment analysis for social media images. In Proceedings of the Twenty-Fourth International Joint Conference on Artificial Intelligence (IJCAI), pages 2378--2379, 2015."},{"key":"e_1_3_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.5555\/2283696.2283856"},{"key":"e_1_3_2_1_39_1","first-page":"2048","volume-title":"ICML","author":"Xu K.","year":"2015","unstructured":"K. Xu , J. Ba , R. Kiros , K. Cho , A. C. Courville , R. Salakhutdinov , R. S. Zemel , and Y. Bengio . Show, attend and tell: Neural image caption generation with visual attention . In ICML , pages 2048 -- 2057 , 2015 . K. Xu, J. Ba, R. Kiros, K. Cho, A. C. Courville, R. Salakhutdinov, R. S. Zemel, and Y. Bengio. Show, attend and tell: Neural image caption generation with visual attention. In ICML, pages 2048--2057, 2015."},{"key":"e_1_3_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2011.5995741"},{"key":"e_1_3_2_1_41_1","first-page":"381","volume-title":"AAAI","author":"You Q.","year":"2015","unstructured":"Q. You , J. Luo , H. Jin , and J. Yang . Robust image sentiment analysis using progressively trained and domain transferred deep networks . In AAAI , pages 381 -- 388 , 2015 . Q. You, J. Luo, H. Jin, and J. Yang. Robust image sentiment analysis using progressively trained and domain transferred deep networks. In AAAI, pages 381--388, 2015."},{"key":"e_1_3_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1145\/2835776.2835779"},{"key":"e_1_3_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1145\/2502069.2502079"},{"key":"e_1_3_2_1_44_1","first-page":"1604","volume-title":"ICML","author":"Zhu X.","year":"2015","unstructured":"X. Zhu , P. Sobhani , and H. Guo . Long short-term memory over recursive structures . In ICML , pages 1604 -- 1612 , 2015 . X. Zhu, P. Sobhani, and H. Guo. Long short-term memory over recursive structures. In ICML, pages 1604--1612, 2015."}],"event":{"name":"MM '16: ACM Multimedia Conference","location":"Amsterdam The Netherlands","acronym":"MM '16","sponsor":["SIGMM ACM Special Interest Group on Multimedia"]},"container-title":["Proceedings of the 24th ACM international conference on Multimedia"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2964284.2964288","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/2964284.2964288","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T19:04:32Z","timestamp":1750273472000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2964284.2964288"}},"subtitle":["When Attention meets Tree-structured Recursive Neural Networks"],"short-title":[],"issued":{"date-parts":[[2016,10]]},"references-count":42,"alternative-id":["10.1145\/2964284.2964288","10.1145\/2964284"],"URL":"https:\/\/doi.org\/10.1145\/2964284.2964288","relation":{},"subject":[],"published":{"date-parts":[[2016,10]]},"assertion":[{"value":"2016-10-01","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}