{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T04:34:45Z","timestamp":1750221285144,"version":"3.41.0"},"publisher-location":"New York, New York, USA","reference-count":20,"publisher":"ACM Press","license":[{"start":{"date-parts":[[2018,1,1]],"date-time":"2018-01-01T00:00:00Z","timestamp":1514764800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Brazilian Coordination for the Improvement of Higher Education Personnel (CAPES)"},{"name":"Project LIMBO (BMVI)","award":["project no. 19F2029C"],"award-info":[{"award-number":["project no. 19F2029C"]}]},{"name":"Project OPAL","award":["project no. 19F20284"],"award-info":[{"award-number":["project no. 19F20284"]}]},{"name":"SOLIDE","award":["project no. 13N14456"],"award-info":[{"award-number":["project no. 13N14456"]}]},{"name":"German Federal Ministry of Education and Research (BMBF)"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2018]]},"DOI":"10.1145\/3184558.3192318","type":"proceedings-article","created":{"date-parts":[[2018,4,18]],"date-time":"2018-04-18T18:04:25Z","timestamp":1524074665000},"page":"1937-1939","source":"Crossref","is-referenced-by-count":1,"title":["Question Answering Mediated by Visual Clues and Knowledge Graphs"],"prefix":"10.1145","author":[{"given":"Fabricio F.","family":"de Faria","sequence":"first","affiliation":[{"name":"Federal University of Rio de Janeiro, Rio de Janeiro, Brazil"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ricardo","family":"Usbeck","sequence":"additional","affiliation":[{"name":"Paderborn University, Paderborn, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Alessio","family":"Sarullo","sequence":"additional","affiliation":[{"name":"University of Manchester, Manchester, United Kingdom"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Tingting","family":"Mu","sequence":"additional","affiliation":[{"name":"University of Manchester, Manchester, United Kingdom"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Andre","family":"Freitas","sequence":"additional","affiliation":[{"name":"University of Manchester, Manchester, United Kingdom"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","reference":[{"key":"key-10.1145\/3184558.3192318-1","doi-asserted-by":"crossref","unstructured":"Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh. 2015. VQA: Visual Question Answering. In 2015 IEEE International Conference on Computer Vision, ICCV 2015, Santiago, Chile, December 7--13, 2015. IEEE Computer Society, 2425--2433. https:\/\/doi.org\/ 10.1109\/ICCV.2015.279","DOI":"10.1109\/ICCV.2015.279"},{"key":"key-10.1145\/3184558.3192318-2","unstructured":"I Cisco. 2012. Cisco visual networking index: Forecast and methodology, 2011-- 2016. CISCO White paper (2012), 2011--2016."},{"key":"key-10.1145\/3184558.3192318-3","unstructured":"Ricardos Usbeck Alessio Sarullo Tingting Mu Fabr&#237;cio Firmino, Andr&#233; Freitas. 2018. Protocol for Question Answering Mediated by Visual Clues and Knowledge Graphs. https:\/\/visual-question-answering-challenge.github.io\/protocol. (2018). [Online; accessed 04-March-2018]."},{"key":"key-10.1145\/3184558.3192318-4","unstructured":"Ricardos Usbeck Alessio Sarullo Tingting Mu Fabr&#237;cio Firmino, Andr&#233; Freitas. 2018. Question Answering Mediated by Visual Clues and Knowledge Graphs - WWW 2018 LYON, FRANCE. https:\/\/visual-question-answering-challenge. github.io. (2018). [Online; accessed 04-March-2018]."},{"key":"key-10.1145\/3184558.3192318-5","unstructured":"Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. 2016. Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering. CoRR abs\/1612.00837 (2016). arXiv:1612.00837 http:\/\/arxiv.org\/abs\/1612.00837"},{"key":"key-10.1145\/3184558.3192318-6","doi-asserted-by":"crossref","unstructured":"Klaus Greff, Rupesh K Srivastava, Jan Koutn&#237;k, Bas R Steunebrink, and J&#252;rgen Schmidhuber. 2017. LSTM: A search space odyssey. IEEE transactions on neural networks and learning systems 28, 10 (2017), 2222--2232.","DOI":"10.1109\/TNNLS.2016.2582924"},{"key":"key-10.1145\/3184558.3192318-7","doi-asserted-by":"crossref","unstructured":"Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition. 770--778.","DOI":"10.1109\/CVPR.2016.90"},{"key":"key-10.1145\/3184558.3192318-8","unstructured":"Justin Johnson, Bharath Hariharan, Laurens van der Maaten, Li Fei-Fei, C. Lawrence Zitnick, and Ross B. Girshick. 2016. CLEVR: A Diagnostic Dataset for Compositional Language and Elementary Visual Reasoning. CoRR abs\/1612.06890 (2016). arXiv:1612.06890 http:\/\/arxiv.org\/abs\/1612.06890"},{"key":"key-10.1145\/3184558.3192318-9","doi-asserted-by":"crossref","unstructured":"Justin Johnson, Ranjay Krishna, Michael Stark, Li-Jia Li, David Shamma, Michael Bernstein, and Li Fei-Fei. 2015. Image retrieval using scene graphs. In Proceedings of the IEEE conference on computer vision and pattern recognition. 3668--3678.","DOI":"10.1109\/CVPR.2015.7298990"},{"key":"key-10.1145\/3184558.3192318-10","unstructured":"Kushal Kafle and Christopher Kanan. 2016. Visual Question Answering: Datasets, Algorithms, and Future Challenges. CoRR abs\/1610.01465 (2016). arXiv:1610.01465 http:\/\/arxiv.org\/abs\/1610.01465"},{"key":"key-10.1145\/3184558.3192318-11","doi-asserted-by":"crossref","unstructured":"Jonathan Krause, Justin Johnson, Ranjay Krishna, and Li Fei-Fei. 2017. A hierarchical approach for generating descriptive image paragraphs. In 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, 3337--3345.","DOI":"10.1109\/CVPR.2017.356"},{"key":"key-10.1145\/3184558.3192318-12","unstructured":"Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, et al. 2017. Visual genome: Connecting language and vision using crowdsourced dense image annotations. International Journal of Computer Vision 123, 1 (2017), 32--73."},{"key":"key-10.1145\/3184558.3192318-13","unstructured":"Wei Liu, Dragomir Anguelov, Dumitru Erhan, Christian Szegedy, Scott Reed, Cheng-Yang Fu, and Alexander C Berg. 2016. Ssd: Single shot multibox detector. In European conference on computer vision. Springer, 21--37."},{"key":"key-10.1145\/3184558.3192318-14","doi-asserted-by":"crossref","unstructured":"A. H. Miller, W. Feng, A. Fisch, J. Lu, D. Batra, A. Bordes, D. Parikh, and J. Weston. 2017. ParlAI: A Dialog Research Software Platform. arXiv preprint arXiv:1705.06476 (2017).","DOI":"10.18653\/v1\/D17-2014"},{"key":"key-10.1145\/3184558.3192318-15","doi-asserted-by":"crossref","unstructured":"Joseph Redmon, Santosh Divvala, Ross Girshick, and Ali Farhadi. 2016. You only look once: Unified, real-time object detection. In Proceedings of the IEEE conference on computer vision and pattern recognition. 779--788.","DOI":"10.1109\/CVPR.2016.91"},{"key":"key-10.1145\/3184558.3192318-16","doi-asserted-by":"crossref","unstructured":"Sebastian Schuster, Ranjay Krishna, Angel Chang, Li Fei-Fei, and Christopher D Manning. 2015. Generating semantically precise scene graphs from textual descriptions for improved image retrieval. In Proceedings of the fourth workshop on vision and language. 70--80.","DOI":"10.18653\/v1\/W15-2812"},{"key":"key-10.1145\/3184558.3192318-17","doi-asserted-by":"crossref","unstructured":"Ji Wan, Dayong Wang, Steven Chu Hong Hoi, Pengcheng Wu, Jianke Zhu, Yongdong Zhang, and Jintao Li. 2014. Deep learning for content-based image retrieval: A comprehensive study. In Proceedings of the 22nd ACM international conference on Multimedia. ACM, 157--166.","DOI":"10.1145\/2647868.2654948"},{"key":"key-10.1145\/3184558.3192318-18","unstructured":"Peng Wang, Qi Wu, Chunhua Shen, Anton van den Hengel, and Anthony R. Dick. 2016. FVQA: Fact-based Visual Question Answering. CoRR abs\/1606.05433 (2016). arXiv:1606.05433 http:\/\/arxiv.org\/abs\/1606.05433"},{"key":"key-10.1145\/3184558.3192318-19","unstructured":"Qi Wu, Damien Teney, Peng Wang, Chunhua Shen, Anthony R. Dick, and Anton van den Hengel. 2016. Visual Question Answering: A Survey of Methods and Datasets. CoRR abs\/1607.05910 (2016). arXiv:1607.05910 http:\/\/arxiv.org\/abs\/ 1607.05910"},{"key":"key-10.1145\/3184558.3192318-20","unstructured":"Danfei Xu, Yuke Zhu, Christopher B Choy, and Li Fei-Fei. 2017. Scene graph generation by iterative message passing. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Vol. 2."}],"event":{"number":"2018","sponsor":["IW3C2, International World Wide Web Conference Committee","SIGWEB, ACM Special Interest Group on Hypertext, Hypermedia, and Web"],"acronym":"WWW '18","name":"Companion of the The Web Conference 2018","start":{"date-parts":[[2018,4,23]]},"location":"Lyon, France","end":{"date-parts":[[2018,4,27]]}},"container-title":["Companion of the The Web Conference 2018 on The Web Conference 2018 - WWW '18"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3184558.3192318","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/dl.acm.org\/ft_gateway.cfm?id=3192318&ftid=1958475&dwn=1","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T02:11:25Z","timestamp":1750212685000},"score":1,"resource":{"primary":{"URL":"http:\/\/dl.acm.org\/citation.cfm?doid=3184558.3192318"}},"subtitle":[],"proceedings-subject":"The Web Conference 2018","short-title":[],"issued":{"date-parts":[[2018]]},"references-count":20,"URL":"https:\/\/doi.org\/10.1145\/3184558.3192318","relation":{},"subject":[],"published":{"date-parts":[[2018]]}}}