{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,6]],"date-time":"2026-07-06T15:59:30Z","timestamp":1783353570422,"version":"3.54.6"},"publisher-location":"New York, NY, USA","reference-count":46,"publisher":"ACM","license":[{"start":{"date-parts":[[2022,10,10]],"date-time":"2022-10-10T00:00:00Z","timestamp":1665360000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"Natural Science Foundation of China","award":["62006061 and 61872107"],"award-info":[{"award-number":["62006061 and 61872107"]}]},{"name":"Strategic Emerging Industry Development Special Funds of Shenzhen","award":["JCYJ20200109113441941"],"award-info":[{"award-number":["JCYJ20200109113441941"]}]},{"name":"Stable Support Program for Higher Education Institutions of Shenzhen","award":["GXWD20201230 155427003-20200824155011001"],"award-info":[{"award-number":["GXWD20201230 155427003-20200824155011001"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2022,10,10]]},"DOI":"10.1145\/3503161.3548284","type":"proceedings-article","created":{"date-parts":[[2022,10,10]],"date-time":"2022-10-10T15:42:46Z","timestamp":1665416566000},"page":"3587-3597","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":9,"title":["Chunk-aware Alignment and Lexical Constraint for Visual Entailment with Natural Language Explanations"],"prefix":"10.1145","author":[{"given":"Qian","family":"Yang","sequence":"first","affiliation":[{"name":"Harbin Institute of Technology, Shenzhen, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yunxin","family":"Li","sequence":"additional","affiliation":[{"name":"Harbin Institute of Technology, Shenzhen, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Baotian","family":"Hu","sequence":"additional","affiliation":[{"name":"Harbin Institute of Technology, Shenzhen, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Lin","family":"Ma","sequence":"additional","affiliation":[{"name":"Meituan, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yuxin","family":"Ding","sequence":"additional","affiliation":[{"name":"Harbin Institute of Technology, Shenzhen, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Min","family":"Zhang","sequence":"additional","affiliation":[{"name":"Harbin Institute of Technology, Shenzhen, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2022,10,10]]},"reference":[{"key":"e_1_3_2_2_1_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46454-1_24"},{"key":"e_1_3_2_2_2_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00636"},{"key":"e_1_3_2_2_3_1","volume-title":"Proceedings of the Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and\/or Summarization@ACL 2005","author":"Banerjee Satanjeev","year":"2005","unstructured":"Satanjeev Banerjee and Alon Lavie . 2005 . METEOR: An Automatic Metric for MT Evaluation with Improved Correlation with Human Judgments . In Proceedings of the Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and\/or Summarization@ACL 2005 , Ann Arbor, Michigan, USA , June 29, 2005, Jade Goldstein, Alon Lavie, Chin-Yew Lin, and Clare R. Voss (Eds.). Association for Computational Linguistics, 65--72. Satanjeev Banerjee and Alon Lavie. 2005. METEOR: An Automatic Metric for MT Evaluation with Improved Correlation with Human Judgments. In Proceedings of the Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and\/or Summarization@ACL 2005, Ann Arbor, Michigan, USA, June 29, 2005, Jade Goldstein, Alon Lavie, Chin-Yew Lin, and Clare R. Voss (Eds.). Association for Computational Linguistics, 65--72."},{"key":"e_1_3_2_2_4_1","volume-title":"Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics","author":"Bowman Samuel R.","unstructured":"Samuel R. Bowman , Gabor Angeli , Christopher Potts , and Christopher D. Manning . 2015. A large annotated corpus for learning natural language inference . In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics , Lisbon, Portugal, 632--642. Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015. A large annotated corpus for learning natural language inference. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, Lisbon, Portugal, 632--642."},{"key":"e_1_3_2_2_5_1","volume-title":"e-snli: Natural language inference with natural language explanations. Advances in Neural Information Processing Systems 31","author":"Camburu Oana-Maria","year":"2018","unstructured":"Oana-Maria Camburu , Tim Rockt\u00e4schel , Thomas Lukasiewicz , and Phil Blunsom . 2018. e-snli: Natural language inference with natural language explanations. Advances in Neural Information Processing Systems 31 ( 2018 ). Oana-Maria Camburu, Tim Rockt\u00e4schel, Thomas Lukasiewicz, and Phil Blunsom. 2018. e-snli: Natural language inference with natural language explanations. Advances in Neural Information Processing Systems 31 (2018)."},{"key":"e_1_3_2_2_6_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58577-8_7"},{"key":"e_1_3_2_2_7_1","volume-title":"Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805","author":"Devlin Jacob","year":"2018","unstructured":"Jacob Devlin , Ming-Wei Chang , Kenton Lee , and Kristina Toutanova . 2018 . Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805 (2018). Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805 (2018)."},{"key":"e_1_3_2_2_8_1","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition Workshops, CVPR Workshops 2021","author":"Dua Radhika","year":"2021","unstructured":"Radhika Dua , Sai Srinivas Kancheti , and Vineeth N. Balasubramanian . 2021. Beyond VQA: Generating Multi-Word Answers and Rationales to Visual Questions . In IEEE Conference on Computer Vision and Pattern Recognition Workshops, CVPR Workshops 2021 , virtual, June 19 --25 , 2021 . Computer Vision Foundation \/ IEEE, 1623--1632. Radhika Dua, Sai Srinivas Kancheti, and Vineeth N. Balasubramanian. 2021. Beyond VQA: Generating Multi-Word Answers and Rationales to Visual Questions. In IEEE Conference on Computer Vision and Pattern Recognition Workshops, CVPR Workshops 2021, virtual, June 19--25, 2021. Computer Vision Foundation \/ IEEE, 1623--1632."},{"key":"e_1_3_2_2_9_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D16-1044"},{"key":"e_1_3_2_2_10_1","volume-title":"Structured Multi-modal Feature Embedding and Alignment for Image- Sentence Retrieval. In MM '21: ACM Multimedia Conference","author":"Ge Xuri","year":"2021","unstructured":"Xuri Ge , Fuhai Chen , Joemon M. Jose , Zhilong Ji , Zhongqin Wu , and Xiao Liu . 2021 . Structured Multi-modal Feature Embedding and Alignment for Image- Sentence Retrieval. In MM '21: ACM Multimedia Conference , Virtual Event, China, October 20 - 24 , 2021, Heng Tao Shen, Yueting Zhuang, John R. Smith, Yang Yang, Pablo Cesar, Florian Metze, and Balakrishnan Prabhakaran (Eds.). ACM, 5185--5193. Xuri Ge, Fuhai Chen, Joemon M. Jose, Zhilong Ji, Zhongqin Wu, and Xiao Liu. 2021. Structured Multi-modal Feature Embedding and Alignment for Image- Sentence Retrieval. In MM '21: ACM Multimedia Conference, Virtual Event, China, October 20 - 24, 2021, Heng Tao Shen, Yueting Zhuang, John R. Smith, Yang Yang, Pablo Cesar, Florian Metze, and Balakrishnan Prabhakaran (Eds.). ACM, 5185--5193."},{"key":"e_1_3_2_2_11_1","volume-title":"Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering. In 2017 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2017","author":"Goyal Yash","year":"2017","unstructured":"Yash Goyal , Tejas Khot , Douglas Summers-Stay , Dhruv Batra , and Devi Parikh . 2017 . Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering. In 2017 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2017 , Honolulu, HI, USA, July 21--26 , 2017. IEEE Computer Society, 6325--6334. Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. 2017. Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering. In 2017 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2017, Honolulu, HI, USA, July 21--26, 2017. IEEE Computer Society, 6325--6334."},{"key":"e_1_3_2_2_12_1","volume-title":"The Curious Case of Neural Text Degeneration. In 8th International Conference on Learning Representations, ICLR 2020","author":"Holtzman Ari","year":"2020","unstructured":"Ari Holtzman , Jan Buys , Li Du , Maxwell Forbes , and Yejin Choi . 2020 . The Curious Case of Neural Text Degeneration. In 8th International Conference on Learning Representations, ICLR 2020 , Addis Ababa, Ethiopia, April 26--30 , 2020. Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2020. The Curious Case of Neural Text Degeneration. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26--30, 2020."},{"key":"e_1_3_2_2_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.01278"},{"key":"e_1_3_2_2_14_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00128"},{"key":"e_1_3_2_2_15_1","volume-title":"Adam: A Method for Stochastic Optimization. In 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7--9, 2015, Conference Track Proceedings.","author":"Diederik","unstructured":"Diederik P. Kingma and Jimmy Ba. 2015 . Adam: A Method for Stochastic Optimization. In 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7--9, 2015, Conference Track Proceedings. Diederik P. Kingma and Jimmy Ba. 2015. Adam: A Method for Stochastic Optimization. In 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7--9, 2015, Conference Track Proceedings."},{"key":"e_1_3_2_2_16_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58577-8_8"},{"key":"e_1_3_2_2_17_1","volume-title":"Rouge: A package for automatic evaluation of summaries. In Text summarization branches out. 74--81.","author":"Lin Chin-Yew","year":"2004","unstructured":"Chin-Yew Lin . 2004 . Rouge: A package for automatic evaluation of summaries. In Text summarization branches out. 74--81. Chin-Yew Lin. 2004. Rouge: A package for automatic evaluation of summaries. In Text summarization branches out. 74--81."},{"key":"e_1_3_2_2_18_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"e_1_3_2_2_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.01093"},{"key":"e_1_3_2_2_20_1","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/P14-5010"},{"key":"e_1_3_2_2_21_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.findings-emnlp.253"},{"key":"e_1_3_2_2_22_1","volume-title":"Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, July 6--12","author":"Papineni Kishore","year":"2002","unstructured":"Kishore Papineni , Salim Roukos , Todd Ward , and Wei-Jing Zhu . 2002 . Bleu: a Method for Automatic Evaluation of Machine Translation . In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, July 6--12 , 2002, Philadelphia, PA, USA. ACL, 311--318. Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: a Method for Automatic Evaluation of Machine Translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, July 6--12, 2002, Philadelphia, PA, USA. ACL, 311--318."},{"key":"e_1_3_2_2_23_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00915"},{"key":"e_1_3_2_2_24_1","doi-asserted-by":"publisher","DOI":"10.1109\/WACV45572.2020.9093295"},{"key":"e_1_3_2_2_25_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.303"},{"key":"e_1_3_2_2_26_1","volume-title":"What to pre-train on efficient intermediate task selection. arXiv preprint arXiv:2104.08247","author":"Poth Clifton","year":"2021","unstructured":"Clifton Poth , Jonas Pfeiffer , Andreas R\u00fcckl\u00e9 , and Iryna Gurevych . 2021. What to pre-train on efficient intermediate task selection. arXiv preprint arXiv:2104.08247 ( 2021 ). Clifton Poth, Jonas Pfeiffer, Andreas R\u00fcckl\u00e9, and Iryna Gurevych. 2021. What to pre-train on efficient intermediate task selection. arXiv preprint arXiv:2104.08247 (2021)."},{"key":"e_1_3_2_2_27_1","volume-title":"Proceedings of the 1st Conference of the Asia-Pacific","author":"Prabhu Nikhil","year":"2021","unstructured":"Nikhil Prabhu and Katharina Kann . 2020. Making a Point: Pointer-Generator Transformers for Disjoint Vocabularies . In Proceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th International Joint Conference on Natural Language Processing: Student Research Workshop, AACL\/IJCNLP 2021 , Suzhou, China, December 4--7, 2020. 85--92. Nikhil Prabhu and Katharina Kann. 2020. Making a Point: Pointer-Generator Transformers for Disjoint Vocabularies. In Proceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th International Joint Conference on Natural Language Processing: Student Research Workshop, AACL\/IJCNLP 2021, Suzhou, China, December 4--7, 2020. 85--92."},{"key":"e_1_3_2_2_28_1","unstructured":"Alec Radford JeffreyWu Rewon Child David Luan Dario Amodei Ilya Sutskever etal 2019. Language models are unsupervised multitask learners. OpenAI blog 1 8 (2019) 9.  Alec Radford JeffreyWu Rewon Child David Luan Dario Amodei Ilya Sutskever et al. 2019. Language models are unsupervised multitask learners. OpenAI blog 1 8 (2019) 9."},{"key":"e_1_3_2_2_29_1","volume-title":"Exploring models and data for image question answering. Advances in neural information processing systems 28","author":"Ren Mengye","year":"2015","unstructured":"Mengye Ren , Ryan Kiros , and Richard Zemel . 2015. Exploring models and data for image question answering. Advances in neural information processing systems 28 ( 2015 ). Mengye Ren, Ryan Kiros, and Richard Zemel. 2015. Exploring models and data for image question answering. Advances in neural information processing systems 28 (2015)."},{"key":"e_1_3_2_2_30_1","volume-title":"Faster r-cnn: Towards real-time object detection with region proposal networks. Advances in neural information processing systems 28","author":"Ren Shaoqing","year":"2015","unstructured":"Shaoqing Ren , Kaiming He , Ross Girshick , and Jian Sun . 2015. Faster r-cnn: Towards real-time object detection with region proposal networks. Advances in neural information processing systems 28 ( 2015 ). Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. 2015. Faster r-cnn: Towards real-time object detection with region proposal networks. Advances in neural information processing systems 28 (2015)."},{"key":"e_1_3_2_2_31_1","volume-title":"NLX-GPT: A Model for Natural Language Explanations in Vision and Vision-Language Tasks. CoRR abs\/2203.05081","author":"Sammani Fawaz","year":"2022","unstructured":"Fawaz Sammani , Tanmoy Mukherjee , and Nikos Deligiannis . 2022. NLX-GPT: A Model for Natural Language Explanations in Vision and Vision-Language Tasks. CoRR abs\/2203.05081 ( 2022 ). Fawaz Sammani, Tanmoy Mukherjee, and Nikos Deligiannis. 2022. NLX-GPT: A Model for Natural Language Explanations in Vision and Vision-Language Tasks. CoRR abs\/2203.05081 (2022)."},{"key":"e_1_3_2_2_32_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/W15-2812"},{"key":"e_1_3_2_2_33_1","volume-title":"Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics, ACL 2017","volume":"1083","author":"Liu Peter J.","unstructured":"Abigail See, Peter J. Liu , and Christopher D. Manning . 2017. Get To The Point: Summarization with Pointer-Generator Networks . In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics, ACL 2017 , Vancouver, Canada, July 30 - August 4 , Volume 1: Long Papers. 1073-- 1083 . Abigail See, Peter J. Liu, and Christopher D. Manning. 2017. Get To The Point: Summarization with Pointer-Generator Networks. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics, ACL 2017, Vancouver, Canada, July 30 - August 4, Volume 1: Long Papers. 1073--1083."},{"key":"e_1_3_2_2_34_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.74"},{"key":"e_1_3_2_2_35_1","volume-title":"Fourth Conference on Computational Natural Language Learning and the Second Learning Language in Logic Workshop.","author":"Erik","unstructured":"Erik F. Tjong Kim Sang and Sabine Buchholz. 2000. Introduction to the CoNLL- 2000 Shared Task Chunking . In Fourth Conference on Computational Natural Language Learning and the Second Learning Language in Logic Workshop. Erik F. Tjong Kim Sang and Sabine Buchholz. 2000. Introduction to the CoNLL- 2000 Shared Task Chunking. In Fourth Conference on Computational Natural Language Learning and the Second Learning Language in Logic Workshop."},{"key":"e_1_3_2_2_36_1","volume-title":"Attention is all you need. Advances in neural information processing systems 30","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani , Noam Shazeer , Niki Parmar , Jakob Uszkoreit , Llion Jones , Aidan N Gomez , Lukasz Kaiser , and Illia Polosukhin . 2017. Attention is all you need. Advances in neural information processing systems 30 ( 2017 ). Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in neural information processing systems 30 (2017)."},{"key":"e_1_3_2_2_37_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7299087"},{"key":"e_1_3_2_2_38_1","volume-title":"CoRR abs\/2202.03052","author":"Wang Peng","year":"2022","unstructured":"Peng Wang , An Yang , Rui Men , Junyang Lin , Shuai Bai , Zhikang Li , Jianxin Ma , Chang Zhou , Jingren Zhou , and Hongxia Yang . 2022. Unifying Architectures , Tasks, and Modalities Through a Simple Sequence-to- Sequence Learning Framework . CoRR abs\/2202.03052 ( 2022 ). arXiv:2202.03052 Peng Wang, An Yang, Rui Men, Junyang Lin, Shuai Bai, Zhikang Li, Jianxin Ma, Chang Zhou, Jingren Zhou, and Hongxia Yang. 2022. Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework. CoRR abs\/2202.03052 (2022). arXiv:2202.03052"},{"key":"e_1_3_2_2_39_1","volume-title":"Zihang Dai, Yulia Tsvetkov, and Yuan Cao.","author":"Wang Zirui","year":"2021","unstructured":"Zirui Wang , Jiahui Yu , Adams Wei Yu , Zihang Dai, Yulia Tsvetkov, and Yuan Cao. 2021 . Simvlm : Simple visual language model pretraining with weak supervision. arXiv preprint arXiv:2108.10904 (2021). Zirui Wang, Jiahui Yu, Adams Wei Yu, Zihang Dai, Yulia Tsvetkov, and Yuan Cao. 2021. Simvlm: Simple visual language model pretraining with weak supervision. arXiv preprint arXiv:2108.10904 (2021)."},{"key":"e_1_3_2_2_40_1","volume-title":"Proceedings of the 2019 ACL Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, BlackboxNLP@ACL 2019","author":"Mooney JialinWu","year":"2019","unstructured":"JialinWu and Raymond J. Mooney . 2019. Faithful Multimodal Explanation for Visual Question Answering . In Proceedings of the 2019 ACL Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, BlackboxNLP@ACL 2019 , Florence, Italy , August 1, 2019 , Tal Linzen, Grzegorz Chrupala, Yonatan Belinkov, and Dieuwke Hupkes (Eds.). Association for Computational Linguistics, 103--112. JialinWu and Raymond J. Mooney. 2019. Faithful Multimodal Explanation for Visual Question Answering. In Proceedings of the 2019 ACL Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, BlackboxNLP@ACL 2019, Florence, Italy, August 1, 2019, Tal Linzen, Grzegorz Chrupala, Yonatan Belinkov, and Dieuwke Hupkes (Eds.). Association for Computational Linguistics, 103--112."},{"key":"e_1_3_2_2_41_1","volume-title":"Visual Entailment: A Novel Task for Fine-Grained Image Understanding. CoRR abs\/1901.06706","author":"Xie Ning","year":"2019","unstructured":"Ning Xie , Farley Lai , Derek Doran , and Asim Kadav . 2019 . Visual Entailment: A Novel Task for Fine-Grained Image Understanding. CoRR abs\/1901.06706 (2019). Ning Xie, Farley Lai, Derek Doran, and Asim Kadav. 2019. Visual Entailment: A Novel Task for Fine-Grained Image Understanding. CoRR abs\/1901.06706 (2019)."},{"key":"e_1_3_2_2_42_1","volume-title":"Proceedings, Part I (Lecture Notes in Computer Science","volume":"706","author":"Yang Jianwei","year":"2018","unstructured":"Jianwei Yang , Jiasen Lu , Stefan Lee , Dhruv Batra , and Devi Parikh . 2018 . Graph R-CNN for Scene Graph Generation. In Computer Vision - ECCV 2018 - 15th European Conference, Munich, Germany, September 8--14, 2018 , Proceedings, Part I (Lecture Notes in Computer Science , Vol. 11205), Vittorio Ferrari, Martial Hebert, Cristian Sminchisescu, and Yair Weiss (Eds.). Springer, 690-- 706 . Jianwei Yang, Jiasen Lu, Stefan Lee, Dhruv Batra, and Devi Parikh. 2018. Graph R-CNN for Scene Graph Generation. In Computer Vision - ECCV 2018 - 15th European Conference, Munich, Germany, September 8--14, 2018, Proceedings, Part I (Lecture Notes in Computer Science, Vol. 11205), Vittorio Ferrari, Martial Hebert, Cristian Sminchisescu, and Yair Weiss (Eds.). Springer, 690--706."},{"key":"e_1_3_2_2_43_1","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00166"},{"key":"e_1_3_2_2_44_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00688"},{"key":"e_1_3_2_2_45_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.01180"},{"key":"e_1_3_2_2_46_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.540"}],"event":{"name":"MM '22: The 30th ACM International Conference on Multimedia","location":"Lisboa Portugal","acronym":"MM '22","sponsor":["SIGMM ACM Special Interest Group on Multimedia"]},"container-title":["Proceedings of the 30th ACM International Conference on Multimedia"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3503161.3548284","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3503161.3548284","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T19:00:43Z","timestamp":1750186843000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3503161.3548284"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,10,10]]},"references-count":46,"alternative-id":["10.1145\/3503161.3548284","10.1145\/3503161"],"URL":"https:\/\/doi.org\/10.1145\/3503161.3548284","relation":{},"subject":[],"published":{"date-parts":[[2022,10,10]]},"assertion":[{"value":"2022-10-10","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}