{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,22]],"date-time":"2026-07-22T03:41:45Z","timestamp":1784691705074,"version":"3.55.0"},"reference-count":50,"publisher":"Association for Computing Machinery (ACM)","issue":"11","license":[{"start":{"date-parts":[[2024,11,14]],"date-time":"2024-11-14T00:00:00Z","timestamp":1731542400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100001809","name":"Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62272018"],"award-info":[{"award-number":["62272018"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2024,11,30]]},"abstract":"<jats:p>\n            This article focuses on reasoning about the location and time behind images. Given that pre-trained vision-language models (VLMs) exhibit excellent image and text understanding capabilities, most existing methods leverage them to match visual cues with location and time-related descriptions. However, these methods cannot look beyond the actual content of an image, failing to produce satisfactory reasoning results, as such reasoning requires connecting visual details with rich external cues (e.g., relevant event contexts). To this end, we propose a novel reasoning method,\n            <jats:italic>QR-CLIP<\/jats:italic>\n            , that aims at enhancing the model\u2019s ability to reason about location and time through interaction with external explicit knowledge such as Wikipedia. Specifically,\n            <jats:italic>QR-CLIP<\/jats:italic>\n            consists of two modules: (1) The\n            <jats:italic>Quantity<\/jats:italic>\n            module abstracts the image into multiple distinct representations and uses them to search and gather external knowledge from different perspectives that are beneficial to model reasoning. (2) The\n            <jats:italic>Relevance<\/jats:italic>\n            module filters the visual features and the searched explicit knowledge and dynamically integrates them to form a comprehensive reasoning result. Extensive experiments demonstrate the effectiveness and generalizability of\n            <jats:italic>QR-CLIP<\/jats:italic>\n            . On the WikiTiLo dataset,\n            <jats:italic>QR-CLIP<\/jats:italic>\n            boosts the accuracy of location (country) and time reasoning by 7.03% and 2.22%, respectively, over previous SOTA methods. On the more challenging TARA dataset, it improves the accuracy for location and time reasoning by 3.05% and 2.45%, respectively. The source code is at\n            <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" ext-link-type=\"url\" xlink:href=\"https:\/\/github.com\/Shi-Wm\/QR-CLIP\">https:\/\/github.com\/Shi-Wm\/QR-CLIP<\/jats:ext-link>\n            .\n          <\/jats:p>","DOI":"10.1145\/3689638","type":"journal-article","created":{"date-parts":[[2024,8,24]],"date-time":"2024-08-24T08:44:42Z","timestamp":1724489082000},"page":"1-22","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":3,"title":["QR-CLIP: Introducing Explicit Knowledge for Location and Time Reasoning"],"prefix":"10.1145","volume":"20","author":[{"ORCID":"https:\/\/orcid.org\/0009-0004-2344-1025","authenticated-orcid":false,"given":"Weimin","family":"Shi","sequence":"first","affiliation":[{"name":"State Key Laboratory of Virtual Reality Technology and Systems, School of Computer Science and Engineering, Beihang University, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6636-5702","authenticated-orcid":false,"given":"Dehong","family":"Gao","sequence":"additional","affiliation":[{"name":"School of Cybersecurity, Northwestern Polytechnical University, Xi\u2019an, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7253-4998","authenticated-orcid":false,"given":"Yuan","family":"Xiong","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Virtual Reality Technology and Systems, School of Computer Science and Engineering, Beihang University, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5825-7517","authenticated-orcid":false,"given":"Zhong","family":"Zhou","sequence":"additional","affiliation":[{"name":"Zhongguancun Laboratory, Beijing, China and State Key Laboratory of Virtual Reality Technology and Systems, School of Computer Science and Engineering, Beihang University, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2024,11,14]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01593"},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1145\/3137597.3137600"},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.1145\/3451215"},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2021.3073267"},{"key":"e_1_3_1_6_2","doi-asserted-by":"publisher","DOI":"10.1145\/3352573"},{"issue":"5","key":"e_1_3_1_7_2","first-page":"4425","article-title":"A survey on accuracy-oriented neural recommendation: From collaborative filtering to information-rich recommendation","volume":"35","author":"Wu Le","year":"2022","unstructured":"Le Wu, Xiangnan He, Xiang Wang, Kun Zhang, and Meng Wang. 2022. A survey on accuracy-oriented neural recommendation: From collaborative filtering to information-rich recommendation. IEEE Transactions on Knowledge and Data Engineering 35, 5 (2022), 4425\u20134445.","journal-title":"IEEE Transactions on Knowledge and Data Engineering"},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2021.3064408"},{"key":"e_1_3_1_9_2","doi-asserted-by":"crossref","unstructured":"Jie Wen Nan Jiang Lang Li Jie Zhou Yanpei Li Hualin Zhan Guang Kou Weihao Gu and Jiahui Zhao. 2024. TA-detector: A GNN-based anomaly detector via trust relationship. ACM Transactions on Multimedia Computing Communications and Applications (2024). Retrieved from https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3672401","DOI":"10.1145\/3672401"},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.ins.2020.11.026"},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.01245"},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i07.6875"},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00499"},{"key":"e_1_3_1_14_2","unstructured":"Rohin Manvi Samar Khanna Gengchen Mai Marshall Burke David Lobell and Stefano Ermon. 2023. Geollm: Extracting geospatial knowledge from large language models. arXiv:2310.06213. Retrieved from https:\/\/arxiv.org\/abs\/2310.06213"},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.00483"},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.1109\/WACV51458.2022.00275"},{"key":"e_1_3_1_17_2","first-page":"8748","volume-title":"International Conference on Machine Learning","author":"Radford Alec","year":"2021","unstructured":"Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning. PMLR, 8748\u20138763."},{"key":"e_1_3_1_18_2","first-page":"12888","volume-title":"International Conference on Machine Learning","author":"Li Junnan","year":"2022","unstructured":"Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. 2022. Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In International Conference on Machine Learning. PMLR, 12888\u201312900."},{"key":"e_1_3_1_19_2","first-page":"1138","volume-title":"Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics","volume":"1","author":"Fu Xingyu","year":"2022","unstructured":"Xingyu Fu, Ben Zhou, Ishaan Chandratreya, Carl Vondrick, and Dan Roth. 2022. There\u2019s a time and place for reasoning beyond the image. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics, Vol. 1, Long Papers, 1138\u20131149."},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1109\/WACV57701.2024.00069"},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.1037\/10096-012"},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.7551\/mitpress\/1881.001.0001"},{"key":"e_1_3_1_23_2","volume-title":"Distributed Cognition","author":"Hutchins Edwin","year":"2000","unstructured":"Edwin Hutchins. 2000. Distributed Cognition. Elsevier Science."},{"key":"e_1_3_1_24_2","first-page":"4171","volume-title":"Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics","author":"Kenton Jacob Devlin Ming-Wei Chang","year":"2019","unstructured":"Jacob Devlin Ming-Wei Chang Kenton and Lee Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics, 4171\u20134186."},{"key":"e_1_3_1_25_2","unstructured":"Alexey Dosovitskiy Lucas Beyer Alexander Kolesnikov Dirk Weissenborn Xiaohua Zhai Thomas Unterthiner Mostafa Dehghani Matthias Minderer Georg Heigold Sylvain Gelly et al. 2020. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv:2010.11929."},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v33i01.3301873"},{"key":"e_1_3_1_27_2","first-page":"7579","volume-title":"Proceedings of the Annual Meeting of the Association for Computational Linguistics","author":"Zhou Ben","year":"2022","unstructured":"Ben Zhou, Qiang Ning, Daniel Khashabi, and Dan Roth. 2022. Temporal common sense acquisition with minimal supervision. In Proceedings of the Annual Meeting of the Association for Computational Linguistics, 7579\u20137589."},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.acl-long.195"},{"key":"e_1_3_1_29_2","doi-asserted-by":"crossref","unstructured":"Yang Chen Hexiang Hu Yi Luan Haitian Sun Soravit Changpinyo Alan Ritter and Ming-Wei Chang. 2023. Can pre-trained vision and language models answer visual information-seeking questions? arXiv:2302.11713. Retrieved from https:\/\/arxiv.org\/abs\/2302.11713","DOI":"10.18653\/v1\/2023.emnlp-main.925"},{"key":"e_1_3_1_30_2","unstructured":"Xiujie Song Mengyue Wu Kenny Q. Zhu Chunhao Zhang and Yanyi Chen. 2024. A cognitive evaluation benchmark of image reasoning and description for large vision language models. arXiv:2402.18409. Retrieved from https:\/\/arxiv.org\/abs\/2402.18409"},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.02080"},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.00654"},{"key":"e_1_3_1_33_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW59228.2023.00593"},{"key":"e_1_3_1_34_2","first-page":"1597","article-title":"A simple framework for contrastive learning of visual representations","author":"Chen Ting","year":"2020","unstructured":"Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. 2020. A simple framework for contrastive learning of visual representations. In International Conference on Machine Learning. PMLR, 1597\u20131607.","journal-title":"International Conference on Machine Learning"},{"key":"e_1_3_1_35_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-19833-5_29"},{"key":"e_1_3_1_36_2","first-page":"4904","volume-title":"International Conference on Machine Learning","author":"Jia Chao","year":"2021","unstructured":"Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. 2021. Scaling up visual and vision-language representation learning with noisy text supervision. In International Conference on Machine Learning. PMLR, 4904\u20134916."},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-19809-0_30"},{"key":"e_1_3_1_38_2","first-page":"35959","volume-title":"Proceedings of the 36th International Conference on Neural Information Processing Systems","author":"Gao Yuting","year":"2022","unstructured":"Yuting Gao, Jinfeng Liu, Zihan Xu, Jun Zhang, Ke Li, Rongrong Ji, and Chunhua Shen. 2022. Pyramidclip: Hierarchical feature alignment for vision-language model pretraining. In Proceedings of the 36th International Conference on Neural Information Processing Systems, 35959\u201335970."},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.00283"},{"key":"e_1_3_1_40_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00061"},{"key":"e_1_3_1_42_2","doi-asserted-by":"publisher","DOI":"10.5555\/3495724.3496297"},{"key":"e_1_3_1_43_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01553"},{"key":"e_1_3_1_44_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00975"},{"key":"e_1_3_1_45_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.acl-long.414"},{"key":"e_1_3_1_46_2","first-page":"1185","article-title":"COFAR: Commonsense and factual reasoning in image search","volume":"1","author":"Gatti Prajwal","year":"2022","unstructured":"Prajwal Gatti, Abhirama Subramanyam Penamakuri, Revant Teotia, Anand Mishra, Shubhashis Sengupta, and Roshni Ramnani. 2022. COFAR: Commonsense and factual reasoning in image search. In Proceedings of the 2nd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 12th International Joint Conference on Natural Language Processing (AACL-IJCNLP), Vol. 1, Long Papers, 1185\u20131199.","journal-title":"Proceedings of the 2nd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 12th International Joint Conference on Natural Language Processing (AACL-IJCNLP)"},{"key":"e_1_3_1_47_2","doi-asserted-by":"publisher","DOI":"10.1145\/3404835.3463257"},{"key":"e_1_3_1_48_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00986"},{"key":"e_1_3_1_49_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"e_1_3_1_50_2","first-page":"21548","article-title":"A contrastive framework for neural text generation","author":"Su Yixuan","year":"2022","unstructured":"Yixuan Su, Tian Lan, Yan Wang, Dani Yogatama, Lingpeng Kong, and Nigel Collier. 2022. A contrastive framework for neural text generation. In Proceedings of the 36th Conference on Neural Information Processing Systems (NeurIPS \u201922), 21548\u201321561.","journal-title":"Proceedings of the 36th Conference on Neural Information Processing Systems (NeurIPS \u201922)"},{"key":"e_1_3_1_51_2","unstructured":"Yixuan Su Tian Lan Yahui Liu Fangyu Liu Dani Yogatama Yan Wang Lingpeng Kong and Nigel Collier. 2022. Language models can see: Plugging visual controls in text generation. arXiv:2205.02655. Retrieved from https:\/\/arxiv.org\/abs\/2205.02655"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3689638","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3689638","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T01:09:47Z","timestamp":1750295387000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3689638"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,11,14]]},"references-count":50,"journal-issue":{"issue":"11","published-print":{"date-parts":[[2024,11,30]]}},"alternative-id":["10.1145\/3689638"],"URL":"https:\/\/doi.org\/10.1145\/3689638","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,11,14]]},"assertion":[{"value":"2023-12-15","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-08-15","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-11-14","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}