{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,18]],"date-time":"2026-08-18T07:24:22Z","timestamp":1787037862193,"version":"build-2736575974"},"reference-count":37,"publisher":"Elsevier BV","license":[{"start":{"date-parts":[[2026,10,1]],"date-time":"2026-10-01T00:00:00Z","timestamp":1790812800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/www.elsevier.com\/tdm\/userlicense\/1.0\/"},{"start":{"date-parts":[[2026,10,1]],"date-time":"2026-10-01T00:00:00Z","timestamp":1790812800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/www.elsevier.com\/legal\/tdmrep-license"},{"start":{"date-parts":[[2026,10,1]],"date-time":"2026-10-01T00:00:00Z","timestamp":1790812800000},"content-version":"stm-asf","delay-in-days":0,"URL":"https:\/\/doi.org\/10.15223\/policy-017"},{"start":{"date-parts":[[2026,10,1]],"date-time":"2026-10-01T00:00:00Z","timestamp":1790812800000},"content-version":"stm-asf","delay-in-days":0,"URL":"https:\/\/doi.org\/10.15223\/policy-037"},{"start":{"date-parts":[[2026,10,1]],"date-time":"2026-10-01T00:00:00Z","timestamp":1790812800000},"content-version":"stm-asf","delay-in-days":0,"URL":"https:\/\/doi.org\/10.15223\/policy-012"},{"start":{"date-parts":[[2026,10,1]],"date-time":"2026-10-01T00:00:00Z","timestamp":1790812800000},"content-version":"stm-asf","delay-in-days":0,"URL":"https:\/\/doi.org\/10.15223\/policy-029"},{"start":{"date-parts":[[2026,10,1]],"date-time":"2026-10-01T00:00:00Z","timestamp":1790812800000},"content-version":"stm-asf","delay-in-days":0,"URL":"https:\/\/doi.org\/10.15223\/policy-004"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100007129","name":"Shandong Province Natural Science Foundation","doi-asserted-by":"publisher","id":[{"id":"10.13039\/501100007129","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["elsevier.com","sciencedirect.com"],"crossmark-restriction":true},"short-container-title":["Knowledge-Based Systems"],"published-print":{"date-parts":[[2026,10]]},"DOI":"10.1016\/j.knosys.2026.116687","type":"journal-article","created":{"date-parts":[[2026,7,28]],"date-time":"2026-07-28T15:42:23Z","timestamp":1785253343000},"page":"116687","update-policy":"https:\/\/doi.org\/10.1016\/elsevier_cm_policy","source":"Crossref","is-referenced-by-count":0,"special_numbering":"PB","title":["TECL: Time-Equivariant Contrastive Learning for weakly-supervised Grounded Video Question Answering"],"prefix":"10.1016","volume":"351","author":[{"given":"Wenzhe","family":"Liu","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7217-3109","authenticated-orcid":false,"given":"Zhenfang","family":"Zhu","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Qiang","family":"Lu","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Shengtai","family":"Zhang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Liang","family":"Zhao","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1812-1316","authenticated-orcid":false,"given":"Dawei","family":"Zhao","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xinyu","family":"Jiang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Menglin","family":"Zhu","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yan","family":"Wang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"78","reference":[{"key":"10.1016\/j.knosys.2026.116687_b1","doi-asserted-by":"crossref","first-page":"76749","DOI":"10.52202\/075280-3354","article-title":"Self-chained image-language model for video localization and question answering","volume":"36","author":"Yu","year":"2023","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"10.1016\/j.knosys.2026.116687_b2","doi-asserted-by":"crossref","unstructured":"Juhong Min, Shyamal Buch, Arsha Nagrani, Minsu Cho, Cordelia Schmid, MoReVQA: Exploring Modular Reasoning Models for Video Question Answering, in: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 13235\u201313245.","DOI":"10.1109\/CVPR52733.2024.01257"},{"key":"10.1016\/j.knosys.2026.116687_b3","doi-asserted-by":"crossref","first-page":"4554","DOI":"10.1109\/TMM.2023.3323878","article-title":"Locate before answering: Answer guided question localization for video question answering","volume":"26","author":"Qian","year":"2023","journal-title":"IEEE Trans. Multimed."},{"key":"10.1016\/j.knosys.2026.116687_b4","doi-asserted-by":"crossref","unstructured":"Qirui Chen, Shangzhe Di, Weidi Xie, Grounded multi-hop videoqa in long-form egocentric videos, in: Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 39, 2025, pp. 2159\u20132167.","DOI":"10.1609\/aaai.v39i2.32214"},{"key":"10.1016\/j.knosys.2026.116687_b5","doi-asserted-by":"crossref","first-page":"5178","DOI":"10.1109\/TIP.2022.3191841","article-title":"HiSA: Hierarchically semantic associating for video temporal grounding","volume":"31","author":"Xu","year":"2022","journal-title":"IEEE Trans. Image Process."},{"key":"10.1016\/j.knosys.2026.116687_b6","doi-asserted-by":"crossref","DOI":"10.1016\/j.knosys.2024.112629","article-title":"Robust visual question answering utilizing bias instances and label imbalance","volume":"305","author":"Zhao","year":"2024","journal-title":"Knowl.-Based Syst."},{"key":"10.1016\/j.knosys.2026.116687_b7","doi-asserted-by":"crossref","unstructured":"Yaoyao Zhong, Wei Ji, Junbin Xiao, Yicong Li, Weihong Deng, Tat-Seng Chua, Video question answering: Datasets, algorithms and challenges, in: Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 2022, pp. 6439\u20136455.","DOI":"10.18653\/v1\/2022.emnlp-main.432"},{"key":"10.1016\/j.knosys.2026.116687_b8","doi-asserted-by":"crossref","first-page":"42748","DOI":"10.52202\/075280-1852","article-title":"Perception test: A diagnostic benchmark for multimodal video models","volume":"36","author":"Patraucean","year":"2023","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"10.1016\/j.knosys.2026.116687_b9","doi-asserted-by":"crossref","unstructured":"Yunseok Jang, Yale Song, Youngjae Yu, Youngjin Kim, Gunhee Kim, Tgif-qa: Toward spatio-temporal reasoning in visual question answering, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017, pp. 2758\u20132766.","DOI":"10.1109\/CVPR.2017.149"},{"key":"10.1016\/j.knosys.2026.116687_b10","series-title":"International Conference on Machine Learning","first-page":"1597","article-title":"A simple framework for contrastive learning of visual representations","author":"Chen","year":"2020"},{"key":"10.1016\/j.knosys.2026.116687_b11","series-title":"Equivariant contrastive learning","author":"Dangovski","year":"2021"},{"key":"10.1016\/j.knosys.2026.116687_b12","doi-asserted-by":"crossref","unstructured":"Junbin Xiao, Angela Yao, Yicong Li, Tat-Seng Chua, Can i trust your answer? visually grounded video question answering, in: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 13204\u201313214.","DOI":"10.1109\/CVPR52733.2024.01254"},{"key":"10.1016\/j.knosys.2026.116687_b13","doi-asserted-by":"crossref","unstructured":"Can Zhang, Tianyu Yang, Junwu Weng, Meng Cao, Jue Wang, Yuexian Zou, Unsupervised pre-training for temporal action localization tasks, in: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 14031\u201314041.","DOI":"10.1109\/CVPR52688.2022.01364"},{"key":"10.1016\/j.knosys.2026.116687_b14","doi-asserted-by":"crossref","unstructured":"Junbin Xiao, Angela Yao, Ziqian Liu, Yicong Li, Wei Ji, Tat-Seng Chua, NExT-QA: Next phase of question answering to explaining temporal actions, in: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 9777\u20139786.","DOI":"10.1109\/CVPR46437.2021.00965"},{"key":"10.1016\/j.knosys.2026.116687_b15","doi-asserted-by":"crossref","unstructured":"Jie Lei, Licheng Yu, Mohit Bansal, Tamara L. Berg, TVQA: Localized, compositional video question answering, in: Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, 2018, pp. 1369\u20131379.","DOI":"10.18653\/v1\/D18-1167"},{"key":"10.1016\/j.knosys.2026.116687_b16","doi-asserted-by":"crossref","unstructured":"Jie Lei, Licheng Yu, Tamara L. Berg, Mohit Bansal, TVQA+: Spatio-temporal grounding for video question answering, in: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 2020, pp. 8211\u20138225.","DOI":"10.18653\/v1\/2020.acl-main.730"},{"key":"10.1016\/j.knosys.2026.116687_b17","doi-asserted-by":"crossref","unstructured":"Antoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev, Cordelia Schmid, Just ask: Learning to answer questions from millions of narrated videos, in: Proceedings of the IEEE\/CVF International Conference on Computer Vision, 2021, pp. 1686\u20131697.","DOI":"10.1109\/ICCV48922.2021.00171"},{"key":"10.1016\/j.knosys.2026.116687_b18","doi-asserted-by":"crossref","unstructured":"Hao Zhang, Aixin Sun, Wei Jing, Joey Tianyi Zhou, Span-based localizing network for natural language video localization, in: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 2020, pp. 6543\u20136554.","DOI":"10.18653\/v1\/2020.acl-main.585"},{"key":"10.1016\/j.knosys.2026.116687_b19","doi-asserted-by":"crossref","unstructured":"Songyang Zhang, Houwen Peng, Jianlong Fu, Jiebo Luo, Learning 2D temporal adjacent networks for moment localization with natural language, in: Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 34, 2020, pp. 12870\u201312877.","DOI":"10.1609\/aaai.v34i07.6984"},{"key":"10.1016\/j.knosys.2026.116687_b20","doi-asserted-by":"crossref","unstructured":"Xinyu Lin, Yucheng Wei, Yue Zhang, Yuxin Peng, UniVTG: Towards unified video-language temporal grounding, in: Proceedings of the IEEE\/CVF International Conference on Computer Vision, 2023, pp. 2792\u20132802.","DOI":"10.1109\/ICCV51070.2023.00262"},{"key":"10.1016\/j.knosys.2026.116687_b21","first-page":"24563","article-title":"End-to-end moment retrieval and highlight detection via query-based transformer","volume":"34","author":"Lei","year":"2021","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"10.1016\/j.knosys.2026.116687_b22","series-title":"European Conference on Computer Vision","first-page":"92","article-title":"Timecraft: Navigate weakly-supervised temporal grounded video question answering via bi-directional reasoning","author":"Liu","year":"2024"},{"key":"10.1016\/j.knosys.2026.116687_b23","doi-asserted-by":"crossref","unstructured":"Haibo Wang, Chenghang Lai, Yixuan Sun, Weifeng Ge, Weakly supervised gaussian contrastive grounding with large multimodal models for video question answering, in: Proceedings of the 32nd ACM International Conference on Multimedia, 2024, pp. 5289\u20135298.","DOI":"10.1145\/3664647.3680826"},{"key":"10.1016\/j.knosys.2026.116687_b24","doi-asserted-by":"crossref","unstructured":"Ayush Gupta, Anirban Roy, Rama Chellappa, Nathaniel D Bastian, Alvaro Velasquez, Susmit Jha, TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision, in: Proceedings of the IEEE\/CVF International Conference on Computer Vision, 2025, pp. 23593\u201323603.","DOI":"10.1109\/ICCV51701.2025.02190"},{"key":"10.1016\/j.knosys.2026.116687_b25","unstructured":"Yicong Li, Xiang Wang, Junbin Xiao, Wei Ji, Tat-Seng Chua, Invariant grounding for video question answering, in: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 2928\u20132937."},{"key":"10.1016\/j.knosys.2026.116687_b26","doi-asserted-by":"crossref","unstructured":"Weixing Chen, Yang Liu, Binglin Chen, Jiandong Su, Yongsen Zheng, Liang Lin, Cross-modal causal relation alignment for video question grounding, in: Proceedings of the Computer Vision and Pattern Recognition Conference, 2025, pp. 24087\u201324096.","DOI":"10.1109\/CVPR52734.2025.02243"},{"key":"10.1016\/j.knosys.2026.116687_b27","doi-asserted-by":"crossref","unstructured":"Rui Qian, Tianjian Meng, Boqing Gong, Ming-Hsuan Yang, Huisheng Wang, Serge Belongie, Yin Cui, Spatiotemporal contrastive video representation learning, in: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 6964\u20136974.","DOI":"10.1109\/CVPR46437.2021.00689"},{"key":"10.1016\/j.knosys.2026.116687_b28","doi-asserted-by":"crossref","unstructured":"Jue Wang, Gedas Bertasius, Du Tran, Lorenzo Torresani, Long-short temporal contrastive learning of video transformers, in: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 14010\u201314020.","DOI":"10.1109\/CVPR52688.2022.01362"},{"key":"10.1016\/j.knosys.2026.116687_b29","doi-asserted-by":"crossref","unstructured":"Ji Lin, Chuang Gan, Song Han, Tsm: Temporal shift module for efficient video understanding, in: Proceedings of the IEEE\/CVF International Conference on Computer Vision, 2019, pp. 7083\u20137093.","DOI":"10.1109\/ICCV.2019.00718"},{"key":"10.1016\/j.knosys.2026.116687_b30","series-title":"International Conference on Machine Learning","first-page":"2990","article-title":"Group equivariant convolutional networks","author":"Cohen","year":"2016"},{"key":"10.1016\/j.knosys.2026.116687_b31","doi-asserted-by":"crossref","unstructured":"Simon Jenni, Hailin Jin, Time-equivariant contrastive video representation learning, in: Proceedings of the IEEE\/CVF International Conference on Computer Vision, 2021, pp. 9970\u20139980.","DOI":"10.1109\/ICCV48922.2021.00982"},{"key":"10.1016\/j.knosys.2026.116687_b32","series-title":"Roberta: A robustly optimized bert pretraining approach","author":"Liu","year":"2019"},{"key":"10.1016\/j.knosys.2026.116687_b33","series-title":"International Conference on Machine Learning","first-page":"8748","article-title":"Learning transferable visual models from natural language supervision","author":"Radford","year":"2021"},{"key":"10.1016\/j.knosys.2026.116687_b34","series-title":"European Conference on Computer Vision","first-page":"39","article-title":"Video graph transformer for video question answering","author":"Xiao","year":"2022"},{"key":"10.1016\/j.knosys.2026.116687_b35","doi-asserted-by":"crossref","unstructured":"Tsu-Jui Fu, Linjie Li, Zhe Gan, Kevin Lin, William Yang Wang, Lijuan Wang, Zicheng Liu, An empirical study of end-to-end video-language transformers with masked visual modeling, in: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 22898\u201322909.","DOI":"10.1109\/CVPR52729.2023.02193"},{"key":"10.1016\/j.knosys.2026.116687_b36","series-title":"Star: A benchmark for situated reasoning in real-world videos","author":"Wu","year":"2024"},{"key":"10.1016\/j.knosys.2026.116687_b37","doi-asserted-by":"crossref","unstructured":"Minghang Zheng, Yanjie Huang, Qingchao Chen, Yuxin Peng, Yang Liu, Weakly supervised temporal sentence grounding with gaussian-based contrastive proposal learning, in: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 15555\u201315564.","DOI":"10.1109\/CVPR52688.2022.01511"}],"container-title":["Knowledge-Based Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/api.elsevier.com\/content\/article\/PII:S0950705126014139?httpAccept=text\/xml","content-type":"text\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/api.elsevier.com\/content\/article\/PII:S0950705126014139?httpAccept=text\/plain","content-type":"text\/plain","content-version":"vor","intended-application":"text-mining"}],"deposited":{"date-parts":[[2026,8,18]],"date-time":"2026-08-18T06:27:36Z","timestamp":1787034456000},"score":1,"resource":{"primary":{"URL":"https:\/\/linkinghub.elsevier.com\/retrieve\/pii\/S0950705126014139"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,10]]},"references-count":37,"alternative-id":["S0950705126014139"],"URL":"https:\/\/doi.org\/10.1016\/j.knosys.2026.116687","relation":{},"ISSN":["0950-7051"],"issn-type":[{"value":"0950-7051","type":"print"}],"subject":[],"published":{"date-parts":[[2026,10]]},"assertion":[{"value":"Elsevier","name":"publisher","label":"This article is maintained by"},{"value":"TECL: Time-Equivariant Contrastive Learning for weakly-supervised Grounded Video Question Answering","name":"articletitle","label":"Article Title"},{"value":"Knowledge-Based Systems","name":"journaltitle","label":"Journal Title"},{"value":"https:\/\/doi.org\/10.1016\/j.knosys.2026.116687","name":"articlelink","label":"CrossRef DOI link to publisher maintained version"},{"value":"article","name":"content_type","label":"Content Type"},{"value":"\u00a9 2026 Elsevier B.V. All rights are reserved, including those for text and data mining, AI training, and similar technologies.","name":"copyright","label":"Copyright"}],"article-number":"116687"}}