{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,8]],"date-time":"2026-07-08T11:51:54Z","timestamp":1783511514701,"version":"3.55.0"},"publisher-location":"California","reference-count":0,"publisher":"International Joint Conferences on Artificial Intelligence Organization","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2023,8]]},"abstract":"<jats:p>Fine-tuning large vision-language models is a challenging task. Prompt tuning approaches have been introduced to learn fixed textual or visual prompts while freezing the pre-trained model in downstream tasks. Despite the effectiveness of prompt tuning, what do those learnable prompts learn remains unexplained. In this work, we explore whether prompts in the fine-tuning can learn knowledge-aware prompts from the pre-training, by designing two different sets of prompts in pre-training and fine-tuning phases respectively. Specifically, we present a Video-Language Prompt tuning (VL-Prompt) approach for video captioning, which first efficiently pre-train a video-language model to extract key information (e.g., actions and objects) with flexibly generated Knowledge-Aware Prompt (KAP). Then, we design a Video-Language Prompt (VLP) to transfer the knowledge from the knowledge-aware prompts and fine-tune the model to generate full captions. Experimental results show the superior performance of our approach over several state-of-the-art baselines. We further demonstrate that the video-language prompts are well learned from the knowledge-aware prompts.<\/jats:p>","DOI":"10.24963\/ijcai.2023\/180","type":"proceedings-article","created":{"date-parts":[[2023,8,11]],"date-time":"2023-08-11T08:31:30Z","timestamp":1691742690000},"page":"1622-1630","source":"Crossref","is-referenced-by-count":37,"title":["Prompt Learns Prompt: Exploring Knowledge-Aware Generative Prompt Collaboration For Video Captioning"],"prefix":"10.24963","author":[{"given":"Liqi","family":"Yan","sequence":"first","affiliation":[{"name":"Fudan University"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Cheng","family":"Han","sequence":"additional","affiliation":[{"name":"Rochester Institute of Technology"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Zenglin","family":"Xu","sequence":"additional","affiliation":[{"name":"Harbin Institute of Technology, Shenzhen"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Dongfang","family":"Liu","sequence":"additional","affiliation":[{"name":"Rochester Institute of Technology"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Qifan","family":"Wang","sequence":"additional","affiliation":[{"name":"Meta AI"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"10584","event":{"name":"Thirty-Second International Joint Conference on Artificial Intelligence {IJCAI-23}","theme":"Artificial Intelligence","location":"Macau, SAR China","acronym":"IJCAI-2023","number":"32","sponsor":["International Joint Conferences on Artificial Intelligence Organization (IJCAI)"],"start":{"date-parts":[[2023,8,19]]},"end":{"date-parts":[[2023,8,25]]}},"container-title":["Proceedings of the Thirty-Second International Joint Conference on Artificial Intelligence"],"original-title":[],"deposited":{"date-parts":[[2023,8,11]],"date-time":"2023-08-11T08:40:03Z","timestamp":1691743203000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.ijcai.org\/proceedings\/2023\/180"}},"subtitle":[],"proceedings-subject":"Artificial Intelligence Research Articles","short-title":[],"issued":{"date-parts":[[2023,8]]},"references-count":0,"URL":"https:\/\/doi.org\/10.24963\/ijcai.2023\/180","relation":{},"subject":[],"published":{"date-parts":[[2023,8]]}}}