{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,2]],"date-time":"2026-06-02T06:14:02Z","timestamp":1780380842467,"version":"3.54.1"},"publisher-location":"New York, NY, USA","reference-count":52,"publisher":"ACM","license":[{"start":{"date-parts":[[2018,10,15]],"date-time":"2018-10-15T00:00:00Z","timestamp":1539561600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2018,10,15]]},"DOI":"10.1145\/3240508.3240527","type":"proceedings-article","created":{"date-parts":[[2018,10,18]],"date-time":"2018-10-18T17:52:08Z","timestamp":1539885128000},"page":"54-62","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":31,"title":["Fast Parameter Adaptation for Few-shot Image Captioning and Visual Question Answering"],"prefix":"10.1145","author":[{"given":"Xuanyi","family":"Dong","sequence":"first","affiliation":[{"name":"Southern University of Science and Technology &amp; University of Technology Sydney, Sydney, Australia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Linchao","family":"Zhu","sequence":"additional","affiliation":[{"name":"Southern University of Science and Technology &amp; University of Technology Sydney, Sydney, Australia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"De","family":"Zhang","sequence":"additional","affiliation":[{"name":"China Electronics Technology Group Corporation, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yi","family":"Yang","sequence":"additional","affiliation":[{"name":"Southern University of Science and Technology &amp; University of Technology Sydney, Sydney, Australia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Fei","family":"Wu","sequence":"additional","affiliation":[{"name":"Zhejiang University, Zhejiang, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2018,10,15]]},"reference":[{"key":"e_1_3_2_1_1_1","doi-asserted-by":"crossref","unstructured":"Lisa Anne Hendricks Subhashini Venugopalan Marcus Rohrbach Raymond Mooney Kate Saenko Trevor Darrell Junhua Mao Jonathan Huang Alexander Toshev Oana Camburu etal 2016. Deep Compositional Captioning: Describing Novel Object Categories without Paired Training Data. In CVPR .  Lisa Anne Hendricks Subhashini Venugopalan Marcus Rohrbach Raymond Mooney Kate Saenko Trevor Darrell Junhua Mao Jonathan Huang Alexander Toshev Oana Camburu et al. 2016. Deep Compositional Captioning: Describing Novel Object Categories without Paired Training Data. In CVPR .","DOI":"10.1109\/CVPR.2016.8"},{"key":"e_1_3_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.279"},{"key":"e_1_3_2_1_3_1","unstructured":"Shaojie Bai J Zico Kolter and Vladlen Koltun. 2018. An Empirical Evaluation of Generic Convolutional and Recurrent Networks for Sequence Modeling. arXiv preprint arXiv:1803.01271 (2018).  Shaojie Bai J Zico Kolter and Vladlen Koltun. 2018. An Empirical Evaluation of Generic Convolutional and Recurrent Networks for Sequence Modeling. arXiv preprint arXiv:1803.01271 (2018)."},{"key":"e_1_3_2_1_4_1","unstructured":"Kyunghyun Cho Bart van Merrienboer Caglar Gulcehre Dzmitry Bahdanau Fethi Bougares Holger Schwenk and Yoshua Bengio. 2014. Learning Phrase Representations using RNN Encoder--Decoder for Statistical Machine Translation. In EMNLP .  Kyunghyun Cho Bart van Merrienboer Caglar Gulcehre Dzmitry Bahdanau Fethi Bougares Holger Schwenk and Yoshua Bengio. 2014. Learning Phrase Representations using RNN Encoder--Decoder for Statistical Machine Translation. In EMNLP ."},{"key":"e_1_3_2_1_5_1","doi-asserted-by":"crossref","unstructured":"Xuanyi Dong Junshi Huang Yi Yang and Shuicheng Yan. 2017a. More Is Less: A More Complicated Network With Less Inference Complexity. In CVPR .  Xuanyi Dong Junshi Huang Yi Yang and Shuicheng Yan. 2017a. More Is Less: A More Complicated Network With Less Inference Complexity. In CVPR .","DOI":"10.1109\/CVPR.2017.205"},{"key":"e_1_3_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/3123266.3123455"},{"key":"e_1_3_2_1_7_1","doi-asserted-by":"crossref","unstructured":"Xuanyi Dong Shoou-I Yu Xinshuo Weng Shih-En Wei Yi Yang and Yaser Sheikh. 2018a. Supervision-by-Registration: An Unsupervised Approach to Improve the Precision of Facial Landmark Detectors. In CVPR .  Xuanyi Dong Shoou-I Yu Xinshuo Weng Shih-En Wei Yi Yang and Yaser Sheikh. 2018a. Supervision-by-Registration: An Unsupervised Approach to Improve the Precision of Facial Landmark Detectors. In CVPR .","DOI":"10.1109\/CVPR.2018.00045"},{"key":"e_1_3_2_1_8_1","unstructured":"Xuanyi Dong Liang Zheng Fan Ma Yi Yang and Deyu Meng. 2018b. Few-Example Object Detection with Model Communication. IEEE Transactions on Pattern Analysis and Machine Intelligence (2018).  Xuanyi Dong Liang Zheng Fan Ma Yi Yang and Deyu Meng. 2018b. Few-Example Object Detection with Model Communication. IEEE Transactions on Pattern Analysis and Machine Intelligence (2018)."},{"key":"e_1_3_2_1_9_1","doi-asserted-by":"crossref","unstructured":"Ali Farhadi Mohsen Hejrati Mohammad Amin Sadeghi Peter Young Cyrus Rashtchian Julia Hockenmaier and David Forsyth. 2010. Every Picture Tells a Story: Generating Sentences from Images. In ECCV .   Ali Farhadi Mohsen Hejrati Mohammad Amin Sadeghi Peter Young Cyrus Rashtchian Julia Hockenmaier and David Forsyth. 2010. Every Picture Tells a Story: Generating Sentences from Images. In ECCV .","DOI":"10.1007\/978-3-642-15561-1_2"},{"key":"e_1_3_2_1_10_1","unstructured":"Chelsea Finn Pieter Abbeel and Sergey Levine. 2017. Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networks. In ICML .  Chelsea Finn Pieter Abbeel and Sergey Levine. 2017. Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networks. In ICML ."},{"key":"e_1_3_2_1_11_1","doi-asserted-by":"crossref","unstructured":"Yash Goyal Tejas Khot Douglas Summers-Stay Dhruv Batra and Devi Parikh. 2017. Making the V in VQA matter: Elevating the role of image understanding in Visual Question Answering. In CVPR .  Yash Goyal Tejas Khot Douglas Summers-Stay Dhruv Batra and Devi Parikh. 2017. Making the V in VQA matter: Elevating the role of image understanding in Visual Question Answering. In CVPR .","DOI":"10.1109\/CVPR.2017.670"},{"key":"e_1_3_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.123"},{"key":"e_1_3_2_1_13_1","unstructured":"Kaiming He Xiangyu Zhang Shaoqing Ren and Jian Sun. 2016. Deep Residual Learning for Image Recognition. In CVPR .  Kaiming He Xiangyu Zhang Shaoqing Ren and Jian Sun. 2016. Deep Residual Learning for Image Recognition. In CVPR ."},{"key":"e_1_3_2_1_14_1","doi-asserted-by":"crossref","unstructured":"Sepp Hochreiter and J\u00fcrgen Schmidhuber. 1997. Long Short-Term Memory. Neural Computation (1997).  Sepp Hochreiter and J\u00fcrgen Schmidhuber. 1997. Long Short-Term Memory. Neural Computation (1997).","DOI":"10.1162\/neco.1997.9.8.1735"},{"key":"e_1_3_2_1_15_1","unstructured":"Ronghang Hu Jacob Andreas Marcus Rohrbach Trevor Darrell and Kate Saenko. 2017. Learning to Reason: End-To-End Module Networks for Visual Question Answering. In ICCV .  Ronghang Hu Jacob Andreas Marcus Rohrbach Trevor Darrell and Kate Saenko. 2017. Learning to Reason: End-To-End Module Networks for Visual Question Answering. In ICCV ."},{"key":"e_1_3_2_1_16_1","doi-asserted-by":"crossref","unstructured":"Gao Huang Zhuang Liu Kilian Q Weinberger and Laurens van der Maaten. 2017. Densely Connected Convolutional Networks. In CVPR .  Gao Huang Zhuang Liu Kilian Q Weinberger and Laurens van der Maaten. 2017. Densely Connected Convolutional Networks. In CVPR .","DOI":"10.1109\/CVPR.2017.243"},{"key":"e_1_3_2_1_17_1","doi-asserted-by":"crossref","unstructured":"Xun Huang and Serge Belongie. 2017. Arbitrary Style Transfer in Real-Time With Adaptive Instance Normalization. In ICCV .  Xun Huang and Serge Belongie. 2017. Arbitrary Style Transfer in Real-Time With Adaptive Instance Normalization. In ICCV .","DOI":"10.1109\/ICCV.2017.167"},{"key":"e_1_3_2_1_18_1","volume-title":"Adam: A Method for Stochastic Optimization. In ICLR .","author":"Kingma Diederik P","year":"2015"},{"key":"e_1_3_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-016-0981-7"},{"key":"e_1_3_2_1_20_1","unstructured":"Alex Krizhevsky Ilya Sutskever and Geoffrey E Hinton. 2012. ImageNet Classification with Deep Convolutional Neural Networks. In NIPS .   Alex Krizhevsky Ilya Sutskever and Geoffrey E Hinton. 2012. ImageNet Classification with Deep Convolutional Neural Networks. In NIPS ."},{"key":"e_1_3_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2012.162"},{"key":"e_1_3_2_1_22_1","doi-asserted-by":"crossref","unstructured":"Angeliki Lazaridou Marco Marelli and Marco Baroni. 2017. Multimodal Word Meaning Induction from Minimal Exposure to Natural Text. Cognitive Science (2017).  Angeliki Lazaridou Marco Marelli and Marco Baroni. 2017. Multimodal Word Meaning Induction from Minimal Exposure to Natural Text. Cognitive Science (2017).","DOI":"10.1111\/cogs.12481"},{"key":"e_1_3_2_1_23_1","doi-asserted-by":"crossref","unstructured":"Colin Lea Michael D. Flynn Rene Vidal Austin Reiter and Gregory D. Hager. 2017. Temporal Convolutional Networks for Action Segmentation and Detection. In CVPR .  Colin Lea Michael D. Flynn Rene Vidal Austin Reiter and Gregory D. Hager. 2017. Temporal Convolutional Networks for Action Segmentation and Detection. In CVPR .","DOI":"10.1109\/CVPR.2017.113"},{"key":"e_1_3_2_1_24_1","unstructured":"Tsung-Yi Lin Michael Maire Serge Belongie James Hays Pietro Perona Deva Ramanan Piotr Doll\u00e1r and C Lawrence Zitnick. 2014. Microsoft COCO: Common objects in Context. In ECCV .  Tsung-Yi Lin Michael Maire Serge Belongie James Hays Pietro Perona Deva Ramanan Piotr Doll\u00e1r and C Lawrence Zitnick. 2014. Microsoft COCO: Common objects in Context. In ECCV ."},{"key":"e_1_3_2_1_25_1","doi-asserted-by":"crossref","unstructured":"Yu Liu Fangyin Wei Jing Shao Lu Sheng Junjie Yan and Xiaogang Wang. 2018. Exploring Disentangled Feature Representation beyond Face Identification. In CVPR .  Yu Liu Fangyin Wei Jing Shao Lu Sheng Junjie Yan and Xiaogang Wang. 2018. Exploring Disentangled Feature Representation beyond Face Identification. In CVPR .","DOI":"10.1109\/CVPR.2018.00222"},{"key":"e_1_3_2_1_26_1","unstructured":"Jiasen Lu Jianwei Yang Dhruv Batra and Devi Parikh. 2018. Neural Baby Talk. In CVPR .  Jiasen Lu Jianwei Yang Dhruv Batra and Devi Parikh. 2018. Neural Baby Talk. In CVPR ."},{"key":"e_1_3_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.291"},{"key":"e_1_3_2_1_28_1","volume-title":"Midge: Generating Image Descriptions from Computer Vision Detections. In EACL .","author":"Mitchell Margaret","year":"2012"},{"key":"e_1_3_2_1_29_1","unstructured":"Tsendsuren Munkhdalai and Hong Yu. 2017. Meta Networks. In ICML .  Tsendsuren Munkhdalai and Hong Yu. 2017. Meta Networks. In ICML ."},{"key":"e_1_3_2_1_30_1","unstructured":"Hyeonwoo Noh Paul Hongsuck Seo and Bohyung Han. 2016. Image Question Answering using Convolutional Neural Network with Dynamic Parameter Prediction. In CVPR .  Hyeonwoo Noh Paul Hongsuck Seo and Bohyung Han. 2016. Image Question Answering using Convolutional Neural Network with Dynamic Parameter Prediction. In CVPR ."},{"key":"e_1_3_2_1_31_1","doi-asserted-by":"crossref","unstructured":"Ethan Perez Florian Strub Harm De Vries Vincent Dumoulin and Aaron Courville. 2018. FiLM: Visual Reasoning with a General Conditioning Layer. In AAAI .  Ethan Perez Florian Strub Harm De Vries Vincent Dumoulin and Aaron Courville. 2018. FiLM: Visual Reasoning with a General Conditioning Layer. In AAAI .","DOI":"10.1609\/aaai.v32i1.11671"},{"key":"e_1_3_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1016\/S0893-6080(98)00116-6"},{"key":"e_1_3_2_1_33_1","doi-asserted-by":"crossref","unstructured":"Santhosh K Ramakrishnan Ambar Pal Gaurav Sharma and Anurag Mittal. 2017. An Empirical Evaluation of Visual Question Answering for Novel Objects. In CVPR .  Santhosh K Ramakrishnan Ambar Pal Gaurav Sharma and Anurag Mittal. 2017. An Empirical Evaluation of Visual Question Answering for Novel Objects. In CVPR .","DOI":"10.1109\/CVPR.2017.773"},{"key":"e_1_3_2_1_34_1","unstructured":"Sachin Ravi and Hugo Larochelle. 2017. Optimization as a Model for Few-shot Learning. In ICLR .  Sachin Ravi and Hugo Larochelle. 2017. Optimization as a Model for Few-shot Learning. In ICLR ."},{"key":"e_1_3_2_1_35_1","unstructured":"Mengye Ren Ryan Kiros and Richard Zemel. 2015. Exploring Models and Data for Image Question Answering. In NIPS .   Mengye Ren Ryan Kiros and Richard Zemel. 2015. Exploring Models and Data for Image Question Answering. In NIPS ."},{"key":"e_1_3_2_1_36_1","doi-asserted-by":"crossref","unstructured":"Kevin J Shih Saurabh Singh and Derek Hoiem. 2016. Where to Look: Focus Regions for Visual Question Answering. In CVPR .  Kevin J Shih Saurabh Singh and Derek Hoiem. 2016. Where to Look: Focus Regions for Visual Question Answering. In CVPR .","DOI":"10.1109\/CVPR.2016.499"},{"key":"e_1_3_2_1_37_1","unstructured":"Jake Snell Kevin Swersky and Richard Zemel. 2017. Prototypical Networks for Few-shot Learning. In NIPS .  Jake Snell Kevin Swersky and Richard Zemel. 2017. Prototypical Networks for Few-shot Learning. In NIPS ."},{"key":"e_1_3_2_1_38_1","unstructured":"Damien Teney and Anton van den Hengel. 2017. Visual Question Answering as a Meta Learning Task. arXiv preprint arXiv:1711.08105 (2017).  Damien Teney and Anton van den Hengel. 2017. Visual Question Answering as a Meta Learning Task. arXiv preprint arXiv:1711.08105 (2017)."},{"key":"e_1_3_2_1_39_1","doi-asserted-by":"crossref","unstructured":"Damien Teney Lingqiao Liu and Anton van den Hengel. 2017. Graph-Structured Representations for Visual Question Answering. In CVPR .  Damien Teney Lingqiao Liu and Anton van den Hengel. 2017. Graph-Structured Representations for Visual Question Answering. In CVPR .","DOI":"10.1109\/CVPR.2017.344"},{"key":"e_1_3_2_1_40_1","unstructured":"Oriol Vinyals Charles Blundell Tim Lillicrap Daan Wierstra etal2016. Matching Networks for One Shot Learning. In NIPS .   Oriol Vinyals Charles Blundell Tim Lillicrap Daan Wierstra et al.2016. Matching Networks for One Shot Learning. In NIPS ."},{"key":"e_1_3_2_1_41_1","unstructured":"Su Wang Stephen Roller and Katrin Erk. 2017. Distributional Modeling on a Diet: One-shot Word Learning from Text only. In IJCNLP .  Su Wang Stephen Roller and Katrin Erk. 2017. Distributional Modeling on a Diet: One-shot Word Learning from Text only. In IJCNLP ."},{"key":"e_1_3_2_1_42_1","doi-asserted-by":"crossref","unstructured":"Yu Wu Yutian Lin Xuanyi Dong Yan Yan Wanli Ouyang and Yi Yang. 2018a. Exploit the Unknown Gradually: One-Shot Video-Based Person Re-Identification by Stepwise Learning. In CVPR .  Yu Wu Yutian Lin Xuanyi Dong Yan Yan Wanli Ouyang and Yi Yang. 2018a. Exploit the Unknown Gradually: One-Shot Video-Based Person Re-Identification by Stepwise Learning. In CVPR .","DOI":"10.1109\/CVPR.2018.00543"},{"key":"e_1_3_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1145\/3240508.3240640"},{"key":"e_1_3_2_1_44_1","unstructured":"Kelvin Xu Jimmy Ba Ryan Kiros Kyunghyun Cho Aaron Courville Ruslan Salakhudinov Rich Zemel and Yoshua Bengio. 2015. Show Attend and Tell: Neural Image Caption Generation with Visual Attention. In ICML .   Kelvin Xu Jimmy Ba Ryan Kiros Kyunghyun Cho Aaron Courville Ruslan Salakhudinov Rich Zemel and Yoshua Bengio. 2015. Show Attend and Tell: Neural Image Caption Generation with Visual Attention. In ICML ."},{"key":"e_1_3_2_1_45_1","unstructured":"Zhongwen Xu Linchao Zhu and Yi Yang. 2017. Few-shot object recognition from machine-labeled web images. In CVPR .  Zhongwen Xu Linchao Zhu and Yi Yang. 2017. Few-shot object recognition from machine-labeled web images. In CVPR ."},{"key":"e_1_3_2_1_46_1","doi-asserted-by":"crossref","unstructured":"Zichao Yang Xiaodong He Jianfeng Gao Li Deng and Alex Smola. 2016. Stacked Attention Networks for Image Question Answering. In CVPR .  Zichao Yang Xiaodong He Jianfeng Gao Li Deng and Alex Smola. 2016. Stacked Attention Networks for Image Question Answering. In CVPR .","DOI":"10.1109\/CVPR.2016.10"},{"key":"e_1_3_2_1_47_1","doi-asserted-by":"crossref","unstructured":"Ting Yao Yingwei Pan Yehao Li and Tao Mei. 2017. Incorporating Copying Mechanism in Image Captioning for Learning Novel Objects. In CVPR .  Ting Yao Yingwei Pan Yehao Li and Tao Mei. 2017. Incorporating Copying Mechanism in Image Captioning for Learning Novel Objects. In CVPR .","DOI":"10.1109\/CVPR.2017.559"},{"key":"e_1_3_2_1_48_1","unstructured":"Quanzeng You Hailin Jin Zhaowen Wang Chen Fang and Jiebo Luo. 2016. Image Captioning with Semantic Attention. In CVPR .  Quanzeng You Hailin Jin Zhaowen Wang Chen Fang and Jiebo Luo. 2016. Image Captioning with Semantic Attention. In CVPR ."},{"key":"e_1_3_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.283"},{"key":"e_1_3_2_1_50_1","unstructured":"Yan Zhang Jonathon Hare and Adam Pr\u00fcgel-Bennett. 2018. Learning to Count Objects in Natural Images for Visual Question Answering. In ICLR .  Yan Zhang Jonathon Hare and Adam Pr\u00fcgel-Bennett. 2018. Learning to Count Objects in Natural Images for Visual Question Answering. In ICLR ."},{"key":"e_1_3_2_1_51_1","doi-asserted-by":"crossref","unstructured":"Zhun Zhong Liang Zheng Zhedong Zheng Shaozi Li and Yi Yang. 2018. Camera Style Adaptation for Person Re-Identification. In CVPR .  Zhun Zhong Liang Zheng Zhedong Zheng Shaozi Li and Yi Yang. 2018. Camera Style Adaptation for Person Re-Identification. In CVPR .","DOI":"10.1109\/CVPR.2018.00541"},{"key":"e_1_3_2_1_52_1","doi-asserted-by":"crossref","unstructured":"Yuke Zhu Oliver Groth Michael Bernstein and Li Fei-Fei. 2016. Visual7w: Grounded question answering in images. In CVPR .  Yuke Zhu Oliver Groth Michael Bernstein and Li Fei-Fei. 2016. Visual7w: Grounded question answering in images. In CVPR .","DOI":"10.1109\/CVPR.2016.540"}],"event":{"name":"MM '18: ACM Multimedia Conference","location":"Seoul Republic of Korea","acronym":"MM '18","sponsor":["SIGMM ACM Special Interest Group on Multimedia"]},"container-title":["Proceedings of the 26th ACM international conference on Multimedia"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3240508.3240527","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3240508.3240527","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T00:44:01Z","timestamp":1750207441000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3240508.3240527"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2018,10,15]]},"references-count":52,"alternative-id":["10.1145\/3240508.3240527","10.1145\/3240508"],"URL":"https:\/\/doi.org\/10.1145\/3240508.3240527","relation":{},"subject":[],"published":{"date-parts":[[2018,10,15]]},"assertion":[{"value":"2018-10-15","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}