{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,12,1]],"date-time":"2025-12-01T11:23:28Z","timestamp":1764588208836,"version":"3.41.0"},"reference-count":50,"publisher":"Association for Computing Machinery (ACM)","issue":"2s","license":[{"start":{"date-parts":[[2021,6,14]],"date-time":"2021-06-14T00:00:00Z","timestamp":1623628800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2021,6,21]]},"abstract":"<jats:p>Video captioning is a challenging task in the field of multimedia processing, which aims to generate informative natural language descriptions\/captions to describe video contents. Previous video captioning approaches mainly focused on capturing visual information in videos using an encoder-decoder structure to generate video captions. Recently, a new encoder-decoder-reconstructor structure was proposed for video captioning, which captured the information in both videos and captions. Based on this, this article proposes a novel multi-instance multi-label dual learning approach (MIMLDL) to generate video captions based on the encoder-decoder-reconstructor structure. Specifically, MIMLDL contains two modules: caption generation and video reconstruction modules. The caption generation module utilizes a lexical fully convolutional neural network (Lexical FCN) with a weakly supervised multi-instance multi-label learning mechanism to learn a translatable mapping between video regions and lexical labels to generate video captions. Then the video reconstruction module synthesizes visual sequences to reproduce raw videos using the outputs of the caption generation module. A dual learning mechanism fine-tunes the two modules according to the gap between the raw and the reproduced videos. Thus, our approach can minimize the semantic gap between raw videos and the generated captions by minimizing the differences between the reproduced and the raw visual sequences. Experimental results on a benchmark dataset demonstrate that MIMLDL can improve the accuracy of video captioning.<\/jats:p>","DOI":"10.1145\/3446792","type":"journal-article","created":{"date-parts":[[2021,6,14]],"date-time":"2021-06-14T12:55:42Z","timestamp":1623675342000},"page":"1-18","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":9,"title":["A Multi-instance Multi-label Dual Learning Approach for Video Captioning"],"prefix":"10.1145","volume":"17","author":[{"given":"Wanting","family":"Ji","sequence":"first","affiliation":[{"name":"School of Natural and Computational Sciences, Massey University, Auckland, New Zealand"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ruili","family":"Wang","sequence":"additional","affiliation":[{"name":"School of Natural and Computational Sciences, Massey University, Auckland, New Zealand"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2021,6,14]]},"reference":[{"volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 7622\u20137631","author":"Wang B.","key":"e_1_2_1_1_1","unstructured":"B. Wang , L. Ma , W. Zhang , and W. Liu . 2018. Reconstruction network for video captioning . In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 7622\u20137631 . B. Wang, L. Ma, W. Zhang, and W. Liu. 2018. Reconstruction network for video captioning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 7622\u20137631."},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00784"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2019.2924576"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1145\/3226037"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/3386725"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/3336495"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1023\/A:1020346032608"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2013.61"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.5555\/2886521.2886647"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11042-018-7142-7"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1007\/s00521-019-04609-8"},{"volume-title":"Proceedings of the Asian Conference on Pattern Recognition. 74\u201384","author":"Liu Z.","key":"e_1_2_1_12_1","unstructured":"Z. Liu , Z. Li , M. Zong , W. Ji , R. Wang , and Y. Tian . 2019. Spatiotemporal saliency based multi-stream networks for action recognition . In Proceedings of the Asian Conference on Pattern Recognition. 74\u201384 . Z. Liu, Z. Li, M. Zong, W. Ji, R. Wang, and Y. Tian. 2019. Spatiotemporal saliency based multi-stream networks for action recognition. In Proceedings of the Asian Conference on Pattern Recognition. 74\u201384."},{"volume-title":"Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 1494\u20131504","author":"Venugopalan S.","key":"e_1_2_1_13_1","unstructured":"S. Venugopalan , H. Xu , J. Donahue , M. Rohrbach , R. Mooney , and K. Saenko . 2015. Translating videos to natural language using deep recurrent neural networks . In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 1494\u20131504 . S. Venugopalan, H. Xu, J. Donahue, M. Rohrbach, R. Mooney, and K. Saenko. 2015. Translating videos to natural language using deep recurrent neural networks. In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 1494\u20131504."},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.515"},{"volume-title":"Proceedings of the 23rd International Conference on Pattern Recognition. 2924\u20132929","author":"Zhang C.","key":"e_1_2_1_15_1","unstructured":"C. Zhang and Y. Tian . 2016. Automatic video description generation via LSTM with joint two-stream encoding . In Proceedings of the 23rd International Conference on Pattern Recognition. 2924\u20132929 . C. Zhang and Y. Tian. 2016. Automatic video description generation via LSTM with joint two-stream encoding. In Proceedings of the 23rd International Conference on Pattern Recognition. 2924\u20132929."},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.512"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/3122865.3122867"},{"volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 1916\u20131924","author":"Shen Z.","key":"e_1_2_1_18_1","unstructured":"Z. Shen , J. Li , Z. Su , M. Li , Y. Chen , Y. Jiang , and X. Xue . 2017. Weakly supervised dense video captioning . In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 1916\u20131924 . Z. Shen, J. Li, Z. Su, M. Li, Y. Chen, Y. Jiang, and X. Xue. 2017. Weakly supervised dense video captioning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 1916\u20131924."},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1016\/S0004-3702(96)00034-3"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/3355612"},{"volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence. 12886\u201312893","author":"Zhang X.","key":"e_1_2_1_21_1","unstructured":"X. Zhang , H. Shi , C. Li , and P. Li . 2020. Multi-instance multi-label action recognition and localization based on spatio-temporal pre-trimming for untrimmed videos . In Proceedings of the AAAI Conference on Artificial Intelligence. 12886\u201312893 . X. Zhang, H. Shi, C. Li, and P. Li. 2020. Multi-instance multi-label action recognition and localization based on spatio-temporal pre-trimming for untrimmed videos. In Proceedings of the AAAI Conference on Artificial Intelligence. 12886\u201312893."},{"volume-title":"Proceedings of the IEEE International Conference on Computer Vision. 2718\u20132726","author":"Luo P.","key":"e_1_2_1_22_1","unstructured":"P. Luo , G. Wang , L. Lin , and X. Wang . 2017. Deep dual learning for semantic image segmentation . In Proceedings of the IEEE International Conference on Computer Vision. 2718\u20132726 . P. Luo, G. Wang, L. Lin, and X. Wang. 2017. Deep dual learning for semantic image segmentation. In Proceedings of the IEEE International Conference on Computer Vision. 2718\u20132726."},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.5555\/3172077.3172323"},{"volume-title":"Proceedings of the IEEE International Conference on Computer Vision. 2849\u20132857","author":"Yi Z.","key":"e_1_2_1_24_1","unstructured":"Z. Yi , H. Zhang , P. Tan , and M. Gong . 2017. Dualgan: Unsupervised dual learning for image-to-image translation . In Proceedings of the IEEE International Conference on Computer Vision. 2849\u20132857 . Z. Yi, H. Zhang, P. Tan, and M. Gong. 2017. Dualgan: Unsupervised dual learning for image-to-image translation. In Proceedings of the IEEE International Conference on Computer Vision. 2849\u20132857."},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.5555\/3305381.3305573"},{"volume-title":"Proceedings of the IEEE International Conference on Computer Vision. 2223\u20132232","author":"Zhu J.","key":"e_1_2_1_26_1","unstructured":"J. Zhu , T. Park , P. Isola , and A. A. Efros . 2017. Unpaired image-to-image translation using cycle-consistent adversarial networks . In Proceedings of the IEEE International Conference on Computer Vision. 2223\u20132232 . J. Zhu, T. Park, P. Isola, and A. A. Efros. 2017. Unpaired image-to-image translation using cycle-consistent adversarial networks. In Proceedings of the IEEE International Conference on Computer Vision. 2223\u20132232."},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.5555\/3157096.3157188"},{"volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence. 1\u20137.","author":"Wang Y.","key":"e_1_2_1_28_1","unstructured":"Y. Wang , Y. Xia , L. Zhao , J. Bian , T. Qin , G. Liu , and T. Liu . 2018. Dual transfer learning for neural machine translation with marginal distribution regularization . In Proceedings of the AAAI Conference on Artificial Intelligence. 1\u20137. Y. Wang, Y. Xia, L. Zhao, J. Bian, T. Qin, G. Liu, and T. Liu. 2018. Dual transfer learning for neural machine translation with marginal distribution regularization. In Proceedings of the AAAI Conference on Artificial Intelligence. 1\u20137."},{"volume-title":"Proceedings of the International Conference on Learning Representations. 1\u201314","author":"Lample G.","key":"e_1_2_1_29_1","unstructured":"G. Lample , A. Conneau , L. Denoyer , and M. A. Ranzato . 2018. Unsupervised machine translation using monolingual corpora only . In Proceedings of the International Conference on Learning Representations. 1\u201314 . G. Lample, A. Conneau, L. Denoyer, and M. A. Ranzato. 2018. Unsupervised machine translation using monolingual corpora only. In Proceedings of the International Conference on Learning Representations. 1\u201314."},{"volume-title":"Proceedings of the International Conference on Learning Representations. 1\u201312","author":"Artetxe M.","key":"e_1_2_1_30_1","unstructured":"M. Artetxe , G. Labaka , E. Agirre , and K. Cho . 2018. Unsupervised neural machine translation . In Proceedings of the International Conference on Learning Representations. 1\u201312 . M. Artetxe, G. Labaka, E. Agirre, and K. Cho. 2018. Unsupervised neural machine translation. In Proceedings of the International Conference on Learning Representations. 1\u201312."},{"volume-title":"Proceedings of the International Conference on Learning Representations. 1\u201315","author":"Wang Y.","key":"e_1_2_1_31_1","unstructured":"Y. Wang , Y. Xia , T. He , F. Tian , T. Qin , C. Zhai , and T. Liu . 2019. Multi-agent dual learning . In Proceedings of the International Conference on Learning Representations. 1\u201315 . Y. Wang, Y. Xia, T. He, F. Tian, T. Qin, C. Zhai, and T. Liu. 2019. Multi-agent dual learning. In Proceedings of the International Conference on Learning Representations. 1\u201315."},{"volume-title":"Proceedings of the International Conference on Learning Representations. 1\u201316","author":"Zhao Z.","key":"e_1_2_1_32_1","unstructured":"Z. Zhao , Y. Xia , T. Qin , and T. Liu . 2019. Dual learning: Theoretical study and algorithmic extensions . In Proceedings of the International Conference on Learning Representations. 1\u201316 . Z. Zhao, Y. Xia, T. Qin, and T. Liu. 2019. Dual learning: Theoretical study and algorithmic extensions. In Proceedings of the International Conference on Learning Representations. 1\u201316."},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.5555\/3305890.3306073"},{"volume-title":"Proceedings of the International Conference on Machine Learning. 5383\u20135392","author":"Xia Y.","key":"e_1_2_1_34_1","unstructured":"Y. Xia , X. Tan , F. Tian , T. Qin , N. Yu , and T. Liu . 2018. Model-level dual learning . In Proceedings of the International Conference on Machine Learning. 5383\u20135392 . Y. Xia, X. Tan, F. Tian, T. Qin, N. Yu, and T. Liu. 2018. Model-level dual learning. In Proceedings of the International Conference on Machine Learning. 5383\u20135392."},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1145\/3132847.3132920"},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.3115\/1073445.1073478"},{"volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 1473\u20131482","author":"Fang H.","key":"e_1_2_1_37_1","unstructured":"H. Fang , S. Gupta , F. Iandola , R. K. Srivastava , L. Deng , P. Doll\u00e1r , J. Gao et al. 2015. From captions to visual concepts and back . In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 1473\u20131482 . H. Fang, S. Gupta, F. Iandola, R. K. Srivastava, L. Deng, P. Doll\u00e1r, J. Gao et al. 2015. From captions to visual concepts and back. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 1473\u20131482."},{"volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 1\u201310","author":"Hendricks L. A.","key":"e_1_2_1_38_1","unstructured":"L. A. Hendricks , S. Venugopalan , M. Rohrbach , R. Mooney , K. Saenko , and T. Darrell . 2016. Deep compositional captioning: Describing novel object categories without paired training data . In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 1\u201310 . L. A. Hendricks, S. Venugopalan, M. Rohrbach, R. Mooney, K. Saenko, and T. Darrell. 2016. Deep compositional captioning: Describing novel object categories without paired training data. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 1\u201310."},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.5555\/647232.719594"},{"volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 3090\u20133098","author":"Gygli M.","key":"e_1_2_1_40_1","unstructured":"M. Gygli , H. Grabner , and L. V. Gool . 2015. Video summarization by learning submodular mixtures of objectives . In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 3090\u20133098 . M. Gygli, H. Grabner, and L. V. Gool. 2015. Video summarization by learning submodular mixtures of objectives. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 3090\u20133098."},{"volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 5288\u20135296","author":"Xu J.","key":"e_1_2_1_41_1","unstructured":"J. Xu , T. Mei , T. Yao , and Y. Rui . 2016. MSR-VTT: A large video description dataset for bridging video and language . In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 5288\u20135296 . J. Xu, T. Mei, T. Yao, and Y. Rui. 2016. MSR-VTT: A large video description dataset for bridging video and language. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 5288\u20135296."},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.5555\/3295222.3295349"},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.5555\/1626355.1626389"},{"key":"e_1_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.3115\/1073083.1073135"},{"volume-title":"Text Summarization Branches Out","author":"Lin C.","key":"e_1_2_1_45_1","unstructured":"C. Lin . 2004. Rouge: A package for automatic evaluation of summaries . In Text Summarization Branches Out . Association for Computational Linguistics , 74\u201381. C. Lin. 2004. Rouge: A package for automatic evaluation of summaries. In Text Summarization Branches Out. Association for Computational Linguistics, 74\u201381."},{"key":"e_1_2_1_46_1","unstructured":"X. Chen H. Fang T. Lin R. Vedantam S. Gupta P. Doll\u00e1r and C. L. Zitnick. 2015. Microsoft COCO captions: Data collection and evaluation server. arXiv preprint arXiv:1504.00325. (2015). X. Chen H. Fang T. Lin R. Vedantam S. Gupta P. Doll\u00e1r and C. L. Zitnick. 2015. Microsoft COCO captions: Data collection and evaluation server. arXiv preprint arXiv:1504.00325. (2015)."},{"key":"e_1_2_1_47_1","article-title":"BR2Net: Defocus blur detection via bidirectional channel attention residual refining network","volume":"10","author":"Tang C.","year":"2020","unstructured":"C. Tang , X. Liu , S. An , and P. Wang . 2020 . BR2Net: Defocus blur detection via bidirectional channel attention residual refining network . IEEE Trans. Multim. DOI : 10 .1109\/TMM.2020.2985541. 10.1109\/TMM.2020.2985541 C. Tang, X. Liu, S. An, and P. Wang. 2020. BR2Net: Defocus blur detection via bidirectional channel attention residual refining network. IEEE Trans. Multim. DOI: 10.1109\/TMM.2020.2985541.","journal-title":"IEEE Trans. Multim. DOI"},{"key":"e_1_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2019.2909860"},{"key":"e_1_2_1_49_1","doi-asserted-by":"crossref","first-page":"1083","DOI":"10.1109\/TNNLS.2013.2287275","article-title":"Global and local structure preservation for feature selection","volume":"25","author":"Liu X.","year":"2013","unstructured":"X. Liu , L. Wang , J. Zhang , J. Yin , and H. Liu . 2013 . Global and local structure preservation for feature selection . IEEE Trans. Neural Netw. Learn. Syst. 25 , 6 (2013), 1083 \u2013 1095 . X. Liu, L. Wang, J. Zhang, J. Yin, and H. Liu. 2013. Global and local structure preservation for feature selection. IEEE Trans. Neural Netw. Learn. Syst. 25, 6 (2013), 1083\u20131095.","journal-title":"IEEE Trans. Neural Netw. Learn. Syst."},{"key":"e_1_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.1613\/jair.1.11338"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3446792","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3446792","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T20:47:31Z","timestamp":1750193251000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3446792"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,6,14]]},"references-count":50,"journal-issue":{"issue":"2s","published-print":{"date-parts":[[2021,6,21]]}},"alternative-id":["10.1145\/3446792"],"URL":"https:\/\/doi.org\/10.1145\/3446792","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"type":"print","value":"1551-6857"},{"type":"electronic","value":"1551-6865"}],"subject":[],"published":{"date-parts":[[2021,6,14]]},"assertion":[{"value":"2020-07-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2021-01-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2021-06-14","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}