{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,13]],"date-time":"2026-01-13T23:19:50Z","timestamp":1768346390497,"version":"3.49.0"},"reference-count":49,"publisher":"Association for Computing Machinery (ACM)","issue":"6","license":[{"start":{"date-parts":[[2024,3,8]],"date-time":"2024-03-08T00:00:00Z","timestamp":1709856000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62076100, 62072188"],"award-info":[{"award-number":["62076100, 62072188"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Fundamental Research Funds for the Central Universities, SCUT","award":["x2rjD2230080"],"award-info":[{"award-number":["x2rjD2230080"]}]},{"DOI":"10.13039\/501100012245","name":"Science and Technology Planning Project of Guangdong Province","doi-asserted-by":"crossref","award":["2020B0101100002"],"award-info":[{"award-number":["2020B0101100002"]}],"id":[{"id":"10.13039\/501100012245","id-type":"DOI","asserted-by":"crossref"}]},{"name":"CAAI-Huawei MindSpore Open FundCCF-Zhipu AI Large Model Fund"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2024,6,30]]},"abstract":"<jats:p>Visual question generation task aims at generating high-quality questions about a given image. To make this tak applicable to various scenarios, e.g., the growing demand for exams, it is important to generate diverse questions. The existing methods for this task control diverse question generation based on different question types, e.g., \u201cwhat\u201d and \u201cwhen.\u201d Although different question types lead to description diversity, they cannot guarantee semantic diversity when asking the same objects. Research in the field of psychology shows that humans pay attention to different objects in an image based on their preferences, which is beneficial to constructing semantically diverse questions. According to the research, we propose a multi-selector visual question generation (MS-VQG) model that aims to focus on different objects to generate diverse questions. Specifically, our MS-VQG model employs multiple selectors to imitate different humans to select different objects in a given image. Based on these different selected objects, our MS-VQG model can generate diverse questions corresponding to each selector. Extensive experiments on two datasets show that our proposed model outperforms the baselines in generating diverse questions.<\/jats:p>","DOI":"10.1145\/3640014","type":"journal-article","created":{"date-parts":[[2024,1,15]],"date-time":"2024-01-15T11:27:44Z","timestamp":1705318064000},"page":"1-22","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":4,"title":["Diverse Visual Question Generation Based on Multiple Objects Selection"],"prefix":"10.1145","volume":"20","author":[{"ORCID":"https:\/\/orcid.org\/0009-0001-2082-6649","authenticated-orcid":false,"given":"Wenhao","family":"Fang","sequence":"first","affiliation":[{"name":"School of Software Engineering, South China University of Technology, Guangzhou, China and Key Laboratory of Big Data and Intelligent Robot (SCUT), MOE of China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6833-7879","authenticated-orcid":false,"given":"Jiayuan","family":"Xie","sequence":"additional","affiliation":[{"name":"School of Software Engineering, South China University of Technology, Guangzhou, China and Key Laboratory of Big Data and Intelligent Robot (SCUT), MOE of China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-5104-3031","authenticated-orcid":false,"given":"Hongfei","family":"Liu","sequence":"additional","affiliation":[{"name":"School of Software Engineering, South China University of Technology, Guangzhou, China and Key Laboratory of Big Data and Intelligent Robot (SCUT), MOE of China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8064-1577","authenticated-orcid":false,"given":"Jiali","family":"Chen","sequence":"additional","affiliation":[{"name":"School of Software Engineering, South China University of Technology, Guangzhou, China and Key Laboratory of Big Data and Intelligent Robot (SCUT), MOE of China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1767-789X","authenticated-orcid":false,"given":"Yi","family":"Cai","sequence":"additional","affiliation":[{"name":"School of Software Engineering, South China University of Technology, Guangzhou, China and Key Laboratory of Big Data and Intelligent Robot (SCUT), MOE of China and Peng Cheng Laboratory, Shenzhen, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2024,3,8]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00636"},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00436"},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.279"},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.3115\/1225403.1225421"},{"key":"e_1_3_1_6_2","first-page":"1929","volume-title":"Proceedings of the Annual Conference on Neural Information Processing Systems (NIPS \u201919)","author":"Chen Fuhai","year":"2019","unstructured":"Fuhai Chen, Rongrong Ji, Jiayi Ji, Xiaoshuai Sun, Baochang Zhang, Xuri Ge, Yongjian Wu, Feiyue Huang, and Yan Wang. 2019. Variational structured semantic inference for diverse image captioning. In Proceedings of the Annual Conference on Neural Information Processing Systems (NIPS \u201919). 1929\u20131939."},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.1145\/3581783.3612536"},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D19-1308"},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/D14-1179"},{"key":"e_1_3_1_10_2","volume-title":"Proceedings of the International Conference on Learning Representations (ICLR \u201914)","author":"Eigen David","year":"2014","unstructured":"David Eigen, Marc\u2019Aurelio Ranzato, and Ilya Sutskever. 2014. Learning factored representations in a deep mixture of experts. In Proceedings of the International Conference on Learning Representations (ICLR \u201914)."},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","DOI":"10.24963\/ijcai.2018\/563"},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.1145\/3282469"},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.670"},{"key":"e_1_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.1162\/neco.1997.9.8.1735"},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.642"},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.1162\/neco.1991.3.1.79"},{"key":"e_1_3_1_17_2","volume-title":"The Principles of Psychology","author":"James William","year":"1890","unstructured":"William James. 1890. The Principles of Psychology. Vol. 1."},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.215"},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","DOI":"10.1080\/02796015.1981.12084904"},{"key":"e_1_3_1_20_2","volume-title":"Proceedings of the International Conference on Learning Representations (ICLR \u201915)","author":"Kingma Diederik P.","year":"2015","unstructured":"Diederik P. Kingma and Jimmy Ba. 2015. Adam: A method for stochastic optimization. In Proceedings of the International Conference on Learning Representations (ICLR \u201915)."},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00211"},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-016-0981-7"},{"key":"e_1_3_1_23_2","volume-title":"Proceedings of the IEEE International Conference on Consumer Electronics (ICCE \u201904)","author":"Kunichika Hidenobu","year":"2004","unstructured":"Hidenobu Kunichika, Tomoki Katayama, Tsukasa Hirashima, and Akira Takeuchi. 2004. Automated question generation methods for intelligent English learning systems and its evaluation. In Proceedings of the IEEE International Conference on Consumer Electronics (ICCE \u201904)."},{"key":"e_1_3_1_24_2","first-page":"2119","volume-title":"Proceedings of the Annual Conference on Neural Information Processing Systems (NIPS \u201916)","author":"Lee Stefan","year":"2016","unstructured":"Stefan Lee, Senthil Purushwalkam, Michael Cogswell, Viresh Ranjan, David J. Crandall, and Dhruv Batra. 2016. Stochastic multiple choice learning for training diverse deep ensembles. In Proceedings of the Annual Conference on Neural Information Processing Systems (NIPS \u201916). 2119\u20132127."},{"key":"e_1_3_1_25_2","first-page":"19730","volume-title":"Proceedings of the International Conference on Machine Learning (ICML \u201923)","author":"Li Junnan","year":"2023","unstructured":"Junnan Li, Dongxu Li, Silvio Savarese, and Steven C. H. Hoi. 2023. BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In Proceedings of the International Conference on Machine Learning (ICML \u201923). 19730\u201319742."},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.01041"},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1145\/3300938"},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00640"},{"key":"e_1_3_1_29_2","first-page":"74","volume-title":"Proceedings of the Text Summarization Branches Out","author":"Lin Chin-Yew","year":"2004","unstructured":"Chin-Yew Lin. 2004. Rouge: A package for automatic evaluation of summaries. In Proceedings of the Text Summarization Branches Out. 74\u201381."},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00434"},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.1145\/3498340"},{"key":"e_1_3_1_32_2","volume-title":"Proceedings of the Annual Conference on Neural Information Processing Systems (NIPS \u201920)","author":"Mahajan Shweta","year":"2020","unstructured":"Shweta Mahajan and Stefan Roth. 2020. Diverse image captioning with context-object split latent spaces. In Proceedings of the Annual Conference on Neural Information Processing Systems (NIPS \u201920)."},{"key":"e_1_3_1_33_2","first-page":"1682","volume-title":"Proceedings of the Annual Conference on Neural Information Processing Systems (NIPS \u201914)","author":"Malinowski Mateusz","year":"2014","unstructured":"Mateusz Malinowski and Mario Fritz. 2014. A multi-world approach to question answering about real-world scenes based on uncertain input. In Proceedings of the Annual Conference on Neural Information Processing Systems (NIPS \u201914). 1682\u20131690."},{"key":"e_1_3_1_34_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P16-1170"},{"key":"e_1_3_1_35_2","first-page":"311","volume-title":"Proceedings of the Annual Meeting of the Association for Computational Linguistics (ACL \u201902)","author":"Papineni Kishore","year":"2002","unstructured":"Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: A method for automatic evaluation of machine translation. In Proceedings of the Annual Meeting of the Association for Computational Linguistics (ACL \u201902). 311\u2013318."},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/D14-1162"},{"key":"e_1_3_1_37_2","first-page":"2953","volume-title":"Proceedings of the Annual Conference on Neural Information Processing Systems (NIPS \u201915)","author":"Ren Mengye","year":"2015","unstructured":"Mengye Ren, Ryan Kiros, and Richard S. Zemel. 2015. Exploring models and data for image question answering. In Proceedings of the Annual Conference on Neural Information Processing Systems (NIPS \u201915). 2953\u20132961."},{"key":"e_1_3_1_38_2","first-page":"91","volume-title":"Proceedings of the Annual Conference on Neural Information Processing Systems (NIPS \u201915)","author":"Ren Shaoqing","year":"2015","unstructured":"Shaoqing Ren, Kaiming He, Ross B. Girshick, and Jian Sun. 2015. Faster R-CNN: Towards real-time object detection with region proposal networks. In Proceedings of the Annual Conference on Neural Information Processing Systems (NIPS \u201915). 91\u201399."},{"key":"e_1_3_1_39_2","volume-title":"Proceedings of the International Conference on Learning Representations (ICLR \u201915)","author":"Simonyan Karen","year":"2015","unstructured":"Karen Simonyan and Andrew Zisserman. 2015. Very deep convolutional networks for large-scale image recognition. In Proceedings of the International Conference on Learning Representations (ICLR \u201915)."},{"issue":"11","key":"e_1_3_1_40_2","article-title":"Visualizing data using t-SNE.","volume":"9","author":"Maaten Laurens Van der","year":"2008","unstructured":"Laurens Van der Maaten and Geoffrey Hinton. 2008. Visualizing data using t-SNE. J. Mach. Learn. Res. 9, 11 (2008).","journal-title":"J. Mach. Learn. Res."},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7299087"},{"key":"e_1_3_1_42_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v32i1.12340"},{"key":"e_1_3_1_43_2","volume-title":"Proceedings of the Annual Conference on Neural Information Processing Systems (NIPS \u201921)","author":"Wu Bo","year":"2021","unstructured":"Bo Wu, Shoubin Yu, Zhenfang Chen, Josh Tenenbaum, and Chuang Gan. 2021. STAR: A benchmark for situated reasoning in real-world videos. In Proceedings of the Annual Conference on Neural Information Processing Systems (NIPS \u201921)."},{"key":"e_1_3_1_44_2","doi-asserted-by":"publisher","DOI":"10.1145\/3474085.3476969"},{"key":"e_1_3_1_45_2","doi-asserted-by":"publisher","DOI":"10.1109\/TASLP.2021.3138706"},{"key":"e_1_3_1_46_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D18-1428"},{"key":"e_1_3_1_47_2","doi-asserted-by":"publisher","DOI":"10.1109\/TNNLS.2020.2986029"},{"key":"e_1_3_1_48_2","doi-asserted-by":"publisher","DOI":"10.1145\/3316767"},{"key":"e_1_3_1_49_2","volume-title":"Proceedings of the International Conference on Learning Representations (ICLR \u201920)","author":"Zhang Tianyi","year":"2020","unstructured":"Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020. BERTScore: Evaluating text generation with BERT. In Proceedings of the International Conference on Learning Representations (ICLR \u201920)."},{"key":"e_1_3_1_50_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D18-1244"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3640014","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3640014","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T22:54:01Z","timestamp":1750287241000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3640014"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,3,8]]},"references-count":49,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2024,6,30]]}},"alternative-id":["10.1145\/3640014"],"URL":"https:\/\/doi.org\/10.1145\/3640014","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,3,8]]},"assertion":[{"value":"2022-10-07","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-12-26","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-03-08","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}