{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,8,23]],"date-time":"2025-08-23T05:23:43Z","timestamp":1755926623806,"version":"3.41.0"},"publisher-location":"New York, NY, USA","reference-count":42,"publisher":"ACM","license":[{"start":{"date-parts":[[2021,10,17]],"date-time":"2021-10-17T00:00:00Z","timestamp":1634428800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2021,10,17]]},"DOI":"10.1145\/3474085.3475222","type":"proceedings-article","created":{"date-parts":[[2021,10,18]],"date-time":"2021-10-18T20:00:05Z","timestamp":1634587205000},"page":"1331-1340","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":19,"title":["Two-stage Visual Cues Enhancement Network for Referring Image Segmentation"],"prefix":"10.1145","author":[{"given":"Yang","family":"Jiao","sequence":"first","affiliation":[{"name":"Fudan University, Shanghai, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zequn","family":"Jie","sequence":"additional","affiliation":[{"name":"Meituan, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Weixin","family":"Luo","sequence":"additional","affiliation":[{"name":"Meituan, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jingjing","family":"Chen","sequence":"additional","affiliation":[{"name":"Fudan University, Shanghai, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yu-Gang","family":"Jiang","sequence":"additional","affiliation":[{"name":"Fudan University, Shanghai, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xiaolin","family":"Wei","sequence":"additional","affiliation":[{"name":"Meituan, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Lin","family":"Ma","sequence":"additional","affiliation":[{"name":"Meituan, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2021,10,17]]},"reference":[{"key":"e_1_3_2_2_1_1","volume-title":"Proceedings of the IEEE International Conference on Computer Vision (ICCV).","author":"Hedi","year":"2017","unstructured":"Hedi Ben-younes, Remi Cadene , Matthieu Cord , and Nicolas Thome . 2017 . MUTAN: Multimodal Tucker Fusion for Visual Question Answering . In Proceedings of the IEEE International Conference on Computer Vision (ICCV). Hedi Ben-younes, Remi Cadene, Matthieu Cord, and Nicolas Thome. 2017. MUTAN: Multimodal Tucker Fusion for Visual Question Answering. In Proceedings of the IEEE International Conference on Computer Vision (ICCV)."},{"key":"e_1_3_2_2_2_1","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/D14-1082"},{"key":"e_1_3_2_2_3_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00755"},{"key":"e_1_3_2_2_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2017.2699184"},{"key":"e_1_3_2_2_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/3394171.3413551"},{"key":"e_1_3_2_2_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"e_1_3_2_2_7_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.cviu.2009.03.008"},{"key":"e_1_3_2_2_8_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-009-0275-4"},{"key":"e_1_3_2_2_9_1","volume-title":"Stacked Latent Attention for Multimodal Reasoning. In 2018 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2018","author":"Fan Haoqi","year":"2018","unstructured":"Haoqi Fan and Jiatong Zhou . 2018 . Stacked Latent Attention for Multimodal Reasoning. In 2018 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2018 , Salt Lake City, UT, USA, June 18--22 , 2018. IEEE Computer Society, 1072--1080. https:\/\/doi.org\/10.1109\/CVPR.2018.00118 10.1109\/CVPR.2018.00118 Haoqi Fan and Jiatong Zhou. 2018. Stacked Latent Attention for Multimodal Reasoning. In 2018 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2018, Salt Lake City, UT, USA, June 18--22, 2018. IEEE Computer Society, 1072--1080. https:\/\/doi.org\/10.1109\/CVPR.2018.00118"},{"key":"e_1_3_2_2_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00326"},{"key":"e_1_3_2_2_11_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46448-0_7"},{"key":"e_1_3_2_2_12_1","volume-title":"Bi-Directional Relationship Inferring Network for Referring Image Segmentation. In 2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2020","author":"Hu Zhiwei","year":"2020","unstructured":"Zhiwei Hu , Guang Feng , Jiayu Sun , Lihe Zhang , and Huchuan Lu . 2020 . Bi-Directional Relationship Inferring Network for Referring Image Segmentation. In 2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2020 , Seattle, WA, USA, June 13--19 , 2020. IEEE, 4423--4432. https:\/\/doi.org\/10.1109\/CVPR42600.2020.00448 10.1109\/CVPR42600.2020.00448 Zhiwei Hu, Guang Feng, Jiayu Sun, Lihe Zhang, and Huchuan Lu. 2020. Bi-Directional Relationship Inferring Network for Referring Image Segmentation. In 2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2020, Seattle, WA, USA, June 13--19, 2020. IEEE, 4423--4432. https:\/\/doi.org\/10.1109\/CVPR42600.2020.00448"},{"key":"e_1_3_2_2_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.01050"},{"key":"e_1_3_2_2_14_1","volume-title":"CCNet: Criss-Cross Attention for Semantic Segmentation. In 2019 IEEE\/CVF International Conference on Computer Vision, ICCV 2019","author":"Huang Zilong","year":"2019","unstructured":"Zilong Huang , Xinggang Wang , Lichao Huang , Chang Huang , Yunchao Wei , and Wenyu Liu . 2019 . CCNet: Criss-Cross Attention for Semantic Segmentation. In 2019 IEEE\/CVF International Conference on Computer Vision, ICCV 2019 , Seoul, Korea (South), October 27 - November 2, 2019. IEEE, 603--612. https:\/\/doi.org\/10.1109\/ICCV.2019.00069 10.1109\/ICCV.2019.00069 Zilong Huang, Xinggang Wang, Lichao Huang, Chang Huang, Yunchao Wei, and Wenyu Liu. 2019. CCNet: Criss-Cross Attention for Semantic Segmentation. In 2019 IEEE\/CVF International Conference on Computer Vision, ICCV 2019, Seoul, Korea (South), October 27 - November 2, 2019. IEEE, 603--612. https:\/\/doi.org\/10.1109\/ICCV.2019.00069"},{"key":"e_1_3_2_2_15_1","volume-title":"Proceedings, Part X (Lecture Notes in Computer Science","volume":"75","author":"Hui Tianrui","year":"2020","unstructured":"Tianrui Hui , Si Liu , Shaofei Huang , Guanbin Li , Sansi Yu , Faxi Zhang , and Jizhong Han . 2020 . Linguistic Structure Guided Context Modeling for Referring Image Segmentation. In Computer Vision - ECCV 2020 - 16th European Conference, Glasgow, UK, August 23--28, 2020 , Proceedings, Part X (Lecture Notes in Computer Science , Vol. 12355), Andrea Vedaldi, Horst Bischof, Thomas Brox, and Jan-Michael Frahm (Eds.). Springer, 59-- 75 . https:\/\/doi.org\/10.1007\/978--3-030--58607--2_4 10.1007\/978--3-030--58607--2_4 Tianrui Hui, Si Liu, Shaofei Huang, Guanbin Li, Sansi Yu, Faxi Zhang, and Jizhong Han. 2020. Linguistic Structure Guided Context Modeling for Referring Image Segmentation. In Computer Vision - ECCV 2020 - 16th European Conference, Glasgow, UK, August 23--28, 2020, Proceedings, Part X (Lecture Notes in Computer Science, Vol. 12355), Andrea Vedaldi, Horst Bischof, Thomas Brox, and Jan-Michael Frahm (Eds.). Springer, 59--75. https:\/\/doi.org\/10.1007\/978--3-030--58607--2_4"},{"key":"e_1_3_2_2_16_1","volume-title":"Berg","author":"Kazemzadeh Sahar","year":"2014","unstructured":"Sahar Kazemzadeh , Vicente Ordonez , Mark Matten , and Tamara L . Berg . 2014 . ReferIt Game: Referring to Objects in Photographs of Natural Scenes. In EMNLP. Sahar Kazemzadeh, Vicente Ordonez, Mark Matten, and Tamara L. Berg. 2014. ReferIt Game: Referring to Objects in Photographs of Natural Scenes. In EMNLP."},{"key":"e_1_3_2_2_17_1","volume-title":"Kingma and Jimmy Ba","author":"Diederik","year":"2015","unstructured":"Diederik P. Kingma and Jimmy Ba . 2015 . Adam : A Method for Stochastic Optimization. In 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7--9, 2015, Conference Track Proceedings, Yoshua Bengio and Yann LeCun (Eds .). http:\/\/arxiv.org\/abs\/1412.6980 Diederik P. Kingma and Jimmy Ba. 2015. Adam: A Method for Stochastic Optimization. In 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7--9, 2015, Conference Track Proceedings, Yoshua Bengio and Yann LeCun (Eds.). http:\/\/arxiv.org\/abs\/1412.6980"},{"key":"e_1_3_2_2_18_1","doi-asserted-by":"publisher","DOI":"10.5555\/2986459.2986472"},{"key":"e_1_3_2_2_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00602"},{"key":"e_1_3_2_2_20_1","volume-title":"European Conference on Computer Vision. https:\/\/www.microsoft.com\/en-us\/research\/publication\/microsoft-coco-common-objects-in-context\/","author":"Lin Tsung-Yi","year":"2014","unstructured":"Tsung-Yi Lin , Michael Maire , Serge Belongie , James Hays , Pietro Perona , Deva Ramanan , Piotr Dollar , and Larry Zitnick . 2014 . Microsoft COCO: Common Objects in Context. In ECCV eccv ed.) . European Conference on Computer Vision. https:\/\/www.microsoft.com\/en-us\/research\/publication\/microsoft-coco-common-objects-in-context\/ Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollar, and Larry Zitnick. 2014. Microsoft COCO: Common Objects in Context. In ECCV eccv ed.). European Conference on Computer Vision. https:\/\/www.microsoft.com\/en-us\/research\/publication\/microsoft-coco-common-objects-in-context\/"},{"key":"e_1_3_2_2_21_1","unstructured":"Chenxi Liu Zhe Lin Xiaohui Shen Jimei Yang Xin Lu and Alan Yuille. 2017. Recurrent Multimodal Interaction for Referring Image Segmentation. In ICCV.  Chenxi Liu Zhe Lin Xiaohui Shen Jimei Yang Xin Lu and Alan Yuille. 2017. Recurrent Multimodal Interaction for Referring Image Segmentation. In ICCV."},{"key":"e_1_3_2_2_22_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00913"},{"key":"e_1_3_2_2_23_1","unstructured":"Junhua Mao Jonathan Huang Alexander Toshev Oana Camburu Alan Yuille and Kevin Murphy. 2016. Generation and Comprehension of Unambiguous Object Descriptions. In CVPR.  Junhua Mao Jonathan Huang Alexander Toshev Oana Camburu Alan Yuille and Kevin Murphy. 2016. Generation and Comprehension of Unambiguous Object Descriptions. In CVPR."},{"key":"e_1_3_2_2_24_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01252-6_39"},{"key":"e_1_3_2_2_25_1","volume-title":"Local-Global Video-Text Interactions for Temporal Grounding. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR).","author":"Mun Jonghwan","year":"2020","unstructured":"Jonghwan Mun , Minsu Cho , and Bohyung Han . 2020 . Local-Global Video-Text Interactions for Temporal Grounding. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Jonghwan Mun, Minsu Cho, and Bohyung Han. 2020. Local-Global Video-Text Interactions for Temporal Grounding. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)."},{"key":"e_1_3_2_2_26_1","volume-title":"Dual Attention Networks for Multimodal Reasoning and Matching. In 2017 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2017","author":"Nam Hyeonseob","year":"2017","unstructured":"Hyeonseob Nam , Jung-Woo Ha , and Jeonghee Kim . 2017 . Dual Attention Networks for Multimodal Reasoning and Matching. In 2017 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2017 , Honolulu, HI, USA, July 21--26 , 2017. IEEE Computer Society, 2156--2164. https:\/\/doi.org\/10.1109\/CVPR.2017.232 10.1109\/CVPR.2017.232 Hyeonseob Nam, Jung-Woo Ha, and Jeonghee Kim. 2017. Dual Attention Networks for Multimodal Reasoning and Matching. In 2017 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2017, Honolulu, HI, USA, July 21--26, 2017. IEEE Computer Society, 2156--2164. https:\/\/doi.org\/10.1109\/CVPR.2017.232"},{"key":"e_1_3_2_2_27_1","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/D14-1162"},{"key":"e_1_3_2_2_28_1","volume-title":"a distilled version of BERT: smaller, faster, cheaper and lighter. CoRR","author":"Sanh Victor","year":"2019","unstructured":"Victor Sanh , Lysandre Debut , Julien Chaumond , and Thomas Wolf . 2019. DistilBERT , a distilled version of BERT: smaller, faster, cheaper and lighter. CoRR , Vol. abs\/ 1910 .01108 ( 2019 ). arxiv: 1910.01108 http:\/\/arxiv.org\/abs\/1910.01108 Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019. DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter. CoRR, Vol. abs\/1910.01108 (2019). arxiv: 1910.01108 http:\/\/arxiv.org\/abs\/1910.01108"},{"key":"e_1_3_2_2_29_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01231-1_3"},{"key":"e_1_3_2_2_30_1","volume-title":"Johannes Totz, Andrew P. Aitken, Rob Bishop, Daniel Rueckert, and Zehan Wang.","author":"Shi Wenzhe","year":"2016","unstructured":"Wenzhe Shi , Jose Caballero , Ferenc Husz\u00e1 r , Johannes Totz, Andrew P. Aitken, Rob Bishop, Daniel Rueckert, and Zehan Wang. 2016 . Real-Time Single Image and Video Super-Resolution Using an Efficient Sub-Pixel Convolutional Neural Network. CoRR , Vol. abs\/ 1609 .05158 (2016). arxiv: 1609.05158 http:\/\/arxiv.org\/abs\/1609.05158 Wenzhe Shi, Jose Caballero, Ferenc Husz\u00e1 r, Johannes Totz, Andrew P. Aitken, Rob Bishop, Daniel Rueckert, and Zehan Wang. 2016. Real-Time Single Image and Video Super-Resolution Using an Efficient Sub-Pixel Convolutional Neural Network. CoRR, Vol. abs\/1609.05158 (2016). arxiv: 1609.05158 http:\/\/arxiv.org\/abs\/1609.05158"},{"key":"e_1_3_2_2_31_1","doi-asserted-by":"publisher","DOI":"10.5555\/2969239.2969329"},{"key":"e_1_3_2_2_32_1","doi-asserted-by":"publisher","DOI":"10.15607\/RSS.2018.XIV.028"},{"volume-title":"Very Deep Convolutional Networks for Large-Scale Image Recognition. In International Conference on Learning Representations.","author":"Simonyan K.","key":"e_1_3_2_2_33_1","unstructured":"K. Simonyan and A. Zisserman . 2015 . Very Deep Convolutional Networks for Large-Scale Image Recognition. In International Conference on Learning Representations. K. Simonyan and A. Zisserman. 2015. Very Deep Convolutional Networks for Large-Scale Image Recognition. In International Conference on Learning Representations."},{"key":"e_1_3_2_2_34_1","doi-asserted-by":"publisher","DOI":"10.5555\/3295222.3295349"},{"key":"e_1_3_2_2_35_1","volume-title":"6th International Conference on Learning Representations, ICLR","author":"Velickovic Petar","year":"2018","unstructured":"Petar Velickovic , Guillem Cucurull , Arantxa Casanova , Adriana Romero , Pietro Li\u00f2 , and Yoshua Bengio . 2018. Graph Attention Networks . In 6th International Conference on Learning Representations, ICLR 2018 , Vancouver, BC , Canada, April 30 - May 3, 2018, Conference Track Proceedings. OpenReview .net. https:\/\/openreview.net\/forum?id=rJXMpikCZ Petar Velickovic, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Li\u00f2, and Yoshua Bengio. 2018. Graph Attention Networks. In 6th International Conference on Learning Representations, ICLR 2018, Vancouver, BC, Canada, April 30 - May 3, 2018, Conference Track Proceedings. OpenReview.net. https:\/\/openreview.net\/forum?id=rJXMpikCZ"},{"key":"e_1_3_2_2_36_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00813"},{"key":"e_1_3_2_2_37_1","volume-title":"Cross-Modal Self-Attention Network for Referring Image Segmentation. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2019","author":"Ye Linwei","year":"2019","unstructured":"Linwei Ye , Mrigank Rochan , Zhi Liu , and Yang Wang . 2019 . Cross-Modal Self-Attention Network for Referring Image Segmentation. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2019 , Long Beach, CA, USA, June 16--20 , 2019. Computer Vision Foundation \/ IEEE, 10502--10511. https:\/\/doi.org\/10.1109\/CVPR.2019.01075 10.1109\/CVPR.2019.01075 Linwei Ye, Mrigank Rochan, Zhi Liu, and Yang Wang. 2019. Cross-Modal Self-Attention Network for Referring Image Segmentation. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2019, Long Beach, CA, USA, June 16--20, 2019. Computer Vision Foundation \/ IEEE, 10502--10511. https:\/\/doi.org\/10.1109\/CVPR.2019.01075"},{"volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR).","author":"Yu Licheng","key":"e_1_3_2_2_38_1","unstructured":"Licheng Yu , Zhe Lin , Xiaohui Shen , Jimei Yang , Xin Lu , Mohit Bansal , and Tamara L. Berg . 2018. MAttNet: Modular Attention Network for Referring Expression Comprehension . In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Licheng Yu, Zhe Lin, Xiaohui Shen, Jimei Yang, Xin Lu, Mohit Bansal, and Tamara L. Berg. 2018. MAttNet: Modular Attention Network for Referring Expression Comprehension. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)."},{"key":"e_1_3_2_2_39_1","volume-title":"Berg","author":"Yu Licheng","year":"2016","unstructured":"Licheng Yu , Patric Poirson , Shan Yang , Alexander C. Berg , and Tamara L . Berg . 2016 . Modeling Context in Referring Expressions. In ECCV. Licheng Yu, Patric Poirson, Shan Yang, Alexander C. Berg, and Tamara L. Berg. 2016. Modeling Context in Referring Expressions. In ECCV."},{"key":"e_1_3_2_2_40_1","volume-title":"Proceedings, Part I (Lecture Notes in Computer Science","volume":"833","author":"Matthew","unstructured":"Matthew D. Zeiler and Rob Fergus. 2014. Visualizing and Understanding Convolutional Networks. In Computer Vision - ECCV 2014 - 13th European Conference, Zurich, Switzerland, September 6--12, 2014 , Proceedings, Part I (Lecture Notes in Computer Science , Vol. 8689), David J. Fleet, Tom\u00e1 s Pajdla, Bernt Schiele, and Tinne Tuytelaars (Eds.). Springer, 818-- 833 . https:\/\/doi.org\/10.1007\/978--3--319--10590--1_53 10.1007\/978--3--319--10590--1_53 Matthew D. Zeiler and Rob Fergus. 2014. Visualizing and Understanding Convolutional Networks. In Computer Vision - ECCV 2014 - 13th European Conference, Zurich, Switzerland, September 6--12, 2014, Proceedings, Part I (Lecture Notes in Computer Science, Vol. 8689), David J. Fleet, Tom\u00e1 s Pajdla, Bernt Schiele, and Tinne Tuytelaars (Eds.). Springer, 818--833. https:\/\/doi.org\/10.1007\/978--3--319--10590--1_53"},{"key":"e_1_3_2_2_41_1","doi-asserted-by":"crossref","unstructured":"Yukun Zhu Ryan Kiros Richard Zemel Ruslan Salakhutdinov Raquel Urtasun Antonio Torralba and Sanja Fidler. 2015. Aligning Books and Movies: Towards Story-like Visual Explanations by Watching Movies and Reading Books. In arXiv preprint arXiv:1506.06724.  Yukun Zhu Ryan Kiros Richard Zemel Ruslan Salakhutdinov Raquel Urtasun Antonio Torralba and Sanja Fidler. 2015. Aligning Books and Movies: Towards Story-like Visual Explanations by Watching Movies and Reading Books. In arXiv preprint arXiv:1506.06724.","DOI":"10.1109\/ICCV.2015.11"},{"key":"e_1_3_2_2_42_1","doi-asserted-by":"crossref","unstructured":"Yuan Zhuang Zhenguang Liu Peng Qian Qi Liu Xiang Wang and Qinming He. 2020. Smart Contract Vulnerability Detection using Graph Neural Network. In IJCAI. 3283--3290.  Yuan Zhuang Zhenguang Liu Peng Qian Qi Liu Xiang Wang and Qinming He. 2020. Smart Contract Vulnerability Detection using Graph Neural Network. In IJCAI. 3283--3290.","DOI":"10.24963\/ijcai.2020\/454"}],"event":{"name":"MM '21: ACM Multimedia Conference","sponsor":["SIGMM ACM Special Interest Group on Multimedia"],"location":"Virtual Event China","acronym":"MM '21"},"container-title":["Proceedings of the 29th ACM International Conference on Multimedia"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3474085.3475222","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3474085.3475222","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T20:48:16Z","timestamp":1750193296000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3474085.3475222"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,10,17]]},"references-count":42,"alternative-id":["10.1145\/3474085.3475222","10.1145\/3474085"],"URL":"https:\/\/doi.org\/10.1145\/3474085.3475222","relation":{},"subject":[],"published":{"date-parts":[[2021,10,17]]},"assertion":[{"value":"2021-10-17","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}