{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,1]],"date-time":"2026-05-01T09:04:48Z","timestamp":1777626288961,"version":"3.51.4"},"reference-count":52,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2021,2,26]],"date-time":"2021-02-26T00:00:00Z","timestamp":1614297600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"Operating Expenses of Basic Scientific Research Project of the People?s Public Security University of China","award":["2019JKF111"],"award-info":[{"award-number":["2019JKF111"]}]},{"DOI":"10.13039\/501100012166","name":"National Key Research and Development Program of China","doi-asserted-by":"publisher","award":["A19808"],"award-info":[{"award-number":["A19808"]}],"id":[{"id":"10.13039\/501100012166","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100012226","name":"Fundamental Research Funds for the Central Universities","doi-asserted-by":"publisher","id":[{"id":"10.13039\/501100012226","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Zhejiang Lab","award":["2019KD0AB04"],"award-info":[{"award-number":["2019KD0AB04"]}]},{"name":"Beijing Natural Science Foundation","award":["4202034"],"award-info":[{"award-number":["4202034"]}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["62076246,61876177"],"award-info":[{"award-number":["62076246,61876177"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Intell. Syst. Technol."],"published-print":{"date-parts":[[2021,4,30]]},"abstract":"<jats:p>The goal of referring image segmentation is to identify the object matched with an input natural language expression. Previous methods only support English descriptions, whereas Chinese is also broadly used around the world, which limits the potential application of this task. Therefore, we propose to extend existing datasets with Chinese descriptions and preprocessing tools for training and evaluating bilingual referring segmentation models. In addition, previous methods also lack the ability to collaboratively learn channel-wise and spatial-wise cross-modal attention to well align visual and linguistic modalities. To tackle these limitations, we propose a Linguistic Excitation module to excite image channels guided by language information and a Linguistic Aggregation module to aggregate multimodal information based on image-language relationships. Since different levels of features from the visual backbone encode rich visual information, we also propose a Cross-Level Attentive Fusion module to fuse multilevel features gated by language information. Extensive experiments on four English and Chinese benchmarks show that our bilingual referring image segmentation model outperforms previous methods.<\/jats:p>","DOI":"10.1145\/3446345","type":"journal-article","created":{"date-parts":[[2021,2,26]],"date-time":"2021-02-26T11:13:43Z","timestamp":1614338023000},"page":"1-17","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":3,"title":["Attentive Excitation and Aggregation for Bilingual Referring Image Segmentation"],"prefix":"10.1145","volume":"12","author":[{"given":"Qianli","family":"Zhou","sequence":"first","affiliation":[{"name":"People\u2019s Public Security University of China, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1172-1554","authenticated-orcid":false,"given":"Tianrui","family":"Hui","sequence":"additional","affiliation":[{"name":"Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Rong","family":"Wang","sequence":"additional","affiliation":[{"name":"People\u2019s Public Security University of China, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Haimiao","family":"Hu","sequence":"additional","affiliation":[{"name":"Beihang University, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Si","family":"Liu","sequence":"additional","affiliation":[{"name":"Beihang University, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2021,2,26]]},"reference":[{"key":"e_1_2_1_1_1","unstructured":"Dzmitry Bahdanau Kyunghyun Cho and Yoshua Bengio. 2014. Neural machine translation by jointly learning to align and translate. arXiv:1409.0473  Dzmitry Bahdanau Kyunghyun Cho and Yoshua Bengio. 2014. Neural machine translation by jointly learning to align and translate. arXiv:1409.0473"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.285"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00755"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2017.2699184"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.5555\/2832415.2832421"},{"key":"e_1_2_1_6_1","unstructured":"Xinpeng Chen Lin Ma Jingyuan Chen Zequn Jie Wei Liu and Jiebo Luo. 2018. Real-time referring expression comprehension by single-stage grounding network. arXiv:1812.03426  Xinpeng Chen Lin Ma Jingyuan Chen Zequn Jie Wei Liu and Jiebo Luo. 2018. Real-time referring expression comprehension by single-stage grounding network. arXiv:1812.03426"},{"key":"e_1_2_1_7_1","unstructured":"Yunpeng Chen Yannis Kalantidis Jianshu Li Shuicheng Yan and Jiashi Feng. 2018. A\u02c62-nets: Double attention networks. In Advances in Neural Information Processing Systems. 352--361.  Yunpeng Chen Yannis Kalantidis Jianshu Li Shuicheng Yan and Jiashi Feng. 2018. A\u02c62-nets: Double attention networks. In Advances in Neural Information Processing Systems. 352--361."},{"key":"e_1_2_1_8_1","unstructured":"Yunpeng Chen Jianan Li Huaxin Xiao Xiaojie Jin Shuicheng Yan and Jiashi Feng. 2017. Dual path networks. arXiv:1707.01629  Yunpeng Chen Jianan Li Huaxin Xiao Xiaojie Jin Shuicheng Yan and Jiashi Feng. 2017. Dual path networks. arXiv:1707.01629"},{"key":"e_1_2_1_9_1","unstructured":"Yi-Wen Chen Yi-Hsuan Tsai Tiantian Wang Yen-Yu Lin and Ming-Hsuan Yang. 2019. Referring expression object segmentation with caption-aware consistency. arXiv:1910.04748  Yi-Wen Chen Yi-Hsuan Tsai Tiantian Wang Yen-Yu Lin and Ming-Hsuan Yang. 2019. Referring expression object segmentation with caption-aware consistency. arXiv:1910.04748"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-014-0733-5"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00326"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00685"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1145\/3394171.3413854"},{"key":"e_1_2_1_14_1","unstructured":"Yoav Goldberg and Omer Levy. 2014. Word2Vec explained: Deriving Mikolov et\u00a0al.\u2019s negative-sampling word-embedding method. arXiv:1402.3722  Yoav Goldberg and Omer Levy. 2014. Word2Vec explained: Deriving Mikolov et\u00a0al.\u2019s negative-sampling word-embedding method. arXiv:1402.3722"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1162\/neco.1997.9.8.1735"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00745"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46448-0_7"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00448"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.01050"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2020.3013142"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58607-2_4"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/D14-1086"},{"key":"e_1_2_1_24_1","unstructured":"Philipp Kr\u00e4henb\u00fchl and Vladlen Koltun. 2011. Efficient inference in fully connected CRFs with Gaussian edge potentials. In Advances in Neural Information Processing Systems. 109--117.  Philipp Kr\u00e4henb\u00fchl and Vladlen Koltun. 2011. Efficient inference in fully connected CRFs with Gaussian edge potentials. In Advances in Neural Information Processing Systems. 109--117."},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00602"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D15-1098"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICME.2019.00112"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.01089"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00056"},{"key":"e_1_2_1_30_1","unstructured":"Chenxi Liu Lin Zhe Xiaohui Shen Jimei Yang and Alan Yuille. 2017. Recurrent multimodal interaction for referring image segmentation. arXiv:1703.07939  Chenxi Liu Lin Zhe Xiaohui Shen Jimei Yang and Alan Yuille. 2017. Recurrent multimodal interaction for referring image segmentation. arXiv:1703.07939"},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.9"},{"key":"e_1_2_1_32_1","doi-asserted-by":"crossref","unstructured":"Edgar A. Margffoy-Tuay Juan C. Perez Emilio Botero and Pablo Arbelaez. 2018. Dynamic multimodal instance segmentation guided by natural language queries. arXiv:1807.02257  Edgar A. Margffoy-Tuay Juan C. Perez Emilio Botero and Pablo Arbelaez. 2018. Dynamic multimodal instance segmentation guided by natural language queries. arXiv:1807.02257","DOI":"10.1007\/978-3-030-01252-6_39"},{"key":"e_1_2_1_33_1","unstructured":"Aditya Mogadala Marimuthu Kalimuthu and Dietrich Klakow. 2019. Trends in integration of vision and language research: A survey of tasks datasets and methods. arXiv:1907.09358  Aditya Mogadala Marimuthu Kalimuthu and Dietrich Klakow. 2019. Trends in integration of vision and language research: A survey of tasks datasets and methods. arXiv:1907.09358"},{"key":"e_1_2_1_34_1","volume-title":"Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP\u201914)","author":"Pennington Jeffrey","unstructured":"Jeffrey Pennington , Richard Socher , and Christopher D. Manning . 2014. Glove: Global vectors for word representation . In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP\u201914) . 1532--1543. Jeffrey Pennington, Richard Socher, and Christopher D. Manning. 2014. Glove: Global vectors for word representation. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP\u201914). 1532--1543."},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2019.2942480"},{"key":"e_1_2_1_36_1","first-page":"1","article-title":"Scene graph generation with hierarchical context","volume":"99","author":"Ren Guanghui","year":"2020","unstructured":"Guanghui Ren , Lejian Ren , Yue Liao , Si Liu , Bo Li , Jizhong Han , and Shuicheng Yan . 2020 . Scene graph generation with hierarchical context . IEEE Transactions on Neural Networks and Learning Systems PP , 99 , 1 -- 7 . Guanghui Ren, Lejian Ren, Yue Liao, Si Liu, Bo Li, Jizhong Han, and Shuicheng Yan. 2020. Scene graph generation with hierarchical context. IEEE Transactions on Neural Networks and Learning Systems PP, 99, 1--7.","journal-title":"IEEE Transactions on Neural Networks and Learning Systems PP"},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01231-1_3"},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.5555\/2986459.2986549"},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-12640-1_34"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-60633-6_3"},{"key":"e_1_2_1_41_1","unstructured":"Zongheng Tang Yue Liao Si Liu Guanbin Li Xiaojie Jin Hongxu Jiang Qian Yu and Dong Xu. 2020. Human-centric spatio-temporal video grounding with visual transformers. arXiv:2011.05049  Zongheng Tang Yue Liao Si Liu Guanbin Li Xiaojie Jin Hongxu Jiang Qian Yu and Dong Xu. 2020. Human-centric spatio-temporal video grounding with visual transformers. arXiv:2011.05049"},{"key":"e_1_2_1_42_1","volume-title":"Proceedings of the 48th Annual Meeting of the Association for Computational Linguistics.","author":"Turian Joseph P.","year":"2010","unstructured":"Joseph P. Turian , Lev Arie Ratinov , and Yoshua Bengio . 2010 . Word representations: A simple and general method for semi-supervised learning . In Proceedings of the 48th Annual Meeting of the Association for Computational Linguistics. Joseph P. Turian, Lev Arie Ratinov, and Yoshua Bengio. 2010. Word representations: A simple and general method for semi-supervised learning. In Proceedings of the 48th Annual Meeting of the Association for Computational Linguistics."},{"key":"e_1_2_1_43_1","unstructured":"Ashish Vaswani Noam Shazeer Niki Parmar Jakob Uszkoreit Llion Jones Aidan N. Gomez \u0141ukasz Kaiser and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems. 5998--6008.  Ashish Vaswani Noam Shazeer Niki Parmar Jakob Uszkoreit Llion Jones Aidan N. Gomez \u0141ukasz Kaiser and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems. 5998--6008."},{"key":"e_1_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01234-2_1"},{"key":"e_1_2_1_45_1","unstructured":"Shi Xingjian Zhourong Chen Hao Wang Dit-Yan Yeung Wai-Kin Wong and Wang-Chun Woo. 2015. Convolutional LSTM network: A machine learning approach for precipitation nowcasting. In Advances in Neural Information Processing Systems. 802--810.  Shi Xingjian Zhourong Chen Hao Wang Dit-Yan Yeung Wai-Kin Wong and Wang-Chun Woo. 2015. Convolutional LSTM network: A machine learning approach for precipitation nowcasting. In Advances in Neural Information Processing Systems. 802--810."},{"key":"e_1_2_1_46_1","unstructured":"Linwei Ye Zhi Liu and Yang Wang. 2020. Dual convolutional LSTM network for referring image segmentation. arXiv:2001.11561  Linwei Ye Zhi Liu and Yang Wang. 2020. Dual convolutional LSTM network for referring image segmentation. arXiv:2001.11561"},{"key":"e_1_2_1_47_1","unstructured":"Linwei Ye Mrigank Rochan Zhi Liu and Yang Wang. 2019. Cross-modal self-attention network for referring image segmentation. arXiv:1904.04745  Linwei Ye Mrigank Rochan Zhi Liu and Yang Wang. 2019. Cross-modal self-attention network for referring image segmentation. arXiv:1904.04745"},{"key":"e_1_2_1_48_1","unstructured":"Licheng Yu Patric Poirson Shan Yang Alexander Berg and Tamara Berg. 2016. Modeling context in referring expressions. arXiv:1608.00272  Licheng Yu Patric Poirson Shan Yang Alexander Berg and Tamara Berg. 2016. Modeling context in referring expressions. arXiv:1608.00272"},{"key":"e_1_2_1_49_1","volume-title":"Berg","author":"Yu Licheng","year":"2018","unstructured":"Licheng Yu , Lin Zhe , Xiaohui Shen , Jimei Yang , and Tamara L . Berg . 2018 . MAttNet : Modular attention network for referring expression comprehension. arXiv:1801.08186 Licheng Yu, Lin Zhe, Xiaohui Shen, Jimei Yang, and Tamara L. Berg. 2018. MAttNet: Modular attention network for referring expression comprehension. arXiv:1801.08186"},{"key":"e_1_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.1145\/3394171.3413846"},{"key":"e_1_2_1_51_1","unstructured":"Yuhui Yuan Xilin Chen and Jingdong Wang. 2019. Object-contextual representations for semantic segmentation. arXiv:1909.11065  Yuhui Yuan Xilin Chen and Jingdong Wang. 2019. Object-contextual representations for semantic segmentation. arXiv:1909.11065"},{"key":"e_1_2_1_52_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00064"}],"container-title":["ACM Transactions on Intelligent Systems and Technology"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3446345","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3446345","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T20:47:05Z","timestamp":1750193225000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3446345"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,2,26]]},"references-count":52,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2021,4,30]]}},"alternative-id":["10.1145\/3446345"],"URL":"https:\/\/doi.org\/10.1145\/3446345","relation":{},"ISSN":["2157-6904","2157-6912"],"issn-type":[{"value":"2157-6904","type":"print"},{"value":"2157-6912","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,2,26]]},"assertion":[{"value":"2020-12-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2020-12-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2021-02-26","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}