{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,17]],"date-time":"2026-04-17T15:55:22Z","timestamp":1776441322440,"version":"3.51.2"},"publisher-location":"New York, NY, USA","reference-count":59,"publisher":"ACM","license":[{"start":{"date-parts":[[2020,10,12]],"date-time":"2020-10-12T00:00:00Z","timestamp":1602460800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2020,10,12]]},"DOI":"10.1145\/3394171.3413583","type":"proceedings-article","created":{"date-parts":[[2020,10,12]],"date-time":"2020-10-12T12:27:35Z","timestamp":1602505655000},"page":"1855-1863","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":25,"title":["Dual Semantic Fusion Network for Video Object Detection"],"prefix":"10.1145","author":[{"given":"Lijian","family":"Lin","sequence":"first","affiliation":[{"name":"Xiamen University, Xiamen, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Haosheng","family":"Chen","sequence":"additional","affiliation":[{"name":"Xiamen University, Xiamen, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Honglun","family":"Zhang","sequence":"additional","affiliation":[{"name":"Applied Research Center (ARC), Tencent PCG, Shenzhen, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jun","family":"Liang","sequence":"additional","affiliation":[{"name":"Xiamen University, Xiamen, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yu","family":"Li","sequence":"additional","affiliation":[{"name":"Applied Research Center (ARC), Tencent PCG, Shenzhen, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ying","family":"Shan","sequence":"additional","affiliation":[{"name":"Applied Research Center (ARC), Tencent PCG, Shenzhen, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hanzi","family":"Wang","sequence":"additional","affiliation":[{"name":"Xiamen University, Xiamen, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2020,10,12]]},"reference":[{"key":"e_1_3_2_2_1_1","volume-title":"Proc. of ECCV. 331--346","author":"Bertasius Gedas","year":"2018","unstructured":"Gedas Bertasius , Lorenzo Torresani , and Jianbo Shi . 2018 . Object detection in video with spatio temporal sampling networks . In Proc. of ECCV. 331--346 . Gedas Bertasius, Lorenzo Torresani, and Jianbo Shi. 2018. Object detection in video with spatio temporal sampling networks. In Proc. of ECCV. 331--346."},{"key":"e_1_3_2_2_2_1","volume-title":"YOLOv4: Optimal speed and accuracy of object detection. arXiv preprint arXiv:2004.10934","author":"Bochkovskiy Alexey","year":"2020","unstructured":"Alexey Bochkovskiy , Chien-YaoWang, and Hong-Yuan Mark Liao . 2020. YOLOv4: Optimal speed and accuracy of object detection. arXiv preprint arXiv:2004.10934 ( 2020 ). Alexey Bochkovskiy, Chien-YaoWang, and Hong-Yuan Mark Liao. 2020. YOLOv4: Optimal speed and accuracy of object detection. arXiv preprint arXiv:2004.10934 (2020)."},{"key":"e_1_3_2_2_3_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00644"},{"key":"e_1_3_2_2_4_1","doi-asserted-by":"publisher","DOI":"10.1145\/3343031.3350581"},{"key":"e_1_3_2_2_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/3240508.3240695"},{"key":"e_1_3_2_2_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00815"},{"key":"e_1_3_2_2_7_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00678"},{"key":"e_1_3_2_2_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"e_1_3_2_2_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00712"},{"key":"e_1_3_2_2_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/3123266.3123455"},{"key":"e_1_3_2_2_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.316"},{"key":"e_1_3_2_2_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00667"},{"key":"e_1_3_2_2_13_1","doi-asserted-by":"publisher","DOI":"10.1145\/3078971.3078990"},{"key":"e_1_3_2_2_14_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.330"},{"key":"e_1_3_2_2_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/3240508.3240544"},{"key":"e_1_3_2_2_16_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2014.81"},{"key":"e_1_3_2_2_17_1","first-page":"2903","article-title":"Distributed and efficient object detection via interactions among devices, edge, and cloud","volume":"21","author":"Guo Yundi","year":"2019","unstructured":"Yundi Guo , Beiji Zou , Ju Ren , Qingqing Liu , Deyu Zhang , and Yaoxue Zhang . 2019 . Distributed and efficient object detection via interactions among devices, edge, and cloud . IEEE TMM 21 , 11 (2019), 2903 -- 2915 . Yundi Guo, Beiji Zou, Ju Ren, Qingqing Liu, Deyu Zhang, and Yaoxue Zhang. 2019. Distributed and efficient object detection via interactions among devices, edge, and cloud. IEEE TMM 21, 11 (2019), 2903--2915.","journal-title":"IEEE TMM"},{"key":"e_1_3_2_2_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/3240508.3240611"},{"key":"e_1_3_2_2_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/3240508.3240693"},{"key":"e_1_3_2_2_20_1","volume-title":"Prajit Ramachandran, Mohammad Babaeizadeh, Honghui Shi, Jianan Li, Shuicheng Yan, and Thomas S Huang.","author":"Han Wei","year":"2016","unstructured":"Wei Han , Pooya Khorrami , Tom Le Paine , Prajit Ramachandran, Mohammad Babaeizadeh, Honghui Shi, Jianan Li, Shuicheng Yan, and Thomas S Huang. 2016 . Seq-NMS for video object detection. arXiv preprint arXiv:1602.08465 (2016). Wei Han, Pooya Khorrami, Tom Le Paine, Prajit Ramachandran, Mohammad Babaeizadeh, Honghui Shi, Jianan Li, Shuicheng Yan, and Thomas S Huang. 2016. Seq-NMS for video object detection. arXiv preprint arXiv:1602.08465 (2016)."},{"key":"e_1_3_2_2_21_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_2_2_22_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00378"},{"key":"e_1_3_2_2_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/3343031.3351064"},{"key":"e_1_3_2_2_24_1","first-page":"2896","article-title":"T-CNN: Tubelets with convolutional neural networks for object detection from videos","volume":"28","author":"Kang Kai","year":"2017","unstructured":"Kai Kang , Hongsheng Li , Junjie Yan , Xingyu Zeng , Bin Yang , Tong Xiao , Cong Zhang , Zhe Wang , Ruohui Wang , Xiaogang Wang , 2017 . T-CNN: Tubelets with convolutional neural networks for object detection from videos . IEEE TCSVT 28 , 10 (2017), 2896 -- 2907 . Kai Kang, Hongsheng Li, Junjie Yan, Xingyu Zeng, Bin Yang, Tong Xiao, Cong Zhang, Zhe Wang, Ruohui Wang, Xiaogang Wang, et al. 2017. T-CNN: Tubelets with convolutional neural networks for object detection from videos. IEEE TCSVT 28, 10 (2017), 2896--2907.","journal-title":"IEEE TCSVT"},{"key":"e_1_3_2_2_25_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01264-9_45"},{"key":"e_1_3_2_2_26_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-48881-3_6"},{"key":"e_1_3_2_2_27_1","first-page":"985","article-title":"Scale-aware fast R-CNN for pedestrian detection","volume":"20","author":"Li Jianan","year":"2017","unstructured":"Jianan Li , Xiaodan Liang , ShengMei Shen , Tingfa Xu , Jiashi Feng , and Shuicheng Yan . 2017 . Scale-aware fast R-CNN for pedestrian detection . IEEE TMM 20 , 4 (2017), 985 -- 996 . Jianan Li, Xiaodan Liang, ShengMei Shen, Tingfa Xu, Jiashi Feng, and Shuicheng Yan. 2017. Scale-aware fast R-CNN for pedestrian detection. IEEE TMM 20, 4 (2017), 985--996.","journal-title":"IEEE TMM"},{"key":"e_1_3_2_2_28_1","first-page":"944","article-title":"Attentive contexts for object detection","volume":"19","author":"Li Jianan","year":"2016","unstructured":"Jianan Li , Yunchao Wei , Xiaodan Liang , Jian Dong , Tingfa Xu , Jiashi Feng , and Shuicheng Yan . 2016 . Attentive contexts for object detection . IEEE TMM 19 , 5 (2016), 944 -- 954 . Jianan Li, Yunchao Wei, Xiaodan Liang, Jian Dong, Tingfa Xu, Jiashi Feng, and Shuicheng Yan. 2016. Attentive contexts for object detection. IEEE TMM 19, 5 (2016), 944--954.","journal-title":"IEEE TMM"},{"key":"e_1_3_2_2_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/3123266.3123343"},{"key":"e_1_3_2_2_30_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.106"},{"key":"e_1_3_2_2_31_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.324"},{"key":"e_1_3_2_2_32_1","volume-title":"Proc. of CVPR. 5686--5695","author":"Liu Mason","year":"2018","unstructured":"Mason Liu and Menglong Zhu . 2018 . Mobile video object detection with temporally-aware feature maps . In Proc. of CVPR. 5686--5695 . Mason Liu and Menglong Zhu. 2018. Mobile video object detection with temporally-aware feature maps. In Proc. of CVPR. 5686--5695."},{"key":"e_1_3_2_2_33_1","volume-title":"Looking fast and slow: Memory-guided mobile video object detection. arXiv preprint arXiv:1903.10172","author":"Liu Mason","year":"2019","unstructured":"Mason Liu , Menglong Zhu , Marie White , Yinxiao Li , and Dmitry Kalenichenko . 2019. Looking fast and slow: Memory-guided mobile video object detection. arXiv preprint arXiv:1903.10172 ( 2019 ). Mason Liu, Menglong Zhu, Marie White, Yinxiao Li, and Dmitry Kalenichenko. 2019. Looking fast and slow: Memory-guided mobile video object detection. arXiv preprint arXiv:1903.10172 (2019)."},{"key":"e_1_3_2_2_34_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46448-0_2"},{"key":"e_1_3_2_2_35_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.257"},{"key":"e_1_3_2_2_36_1","first-page":"761","article-title":"Performance evaluation of object detection algorithms for video surveillance","volume":"8","author":"Nascimento Jacinto C","year":"2006","unstructured":"Jacinto C Nascimento and Jorge S Marques . 2006 . Performance evaluation of object detection algorithms for video surveillance . IEEE TMM 8 , 4 (2006), 761 -- 774 . Jacinto C Nascimento and Jorge S Marques. 2006. Performance evaluation of object detection algorithms for video surveillance. IEEE TMM 8, 4 (2006), 761--774.","journal-title":"IEEE TMM"},{"key":"e_1_3_2_2_37_1","first-page":"1","article-title":"Deep learning for mobile multimedia: A survey","volume":"13","author":"Ota Kaoru","year":"2017","unstructured":"Kaoru Ota , Minh Son Dao , Vasileios Mezaris , and Francesco GB De Natale . 2017 . Deep learning for mobile multimedia: A survey . ACM TOMM 13 , 3s (2017), 1 -- 22 . Kaoru Ota, Minh Son Dao, Vasileios Mezaris, and Francesco GB De Natale. 2017. Deep learning for mobile multimedia: A survey. ACM TOMM 13, 3s (2017), 1--22.","journal-title":"ACM TOMM"},{"key":"e_1_3_2_2_38_1","volume-title":"Hierarchical context features embedding for object Detection","author":"Qiu Heqian","year":"2020","unstructured":"Heqian Qiu , Hongliang Li , Qingbo Wu , Fanman Meng , Linfeng Xu , King N Ngan , and Hengcan Shi . 2020. Hierarchical context features embedding for object Detection . IEEE TMM ( 2020 ). Heqian Qiu, Hongliang Li, Qingbo Wu, Fanman Meng, Linfeng Xu, King N Ngan, and Hengcan Shi. 2020. Hierarchical context features embedding for object Detection. IEEE TMM (2020)."},{"key":"e_1_3_2_2_39_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.91"},{"key":"e_1_3_2_2_40_1","volume-title":"Proc. of NIPS. 91--99","author":"Ren Shaoqing","year":"2015","unstructured":"Shaoqing Ren , Kaiming He , Ross Girshick , and Jian Sun . 2015 . Faster R-CNN: Towards real-time object detection with region proposal networks . In Proc. of NIPS. 91--99 . Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. 2015. Faster R-CNN: Towards real-time object detection with region proposal networks. In Proc. of NIPS. 91--99."},{"key":"e_1_3_2_2_41_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-015-0816-y"},{"key":"e_1_3_2_2_42_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00985"},{"key":"e_1_3_2_2_43_1","volume-title":"Opening the black box of deep neural networks via information. arXiv preprint arXiv:1703.00810","author":"Shwartz-Ziv Ravid","year":"2017","unstructured":"Ravid Shwartz-Ziv and Naftali Tishby . 2017. Opening the black box of deep neural networks via information. arXiv preprint arXiv:1703.00810 ( 2017 ). Ravid Shwartz-Ziv and Naftali Tishby. 2017. Opening the black box of deep neural networks via information. arXiv preprint arXiv:1703.00810 (2017)."},{"key":"e_1_3_2_2_44_1","doi-asserted-by":"publisher","DOI":"10.1145\/3343031.3356076"},{"key":"e_1_3_2_2_45_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2019.2910529"},{"key":"e_1_3_2_2_46_1","first-page":"393","article-title":"Weakly supervised learning of deformable part-based models for object detection via region proposals","volume":"19","author":"Tang Yuxing","year":"2016","unstructured":"Yuxing Tang , Xiaofang Wang , Emmanuel Dellandr\u00e9a , and Liming Chen . 2016 . Weakly supervised learning of deformable part-based models for object detection via region proposals . IEEE TMM 19 , 2 (2016), 393 -- 407 . Yuxing Tang, Xiaofang Wang, Emmanuel Dellandr\u00e9a, and Liming Chen. 2016. Weakly supervised learning of deformable part-based models for object detection via region proposals. IEEE TMM 19, 2 (2016), 393--407.","journal-title":"IEEE TMM"},{"key":"e_1_3_2_2_47_1","first-page":"1135","article-title":"Exploiting web images for weakly supervised object detection","volume":"21","author":"Tao Qingyi","year":"2018","unstructured":"Qingyi Tao , Hao Yang , and Jianfei Cai . 2018 . Exploiting web images for weakly supervised object detection . IEEE TMM 21 , 5 (2018), 1135 -- 1146 . Qingyi Tao, Hao Yang, and Jianfei Cai. 2018. Exploiting web images for weakly supervised object detection. IEEE TMM 21, 5 (2018), 1135--1146.","journal-title":"IEEE TMM"},{"key":"e_1_3_2_2_48_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00972"},{"key":"e_1_3_2_2_49_1","volume-title":"The information bottleneck method. arXiv preprint physics\/0004057","author":"Tishby Naftali","year":"2000","unstructured":"Naftali Tishby , Fernando C Pereira , and William Bialek . 2000. The information bottleneck method. arXiv preprint physics\/0004057 ( 2000 ). Naftali Tishby, Fernando C Pereira, and William Bialek. 2000. The information bottleneck method. arXiv preprint physics\/0004057 (2000)."},{"key":"e_1_3_2_2_50_1","volume-title":"Proc. of ECCV. 542--557","author":"Wang Shiyao","year":"2018","unstructured":"Shiyao Wang , Yucong Zhou , Junjie Yan , and Zhidong Deng . 2018 . Fully motionaware network for video object detection . In Proc. of ECCV. 542--557 . Shiyao Wang, Yucong Zhou, Junjie Yan, and Zhidong Deng. 2018. Fully motionaware network for video object detection. In Proc. of ECCV. 542--557."},{"key":"e_1_3_2_2_51_1","volume-title":"Proc. of CVPR. 7794--7803","author":"Girshick Ross","year":"2018","unstructured":"XiaolongWang, Ross Girshick , Abhinav Gupta , and Kaiming He . 2018 . Non-local neural networks . In Proc. of CVPR. 7794--7803 . XiaolongWang, Ross Girshick, Abhinav Gupta, and Kaiming He. 2018. Non-local neural networks. In Proc. of CVPR. 7794--7803."},{"key":"e_1_3_2_2_52_1","volume-title":"Proc. of CVPR.","author":"Lu Jiwen","year":"2020","unstructured":"ZiweiWang, ZiyiWu, Jiwen Lu , and Jie Zhou . 2020 . BiDet: An efficient binarized object detector . In Proc. of CVPR. ZiweiWang, ZiyiWu, Jiwen Lu, and Jie Zhou. 2020. BiDet: An efficient binarized object detector. In Proc. of CVPR."},{"key":"e_1_3_2_2_53_1","volume-title":"Proc. of ICCV. 9217-- 9225","author":"Chen Yuntao","year":"2019","unstructured":"HaipingWu, Yuntao Chen , NaiyanWang, and Zhaoxiang Zhang . 2019 . Sequence level semantics aggregation for video object detection . In Proc. of ICCV. 9217-- 9225 . HaipingWu, Yuntao Chen, NaiyanWang, and Zhaoxiang Zhang. 2019. Sequence level semantics aggregation for video object detection. In Proc. of ICCV. 9217-- 9225."},{"key":"e_1_3_2_2_54_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01237-3_30"},{"key":"e_1_3_2_2_55_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.634"},{"key":"e_1_3_2_2_56_1","doi-asserted-by":"publisher","DOI":"10.1145\/2964284.2967274"},{"key":"e_1_3_2_2_57_1","doi-asserted-by":"publisher","DOI":"10.1145\/3343031.3351024"},{"key":"e_1_3_2_2_58_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00753"},{"key":"e_1_3_2_2_59_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.52"}],"event":{"name":"MM '20: The 28th ACM International Conference on Multimedia","location":"Seattle WA USA","acronym":"MM '20","sponsor":["SIGMM ACM Special Interest Group on Multimedia"]},"container-title":["Proceedings of the 28th ACM International Conference on Multimedia"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3394171.3413583","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3394171.3413583","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T20:47:14Z","timestamp":1750193234000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3394171.3413583"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,10,12]]},"references-count":59,"alternative-id":["10.1145\/3394171.3413583","10.1145\/3394171"],"URL":"https:\/\/doi.org\/10.1145\/3394171.3413583","relation":{},"subject":[],"published":{"date-parts":[[2020,10,12]]},"assertion":[{"value":"2020-10-12","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}