{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T04:17:25Z","timestamp":1750220245237,"version":"3.41.0"},"publisher-location":"New York, NY, USA","reference-count":29,"publisher":"ACM","license":[{"start":{"date-parts":[[2021,11,19]],"date-time":"2021-11-19T00:00:00Z","timestamp":1637280000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"State Grid Shandong Electric Power Company Laiwu Power Supply Company","award":["SGSDLW00SDJS2100417"],"award-info":[{"award-number":["SGSDLW00SDJS2100417"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2021,11,19]]},"DOI":"10.1145\/3503961.3503972","type":"proceedings-article","created":{"date-parts":[[2022,3,8]],"date-time":"2022-03-08T22:04:56Z","timestamp":1646777096000},"page":"65-69","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["A Multi-scale Framework for Visual Grounding"],"prefix":"10.1145","author":[{"given":"Baosheng","family":"Li","sequence":"first","affiliation":[{"name":"State Grid Shandong Electric Power Company Laiwu Power Supply Company, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Peng","family":"Qi","sequence":"additional","affiliation":[{"name":"State Grid Shandong Electric Power Company Laiwu Power Supply Company, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jian","family":"Wang","sequence":"additional","affiliation":[{"name":"State Grid Shandong Electric Power Company Laiwu Power Supply Company, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Chong","family":"Tan","sequence":"additional","affiliation":[{"name":"State Grid Shandong Electric Power Company Laiwu Power Supply Company, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Runze","family":"Qi","sequence":"additional","affiliation":[{"name":"State Grid Shandong Electric Power Company Laiwu Power Supply Company, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xiaolei","family":"Li","sequence":"additional","affiliation":[{"name":"School of Control Science and Engineering, Shandong University, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2022,3,8]]},"reference":[{"key":"e_1_3_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00477"},{"key":"e_1_3_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00205"},{"volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (pp. 1960-1968)","author":"Wang P.","key":"e_1_3_2_1_3_1","unstructured":"Wang , P. , Wu , Q. , Cao , J. , Shen , C. , Gao , L. , and Hengel , A. V. D. (2019). Neighbourhood watch: Referring expression comprehension via language-guided graph attention networks . In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (pp. 1960-1968) . Wang, P., Wu, Q., Cao, J., Shen, C., Gao, L., and Hengel, A. V. D. (2019). Neighbourhood watch: Referring expression comprehension via language-guided graph attention networks. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (pp. 1960-1968)."},{"key":"e_1_3_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00427"},{"key":"e_1_3_2_1_5_1","first-page":"11778","article-title":"Online decision based visual tracking via reinforcement learning","volume":"3","author":"Song K.","year":"2020","unstructured":"Song , K. , Zhang , W. , Song , R. and Li , Y. ( 2020 ). Online decision based visual tracking via reinforcement learning . Advances in Neural Information Processing Systems (NeurIPS) , 3 , pp. 11778 - 11788 .. Song, K., Zhang, W., Song, R. and Li, Y. (2020). Online decision based visual tracking via reinforcement learning. Advances in Neural Information Processing Systems (NeurIPS), 3, pp. 11778-11788..","journal-title":"Advances in Neural Information Processing Systems (NeurIPS)"},{"key":"e_1_3_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2021.3104183"},{"key":"e_1_3_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00438"},{"key":"e_1_3_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.95"},{"key":"e_1_3_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00479"},{"key":"e_1_3_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00478"},{"key":"e_1_3_2_1_11_1","volume-title":"Proceedings, Part XIV 16 (pp. 387-404)","author":"Yang Z.","year":"2020","unstructured":"Yang , Z. , Chen , T. , Wang , L. , and Luo , J . ( 2020 ). Improving one-stage visual grounding by recursive sub-query construction. In Computer Vision\u2013ECCV 2020: 16th European Conference, Glasgow, UK, August 23\u201328, 2020 , Proceedings, Part XIV 16 (pp. 387-404) . Springer International Publishing. Yang, Z., Chen, T., Wang, L., and Luo, J. (2020). Improving one-stage visual grounding by recursive sub-query construction. In Computer Vision\u2013ECCV 2020: 16th European Conference, Glasgow, UK, August 23\u201328, 2020, Proceedings, Part XIV 16 (pp. 387-404). Springer International Publishing."},{"key":"e_1_3_2_1_12_1","volume-title":"Learning Landmark Features for One-Stage Visual Grounding.\" Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Huang","year":"2021","unstructured":"Huang , Binbin, \" Look Before You Leap : Learning Landmark Features for One-Stage Visual Grounding.\" Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition . 2021 . Huang, Binbin, \"Look Before You Leap: Learning Landmark Features for One-Stage Visual Grounding.\" Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 2021."},{"key":"e_1_3_2_1_13_1","volume-title":"An incremental improvement.\" arXiv preprint arXiv:1804.02767","author":"Redmon","year":"2018","unstructured":"Redmon , Joseph, and Ali Farhadi . \"Yolov3 : An incremental improvement.\" arXiv preprint arXiv:1804.02767 ( 2018 ). Redmon, Joseph, and Ali Farhadi. \"Yolov3: An incremental improvement.\" arXiv preprint arXiv:1804.02767 (2018)."},{"key":"e_1_3_2_1_14_1","volume-title":"Single shot multibox detector.\" European conference on computer vision","author":"Liu","year":"2016","unstructured":"Liu , Wei, \" Ssd : Single shot multibox detector.\" European conference on computer vision . Springer , Cham , 2016 . Liu, Wei, \"Ssd: Single shot multibox detector.\" European conference on computer vision. Springer, Cham, 2016."},{"key":"e_1_3_2_1_15_1","doi-asserted-by":"crossref","unstructured":"Girshick Ross \"Rich feature hierarchies for accurate object detection and semantic segmentation.\" Proceedings of the IEEE conference on computer vision and pattern recognition. 2014.  Girshick Ross \"Rich feature hierarchies for accurate object detection and semantic segmentation.\" Proceedings of the IEEE conference on computer vision and pattern recognition. 2014.","DOI":"10.1109\/CVPR.2014.81"},{"key":"e_1_3_2_1_16_1","unstructured":"Ren Shaoqing \"Faster r-cnn: Towards real-time object detection with region proposal networks.\" Advances in neural information processing systems 28 (2015): 91-99.  Ren Shaoqing \"Faster r-cnn: Towards real-time object detection with region proposal networks.\" Advances in neural information processing systems 28 (2015): 91-99."},{"key":"e_1_3_2_1_17_1","volume-title":"Unified, real-time object detection.\" Proceedings of the IEEE conference on computer vision and pattern recognition","author":"Redmon","year":"2016","unstructured":"Redmon , Joseph, \" You only look once : Unified, real-time object detection.\" Proceedings of the IEEE conference on computer vision and pattern recognition . 2016 . Redmon, Joseph, \"You only look once: Unified, real-time object detection.\" Proceedings of the IEEE conference on computer vision and pattern recognition. 2016."},{"key":"e_1_3_2_1_18_1","volume-title":"better, faster, stronger.\" Proceedings of the IEEE conference on computer vision and pattern recognition","author":"Redmon","year":"2017","unstructured":"Redmon , Joseph, and Ali Farhadi . \"YOLO900 0 : better, faster, stronger.\" Proceedings of the IEEE conference on computer vision and pattern recognition . 2017 . Redmon, Joseph, and Ali Farhadi. \"YOLO9000: better, faster, stronger.\" Proceedings of the IEEE conference on computer vision and pattern recognition. 2017."},{"key":"e_1_3_2_1_19_1","volume-title":"Ieee","author":"Dalal","year":"2005","unstructured":"Dalal , Navneet, and Bill Triggs . \"Histograms of oriented gradients for human detection.\" 2005 IEEE computer society conference on computer vision and pattern recognition (CVPR'05). Vol. 1 . Ieee , 2005 . Dalal, Navneet, and Bill Triggs. \"Histograms of oriented gradients for human detection.\" 2005 IEEE computer society conference on computer vision and pattern recognition (CVPR'05). Vol. 1. Ieee, 2005."},{"key":"e_1_3_2_1_20_1","volume-title":"10491","author":"Lindeberg","year":"2012","unstructured":"Lindeberg , Tony. \"Scale invariant feature transform.\" ( 2012 ): 10491 . Lindeberg, Tony. \"Scale invariant feature transform.\" (2012): 10491."},{"key":"e_1_3_2_1_21_1","volume-title":"154-171","author":"Uijlings RR","year":"2013","unstructured":"Uijlings , Jasper RR , \" Selective search for object recognition.\" International journal of computer vision 104.2 ( 2013 ): 154-171 . Uijlings, Jasper RR, \"Selective search for object recognition.\" International journal of computer vision 104.2 (2013): 154-171."},{"key":"e_1_3_2_1_22_1","doi-asserted-by":"crossref","unstructured":"Dogan Pelin Leonid Sigal and Markus Gross. \"Neural sequential phrase grounding (seqground).\" Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 2019.  Dogan Pelin Leonid Sigal and Markus Gross. \"Neural sequential phrase grounding (seqground).\" Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 2019.","DOI":"10.1109\/CVPR.2019.00430"},{"key":"e_1_3_2_1_23_1","doi-asserted-by":"crossref","unstructured":"He Kaiming \"Deep residual learning for image recognition.\" Proceedings of the IEEE conference on computer vision and pattern recognition. 2016.  He Kaiming \"Deep residual learning for image recognition.\" Proceedings of the IEEE conference on computer vision and pattern recognition. 2016.","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_2_1_24_1","volume-title":"Referring to objects in photographs of natural scenes.\" Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP)","author":"Kazemzadeh","year":"2014","unstructured":"Kazemzadeh , Sahar, \" Referitgame : Referring to objects in photographs of natural scenes.\" Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP) . 2014 . Kazemzadeh, Sahar, \"Referitgame: Referring to objects in photographs of natural scenes.\" Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP). 2014."},{"key":"e_1_3_2_1_25_1","doi-asserted-by":"crossref","unstructured":"Hu Ronghang \"Natural language object retrieval.\" Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2016.  Hu Ronghang \"Natural language object retrieval.\" Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2016.","DOI":"10.1109\/CVPR.2016.493"},{"key":"e_1_3_2_1_26_1","unstructured":"BERT\n  : Pre-training of Deep Bidirectional Transformers for Language Understanding  BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding"},{"key":"e_1_3_2_1_27_1","unstructured":"Zaremba Wojciech Ilya Sutskever and Oriol Vinyals. \"Recurrent neural network regularization.\" arXiv preprint arXiv:1409.2329 (2014).  Zaremba Wojciech Ilya Sutskever and Oriol Vinyals. \"Recurrent neural network regularization.\" arXiv preprint arXiv:1409.2329 (2014)."},{"key":"e_1_3_2_1_28_1","unstructured":"Olah Christopher. \"Understanding lstm networks.\" (2015).  Olah Christopher. \"Understanding lstm networks.\" (2015)."},{"key":"e_1_3_2_1_29_1","volume-title":"PMLR","author":"Mukkamala","year":"2017","unstructured":"Mukkamala , Mahesh Chandra, and Matthias Hein . \"Variants of rmsprop and adagrad with logarithmic regret bounds.\" International Conference on Machine Learning . PMLR , 2017 . Mukkamala, Mahesh Chandra, and Matthias Hein. \"Variants of rmsprop and adagrad with logarithmic regret bounds.\" International Conference on Machine Learning. PMLR, 2017."}],"event":{"name":"VSIP 2021: 2021 3rd International Conference on Video, Signal and Image Processing","acronym":"VSIP 2021","location":"Wuhan China"},"container-title":["2021 3rd International Conference on Video, Signal and Image Processing"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3503961.3503972","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3503961.3503972","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T19:30:32Z","timestamp":1750188632000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3503961.3503972"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,11,19]]},"references-count":29,"alternative-id":["10.1145\/3503961.3503972","10.1145\/3503961"],"URL":"https:\/\/doi.org\/10.1145\/3503961.3503972","relation":{},"subject":[],"published":{"date-parts":[[2021,11,19]]},"assertion":[{"value":"2022-03-08","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}