{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,11,6]],"date-time":"2025-11-06T12:27:20Z","timestamp":1762432040814,"version":"3.41.0"},"reference-count":77,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2021,7,22]],"date-time":"2021-07-22T00:00:00Z","timestamp":1626912000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/100009110","name":"Natural Science Foundation of Xinjiang Province","doi-asserted-by":"crossref","award":["2020D01A73"],"award-info":[{"award-number":["2020D01A73"]}],"id":[{"id":"10.13039\/100009110","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100001809","name":"NSFC","doi-asserted-by":"crossref","award":["61662082, U1703261, 61822208, 61632019, and 61836011"],"award-info":[{"award-number":["61662082, U1703261, 61822208, 61632019, and 61836011"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100004739","name":"Youth Innovation Promotion Association CAS","doi-asserted-by":"crossref","award":["2018497"],"award-info":[{"award-number":["2018497"]}],"id":[{"id":"10.13039\/501100004739","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2021,8,31]]},"abstract":"<jats:p>\n            Recently, many scene text detection algorithms have achieved impressive performance by using convolutional neural networks. However, most of them do not make full use of the context among the hierarchical multi-level features to improve the performance of scene text detection. In this article, we present an efficient multi-level features enhanced cumulative framework based on instance segmentation for scene text detection. At first, we adopt a Multi-Level Features Enhanced Cumulative (\n            <jats:italic>MFEC<\/jats:italic>\n            ) module to capture features of cumulative enhancement of representational ability. Then, a Multi-Level Features Fusion (\n            <jats:italic>MFF<\/jats:italic>\n            ) module is designed to fully integrate both high-level and low-level MFEC features, which can adaptively encode scene text information. To verify the effectiveness of the proposed method, we perform experiments on six public datasets (namely, CTW1500, Total-text, MSRA-TD500, ICDAR2013, ICDAR2015, and MLT2017), and make comparisons with other state-of-the-art methods. Experimental results demonstrate that the proposed Multi-Level Features Enhanced Cumulative Network (MFECN) detector can well handle scene text instances with irregular shapes (i.e., curved, oriented, and horizontal) and achieves better or comparable results.\n          <\/jats:p>","DOI":"10.1145\/3440087","type":"journal-article","created":{"date-parts":[[2021,7,22]],"date-time":"2021-07-22T14:44:29Z","timestamp":1626965069000},"page":"1-22","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":3,"title":["MFECN: Multi-level Feature Enhanced Cumulative Network for Scene Text Detection"],"prefix":"10.1145","volume":"17","author":[{"given":"Zhandong","family":"Liu","sequence":"first","affiliation":[{"name":"University of Science and Technology of China, Hefei, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Wengang","family":"Zhou","sequence":"additional","affiliation":[{"name":"University of Science and Technology of China, Hefei, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Houqiang","family":"Li","sequence":"additional","affiliation":[{"name":"University of Science and Technology of China, Hefei, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2021,7,22]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2017.2699184"},{"key":"e_1_2_1_2_1","volume-title":"Rethinking atrous convolution for semantic image segmentation. arXiv preprint arXiv:1706.05587","author":"Chen Liang-Chieh","year":"2017","unstructured":"Liang-Chieh Chen , George Papandreou , Florian Schroff , and Hartwig Adam . 2017. Rethinking atrous convolution for semantic image segmentation. arXiv preprint arXiv:1706.05587 ( 2017 ). Liang-Chieh Chen, George Papandreou, Florian Schroff, and Hartwig Adam. 2017. Rethinking atrous convolution for semantic image segmentation. arXiv preprint arXiv:1706.05587 (2017)."},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/3231742"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDAR.2017.157"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/3177745"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICPR.2018.8546066"},{"key":"e_1_2_1_7_1","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence (AAAI). 6773\u20136780","author":"Deng Dan","year":"2018","unstructured":"Dan Deng , Haifeng Liu , Xuelong Li , and Deng Cai . 2018 . PixelLink: Detecting scene text via instance segmentation . In Proceedings of the AAAI Conference on Artificial Intelligence (AAAI). 6773\u20136780 . Dan Deng, Haifeng Liu, Xuelong Li, and Deng Cai. 2018. PixelLink: Detecting scene text via instance segmentation. In Proceedings of the AAAI Conference on Artificial Intelligence (AAAI). 6773\u20136780."},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-009-0275-4"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00044"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.254"},{"volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 3519\u20133528","author":"He Dafang","key":"e_1_2_1_11_1","unstructured":"Dafang He , Xiao Yang , Chen Liang , Zihan Zhou , Alexander G. Ororbi , Daniel Kifer , and C. Lee Giles . 2017. Multi-scale FCN with cascaded instance aware segmentation for arbitrary oriented word spotting in the wild . In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 3519\u20133528 . Dafang He, Xiao Yang, Chen Liang, Zihan Zhou, Alexander G. Ororbi, Daniel Kifer, and C. Lee Giles. 2017. Multi-scale FCN with cascaded instance aware segmentation for arbitrary oriented word spotting in the wild. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 3519\u20133528."},{"key":"e_1_2_1_12_1","volume-title":"Mask R-CNN. In Proceedings of the IEEE International Conference on Computer Vision (ICCV). 2961\u20132969","author":"He Kaiming","year":"2017","unstructured":"Kaiming He , Georgia Gkioxari , Piotr Doll\u00e1r , and Ross Girshick . 2017 . Mask R-CNN. In Proceedings of the IEEE International Conference on Computer Vision (ICCV). 2961\u20132969 . Kaiming He, Georgia Gkioxari, Piotr Doll\u00e1r, and Ross Girshick. 2017. Mask R-CNN. In Proceedings of the IEEE International Conference on Computer Vision (ICCV). 2961\u20132969."},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.331"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00527"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.87"},{"volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 3203\u20133212","author":"Hou Qibin","key":"e_1_2_1_17_1","unstructured":"Qibin Hou , Ming-Ming Cheng , Xiaowei Hu , Ali Borji , Zhuowen Tu , and Philip H. S. Torr . 2017. Deeply supervised salient object detection with short connections . In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 3203\u20133212 . Qibin Hou, Ming-Ming Cheng, Xiaowei Hu, Ali Borji, Zhuowen Tu, and Philip H. S. Torr. 2017. Deeply supervised salient object detection with short connections. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 3203\u20133212."},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.529"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00745"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00445"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/3152129"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1109\/WACV.2019.00086"},{"key":"e_1_2_1_23_1","volume-title":"Enhancement of SSD by concatenating feature maps for object detection. arXiv preprint arXiv:1705.09587","author":"Jeong Jisoo","year":"2017","unstructured":"Jisoo Jeong , Hyojin Park , and Nojun Kwak . 2017. Enhancement of SSD by concatenating feature maps for object detection. arXiv preprint arXiv:1705.09587 ( 2017 ). Jisoo Jeong, Hyojin Park, and Nojun Kwak. 2017. Enhancement of SSD by concatenating feature maps for object detection. arXiv preprint arXiv:1705.09587 (2017)."},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDAR.2015.7333942"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDAR.2013.221"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.40"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2017.2725580"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.472"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.5555\/3298023.3298172"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00619"},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.106"},{"key":"e_1_2_1_32_1","volume-title":"Pyramid mask text detector. arXiv preprint arXiv:1903.11800","author":"Liu Jingchao","year":"2019","unstructured":"Jingchao Liu , Xuebo Liu , Jie Sheng , Ding Liang , Xin Li , and Qingjie Liu . 2019. Pyramid mask text detector. arXiv preprint arXiv:1903.11800 ( 2019 ). Jingchao Liu, Xuebo Liu, Jie Sheng, Ding Liang, Xin Li, and Qingjie Liu. 2019. Pyramid mask text detector. arXiv preprint arXiv:1903.11800 (2019)."},{"key":"e_1_2_1_33_1","volume-title":"Detecting text in the wild with deep character embedding network. arXiv preprint arXiv:1901.00363","author":"Liu Jiaming","year":"2019","unstructured":"Jiaming Liu , Chengquan Zhang , Yipeng Sun , Junyu Han , and Errui Ding . 2019. Detecting text in the wild with deep character embedding network. arXiv preprint arXiv:1901.00363 ( 2019 ). Jiaming Liu, Chengquan Zhang, Yipeng Sun, Junyu Han, and Errui Ding. 2019. Detecting text in the wild with deep character embedding network. arXiv preprint arXiv:1901.00363 (2019)."},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00913"},{"volume-title":"Proceedings of the European Conference on Computer Vision (ECCV). 21\u201337","author":"Liu Wei","key":"e_1_2_1_35_1","unstructured":"Wei Liu , Dragomir Anguelov , Dumitru Erhan , Christian Szegedy , Scott Reed , Cheng-Yang Fu , and Alexander C. Berg . 2016. SSD: Single shot multibox detector . In Proceedings of the European Conference on Computer Vision (ECCV). 21\u201337 . Wei Liu, Dragomir Anguelov, Dumitru Erhan, Christian Szegedy, Scott Reed, Cheng-Yang Fu, and Alexander C. Berg. 2016. SSD: Single shot multibox detector. In Proceedings of the European Conference on Computer Vision (ECCV). 21\u201337."},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00595"},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.622"},{"key":"e_1_2_1_38_1","volume-title":"Detecting curve text in the wild: New dataset and new solution. arXiv preprint arXiv:1712.02170","author":"Liu Yuliang","year":"2017","unstructured":"Yuliang Liu , Lianwen Jin , Shuaitao Zhang , and Sheng Zhang . 2017. Detecting curve text in the wild: New dataset and new solution. arXiv preprint arXiv:1712.02170 ( 2017 ). Yuliang Liu, Lianwen Jin, Shuaitao Zhang, and Sheng Zhang. 2017. Detecting curve text in the wild: New dataset and new solution. arXiv preprint arXiv:1712.02170 (2017)."},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00744"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1145\/3356728"},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11042-019-7177-4"},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298965"},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01216-8_2"},{"key":"e_1_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01264-9_5"},{"key":"e_1_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00788"},{"key":"e_1_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2018.2818020"},{"key":"e_1_2_1_47_1","volume-title":"Joseph Chazalon et\u00a0al","author":"Nayef Nibal","year":"2017","unstructured":"Nibal Nayef , Fei Yin , Imen Bizid , Hyunsoo Choi , Yuan Feng , Dimosthenis Karatzas , Zhenbo Luo , Umapada Pal , Christophe Rigaud , Joseph Chazalon et\u00a0al . 2017 . ICDAR2017 robust reading challenge on multi-lingual scene text detection and script identification-RRC-MLT. In Proceedings of the International Conference on Document Analysis and Recognition (ICDAR) . 1454\u20131459. Nibal Nayef, Fei Yin, Imen Bizid, Hyunsoo Choi, Yuan Feng, Dimosthenis Karatzas, Zhenbo Luo, Umapada Pal, Christophe Rigaud, Joseph Chazalon et\u00a0al. 2017. ICDAR2017 robust reading challenge on multi-lingual scene text detection and script identification-RRC-MLT. In Proceedings of the International Conference on Document Analysis and Recognition (ICDAR). 1454\u20131459."},{"key":"e_1_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2016.2577031"},{"key":"e_1_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMI.2018.2867261"},{"key":"e_1_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.371"},{"key":"e_1_2_1_51_1","volume-title":"Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556","author":"Simonyan Karen","year":"2014","unstructured":"Karen Simonyan and Andrew Zisserman . 2014. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556 ( 2014 ). Karen Simonyan and Andrew Zisserman. 2014. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556 (2014)."},{"key":"e_1_2_1_52_1","doi-asserted-by":"publisher","DOI":"10.1145\/3240508.3240688"},{"volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 10781\u201310790","author":"Tan Mingxing","key":"e_1_2_1_53_1","unstructured":"Mingxing Tan , Ruoming Pang , and Quoc V. Le . 2020. EfficientDet: Scalable and efficient object detection . In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 10781\u201310790 . Mingxing Tan, Ruoming Pang, and Quoc V. Le. 2020. EfficientDet: Scalable and efficient object detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 10781\u201310790."},{"key":"e_1_2_1_54_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46484-8_4"},{"key":"e_1_2_1_55_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00150"},{"key":"e_1_2_1_56_1","doi-asserted-by":"publisher","DOI":"10.1145\/3343031.3350988"},{"key":"e_1_2_1_57_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00956"},{"key":"e_1_2_1_58_1","doi-asserted-by":"publisher","DOI":"10.5555\/2722900.2723092"},{"key":"e_1_2_1_59_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v33i01.33019038"},{"key":"e_1_2_1_60_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.634"},{"key":"e_1_2_1_61_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.164"},{"key":"e_1_2_1_62_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2019.2900589"},{"key":"e_1_2_1_63_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01270-0_22"},{"key":"e_1_2_1_64_1","doi-asserted-by":"publisher","DOI":"10.5555\/3367032.3367173"},{"key":"e_1_2_1_65_1","doi-asserted-by":"publisher","DOI":"10.5555\/3304415.3304567"},{"key":"e_1_2_1_66_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2014.2353813"},{"key":"e_1_2_1_67_1","doi-asserted-by":"publisher","DOI":"10.5555\/2354409.2354851"},{"key":"e_1_2_1_68_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2016.2554321"},{"key":"e_1_2_1_69_1","doi-asserted-by":"publisher","DOI":"10.1145\/2964284.2967274"},{"key":"e_1_2_1_70_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46478-7_22"},{"key":"e_1_2_1_71_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00187"},{"key":"e_1_2_1_72_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.31"},{"key":"e_1_2_1_73_1","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence (AAAI). 2612\u20132619","author":"Zhang Sheng","year":"2018","unstructured":"Sheng Zhang , Yuliang Liu , Lianwen Jin , and Canjie Luo . 2018 . Feature enhancement network: A refined scene text detector . In Proceedings of the AAAI Conference on Artificial Intelligence (AAAI). 2612\u20132619 . Sheng Zhang, Yuliang Liu, Lianwen Jin, and Canjie Luo. 2018. Feature enhancement network: A refined scene text detector. In Proceedings of the AAAI Conference on Artificial Intelligence (AAAI). 2612\u20132619."},{"key":"e_1_2_1_74_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.451"},{"key":"e_1_2_1_75_1","doi-asserted-by":"publisher","DOI":"10.5555\/3304415.3304584"},{"key":"e_1_2_1_76_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.283"},{"key":"e_1_2_1_77_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11704-015-4488-0"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3440087","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3440087","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T22:02:18Z","timestamp":1750197738000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3440087"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,7,22]]},"references-count":77,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2021,8,31]]}},"alternative-id":["10.1145\/3440087"],"URL":"https:\/\/doi.org\/10.1145\/3440087","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"type":"print","value":"1551-6857"},{"type":"electronic","value":"1551-6865"}],"subject":[],"published":{"date-parts":[[2021,7,22]]},"assertion":[{"value":"2019-08-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2021-11-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2021-07-22","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}