{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,12]],"date-time":"2025-10-12T03:24:58Z","timestamp":1760239498921,"version":"build-2065373602"},"reference-count":39,"publisher":"MDPI AG","issue":"11","license":[{"start":{"date-parts":[[2020,11,17]],"date-time":"2020-11-17T00:00:00Z","timestamp":1605571200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Scientific Research Program of the Higher Education Institution of Xinjiang","award":["XJEDU2017S043"],"award-info":[{"award-number":["XJEDU2017S043"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Future Internet"],"abstract":"<jats:p>Text detection is a prerequisite for text recognition in scene images. Previous segmentation-based methods for detecting scene text have already achieved a promising performance. However, these kinds of approaches may produce spurious text instances, as they usually confuse the boundary of dense text instances, and then infer word\/text line instances relying heavily on meticulous heuristic rules. We propose a novel Assembling Text Components (AT-text) that accurately detects dense text in scene images. The AT-text localizes word\/text line instances in a bottom-up mechanism by assembling a parsimonious component set. We employ a segmentation model that encodes multi-scale text features, considerably improving the classification accuracy of text\/non-text pixels. The text candidate components are finely classified and selected via discriminate segmentation results. This allows the AT-text to efficiently filter out false-positive candidate components, and then to assemble the remaining text components into different text instances. The AT-text works well on multi-oriented and multi-language text without complex post-processing and character-level annotation. Compared with the existing works, it achieves satisfactory results and a considerable balance between precision and recall without a large margin in ICDAR2013 and MSRA-TD 500 public benchmark datasets.<\/jats:p>","DOI":"10.3390\/fi12110200","type":"journal-article","created":{"date-parts":[[2020,11,16]],"date-time":"2020-11-16T21:48:52Z","timestamp":1605563332000},"page":"200","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":5,"title":["AT-Text: Assembling Text Components for Efficient Dense Scene Text Detection"],"prefix":"10.3390","volume":"12","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-4674-9108","authenticated-orcid":false,"given":"Haiyan","family":"Li","sequence":"first","affiliation":[{"name":"Department of Computer Science and Engineering, Shanghai Jiao Tong University, Shanghai 200240, China"},{"name":"School of Computer Science and Technology, Kashi University, Kashi 844000, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hongtao","family":"Lu","sequence":"additional","affiliation":[{"name":"Department of Computer Science and Engineering, Shanghai Jiao Tong University, Shanghai 200240, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2020,11,17]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Zhang, Z., Shen, W., Yao, C., and Bai, X. (2015, January 7\u201312). Symmetry-Based Text Line Detection in Natural Scenes. Proceedings of the International Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298871"},{"key":"ref_2","unstructured":"Yao, C., Bai, X., Liu, W.Y., Ma, Y., and Tu, Z.W. (2012, January 16\u201321). Detecting Texts of Arbitrary Orientations in Natural Images. Proceedings of the International Conference on Computer Vision and Pattern Recognition, Providence, RI, USA."},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Kang, L., Li, Y., and Doermann, D. (2014, January 23\u201328). Orientation Robust Text Line Detection in Natural Images. Proceedings of the International Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA.","DOI":"10.1109\/CVPR.2014.514"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Jaderberg, M., Vedaldi, A., and Zisserman, A. (2014). Deep features for text spotting. European Conference on Computer Vision, Springer.","DOI":"10.1007\/978-3-319-10593-2_34"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"1930","DOI":"10.1109\/TPAMI.2014.2388210","article-title":"Multi-Orientation Scene Text Detection with Adaptive Clustering","volume":"9","author":"Yin","year":"2015","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"112","DOI":"10.1016\/j.neucom.2017.03.078","article-title":"Natural scene text detection with MC-MR candidate extraction and coarse-to-fine filtering","volume":"260","author":"Tian","year":"2017","journal-title":"Neurocomputing"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Ma, J., Wang, W., Lu, K., and Zhou, J. (2017, January 10\u201314). Scene text detection based on pruning strategy of MSER-trees and Linkage-trees. Proceedings of the International Conference on Multimedia and Expo (ICME), Hong Kong, China.","DOI":"10.1109\/ICME.2017.8019440"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Huang, W., Qiao, Y., and Tang, X. (2014). Robust Scene Text Detection with Convolution Neural Network Induced MSER Trees. The International ECCV 2014, Springer.","DOI":"10.1007\/978-3-319-10593-2_33"},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1016\/j.patcog.2019.06.020","article-title":"SegLink++: Detecting Dense and Arbitrary-shaped Scene Text by Instance-aware Component Grouping","volume":"96","author":"Tang","year":"2019","journal-title":"Pattern Recognit."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"1295","DOI":"10.1109\/LSP.2018.2852954","article-title":"A New Anchor-Labeling Method for Oriented Text Detection Using Dense Detection Framework","volume":"25","author":"Yan","year":"2018","journal-title":"Signal Process. Lett."},{"key":"ref_11","unstructured":"Zhu, A., Du, H., and Xiong, S.W. (2020). Scene Text Detection with Selected Anchor. arXiv."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"2529","DOI":"10.1109\/TIP.2016.2547588","article-title":"Text-Attentional Convolutional Neural Network for Scene Text Detection","volume":"25","author":"He","year":"2016","journal-title":"IEEE Trans. Image Process."},{"key":"ref_13","unstructured":"He, T., Huang, W.L., Qiao, Y., and Yao, J. (2016). Accurate Text Localization in Natural Image with Cascaded Convolutional Text Network. arXiv."},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"970","DOI":"10.1109\/TPAMI.2013.182","article-title":"Robust Text Detection in Natural Scene Images","volume":"36","author":"Yin","year":"2014","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_15","unstructured":"Yao, C., Bai, X., Sang, N., Zhou, X.Y., Zhou, S.C., and Cao, Z.M. (2016). Scene Text Detection via Holistic. Multi-Channel Prediction. arXiv."},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"307","DOI":"10.1016\/j.neucom.2017.01.066","article-title":"A cascaded method for text detection in natural scene images","volume":"238","author":"Zheng","year":"2017","journal-title":"Neurocomputing"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Turki, H., Halima, M.B., and Alimi, A.M. (2017, January 9\u201315). Text Detection Based on MSER and CNN Features. Proceedings of the International Conference on Document Analysis and Recognition (ICDAR), Kyoto, Japan.","DOI":"10.1109\/ICDAR.2017.159"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Cho, H., Sung, M., and Jun, B. (2016, January 27\u201330). Canny Text Detector: Fast and Robust Scene Text Localization Algorithm. Proceedings of the International Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.388"},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"10821","DOI":"10.1007\/s11042-018-6613-1","article-title":"A robust model for salient text detection in natural scene images using MSER feature detector and Grabcut","volume":"78","author":"Gupta","year":"2018","journal-title":"Multimed. Tools Appl."},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"12815","DOI":"10.1007\/s11042-015-3237-6","article-title":"Texture feature-based text region segmentation in social multimedia data","volume":"75","author":"Kim","year":"2016","journal-title":"Multimed. Tools Appl."},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"2906","DOI":"10.1016\/j.patcog.2015.04.002","article-title":"A robust approach for text detection from natural scene images","volume":"48","author":"Sun","year":"2017","journal-title":"Pattern Recognit."},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"7033","DOI":"10.1007\/s11042-017-4619-8","article-title":"Automatic video superimposed text detection based on Nonsubsampled Contourlet Transform","volume":"77","author":"Huang","year":"2017","journal-title":"Multimed. Tools Appl."},{"key":"ref_23","unstructured":"Wang, T., Wu, D.J., Coates, A., and Ng, A.Y. (2012, January 11\u201315). End-to-end text recognition with convolutional neural networks. Proceedings of the International Conference on Pattern Recognition (ICPR), Tsukuba, Japan."},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Epshtein, B., Ofek, E., and Wexler, Y. (2010, January 13\u201318). Detecting Text in Natural Scenes with Stroke Width Transform. Proceedings of the International Conference on Computer Vision and Pattern Recognition, San Francisco, CA, USA.","DOI":"10.1109\/CVPR.2010.5540041"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Subramanian, K., Natarajan, P., Decerbo, M., and Casta\u00f1\u00f2n, D. (2007, January 23\u201326). Character Stroke Detection for Text-Localization and Extraction. Proceedings of the International Conference on Document Analysis and Recognition, Parana, Brazil.","DOI":"10.1109\/ICDAR.2007.4378671"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Dinh, V., Chun, S., Cha, S., Ryu, H., and Sull, S. (2007). An Efficient Method for Text Detection in Video Based on Stroke Width Similarity. ACCV 2007, Springer.","DOI":"10.1007\/978-3-540-76386-4_18"},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"761","DOI":"10.1016\/j.imavis.2004.02.006","article-title":"Robust Wide Baseline Stereo from Maximally Stable Extremal Regions","volume":"22","author":"Matas","year":"2004","journal-title":"Image Vision Comput."},{"key":"ref_28","unstructured":"LeCun, Y., Boser, B., Denker, J.S., Howard, R.E., Habbard, W., Jackel, L.D., and Henderson, D. (1997). Handwritten digit recognition with a back-propagation network. The International Conference on Neural Information Processing Systems, Morgan Kaufman."},{"key":"ref_29","unstructured":"Krizhevsky, A., Sutskever, I., and Hinton, G. (2012). ImageNet Classification with Deep Convolutional Neural Networks. The International Conference on Neural Information Processing Systems, ACM."},{"key":"ref_30","unstructured":"Dalal, N., and Triggs, B. (2005, January 20\u201326). Histograms of oriented gradients for human detection. Proceedings of the International Conference on Computer Vision and Pattern Recognition, San Diego, CA, USA."},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"640","DOI":"10.1109\/TPAMI.2016.2572683","article-title":"Fully Convolutional Networks for Semantic Segmentation","volume":"39","author":"Shelhamer","year":"2016","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_32","unstructured":"Simonyan, K., and Zisserman, A. (2014). Very deep convolu-tional networks for large-scale image recognition. arXiv."},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"105","DOI":"10.1007\/s10032-004-0134-3","article-title":"ICDAR 2003 robust reading competitions: Entries, results, and future directions","volume":"7","author":"Lucas","year":"2005","journal-title":"Int. J. Doc. Anal. Recognit. (IJDAR)"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Karatzas, D., Shafait, F., Uchida, S., Iwamura, M., Bigorda, L.G., Mestre, S.R., Mas, J., Mota, D.F., Almaz\u00e0n, J.A., and Heras, L.P. (2013, January 25\u201328). ICDAR 2013 Robust Reading Competition. Proceedings of the International Conference on Document Analysis and Recognition, Washington, DC, USA.","DOI":"10.1109\/ICDAR.2013.221"},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"280","DOI":"10.1007\/s10032-006-0014-0","article-title":"Object count\/area graphs for the evaluation of object detection and segmentation algorithms","volume":"8","author":"Wolf","year":"2006","journal-title":"Int. J. Doc. Anal. Recognit. (IJDAR)"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Tian, S., Pan, Y., Huang, C., Lu, S., Yu, K., and Tan, C.L. (2015, January 7\u201313). Text Flow: A Unified Text Detection System in Natural Scene Images. Proceedings of the International Conference on Computer Vision (ICCV), Santiago, Chile.","DOI":"10.1109\/ICCV.2015.528"},{"key":"ref_37","doi-asserted-by":"crossref","first-page":"204","DOI":"10.1016\/j.patcog.2016.04.011","article-title":"Could scene context be beneficial for scene text detection?","volume":"58","author":"Zhu","year":"2016","journal-title":"Pattern Recognit."},{"key":"ref_38","doi-asserted-by":"crossref","first-page":"3235","DOI":"10.1109\/TIP.2017.2695104","article-title":"Tracking Based Multi-Orientation Scene Text Detection: A Unified Framework with Dynamic Programming","volume":"26","author":"Yang","year":"2017","journal-title":"IEEE Trans. Image Process."},{"key":"ref_39","doi-asserted-by":"crossref","first-page":"28049","DOI":"10.1007\/s11042-018-5975-8","article-title":"Sign text detection in street view images using an integrated feature","volume":"77","author":"Zhao","year":"2018","journal-title":"Multimed. Tools Appl."}],"container-title":["Future Internet"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1999-5903\/12\/11\/200\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T10:34:20Z","timestamp":1760178860000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1999-5903\/12\/11\/200"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,11,17]]},"references-count":39,"journal-issue":{"issue":"11","published-online":{"date-parts":[[2020,11]]}},"alternative-id":["fi12110200"],"URL":"https:\/\/doi.org\/10.3390\/fi12110200","relation":{},"ISSN":["1999-5903"],"issn-type":[{"type":"electronic","value":"1999-5903"}],"subject":[],"published":{"date-parts":[[2020,11,17]]}}}