{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,11]],"date-time":"2026-04-11T18:37:22Z","timestamp":1775932642486,"version":"3.50.1"},"reference-count":25,"publisher":"MDPI AG","issue":"14","license":[{"start":{"date-parts":[[2021,7,16]],"date-time":"2021-07-16T00:00:00Z","timestamp":1626393600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["No.51874022"],"award-info":[{"award-number":["No.51874022"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"name":"National Key R&amp;D Program of China","award":["no.2018YFB0704304"],"award-info":[{"award-number":["no.2018YFB0704304"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>To address the problem of low detection rate caused by the close alignment and multi-directional position of text words in practical application and the need to improve the detection speed of the algorithm, this paper proposes a multi-directional text detection algorithm based on improved YOLOv3, and applies it to natural text detection. To detect text in multiple directions, this paper introduces a method of box definition based on sliding vertices. Then, a new rotating box loss function MD-Closs based on CIOU is proposed to improve the detection accuracy. In addition, a step-by-step NMS method is used to further reduce the amount of calculation. Experimental results show that on the ICDAR 2015 data set, the accuracy rate is 86.2%, the recall rate is 81.9%, and the timeliness is 21.3 fps, which shows that the proposed algorithm has a good detection effect on text detection in natural scenes.<\/jats:p>","DOI":"10.3390\/s21144870","type":"journal-article","created":{"date-parts":[[2021,7,18]],"date-time":"2021-07-18T21:18:52Z","timestamp":1626643132000},"page":"4870","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":17,"title":["Multi-Directional Scene Text Detection Based on Improved YOLOv3"],"prefix":"10.3390","volume":"21","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-4065-2567","authenticated-orcid":false,"given":"Liyun","family":"Xiao","sequence":"first","affiliation":[{"name":"Institute of Engineering Technology, University of Science and Technology Beijing, Beijing 100083, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0055-3065","authenticated-orcid":false,"given":"Peng","family":"Zhou","sequence":"additional","affiliation":[{"name":"Beijing Key Laboratory of Knowledge Engineering for Materials Science, Institute of Artificial Intelligence, University of Science and Technology Beijing, Beijing 100083, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1809-7413","authenticated-orcid":false,"given":"Ke","family":"Xu","sequence":"additional","affiliation":[{"name":"Collaborative Innovation Center of Steel Technology, University of Science and Technology Beijing, Beijing 100083, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1898-1359","authenticated-orcid":false,"given":"Xiaofang","family":"Zhao","sequence":"additional","affiliation":[{"name":"Institute of Engineering Technology, University of Science and Technology Beijing, Beijing 100083, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2021,7,16]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Li, S., and Cao, W. (2021). SEMPANet: A modified path aggregation network with squeeze-excitation for scene text detection. Sensors, 21.","DOI":"10.3390\/s21082657"},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"7068349","DOI":"10.1155\/2018\/7068349","article-title":"Deep learning for computer vision: A brief review","volume":"2018","author":"Voulodimos","year":"2018","journal-title":"Comput. Intell. Neurosci."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"27","DOI":"10.1016\/j.neucom.2015.09.116","article-title":"Deep learning for visual understanding: A review","volume":"187","author":"Guo","year":"2016","journal-title":"Neurocomputing"},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"16423","DOI":"10.1109\/ACCESS.2018.2813319","article-title":"Rail profile measurement based on line-structured light vision","volume":"6","author":"Zhou","year":"2018","journal-title":"IEEE Access"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Redmon, J., Divvala, S., Girshick, R., and Farhadi, A. (2016, January 27\u201330). You only look once: Unified, real-time object detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.91"},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"1137","DOI":"10.1109\/TPAMI.2016.2577031","article-title":"Faster R-CNN: Towards real-time object detection with region proposal networks","volume":"39","author":"Ren","year":"2016","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Redmon, J., and Farhadi, A. (2017, January 21\u201326). YOLO9000: Better, faster, stronger. Proceedings of the IEEE Conference on Computer Vision & Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.690"},{"key":"ref_8","unstructured":"Redmon, J., and Farhadi, A. (2018). Yolov3: An incremental improvement. arXiv."},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Liu, W., Anguelov, D., Erhan, D., Szegedy, C., Reed, S., Fu, C.-Y., and Berg, A.C. (2016). SSD: Single Shot Multibox Detector, Springer.","DOI":"10.1007\/978-3-319-46448-0_2"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Tian, Z., Huang, W., He, T., He, P., and Qiao, Y. (2016). Detecting Text in Natural Image with Connectionist Text Proposal Network, Springer.","DOI":"10.1007\/978-3-319-46484-8_4"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Shi, B., Bai, X., and Belongie, S. (2017, January 21\u201326). Detecting oriented text in natural images by linking segments. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition CVPR, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.371"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Ye, J., Chen, Z., Liu, J., and Du, B. (2020, January 11\u201317). TextFuseNet: Scene text detection with richer fused features. Proceedings of the Twenty-Ninth International Joint Conference on Artificial Intelligence and Seventeenth Pacific Rim International Conference on Artificial Intelligence, Yokohama, Japan.","DOI":"10.24963\/ijcai.2020\/72"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Xing, L., Tian, Z., Huang, W., and Scott, M. (November, January 27). Convolutional character networks. Proceedings of the 2019 IEEE\/CVF International Conference on Computer Vision (ICCV), Seoul, Korea.","DOI":"10.1109\/ICCV.2019.00922"},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"1972","DOI":"10.1007\/s11263-021-01459-7","article-title":"Exploring the capacity of an orderless box discretization network for multi-orientation scene text detection","volume":"129","author":"Liu","year":"2021","journal-title":"Int. J. Comput. Vis."},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"3111","DOI":"10.1109\/TMM.2018.2818020","article-title":"Arbitrary-oriented scene text detection via rotation proposals","volume":"20","author":"Ma","year":"2018","journal-title":"IEEE Trans. Multimed."},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"3676","DOI":"10.1109\/TIP.2018.2825107","article-title":"Textboxes++: A single-shot oriented scene text detector","volume":"27","author":"Liao","year":"2018","journal-title":"IEEE Trans. Image Process."},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"1452","DOI":"10.1109\/TPAMI.2020.2974745","article-title":"Gliding vertex on the horizontal bounding box for multi-oriented object detection","volume":"43","author":"Xu","year":"2020","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"12","DOI":"10.1016\/j.imavis.2019.06.008","article-title":"Design of multi-scale receptive field convolutional neural network for surface inspection of hot rolled steels","volume":"89","author":"He","year":"2019","journal-title":"Image Vis. Comput."},{"key":"ref_19","unstructured":"Zheng, Z., Wang, P., Ren, D., Liu, W., Ye, R., Hu, Q., and Zuo, W. (2020). Enhancing geometric factors in model learning and inference for object detection and instance segmentation. arXiv."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Karatzas, D., Gomez-Bigorda, L., Nicolaou, A., Ghosh, S., Bagdanov, A., Iwamura, M., Matas, J., Neumann, L., Chandrasekhar, V.R., and Lu, S. (2015, January 23\u201326). ICDAR 2015 competition on robust reading. Proceedings of the 2015 13th International Conference on Document Analysis and Recognition, Tunis, Tunisia.","DOI":"10.1109\/ICDAR.2015.7333942"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Hu, H., Zhang, C.Q., Luo, Y.X., Wang, Y.Z., Han, J., and Ding, E. (2017, January 21\u201326). Wordsup: Exploiting word annotations for character based text detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition CVPR, Honolulu, HI, USA.","DOI":"10.1109\/ICCV.2017.529"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Jiang, Y., Zhu, X., Wang, X., Yang, S., Li, W., Wang, H., Fu, P., and Luo, Z. (2017, January 21\u201326). R2CNN: Rotational Region CNN for Orientation Robust Scene Text Detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition CVPR, Honolulu, HI, USA.","DOI":"10.1109\/ICPR.2018.8545598"},{"key":"ref_23","unstructured":"Dan, D., Liu, H., Li, X., and Cai, D. (2017, January 2\u20137). PixelLink: Detecting Scene Text via Instance Segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition CVPR, New Orleans, LA, USA."},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"2013","DOI":"10.1109\/TIP.2019.2946975","article-title":"Learning sparse and identity-preserved hidden attributes for person re-identification","volume":"29","author":"Wang","year":"2019","journal-title":"IEEE Trans. Image Process."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Liao, M., Wan, Z., Yao, C., Chen, K., and Bai, X. (2020, January 7\u201312). Real-time scene text detection with differentiable binarization. Proceedings of the Thirty-Fourth AAAI Conference on Artificial Intel-Ligence, New York, NY, USA.","DOI":"10.1609\/aaai.v34i07.6812"}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/21\/14\/4870\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T06:31:00Z","timestamp":1760164260000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/21\/14\/4870"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,7,16]]},"references-count":25,"journal-issue":{"issue":"14","published-online":{"date-parts":[[2021,7]]}},"alternative-id":["s21144870"],"URL":"https:\/\/doi.org\/10.3390\/s21144870","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,7,16]]}}}