{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,28]],"date-time":"2025-10-28T18:43:03Z","timestamp":1761676983398,"version":"build-2065373602"},"reference-count":23,"publisher":"MDPI AG","issue":"13","license":[{"start":{"date-parts":[[2021,7,5]],"date-time":"2021-07-05T00:00:00Z","timestamp":1625443200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["No.51874022"],"award-info":[{"award-number":["No.51874022"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"name":"National Key R&amp;D Program of China","award":["No.2018YFB0704304"],"award-info":[{"award-number":["No.2018YFB0704304"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>An improved DETR (detection with transformers) object detection framework is proposed to realize accurate detection and recognition of characters on shipping containers. ResneSt is used as a backbone network with split attention to extract features of different dimensions by multi-channel weight convolution operation, thus increasing the overall feature acquisition ability of the backbone. In addition, multi-scale location encoding is introduced on the basis of the original sinusoidal position encoding model, improving the sensitivity of input position information for the transformer structure. Compared with the original DETR framework, our model has higher confidence regarding accurate detection, with detection accuracy being improved by 2.6%. In a test of character detection and recognition with a self-built dataset, the overall accuracy can reach 98.6%, which meets the requirements of logistics information identification acquisition.<\/jats:p>","DOI":"10.3390\/s21134612","type":"journal-article","created":{"date-parts":[[2021,7,5]],"date-time":"2021-07-05T22:02:04Z","timestamp":1625522524000},"page":"4612","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":11,"title":["An Improved Character Recognition Framework for Containers Based on DETR Algorithm"],"prefix":"10.3390","volume":"21","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-1898-1359","authenticated-orcid":false,"given":"Xiaofang","family":"Zhao","sequence":"first","affiliation":[{"name":"Institute of Cognitive Computing and Intelligent Information, University of Science and Technology Beijing, Beijing 100083, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0055-3065","authenticated-orcid":false,"given":"Peng","family":"Zhou","sequence":"additional","affiliation":[{"name":"Research Institute of Artificial Intelligence, University of Science and Technology Beijing, Beijing 100083, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1809-7413","authenticated-orcid":false,"given":"Ke","family":"Xu","sequence":"additional","affiliation":[{"name":"Collaborative Innovation Center of Steel Technology, University of Science and Technology Beijing, Beijing 100083, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Liyun","family":"Xiao","sequence":"additional","affiliation":[{"name":"Institute of Cognitive Computing and Intelligent Information, University of Science and Technology Beijing, Beijing 100083, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2021,7,5]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"9","DOI":"10.1134\/S1054661816010065","article-title":"A survey of deep learning methods and software tools for image classification and object detection","volume":"26","author":"Druzhkov","year":"2016","journal-title":"Pattern Recognit. Image Anal."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"143","DOI":"10.1007\/s10032-019-00320-5","article-title":"Scene text detection and recognition with advances in deep learning: A survey","volume":"22","author":"Liu","year":"2019","journal-title":"Int. J. Doc. Anal. Recognit."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"1137","DOI":"10.1109\/TPAMI.2016.2577031","article-title":"Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks","volume":"39","author":"Ren","year":"2017","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_4","unstructured":"Redmon, J., and Farhadi, A. (2018). YOLOv3: An Incremental Improvement. arXiv, Available online: https:\/\/arxiv.org\/abs\/1804.02767."},{"key":"ref_5","unstructured":"Bochkovskiy, A., Wang, C.Y., and Liao, H. (2020). YOLOv4: Optimal Speed and Accuracy of Object Detection. arXiv."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Liu, W., Anguelov, D., Erhan, D., Szegedy, C., Reed, S., Fu, C.Y., and Berg, A.C. (2016, January 11\u201314). SSD: Single Shot MultiBox Detector. Proceedings of the European Conference on Computer Vision, Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46448-0_2"},{"key":"ref_7","unstructured":"Feng, F., Yang, Y., Cer, D., Arivazhagan, N., and Wang, W. (2020). Language-agnostic BERT Sentence Embedding. arXiv."},{"key":"ref_8","unstructured":"Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., and Houlsby, N. (2020). An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. arXiv."},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Valanarasu, J., Oza, P., Hacihaliloglu, I., and Patel, V.M. (2021). Medical Transformer: Gated Axial-Attention for Medical Image Segmentation. arXiv.","DOI":"10.1007\/978-3-030-87193-2_4"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Carion, N., Massa, F., Synnaeve, G., Usunier, N., Kirillov, A., and Zagoruyko, S. (2020, January 23\u201328). End-to-End Object Detection with Transformers. Proceedings of the European Conference on Computer Vision, Glasgow, UK.","DOI":"10.1007\/978-3-030-58452-8_13"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep Residual Learning for Image Recognition. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_12","unstructured":"Zhang, H., Wu, C., Zhang, Z., Zhu, Y., Zhang, Z., Lin, H., Sun, Y., He, T., Mueller, J., and Manmatha, R. (2020). ResNeSt: Split-Attention Networks. arXiv."},{"key":"ref_13","unstructured":"Jie, H., Li, S., Gang, S., and Albanie, S. (2018, January 18\u201322). Squeeze-and-Excitation Networks. Proceedings of the IEEE Transactions on Pattern Analysis and Machine Intelligence, Salt Lake City, UT, USA. Available online: https:\/\/openaccess.thecvf.com\/content_cvpr_2018\/html\/Hu_Squeeze-and-Excitation_Networks_CVPR_2018_paper."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Li, X., Wang, W., Hu, X., and Yang, J. (2019, January 15\u201320). Selective Kernel Networks. Proceedings of the 2019 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00060"},{"key":"ref_15","first-page":"5987","article-title":"Aggregated Residual Transformations for Deep Neural Networks","volume":"1","author":"Xie","year":"2017","journal-title":"Comput. Vis. Pattern Recognit."},{"key":"ref_16","unstructured":"Zhu, X., Su, W., Lu, L., Li, B., and Dai, J. (2020). Deformable DETR: Deformable Transformers for End-to-End Object Detection. arXiv."},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"12","DOI":"10.1016\/j.imavis.2019.06.008","article-title":"Design of multi-scale receptive field convolutional neural network for surface inspection of hot rolled steels","volume":"89","author":"Di","year":"2019","journal-title":"Image Vis. Comput."},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"290","DOI":"10.1016\/j.cie.2018.12.043","article-title":"Defect detection of hot rolled steels with a new object detection framework called classification priority network","volume":"128","author":"He","year":"2019","journal-title":"Comput. Ind. Eng."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Tang, G., M\u00fcller, M., Rios, A., and Sennrich, R. (November, January 31). Why Self-Attention? A Targeted Evaluation of Neural Machine Translation Architectures. Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Brussels, Belgium.","DOI":"10.18653\/v1\/D18-1458"},{"key":"ref_20","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, L., and Polosukhin, I. (2017). Attention Is All You Need. arXiv."},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"7","DOI":"10.1002\/nav.20053","article-title":"The Hungarian method for the assignment problem","volume":"52","author":"Kuhn","year":"2010","journal-title":"Nav. Res. Logist."},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Rezatofighi, H., Tsoi, N., Gwak, J.Y., Sadeghian, A., and Savarese, S. (2019, January 15\u201320). Generalized Intersection Over Union: A Metric and a Loss for Bounding Box Regression. Proceedings of the 2019 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00075"},{"key":"ref_23","first-page":"16","article-title":"Simulation algorithm of outdoor natural fog scene based on monocular video","volume":"6","author":"Dong","year":"2012","journal-title":"J. Hebei Univ. Technol."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/21\/13\/4612\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T06:26:20Z","timestamp":1760163980000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/21\/13\/4612"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,7,5]]},"references-count":23,"journal-issue":{"issue":"13","published-online":{"date-parts":[[2021,7]]}},"alternative-id":["s21134612"],"URL":"https:\/\/doi.org\/10.3390\/s21134612","relation":{},"ISSN":["1424-8220"],"issn-type":[{"type":"electronic","value":"1424-8220"}],"subject":[],"published":{"date-parts":[[2021,7,5]]}}}