{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,19]],"date-time":"2026-08-19T18:57:51Z","timestamp":1787165871489,"version":"build-2736575974"},"reference-count":56,"publisher":"Oxford University Press (OUP)","issue":"6","license":[{"start":{"date-parts":[[2025,6,4]],"date-time":"2025-06-04T00:00:00Z","timestamp":1748995200000},"content-version":"vor","delay-in-days":3,"URL":"https:\/\/creativecommons.org\/licenses\/by-nc\/4.0\/"}],"funder":[{"DOI":"10.13039\/100009950","name":"Ministry of Education","doi-asserted-by":"publisher","id":[{"id":"10.13039\/100009950","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100003725","name":"National Research Foundation of Korea","doi-asserted-by":"publisher","id":[{"id":"10.13039\/501100003725","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100014188","name":"Ministry of Science and ICT","doi-asserted-by":"publisher","award":["NRF-2022R1A2C2006170"],"award-info":[{"award-number":["NRF-2022R1A2C2006170"]}],"id":[{"id":"10.13039\/501100014188","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2025,6,4]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>Recent studies propose deep learning-based methods to recognize symbols and text in Piping and Instrumentation Diagrams (P&amp;ID). However, existing approaches use complex processes with separate models for symbol detection, text detection, and text recognition. We propose an integrated model combining symbol-text detection and text recognition modules using a text spotting method. Our model extracts text region features encoded with local character information, enabling a lightweight text recognition module that reduces processing time. The integrated approach allows end-to-end learning between modules, facilitating semantic information transmission and improving overall performance compared to multi-model architecture. When tested on industrial P&amp;ID images, our model achieved high performance with an IoU threshold of 0.5: maximum precision of 0.9763\/0.9527, recall of 0.9521\/0.9075, and F1 score of 0.9640\/0.9295 for symbol-text detection\/text recognition.<\/jats:p>","DOI":"10.1093\/jcde\/qwaf053","type":"journal-article","created":{"date-parts":[[2025,6,3]],"date-time":"2025-06-03T07:34:24Z","timestamp":1748936064000},"page":"55-72","source":"Crossref","is-referenced-by-count":2,"title":["Optimizing image format piping and instrumentation diagram recognition: Integrating symbol and text recognition with a single backbone architecture"],"prefix":"10.1093","volume":"12","author":[{"given":"Junhyung","family":"Byun","sequence":"first","affiliation":[{"name":"Jeonbuk National University Department of Computer Science and Artificial Intelligence, , 567 Baekje-daero, Deokjin-gu, Jeonju 54896 ,","place":["Republic of Korea"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Bonggu","family":"Kang","sequence":"additional","affiliation":[{"name":"Jeonbuk National University Department of Computer Science and Artificial Intelligence, , 567 Baekje-daero, Deokjin-gu, Jeonju 54896 ,","place":["Republic of Korea"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5477-0671","authenticated-orcid":false,"given":"Duhwan","family":"Mun","sequence":"additional","affiliation":[{"name":"School of Mechanical Engineering, Korea University , 145 Anam-ro, Seongbuk-gu, Seoul 02841 ,","place":["Republic of Korea"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Gwang","family":"Lee","sequence":"additional","affiliation":[{"name":"Aster Industry Co., Ltd Department of CAD Development, , 491, Nohae-ro, Nowon-gu, Seoul 01695 ,","place":["Republic of Korea"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9013-2338","authenticated-orcid":false,"given":"Hyungki","family":"Kim","sequence":"additional","affiliation":[{"name":"Division of Computer Convergence, Chungnam National University , 99 Daehak-ro, Yuseong-gu, Daejeon 34134 ,","place":["Republic of Korea"]}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"286","published-online":{"date-parts":[[2025,6,4]]},"reference":[{"key":"2025102210120720100_bib1","doi-asserted-by":"publisher","first-page":"9365","DOI":"10.1109\/CVPR.2019.00959","article-title":"Character region awareness for text detection","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Baek","year":"2019"},{"key":"2025102210120720100_bib2","doi-asserted-by":"publisher","first-page":"2404","DOI":"10.1109\/CVPRW50498.2020.00290","article-title":"CLEval: Character-level evaluation for text detection and recognition tasks","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops","author":"Baek","year":"2020"},{"key":"2025102210120720100_bib3","doi-asserted-by":"publisher","first-page":"71","DOI":"10.1145\/3219819.3219861","article-title":"Rosetta: Large scale system for text detection and recognition in images","volume-title":"Proceedings of the Association for Computing Machinery's Special Interest Group on Knowledge Discovery and Data Mining (ACM SIGKDD) International Conference on Knowledge Discovery & Data Mining","author":"Borisyuk","year":"2018"},{"key":"2025102210120720100_bib4","doi-asserted-by":"publisher","first-page":"1483","DOI":"10.1109\/tpami.2019.2956516","article-title":"Cascade R-CNN: Delving into high quality object detection","volume":"43","author":"Cai","year":"2019","journal-title":"IEEE transactions on pattern analysis and machine intelligence"},{"key":"2025102210120720100_bib5","doi-asserted-by":"publisher","first-page":"213","DOI":"10.1007\/978-3-030-58452-8_13","article-title":"End-to-End object detection with transformers","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Carion","year":"2020"},{"key":"2025102210120720100_bib6","doi-asserted-by":"publisher","first-page":"4939","DOI":"10.1145\/3474085.3475351","article-title":"Disentangle your dense object detector","volume-title":"Proceedings of the Association for Computing Machinery (ACM) international conference on multimedia","author":"Chen","year":"2021"},{"key":"2025102210120720100_bib7","doi-asserted-by":"publisher","first-page":"793","DOI":"10.1093\/jcde\/qwac027","article-title":"A video-based SlowFastMTB model for detection of small amounts of smoke from incipient forest fires","volume":"9","author":"Choi","year":"2022","journal-title":"Journal of Computational Design and Engineering"},{"key":"2025102210120720100_bib8","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2010.11929","article-title":"An image is worth 16\u00d716 words: Transformers for image recognition at scale","author":"Dosovitskiy","year":"2020"},{"key":"2025102210120720100_bib9","doi-asserted-by":"publisher","first-page":"884","DOI":"10.24963\/ijcai.2022\/124","article-title":"SVTR: Scene Text Recognition with a single visual model","author":"Du","year":"2022","journal-title":"International Joint Conference on Artificial Intelligence (IJCAI)"},{"key":"2025102210120720100_bib10","doi-asserted-by":"publisher","first-page":"642","DOI":"10.1007\/s11263-019-01204-1","article-title":"CenterNet: Keypoint triplets for object detection","volume-title":"Internation Journal of Computer Vision","author":"Duan","year":"2020"},{"key":"2025102210120720100_bib11","doi-asserted-by":"publisher","first-page":"7098","DOI":"10.1109\/CVPR46437.2021.00702","article-title":"Read like humans: Autonomous, bidirectional and iterative language modeling for Scene Text Recognition","author":"Fang","year":"2021","journal-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition"},{"key":"2025102210120720100_bib12","doi-asserted-by":"publisher","first-page":"369","DOI":"10.1145\/1143844.1143891","article-title":"Connectionist temporal Classification: Labelling unsegmented sequence data with recurrent neural networks","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Graves","year":"2006"},{"key":"2025102210120720100_bib13","doi-asserted-by":"publisher","first-page":"2961","DOI":"10.1109\/ICCV.2017.322","article-title":"Mask R-CNN","volume-title":"Proceedings of the IEEE international conference on computer vision","author":"He","year":"2017"},{"key":"2025102210120720100_bib14","doi-asserted-by":"publisher","first-page":"770","DOI":"10.1109\/CVPR.2016.90","article-title":"Deep residual learning for image recognition","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"He","year":"2016"},{"key":"2025102210120720100_bib15","doi-asserted-by":"publisher","first-page":"5020","DOI":"10.1109\/CVPR.2018.00527","article-title":"An end-to-end TextSpotter with explicit alignment and attention","author":"He","year":"2018","journal-title":"Proceedings of the IEEE conference on computer vision and pattern recognition"},{"key":"2025102210120720100_bib16","doi-asserted-by":"publisher","first-page":"1735","DOI":"10.1162\/neco.1997.9.8.1735","article-title":"Long short-term memory","volume":"9","author":"Hochreiter","year":"1997","journal-title":"Neural Computation"},{"key":"2025102210120720100_bib17","doi-asserted-by":"publisher","first-page":"1314","DOI":"10.1109\/ICCV.2019.00140","article-title":"Searching for MobileNetV3","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Howard","year":"2019"},{"key":"2025102210120720100_bib18","doi-asserted-by":"publisher","first-page":"2593","DOI":"10.3390\/en12132593","article-title":"A digitization and conversion tool for imaged drawings to intelligent piping and instrumentation diagrams (P&ID)","volume":"12","author":"Kang","year":"2019","journal-title":"Energies"},{"key":"2025102210120720100_bib19","doi-asserted-by":"publisher","first-page":"1298","DOI":"10.1093\/jcde\/qwac056","article-title":"End-to-end digitization of image format piping and instrumentation diagrams at an industrially applicable level","volume":"9","author":"Kim","year":"2022","journal-title":"Journal of Computational Design and Engineering"},{"key":"2025102210120720100_bib20","doi-asserted-by":"publisher","first-page":"115337","DOI":"10.1016\/j.eswa.2021.115337","article-title":"Deep-learning-based recognition of symbols and texts at an industrially applicable level from images of high-density piping and instrumentation diagrams","volume":"183","author":"Kim","year":"2021","journal-title":"Expert Systems with Applications"},{"key":"2025102210120720100_bib21","doi-asserted-by":"publisher","first-page":"84","DOI":"10.1145\/3065386","article-title":"ImageNet classification with deep convolutional neural networks","volume":"60","author":"Krizhevsky","year":"2012","journal-title":"Communications of the Association for Computing Machinery (ACM)"},{"key":"2025102210120720100_bib22","doi-asserted-by":"publisher","first-page":"642","DOI":"10.1007\/S11263-019-01204-1","article-title":"CornerNet: Detecting objects as paired keypoints","volume":"128","author":"Law","year":"2020","journal-title":"International Journal of Computer Vision"},{"key":"2025102210120720100_bib23","doi-asserted-by":"publisher","first-page":"355","DOI":"10.7315\/CDE.2021.355","article-title":"Image format P&ID recognition technique using synthetic data and text-symbol integrated detection","volume":"26","author":"Lee","year":"2021","journal-title":"Korean Journal of Computational Design and Engineering"},{"key":"2025102210120720100_bib24","doi-asserted-by":"publisher","first-page":"5238","DOI":"10.1109\/ICCV.2017.560","article-title":"Towards end-to-end text spotting with convolutional recurrent neural networks","author":"Li","year":"2017","journal-title":"Proceedings of the IEEE international conference on computer vision"},{"key":"2025102210120720100_bib25","doi-asserted-by":"publisher","first-page":"3139","DOI":"10.1109\/tpami.2022.3180392","article-title":"Generalized Focal loss: Learning qualified and distributed bounding boxes for dense object detection","volume":"45","author":"Li","year":"2022","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"2025102210120720100_bib26","doi-asserted-by":"publisher","first-page":"3676","DOI":"10.1109\/TIP.2018.2825107","article-title":"TextBoxes++: A single-shot oriented scene text detector","volume":"27","author":"Liao","year":"2018","journal-title":"IEEE Transactions on Image Processing"},{"key":"2025102210120720100_bib27","doi-asserted-by":"publisher","first-page":"4161","DOI":"10.1609\/aaai.v31i1.11196","article-title":"TextBoxes: A fast text detector with a single deep neural network","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","author":"Liao","year":"2017"},{"key":"2025102210120720100_bib28","doi-asserted-by":"publisher","first-page":"2117","DOI":"10.1109\/CVPR.2017.106","article-title":"Feature Pyramid Networks for object detection","author":"Lin","year":"2017","journal-title":"Proceedings of the IEEE conference on computer vision and pattern recognition"},{"key":"2025102210120720100_bib29","doi-asserted-by":"publisher","first-page":"2999","DOI":"10.1109\/ICCV.2017.324","article-title":"Focal loss for dense object detection","author":"Lin","year":"2017","journal-title":"Proceedings of the IEEE international conference on computer vision"},{"key":"2025102210120720100_bib30","doi-asserted-by":"crossref","first-page":"740","DOI":"10.1007\/978-3-319-10602-1_48","article-title":"Microsoft COCO: Common objects in context","author":"Lin","year":"2014","journal-title":"Computer Vision\u2013ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12"},{"key":"2025102210120720100_bib31","doi-asserted-by":"publisher","first-page":"6459","DOI":"10.1109\/CVPR.2019.00662","article-title":"Adaptive NMS: Refining pedestrian detection in a crowd","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Liu","year":"2019"},{"key":"2025102210120720100_bib32","doi-asserted-by":"publisher","first-page":"640","DOI":"10.1109\/TPAMI.2016.2572683","article-title":"Fully convolutional networks for semantic segmentation","volume":"39","author":"Long","year":"2016","journal-title":"IEEE transactions on pattern analysis and machine intelligence"},{"key":"2025102210120720100_bib33","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.1711.05101","article-title":"Decoupled weight decay regularization","author":"Loshchilov","year":"2017"},{"key":"2025102210120720100_bib34","doi-asserted-by":"publisher","first-page":"176","DOI":"10.1109\/CVPRW50498.2020.00096","article-title":"Automatic digitization of engineering diagrams using Deep Learning and graph search","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops","author":"Mani","year":"2020"},{"key":"2025102210120720100_bib35","doi-asserted-by":"publisher","first-page":"412","DOI":"10.1145\/3453892.3461323","article-title":"Multi-class confusion matrix reduction method and its application on net promoter score classification problem","volume-title":"Proceedings of the 14th PErvasive Technologies Related to Assistive Environments Conference","author":"Markoulidakis","year":"2021"},{"key":"2025102210120720100_bib36","doi-asserted-by":"publisher","first-page":"13528","DOI":"10.1109\/CVPR42600.2020.01354","article-title":"SEED: Semantics enhanced encoder-decoder Framework for scene text recognition","author":"Qiao","year":"2020","journal-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition"},{"key":"2025102210120720100_bib37","doi-asserted-by":"publisher","first-page":"163","DOI":"10.5220\/0007376401630172","article-title":"Automatic information extraction from piping and instrumentation diagrams","volume":"1","author":"Rahul","year":"2019","journal-title":"Proceedings of the International Conference on Pattern Recognition Applications and Methods"},{"key":"2025102210120720100_bib38","doi-asserted-by":"publisher","first-page":"1137","DOI":"10.1109\/TPAMI.2016.2577031","article-title":"Faster R-CNN: Towards real-time object detection with Region Proposal Networks","volume":"39","author":"Ren","year":"2016","journal-title":"\u00a0IEEE transactions on pattern analysis and machine intelligence"},{"key":"2025102210120720100_bib39","doi-asserted-by":"publisher","first-page":"658","DOI":"10.1109\/CVPR.2019.00075","article-title":"Generalized intersection over union: A metric and a loss for bounding box regression","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Rezatofighi","year":"2019"},{"key":"2025102210120720100_bib40","doi-asserted-by":"publisher","first-page":"307","DOI":"10.3233\/ICA-230726","article-title":"Enhancing smart home appliance recognition with wavelet and scalogram analysis using data augmentation","volume":"31","author":"Salazar-Gonz\u00e1lez","year":"2024","journal-title":"Integrated Computer-Aided Engineering"},{"key":"2025102210120720100_bib41","doi-asserted-by":"publisher","first-page":"2298","DOI":"10.1109\/TPAMI.2016.2646371","article-title":"An end-to-end trainable neural network for image-based sequence recognition and its application to scene text recognition","volume":"39","author":"Shi","year":"2016","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"2025102210120720100_bib42","doi-asserted-by":"publisher","first-page":"2035","DOI":"10.1109\/TPAMI.2018.2848939","article-title":"ASTER: An attentional scene text recognizer with flexible rectification","volume":"41","author":"Shi","year":"2018","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"2025102210120720100_bib43","doi-asserted-by":"publisher","first-page":"629","DOI":"10.1109\/ICDAR.2007.4376991","article-title":"An overview of the Tesseract OCR Engine","volume":"2","author":"Smith","year":"2007","journal-title":"Proceedings of the International Conference on Document Analysis and Recognition"},{"key":"2025102210120720100_bib44","doi-asserted-by":"publisher","first-page":"14454","DOI":"10.1109\/CVPR46437.2021.01422","article-title":"Sparse R-CNN: End-to-end object detection with learnable proposals","author":"Sun","year":"2021","journal-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition"},{"key":"2025102210120720100_bib45","doi-asserted-by":"publisher","first-page":"110822","DOI":"10.1016\/j.patcog.2024.110822","article-title":"ITFuse: An interactive transformer for infrared and visible image fusion","volume":"156","author":"Tang","year":"2024","journal-title":"Pattern Recognition"},{"key":"2025102210120720100_bib46","doi-asserted-by":"publisher","first-page":"56","DOI":"10.1007\/978-3-319-46484-8_4","article-title":"Detecting text in natural image with connectionist text proposal network","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Tian","year":"2016"},{"key":"2025102210120720100_bib47","doi-asserted-by":"publisher","first-page":"1922","DOI":"10.1109\/TPAMI.2020.3032166","article-title":"FCOS: A simple and strong anchor-free object detector","volume":"44","author":"Tian","year":"2020","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"2025102210120720100_bib48","first-page":"6000","article-title":"Attention is all you need","author":"Vaswani","year":"2017","journal-title":"Proceedings of the 31st International Conference on Neural Information Processing Systems"},{"key":"2025102210120720100_bib49","doi-asserted-by":"publisher","first-page":"1158","DOI":"10.1093\/jcde\/qwad042","article-title":"An improved YOLOX approach for low-light and small object detection: PPE on tunnel construction sites","volume":"10","author":"Wang","year":"2023","journal-title":"Journal of Computational Design and Engineering"},{"key":"2025102210120720100_bib50","doi-asserted-by":"publisher","first-page":"12113","DOI":"10.1109\/cvpr42600.2020.01213","article-title":"Towards accurate scene text recognition with semantic reasoning networks","author":"Yu","year":"2020","journal-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition"},{"key":"2025102210120720100_bib51","doi-asserted-by":"publisher","first-page":"1616","DOI":"10.1093\/jcde\/qwac071","article-title":"A novel deep convolutional neural network algorithm for surface defect detection","volume":"9","author":"Zhang","year":"2022","journal-title":"Journal of Computational Design and Engineering"},{"key":"2025102210120720100_bib52","doi-asserted-by":"publisher","first-page":"8514","DOI":"10.1109\/cvpr46437.2021.00841","article-title":"VarifocalNet: An IoU-aware dense object detector","author":"Zhang","year":"2021","journal-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition"},{"key":"2025102210120720100_bib53","doi-asserted-by":"publisher","first-page":"9759","DOI":"10.1109\/cvpr42600.2020.00978","article-title":"Bridging the gap between anchor-based and anchor-free detection via adaptive training sample selection","author":"Zhang","year":"2020","journal-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition"},{"key":"2025102210120720100_bib54","first-page":"5551","article-title":"EAST: An efficient and accurate scene text detector","author":"Zhou","year":"2017","journal-title":"Proceedings of the IEEE conference on Computer Vision and Pattern Recognition"},{"key":"2025102210120720100_bib55","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.1506.03184","article-title":"ICDAR 2015 text Reading in the Wild competition","author":"Zhou","year":"2015"},{"key":"2025102210120720100_bib56","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2010.04159","article-title":"Deformable DETR: Deformable transformers for end-to-end object detection","author":"Zhu","year":"2020"}],"container-title":["Journal of Computational Design and Engineering"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/jcde\/advance-article-pdf\/doi\/10.1093\/jcde\/qwaf053\/63436453\/qwaf053.pdf","content-type":"application\/pdf","content-version":"am","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/jcde\/article-pdf\/12\/6\/55\/63436453\/qwaf053.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/jcde\/article-pdf\/12\/6\/55\/63436453\/qwaf053.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,22]],"date-time":"2025-10-22T14:12:15Z","timestamp":1761142335000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/jcde\/article\/12\/6\/55\/8156798"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,6]]},"references-count":56,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2025,6,4]]}},"URL":"https:\/\/doi.org\/10.1093\/jcde\/qwaf053","relation":{},"ISSN":["2288-5048"],"issn-type":[{"value":"2288-5048","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2025,6]]},"published":{"date-parts":[[2025,6]]}}}