{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,1]],"date-time":"2026-04-01T11:46:02Z","timestamp":1775043962176,"version":"3.50.1"},"reference-count":39,"publisher":"MDPI AG","issue":"4","license":[{"start":{"date-parts":[[2026,4,1]],"date-time":"2026-04-01T00:00:00Z","timestamp":1775001600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Algorithms"],"abstract":"<jats:p>Wireless capsule endoscopy (WCE) plays a vital role in non-invasive screening of small intestinal lesions. However, the automated detection of lesions remains challenging due to low contrast, uneven illumination, and severe visual variability across images. Existing convolutional detectors rely heavily on manually designed anchors and post-processing, while end-to-end detection transformers developed for natural images exhibit limited adaptability to the complex texture and spectral characteristics of WCE data. To overcome these limitations, this study proposes a deep learning-based detection transformer with enhanced relational-zone aggregation for WCE lesion detection, termed ERZA-DETR, specifically tailored for WCE lesion detection. The framework integrates three complementary modules: a Dual-Band Adaptive Fourier Spectral module (DBFS) that recalibrates frequency responses to suppress illumination artifacts and highlight lesion boundaries; a Fused Dual-scale Gated Convolutional module (FD-gConv) that selectively fuses multi-scale texture features; and a Graph-Linked Embedding at Semantic Scales module (GLES) that preserves local topological relationships through coordinate-gated aggregation. Experimental evaluations on the SEE-AI small intestine dataset demonstrate that ERZA-DETR achieves a 3.2% improvement in mAP@50 and a 12.4% reduction in parameters compared with RT-DETRv2, achieving a superior balance between detection accuracy, computational efficiency, and clinical applicability.<\/jats:p>","DOI":"10.3390\/a19040268","type":"journal-article","created":{"date-parts":[[2026,4,1]],"date-time":"2026-04-01T10:09:21Z","timestamp":1775038161000},"page":"268","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["ERZA-DETR: A Deep Learning-Based Detection Transformer with Enhanced Relational-Zone Aggregation for WCE Lesion Detection"],"prefix":"10.3390","volume":"19","author":[{"given":"Shiren","family":"Ye","sequence":"first","affiliation":[{"name":"School of Computer Science and Artificial Intelligence, Changzhou University, Changzhou 213168, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Haipeng","family":"Ma","sequence":"additional","affiliation":[{"name":"School of Computer Science and Artificial Intelligence, Changzhou University, Changzhou 213168, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zetong","family":"Zhang","sequence":"additional","affiliation":[{"name":"School of Computer Science and Artificial Intelligence, Changzhou University, Changzhou 213168, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Liangjing","family":"Li","sequence":"additional","affiliation":[{"name":"School of Computer Science and Artificial Intelligence, Changzhou University, Changzhou 213168, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2026,4,1]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"417","DOI":"10.1038\/35013140","article-title":"Wireless capsule endoscopy","volume":"405","author":"Iddan","year":"2000","journal-title":"Nature"},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"58","DOI":"10.1055\/a-1973-3796","article-title":"Small-bowel capsule endoscopy and device-assisted enteroscopy for diagnosis and treatment of small-bowel disorders: European Society of Gastrointestinal Endoscopy (ESGE) Guideline\u2013Update 2022","volume":"55","author":"Pennazio","year":"2023","journal-title":"Endoscopy"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"e56361","DOI":"10.2196\/56361","article-title":"Diagnostic Accuracy of Artificial Intelligence in Endoscopy: Umbrella Review","volume":"12","author":"Zha","year":"2024","journal-title":"JMIR Med. Inform."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"290","DOI":"10.1111\/den.13896","article-title":"Artificial intelligence and deep learning for small bowel capsule endoscopy","volume":"33","author":"Trasolini","year":"2021","journal-title":"Dig. Endosc."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"4597","DOI":"10.1038\/s41467-024-49019-0","article-title":"Robotic wireless capsule endoscopy: Recent advances and upcoming technologies","volume":"15","author":"Cao","year":"2024","journal-title":"Nat. Commun."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"261","DOI":"10.1016\/j.crohns.2013.09.004","article-title":"Advanced endoscopic imaging techniques in Crohn\u2019s disease","volume":"8","author":"Tontini","year":"2014","journal-title":"J. Crohn\u2019s Colitis"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Son, G., Eo, T., An, J., Oh, D.J., Shin, Y., Rha, H., Kim, Y.J., Lim, Y.J., and Hwang, D. (2022). Small bowel detection for wireless capsule endoscopy using convolutional neural networks with temporal filtering. Diagnostics, 12.","DOI":"10.3390\/diagnostics12081858"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Zhao, X., Fang, C., Gao, F., Fan, D.J., Lin, X., and Li, G. (2021). Deep transformers for fast small intestine grounding in capsule endoscope video. Proceedings of the 2021 IEEE 18th International Symposium on Biomedical Imaging (ISBI), IEEE.","DOI":"10.1109\/ISBI48211.2021.9433921"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Redmon, J., Divvala, S., Girshick, R., and Farhadi, A. (2016). You only look once: Unified, real-time object detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, IEEE.","DOI":"10.1109\/CVPR.2016.91"},{"key":"ref_10","first-page":"2043","article-title":"Multi-Scale Feature Fusion Network Model for Wireless Capsule Endoscopic Intestinal Lesion Detection","volume":"82","author":"Ye","year":"2025","journal-title":"Comput. Mater. Contin."},{"key":"ref_11","unstructured":"Tian, Y., Ye, Q., and Doermann, D. (2025). Yolov12: Attention-centric real-time object detectors. arXiv."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"1137","DOI":"10.1109\/TPAMI.2016.2577031","article-title":"Faster R-CNN: Towards real-time object detection with region proposal networks","volume":"39","author":"Ren","year":"2016","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"He, K., Gkioxari, G., Doll\u00e1r, P., and Girshick, R. (2017). Mask r-cnn. Proceedings of the IEEE International Conference on Computer Vision, IEEE.","DOI":"10.1109\/ICCV.2017.322"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Liu, S., Qi, L., Qin, H., Shi, J., and Jia, J. (2018). Path aggregation network for instance segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, IEEE.","DOI":"10.1109\/CVPR.2018.00913"},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"102141","DOI":"10.1016\/j.artmed.2021.102141","article-title":"Multi-pathology detection and lesion localization in WCE videos by using the instance segmentation approach","volume":"119","author":"Vieira","year":"2021","journal-title":"Artif. Intell. Med."},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1109\/JTEHM.2022.3198819","article-title":"Rat-capsnet: A deep learning network utilizing attention and regional information for abnormality detection in wireless capsule endoscopy","volume":"10","author":"Alam","year":"2022","journal-title":"IEEE J. Transl. Eng. Health Med."},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"104789","DOI":"10.1016\/j.compbiomed.2021.104789","article-title":"A deep CNN model for anomaly detection and localization in wireless capsule endoscopy images","volume":"137","author":"Jain","year":"2021","journal-title":"Comput. Biol. Med."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Carion, N., Massa, F., Synnaeve, G., Usunier, N., Kirillov, A., and Zagoruyko, S. (2020). End-to-end object detection with transformers. Proceedings of the European Conference on Computer Vision, Springer.","DOI":"10.1007\/978-3-030-58452-8_13"},{"key":"ref_19","unstructured":"Zhu, X., Su, W., Lu, L., Li, B., Wang, X., and Dai, J. (2020). Deformable detr: Deformable transformers for end-to-end object detection. arXiv."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Zhao, Y., Lv, W., Xu, S., Wei, J., Wang, G., Dang, Q., Liu, Y., and Chen, J. (2024). Detrs beat yolos on real-time object detection. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, IEEE.","DOI":"10.1109\/CVPR52733.2024.01605"},{"key":"ref_21","unstructured":"Lv, W., Zhao, Y., Chang, Q., Huang, K., Wang, G., and Liu, Y. (2024). RT-DETRv2: Improved Baseline with Bag-of-Freebies for Real-Time Detection Transformer. arXiv."},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Hosain, A.S., Islam, M., Mehedi, M.H.K., Kabir, I.E., and Khan, Z.T. (2022). Gastrointestinal disorder detection with a transformer based approach. Proceedings of the 2022 IEEE 13th Annual Information Technology, Electronics and Mobile Communication Conference (IEMCON), IEEE.","DOI":"10.1109\/IEMCON56893.2022.9946531"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Habe, T.T., Haataja, K., and Toivanen, P. (2025). Precision enhancement in wireless capsule endoscopy: A novel transformer-based approach for real-time video object detection. Front. Artif. Intell., 8.","DOI":"10.3389\/frai.2025.1529814"},{"key":"ref_24","unstructured":"Gu, A., and Dao, T. (2023). Mamba: Linear-Time Sequence Modeling with Selective State Spaces. arXiv."},{"key":"ref_25","unstructured":"Dao, T., and Gu, A. (2024). Transformers are ssms: Generalized models and efficient algorithms through structured state space duality. arXiv."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Nam, J.H., Syazwany, N.S., Kim, S.J., and Lee, S.C. (2024). Modality-agnostic domain generalizable medical image segmentation by multi-frequency in multi-scale attention. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, IEEE.","DOI":"10.1109\/CVPR52733.2024.01091"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Chen, L., Gu, L., Li, L., Yan, C., and Fu, Y. (2025). Frequency Dynamic Convolution for Dense Image Prediction. Proceedings of the Computer Vision and Pattern Recognition Conference, IEEE.","DOI":"10.1109\/CVPR52734.2025.02809"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Liu, Y., Wang, J., Huang, C., Wang, Y., and Xu, Y. (2023). Cigar: Cross-modality graph reasoning for domain adaptive object detection. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, IEEE.","DOI":"10.1109\/CVPR52729.2023.02277"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Im, J., Nam, J., Park, N., Lee, H., and Park, S. (2024). Egtr: Extracting graph from transformer for scene graph generation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, IEEE.","DOI":"10.1109\/CVPR52733.2024.02287"},{"key":"ref_30","unstructured":"Lei, M., Li, S., Wu, Y., Hu, H., Zhou, Y., Zheng, X., Ding, G., Du, S., Wu, Z., and Gao, Y. (2025). Yolov13: Real-time object detection with hypergraph-enhanced adaptive visual perception. arXiv."},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"6896","DOI":"10.1609\/aaai.v39i7.32740","article-title":"HS-FPN: High frequency and spatial perception FPN for tiny object detection","volume":"Volume 39","author":"Shi","year":"2025","journal-title":"Proceedings of the AAAI Conference on Artificial Intelligence"},{"key":"ref_32","first-page":"625","article-title":"RevBiFPN: The fully reversible bidirectional feature pyramid network","volume":"5","author":"Chiley","year":"2023","journal-title":"Proc. Mach. Learn. Syst."},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Lin, T.Y., Doll\u00e1r, P., Girshick, R., He, K., Hariharan, B., and Belongie, S. (2017). Feature pyramid networks for object detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, IEEE.","DOI":"10.1109\/CVPR.2017.106"},{"key":"ref_34","unstructured":"Kipf, T.N., and Welling, M. (2016). Semi-supervised classification with graph convolutional networks. arXiv."},{"key":"ref_35","unstructured":"Veli\u010dkovi\u0107, P., Cucurull, G., Casanova, A., Romero, A., Lio, P., and Bengio, Y. (2017). Graph attention networks. arXiv."},{"key":"ref_36","doi-asserted-by":"crossref","first-page":"505","DOI":"10.1016\/j.neunet.2023.01.051","article-title":"SP-GNN: Learning structure and position information from graphs","volume":"161","author":"Chen","year":"2023","journal-title":"Neural Netw."},{"key":"ref_37","unstructured":"You, J., Ying, R., and Leskovec, J. (2019). Position-aware graph neural networks. Proceedings of the International Conference on Machine Learning, PMLR."},{"key":"ref_38","doi-asserted-by":"crossref","first-page":"e258","DOI":"10.1002\/deo2.258","article-title":"Small bowel capsule endoscopy examination and open access database with artificial intelligence: The SEE-artificial intelligence project","volume":"4","author":"Yokote","year":"2024","journal-title":"DEN Open"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Liu, W., Anguelov, D., Erhan, D., Szegedy, C., Reed, S., Fu, C.Y., and Berg, A.C. (2016). Ssd: Single shot multibox detector. Proceedings of the European Conference on Computer Vision, Springer.","DOI":"10.1007\/978-3-319-46448-0_2"}],"container-title":["Algorithms"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1999-4893\/19\/4\/268\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,1]],"date-time":"2026-04-01T10:14:40Z","timestamp":1775038480000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1999-4893\/19\/4\/268"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,4,1]]},"references-count":39,"journal-issue":{"issue":"4","published-online":{"date-parts":[[2026,4]]}},"alternative-id":["a19040268"],"URL":"https:\/\/doi.org\/10.3390\/a19040268","relation":{},"ISSN":["1999-4893"],"issn-type":[{"value":"1999-4893","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,4,1]]}}}