{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,30]],"date-time":"2026-05-30T00:00:22Z","timestamp":1780099222120,"version":"3.54.0"},"reference-count":53,"publisher":"MDPI AG","issue":"12","license":[{"start":{"date-parts":[[2024,11,21]],"date-time":"2024-11-21T00:00:00Z","timestamp":1732147200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["J. Imaging"],"abstract":"<jats:p>In the maritime environment, the instance segmentation of small ships is crucial. Small ships are characterized by their limited appearance, smaller size, and ships in distant locations in marine scenes. However, existing instance segmentation algorithms do not detect and segment them, resulting in inaccurate ship segmentation. To address this, we propose a novel solution called enhanced Atrous Spatial Pyramid Pooling (ASPP) feature fusion for small ship instance segmentation. The enhanced ASPP feature fusion module focuses on small objects by refining them and fusing important features. The framework consistently outperforms state-of-the-art models, including Mask R-CNN, Cascade Mask R-CNN, YOLACT, SOLO, and SOLOv2, in three diverse datasets, achieving an average precision (mask AP) score of 75.8% for ShipSG, 69.5% for ShipInsSeg, and 54.5% for the MariBoats datasets.<\/jats:p>","DOI":"10.3390\/jimaging10120299","type":"journal-article","created":{"date-parts":[[2024,11,21]],"date-time":"2024-11-21T12:25:42Z","timestamp":1732191942000},"page":"299","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":5,"title":["Enhanced Atrous Spatial Pyramid Pooling Feature Fusion for Small Ship Instance Segmentation"],"prefix":"10.3390","volume":"10","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-6198-7924","authenticated-orcid":false,"given":"Rabi","family":"Sharma","sequence":"first","affiliation":[{"name":"School of Computer Science, University of Technology Sydney, Broadway, Sydney 2007, Australia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4374-0888","authenticated-orcid":false,"given":"Muhammad","family":"Saqib","sequence":"additional","affiliation":[{"name":"School of Computer Science, University of Technology Sydney, Broadway, Sydney 2007, Australia"},{"name":"National Collections & Marine Infrastructure, CSIRO, Sydney 2007, Australia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8371-8197","authenticated-orcid":false,"given":"C. T.","family":"Lin","sequence":"additional","affiliation":[{"name":"School of Computer Science, University of Technology Sydney, Broadway, Sydney 2007, Australia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9908-3744","authenticated-orcid":false,"given":"Michael","family":"Blumenstein","sequence":"additional","affiliation":[{"name":"School of Computer Science, University of Technology Sydney, Broadway, Sydney 2007, Australia"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2024,11,21]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"He, K., Gkioxari, G., Dollar, P., and Girshick, R. (2017, January 22\u201329). Mask r-cnn. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.322"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Wang, Y., Xu, Z., Wang, X., Shen, C., Cheng, B., Shen, H., and Xia, H. (2021, January 20\u201325). End-to-end video instance segmentation with transformers. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.00863"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"147","DOI":"10.1109\/RBME.2009.2034865","article-title":"Histopathological Image Analysis: A Review","volume":"2","author":"Gurcan","year":"2009","journal-title":"IEEE Rev. Biomed. Eng."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"257","DOI":"10.1016\/j.neucom.2022.01.017","article-title":"Global Mask R-CNN for marine ship instance segmentation","volume":"480","author":"Sun","year":"2022","journal-title":"Neurocomputing"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"127830","DOI":"10.1016\/j.neucom.2024.127830","article-title":"MASSNet: Multiscale Attention for Single-Stage Ship Instance Segmentation","volume":"594","author":"Sharma","year":"2024","journal-title":"Neurocomputing"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Wu, Z., Hou, B., Ren, B., Ren, Z., Wang, S., and Jiao, L. (2021). A deep detection network based on interaction of instance segmentation and object detection for SAR images. Remote Sens., 13.","DOI":"10.3390\/rs13132582"},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"103824","DOI":"10.1016\/j.imavis.2019.11.002","article-title":"An integrated ship segmentation method based on discriminator and extractor","volume":"93","author":"Zhang","year":"2020","journal-title":"Image Vis. Comput."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"1382","DOI":"10.1038\/ng.2452","article-title":"Vertebrate kidney tubules elongate using a planar cell polarity-dependent, rosette-based mechanism of convergent extension","volume":"44","author":"Lienkamp","year":"2012","journal-title":"Nat. Genet."},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Wang, Y., Xu, Z., Shen, H., Cheng, B., and Yang, L. (2020, January 13\u201319). Centermask: Single shot instance segmentation with point representation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00933"},{"key":"ref_10","first-page":"1","article-title":"FactSeg: Foreground activation-driven small object semantic segmentation in large-scale remote sensing imagery","volume":"60","author":"Ma","year":"2021","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_11","first-page":"1","article-title":"Class-incremental learning network for small objects enhancing of semantic segmentation in aerial imagery","volume":"60","author":"Li","year":"2021","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"6708","DOI":"10.1109\/TCSVT.2023.3267127","article-title":"DANet: Dual-branch activation network for small object instance segmentation of ship images","volume":"33","author":"Sun","year":"2023","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Lin, T.Y., Goyal, P., Girshick, R., He, K., and Dollar, P. (2017, January 22\u201329). Focal loss for dense object detection. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.324"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Law, H., and Deng, J. (2018, January 8\u201314). Cornernet: Detecting objects as paired keypoints. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01264-9_45"},{"key":"ref_15","unstructured":"Zhou, X., Wang, D., and Kr\u00e4henbuhl, P. (2019). Objects as points. arXiv."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Lee, Y., and Park, J. (2020, January 13\u201319). Centermask: Real-time anchor-free instance segmentation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.01392"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Xie, E., Sun, P., Song, X., Wang, W., Liu, X., Liang, D., Shen, C., and Luo, P. (2020, January 13\u201319). Polarmask: Single shot instance segmentation with polar representation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.01221"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Chen, H., Sun, K., Tian, Z., Shen, C., Huang, Y., and Yan, Y. (2020, January 13\u201319). Blendmask: Top-down meets bottom-up for instance segmentation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00860"},{"key":"ref_19","unstructured":"Wang, X., Kong, T., Shen, C., Jiang, Y., and Li, L. (2020, January 23\u201328). Solo: Segmenting objects by locations. Proceedings of the Computer Vision\u2013ECCV 2020: 16th European Conference, Glasgow, UK. Proceedings, Part XVIII 16."},{"key":"ref_20","unstructured":"Ren, S., He, K., Girshick, R., and Sun, J. (2015, January 7\u201312). Faster r-cnn: Towards real-time object detection with region proposal networks. Proceedings of the Annual Conference on Neural Information Processing Systems 2015, Montreal, QC, Canada. Advances in Neural Information Processing Systems."},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Liu, S., Qi, L., Qin, H., Shi, J., and Jia, J. (2018, January 18\u201323). Path aggregation network for instance segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00913"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Huang, Z., Huang, L., Gong, Y., Huang, C., and Wang, X. (2019, January 15\u201320). Mask scoring r-cnn. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00657"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Cai, Z., and Vasconcelos, N. (2018, January 18\u201323). Cascade r-cnn: Delving into high quality object detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00644"},{"key":"ref_24","unstructured":"Chen, X., Girshick, R., He, K., and Dollar, P. (November, January 27). Tensormask: A foundation for dense object segmentation. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Seoul, Republic of Korea."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Liu, S., Jia, J., Fidler, S., and Urtasun, R. (2017, January 22\u201329). Sgn: Sequential grouping networks for instance segmentation. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.378"},{"key":"ref_26","unstructured":"Gao, N., Shan, Y., Wang, Y., Zhao, X., Yu, Y., Yang, M., and Huang, K. (November, January 27). Ssap: Single-shot instance segmentation with affinity pyramid. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Seoul, Republic of Korea."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"De Brabandere, B., Neven, D., and Van Gool, L. (2017). Semantic instance segmentation with a discriminative loss function. arXiv.","DOI":"10.1109\/CVPRW.2017.66"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Hu, J., Shen, L., and Sun, G. (2018, January 18\u201323). Squeeze-and-excitation networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00745"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Wang, Q., Wu, B., Zhu, P., Li, P., Zuo, W., and Hu, Q. (2020, January 13\u201319). ECA-Net: Efficient channel attention for deep convolutional neural networks. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.01155"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Woo, S., Park, J., Lee, J.Y., and Kweon, I. (2018, January 8\u201314). Cbam: Convolutional block attention module. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01234-2_1"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Hou, Q., Zhou, D., and Feng, J. (2021, January 20\u201325). Coordinate attention for efficient mobile network design. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.01350"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Nikolio, D., Popovic, Z., Borenovi\u00f3, M., Stojkovi\u00f3, N., Orli\u0107, V., Dzvonkovskaya, A., and Todorovic, B. (2016, January 10\u201312). Multi-radar multi-target tracking algorithm for maritime surveillance at OTH distances. Proceedings of the 2016 17th International Radar Symposium (IRS), Krakow, Poland.","DOI":"10.1109\/IRS.2016.7497299"},{"key":"ref_33","unstructured":"Schwehr, K. (2011). Vessel Tracking Using the Automatic Identification System (AIS) During Emergency Response: Lessons from the Deepwater Horizon Incident, Centre for Coastal and Ocean Mapping\/Joint Hydrographic Centre."},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Sharma, N., Scully-Power, P., and Blumenstein, M. (2018, January 11\u201314). Shark detection from aerial imagery using region-based CNN, a study. Proceedings of the AI 2018: Advances in Artificial Intelligence: 31st Australasian Joint Conference, Wellington, New Zealand. Proceedings 31.","DOI":"10.1007\/978-3-030-03991-2_23"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Saqib, M., Khan, S., Sharma, N., Scully-Power, P., Butcher, P., Colefax, A., and Blumenstein, M. (2018, January 19\u201321). Real-time drone surveillance and population estimation of marine animals from aerial imagery. Proceedings of the 2018 International Conference on Image and Vision Computing New Zealand (IVCNZ), Auckland, New Zealand.","DOI":"10.1109\/IVCNZ.2018.8634661"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Nalamati, M., Kapoor, A., Saqib, M., Sharma, N., and Blumenstein, M. (2019, January 18\u201321). Drone detection in long-range surveillance videos. Proceedings of the 2019 16th IEEE International Conference on Advanced Video and Signal Based Surveillance (AVSS), Taipei, Taiwan.","DOI":"10.1109\/AVSS.2019.8909830"},{"key":"ref_37","unstructured":"Prasad, D., Prasath, C., Rajan, D., Rachmawati, L., Rajabaly, E., and Quek, C. (2016). Challenges in video based object detection in maritime scenario using computer vision. arXiv."},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Zou, Y., Zhao, L., Qin, S., Pan, M., and Li, Z. (2020, January 12\u201314). Ship target detection and identification based on SSD_MobilenetV2. Proceedings of the 2020 IEEE 5th Information Technology and Mechatronics Engineering Conference (ITOEC), Chongqing, China.","DOI":"10.1109\/ITOEC49072.2020.9141734"},{"key":"ref_39","doi-asserted-by":"crossref","first-page":"145","DOI":"10.1016\/j.cja.2020.12.013","article-title":"Ship detection and classification from optical remote sensing images: A survey","volume":"34","author":"Li","year":"2021","journal-title":"Chin. J. Aeronaut."},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Sun, Z., Meng, C., Huang, T., Zhang, Z., and Chang, S. (2023). Marine ship instance segmentation by deep neural networks using a global and local attention (GALA) mechanism. PLoS ONE, 18.","DOI":"10.1371\/journal.pone.0279248"},{"key":"ref_41","doi-asserted-by":"crossref","first-page":"5175","DOI":"10.1109\/TIP.2020.2976856","article-title":"Small object augmentation of urban scenes for real-time semantic segmentation","volume":"29","author":"Yang","year":"2020","journal-title":"IEEE Trans. Image Process."},{"key":"ref_42","first-page":"1","article-title":"A multi-task framework for infrared small target detection and segmentation","volume":"60","author":"Chen","year":"2022","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Li, J., Liang, X., Wei, Y., Xu, T., Feng, J., and Yan, S. (2017, January 21\u201326). Perceptual generative adversarial networks for small object detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.211"},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"Chen, L., Fu, Y., You, S., and Liu, H. (2021). Efficient hybrid supervision for instance segmentation in aerial images. Remote Sens., 13.","DOI":"10.3390\/rs13020252"},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Lin, T.Y., Doll\u00e1r, P., Girshick, R., He, K., Hariharan, B., and Belongie, S. (2017, January 21\u201326). Feature pyramid networks for object detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.106"},{"key":"ref_46","doi-asserted-by":"crossref","first-page":"834","DOI":"10.1109\/TPAMI.2017.2699184","article-title":"Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs","volume":"40","author":"Chen","year":"2017","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_47","doi-asserted-by":"crossref","unstructured":"Grauman, K., and Darrell, T. (2005, January 17\u201321). The pyramid match kernel: Discriminative classification with sets of image features. Proceedings of the Tenth IEEE International Conference on Computer Vision (ICCV\u201905), Beijing, China.","DOI":"10.1109\/ICCV.2005.239"},{"key":"ref_48","doi-asserted-by":"crossref","first-page":"1904","DOI":"10.1109\/TPAMI.2015.2389824","article-title":"Spatial pyramid pooling in deep convolutional networks for visual recognition","volume":"37","author":"He","year":"2015","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_49","doi-asserted-by":"crossref","unstructured":"Carrillo-Perez, B., Barnes, S., and Stephan, M. (2022). Ship segmentation and georeferencing from static oblique view images. Sensors, 22.","DOI":"10.3390\/s22072713"},{"key":"ref_50","unstructured":"Sharma, R., Saqib, M., Lin, C., and Blumenstein, M. (2023, January 23\u201324). Maritime Surveillance Using Instance Segmentation Techniques. Proceedings of the International Conference on Data Science and Communication, Siliguri, India."},{"key":"ref_51","doi-asserted-by":"crossref","first-page":"157","DOI":"10.1007\/s11263-007-0090-8","article-title":"LabelMe: A database and web-based tool for image annotation","volume":"77","author":"Russell","year":"2008","journal-title":"Int. J. Comput. Vis."},{"key":"ref_52","unstructured":"Bolya, D., Zhou, C., Xiao, F., and Lee, Y. (November, January 27). Yolact: Real-time instance segmentation. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Seoul, Republic of Korea."},{"key":"ref_53","first-page":"17721","article-title":"Solov2: Dynamic and fast instance segmentation","volume":"33","author":"Wang","year":"2020","journal-title":"Adv. Neural Inf. Process. Syst."}],"container-title":["Journal of Imaging"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2313-433X\/10\/12\/299\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T16:37:08Z","timestamp":1760114228000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2313-433X\/10\/12\/299"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,11,21]]},"references-count":53,"journal-issue":{"issue":"12","published-online":{"date-parts":[[2024,12]]}},"alternative-id":["jimaging10120299"],"URL":"https:\/\/doi.org\/10.3390\/jimaging10120299","relation":{},"ISSN":["2313-433X"],"issn-type":[{"value":"2313-433X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,11,21]]}}}