{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,21]],"date-time":"2026-04-21T20:50:46Z","timestamp":1776804646479,"version":"3.51.2"},"reference-count":62,"publisher":"MDPI AG","issue":"2","license":[{"start":{"date-parts":[[2021,1,13]],"date-time":"2021-01-13T00:00:00Z","timestamp":1610496000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Remote Sensing"],"abstract":"<jats:p>Instance segmentation in aerial images is of great significance for remote sensing applications, and it is inherently more challenging because of cluttered background, extremely dense and small objects, and objects with arbitrary orientations. Besides, current mainstream CNN-based methods often suffer from the trade-off between labeling cost and performance. To address these problems, we present a pipeline of hybrid supervision. In the pipeline, we design an ancillary segmentation model with the bounding box attention module and bounding box filter module. It is able to generate accurate pseudo pixel-wise labels from real-world aerial images for training any instance segmentation models. Specifically, bounding box attention module can effectively suppress the noise in cluttered background and improve the capability of segmenting small objects. Bounding box filter module works as a filter which removes the false positives caused by cluttered background and densely distributed objects. Our ancillary segmentation model can locate object pixel-wisely instead of relying on horizontal bounding box prediction, which has better adaptability to arbitrary oriented objects. Furthermore, oriented bounding box labels are utilized for handling arbitrary oriented objects. Experiments on iSAID dataset show that the proposed method can achieve comparable performance (32.1 AP) to fully supervised methods (33.9 AP), which is obviously higher than weakly supervised setting (26.5 AP), when using only 10% pixel-wise labels.<\/jats:p>","DOI":"10.3390\/rs13020252","type":"journal-article","created":{"date-parts":[[2021,1,13]],"date-time":"2021-01-13T21:50:54Z","timestamp":1610574654000},"page":"252","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":18,"title":["Efficient Hybrid Supervision for Instance Segmentation in Aerial Images"],"prefix":"10.3390","volume":"13","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-9639-3237","authenticated-orcid":false,"given":"Linwei","family":"Chen","sequence":"first","affiliation":[{"name":"School of Computer Science and Technology, Beijing Institute of Technology, Beijing 100081, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6677-694X","authenticated-orcid":false,"given":"Ying","family":"Fu","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Beijing Institute of Technology, Beijing 100081, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8973-645X","authenticated-orcid":false,"given":"Shaodi","family":"You","sequence":"additional","affiliation":[{"name":"Computer Vision Research Group, Institute of Informatics, University of Amsterdam, 1098 XH Amsterdam, The Netherlands"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2314-5272","authenticated-orcid":false,"given":"Hongzhe","family":"Liu","sequence":"additional","affiliation":[{"name":"Beijing Key Laboratory of Information Service Engineering, Beijing Union University, Beijing 100101, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2021,1,13]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"609","DOI":"10.1109\/TGRS.2015.2463075","article-title":"A novel automatic change detection method for urban high-resolution remotely sensed imagery based on multiindex scene representation","volume":"54","author":"Wen","year":"2015","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"881","DOI":"10.1109\/TGRS.2016.2616585","article-title":"Dense semantic labeling of subdecimeter resolution images with convolutional neural networks","volume":"55","author":"Volpi","year":"2016","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Kopsiaftis, G., and Karantzalos, K. (2015, January 26\u201331). Vehicle detection and traffic density monitoring from very high resolution satellite video data. Proceedings of the 2015 IEEE International Geoscience and Remote Sensing Symposium (IGARSS), Milan, Italy.","DOI":"10.1109\/IGARSS.2015.7326160"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"He, K., Gkioxari, G., Doll\u00e1r, P., and Girshick, R. (2017, January 22\u201329). Mask R-CNN. Proceedings of the International Conference on Computer Vision (ICCV), Venice, Italy.","DOI":"10.1109\/ICCV.2017.322"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Liu, S., Qi, L., Qin, H., Shi, J., and Jia, J. (2018, January 18\u201322). Path aggregation network for instance segmentation. Proceedings of the Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00913"},{"key":"ref_6","unstructured":"Chen, K., Pang, J., Wang, J., Xiong, Y., Li, X., Sun, S., Feng, W., Liu, Z., Shi, J., and Ouyang, W. (November, January 27). Hybrid task cascade for instance segmentation. Proceedings of the International Conference on Computer Vision, Seoul, Korea."},{"key":"ref_7","unstructured":"Bolya, D., Zhou, C., Xiao, F., and Lee, Y.J. (November, January 27). YOLACT: Real-time instance segmentation. Proceedings of the International Conference on Computer Vision, Seoul, Korea."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Lee, Y., and Park, J. (2019, January 15\u201320). CenterMask: Real-Time Anchor-Free Instance Segmentation. Proceedings of the Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR42600.2020.01392"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Chen, H., Sun, K., Tian, Z., Shen, C., Huang, Y., and Yan, Y. (2020, January 13\u201319). BlendMask: Top-down meets bottom-up for instance segmentation. Proceedings of the Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00860"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Bearman, A., Russakovsky, O., Ferrari, V., and Li, F.-F. (2016, January 8\u201316). What\u2019s the point: Semantic segmentation with point supervision. Proceedings of the European Conference on Computer Vision, Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46478-7_34"},{"key":"ref_11","unstructured":"Waqas Zamir, S., Arora, A., Gupta, A., Khan, S., Sun, G., Shahbaz Khan, F., Zhu, F., Shao, L., Xia, G.S., and Bai, X. (2019, January 27\u201328). iSAID: A Large-scale Dataset for Instance Segmentation in Aerial Images. Proceedings of the International Conference on Computer Vision Workshop, Seoul, Korea."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"303","DOI":"10.1007\/s11263-009-0275-4","article-title":"The pascal visual object classes (voc) challenge","volume":"88","author":"Everingham","year":"2010","journal-title":"Int. J. Comput. Vision"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Zhou, Y., Zhu, Y., Ye, Q., Qiu, Q., and Jiao, J. (2018, January 18\u201322). Weakly supervised instance segmentation using class peak response. Proceedings of the Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00399"},{"key":"ref_14","unstructured":"Ahn, J., Cho, S., and Kwak, S. (November, January 27). Weakly Supervised Learning of Instance Segmentation with Inter-pixel Relations. Proceedings of the International Conference on Computer Vision, Seoul, Korea."},{"key":"ref_15","unstructured":"Ge, W., Guo, S., Huang, W., and Scott, M.R. (November, January 27). Label-PEnet: Sequential Label Propagation and Enhancement Networks for Weakly Supervised Instance Segmentation. Proceedings of the International Conference on Computer Vision, Seoul, Korea."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Khoreva, A., Benenson, R., Hosang, J., Hein, M., and Schiele, B. (2017, January 22\u201329). Simple does it: Weakly supervised instance and semantic segmentation. Proceedings of the International Conference on Computer Vision (ICCV), Venice, Italy.","DOI":"10.1109\/CVPR.2017.181"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Li, Q., Arnab, A., and Torr, P.H. (2018, January 8\u201314). Weakly-and semi-supervised panoptic segmentation. Proceedings of the European Conference on Computer Vision, Munich, Germany.","DOI":"10.1007\/978-3-030-01267-0_7"},{"key":"ref_18","unstructured":"Hsu, C.C., Hsu, K.J., Tsai, C.C., Lin, Y.Y., and Chuang, Y.Y. (2019, January 8\u201314). Weakly Supervised Instance Segmentation using the Bounding Box Tightness Prior. Proceedings of the Advances in Neural Information Processing Systems, Vancouver, BC, Canada."},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"309","DOI":"10.1145\/1015706.1015720","article-title":"\u201cGrabCut\u201d interactive foreground extraction using iterated graph cuts","volume":"23","author":"Rother","year":"2004","journal-title":"ACM Trans. Graph."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Arbel\u00e1ez, P., Pont-Tuset, J., Barron, J.T., Marques, F., and Malik, J. (2014, January 23\u201328). Multiscale combinatorial grouping. Proceedings of the Computer Vision and Pattern Recognition (CVPR), Columbus, OH, USA.","DOI":"10.1109\/CVPR.2014.49"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Long, J., Shelhamer, E., and Darrell, T. (2015, January 7\u201313). Fully convolutional networks for semantic segmentation. Proceedings of the International Conference on Computer Vision, Santiago, Chile.","DOI":"10.1109\/CVPR.2015.7298965"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Chen, L.C., Zhu, Y., Papandreou, G., Schroff, F., and Adam, H. (2018, January 8\u201314). Encoder-decoder with atrous separable convolution for semantic image segmentation. Proceedings of the European Conference on Computer Vision, Munich, Germany.","DOI":"10.1007\/978-3-030-01234-2_49"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Yang, M., Yu, K., Zhang, C., Li, Z., and Yang, K. (2018, January 18\u201322). Denseaspp for semantic segmentation in street scenes. Proceedings of the Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00388"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Wu, T., Tang, S., Zhang, R., Cao, J., and Li, J. (2019, January 8\u201312). Tree-structured kronecker convolutional network for semantic segmentation. Proceedings of the International Conference on Multimedia and Expo, Shanghai, China.","DOI":"10.1109\/ICME.2019.00166"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Zhao, H., Shi, J., Qi, X., Wang, X., and Jia, J. (2017, January 22\u201329). Pyramid scene parsing network. Proceedings of the International Conference on Computer Vision (ICCV), Venice, Italy.","DOI":"10.1109\/CVPR.2017.660"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Lin, G., Milan, A., Shen, C., and Reid, I. (2017, January 22\u201329). Refinenet: Multi-path refinement networks for high-resolution semantic segmentation. Proceedings of the International Conference on Computer Vision (ICCV), Venice, Italy.","DOI":"10.1109\/CVPR.2017.549"},{"key":"ref_27","unstructured":"Fu, J., Liu, J., Tian, H., Li, Y., Bao, Y., Fang, Z., and Lu, H. (November, January 27). Dual attention network for scene segmentation. Proceedings of the International Conference on Computer Vision, Seoul, Korea."},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Hung, W.C., Tsai, Y.H., Shen, X., Lin, Z., Sunkavalli, K., Lu, X., and Yang, M.H. (2017, January 22\u201329). Scene parsing with global context embedding. Proceedings of the International Conference on Computer Vision (ICCV), Venice, Italy.","DOI":"10.1109\/ICCV.2017.287"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Zhang, H., Dana, K., Shi, J., Zhang, Z., Wang, X., Tyagi, A., and Agrawal, A. (2018, January 18\u201322). Context encoding for semantic segmentation. Proceedings of the Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00747"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Zhang, Z., Zhang, X., Peng, C., Xue, X., and Sun, J. (2018, January 18\u201322). Exfuse: Enhancing feature fusion for semantic segmentation. Proceedings of the Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1007\/978-3-030-01249-6_17"},{"key":"ref_31","unstructured":"Huang, Z., Huang, L., Gong, Y., Huang, C., and Wang, X. (November, January 27). Mask scoring R-CNN. Proceedings of the International Conference on Computer Vision, Seoul, Korea."},{"key":"ref_32","unstructured":"Ying, H., Huang, Z., Liu, S., Shao, T., and Zhou, K. (2019). EmbedMask: Embedding Coupling for One-stage Instance Segmentation. arXiv."},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"1137","DOI":"10.1109\/TPAMI.2016.2577031","article-title":"Faster R-CNN: Towards real-time object detection with region proposal networks","volume":"39","author":"Ren","year":"2016","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Lin, T.Y., Doll\u00e1r, P., Girshick, R., He, K., Hariharan, B., and Belongie, S. (2017, January 21\u201326). Feature pyramid networks for object detection. Proceedings of the Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.106"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Liu, W., Anguelov, D., Erhan, D., Szegedy, C., Reed, S., Fu, C.Y., and Berg, A.C. (2016, January 8\u201316). Ssd: Single shot multibox detector. Proceedings of the European Conference on Computer Vision, Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46448-0_2"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Redmon, J., Divvala, S., Girshick, R., and Farhadi, A. (2016, January 27\u201330). You only look once: Unified, real-time object detection. Proceedings of the Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.91"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Lin, T.Y., Goyal, P., Girshick, R., He, K., and Doll\u00e1r, P. (2017, January 22\u201329). Focal loss for dense object detection. Proceedings of the International Conference on Computer Vision (ICCV), Venice, Italy.","DOI":"10.1109\/ICCV.2017.324"},{"key":"ref_38","unstructured":"Tian, Z., Shen, C., Chen, H., and He, T. (November, January 27). Fcos: Fully convolutional one-stage object detection. Proceedings of the International Conference on Computer Vision, Seoul, Korea."},{"key":"ref_39","unstructured":"Sherrah, J. (2016). Fully convolutional networks for dense semantic labelling of high-resolution aerial imagery. arXiv."},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Ghosh, A., Ehrlich, M., Shah, S., Davis, L.S., and Chellappa, R. (2018, January 18\u201322). Stacked U-Nets for Ground Material Segmentation in Remote Sensing Imagery. Proceedings of the Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPRW.2018.00047"},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Hamaguchi, R., Fujita, A., Nemoto, K., Imaizumi, T., and Hikosaka, S. (2018, January 12\u201315). Effective use of dilated convolutions for segmenting small object instances in remote sensing imagery. Proceedings of the IEEE Winter Conference on Applications of Computer Vision, Lake Tahoe, CA, USA.","DOI":"10.1109\/WACV.2018.00162"},{"key":"ref_42","doi-asserted-by":"crossref","first-page":"6699","DOI":"10.1109\/TGRS.2018.2841808","article-title":"Vehicle instance segmentation from aerial image and video using a multitask learning residual fully convolutional network","volume":"56","author":"Mou","year":"2018","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_43","unstructured":"Feng, Y., Diao, W., Zhang, Y., Li, H., Chang, Z., Yan, M., Sun, X., and Gao, X. (August, January 28). Ship Instance Segmentation from Remote Sensing Images Using Sequence Local Context Module. Proceedings of the IEEE International Geoscience and Remote Sensing Symposium, Yokohama, Japan."},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"Xia, G.S., Bai, X., Ding, J., Zhu, Z., Belongie, S., Luo, J., Datcu, M., Pelillo, M., and Zhang, L. (2018, January 18\u201322). DOTA: A large-scale dataset for object detection in aerial images. Proceedings of the Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00418"},{"key":"ref_45","unstructured":"Lam, D., Kuzma, R., McGee, K., Dooley, S., Laielli, M., Klaric, M., Bulatov, Y., and McCord, B. (2018). xview: Objects in context in overhead imagery. arXiv."},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Maggiori, E., Tarabalka, Y., Charpiat, G., and Alliez, P. (2017, January 23\u201328). Can semantic labeling methods generalize to any city? The inria aerial image labeling benchmark. Proceedings of the IEEE International Geoscience and Remote Sensing Symposium, Fort Worth, TX, USA.","DOI":"10.1109\/IGARSS.2017.8127684"},{"key":"ref_47","doi-asserted-by":"crossref","unstructured":"Goldberg, H., Brown, M., and Wang, S. (2017, January 10\u201312). A benchmark for building footprint classification using orthorectified rgb imagery and digital surface models from commercial satellites. Proceedings of the IEEE Applied Imagery Pattern Recognition Workshop, Washington, DC, USA.","DOI":"10.1109\/AIPR.2017.8457973"},{"key":"ref_48","doi-asserted-by":"crossref","unstructured":"Weir, N., Lindenbaum, D., Bastidas, A., Van Etten, A., McPherson, S., Shermeyer, J., Kumar, V., and Tang, H. (2019). SpaceNet MVOI: A Multi-View Overhead Imagery Dataset Supplementary Material. arXiv.","DOI":"10.1109\/ICCV.2019.00108"},{"key":"ref_49","doi-asserted-by":"crossref","first-page":"98","DOI":"10.1007\/s11263-014-0733-5","article-title":"The pascal visual object classes challenge: A retrospective","volume":"111","author":"Everingham","year":"2015","journal-title":"Int. J. Comput. Vision"},{"key":"ref_50","doi-asserted-by":"crossref","unstructured":"Cordts, M., Omran, M., Ramos, S., Rehfeld, T., Enzweiler, M., Benenson, R., Franke, U., Roth, S., and Schiele, B. (2016, January 27\u201330). The cityscapes dataset for semantic urban scene understanding. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.350"},{"key":"ref_51","doi-asserted-by":"crossref","unstructured":"Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Doll\u00e1r, P., and Zitnick, C.L. (2014, January 6\u201312). Microsoft coco: Common objects in context. Proceedings of the European Conference on Computer Vision, Zurich, Switzerland.","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"ref_52","doi-asserted-by":"crossref","first-page":"211","DOI":"10.1007\/s11263-015-0816-y","article-title":"Imagenet large scale visual recognition challenge","volume":"115","author":"Russakovsky","year":"2015","journal-title":"Int. J. Comput. Vision"},{"key":"ref_53","unstructured":"Bellver, M., Salvador, A., Torres, J., and Giro-i-Nieto, X. (2019). Budget-aware Semi-Supervised Semantic and Instance Segmentation. arXiv."},{"key":"ref_54","doi-asserted-by":"crossref","unstructured":"Wei, Y., Xiao, H., Shi, H., Jie, Z., Feng, J., and Huang, T.S. (2018, January 18\u201322). Revisiting dilated convolution: A simple approach for weakly-and semi-supervised semantic segmentation. Proceedings of the Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00759"},{"key":"ref_55","doi-asserted-by":"crossref","unstructured":"Ibrahim, M.S., Vahdat, A., and Macready, W.G. (2020, January 14\u201319). Semi-Supervised Semantic Image Segmentation with Self-correcting Networks. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.01273"},{"key":"ref_56","unstructured":"Neven, D., Brabandere, B.D., Proesmans, M., and Gool, L.V. (November, January 27). Instance segmentation by jointly optimizing spatial embeddings and clustering bandwidth. Proceedings of the International Conference on Computer Vision, Seoul, Korea."},{"key":"ref_57","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep residual learning for image recognition. Proceedings of the Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_58","doi-asserted-by":"crossref","unstructured":"Jiang, Y., Zhu, X., Wang, X., Yang, S., Li, W., Wang, H., Fu, P., and Luo, Z. (2017). R2CNN: Rotational region cnn for orientation robust scene text detection. arXiv.","DOI":"10.1109\/ICPR.2018.8545598"},{"key":"ref_59","unstructured":"Yang, X., Yang, J., Yan, J., Zhang, Y., Zhang, T., Guo, Z., Sun, X., and Fu, K. (November, January 27). SCRDet: Towards more robust detection for small, cluttered and rotated objects. Proceedings of the International Conference on Computer Vision, Seoul, Korea."},{"key":"ref_60","doi-asserted-by":"crossref","unstructured":"Berman, M., Rannen Triki, A., and Blaschko, M.B. (2018, January 18\u201322). The Lov\u00e1sz-Softmax loss: A tractable surrogate for the optimization of the intersection-over-union measure in neural networks. Proceedings of the Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00464"},{"key":"ref_61","unstructured":"Kingma, D.P., and Ba, J. (2014, January 14\u201316). Adam: A method for stochastic optimization. Proceedings of the PInternational Conference on Learning Representations, Banff, AB, Canada."},{"key":"ref_62","doi-asserted-by":"crossref","unstructured":"Xie, S., Girshick, R., Doll\u00e1r, P., Tu, Z., and He, K. (2017, January 21\u201326). Aggregated residual transformations for deep neural networks. Proceedings of the Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.634"}],"container-title":["Remote Sensing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2072-4292\/13\/2\/252\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T05:10:44Z","timestamp":1760159444000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2072-4292\/13\/2\/252"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,1,13]]},"references-count":62,"journal-issue":{"issue":"2","published-online":{"date-parts":[[2021,1]]}},"alternative-id":["rs13020252"],"URL":"https:\/\/doi.org\/10.3390\/rs13020252","relation":{},"ISSN":["2072-4292"],"issn-type":[{"value":"2072-4292","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,1,13]]}}}