{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,7]],"date-time":"2026-07-07T04:22:04Z","timestamp":1783398124613,"version":"3.54.6"},"reference-count":49,"publisher":"MDPI AG","issue":"4","license":[{"start":{"date-parts":[[2020,2,13]],"date-time":"2020-02-13T00:00:00Z","timestamp":1581552000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["61663031, 61866028, 61661036, 61763033, 61662049, 61741312, 61881340421, and 61866025"],"award-info":[{"award-number":["61663031, 61866028, 61661036, 61763033, 61662049, 61741312, 61881340421, and 61866025"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Key Program Project of Research and Development (Jiangxi Provincial Department of Science and Technology)","award":["20171ACE50024 and 20192BBE50073"],"award-info":[{"award-number":["20171ACE50024 and 20192BBE50073"]}]},{"name":"Construction Project of Advantageous Science and Technology Innovation Team in Jiangxi Province","award":["20165BCB19007"],"award-info":[{"award-number":["20165BCB19007"]}]},{"name":"Application Innovation Plan (Ministry of Public Security of P. R. China)","award":["2017YYCXJXST048"],"award-info":[{"award-number":["2017YYCXJXST048"]}]},{"name":"Open Foundation of Key Laboratory of Jiangxi Province for Image Processing and Pattern Recognition","award":["ET201680245, TX201604002"],"award-info":[{"award-number":["ET201680245, TX201604002"]}]},{"name":"Foundation of China Scholarship Council","award":["CSC201908360075"],"award-info":[{"award-number":["CSC201908360075"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>With the rapid development of flexible vision sensors and visual sensor networks, computer vision tasks, such as object detection and tracking, are entering a new phase. Accordingly, the more challenging comprehensive task, including instance segmentation, can develop rapidly. Most state-of-the-art network frameworks, for instance, segmentation, are based on Mask R-CNN (mask region-convolutional neural network). However, the experimental results confirm that Mask R-CNN does not always successfully predict instance details. The scale-invariant fully convolutional network structure of Mask R-CNN ignores the difference in spatial information between receptive fields of different sizes. A large-scale receptive field focuses more on detailed information, whereas a small-scale receptive field focuses more on semantic information. So the network cannot consider the relationship between the pixels at the object edge, and these pixels will be misclassified. To overcome this problem, Mask-Refined R-CNN (MR R-CNN) is proposed, in which the stride of ROIAlign (region of interest align) is adjusted. In addition, the original fully convolutional layer is replaced with a new semantic segmentation layer that realizes feature fusion by constructing a feature pyramid network and summing the forward and backward transmissions of feature maps of the same resolution. The segmentation accuracy is substantially improved by combining the feature layers that focus on the global and detailed information. The experimental results on the COCO (Common Objects in Context) and Cityscapes datasets demonstrate that the segmentation accuracy of MR R-CNN is about 2% higher than that of Mask R-CNN using the same backbone. The average precision of large instances reaches 56.6%, which is higher than those of all state-of-the-art methods. In addition, the proposed method requires low time cost and is easily implemented. The experiments on the Cityscapes dataset also prove that the proposed method has great generalization ability.<\/jats:p>","DOI":"10.3390\/s20041010","type":"journal-article","created":{"date-parts":[[2020,2,20]],"date-time":"2020-02-20T03:20:03Z","timestamp":1582168803000},"page":"1010","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":195,"title":["Mask-Refined R-CNN: A Network for Refining Object Details in Instance Segmentation"],"prefix":"10.3390","volume":"20","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-1997-4871","authenticated-orcid":false,"given":"Yiqing","family":"Zhang","sequence":"first","affiliation":[{"name":"Department of Key Laboratory of Jiangxi Province for Image Processing and Pattern Recognition, Nanchang Hangkong University, Nanchang 330063, China"},{"name":"School of software, Nanchang Hangkong University, Nanchang 330063, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jun","family":"Chu","sequence":"additional","affiliation":[{"name":"Department of Key Laboratory of Jiangxi Province for Image Processing and Pattern Recognition, Nanchang Hangkong University, Nanchang 330063, China"},{"name":"School of software, Nanchang Hangkong University, Nanchang 330063, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Lu","family":"Leng","sequence":"additional","affiliation":[{"name":"Department of Key Laboratory of Jiangxi Province for Image Processing and Pattern Recognition, Nanchang Hangkong University, Nanchang 330063, China"},{"name":"School of software, Nanchang Hangkong University, Nanchang 330063, China"},{"name":"School of Electrical and Electronic Engineering, College of Engineering, Yonsei University, Seoul 120749, Korea"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jun","family":"Miao","sequence":"additional","affiliation":[{"name":"Department of Key Laboratory of Jiangxi Province for Image Processing and Pattern Recognition, Nanchang Hangkong University, Nanchang 330063, China"},{"name":"School of Aeronautical Manufacturing Engineering, Nanchang Hangkong University, Nanchang 330063, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2020,2,13]]},"reference":[{"key":"ref_1","first-page":"285","article-title":"An Efficient Vision-based Object Detection and Tracking using Online Learning","volume":"4","author":"Kim","year":"2017","journal-title":"J. Multimed. Inf. Syst."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"41273","DOI":"10.1109\/ACCESS.2019.2907327","article-title":"Efficient Facial Expression Recognition Algorithm Based on Hierarchical Deep Neural Network Structure","volume":"7","author":"Kim","year":"2019","journal-title":"IEEE Access."},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Kahaki, S.M., Nordin, M.J., Ahmad, N.S., Arzoky, M., and Ismail, W. (2019). Deep convolutional neural network designed for age assessment based on orthopantomography data. Neural Comput. Appl., 1\u201312.","DOI":"10.1007\/s00521-019-04449-6"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Lee, Y.-W., Kim, J.-H., Choi, Y.-J., and Kim, B.-G. (2018, January 12\u201314). CNN-based approach for visual quality improvement on HEVC. Proceedings of the IEEE Conference on Consumer Electronics (ICCE), Las Vegas, NV, USA.","DOI":"10.1109\/ICCE.2018.8326088"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Yuan, B., Li, Y., Jiang, F., Xu, X., Guo, Y., Zhao, J., Zhang, D., Guo, J., and Shen, X. (2019). MU R-CNN: A Two-Dimensional Code Instance Segmentation Network Based on Deep Learning. Future Internet., 11.","DOI":"10.3390\/fi11090197"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Liu, G., He, B., Liu, S., and Huang, J. (2019). Chassis Assembly Detection and Identification Based on Deep Learning Component Instance Segmentation. Symmetry, 11.","DOI":"10.3390\/sym11081001"},{"key":"ref_7","unstructured":"Krizhevsky, A., Sutskever, I., and Hinton, G. (2012, January 3\u20138). Imagenet Classification with Deep Convolutional Neural Networks. Proceedings of the Twenty-sixth Annual Conference on Neural Information Processing Systems (NIPS), Lake Tahoe, NV, USA."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S., Anguelov, D., Erhan, D., Vanhoucke, V., and Rabinovich, A. (2015, January 7\u201312). Going Deeper with Convolutions. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298594"},{"key":"ref_9","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (July, January 26). Deep residual learning for image recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Lin, T., Goyal, P., Girshick, R., He, K., and Dollar, P. (2017, January 22\u201329). Focal Loss for Dense Object Detection. Proceedings of the IEEE International Conference on Computer Vision (ICCV), Venice, Italy.","DOI":"10.1109\/ICCV.2017.324"},{"key":"ref_11","unstructured":"Ren, S., He, K., Girshick, R., and Sun, J. (2015, January 7\u201312). Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks. Proceedings of the 29th Annual Conference on Neural Information Processing Systems (NIPS), Montr\u00e9al, QC, Canada."},{"key":"ref_12","unstructured":"Redmon, J., Divvala, S., Girshick, R., and Farhadi, A. (July, January 26). You Only Look Once: Unified, Real-Time Object Detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Liu, W., Anguelov, D., Erhan, D., Szegedy, C., Reed, S., Fu, C., and Berg, A. (2016, January 8\u201316). SSD: Single Shot Multibox Detector. Proceedings of the European Conference on Computer Vision (ECCV), Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46448-0_2"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Peng, C., Zhang, X., Yu, G., Luo, G., and Sun, J. (2017, January 21\u201326). Large Kernel Matters\u2014Improve Semantic Segmentation by Global Convolutional Network. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.189"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Long, J., Shelhamer, E., and Darrell, T. (2015, January 7\u201312). Fully Convolutional Networks for Semantic Segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298965"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Zhao, H., Shi, J., Qi, X., Wang, X., and Jia, J. (2017, January 21\u201326). Pyramid Scene Parsing Network. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.660"},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"834","DOI":"10.1109\/TPAMI.2017.2699184","article-title":"DeepLab: Semantic Image Segmentation with Deep Convolutional Nets, Atrous Convolution, and Fully Connected CRFs","volume":"40","author":"Chen","year":"2018","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Wang, R., Xu, Y., Sotelo, M., Ma, Y., Sarkodie, T., Li, Z., and Li, W. (2019). A Robust Registration Method for Autonomous Driving Pose Estimation in Urban Dynamic Environment Using LiDAR. Electronics, 8.","DOI":"10.3390\/electronics8010043"},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"19959","DOI":"10.1109\/ACCESS.2018.2815149","article-title":"Object Detection Based on Multi-Layer Convolution Feature Fusion and Online Hard Example Mining","volume":"6","author":"Chu","year":"2018","journal-title":"IEEE Access."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Lin, T., Doll\u00e1r, P., Girshick, R., He, K., Hariharan, B., and Belongie, S. (2017, January 21\u201326). Feature Pyramid Networks for Object Detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.106"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"He, K., Gkioxari, G., Doll\u00e1r, P., and Girshick, R. (2017, January 22\u201329). Mask R-CNN. Proceedings of the IEEE International Conference on Computer Vision (ICCV), Venice, Italy.","DOI":"10.1109\/ICCV.2017.322"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Liu, S., Qi, L., Qin, H., Shi, J., and Jia, J. (2018, January 18\u201322). Path Aggregation Network for Instance Segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, SU, USA.","DOI":"10.1109\/CVPR.2018.00913"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Huang, Z., Huang, L., Gong, Y., Huang, C., and Wang, X. (2019). Mask Scor-ing R-CNN. arXiv.","DOI":"10.1109\/CVPR.2019.00657"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Lin, T., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Doll\u00e1r, P., and Zitnick, C. (2014, January 6\u201312). Microsoft COCO: Common Objects in Context. Proceedings of the European Conference on Computer Vision (ECCV), Zurich, Switzerland.","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Bai, M., and Urtasun, R. (2017, January 21\u201326). Deep Watershed Transform for Instance Segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.305"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Uhrig, J., Cordts, M., Franke, U., and Brox, T. (2016, January 14\u201325). Pixel-Level Encoding and Depth Layering for Instance-Level Semantic Labeling. Proceedings of the IEEE Conference on German Conference on Pattern Recognition (GCPR), Hannover, Germany.","DOI":"10.1007\/978-3-319-45886-1_2"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Liu, S., Jia, J., Fidler, S., and Urtasun, R. (2017, January 22\u201329). SGN: Sequential Grouping Networks for Instance Segmentation. Proceedings of the IEEE International Conference on Computer Vision (ICCV), Venice, Italy.","DOI":"10.1109\/ICCV.2017.378"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Kirillov, A., Levinkov, E., Andres, B., Savchynskyy, B., and Rother, C. (2017, January 21\u201326). InstanceCut: From Edges to Instances with MultiCut. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.774"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Arnab, A., and Torr, P. (2017, January 21\u201326). Pixelwise Instance Segmentation with a Dynamically Instantiated Network. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.100"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Romera-Paredes, B., and Torr, P. (2016, January 8\u201316). Recurrent Instance Segmentation. Proceedings of the European Conference on Computer Vision (ECCV), Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46466-4_19"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Ren, M., and Zemel, R. (2017, January 21\u201326). End-To-End Instance Segmentation with Recurrent Attention. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.39"},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"84","DOI":"10.1109\/MSP.2017.2749125","article-title":"Advanced deep-learning techniques for salient and category-specific object detection: A survey","volume":"35","author":"Han","year":"2018","journal-title":"IEEE Signal Process. Mag."},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"265","DOI":"10.1109\/TIP.2018.2867198","article-title":"Learning rotation-invariant and fisher discriminative convolutional neural networks for object detection","volume":"28","author":"Cheng","year":"2018","journal-title":"IEEE Trans. Image Process."},{"key":"ref_34","unstructured":"Dai, J., Li, Y., He, K., and Sun, J. (2016, January 5\u201310). R-FCN: Object Detection via Region-Based Fully Convolutional Networks. Proceedings of the 31st Conference on Neural Information Processing Systems (NIPS), Barcelona, Spain."},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Li, Y., Qi, H., Dai, J., Ji, X., and Wei, Y. (2017, January 21\u201326). Fully Convolutional Instance-Aware Semantic Segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.472"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Dai, J., He, K., Li, Y., Ren, S., and Sun, J. (2016, January 8\u201316). Instance-sensitive fully convolutional networks. Proceedings of the European Conference on Computer Vision (ECCV), Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46466-4_32"},{"key":"ref_37","doi-asserted-by":"crossref","first-page":"333","DOI":"10.1007\/s11042-015-3058-7","article-title":"Dual-source discrimination power analysis for multi-instance contactless palmprint recognition","volume":"76","author":"Leng","year":"2017","journal-title":"Multimed. Tools Applic."},{"key":"ref_38","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1016\/j.neucom.2012.08.028","article-title":"PalmHash Code vs. PalmPhasor Code","volume":"108","author":"Leng","year":"2013","journal-title":"Neurocomputing"},{"key":"ref_39","doi-asserted-by":"crossref","first-page":"1860","DOI":"10.1002\/sec.900","article-title":"A remote cancelable palmprint authentication protocol based on multi-directional two-dimensional PalmPhasor-fusion","volume":"7","author":"Leng","year":"2014","journal-title":"Securit. Commun. Netw."},{"key":"ref_40","unstructured":"Lin, T., Collobert, R., and Doll\u00e1r, P. (2016, January 8\u201316). Learning to Refine Object Segments. Proceedings of the European Conference on Computer Vision (ECCV), Amsterdam, The Netherlands."},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Ghiasi, G., and Fowlkes, C. (2016, January 8\u201316). Laplacian Pyramid Reconstruction and Refinement for Semantic Segmentation. Proceedings of the European Conference on Computer Vision (ECCV), Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46487-9_32"},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Ronneberger, O., Fischer, P., and Brox, T. (2015, January 5\u20139). U-Net: Convolutional Networks for Biomedical Image Segmentation. Proceedings of the IEEE Conference on Medical Image Computing and Computer Assisted Intervention (MICCAI), Munich, Germany.","DOI":"10.1007\/978-3-319-24574-4_28"},{"key":"ref_43","unstructured":"Fu, C., Liu, W., Ranga, A., Tyagi, A., and Berg, A. (2017, January 21\u201326). DSSD: Deconvolutional Single Shot Detector. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA."},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"Zeiler, M., and Fergus, G. (2014, January 6\u201312). Visualizing and Understanding Convolutional Networks. Proceedings of the European Conference on Computer Vision (ECCV), Zurich, Switzerland.","DOI":"10.1007\/978-3-319-10590-1_53"},{"key":"ref_45","unstructured":"Leng, L., Zhang, J., Xu, J., Khan, M.K., and Alghathbar, K. (2010, January 17\u201319). Dynamic weighted discrimination power analysis in DCT domain for face and palmprint recognition. Proceedings of the International Conference on Information and Communication Technology Convergence IEEE(ICTC), Jeju, Korea."},{"key":"ref_46","first-page":"2543","article-title":"Dynamic weighted discrimination power analysis: A novel approach for face and palmprint recognition in DCT domain","volume":"5","author":"Leng","year":"2010","journal-title":"Int. J. Phys. Sci."},{"key":"ref_47","doi-asserted-by":"crossref","unstructured":"Rahman, M.A., and Wang, Y. (2016). Optimizing intersection-over-union in deep neural networks for image segmentation. International Symposium on Visual Computing, Springer.","DOI":"10.1007\/978-3-319-50835-1_22"},{"key":"ref_48","doi-asserted-by":"crossref","unstructured":"Dai, J., Qi, H., Xiong, Y., Li, Y., Zhang, G., Hu, H., and Wei, Y. (2017, January 22\u201329). Deformable Convolutional Networks. Proceedings of the IEEE International Conference on Computer Vision (ICCV), Venice, Italy.","DOI":"10.1109\/ICCV.2017.89"},{"key":"ref_49","unstructured":"Cordts, M., Omran, M., Ramos, S., Rehfeld, T., Enzweiler, M., Benenson, R., Franke, U., Roth, S., and Schiele, B. (July, January 26). The Cityscapes Dataset for Semantic Urban Scene Understanding. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/20\/4\/1010\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T08:57:29Z","timestamp":1760173049000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/20\/4\/1010"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,2,13]]},"references-count":49,"journal-issue":{"issue":"4","published-online":{"date-parts":[[2020,2]]}},"alternative-id":["s20041010"],"URL":"https:\/\/doi.org\/10.3390\/s20041010","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,2,13]]}}}