{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,12]],"date-time":"2025-10-12T01:59:47Z","timestamp":1760234387667,"version":"build-2065373602"},"reference-count":50,"publisher":"MDPI AG","issue":"9","license":[{"start":{"date-parts":[[2021,5,9]],"date-time":"2021-05-09T00:00:00Z","timestamp":1620518400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"National Natural Science Foundation of China","award":["61976094"],"award-info":[{"award-number":["61976094"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>The main challenges of semantic segmentation in vehicle-mounted scenes are object scale variation and trading off model accuracy and efficiency. Lightweight backbone networks for semantic segmentation usually extract single-scale features layer-by-layer only by using a fixed receptive field. Most modern real-time semantic segmentation networks heavily compromise spatial details when encoding semantics, and sacrifice accuracy for speed. Many improving strategies adopt dilated convolution and add a sub-network, in which either intensive computation or redundant parameters are brought. We propose a multi-level and multi-scale feature aggregation network (MMFANet). A spatial pyramid module is designed by cascading dilated convolutions with different receptive fields to extract multi-scale features layer-by-layer. Subseqently, a lightweight backbone network is built by reducing the feature channel capacity of the module. To improve the accuracy of our network, we design two additional modules to separately capture spatial details and high-level semantics from the backbone network without significantly increasing the computation cost. Comprehensive experimental results show that our model achieves 79.3% MIoU on the Cityscapes test dataset at a speed of 58.5 FPS, and it is more accurate than SwiftNet (75.5% MIoU). Furthermore, the number of parameters of our model is at least 53.38% less than that of other state-of-the-art models.<\/jats:p>","DOI":"10.3390\/s21093270","type":"journal-article","created":{"date-parts":[[2021,5,10]],"date-time":"2021-05-10T02:54:58Z","timestamp":1620615298000},"page":"3270","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":7,"title":["Multi-Level and Multi-Scale Feature Aggregation Network for Semantic Segmentation in Vehicle-Mounted Scenes"],"prefix":"10.3390","volume":"21","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-1989-8469","authenticated-orcid":false,"given":"Yong","family":"Liao","sequence":"first","affiliation":[{"name":"School of Software Engineering, South China University of Technology, Guangzhou 510006, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Qiong","family":"Liu","sequence":"additional","affiliation":[{"name":"School of Software Engineering, South China University of Technology, Guangzhou 510006, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2021,5,9]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"137","DOI":"10.1109\/TITS.2018.2801309","article-title":"Importance-Aware Semantic Segmentation for Autonomous Vehicles","volume":"20","author":"Chen","year":"2019","journal-title":"IEEE Trans. Intell. Transp. Syst."},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Zhang, Y., Liu, H., and Hu, Q. (2021). TransFuse: Fusing Transformers and CNNs for Medical Image Segmentation. arXiv.","DOI":"10.1007\/978-3-030-87193-2_2"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Lumini, A., Nanni, L., and Maguolo, G. (2021). Deep ensembles based on Stochastic Activation Selection for Polyp Segmentation. arXiv.","DOI":"10.20944\/preprints202107.0691.v1"},{"key":"ref_4","unstructured":"Dhere, A., and Sivaswamy, J. (2021). Self-Supervised Learning for Segmentation. arXiv."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"834","DOI":"10.1109\/TPAMI.2017.2699184","article-title":"DeepLab: Semantic Image Segmentation with Deep Convolutional Nets, Atrous Convolution, and Fully Connected CRFs","volume":"40","author":"Chen","year":"2018","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_6","unstructured":"Chen, L., Papandreou, G., Schroff, F., and Adam, H. (2017). Rethinking Atrous Convolution for Semantic Image Segmentation. arXiv."},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Chen, L., Yang, Y., Wang, J., Xu, W., and Yuille, A.L. (2016, January 27\u201330). Attention to Scale: Scale-Aware Semantic Image Segmentation. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.396"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Zhao, H., Shi, J., Qi, X., Wang, X., and Jia, J. (2017, January 21\u201326). Pyramid Scene Parsing Network. Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.660"},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"1915","DOI":"10.1109\/TPAMI.2012.231","article-title":"Learning Hierarchical Features for Scene Labeling","volume":"35","author":"LeCun","year":"2013","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Eigen, D., and Fergus, R. (2014). Predicting Depth, Surface Normals and Semantic Labels with a Common Multi-Scale Convolutional Architecture. arXiv.","DOI":"10.1109\/ICCV.2015.304"},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"2481","DOI":"10.1109\/TPAMI.2016.2644615","article-title":"SegNet: A Deep Convolutional Encoder-Decoder Architecture for Image Segmentation","volume":"39","author":"Badrinarayanan","year":"2017","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_12","first-page":"263","article-title":"U-Net: Convolutional Networks for Biomedical Image Segmentation","volume":"9351","author":"Ronneberger","year":"2015","journal-title":"Med Image Comput. Comput. Assist. Interv."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Lin, G., Milan, A., Shen, C., and Reid, I. (2017, January 21\u201326). RefineNet: Multi-path Refinement Networks for High-Resolution Semantic Segmentation. Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.549"},{"key":"ref_14","unstructured":"Simonyan, K., and Zisserman, A. (2014). Very deep convolutional networks for large-scale image recognition. arXiv."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep Residual Learning for Image Recognition. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_16","unstructured":"Han, S., Pool, J., Tran, J., and Dally, W. (2015, January 7\u201312). Learning both Weights and Connections for Efficient Neural Network. Proceedings of the 29th Conference on Neural Information Processing Systems (NIPS 2015), Montreal, QC, Canada."},{"key":"ref_17","first-page":"1269","article-title":"Exploiting Linear Structure Within Convolutional Networks for Efficient Evaluation","volume":"Volume 27","author":"Ghahramani","year":"2014","journal-title":"Advances in Neural Information Processing Systems"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Cao, X., Chen, Y., Zhao, Q., Meng, D., Wang, Y., Wang, D., and Xu, Z. (2015, January 7\u201313). Low-Rank Matrix Factorization under General Mixture Noise Distributions. Proceedings of the 2015 IEEE International Conference on Computer Vision (ICCV), Santiago, Chile.","DOI":"10.1109\/ICCV.2015.175"},{"key":"ref_19","unstructured":"Li, G., and Kim, J. (2019, January 9\u201312). DABNet: Depth-wise Asymmetric Bottleneck for Real-time Semantic Segmentation. Proceedings of the British Machine Vision Conference 2019, Cardiff, UK."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Wang, Y., Zhou, Q., Liu, J., Xiong, J., Gao, G., Wu, X., and Latecki, L.J. (2019, January 22\u201325). Lednet: A Lightweight Encoder-Decoder Network for Real-Time Semantic Segmentation. Proceedings of the 2019 IEEE International Conference on Image Processing (ICIP), Taipei, Taiwan.","DOI":"10.1109\/ICIP.2019.8803154"},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"1045","DOI":"10.1007\/s10489-019-01587-1","article-title":"ADSCNet: Asymmetric depthwise separable convolution for semantic segmentation in real-time","volume":"115","author":"Wang","year":"2020","journal-title":"Appl. Intell."},{"key":"ref_22","unstructured":"Yang, Z., Yu, H., Fu, Q., Sun, W., Jia, W., Sun, M., and Mao, Z. (2020). NDNet: Narrow While Deep Network for Real-Time Semantic Segmentation. IEEE Trans. Intell. Transp. Syst., 1\u201312."},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"263","DOI":"10.1109\/TITS.2017.2750080","article-title":"ERFNet: Efficient Residual Factorized ConvNet for Real-Time Semantic Segmentation","volume":"19","author":"Romera","year":"2018","journal-title":"IEEE Trans. Intell. Transp. Syst."},{"key":"ref_24","unstructured":"Ferrari, V., Hebert, M., Sminchisescu, C., and Weiss, Y. (2018). ESPNet: Efficient Spatial Pyramid of Dilated Convolutions for Semantic Segmentation. Computer Vision\u2014ECCV 2018, Springer International Publishing."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Mehta, S., Rastegari, M., Shapiro, L., and Hajishirzi, H. (2019, January 15\u201320). ESPNetv2: A Light-Weight, Power Efficient, and General Purpose Convolutional Neural Network. Proceedings of the 2019 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00941"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Orsic, M., Kreso, I., Bevandic, P., and Segvic, S. (2019, January 15\u201320). In Defense of Pre-Trained ImageNet Architectures for Real-Time Semantic Segmentation of Road-Driving Images. Proceedings of the 2019 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.01289"},{"key":"ref_27","unstructured":"Ferrari, V., Hebert, M., Sminchisescu, C., and Weiss, Y. (2018). BiSeNet: Bilateral Segmentation Network for Real-Time Semantic Segmentation. Computer Vision\u2014ECCV 2018, Springer International Publishing."},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Chaurasia, A., and Culurciello, E. (2017, January 10\u201313). LinkNet: Exploiting encoder representations for efficient semantic segmentation. Proceedings of the 2017 IEEE Visual Communications and Image Processing (VCIP), St. Petersburg, FL, USA.","DOI":"10.1109\/VCIP.2017.8305148"},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"54","DOI":"10.1016\/j.neucom.2020.02.019","article-title":"MFENet: Multi-level feature enhancement network for real-time semantic segmentation","volume":"393","author":"Zhang","year":"2020","journal-title":"Neurocomputing"},{"key":"ref_30","unstructured":"Ferrari, V., Hebert, M., Sminchisescu, C., and Weiss, Y. (2018). ICNet for Real-Time Semantic Segmentation on High-Resolution Images. Computer Vision\u2014ECCV 2018, Springer International Publishing."},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Li, H., Xiong, P., Fan, H., and Sun, J. (2019, January 15\u201320). DFANet: Deep Feature Aggregation for Real-Time Semantic Segmentation. Proceedings of the 2019 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00975"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Chen, L.C., Zhu, Y., Papandreou, G., Schroff, F., and Adam, H. (2018, January 8\u201314). Encoder-Decoder with Atrous Separable Convolution for Semantic Image Segmentation. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germay.","DOI":"10.1007\/978-3-030-01234-2_49"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Fu, J., Liu, J., Tian, H., Li, Y., Bao, Y., Fang, Z., and Lu, H. (2019, January 15\u201320). Dual Attention Network for Scene Segmentation. Proceedings of the 2019 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00326"},{"key":"ref_34","doi-asserted-by":"crossref","first-page":"640","DOI":"10.1109\/TPAMI.2016.2572683","article-title":"Fully Convolutional Networks for Semantic Segmentation","volume":"39","author":"Shelhamer","year":"2017","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_35","unstructured":"Wu, H., Zhang, J., Huang, K., Liang, K., and Yizhou, Y. (2019). FastFCN: Rethinking Dilated Convolution in the Backbone for Semantic Segmentation. arXiv."},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Lin, T., Dollar, P., Girshick, R., He, K., Hariharan, B., and Belongie, S. (2017, January 21\u201326). Feature Pyramid Networks for Object Detection. Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.106"},{"key":"ref_37","unstructured":"Howard, A.G., Zhu, M., Chen, B., Kalenichenko, D., Wang, W., Weyand, T., Andreetto, M., and Adam, H. (2017). MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications. arXiv."},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Zhang, X., Zhou, X., Lin, M., and Sun, J. (2018, January 18\u201323). ShuffleNet: An Extremely Efficient Convolutional Neural Network for Mobile Devices. Proceedings of the 2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00716"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Vedaldi, A., Bischof, H., Brox, T., and Frahm, J.M. (2020). Semantic Flow for Fast and Accurate Scene Parsing. Computer Vision\u2014ECCV 2020, Springer International Publishing.","DOI":"10.1007\/978-3-030-58604-1"},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Wang, P., Chen, P., Yuan, Y., Liu, D., Huang, Z., Hou, X., and Cottrell, G.W. (2017). Understanding Convolution for Semantic Segmentation. arXiv.","DOI":"10.1109\/WACV.2018.00163"},{"key":"ref_41","doi-asserted-by":"crossref","first-page":"652","DOI":"10.1109\/TPAMI.2019.2938758","article-title":"Res2Net: A New Multi-scale Backbone Architecture","volume":"43","author":"Gao","year":"2021","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Sandler, M., Howard, A.G., Zhu, M., Zhmoginov, A., and Chen, L. (2018). Inverted Residuals and Linear Bottlenecks: Mobile Networks for Classification, Detection and Segmentation. arXiv.","DOI":"10.1109\/CVPR.2018.00474"},{"key":"ref_43","first-page":"448","article-title":"Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift","volume":"Volume 37","author":"Ioffe","year":"2015","journal-title":"Proceedings of Machine Learning Research"},{"key":"ref_44","unstructured":"Maas, A.L. (2013, January 16\u201321). Rectifier Nonlinearities Improve Neural Network Acoustic Models. Proceedings of the 30th International Conference on Machine Learning, Atlanta, GA, USA."},{"key":"ref_45","unstructured":"Dong, G., Yan, Y., Shen, C., and Wang, H. (2020). Real-Time High-Performance Semantic Image Segmentation of Urban Street Scenes. IEEE Trans. Intell. Transp. Syst., 1\u201317."},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Hu, J., Shen, L., and Sun, G. (2018, January 18\u201323). Squeeze-and-Excitation Networks. Proceedings of the 2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00745"},{"key":"ref_47","doi-asserted-by":"crossref","unstructured":"Cordts, M., Omran, M., Ramos, S., Rehfeld, T., Enzweiler, M., Benenson, R., Franke, U., Roth, S., and Schiele, B. (2016, January 27\u201330). The Cityscapes Dataset for Semantic Urban Scene Understanding. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.350"},{"key":"ref_48","unstructured":"Forsyth, D., Torr, P., and Zisserman, A. (2008). Segmentation and Recognition Using Structure from Motion Point Clouds. Computer Vision\u2014ECCV 2008, Springer Berlin Heidelberg."},{"key":"ref_49","unstructured":"Paszke, A., Chaurasia, A., Kim, S., and Culurciello, E. (2016). ENet: A Deep Neural Network Architecture for Real-Time Semantic Segmentation. arXiv."},{"key":"ref_50","doi-asserted-by":"crossref","first-page":"211","DOI":"10.1007\/s11263-015-0816-y","article-title":"ImageNet Large Scale Visual Recognition Challenge","volume":"115","author":"Russakovsky","year":"2015","journal-title":"Int. J. Comput. Vis."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/21\/9\/3270\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T05:58:32Z","timestamp":1760162312000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/21\/9\/3270"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,5,9]]},"references-count":50,"journal-issue":{"issue":"9","published-online":{"date-parts":[[2021,5]]}},"alternative-id":["s21093270"],"URL":"https:\/\/doi.org\/10.3390\/s21093270","relation":{},"ISSN":["1424-8220"],"issn-type":[{"type":"electronic","value":"1424-8220"}],"subject":[],"published":{"date-parts":[[2021,5,9]]}}}