{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,18]],"date-time":"2025-10-18T15:17:44Z","timestamp":1760800664955,"version":"3.41.2"},"reference-count":34,"publisher":"Frontiers Media SA","license":[{"start":{"date-parts":[[2023,8,31]],"date-time":"2023-08-31T00:00:00Z","timestamp":1693440000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["frontiersin.org"],"crossmark-restriction":true},"short-container-title":["Front. Neurorobot."],"abstract":"<jats:p>Semantic segmentation, which is a fundamental task in computer vision. Every pixel will have a specific semantic class assigned to it through semantic segmentation methods. Embedded systems and mobile devices are difficult to deploy high-accuracy segmentation algorithms. Despite the rapid development of semantic segmentation, the balance between speed and accuracy must be improved. As a solution to the above problems, we created a cross-scale fusion attention mechanism network called CFANet, which fuses feature maps from different scales. We first design a novel efficient residual module (ERM), which applies both dilation convolution and factorized convolution. Our CFANet is mainly constructed from ERM. Subsequently, we designed a new multi-branch channel attention mechanism (MCAM) to refine the feature maps at different levels. Experiment results show that CFANet achieved 70.6% mean intersection over union (mIoU) and 67.7% mIoU on Cityscapes and CamVid datasets, respectively, with inference speeds of 118 FPS and 105 FPS on NVIDIA RTX2080Ti GPU cards with 0.84M parameters.<\/jats:p>","DOI":"10.3389\/fnbot.2023.1204418","type":"journal-article","created":{"date-parts":[[2023,8,31]],"date-time":"2023-08-31T16:09:15Z","timestamp":1693498155000},"update-policy":"https:\/\/doi.org\/10.3389\/crossmark-policy","source":"Crossref","is-referenced-by-count":3,"title":["Based on cross-scale fusion attention mechanism network for semantic segmentation for street scenes"],"prefix":"10.3389","volume":"17","author":[{"given":"Xin","family":"Ye","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Lang","family":"Gao","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jichen","family":"Chen","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Mingyue","family":"Lei","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1965","published-online":{"date-parts":[[2023,8,31]]},"reference":[{"key":"B1","doi-asserted-by":"publisher","first-page":"2481","DOI":"10.1109\/TPAMI.2016.2644615","article-title":"Segnet: A deep convolutional encoder-decoder architecture for image segmentation","volume":"39","author":"Badrinarayanan","year":"2017","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell"},{"key":"B2","doi-asserted-by":"crossref","first-page":"177","DOI":"10.1007\/978-3-7908-2604-3_16","article-title":"\u201cLarge scale machine learning with stochastic gradient descent,\u201d","volume-title":"Proceedings of COMPSTAT'2010","author":"Bottou","year":"2010"},{"key":"B3","doi-asserted-by":"publisher","first-page":"88","DOI":"10.1016\/j.patrec.2008.04.005","article-title":"Semantic object classes in video: a high-definition ground truth database","volume":"30","author":"Brostow","year":"2009","journal-title":"Pattern Recognit. Lett"},{"key":"B4","doi-asserted-by":"publisher","first-page":"834","DOI":"10.1109\/TPAMI.2017.2699184","article-title":"Deeplab: semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs","volume":"40","author":"Chen","year":"2017","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell"},{"key":"B5","first-page":"3213","article-title":"\u201cThe cityscapes dataset for semantic urban scene understanding,\u201d","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Cordts","year":"2016"},{"key":"B6","doi-asserted-by":"publisher","first-page":"725","DOI":"10.1108\/AA-06-2021-0078","article-title":"MDRNet: a lightweight network for real-time semantic segmentation in street scenes","volume":"46","author":"Dai","year":"2021","journal-title":"Assembly Automat"},{"key":"B7","first-page":"503","article-title":"\u201cEdgenet: semantic scene completion from rgb-d image,\u201d","volume-title":"2020 25th International Conference on Pattern Recognition (ICPR)","author":"Dourado","year":"2020"},{"key":"B8","first-page":"42","article-title":"\u201cSanet: structure-aware network for visual trackin,\u201d","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops","author":"Fan","year":"2017"},{"key":"B9","doi-asserted-by":"publisher","first-page":"25489","DOI":"10.1109\/TITS.2021.3098355","article-title":"MSCFNet: a lightweight network with multi-scale context fusion for real-time semantic segmentation","volume":"23","author":"Gao","year":"2021","journal-title":"IEEE Transact. Intell. Transport. Syst"},{"key":"B10","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2020-2059","article-title":"Contextnet: Improving convolutional neural networks for automatic speech recognition with global context","author":"Han","year":"2020","journal-title":"arXiv"},{"key":"B11","first-page":"7132","article-title":"\u201cSqueeze-and-excitation networks,\u201d","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Hu","year":"2018"},{"key":"B12","doi-asserted-by":"publisher","first-page":"580","DOI":"10.1007\/s10489-021-02446-8","article-title":"Joint pyramid attention network for real-time semantic segmentation of urban scenes","volume":"52","author":"Hu","year":"2022","journal-title":"Appl. Intell"},{"key":"B13","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.1412.6980","article-title":"Adam: A method for stochastic optimization","author":"Kingma","year":"2014","journal-title":"arXiv [Preprint]."},{"key":"B14","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.1907.11357","article-title":"Dabnet: depth-wise asymmetric bottleneck for real-time semantic segmentation","author":"Li","year":"2019","journal-title":"arXiv [Preprint]."},{"key":"B15","first-page":"9522","article-title":"\u201cDfanet: deep feature aggregation for real-time semantic segmentation,\u201d","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Li","year":""},{"key":"B16","doi-asserted-by":"publisher","first-page":"115","DOI":"10.1016\/j.neucom.2021.12.003","article-title":"RELAXNet: residual efficient learning and attention expected fusion network for real-time semantic segmentation","volume":"474","author":"Liu","year":"2022","journal-title":"Neurocomputing"},{"key":"B17","first-page":"2373","article-title":"\u201cFDDWNet: a lightweight convolutional neural network for real-time semantic segmentation,\u201d","volume-title":"Proceedings of the ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)","author":"Liu","year":"2019"},{"key":"B18","first-page":"3431","article-title":"\u201cFully convolutional networks for semantic segmentation,\u201d","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Long","year":"2015"},{"key":"B19","doi-asserted-by":"publisher","first-page":"65","DOI":"10.1109\/MNET.2019.1800339","article-title":"The cognitive internet of vehicles for autonomous driving","volume":"33","author":"Lu","year":"2019","journal-title":"IEEE Netw"},{"key":"B20","first-page":"552","article-title":"\u201cEspnet: efficient spatial pyramid of dilated convolutions for semantic segmentation,\u201d","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV)","author":"Mehta","year":"2018"},{"key":"B21","first-page":"9190","article-title":"\u201cEspnetv2: a light-weight, power efficient, and general purpose convolu-tional neural network,\u201d","author":"Mehta","year":"2019","journal-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition"},{"key":"B22","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.1606.02147","article-title":"Enet: a deep neural network architecture for real-time semantic segmentation","author":"Paszke","year":"2016","journal-title":"arXiv [Preprint]."},{"key":"B23","doi-asserted-by":"publisher","first-page":"263","DOI":"10.1109\/TITS.2017.2750080","article-title":"Erfnet: efficient residual factorized convnet for real-time semantic segmentation","volume":"19","author":"Romera","year":"2017","journal-title":"IEEE Transact. Intell. Transport. Syst"},{"key":"B24","doi-asserted-by":"publisher","first-page":"14339","DOI":"10.1109\/ICPR48806.2021.9413176","article-title":"FASSD-Net: fast and accurate real-time semantic segmentation for embedded systems","volume":"23","author":"Rosas-Arias","year":"2021","journal-title":"IEEE Transact. Intell. Transport. Syst"},{"key":"B25","doi-asserted-by":"publisher","first-page":"6000","DOI":"10.5555\/3295222.3295349","article-title":"Attention is all you need","volume":"30","author":"Vaswani","year":"2017","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"B26","doi-asserted-by":"crossref","first-page":"1860","DOI":"10.1109\/ICIP.2019.8803154","article-title":"\u201cLednet: a lightweight encoder-decoder network for real-timesemantic segmentation,\u201d","volume-title":"Proceedings of the 2019 IEEE International Conference on Image Processing (ICIP)","author":"Wang","year":"2019"},{"key":"B27","first-page":"3","article-title":"\u201cCbam: convolutional block attention module,\u201d","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV)","author":"Woo","year":"2018"},{"key":"B28","doi-asserted-by":"publisher","first-page":"1169","DOI":"10.1109\/TIP.2020.3042065","article-title":"Cgnet: a light-weight context guided network for semantic segmentation","volume":"30","author":"Wu","year":"2020","journal-title":"IEEE Transact. Image Process"},{"key":"B29","first-page":"246","article-title":"\u201cEDA-Net: dense aggregation of deep and shallow information achieves quantitative photoacoustic blood oxygenation imaging deep in human breast,\u201d","volume-title":"Proceedings of the International Conference on Medical Image Computing and Computer-Assisted Intervention","author":"Yang","year":"2019"},{"key":"B30","doi-asserted-by":"publisher","first-page":"124","DOI":"10.1016\/j.isprsjprs.2021.06.006","article-title":"Real-time semantic segmentation with context aggregation network","volume":"178","author":"Yang","year":"2021","journal-title":"ISPRS J. Photogr. Remote Sens"},{"key":"B31","doi-asserted-by":"publisher","first-page":"5508","DOI":"10.1109\/TITS.2020.2987816","article-title":"NDNet: Narrow while deep network for real-time semantic segmentation","volume":"22","author":"Yang","year":"2020","journal-title":"IEEE Transact. Intell. Transport. Syst."},{"key":"B32","first-page":"325","article-title":"\u201cBisenet: bilateral segmentation network for real-time semantic seg-mentation,\u201d","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Yu","year":"2018"},{"key":"B33","doi-asserted-by":"publisher","first-page":"1183","DOI":"10.1109\/TII.2018.2849348","article-title":"Fast semantic segmentation for scene perception","volume":"15","author":"Zhang","year":"2018","journal-title":"IEEE Transact. Ind. Informat"},{"key":"B34","first-page":"405","article-title":"\u201cIcnet for real-time semantic segmentation on high-resolution images,\u201d","author":"Zhao","year":"2017","journal-title":"Proceedings of the European Conference on Computer Vision (ECCV"}],"container-title":["Frontiers in Neurorobotics"],"original-title":[],"link":[{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/fnbot.2023.1204418\/full","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,8,31]],"date-time":"2023-08-31T16:09:27Z","timestamp":1693498167000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/fnbot.2023.1204418\/full"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,8,31]]},"references-count":34,"alternative-id":["10.3389\/fnbot.2023.1204418"],"URL":"https:\/\/doi.org\/10.3389\/fnbot.2023.1204418","relation":{},"ISSN":["1662-5218"],"issn-type":[{"type":"electronic","value":"1662-5218"}],"subject":[],"published":{"date-parts":[[2023,8,31]]},"article-number":"1204418"}}