{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,12]],"date-time":"2026-02-12T08:02:28Z","timestamp":1770883348718,"version":"3.50.1"},"reference-count":67,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2026,2,12]],"date-time":"2026-02-12T00:00:00Z","timestamp":1770854400000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2026,2,12]],"date-time":"2026-02-12T00:00:00Z","timestamp":1770854400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["82441024"],"award-info":[{"award-number":["82441024"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["62202034"],"award-info":[{"award-number":["62202034"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Beijing Natural Science Foundation","award":["4242044"],"award-info":[{"award-number":["4242044"]}]},{"DOI":"10.13039\/501100004750","name":"Aeronautical Science Foundation of China","doi-asserted-by":"publisher","award":["2023Z071051002"],"award-info":[{"award-number":["2023Z071051002"]}],"id":[{"id":"10.13039\/501100004750","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100012226","name":"Fundamental Research Funds for the Central Universities","doi-asserted-by":"publisher","award":["None"],"award-info":[{"award-number":["None"]}],"id":[{"id":"10.13039\/501100012226","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Vis. Intell."],"published-print":{"date-parts":[[2026,12]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>Drivable area segmentation (DAS) plays an important role in autonomous driving. Segment anything model (SAM) has recently emerged as a powerful foundation model, demonstrating remarkable potential across diverse downstream segmentation tasks through domain-specific parameter-efficient fine-tuning (PEFT). This paper explores effective adaptation strategies for applying SAM to DAS. However, existing approaches suffer from the following two limitations: 1) SAM employs a vanilla vision transformer (ViT) as its image encoder. However, the ViT struggles to extract multi-scale features without incurring substantial computational overhead; 2) current fine-tuning approaches for SAM have been found to inadequately explore traffic scene context. Thus, they are not fully optimized for DAS and leave much room for improvement. To address the above issues, we propose segment anything model for drivable area segmentation termed as DAS-SAM, a novel efficient adaption framework that fine-tunes SAM towards DAS. Our approach incorporates a lightweight, learnable network to extract multi-scale features and introduces three auxiliary learning objectives to incorporate traffic scene context. Furthermore, DAS-SAM employs mosaic image augmentation to improve robustness and generalization. Our framework is compatible with most of the existing PEFT methods, allowing for flexible integration that boosts performance. Extensive experiments on the BDD100k and Cityscapes datasets demonstrate that DAS-SAM outperforms both full fine-tuning and state-of-the-art PEFT methods.<\/jats:p>","DOI":"10.1007\/s44267-026-00109-1","type":"journal-article","created":{"date-parts":[[2026,2,12]],"date-time":"2026-02-12T06:11:55Z","timestamp":1770876715000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["DAS-SAM: fine-tuning SAM towards drivable area segmentation via efficient multi-scale traffic scene-aware adaptation"],"prefix":"10.1007","volume":"4","author":[{"ORCID":"https:\/\/orcid.org\/0009-0009-1646-1657","authenticated-orcid":false,"given":"Zhenghao","family":"Chen","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3443-6171","authenticated-orcid":false,"given":"Nan","family":"Zhou","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yi","family":"Fan","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Lina","family":"Zhou","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yubao","family":"Xie","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0112-4166","authenticated-orcid":false,"given":"Jiaxin","family":"Chen","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2412-9330","authenticated-orcid":false,"given":"Di","family":"Huang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2026,2,12]]},"reference":[{"key":"109_CR1","unstructured":"Che, Q.-H., Le, D.-T., Pham, M.-Q., Nguyen, V.-T., & Lam, D.-K. (2024). TwinLiteNet+: a stronger model for real-time drivable area and lane segmentation. arXiv preprint. arXiv:2403.16958."},{"key":"109_CR2","doi-asserted-by":"publisher","first-page":"2530","DOI":"10.1109\/TITS.2024.3492383","volume":"26","author":"X. Sheng","year":"2025","unstructured":"Sheng, X., Zhang, J.-Z., Wang, Z., & Duan, Z.-T. (2025). Edgeunet: edge-guided multi-loss network for drivable area and lane segmentation in autonomous vehicles. IEEE Transactions on Intelligent Transportation Systems, 26, 2530\u20132542.","journal-title":"IEEE Transactions on Intelligent Transportation Systems"},{"key":"109_CR3","first-page":"1","volume-title":"Proceedings of the IEEE international conference on computing and machine intelligence","author":"H. Wang","year":"2024","unstructured":"Wang, H., Wang, J., Xiao, B., Jiao, Y., & Guo, J. (2024). Drivable area and lane line detection model based on semantic segmentation. In Proceedings of the IEEE international conference on computing and machine intelligence (pp. 1\u20137). Piscataway: IEEE."},{"key":"109_CR4","first-page":"248","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition","author":"J. Deng","year":"2009","unstructured":"Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., & Li, F.-F. (2009). ImageNet: a large-scale hierarchical image database. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 248\u2013255). Piscataway: IEEE."},{"key":"109_CR5","first-page":"4015","volume-title":"Proceedings of the IEEE\/CVF international conference on computer vision","author":"A. Kirillov","year":"2023","unstructured":"Kirillov, A., Mintun, E., Ravi, N., Mao, H., Rolland, C., Gustafson, L., Xiao, T., Whitehead, S., Berg, A. C., Lo, W.-Y., et al. (2023). Segment anything. In Proceedings of the IEEE\/CVF international conference on computer vision (pp. 4015\u20134026). Piscataway: IEEE."},{"key":"109_CR6","first-page":"1","volume-title":"Proceedings of the 61st annual meeting of the Association for Computational Linguistics","author":"E. B. Zaken","year":"2022","unstructured":"Zaken, E. B., Ravfogel, S., & Goldberg, Y. (2022). BitFit: simple parameter-efficient fine-tuning for transformer-based masked language-models. In S. Muresan, P. Nakov, & A. Villavicencio (Eds.), Proceedings of the 61st annual meeting of the Association for Computational Linguistics (pp. 1\u20139). Stroudsburg: ACL."},{"key":"109_CR7","first-page":"418","volume-title":"Proceedings of the 17th European conference on computer vision","author":"M. Jia","year":"2022","unstructured":"Jia, M., Tang, L., Chen, B.-C., Cardie, C., Belongie, S., Hariharan, B., & Lim, S.-N. (2022). Visual prompt tuning. In S. Avidan, G. J. Brostow, M. Ciss\u00e9, G. M. Farinella, & T. Hassner (Eds.), Proceedings of the 17th European conference on computer vision (pp. 418\u2013434). Cham: Springer."},{"key":"109_CR8","first-page":"109","volume-title":"Proceedings of the 36th international conference on neural information processing systems","author":"D. Lian","year":"2022","unstructured":"Lian, D., Zhou, D., Feng, J., & Wang, X. (2022). Scaling & shifting your features: a new baseline for efficient model tuning. In S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, & A. Oh (Eds.), Proceedings of the 36th international conference on neural information processing systems (pp. 109\u2013123). Red Hook: Curran Associates."},{"key":"109_CR9","first-page":"12513","volume-title":"Proceedings of the 10th international conference on learning representations","author":"E. J. Hu","year":"2022","unstructured":"Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., & Chen, W. (2022). LoRA: low-rank adaptation of large language models. In Proceedings of the 10th international conference on learning representations (pp. 12513\u201312525). Retrieved December 5, 2025, from https:\/\/openreview.net\/forum?id=nZeVKeeFYf9."},{"key":"109_CR10","first-page":"32100","volume-title":"Proceedings of the international conference on machine learning","author":"S.-Y. Liu","year":"2024","unstructured":"Liu, S.-Y., Wang, C.-Y., Yin, H., Molchanov, P., Wang, Y.-C. F., Cheng, K.-T., & Chen, M.-H. (2024). DoRA: weight-decomposed low-rank adaptation. In R. Salakhutdinov, Z. Kolter, K. Heller, A. Weller, N. Oliver, J. Scarlett, & F. Berkenkamp (Eds.), Proceedings of the international conference on machine learning (pp. 32100\u201332121). Retrieved December 5, 2025, from https:\/\/proceedings.mlr.press\/v235\/liu24bn.html."},{"key":"109_CR11","volume-title":"Proceedings of the AAAI conference on artificial intelligence","author":"Y. Zhang","year":"2026","unstructured":"Zhang, Y., Wu, Z., Liao, S., Wu, S., & Chen, J. (2026). MP-ISMoE: mixed-precision interactive side mixture-of-experts for efficient transfer learning. In Proceedings of the AAAI conference on artificial intelligence."},{"key":"109_CR12","doi-asserted-by":"publisher","first-page":"4527","DOI":"10.1109\/TIP.2025.3587587","volume":"34","author":"N. Zhou","year":"2025","unstructured":"Zhou, N., Chen, J., & Huang, D. (2025). Sharing task-relevant information in visual prompt tuning by cross-layer dynamic connection. IEEE Transactions on Image Processing, 34, 4527\u20134540.","journal-title":"IEEE Transactions on Image Processing"},{"key":"109_CR13","volume-title":"Proceedings of the annual conference on neural information processing systems","author":"H. Zhong","year":"2024","unstructured":"Zhong, H., Chen, J., Zhang, Y., Huang, D., & Wang, Y. (2024). Transforming vision transformer: towards efficient multi-task asynchronous learner. In Proceedings of the annual conference on neural information processing systems."},{"key":"109_CR14","first-page":"611","volume-title":"Proceedings of the 9th international conference on learning representations","author":"A. Dosovitskiy","year":"2021","unstructured":"Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al. (2021). An image is worth 16x16 words: transformers for image recognition at scale. In Proceedings of the 9th international conference on learning representations (pp. 611\u2013631). Retrieved December 5, 2025, from https:\/\/openreview.net\/forum?id=YicbFdNTTy."},{"key":"109_CR15","first-page":"2881","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition","author":"H. Zhao","year":"2017","unstructured":"Zhao, H., Shi, J., Qi, X., Wang, X., & Jia, J. (2017). Pyramid scene parsing network. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 2881\u20132890). Piscataway: IEEE."},{"key":"109_CR16","first-page":"4604","volume-title":"Proceedings of the 11th international conference on learning representations","author":"Z. Chen","year":"2023","unstructured":"Chen, Z., Duan, Y., Wang, W., He, J., Lu, T., Dai, J., & Qiao, Y. (2023). Vision transformer adapter for dense predictions. In Proceedings of the 11th international conference on learning representations (pp. 4604\u20134623). Retrieved December 5, 2025, from https:\/\/openreview.net\/forum?id=plKu2GByCNW."},{"key":"109_CR17","first-page":"2636","volume-title":"Proceedings of the IEEE\/CVF international conference on computer vision","author":"F. Yu","year":"2020","unstructured":"Yu, F., Chen, H., Wang, X., Xian, W., Chen, Y., Liu, F., Madhavan, V., & Darrell, T. (2020). Bdd100k: a diverse driving dataset for heterogeneous multitask learning. In Proceedings of the IEEE\/CVF international conference on computer vision (pp. 2636\u20132645). Piscataway: IEEE."},{"key":"109_CR18","first-page":"3213","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition","author":"M. Cordts","year":"2016","unstructured":"Cordts, M., Omran, M., Ramos, S., Rehfeld, T., Enzweiler, M., Benenson, R., Franke, U., Roth, S., & Schiele, B. (2016). The cityscapes dataset for semantic urban scene understanding. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 3213\u20133223). Piscataway: IEEE."},{"key":"109_CR19","first-page":"5998","volume-title":"Proceedings of the 31st international conference on neural information processing systems","author":"A. Vaswani","year":"2017","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, \u0141., & Polosukhin, I. (2017). Attention is all you need. In I. Guyon, U. von Luxburg, S. Bengio, H. M. Wallach, R. Fergus, S. V. N. Vishwanathan, & R. Garnett (Eds.), Proceedings of the 31st international conference on neural information processing systems (pp. 5998\u20136008). Red Hook: Curran Associates."},{"key":"109_CR20","first-page":"3431","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition","author":"J. Long","year":"2015","unstructured":"Long, J., Shelhamer, E., & Darrell, T. (2015). Fully convolutional networks for semantic segmentation. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 3431\u20133440). Piscataway: IEEE."},{"key":"109_CR21","unstructured":"Simonyan, K., & Zisserman, A. (2014). Very deep convolutional networks for large-scale image recognition. arXiv preprint. arXiv:1409.1556."},{"key":"109_CR22","first-page":"770","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition","author":"K. He","year":"2016","unstructured":"He, K., Zhang, X., Ren, S., & Sun, J. (2016). Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 770\u2013778). Piscataway: IEEE."},{"key":"109_CR23","first-page":"1520","volume-title":"Proceedings of the IEEE international conference on computer vision","author":"H. Noh","year":"2015","unstructured":"Noh, H., Hong, S., & Han, B. (2015). Learning deconvolution network for semantic segmentation. In Proceedings of the IEEE international conference on computer vision (pp. 1520\u20131528). Piscataway: IEEE."},{"key":"109_CR24","first-page":"234","volume-title":"Proceedings of the 18th international conference on medical image computing and computer-assisted intervention","author":"O. Ronneberger","year":"2015","unstructured":"Ronneberger, O., Fischer, P., & Brox, T. (2015). U-Net: convolutional networks for biomedical image segmentation. In N. Navab, J. Hornegger, W. M. Wells III, & A. F. Frangi (Eds.), Proceedings of the 18th international conference on medical image computing and computer-assisted intervention (pp. 234\u2013241). Cham: Springer."},{"issue":"12","key":"109_CR25","doi-asserted-by":"publisher","first-page":"2481","DOI":"10.1109\/TPAMI.2016.2644615","volume":"39","author":"V. Badrinarayanan","year":"2017","unstructured":"Badrinarayanan, V., Kendall, A., & Cipolla, R. (2017). SegNet: a deep convolutional encoder-decoder architecture for image segmentation. IEEE Transactions on Pattern Analysis and Machine Intelligence, 39(12), 2481\u20132495.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"109_CR26","unstructured":"Chen, L.-C., Papandreou, G., Kokkinos, I., Murphy, K., & Yuille, A. L. (2014). Semantic image segmentation with deep convolutional nets and fully connected CRFs. arXiv preprint. arXiv:1412.7062."},{"issue":"4","key":"109_CR27","doi-asserted-by":"publisher","first-page":"834","DOI":"10.1109\/TPAMI.2017.2699184","volume":"40","author":"L.-C. Chen","year":"2017","unstructured":"Chen, L.-C., Papandreou, G., Kokkinos, I., Murphy, K., & Yuille, A. L. (2017). DeepLab: semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected CRFs. IEEE Transactions on Pattern Analysis and Machine Intelligence, 40(4), 834\u2013848.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"109_CR28","unstructured":"Chen, L.-C., Papandreou, G., Schroff, F., & Adam, H. (2017). Rethinking atrous convolution for semantic image segmentation. arXiv preprint. arXiv:1706.05587."},{"key":"109_CR29","first-page":"801","volume-title":"Proceedings of the 15th European conference on computer vision","author":"L.-C. Chen","year":"2018","unstructured":"Chen, L.-C., Zhu, Y., Papandreou, G., Schroff, F., & Adam, H. (2018). Encoder-decoder with atrous separable convolution for semantic image segmentation. In V. Ferrari, M. Hebert, C. Sminchisescu, & Y. Weiss (Eds.), Proceedings of the 15th European conference on computer vision (pp. 801\u2013818). Cham: Springer."},{"key":"109_CR30","first-page":"1925","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition","author":"G. Lin","year":"2017","unstructured":"Lin, G., Milan, A., Shen, C., & Reid, I. (2017). RefineNet: multi-path refinement networks for high-resolution semantic segmentation. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 1925\u20131934). Piscataway: IEEE."},{"key":"109_CR31","first-page":"432","volume-title":"Proceedings of the 15th European conference on computer vision","author":"T. Xiao","year":"2018","unstructured":"Xiao, T., Liu, Y., Zhou, B., Jiang, Y., & Sun, J. (2018). Unified perceptual parsing for scene understanding. In V. Ferrari, M. Hebert, C. Sminchisescu, & Y. Weiss (Eds.), Proceedings of the 15th European conference on computer vision (pp. 432\u2013448). Cham: Springer."},{"key":"109_CR32","first-page":"6881","volume-title":"Proceedings of the IEEE\/CVF international conference on computer vision","author":"S. Zheng","year":"2021","unstructured":"Zheng, S., Lu, J., Zhao, H., Zhu, X., Luo, Z., Wang, Y., Fu, Y., Feng, J., Xiang, T., Torr, P. H., et al. (2021). Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers. In Proceedings of the IEEE\/CVF international conference on computer vision (pp. 6881\u20136890). Piscataway: IEEE."},{"key":"109_CR33","unstructured":"Chen, J., Lu, Y., Yu, Q., Luo, X., Adeli, E., Wang, Y., Lu, L., Yuille, A. L., & Zhou, Y. (2021). TransUNet: transformers make strong encoders for medical image segmentation. arXiv preprint. arXiv:2102.04306."},{"key":"109_CR34","first-page":"12077","volume-title":"Proceedings of the 35th international conference on neural information processing systems","author":"E. Xie","year":"2021","unstructured":"Xie, E., Wang, W., Yu, Z., Anandkumar, A., Alvarez, J. M., & Luo, P. (2021). SegFormer: simple and efficient design for semantic segmentation with transformers. In M. Ranzato, A. Beygelzimer, Y. N. Dauphin, P. Liang, & J. W. Vaughan (Eds.), Proceedings of the 35th international conference on neural information processing systems (pp. 12077\u201312090). Red Hook: Curran Associates."},{"key":"109_CR35","first-page":"568","volume-title":"Proceedings of the IEEE\/CVF international conference on computer vision","author":"W. Wang","year":"2021","unstructured":"Wang, W., Xie, E., Li, X., Fan, D.-P., Song, K., Liang, D., Lu, T., Luo, P., & Shao, L. (2021). Pyramid vision transformer: a versatile backbone for dense prediction without convolutions. In Proceedings of the IEEE\/CVF international conference on computer vision (pp. 568\u2013578). Piscataway: IEEE."},{"key":"109_CR36","first-page":"15804","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"N. Cavagnero","year":"2024","unstructured":"Cavagnero, N., Rosi, G., Cuttano, C., Pistilli, F., Ciccone, M., Averta, G., & Cermelli, F. (2024). PEM: prototype-based efficient maskformer for image segmentation. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 15804\u201315813). Piscataway: IEEE."},{"key":"109_CR37","first-page":"10012","volume-title":"Proceedings of the IEEE\/CVF international conference on computer vision","author":"Z. Liu","year":"2021","unstructured":"Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., & Guo, B. (2021). Swin transformer: hierarchical vision transformer using shifted windows. In Proceedings of the IEEE\/CVF international conference on computer vision (pp. 10012\u201310022). Piscataway: IEEE."},{"key":"109_CR38","unstructured":"Oquab, M., Darcet, T., Moutakanni, T., Vo, H., Szafraniec, M., Khalidov, V., Fernandez, P., Haziza, D., Massa, F., El-Nouby, A., et\u00a0al. (2024). DINOv2: learning robust visual features without supervision. arXiv preprint. arXiv:2304.07193."},{"key":"109_CR39","unstructured":"Bai, S., Chen, K., Liu, X., Wang, J., Ge, W., Song, S., Dang, K., Wang, P., Wang, S., Tang, J., et\u00a0al. (2025). Qwen2.5-VL technical report. arXiv preprint. arXiv:2502.13923."},{"key":"109_CR40","unstructured":"Lialin, V., Deshpande, V., & Rumshisky, A. (2023). Scaling down to scale up: a guide to parameter-efficient fine-tuning. arXiv preprint. arXiv:2303.15647."},{"key":"109_CR41","first-page":"4884","volume-title":"Proceedings of the 59th annual meeting of the association for computational linguistics","author":"D. Guo","year":"2021","unstructured":"Guo, D., Rush, A. M., & Kim, Y. (2021). Parameter-efficient transfer learning with diff pruning. In C. Zong, F. Xia, W. Li, & R. Navigli (Eds.), Proceedings of the 59th annual meeting of the association for computational linguistics (pp. 4884\u20134896). Stroudsburg: ACL."},{"key":"109_CR42","first-page":"11825","volume-title":"Proceedings of the IEEE\/CVF international conference on computer vision","author":"H. He","year":"2023","unstructured":"He, H., Cai, J., Zhang, J., Tao, D., & Zhuang, B. (2023). Sensitivity-aware visual parameter-efficient fine-tuning. In Proceedings of the IEEE\/CVF international conference on computer vision (pp. 11825\u201311835). Piscataway: IEEE."},{"key":"109_CR43","first-page":"2790","volume-title":"Proceedings of the international conference on machine learning","author":"N. Houlsby","year":"2019","unstructured":"Houlsby, N., Giurgiu, A., Jastrzebski, S., Morrone, B., De Laroussilhe, Q., Gesmundo, A., Attariyan, M., & Gelly, S. (2019). Parameter-efficient transfer learning for NLP. In K. Chaudhuri & R. Salakhutdinov (Eds.), Proceedings of the international conference on machine learning (pp. 2790\u20132799). Retrieved December 5, 2025, from https:\/\/proceedings.mlr.press\/v162\/li22n.html."},{"key":"109_CR44","doi-asserted-by":"crossref","unstructured":"Wang, Y., Mukherjee, S., Liu, X., Gao, J., Awadallah, A. H., & Gao, J. (2022). AdaMix: mixture-of-adapter for parameter-efficient tuning of large language models. arXiv preprint. arXiv:2205.12410.","DOI":"10.18653\/v1\/2022.emnlp-main.388"},{"key":"109_CR45","first-page":"19229","volume-title":"Proceedings of the 10th international conference on learning representations","author":"J. He","year":"2022","unstructured":"He, J., Zhou, C., Ma, X., Berg-Kirkpatrick, T., & Neubig, G. (2022). Towards a unified view of parameter-efficient transfer learning. In Proceedings of the 10th international conference on learning representations (pp. 19229\u201319243). Retrieved December 5, 2025, from https:\/\/openreview.net\/forum?id=0RDcd5Axok."},{"key":"109_CR46","first-page":"16664","volume-title":"Proceedings of the 36th international conference on neural information processing systems","author":"S. Chen","year":"2022","unstructured":"Chen, S., Ge, C., Tong, Z., Wang, J., Song, Y., Wang, J., & Luo, P. (2022). AdaptFormer: adapting vision transformers for scalable visual recognition. In S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, & A. Oh (Eds.), Proceedings of the 36th international conference on neural information processing systems (pp. 16664\u201316678). Red Hook: Curran Associates."},{"key":"109_CR47","doi-asserted-by":"publisher","first-page":"3045","DOI":"10.18653\/v1\/2021.emnlp-main.243","volume-title":"Proceedings of the 2021 conference on empirical methods in natural language processing","author":"B. Lester","year":"2021","unstructured":"Lester, B., Al-Rfou, R., & Constant, N. (2021). The power of scale for parameter-efficient prompt tuning. In M. Moens, X. Huang, L. Specia, & S. W. Yih (Eds.), Proceedings of the 2021 conference on empirical methods in natural language processing (pp. 3045\u20133059). Stroudsburg: ACL."},{"key":"109_CR48","first-page":"4582","volume-title":"Proceedings of the 59th annual meeting of the association for computational linguistics","author":"X. L. Li","year":"2021","unstructured":"Li, X. L., & Liang, P. (2021). Prefix-tuning: optimizing continuous prompts for generation. In C. Zong, F. Xia, W. Li, & R. Navigli (Eds.), Proceedings of the 59th annual meeting of the association for computational linguistics (pp. 4582\u20134597). Stroudsburg: ACL."},{"key":"109_CR49","unstructured":"Liu, X., Zheng, Y., Du, Z., Ding, M., Qian, Y., Yang, Z., & Tang, J. (2021). GPT understands, too. arXiv preprint. arXiv:2103.10385."},{"key":"109_CR50","doi-asserted-by":"crossref","unstructured":"Liu, X., Ji, K., Fu, Y., Tam, W. L., Du, Z., Yang, Z., & Tang, J. (2021). P-Tuning v2: prompt tuning can be comparable to fine-tuning universally across scales and tasks. arXiv preprint. arXiv:2110.07602.","DOI":"10.18653\/v1\/2022.acl-short.8"},{"key":"109_CR51","unstructured":"Edalati, A., Tahaei, M., Kobyzev, I., Nia, V. P., Clark, J. J., & Rezagholizadeh, M. (2022). KronA: parameter efficient tuning with Kronecker adapter. arXiv preprint. arXiv:2212.10650."},{"key":"109_CR52","first-page":"9082","volume-title":"Proceedings of the 12th international conference on learning representations","author":"Q. Zhang","year":"2023","unstructured":"Zhang, Q., Chen, M., Bukharin, A., Karampatziakis, N., He, P., Cheng, Y., Chen, W., & Zhao, T. (2023). Adaptive budget allocation for parameter-efficient fine-tuning. In Proceedings of the 12th international conference on learning representations (pp. 9082\u20139098). Retrieved December 5, 2025, from https:\/\/openreview.net\/forum?id=lq62uWRJjiY."},{"key":"109_CR53","first-page":"448","volume-title":"Proceedings of the international conference on machine learning","author":"S. Ioffe","year":"2015","unstructured":"Ioffe, S., & Szegedy, C. (2015). Batch normalization: accelerating deep network training by reducing internal covariate shift. In F. R. Bach & D. M. Blei (Eds.), Proceedings of the international conference on machine learning (pp. 448\u2013456). Retrieved December 5, 2025, from http:\/\/proceedings.mlr.press\/v37\/ioffe15.html."},{"key":"109_CR54","first-page":"2117","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition","author":"T.-Y. Lin","year":"2017","unstructured":"Lin, T.-Y., Doll\u00e1r, P., Girshick, R., He, K., Hariharan, B., & Belongie, S. (2017). Feature pyramid networks for object detection. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 2117\u20132125). Piscataway: IEEE."},{"key":"109_CR55","unstructured":"Bochkovskiy, A., Wang, C.-Y., & Liao, H.-Y. M. (2020). YOLOv4: optimal speed and accuracy of object detection. arXiv preprint. arXiv:2004.10934."},{"key":"109_CR56","first-page":"550","volume":"19","author":"D. Wu","year":"2022","unstructured":"Wu, D., Liao, M.-W., Zhang, W.-T., Wang, X.-G., Bai, X., Cheng, W.-Q., & Liu, W.-Y. (2022). YOLOP: you only look once for panoptic driving perception. International Journal of Automation and Computing, 19, 550\u2013562.","journal-title":"International Journal of Automation and Computing"},{"key":"109_CR57","unstructured":"Han, C., Zhao, Q., Zhang, S., Chen, Y., Zhang, Z., & Yuan, J. (2022). YOLOPv2: better, faster, stronger for panoptic driving perception. arXiv preprint. arXiv:2208.11434."},{"key":"109_CR58","first-page":"106","volume":"35","author":"V. T. Dat","year":"2025","unstructured":"Dat, V. T., Bao, N. V. H., & Hung, P. D. (2025). HybridNets: end-to-end perception network. IEEE Transactions on Pattern Analysis and Machine Intelligence, 35, 106\u2013118.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"109_CR59","first-page":"1013","volume-title":"Proceedings of the IEEE\/CVF international conference on computer vision","author":"Y. Hou","year":"2019","unstructured":"Hou, Y., Ma, Z., Liu, C., & Loy, C. C. (2019). Learning lightweight lane detection CNNs by self attention distillation. In Proceedings of the IEEE\/CVF international conference on computer vision (pp. 1013\u20131021). Piscataway: IEEE."},{"key":"109_CR60","first-page":"898","volume-title":"Proceedings of the IEEE\/CVF international conference on computer vision","author":"T. Zheng","year":"2022","unstructured":"Zheng, T., Huang, Y., Liu, Y., Tang, W., Yang, Z., Cai, D., & He, X. (2022). CLRNet: cross layer refinement network for lane detection. In Proceedings of the IEEE\/CVF international conference on computer vision (pp. 898\u2013907). Piscataway: IEEE."},{"key":"109_CR61","first-page":"12486","volume-title":"Proceedings of the IEEE\/CVF international conference on computer vision","author":"Y. Hou","year":"2020","unstructured":"Hou, Y., Ma, Z., Liu, C., Hui, T.-W., & Loy, C. C. (2020). Inter-region affinity distillation for road marking segmentation. In Proceedings of the IEEE\/CVF international conference on computer vision (pp. 12486\u201312495). Piscataway: IEEE."},{"key":"109_CR62","first-page":"14408","volume-title":"Proceedings of the IEEE\/CVF international conference on computer vision","author":"W. Wang","year":"2023","unstructured":"Wang, W., Dai, J., Chen, Z., Huang, Z., Li, Z., Zhu, X., Hu, X., Lu, T., Lu, L., Li, H., et al. (2023). InternImage: exploring large-scale vision foundation models with deformable convolutions. In Proceedings of the IEEE\/CVF international conference on computer vision (pp. 14408\u201314419). Piscataway: IEEE."},{"key":"109_CR63","unstructured":"MMSegmentation Contributors. MMSegmentation: Openmmlab semantic segmentation toolbox and benchmark. Retrieved December 5, 2025, from https:\/\/github.com\/open-mmlab\/mmsegmentation."},{"key":"109_CR64","first-page":"1","volume-title":"Proceedings of the 3rd international conference on learning representations","author":"D. P. Kingma","year":"2015","unstructured":"Kingma, D. P., & Ba, J. (2015). Adam: a method for stochastic optimization. In Proceedings of the 3rd international conference on learning representations (pp. 1\u201315). Retrieved December 5, 2025, from https:\/\/openreview.net\/forum?id=szXGN2CLjwf."},{"key":"109_CR65","first-page":"6023","volume-title":"Proceedings of the IEEE\/CVF international conference on computer vision","author":"S. Yun","year":"2019","unstructured":"Yun, S., Han, D., Oh, S. J., Chun, S., Choe, J., & Yoo, Y. (2019). CutMix: regularization strategy to train strong classifiers with localizable features. In Proceedings of the IEEE\/CVF international conference on computer vision (pp. 6023\u20136032). Piscataway: IEEE."},{"key":"109_CR66","first-page":"2866","volume-title":"Proceedings of the 6th international conference on learning representations","author":"H. Zhang","year":"2018","unstructured":"Zhang, H., Cisse, M., Dauphin, Y. N., & Lopez-Paz, D. (2018). mixup: beyond empirical risk minimization. In Proceedings of the 6th international conference on learning representations (pp. 2866\u20132878). Retrieved December 5, 2025, from https:\/\/openreview.net\/forum?id=r1Ddp1-Rb."},{"key":"109_CR67","first-page":"4990","volume-title":"Proceedings of the IEEE international conference on computer vision","author":"G. Neuhold","year":"2017","unstructured":"Neuhold, G., Ollmann, T., Rota Bulo, S., & Kontschieder, P. (2017). The mapillary vistas dataset for semantic understanding of street scenes. In Proceedings of the IEEE international conference on computer vision (pp. 4990\u20134999). Piscataway: IEEE."}],"container-title":["Visual Intelligence"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s44267-026-00109-1.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s44267-026-00109-1","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s44267-026-00109-1.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,2,12]],"date-time":"2026-02-12T07:03:44Z","timestamp":1770879824000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s44267-026-00109-1"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,2,12]]},"references-count":67,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2026,12]]}},"alternative-id":["109"],"URL":"https:\/\/doi.org\/10.1007\/s44267-026-00109-1","relation":{},"ISSN":["2097-3330","2731-9008"],"issn-type":[{"value":"2097-3330","type":"print"},{"value":"2731-9008","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,2,12]]},"assertion":[{"value":"17 August 2025","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"28 January 2026","order":2,"name":"revised","label":"Revised","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"28 January 2026","order":3,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"12 February 2026","order":4,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors have no relevant financial or non-financial interests to disclose.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"6"}}