{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,8]],"date-time":"2026-07-08T14:13:55Z","timestamp":1783520035860,"version":"3.55.0"},"reference-count":53,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2025,2,18]],"date-time":"2025-02-18T00:00:00Z","timestamp":1739836800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2025,3,31]]},"abstract":"<jats:p>\n            In autonomous driving, deep models have shown remarkable performance across various visual perception tasks with the demand of high-quality and huge-diversity training datasets. Such datasets are expected to cover various driving scenarios with adverse weather, lighting conditions, and diverse moving objects. However, manually collecting these data presents huge challenges and is expensive. With the rapid development of large generative models, we propose DriveDiTFit, a novel method for efficiently generating autonomous\n            <jats:italic>Driv<\/jats:italic>\n            ing data by\n            <jats:italic>Fi<\/jats:italic>\n            ne-\n            <jats:italic>t<\/jats:italic>\n            uning pre-trained\n            <jats:italic>Di<\/jats:italic>\n            ffusion\n            <jats:italic>T<\/jats:italic>\n            ransformers (DiTs). Specifically, DriveDiTFit utilizes a gap-driven modulation technique to carefully select and efficiently fine-tune a few parameters in DiTs according to the discrepancy between the pre-trained source data and the target driving data. Additionally, DriveDiTFit develops an effective weather and lighting condition embedding module to ensure diversity in the generated data, which is initialized by a nearest-semantic-similarity initialization approach. Through progressive tuning scheme to refine the process of detail generation in early diffusion process and enlarging the weights corresponding to small objects in training loss, DriveDiTFit ensures high-quality generation of small moving objects in the generated data. Extensive experiments conducted on driving datasets confirm that our method could efficiently produce diverse real driving data.\n          <\/jats:p>","DOI":"10.1145\/3712064","type":"journal-article","created":{"date-parts":[[2025,1,13]],"date-time":"2025-01-13T15:37:23Z","timestamp":1736782643000},"page":"1-29","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":5,"title":["DriveDiTFit: Fine-tuning Diffusion Transformers for Autonomous Driving Data Generation"],"prefix":"10.1145","volume":"21","author":[{"ORCID":"https:\/\/orcid.org\/0009-0000-2763-6840","authenticated-orcid":false,"given":"Jiahang","family":"Tu","sequence":"first","affiliation":[{"name":"Zhejiang University, Hangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8106-9768","authenticated-orcid":false,"given":"Wei","family":"Ji","sequence":"additional","affiliation":[{"name":"National University of Singapore, Singapore, Singapore"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8906-4534","authenticated-orcid":false,"given":"Hanbin","family":"Zhao","sequence":"additional","affiliation":[{"name":"Zhejiang University, Hangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7174-8663","authenticated-orcid":false,"given":"Chao","family":"Zhang","sequence":"additional","affiliation":[{"name":"Zhejiang University, Hangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7410-2590","authenticated-orcid":false,"given":"Roger","family":"Zimmermann","sequence":"additional","affiliation":[{"name":"National University of Singapore, Singapore, Singapore"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0293-2656","authenticated-orcid":false,"given":"Hui","family":"Qian","sequence":"additional","affiliation":[{"name":"Zhejiang University, Hangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,2,18]]},"reference":[{"key":"e_1_3_1_2_2","first-page":"20178","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Alibeigi Mina","year":"2023","unstructured":"Mina Alibeigi, William Ljungbergh, Adam Tonderski, Georg Hess, Adam Lilja, Carl Lindstr\u00f6m, Daria Motorniuk, Junsheng Fu, Jenny Widahl, and Christoffer Petersson. 2023. Zenseact Open Dataset: A large-scale and diverse multimodal dataset for autonomous driving. In Proceedings of the IEEE\/CVF International Conference on Computer Vision, 20178\u201320188."},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.02171"},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.02161"},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.01164"},{"key":"e_1_3_1_6_2","first-page":"461","volume-title":"Proceedings of the 37th AAAI Conference on Artificial Intelligence","author":"Cong Peishan","year":"2023","unstructured":"Peishan Cong, Yiteng Xu, Yiming Ren, Juze Zhang, Lan Xu, Jingya Wang, Jingyi Yu, and Yuexin Ma. 2023. Weakly supervised 3d multi-person pose estimation for large-scale scenes based on monocular camera and single LiDAR. In Proceedings of the 37th AAAI Conference on Artificial Intelligence, 461\u2013469."},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"e_1_3_1_8_2","unstructured":"Jacob Devlin Ming-Wei Chang Kenton Lee and Kristina Toutanova. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv:1810.04805. Retrieved from https:\/\/arxiv.org\/abs\/1810.04805"},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.02069"},{"key":"e_1_3_1_10_2","unstructured":"Alexey Dosovitskiy Lucas Beyer Alexander Kolesnikov Dirk Weissenborn Xiaohua Zhai Thomas Unterthiner Mostafa Dehghani Matthias Minderer Georg Heigold Sylvain Gelly et al. 2020. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv:2010.11929. Retrieved from https:\/\/arxiv.org\/abs\/2010.11929"},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","DOI":"10.5555\/3692070.3692573"},{"key":"e_1_3_1_12_2","unstructured":"Qian Feng Hanbin Zhao Chao Zhang Jiahua Dong Henghui Ding Yu-Gang Jiang and Hui Qian. 2024. PECTP: Parameter-efficient cross-task prompts for incremental vision transformer. arXiv:2407.03813. Retrieved from https:\/\/arxiv.org\/abs\/2407.03813"},{"key":"e_1_3_1_13_2","unstructured":"Rinon Gal Yuval Alaluf Yuval Atzmon Or Patashnik Amit H. Bermano Gal Chechik and Daniel Cohen-Or. 2022. An image is worth one word: Personalizing text-to-image generation using textual inversion. arXiv:2208.01618. Retrieved from https:\/\/arxiv.org\/abs\/2208.01618"},{"key":"e_1_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.5555\/2354409.2354978"},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.1145\/3422622"},{"key":"e_1_3_1_17_2","first-page":"6629","article-title":"GANs trained by a two time-scale update rule converge to a local Nash equilibrium","volume":"30","author":"Heusel Martin","year":"2017","unstructured":"Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. 2017. GANs trained by a two time-scale update rule converge to a local Nash equilibrium. Advances in Neural Information Processing Systems 30 (2017), 6629\u20136640.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_18_2","unstructured":"Jonathan Ho William Chan Chitwan Saharia Jay Whang Ruiqi Gao Alexey Gritsenko Diederik P. Kingma Ben Poole Mohammad Norouzi David J. Fleet et al. 2022. Imagen video: High definition video generation with diffusion models. arXiv:2210.02303. Retrieved from https:\/\/arxiv.org\/abs\/2210.02303"},{"key":"e_1_3_1_19_2","first-page":"6840","article-title":"Denoising diffusion probabilistic models","volume":"33","author":"Ho Jonathan","year":"2020","unstructured":"Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020. Denoising diffusion probabilistic models. Advances in Neural Information Processing Systems 33 (2020), 6840\u20136851.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_20_2","unstructured":"Emiel Hoogeboom Jonathan Heek and Tim Salimans. 2023. Simple diffusion: End-to-end diffusion for high resolution images. arXiv:2301.11093. Retrieved from https:\/\/arxiv.org\/abs\/2301.11093"},{"key":"e_1_3_1_21_2","unstructured":"Edward J. Hu Yelong Shen Phillip Wallis Zeyuan Allen-Zhu Yuanzhi Li Shean Wang Lu Wang and Weizhu Chen. 2021. LoRA: Low-rank adaptation of large language models. arXiv:2106.09685. Retrieved from https:\/\/arxiv.org\/abs\/2106.09685"},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.167"},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-19827-4_41"},{"key":"e_1_3_1_24_2","unstructured":"Diederik P. Kingma and Max Welling. 2013. Auto-encoding variational bayes. arXiv:1312.6114. Retrieved from https:\/\/arxiv.org\/abs\/1312.6114"},{"key":"e_1_3_1_25_2","unstructured":"Alex Krizhevsky and Geoffrey Hinton. 2009.\u00a0Learning Multiple Layers of Features from Tiny Images. Master's thesis University of Tront."},{"key":"e_1_3_1_26_2","first-page":"3927","article-title":"Improved precision and recall metric for assessing generative models","volume":"32","author":"Kynk\u00e4\u00e4nniemi Tuomas","year":"2019","unstructured":"Tuomas Kynk\u00e4\u00e4nniemi, Tero Karras, Samuli Laine, Jaakko Lehtinen, and Timo Aila. 2019. Improved precision and recall metric for assessing generative models. Advances in Neural Information Processing Systems 32 (2019), 3927\u20133936.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_27_2","unstructured":"Chenlin Meng Yutong He Yang Song Jiaming Song Jiajun Wu Jun-Yan Zhu and Stefano Ermon. 2021. Sdedit: Guided image synthesis and editing with stochastic differential equations. arXiv:2108.01073. Retrieved from https:\/\/arxiv.org\/abs\/2108.01073"},{"key":"e_1_3_1_28_2","volume-title":"Proceedings of the NeurIPS 2022 Workshop on Score-Based Methods","author":"Moon Taehong","year":"2022","unstructured":"Taehong Moon, Moonseok Choi, Gayoung Lee, Jung-Woo Ha, and Juho Lee. 2022. Fine-tuning diffusion models with limited data. In Proceedings of the NeurIPS 2022 Workshop on Score-Based Methods."},{"key":"e_1_3_1_29_2","unstructured":"Charlie Nash Jacob Menick Sander Dieleman and Peter W. Battaglia. 2021. Generating images with sparse representations. arXiv:2103.03841. Retrieved from https:\/\/arxiv.org\/abs\/2103.03841"},{"key":"e_1_3_1_30_2","first-page":"8162","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Quinn Nichol Alexander","year":"2021","unstructured":"Alexander Quinn Nichol and Prafulla Dhariwal. 2021. Improved denoising diffusion probabilistic models. In Proceedings of the International Conference on Machine Learning. PMLR, 8162\u20138171."},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICVGIP.2008.47"},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.00387"},{"key":"e_1_3_1_33_2","first-page":"8748","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Radford Alec","year":"2021","unstructured":"Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. In Proceedings of the International Conference on Machine Learning. PMLR, 8748\u20138763."},{"key":"e_1_3_1_34_2","unstructured":"Aditya Ramesh Prafulla Dhariwal Alex Nichol Casey Chu and Mark Chen. 2022. Hierarchical text-conditional image generation with clip latents. arXiv:2204.06125. Retrieved from https:\/\/arxiv.org\/abs\/2204.06125"},{"key":"e_1_3_1_35_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.91"},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01042"},{"key":"e_1_3_1_37_2","first-page":"234","volume-title":"Proceedings of the 18th International Conference of Medical Image Computing and Computer-Assisted Intervention (MICCAI \u201915), part III","author":"Ronneberger Olaf","year":"2015","unstructured":"Olaf Ronneberger, Philipp Fischer, and Thomas Brox. 2015. U-Net: Convolutional networks for biomedical image segmentation. In Proceedings of the 18th International Conference of Medical Image Computing and Computer-Assisted Intervention (MICCAI \u201915), part III. Springer, 234\u2013241."},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.02155"},{"key":"e_1_3_1_39_2","unstructured":"Jiaming Song Chenlin Meng and Stefano Ermon. 2020. Denoising diffusion implicit models. arXiv:2010.02502. Retrieved from https:\/\/arxiv.org\/abs\/2010.02502"},{"key":"e_1_3_1_40_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2023.127063"},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00252"},{"key":"e_1_3_1_42_2","first-page":"17482","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Tao Runzhou","year":"2023","unstructured":"Runzhou Tao, Wencheng Han, Zhongying Qiu, Cheng-zhong Xu, and Jianbing Shen. 2023. Weakly supervised monocular 3D object detection using multi-view projection and direction consistency. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 17482\u201317492."},{"key":"e_1_3_1_43_2","unstructured":"Jiahang Tu Hao Fu Fengyu Yang Hanbin Zhao Chao Zhang and Hui Qian. 2024. Texttoucher: Fine-grained text-to-touch generation. arXiv:2409.05427. Retrieved from https:\/\/arxiv.org\/abs\/2409.05427"},{"key":"e_1_3_1_44_2","doi-asserted-by":"publisher","DOI":"10.1109\/ADICS58448.2024.10533619"},{"issue":"2017","key":"e_1_3_1_45_2","first-page":"6000","article-title":"Attention is all you need","volume":"30","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in Neural Information Processing Systems 30 (2017), 6000\u20136010.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_46_2","first-page":"4302","volume-title":"Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision","author":"Wang Boyang","year":"2024","unstructured":"Boyang Wang, Bowen Liu, Shiyu Liu, and Fengyu Yang. 2024. VCISR: Blind single image super-resolution with video compression synthetic data. In Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision, 4302\u20134312."},{"key":"e_1_3_1_47_2","first-page":"25574","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Wang Boyang","year":"2024","unstructured":"Boyang Wang, Fengyu Yang, Xihang Yu, Chao Zhang, and Hanbin Zhao. 2024. APISR: Anime production inspired real-world anime super-resolution. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 25574\u201325584."},{"key":"e_1_3_1_48_2","doi-asserted-by":"crossref","unstructured":"Fangyikang Wang Huminhao Zhu Chao Zhang Hanbin Zhao and Hui Qian. 2024. GAD-PVI: A general accelerated dynamic-weight particle-based variational inference framework. In Proceedings of the 38th AAAI Conference on Artificial Intelligence 15466\u201315473.","DOI":"10.1609\/aaai.v38i14.29472"},{"key":"e_1_3_1_49_2","unstructured":"Yuqi Wang Jiawei He Lue Fan Hongxin Li Yuntao Chen and Zhaoxiang Zhang. 2023. Driving into the future: Multiview visual forecasting and planning with world model for autonomous driving. arXiv:2311.17918. Retrieved from https:\/\/arxiv.org\/abs\/2311.17918"},{"key":"e_1_3_1_50_2","unstructured":"Yuqing Wen Yucheng Zhao Yingfei Liu Fan Jia Yanhui Wang Chong Luo Chi Zhang Tiancai Wang Xiaoyan Sun and Xiangyu Zhang. 2023. Panacea: Panoramic and controllable video generation for autonomous driving. arXiv:2311.16813. Retrieved from https:\/\/arxiv.org\/abs\/2311.16813"},{"key":"e_1_3_1_51_2","doi-asserted-by":"publisher","DOI":"10.1145\/3503161.3547821"},{"key":"e_1_3_1_52_2","first-page":"4230","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV)","author":"Xie Enze","year":"2023","unstructured":"Enze Xie, Lewei Yao, Han Shi, Zhili Liu, Daquan Zhou, Zhaoqiang Liu, Jiawei Li, and Zhenguo Li. 2023. DiffFit: Unlocking transferability of large diffusion models via simple parameter-efficient fine-tuning. In Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV), 4230\u20134239."},{"key":"e_1_3_1_53_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00271"},{"key":"e_1_3_1_54_2","doi-asserted-by":"crossref","unstructured":"Elad Ben Zaken Shauli Ravfogel and Yoav Goldberg. 2021. BitFit: Simple parameter-efficient fine-tuning for transformer-based masked language-models. arXiv:2106.10199. Retrieved from https:\/\/arxiv.org\/abs\/2106.10199","DOI":"10.18653\/v1\/2022.acl-short.1"},{"key":"e_1_3_1_55_2","unstructured":"Huminhao Zhu Fangyikang Wang Chao Zhang Hanbin Zhao and Hui Qian. 2024. Neural Sinkhorn gradient flow. arXiv:2401.14069. Retrieved from https:\/\/arxiv.org\/abs\/2401.14069"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3712064","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3712064","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T01:18:11Z","timestamp":1750295891000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3712064"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,2,18]]},"references-count":53,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2025,3,31]]}},"alternative-id":["10.1145\/3712064"],"URL":"https:\/\/doi.org\/10.1145\/3712064","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,2,18]]},"assertion":[{"value":"2024-07-23","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-01-05","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-02-18","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}