{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,27]],"date-time":"2026-03-27T14:24:49Z","timestamp":1774621489545,"version":"3.50.1"},"reference-count":61,"publisher":"Wiley","license":[{"start":{"date-parts":[[2026,3,27]],"date-time":"2026-03-27T00:00:00Z","timestamp":1774569600000},"content-version":"vor","delay-in-days":0,"URL":"http:\/\/onlinelibrary.wiley.com\/termsAndConditions#vor"},{"start":{"date-parts":[[2026,3,27]],"date-time":"2026-03-27T00:00:00Z","timestamp":1774569600000},"content-version":"tdm","delay-in-days":0,"URL":"http:\/\/doi.wiley.com\/10.1002\/tdm_license_1.1"}],"content-domain":{"domain":["onlinelibrary.wiley.com"],"crossmark-restriction":true},"short-container-title":["Computer Graphics Forum"],"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>\n                    Diffusion\u2010based scene text synthesis has progressed rapidly, yet existing methods commonly rely on additional visual conditioning modules and require large\u2010scale annotated data to support multilingual generation. In this work, we revisit the necessity of complex auxiliary modules and further explore an approach that simultaneously ensures glyph accuracy and achieves high\u2010fidelity scene integration, by leveraging diffusion models' inherent capabilities for contextual reasoning. To this end, we introduce TextFlux, a DiT\u2010based framework that enables multilingual scene text synthesis. The advantages of TextFlux can be summarized as follows: (1) OCR\u2010free model architecture. TextFlux eliminates the need for OCR encoders that are specifically used to extract visual text\u2010related features. (2) Strong multilingual scalability. TextFlux is effective in low\u2010resource multilingual settings, and achieves strong performance in newly added languages with fewer than 1,000 samples. (3) Streamlined training setup. TextFlux is trained with only 1% of the training data required by competing methods. (4) Controllable multi\u2010line text generation. TextFlux offers flexible multi\u2010line synthesis with precise line\u2010level control, outperforming methods restricted to single\u2010line or rigid layouts. Extensive experiments and visualizations demonstrate that TextFlux outperforms previous methods in both qualitative and quantitative evaluations. Our code is available at\n                    <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"https:\/\/github.com\/yyyyyxie\/textflux\">https:\/\/github.com\/yyyyyxie\/textflux<\/jats:ext-link>\n                    .\n                  <\/jats:p>","DOI":"10.1111\/cgf.70342","type":"journal-article","created":{"date-parts":[[2026,3,27]],"date-time":"2026-03-27T13:29:04Z","timestamp":1774618144000},"update-policy":"https:\/\/doi.org\/10.1002\/crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["TextFlux: An OCR\u2010Free DiT Model for High\u2010Fidelity Multilingual Scene Text Synthesis"],"prefix":"10.1111","author":[{"ORCID":"https:\/\/orcid.org\/0009-0007-1695-9370","authenticated-orcid":false,"given":"Yu","family":"Xie","sequence":"first","affiliation":[{"name":"Bilibili Inc.  China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0005-9449-7189","authenticated-orcid":false,"given":"Jielei","family":"Zhang","sequence":"additional","affiliation":[{"name":"Bilibili Inc.  China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-0587-3771","authenticated-orcid":false,"given":"Pengyu","family":"Chen","sequence":"additional","affiliation":[{"name":"Bilibili Inc.  China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-4504-4441","authenticated-orcid":false,"given":"Weihang","family":"Wang","sequence":"additional","affiliation":[{"name":"Bilibili Inc.  China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0000-0621-2950","authenticated-orcid":false,"given":"Longwen","family":"Gao","sequence":"additional","affiliation":[{"name":"Bilibili Inc.  China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-7351-1986","authenticated-orcid":false,"given":"Peiyi","family":"Li","sequence":"additional","affiliation":[{"name":"Bilibili Inc.  China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0005-6154-1411","authenticated-orcid":false,"given":"Qian","family":"Qiao","sequence":"additional","affiliation":[{"name":"Bilibili Inc.  China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2683-7170","authenticated-orcid":false,"given":"Zhouhui","family":"Lian","sequence":"additional","affiliation":[{"name":"Wangxuan Institute of ComputerTechnology Peking University  China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"311","published-online":{"date-parts":[[2026,3,27]]},"reference":[{"key":"e_1_2_6_2_2","doi-asserted-by":"crossref","unstructured":"BrooksT. HolynskiA. EfrosA. A.: Instruetpix2pix: Learning to follow image editing instructions. InProceedings of the IEEE\/CVF conference on computer vision and pattern recognition(2023) pp.18392\u201318402. 3","DOI":"10.1109\/CVPR52729.2023.01764"},{"key":"e_1_2_6_3_2","unstructured":"BalahY. NahS. HuangX. VahdatA. SongJ. ZhangQ. KreisK. AittalaM. AilaT. LaineS. et al.: ediffi: Text-to-image diffusion models with an ensemble of expert denoisers.arXiv preprint arXiv:2211.01324(2022). 3"},{"issue":"1","key":"e_1_2_6_4_2","doi-asserted-by":"crossref","first-page":"31","DOI":"10.1007\/s10032-019-00334-z","article-title":"Total-text: toward orientation robustness in scene text detection","volume":"23","author":"Ch'ng C.-K.","year":"2020","journal-title":"International Journal on Document Analysis and Recognition (IJDAR)"},{"key":"e_1_2_6_5_2","doi-asserted-by":"crossref","first-page":"9353","DOI":"10.52202\/075280-0410","article-title":"Textdiffuser: Diffusion models as text painters","volume":"36","author":"Chen J.","year":"2023","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_2_6_6_2","first-page":"386","volume-title":"European Conference on Computer Vision","author":"Chen J.","year":"2024"},{"issue":"70","key":"e_1_2_6_7_2","first-page":"1","article-title":"Scaling instruction-finetuned language models","volume":"25","author":"Chung H. W.","year":"2024","journal-title":"Journal of Machine Learning Research"},{"key":"e_1_2_6_8_2","unstructured":"ChenL. MaoQ. GuY. ShouM. Z.: Edit transfer: Learning image editing via vision in-context relations.arXiv preprint arXiv:2503.13327(2025). 3"},{"key":"e_1_2_6_9_2","doi-asserted-by":"crossref","unstructured":"ChenH. XuX. LiW. RenJ. YeT. LiuS. ChenY.-C. ZhuL. WangX.: Posta: A go-to framework for customized artistic poster generation.arXiv preprint arXiv:2503.14908(2025). 3","DOI":"10.1109\/CVPR52734.2025.02672"},{"key":"e_1_2_6_10_2","unstructured":"ChenJ. YuJ. GeC. YaoL. XieE. WuY. WangZ. KwokJ. LuoP. LuH. LiZ.: Pixart-\u03b1: Fast training of diffusion transformer for photorealistic text-to-image synthesis.arXiv preprint arXiv:2310.00426(2023). 3"},{"key":"e_1_2_6_11_2","unstructured":"DuN. ChenZ. ChenZ. GaoS. ChenX. JiangZ. Yangj. TaiY.: Texterafter: Accurately rendering multiple texts in complex visual scenes.arXiv preprint arXiv:2503.23461(2025). 2"},{"key":"e_1_2_6_12_2","unstructured":"DeepFloyD:Github link:https:\/\/githubcom\/deep-floyd\/if 2023. URL:https:\/\/github.com\/deep-floyd\/IF. 3"},{"key":"e_1_2_6_13_2","first-page":"8780","article-title":"Diffusion models beat gans on image synthesis","volume":"34","author":"Dhariwal P.","year":"2021","journal-title":"Advances in neural information processing systems"},{"key":"e_1_2_6_14_2","unstructured":"FangZ. LyuP. WuJ. ZhangC. YuJ. LuG. PeiW.: Recognition-synergistic scene text editing.arXiv preprint arXiv:2503.08387(2025). 3"},{"key":"e_1_2_6_15_2","unstructured":"FengK. MaY. WangB. QiC. ChenH. ChenQ. WangZ.: Dit4edit: Diffusion transformer for image editing.arXiv preprint arXiv:2411.03286(2024). 3"},{"key":"e_1_2_6_16_2","first-page":"18225","article-title":"Layoutgpt: Compositional visual planning and generation with large language models","volume":"36","author":"Feng W.","year":"2023","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_2_6_17_2","unstructured":"GaoY. GongL. GuoQ. HouX. LaiZ. LiF. LiL. LianX. LiaoC. LiuL. et al.: Seedream 3.0 technical reportarXiv preprint arXiv:2504.11346(2025). 3"},{"key":"e_1_2_6_18_2","unstructured":"GongL. HouX. LiF. LiL. LianX. LiuF. LiuL. LiuW. LuW. ShiY. et al.: Seedream 2.0: A native chinese-english bilingual image generation foundation model.arXiv preprint arXiv:2503.07703(2025). 3"},{"key":"e_1_2_6_19_2","doi-asserted-by":"crossref","unstructured":"GaoY. LinZ. LiuC. ZhouM. GeT. ZhengB. XieH.: Postermaker: Towards high-quality product poster generation with accurate text rendering.arXiv preprint arXiv:2504.06632(2025). 3","DOI":"10.1109\/CVPR52734.2025.00757"},{"key":"e_1_2_6_20_2","unstructured":"Google:Gemini 2.5 flash image.https:\/\/ai.google.dev\/gemini-api\/docs\/image-generation#gemini 2025. 2 3 6 8 9"},{"key":"e_1_2_6_21_2","unstructured":"HertzA. MokadyR. TenenbaumJ. AbermanK. PritchY. Cohen-OrD.: Prompt-to-promptimage editing with cross attention control.arXiv preprint arXiv:2208.01626(2022). 3"},{"key":"e_1_2_6_22_2","unstructured":"HuangL. WangW. WuZ.-F. ShiY. DouH. LiangC. FengY. LiuY. ZhouJ.: In-context lora for diffusion transformers.arXiv preprint arXiv:2410.23775(2024). 3 5"},{"key":"e_1_2_6_23_2","unstructured":"HuX. XuK. LiuB. LiuQ. FeiH.: Amo sampler: Enhancing text rendering with overshooting.arXiv preprint arXiv:2411.19415(2024). 3"},{"key":"e_1_2_6_24_2","unstructured":"LabsB. F.:Flux: Official inference repository for flux.1 models 2024. Accessed: 2024-11-12. URL:https:\/\/github.com\/black-forest-labs\/flux. 2 3 4 5 6 10"},{"key":"e_1_2_6_25_2","unstructured":"LanR. BaiY. DuanX. LiM. JinD. XuR. NieD. SunL. ChuX.: Flux-text: A simple and advanced diffusion transformer baseline for scene text editing.arXiv preprint arXiv:2505.03329(2025). 6 7"},{"key":"e_1_2_6_26_2","unstructured":"LipmanY. ChenR. T. Ben-HamuH. NickelM. LeM.: Flow matching for generative modeling.arXiv preprint arXiv:2210.02747(2022). 3"},{"key":"e_1_2_6_27_2","unstructured":"LiuZ. LiangW. ZhaoY. ChenB. LiangL. WangL. LiJ. YuanY.: Glyph-byt5-v2: A strong aesthetic baseline for accurate multilingual visual text rendering.arXiv preprint arXiv:2406.10208(2024). 3"},{"key":"e_1_2_6_28_2","doi-asserted-by":"crossref","unstructured":"LiH. ShenC. TorrP. TrespV. GuJ.: Self-disco vering interpretable diffusion latent directions for responsible text-to-image generation. InProceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(2024) pp.12006\u201312016. 3","DOI":"10.1109\/CVPR52733.2024.01141"},{"key":"e_1_2_6_29_2","first-page":"129","volume-title":"European Conference on Computer Vision","author":"Li M.","year":"2024"},{"key":"e_1_2_6_30_2","unstructured":"MaJ. DengY. ChenC. DuN. LuH. YangZ.: Glyphdraw2: Automatic generation of complex glyph posters with diffusion models and large language models.arXiv preprint arXiv:2407.02252(2024). 2 3"},{"key":"e_1_2_6_31_2","unstructured":"Icdar 2019 robust reading challenge on multi-lingual scene text detection and recognition.https:\/\/rrc.cvc.uab.es\/?ch=15 2019. 6"},{"key":"e_1_2_6_32_2","unstructured":"MaN. TongS. JiaH. HuH. SuY.-C. ZhangM. YangX. LiY. JaakkolaT. JiaX. et al.: Inference-time scaling for diffusion models beyond scaling denoising steps.arXiv preprint arXiv:2501.09732(2025). 3"},{"key":"e_1_2_6_33_2","doi-asserted-by":"crossref","first-page":"4296","DOI":"10.1609\/aaai.v38i5.28226","article-title":"T2i-adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models","volume":"38","author":"Mou C.","year":"2024","journal-title":"Proceedings of the AAAI conference on artificial intelligence"},{"key":"e_1_2_6_34_2","unstructured":"MaJ. ZhaoM. ChenC. WangR. NiuD. LuH. LinX.: Glyphdraw: Seamlessly rendering text with intricate spatial structures in text-to-image generation.arXiv preprint arXiv:2303.17870(2023). 3"},{"key":"e_1_2_6_35_2","unstructured":"OpenAI:Gpt image-1.https:\/\/openai.com\/index\/introducing-4o-image-generation\/ 2025. 2 3 6 8 9"},{"key":"e_1_2_6_36_2","unstructured":"PodellD. EnglishZ. LaceyK. BlattmannA. DockhornT. M\u00fcllerJ. PennaJ. RombachR.: Sdxl: Improving latent diffusion models for high-resolution image synthesis.arXiv preprint arXiv:2307.01952(2023). 2"},{"key":"e_1_2_6_37_2","unstructured":"PeeblesW. XieS.: Scalable diffusion models with transformers. InProceedings of the IEEE\/CVF international conference on computer vision(2023) pp.4195\u20134205. 2 3"},{"key":"e_1_2_6_38_2","unstructured":"RombachR. BlattmannA. LorenzD. EsserP. OmmerB.:High-resolution image synthesis with latent diffusion models 2021. arXiv:2112.10752. 2 3"},{"key":"e_1_2_6_39_2","unstructured":"Icdar2017 competition on reading chinese text in the wild.https:\/\/rctw.vlrlab.net\/dataset 2017. 6"},{"key":"e_1_2_6_40_2","unstructured":"Icdar 2019 robust reading challenge on reading chinese text on signboard.https:\/\/rrc.cvc.uab.es\/?ch=12 2019. 6"},{"key":"e_1_2_6_41_2","first-page":"8748","volume-title":"International conference on machine learning","author":"Radford A.","year":"2021"},{"key":"e_1_2_6_42_2","doi-asserted-by":"crossref","unstructured":"RuizN. LiY. JampaniV. PritchY. RubinsteinM. AbermanK.: Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation. InProceedings of the IEEE\/CVF conference on computer vision and pattern recognition(2023) pp.22500\u201322510. 3","DOI":"10.1109\/CVPR52729.2023.02155"},{"key":"e_1_2_6_43_2","doi-asserted-by":"crossref","unstructured":"RuizN. LiY. JampaniV. WeiW. HouT. PritchY. WadhwaN. RubinsteinM. AbermanK.: Hyperdreambooth:Hypernetworks for fast personalization of text-to-image models. InProceedings of the IEEE\/CVF conference on computer vision and pattern recognition(2024) pp.6527\u20136536. 3","DOI":"10.1109\/CVPR52733.2024.00624"},{"key":"e_1_2_6_44_2","doi-asserted-by":"crossref","unstructured":"SahariaC. ChanW. ChangH. LeeC. HoJ. SalimansT. FleetD. NorouziM.: Palette: Image-to-image diffusion models. InACM SIGGRAPH 2022 conference proceedings(2022) pp.1\u201310. 3","DOI":"10.1145\/3528233.3530757"},{"key":"e_1_2_6_45_2","first-page":"36479","article-title":"Photorealistic text-to-image diffusion models with deep language understanding","volume":"35","author":"Saharia C.","year":"2022","journal-title":"Advances in neural information processing systems"},{"key":"e_1_2_6_46_2","unstructured":"TuoY. GengY. BoL.: Anytext2: Visual text generation and editing with customizable attributes.arXiv preprint arXiv:2411.15245(2024). 2 6"},{"key":"e_1_2_6_47_2","unstructured":"TanZ. LiuS. YangX. XueQ. WangX.: Ominicontrol: Minimal and universal control for diffusion transformer.arXiv preprint arXiv:2411.15098(2024). 3"},{"key":"e_1_2_6_48_2","unstructured":"TuoY. XiangW. HeJ.-Y. GengY. XieX.: Anytext: Multilingual visual text generation and editing.arXiv preprint arXiv:2311.03054(2023). 2 3 4 5 6"},{"key":"e_1_2_6_49_2","unstructured":"TanZ. XueQ. YangX. LiuS. WangX.: Ominicontrol2: Efficient conditioning for diffusion transformers.arXiv preprint arXiv:2503.08280(2025). 3"},{"key":"e_1_2_6_50_2","unstructured":"WuZ. GaoH. WangY. ZhangX. WangS.: Universal prompt optimizer for safe text-to-image generation.arXiv preprint arXiv:2402.10882(2024). 3"},{"key":"e_1_2_6_51_2","doi-asserted-by":"crossref","unstructured":"WuX. HuZ. ShengL. XuD.: Styleformer: Real-time arbitrary style transfer via parametric style composition. InProceedings of the IEEE\/CVF international conference on computer vision(2021) pp.14618\u201314627. 3","DOI":"10.1109\/ICCV48922.2021.01435"},{"key":"e_1_2_6_52_2","doi-asserted-by":"crossref","unstructured":"WangT. LiuT. QuX. WuC. LiuL. HuX.: Glyphmastero: A glyph encoder for high-fidelity scene text editing.arXiv preprint arXiv:2505.04915(2025). 3","DOI":"10.1109\/CVPR52734.2025.02656"},{"key":"e_1_2_6_53_2","unstructured":"WuC. LiJ. ZhouJ. LinJ. GaoK. YanK. YinS.-M. BaiS. XuX. ChenY. et al.: Qwen-image technical report.arXiv preprint arXiv:2508.02324(2025). 2 3 6 8 9"},{"key":"e_1_2_6_54_2","unstructured":"WangY. ZhangW. ZhouC. JinC.: High fidelity scene text synthesis.arXiv preprint arXiv:2405.14701(2024). 2 3 6 7"},{"key":"e_1_2_6_55_2","doi-asserted-by":"crossref","unstructured":"XieY. QiaoQ. GaoJ. WuT. FanJ. ZhangY. ZhangJ. SunH.: Dntextspotter: Arbitrary-shaped scene text spotting via improved denoising training.arXiv preprint arXiv:2408.00355(2024). 6","DOI":"10.1145\/3664647.3680981"},{"key":"e_1_2_6_56_2","first-page":"44050","article-title":"Glyphcontrol: glyph conditional control for visual text generation","volume":"36","author":"Yang Y.","year":"2023","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_2_6_57_2","unstructured":"YuS. KwakS. JangH. JeongJ. HuangJ. ShinJ. XieS.: Representation alignment for generation: Training diffusion transformers is easier than you think.arXiv preprint arXiv:2410.06940(2024). 3"},{"key":"e_1_2_6_58_2","doi-asserted-by":"crossref","unstructured":"YeM. ZhangJ. ZhaoS. LiuJ. LiuT. DuB. TaoD.: Deepsolo: Let transformer decoder with explicit points solo for text spotting. InProceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(2023) pp.19348\u201319357. 6","DOI":"10.1109\/CVPR52729.2023.01854"},{"key":"e_1_2_6_59_2","doi-asserted-by":"crossref","first-page":"7215","DOI":"10.1609\/aaai.v38i7.28550","article-title":"Brush your text: Synthesize any scene text on images via diffusion model","volume":"38","author":"Zhang L.","year":"2024","journal-title":"Proceedings of the AAAI Conference on Artificial Intelligence"},{"key":"e_1_2_6_60_2","unstructured":"ZhaoY. LianZ.: Udifftext: A unified framework for high-quality text synthesis in arbitrary images via character-aware diffusion models.arXiv preprint arXiv:2312.04884(2023). 2 3 6 7"},{"key":"e_1_2_6_61_2","unstructured":"ZhangL. RaoA. AgrawalaM.: Adding conditional control to text-to-image diffusion models. InProceedings of the IEEE\/CVF international conference on computer vision(2023) pp.3836\u20133847. 3"},{"key":"e_1_2_6_62_2","first-page":"138569","article-title":"Textctrl: Diffusion-based scene text editing with prior guidance control","volume":"37","author":"Zeng W.","year":"2024","journal-title":"Advances in Neural Information Processing Systems"}],"container-title":["Computer Graphics Forum"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/pdf\/10.1111\/cgf.70342","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/full-xml\/10.1111\/cgf.70342","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/pdf\/10.1111\/cgf.70342","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,3,27]],"date-time":"2026-03-27T13:29:22Z","timestamp":1774618162000},"score":1,"resource":{"primary":{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/10.1111\/cgf.70342"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,3,27]]},"references-count":61,"alternative-id":["10.1111\/cgf.70342"],"URL":"https:\/\/doi.org\/10.1111\/cgf.70342","archive":["Portico"],"relation":{},"ISSN":["0167-7055","1467-8659"],"issn-type":[{"value":"0167-7055","type":"print"},{"value":"1467-8659","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,3,27]]},"assertion":[{"value":"2026-03-27","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}],"article-number":"e70342"}}