{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,16]],"date-time":"2026-02-16T16:02:29Z","timestamp":1771257749661,"version":"3.50.1"},"reference-count":64,"publisher":"Institution of Engineering and Technology (IET)","issue":"1","license":[{"start":{"date-parts":[[2026,2,16]],"date-time":"2026-02-16T00:00:00Z","timestamp":1771200000000},"content-version":"vor","delay-in-days":46,"URL":"http:\/\/creativecommons.org\/licenses\/by-nc\/4.0\/"},{"start":{"date-parts":[[2026,1,1]],"date-time":"2026-01-01T00:00:00Z","timestamp":1767225600000},"content-version":"tdm","delay-in-days":0,"URL":"http:\/\/doi.wiley.com\/10.1002\/tdm_license_1.1"}],"funder":[{"DOI":"10.13039\/501100006465","name":"Korea Creative Content Agency","doi-asserted-by":"publisher","award":["RS\u20102024\u201000398320"],"award-info":[{"award-number":["RS\u20102024\u201000398320"]}],"id":[{"id":"10.13039\/501100006465","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["ietresearch.onlinelibrary.wiley.com"],"crossmark-restriction":true},"short-container-title":["IET Image Processing"],"published-print":{"date-parts":[[2026,1]]},"abstract":"<jats:title>ABSTRACT<\/jats:title>\n                  <jats:p>Recent advancements in text\u2010to\u2010image (T2I) models have enabled the synthesis of personalized images that align closely with user\u2010specified prompts, especially through the use of modifiers. However, generating multiple detailed objects with distinct modifiers in a single image remains challenging due to concept\u2010mixing, resulting from the difficulty of capturing interactions among text tokens. This paper proposes a modifier\u2010based approach to mitigate concept\u2010mixing by addressing the interaction among text tokens. Our method enables practical multi\u2010personalization while preserving the original T2I model's straightforward inference pipeline. Without structural guidance, it ensures seamless object interaction with enhanced consistency. Through a loss\u2010based finetuning approach, our method is adaptable to various concept\u2010learning algorithms, enabling plug\u2010and\u2010play functionality. Through both qualitative and quantitative evaluations, we demonstrate that our method effectively resolves concept\u2010mixing issues to better preserve concepts' identities and outperforms recent baselines in both quantitative and qualitative results. Our code will be publicly\u00a0available.<\/jats:p>","DOI":"10.1049\/ipr2.70306","type":"journal-article","created":{"date-parts":[[2026,2,16]],"date-time":"2026-02-16T15:03:36Z","timestamp":1771254216000},"update-policy":"https:\/\/doi.org\/10.1002\/crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["End\u2010to\u2010End Multi\u2010Entity Customization"],"prefix":"10.1049","volume":"20","author":[{"ORCID":"https:\/\/orcid.org\/0009-0000-9346-3327","authenticated-orcid":false,"given":"Wonhark","family":"Park","sequence":"first","affiliation":[{"name":"Department of Intelligence and Information Seoul National University Seoul Korea (the Republic of)"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jaehyun","family":"Lee","sequence":"additional","affiliation":[{"name":"Department of Intelligence and Information Seoul National University Seoul Korea (the Republic of)"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Wonsik","family":"Shin","sequence":"additional","affiliation":[{"name":"Department of Artificial Intelligence Seoul National University Seoul Korea (the Republic of)"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Junhoo","family":"Lee","sequence":"additional","affiliation":[{"name":"Department of Intelligence and Information Seoul National University Seoul Korea (the Republic of)"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Nojun","family":"Kwak","sequence":"additional","affiliation":[{"name":"Department of Intelligence and Information Seoul National University Seoul Korea (the Republic of)"},{"name":"Department of Artificial Intelligence Seoul National University Seoul Korea (the Republic of)"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"265","published-online":{"date-parts":[[2026,2,16]]},"reference":[{"key":"e_1_2_12_2_1","doi-asserted-by":"crossref","unstructured":"R.Rombach A.Blattmann D.Lorenz P.Esser andB.Ommer \u201cHigh\u2010Resolution Image Synthesis With Latent Diffusion Models \u201d inIEEE\/CVF Conference on Computer Vision and Pattern Recognition CVPR 2022(IEEE 2022) 10674\u201310685.","DOI":"10.1109\/CVPR52688.2022.01042"},{"key":"e_1_2_12_3_1","unstructured":"A.Ramesh M.Pavlov G.Goh et\u00a0al. \u201cZero\u2010Shot Text\u2010to\u2010Image Generation \u201d inProceedings of the 38th International Conference on Machine Learning Vol.139 ed.M.MeilaandT.Zhang(PMLR 2021) 8821\u20138831."},{"key":"e_1_2_12_4_1","volume-title":"Advances in Neural Information Processing Systems","author":"Saharia C.","year":"2022"},{"key":"e_1_2_12_5_1","unstructured":"R.Gal Y.Alaluf Y.Atzmon et\u00a0al. \u201cAn Image is Worth One Word: Personalizing Text\u2010to\u2010Image Generation Using Textual Inversion \u201d inThe Eleventh International Conference on Learning Representations ICLR 2023(2023)."},{"key":"e_1_2_12_6_1","doi-asserted-by":"crossref","unstructured":"N.Ruiz Y.Li V.Jampani Y.Pritch M.Rubinstein andK.Aberman \u201cDreambooth: Fine Tuning Text\u2010to\u2010Image Diffusion Models for Subject\u2010Driven Generation \u201d in2023 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)(IEEE 2022) 22500\u201322510.","DOI":"10.1109\/CVPR52729.2023.02155"},{"key":"e_1_2_12_7_1","doi-asserted-by":"crossref","unstructured":"O.Avrahami K.Aberman O.Fried D.Cohen\u2010Or andD.Lischinski \u201cBreak\u2010a\u2010Dcene: Extracting Multiple Concepts from a Single Image \u201d inSIGGRAPH Asia 2023 Conference Papers ser. SA '23 (ACM 2023) 1\u201312.","DOI":"10.1145\/3610548.3618154"},{"key":"e_1_2_12_8_1","unstructured":"Y.Gu X.Wang J. Z.Wu et\u00a0al. \u201cMix\u2010of\u2010Show: Decentralized Low\u2010Rank Adaptation for Multi\u2010Concept Customization of Diffusion Models \u201d inThirty\u2010seventh Conference on Neural Information Processing Systems(Curran Associates Inc. 2023)."},{"key":"e_1_2_12_9_1","doi-asserted-by":"crossref","unstructured":"Y.Zhang M.Yang Q.Zhou andZ.Wang \u201cAttention Calibration for Disentangled Text\u2010to\u2010Image Personalization \u201d inProceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(IEEE 2024).","DOI":"10.1109\/CVPR52733.2024.00456"},{"key":"e_1_2_12_10_1","unstructured":"S.Jang J.Jo K.Lee andS. J.Hwang \u201cIdentity Decoupling for Multi\u2010Subject Personalization of Text\u2010to\u2010Image Models \u201d inThe Thirty\u2010Eighth Annual Conference on Neural Information Processing Systems(Curran Associates Inc. 2024)."},{"key":"e_1_2_12_11_1","unstructured":"G.KwonandJ. C.Ye \u201cTweediemix: Improving Multi\u2010Concept Fusion for Diffusion\u2010Based Image\/Video Generation \u201d inThe Thirteenth International Conference on Learning Representations(IEEE Information Theory Society 2025)."},{"key":"e_1_2_12_12_1","first-page":"57500","article-title":"Customizable Image Synthesis with Multiple Subjects","volume":"36","author":"Liu Z.","year":"2023","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_2_12_13_1","unstructured":"X.Wang S.Fu Q.Huang W.He andH.Jiang \u201cMS\u2010Diffusion: Multi\u2010Subject Zero\u2010Shot Image Personalization with Layout Guidance \u201d inThe Thirteenth International Conference on Learning Representations(IEEE Information Theory Society 2025)."},{"key":"e_1_2_12_14_1","doi-asserted-by":"crossref","unstructured":"G.Ding C.Zhao W.Wang et\u00a0al. \u201cFreecustom: Tuning\u2010Free Customized Image Generation for Multi\u2010Concept Composition \u201d inProceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(IEEE 2024).","DOI":"10.1109\/CVPR52733.2024.00868"},{"key":"e_1_2_12_15_1","doi-asserted-by":"crossref","unstructured":"Z.Kong Y.Zhang T.Yang et\u00a0al. \u201cOmg: Occlusion\u2010Friendly Personalized Multi\u2010Concept Generation in Diffusion Models \u201d inComputer Vision \u2013 ECCV 2024: 18th European Conference(Springer 2024) 253\u2013270.","DOI":"10.1007\/978-3-031-72751-1_15"},{"key":"e_1_2_12_16_1","unstructured":"E. J.Hu Y.Shen P.Wallis et\u00a0al. \u201cLoRA: Low\u2010Rank Adaptation of Large Language Models \u201d inInternational Conference on Learning Representations(IEEE Information Theory Society 2022)."},{"key":"e_1_2_12_17_1","doi-asserted-by":"crossref","unstructured":"N.Kumari B.Zhang R.Zhang E.Shechtman andJ.\u2010Y.Zhu \u201cMulti\u2010Concept Customization of Text\u2010to\u2010Image Diffusion \u201d inProceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)(IEEE 2023).","DOI":"10.1109\/CVPR52729.2023.00192"},{"key":"e_1_2_12_18_1","doi-asserted-by":"crossref","unstructured":"H.Zhang T.Xu andH.Li \u201cStackgan: Text to Photo\u2010Realistic Image Synthesis with Stacked Generative Adversarial Networks \u201d inIEEE International Conference on Computer Vision ICCV 2017(IEEE Computer Society 2017) 5908\u20135916.","DOI":"10.1109\/ICCV.2017.629"},{"key":"e_1_2_12_19_1","doi-asserted-by":"crossref","unstructured":"M.Kang J.\u2010Y.Zhu R.Zhang et\u00a0al. \u201cScaling Up GANs for Text\u2010to\u2010Image Synthesis \u201d inProceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)(IEEE 2023).","DOI":"10.1109\/CVPR52729.2023.00976"},{"key":"e_1_2_12_20_1","doi-asserted-by":"crossref","unstructured":"Z.Li M. R.Min K.Li andC.Xu \u201cStylet2i: Toward Compositional and High\u2010Fidelity Text\u2010to\u2010Image Synthesis \u201d inProceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)(IEEE 2022).","DOI":"10.1109\/CVPR52688.2022.01766"},{"key":"e_1_2_12_21_1","unstructured":"S.Reed Z.Akata X.Yan L.Logeswaran B.Schiele andH.Lee \u201cGenerative Adversarial Text to Image Synthesis \u201d inProceedings of the 33rd International Conference on Machine Learning Vol.48 ed.M. F.BalcanandK. Q.Weinberger(PMLR 2016) 1060\u20131069."},{"key":"e_1_2_12_22_1","doi-asserted-by":"crossref","unstructured":"H.Zhang J. Y.Koh J.Baldridge H.Lee andY.Yang \u201cCross\u2010Modal Contrastive Learning for Text\u2010to\u2010Image Generation \u201d inProceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(IEEE 2021) 833\u2013842.","DOI":"10.1109\/CVPR46437.2021.00089"},{"key":"e_1_2_12_23_1","doi-asserted-by":"crossref","unstructured":"T.Xu P.Zhang Q.Huang et\u00a0al. \u201cAttngan: Fine\u2010Grained Text to Image Generation with Attentional Generative Adversarial Networks \u201d inProceedings of the IEEE Conference on Computer Vision and Pattern Recognition(IEEE 2018) 1316\u20131324.","DOI":"10.1109\/CVPR.2018.00143"},{"key":"e_1_2_12_24_1","unstructured":"J.Ho A.Jain andP.Abbeel \u201cDenoising diffusion probabilistic models \u201d inAdvances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020 NeurIPS 2020 ed.H.Larochelle M.Ranzato R.Hadsell M.Balcan andH.Lin(Curran Associates Inc. 2020)."},{"key":"e_1_2_12_25_1","unstructured":"J.Song C.Meng andS.Ermon \u201cDenoising Diffusion Implicit Models \u201d inInternational Conference on Learning Representations(IEEE Information Theory Society 2021)."},{"key":"e_1_2_12_26_1","doi-asserted-by":"crossref","unstructured":"L.Zhang A.Rao andM.Agrawala \u201cAdding Conditional Control to Text\u2010to\u2010Image Diffusion Models \u201d inIEEE International Conference on Computer Vision (ICCV)(IEEE 2023).","DOI":"10.1109\/ICCV51070.2023.00355"},{"key":"e_1_2_12_27_1","unstructured":"J.Gu S.Zhai Y.Zhang J. M.Susskind andN.Jaitly \u201cMatryoshka Diffusion Models \u201d inThe Twelfth International Conference on Learning Representations(IEEE Information Theory Society 2024)."},{"key":"e_1_2_12_28_1","volume-title":"Advances in Neural Information Processing Systems","author":"Dhariwal P.","year":"2021"},{"key":"e_1_2_12_29_1","unstructured":"J.HoandT.Salimans \u201cClassifier\u2010Free Diffusion Guidance \u201d inNeurIPS 2021 Workshop on Deep Generative Models and Downstream Applications(Curran Associates Inc. 2021)."},{"key":"e_1_2_12_30_1","unstructured":"Y.Balaji S.Nah X.Huang et\u00a0al. \u201ceDiff\u2010I: Text\u2010to\u2010Image Diffusion Models with an Ensemble of Expert Denoisers \u201dCoRR abs\/2211.01324(2022)."},{"key":"e_1_2_12_31_1","first-page":"47:1","article-title":"Cascaded Diffusion Models for High Fidelity Image Generation","volume":"23","author":"Ho J.","year":"2022","journal-title":"Journal of Machine Learning Research"},{"key":"e_1_2_12_32_1","unstructured":"A.Nichol P.Dhariwal A.Ramesh et\u00a0al. \u201cGLIDE: Towards Photorealistic Image Generation and Editing With Text\u2010Guided Diffusion Models \u201d inInternational Conference on Machine Learning ICML 2022 Vol.162 ed.K.Chaudhuri S.Jegelka L.Song C.Szepesv\u00e1ri G.Niu andS.Sabato(PMLR 2022) 16784\u201316804."},{"key":"e_1_2_12_33_1","doi-asserted-by":"crossref","unstructured":"S.Gu D.Chen J.Bao et\u00a0al. \u201cVector Quantized Diffusion Model for Text\u2010to\u2010Image Synthesis \u201darXiv preprint arXiv:2111.14822(2021).","DOI":"10.1109\/CVPR52688.2022.01043"},{"key":"e_1_2_12_34_1","unstructured":"A.Ramesh P.Dhariwal A.Nichol C.Chu andM.Chen \u201cHierarchical Text\u2010Conditional Image Generation with CLIP Latents \u201dCoRR abs\/2204.06125(2022)."},{"key":"e_1_2_12_35_1","unstructured":"P.Esser S.Kulal A.Blattmann et\u00a0al. \u201cScaling Rectified Flow Transformers for High\u2010Resolution Image Synthesis \u201d inForty\u2010First International Conference on Machine Learning ICML 2024(International Machine Learning Society 2024)."},{"key":"e_1_2_12_36_1","doi-asserted-by":"crossref","unstructured":"Y.Tewel R.Gal G.Chechik andY.Atzmon \u201cKey\u2010Locked Rank One Editing for Text\u2010to\u2010Image Personalization \u201d inACM SIGGRAPH 2023 Conference Proceedings SIGGRAPH 2023 ed.E.Brunvand A.Sheffer andM.Wimmer(ACM 2023) 12:1\u201312:11.","DOI":"10.1145\/3588432.3591506"},{"key":"e_1_2_12_37_1","volume-title":"ACM Multimedia 2024","author":"Weili Z.","year":"2024"},{"key":"e_1_2_12_38_1","unstructured":"J.Shentu M.Watson andN. A.Moubayed \u201cTextual Localization: Decomposing Multi\u2010Concept Images for Subject\u2010Driven Text\u2010to\u2010Image Generation \u201d (2024) https:\/\/arxiv.org\/abs\/2402.09966."},{"key":"e_1_2_12_39_1","unstructured":"J.Lu C.Xie andH.Guo \u201cObject\u2010Driven One\u2010Shot Fine\u2010Tuning of Text\u2010to\u2010Image Diffusion With Prototypical Embedding \u201d (2024) https:\/\/arxiv.org\/abs\/2401.15708."},{"key":"e_1_2_12_40_1","doi-asserted-by":"crossref","unstructured":"T.Rahman S.Mahajan H.\u2010Y.Lee J.Ren S.Tulyakov andL.Sigal \u201cVisual Concept\u2010Driven Image Generation With Text\u2010to\u2010Image Diffusion Model \u201d (2024) https:\/\/arxiv.org\/abs\/2402.11487.","DOI":"10.21428\/d82e957c.cc57c54c"},{"key":"e_1_2_12_41_1","unstructured":"J.Huang J. H.Liew H.Yan et\u00a0al. \u201cClassdiffusion: More Aligned Personalization Tuning with Explicit Class Guidance \u201d inThe Thirteenth International Conference on Learning Representations(International Machine Learning Society 2025)."},{"issue":"6","key":"e_1_2_12_42_1","doi-asserted-by":"crossref","first-page":"415","DOI":"10.1007\/s00530-025-02008-9","article-title":"Multi\u2010SBoRA: Regional and Non\u2010Overlapping Weight Updates for Multi\u2010Concept Customization of Diffusion Models","volume":"31","author":"Wu H.","year":"2025","journal-title":"Multimedia Systems"},{"key":"e_1_2_12_43_1","unstructured":"B.Chen M.Zhao H.Sun et\u00a0al. \u201cXverse: Consistent Multi\u2010Subject Control of Identity and Semantic Attributes via DiT Modulation \u201darXiv:2506.21416(2025)."},{"key":"e_1_2_12_44_1","doi-asserted-by":"crossref","unstructured":"L.Han Y.Li H.Zhang P.Milanfar D.Metaxas andF.Yang \u201cSVDiff: Compact Parameter Space for Diffusion Fine\u2010Tuning \u201d inProceedings of the International Conference on Computer Vision (ICCV)(IEEE 2023).","DOI":"10.1109\/ICCV51070.2023.00673"},{"key":"e_1_2_12_45_1","doi-asserted-by":"crossref","unstructured":"R.Po G.Yang K.Aberman andG.Wetzstein \u201cOrthogonal Adaptation for Modular Customization of Diffusion Models \u201darXiv:2312.02432(2024).","DOI":"10.1109\/CVPR52733.2024.00761"},{"key":"e_1_2_12_46_1","unstructured":"M.Zhong Y.Shen S.Wang et\u00a0al. \u201cMulti\u2010Lora Composition for Image Generation \u201darXiv:2402.16843(2024)."},{"key":"e_1_2_12_47_1","volume-title":"Advances in Neural Information Processing Systems","author":"Li B.","year":"2019"},{"key":"e_1_2_12_48_1","unstructured":"CompVis \u201cStable Diffusion \u201d (2022) https:\/\/huggingface.co\/CompVis\/stable\u2010diffusion\u2010v1\u20104."},{"key":"e_1_2_12_49_1","unstructured":"D. P.KingmaandM.Welling \u201cAuto\u2010Encoding Variational Bayes \u201d in2nd International Conference on Learning Representations ICLR 2014 ed.Y.BengioandY.LeCun(IEEE Information Theory Society 2014)."},{"key":"e_1_2_12_50_1","unstructured":"OpenAI \u201cGPT\u20104 Technical Report \u201dCoRRabs\/2303.08774(2023)."},{"key":"e_1_2_12_51_1","unstructured":"T.Ren S.Liu A.Zeng et\u00a0al. \u201cGrounded SAM: Assembling Open\u2010World Models for Diverse Visual Tasks \u201dCoRRabs\/2401.14159(2024)."},{"key":"e_1_2_12_52_1","doi-asserted-by":"crossref","unstructured":"J.Shi W.Xiong Z.Lin andH. J.Jung \u201cInstantbooth: Personalized Text\u2010to\u2010Image Generation Without Test\u2010Time Finetuning \u201d inProceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)(IEEE 2024) 8543\u20138552.","DOI":"10.1109\/CVPR52733.2024.00816"},{"key":"e_1_2_12_53_1","doi-asserted-by":"crossref","unstructured":"J.Liang H.Zeng M.Cui X.Xie andL.Zhang \u201cPPR10K: A Large\u2010Scale Portrait Photo Retouching Dataset with Human\u2010Region Mask and Group\u2010Level Consistency \u201d inProceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)(IEEE 2021) 653\u2013661.","DOI":"10.1109\/CVPR46437.2021.00071"},{"key":"e_1_2_12_54_1","unstructured":"I.LoshchilovandF.Hutter \u201cDecoupled Weight Decay Regularization \u201d inInternational Conference on Learning Representations(IEEE Information Theory Society 2019)."},{"key":"e_1_2_12_55_1","volume-title":"Advances in Neural Information Processing Systems","author":"Lu C.","year":"2022"},{"key":"e_1_2_12_56_1","unstructured":"D.Podell Z.English K.Lacey et\u00a0al. \u201cSDXL: Improving Latent Diffusion Models for High\u2010Resolution Image Synthesis \u201d inThe Twelfth International Conference on Learning Representations ICLR 2024(IEEE Information Theory Society 2024)."},{"key":"e_1_2_12_57_1","doi-asserted-by":"crossref","unstructured":"R.Zhang P.Isola A. A.Efros E.Shechtman andO.Wang \u201cThe Unreasonable Effectiveness of Deep Features as a Perceptual Metric \u201d inProceedings of the IEEE Conference on Computer Vision and Pattern Recognition(IEEE 2018) 586\u2013595.","DOI":"10.1109\/CVPR.2018.00068"},{"key":"e_1_2_12_58_1","unstructured":"A.Radford J. W.Kim C.Hallacy et\u00a0al. \u201cLearning Transferable Visual Models from Natural Language Supervision \u201d inProceedings of the 38th International Conference on Machine Learning ICML 2021 Vol.139 ed.M.MeilaandT.Zhang(PMLR 2021) 8748\u20138763."},{"key":"e_1_2_12_59_1","doi-asserted-by":"crossref","unstructured":"M.Caron H.Touvron I.Misra et\u00a0al. \u201cEmerging Properties in Self\u2010Supervised Vision Transformers \u201d inProceedings of the International Conference on Computer Vision (ICCV)(IEEE 2021).","DOI":"10.1109\/ICCV48922.2021.00951"},{"key":"e_1_2_12_60_1","doi-asserted-by":"crossref","unstructured":"B.Li Z.Lin D.Pathak et\u00a0al. \u201cGenAI\u2010Bench: Evaluating and Improving Compositional Text\u2010to\u2010Visual Generation \u201darXiv:2406.13743(2024).","DOI":"10.1109\/CVPRW63382.2024.00538"},{"key":"e_1_2_12_61_1","first-page":"34892","volume-title":"Advances in Neural Information Processing Systems","author":"Liu H.","year":"2023"},{"key":"e_1_2_12_62_1","unstructured":"C.Raffel N.Shazeer A.Roberts et\u00a0al. \u201cExploring the Limits of Transfer Learning with a Unified Text\u2010to\u2010Text Transformer \u201d (2023) https:\/\/arxiv.org\/abs\/1910.10683."},{"key":"e_1_2_12_63_1","unstructured":"D.Chae N.Park J.Kim andK.Lee \u201cInstructbooth: Instruction\u2010following Personalized Text\u2010to\u2010Image Generation \u201d inICML 2024 Workshop on Foundation Models in the Wild(International Machine Learning Society 2024)."},{"key":"e_1_2_12_64_1","unstructured":"B. F.Labs S.Batifol A.Blattmann et\u00a0al. \u201cFLUX. 1 Kontext: Flow Matching for In\u2010Context Image Generation and Editing in Latent Space \u201darXiv preprint arXiv:2506.15742(2025)."},{"key":"e_1_2_12_65_1","doi-asserted-by":"crossref","unstructured":"W.PeeblesandS.Xie \u201cScalable Diffusion Models with Transformers \u201d inProceedings of the IEEE\/CVF International Conference on Computer Vision(IEEE 2023) 4195\u20134205.","DOI":"10.1109\/ICCV51070.2023.00387"}],"container-title":["IET Image Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/ietresearch.onlinelibrary.wiley.com\/doi\/pdf\/10.1049\/ipr2.70306","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/ietresearch.onlinelibrary.wiley.com\/doi\/full-xml\/10.1049\/ipr2.70306","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/ietresearch.onlinelibrary.wiley.com\/doi\/pdf\/10.1049\/ipr2.70306","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,2,16]],"date-time":"2026-02-16T15:03:49Z","timestamp":1771254229000},"score":1,"resource":{"primary":{"URL":"https:\/\/ietresearch.onlinelibrary.wiley.com\/doi\/10.1049\/ipr2.70306"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,1]]},"references-count":64,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2026,1]]}},"alternative-id":["10.1049\/ipr2.70306"],"URL":"https:\/\/doi.org\/10.1049\/ipr2.70306","archive":["Portico"],"relation":{},"ISSN":["1751-9659","1751-9667"],"issn-type":[{"value":"1751-9659","type":"print"},{"value":"1751-9667","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,1]]},"assertion":[{"value":"2025-10-04","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-02-06","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-02-16","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}],"article-number":"e70306"}}