{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,9]],"date-time":"2026-05-09T00:14:30Z","timestamp":1778285670982,"version":"3.51.4"},"reference-count":49,"publisher":"Institution of Engineering and Technology (IET)","issue":"1","license":[{"start":{"date-parts":[[2025,8,28]],"date-time":"2025-08-28T00:00:00Z","timestamp":1756339200000},"content-version":"vor","delay-in-days":239,"URL":"http:\/\/creativecommons.org\/licenses\/by-nd\/4.0\/"},{"start":{"date-parts":[[2025,1,1]],"date-time":"2025-01-01T00:00:00Z","timestamp":1735689600000},"content-version":"tdm","delay-in-days":0,"URL":"http:\/\/doi.wiley.com\/10.1002\/tdm_license_1.1"}],"content-domain":{"domain":["ietresearch.onlinelibrary.wiley.com"],"crossmark-restriction":true},"short-container-title":["IET Image Processing"],"published-print":{"date-parts":[[2025,1]]},"abstract":"<jats:title>ABSTRACT<\/jats:title>\n                  <jats:p>Generating photo\u2010realistic images from natural language descriptions is a challenging task at the intersection of natural language processing and computer vision. Text\u2010to\u2010image synthesis involves generating visual images in a way which naturally matches the semantic meaning of the input text. Recent diffusion\u2010based models have demonstrated strong performance in image fidelity but are slow in inference and exhibit coarse semantic alignment. To overcome the two problems above and allow images to be more faithful (realistic) to texts and semantics in the wild, we propose a novel hybrid architecture called DM\u2010GAN+ATT+CL (dynamic memory GAN + contrastive learning and attention mechanisms). Our method proceeds in a two\u2010step manner: we first produce low\u2010resolution images based on the DM\u2010GAN model with dual attention modules and then refine the results through a memory\u2010based feature refinement mechanism. Contrastive learning was then utilized on a separate dataset with high resolution image\u2010text pairs to enhance feature discrimination and strengthen semantic consistency. The result is richer semantic relevance, stronger image variation and better visual quality. Extensive experimental results across multiple benchmark datasets\u2014CUB, Oxford\u2010102, MS\u2010COCO, and MM\u2010CelebA\u2010HQ\u2014demonstrate that the proposed DM\u2010GAN+ATT+CL framework consistently outperforms state\u2010of\u2010the\u2010art baselines. Notably, it achieved an R\u2010precision of 95.24, an inception score (IS) of 38.43, and a Fr\u00e9chet inception distance (FID) of 11.30 on the MS\u2010COCO dataset, with similarly strong and consistent performance observed across the other datasets. These findings indicate that our approach substantially enriches the diversity and reality of synthetic images, promising a better future for text\u2010image\u00a0matching.<\/jats:p>","DOI":"10.1049\/ipr2.70185","type":"journal-article","created":{"date-parts":[[2025,8,28]],"date-time":"2025-08-28T13:46:19Z","timestamp":1756388779000},"update-policy":"https:\/\/doi.org\/10.1002\/crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["MemAttn\u2010CL: Unified Memory, Attention, and Contrastive Learning for Enhanced Text\u2010to\u2010Image Generation"],"prefix":"10.1049","volume":"19","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-5498-8339","authenticated-orcid":false,"given":"Md. Ahsan","family":"Habib","sequence":"first","affiliation":[{"name":"Department of Software Engineering Gazipur Digital University Kaliakair Bangladesh"},{"name":"Department of Computer Science and Engineering Mawlana Bhashani Science and Technology University, Santosh Tangail Bangladesh"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Md. Anwar Hussen","family":"Wadud","sequence":"additional","affiliation":[{"name":"Department of Computer Science and Engineering Sunamgonj Science and Technology University Sunamgonj Bangladesh"},{"name":"Department of Computer Science and Engineering Mawlana Bhashani Science and Technology University, Santosh Tangail Bangladesh"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Mohammad Motiur","family":"Rahman","sequence":"additional","affiliation":[{"name":"Department of Computer Science and Engineering Mawlana Bhashani Science and Technology University, Santosh Tangail Bangladesh"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5738-1631","authenticated-orcid":false,"given":"M. F.","family":"Mridha","sequence":"additional","affiliation":[{"name":"Department of Computer Science and Engineering American International University Dhaka Bangaladesh"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"265","published-online":{"date-parts":[[2025,8,28]]},"reference":[{"key":"e_1_2_10_2_1","unstructured":"M.MirzaandS.Osindero \u201cConditional Generative Adversarial Nets \u201darXiv preprint arXiv:1411.1784(2014)."},{"key":"e_1_2_10_3_1","doi-asserted-by":"crossref","unstructured":"H.Zhang T.Xu H.Li et\u00a0al. \u201cStackgan: Text to Photo\u2010Realistic Image Synthesis With Stacked Generative Adversarial Networks \u201d inProceedings of the IEEE International Conference on Computer Vision(IEEE 2017) 5907\u20135915.","DOI":"10.1109\/ICCV.2017.629"},{"key":"e_1_2_10_4_1","doi-asserted-by":"crossref","unstructured":"T.Xu P.Zhang Q.Huang et\u00a0al. \u201cAttngan: Fine\u2010Grained Text to Image Generation With Attentional Generative Adversarial Networks \u201d inProceedings of the IEEE Conference on Computer Vision and Pattern Recognition(IEEE 2018) 1316\u20131324.","DOI":"10.1109\/CVPR.2018.00143"},{"key":"e_1_2_10_5_1","doi-asserted-by":"crossref","unstructured":"M.Zhu P.Pan W.Chen andY.Yang \u201cDm\u2010gan: Dynamic Memory Generative Adversarial Networks for Text\u2010To\u2010Image Synthesis \u201d inProceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(IEEE 2019) 5802\u20135810.","DOI":"10.1109\/CVPR.2019.00595"},{"key":"e_1_2_10_6_1","unstructured":"A.Radford J. W.Kim C.Hallacy et\u00a0al. \u201cLearning Transferable Visual Models From Natural Language Supervision \u201d inInternational Conference on Machine Learning(International Machine Learning Society 2021) 8748\u20138763."},{"key":"e_1_2_10_7_1","unstructured":"K.Masui M.Otani M.Nomura andH.Nakayama \u201cHarnessing the Latent Diffusion Model for Training\u2010Free Image Style Transfer \u201darXiv preprint arXiv:2410.01366(2024)."},{"key":"e_1_2_10_8_1","doi-asserted-by":"crossref","unstructured":"J.Liao Z.Yang L.Li et\u00a0al. \u201cImagegen\u2010cot: Enhancing Text\u2010to\u2010Image In\u2010Context Learning with Chain\u2010of\u2010Thought Reasoning \u201darXiv preprint arXiv:2503.19312(2025).","DOI":"10.1109\/ICCV51701.2025.01599"},{"key":"e_1_2_10_9_1","doi-asserted-by":"crossref","unstructured":"C.Saharia W.Chan S.Saxena et\u00a0al. \u201cPhotorealistic Text\u2010to\u2010Image Diffusion Models with Deep Language Understanding \u201d inAdvances in Neural Information Processing Systems Vol.35(Curran Associates Inc. 2022) 36479\u201336494.","DOI":"10.52202\/068431-2643"},{"key":"e_1_2_10_10_1","unstructured":"J. S.Fischer M.Gui P.Ma N.Stracke S. A.Baumann andB.Ommer \u201cBoosting Latent Diffusion with Flow Matching \u201darXiv preprint arXiv:2312.07360(2023)."},{"key":"e_1_2_10_11_1","doi-asserted-by":"publisher","DOI":"10.3390\/app13085098"},{"key":"e_1_2_10_12_1","doi-asserted-by":"crossref","unstructured":"T.Qiao J.Zhang D.Xu andD.Tao \u201cMirrorgan: Learning Text\u2010to\u2010Image Generation by Redescription \u201d inProceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(IEEE 2019) 1505\u20131514.","DOI":"10.1109\/CVPR.2019.00160"},{"key":"e_1_2_10_13_1","unstructured":"S. E.Reed Z.Akata S.Mohan S.Tenka B.Schiele andH.Lee \u201cLearning What and Where to Draw \u201d inAdvances in Neural Information Processing Systems Vol.29(Curran Associates Inc. 2016)."},{"key":"e_1_2_10_14_1","doi-asserted-by":"crossref","unstructured":"H.Zhang J. Y.Koh J.Baldridge H.Lee andY.Yang \u201cCross\u2010Modal Contrastive Learning for Text\u2010to\u2010Image Generation \u201d inProceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(IEEE 2021) 833\u2013842.","DOI":"10.1109\/CVPR46437.2021.00089"},{"key":"e_1_2_10_15_1","first-page":"280","volume-title":"Advances in Neural Information Processing Systems","author":"Peng S.","year":"2022"},{"key":"e_1_2_10_16_1","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2023.3342866"},{"key":"e_1_2_10_17_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.neunet.2021.01.023"},{"key":"e_1_2_10_18_1","unstructured":"R.MishraandA.Subramanyam \u201cImage Synthesis with Graph Conditioning: Clip\u2010Guided Diffusion Models for Scene Graphs \u201darXiv preprint arXiv:2401.14111(2024)."},{"key":"e_1_2_10_19_1","doi-asserted-by":"crossref","unstructured":"M.Tao H.Tang F.Wu X.\u2010Y.Jing B.\u2010K.Bao andC.Xu \u201cDf\u2010gan: A simple and effective baseline for text\u2010to\u2010image synthesis \u201d inProceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(IEEE 2022) 16515\u201316525.","DOI":"10.1109\/CVPR52688.2022.01602"},{"key":"e_1_2_10_20_1","unstructured":"Y.Ge J.Xu B. N.Zhao N.Joshi L.Itti andV.Vineet \u201cBeyond generation: Harnessing text to image models for object detection and segmentation \u201darXiv preprint arXiv:2309.05956(2023)."},{"key":"e_1_2_10_21_1","unstructured":"N.Siddiqui \u201cComparative Study of Generative Models for Text\u2010to\u2010Image Generation\u201d (Master's thesis University of Windsor 2023)."},{"key":"e_1_2_10_22_1","first-page":"784","article-title":"Advancing Text\u2010to\u2010Image Generation: A Comparative Study of stylegan\u2010t and stable Diffusion 3 under Neutrosophic Sets","volume":"85","author":"Sadek M. G.","year":"2025","journal-title":"Neutrosophic Sets and Systems"},{"key":"e_1_2_10_23_1","doi-asserted-by":"crossref","unstructured":"J.Deng Z.Yang T.Chen W.Zhou andH.Li \u201cTransvg: End\u2010to\u2010End Visual Grounding with Transformers \u201d inProceedings of the IEEE\/CVF International Conference on Computer Vision(IEEE 2021) 1769\u20131779.","DOI":"10.1109\/ICCV48922.2021.00179"},{"key":"e_1_2_10_24_1","doi-asserted-by":"crossref","unstructured":"R.Rombach A.Blattmann D.Lorenz P.Esser andB.Ommer \u201cHigh\u2010Resolution Image Synthesis with Latent Diffusion Models \u201d inProceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(IEEE 2022) 10684\u201310695.","DOI":"10.1109\/CVPR52688.2022.01042"},{"key":"e_1_2_10_25_1","doi-asserted-by":"crossref","unstructured":"S.Koley A. K.Bhunia A.Sain P. N.Chowdhury T.Xiang andY.\u2010Z.Song \u201cText\u2010to\u2010Image Diffusion Models are Great Sketch\u2010Photo MatchMakers \u201d inProceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(IEEE 2024) 16826\u201316837.","DOI":"10.1109\/CVPR52733.2024.01592"},{"key":"e_1_2_10_26_1","unstructured":"Z.Zhan D.Chen J.\u2010P.Mei et\u00a0al. \u201cConditional Image Synthesis with Diffusion Models: A Survey \u201darXiv preprint arXiv:2409.19365(2024)."},{"key":"e_1_2_10_27_1","unstructured":"J.Sohl\u2010Dickstein E.Weiss N.Maheswaranathan andS.Ganguli \u201cDeep Unsupervised Learning Using Nonequilibrium Thermodynamics \u201d inInternational Conference on Machine Learning(International Machine Learning Society 2015) 2256\u20132265."},{"key":"e_1_2_10_28_1","first-page":"6840","volume-title":"Advances in Neural Information Processing Systems","author":"Ho J.","year":"2020"},{"key":"e_1_2_10_29_1","first-page":"8780","volume-title":"Advances in Neural Information Processing Systems","author":"Dhariwal P.","year":"2021"},{"key":"e_1_2_10_30_1","unstructured":"A.Nichol P.Dhariwal A.Ramesh et\u00a0al. \u201cGlide: Towards Photorealistic Image Generation and Editing with Text\u2010Guided Diffusion Models \u201darXiv preprint arXiv:2112.10741(2021)."},{"key":"e_1_2_10_31_1","unstructured":"A.Ramesh P.Dhariwal A.Nichol C.Chu andM.Chen \u201cHierarchical Text\u2010Conditional Image Generation with Clip Latents \u201darXiv preprint arXiv:2204.06125(2022)."},{"key":"e_1_2_10_32_1","first-page":"26565","article-title":"Elucidating the Design Space of Diffusion\u2010Based Generative Models","volume":"35","author":"Karras T.","year":"2022","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_2_10_33_1","volume-title":"Advances in Neural Information Processing Systems","author":"Goodfellow I. J.","year":"2014"},{"key":"e_1_2_10_34_1","doi-asserted-by":"crossref","unstructured":"S.LatifiandS. N.Esfahani \u201cA Survey of State\u2010of\u2010the\u2010Art GAN\u2010Based Approaches to Image Synthesis\u201d2019.","DOI":"10.5121\/csit.2019.90906"},{"key":"e_1_2_10_35_1","unstructured":"M.\u017belaszczykandJ.Ma\u0144dziuk \u201cText\u2010to\u2010Image Cross\u2010Modal Generation: A Systematic Review \u201darXiv preprint arXiv:2401.11631(2024)."},{"key":"e_1_2_10_36_1","doi-asserted-by":"crossref","unstructured":"N.Dufour V.Besnier V.Kalogeiton andD.Picard \u201cDon't Drop Your Samples! Coherence\u2010Aware Training Benefits Conditional Diffusion \u201d inProceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(IEEE 2024) 6264\u20136273.","DOI":"10.1109\/CVPR52733.2024.00599"},{"key":"e_1_2_10_37_1","doi-asserted-by":"crossref","unstructured":"H.Nam J.\u2010W.Ha andJ.Kim \u201cDual Attention Networks for Multimodal Reasoning and Matching \u201d inProceedings of the IEEE Conference on Computer Vision and Pattern Recognition(IEEE 2017) 299\u2013307.","DOI":"10.1109\/CVPR.2017.232"},{"key":"e_1_2_10_38_1","volume-title":"Advances in Neural Information Processing Systems","author":"Salimans T.","year":"2016"},{"key":"e_1_2_10_39_1","unstructured":"A.Ramesh M.Pavlov G.Goh et\u00a0al. \u201cZero\u2010Shot Text\u2010to\u2010Image Generation \u201d inInternational Conference on Machine Learning(International Machine Learning Society 2021) 8821\u20138831."},{"key":"e_1_2_10_40_1","doi-asserted-by":"crossref","unstructured":"W.Li P.Zhang L.Zhang Q.Huang X.He S.Lyu andJ.Gao \u201cObject\u2010Driven Text\u2010to\u2010Image Synthesis via Adversarial Training \u201d inProceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(IEEE 2019) 12174\u201312182.","DOI":"10.1109\/CVPR.2019.01245"},{"key":"e_1_2_10_41_1","doi-asserted-by":"publisher","DOI":"10.14569\/IJACSA.2021.0120124"},{"key":"e_1_2_10_42_1","doi-asserted-by":"crossref","unstructured":"T.\u2010Y.Lin M.Maire S.Belongie et\u00a0al. \u201cMicrosoft coco: Common objects in context \u201d inComputer Vision\u2013ECCV 2014: 13th European Conference Proceedings Part V 13 Vol.8693(Springer 2014) 740\u2013755.","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"e_1_2_10_43_1","doi-asserted-by":"crossref","unstructured":"M.\u2010E.NilsbackandA.Zisserman \u201cAutomated Flower Classification over a Large Number of Classes \u201d in2008 Sixth Indian Conference on Computer Vision Graphics & Image Processing(IEEE 2008) 722\u2013729.","DOI":"10.1109\/ICVGIP.2008.47"},{"key":"e_1_2_10_44_1","doi-asserted-by":"crossref","unstructured":"P.GavaliandJ. S.Banu \u201cBird species identification using deep learning on gpu platform \u201d in2020 International Conference on Emerging Trends in Information Technology and Engineering (ic\u2010ETITE)(IEEE 2020) 1\u20136.","DOI":"10.1109\/ic-ETITE47903.2020.85"},{"key":"e_1_2_10_45_1","unstructured":"J.Li L.Wei Z.Zhan et\u00a0al. \u201cLformer: Text\u2010to\u2010Image Generation with l\u2010Shape Block Parallel Decoding \u201darXiv preprint arXiv:2303.03800(2023)."},{"key":"e_1_2_10_46_1","doi-asserted-by":"crossref","unstructured":"H.Ye X.Yang M.Takac R.Sunderraman andS.Ji \u201cImproving Text\u2010to\u2010Image Synthesis Using Contrastive Learning \u201darXiv preprint arXiv:2107.02423(2021) https:\/\/doi.org\/10.48550\/arXiv.2107.02423.","DOI":"10.5244\/C.35.137"},{"key":"e_1_2_10_47_1","unstructured":"M.Tao H.Tang S.Wu et\u00a0al. \u201cDf\u2010gan: Deep Fusion Generative Adversarial Networks for Text\u2010to\u2010Image Synthesis \u201darXiv preprint arXiv:2008.05865(2020) https:\/\/doi.org\/10.48550\/arXiv.2008.05865."},{"key":"e_1_2_10_48_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11063022\u201010866\u2010x"},{"key":"e_1_2_10_49_1","unstructured":"J.Li X.Liu andL.Zheng \u201cFactor Decomposed Generative Adversarial Networks for Text\u2010to\u2010Image Synthesis \u201darXiv preprint arXiv:2303.13821(2023) https:\/\/doi.org\/10.48550\/arXiv.2303.13821."},{"key":"e_1_2_10_50_1","first-page":"21357","volume-title":"Advances in Neural Information Processing Systems","author":"Kang M.","year":"2020"}],"container-title":["IET Image Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/ietresearch.onlinelibrary.wiley.com\/doi\/pdf\/10.1049\/ipr2.70185","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/ietresearch.onlinelibrary.wiley.com\/doi\/full-xml\/10.1049\/ipr2.70185","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/ietresearch.onlinelibrary.wiley.com\/doi\/pdf\/10.1049\/ipr2.70185","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,5,8]],"date-time":"2026-05-08T23:41:53Z","timestamp":1778283713000},"score":1,"resource":{"primary":{"URL":"https:\/\/ietresearch.onlinelibrary.wiley.com\/doi\/10.1049\/ipr2.70185"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,1]]},"references-count":49,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2025,1]]}},"alternative-id":["10.1049\/ipr2.70185"],"URL":"https:\/\/doi.org\/10.1049\/ipr2.70185","archive":["Portico"],"relation":{},"ISSN":["1751-9659","1751-9667"],"issn-type":[{"value":"1751-9659","type":"print"},{"value":"1751-9667","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,1]]},"assertion":[{"value":"2025-07-02","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-08-09","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-08-28","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}],"article-number":"e70185"}}