{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,8]],"date-time":"2026-04-08T11:43:47Z","timestamp":1775648627186,"version":"3.50.1"},"reference-count":37,"publisher":"Springer Science and Business Media LLC","issue":"14","license":[{"start":{"date-parts":[[2023,10,11]],"date-time":"2023-10-11T00:00:00Z","timestamp":1696982400000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2023,10,11]],"date-time":"2023-10-11T00:00:00Z","timestamp":1696982400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100001775","name":"University of Technology Sydney","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100001775","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Multimed Tools Appl"],"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Image colorization is a well-known problem in computer vision. However, due to the ill-posed nature of the task, image colorization is inherently challenging. Though several attempts have been made by researchers to make the colorization pipeline automatic, these processes often produce unrealistic results due to a lack of conditioning. In this work, we attempt to integrate textual descriptions as an auxiliary condition, along with the grayscale image that is to be colorized, to improve the fidelity of the colorization process. To the best of our knowledge, this is one of the first attempts to incorporate textual conditioning in the colorization pipeline. To do so, a novel deep network has been proposed that takes two inputs (the grayscale image and the respective encoded text description) and tries to predict the relevant color gamut. As the respective textual descriptions contain color information of the objects present in the scene, the text encoding helps to improve the overall quality of the predicted colors. The proposed model has been evaluated using different metrics like SSIM, PSNR, LPISPS and achieved scores of 0.917, 23.27,0.223, respectively. These quantitative metrics have shown that the proposed method outperforms the SOTA techniques in most of the cases.<\/jats:p>","DOI":"10.1007\/s11042-023-15330-z","type":"journal-article","created":{"date-parts":[[2023,10,24]],"date-time":"2023-10-24T09:04:33Z","timestamp":1698138273000},"page":"41121-41136","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":6,"title":["TIC: text-guided image colorization using conditional generative model"],"prefix":"10.1007","volume":"83","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-3242-3406","authenticated-orcid":false,"given":"Subhankar","family":"Ghosh","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Prasun","family":"Roy","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Saumik","family":"Bhattacharya","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Umapada","family":"Pal","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Michael","family":"Blumenstein","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2023,10,11]]},"reference":[{"key":"15330_CR1","doi-asserted-by":"crossref","unstructured":"Ali A, Zhu Y, Chen Q, Yu J, Cai H (2019) Leveraging spatio-temporal patterns for predicting citywide traffic crowd flows using deep hybrid neural networks. In: 2019 IEEE 25th International Conference on Parallel and Distributed Systems (ICPADS). IEEE, pp 125\u2013132","DOI":"10.1109\/ICPADS47876.2019.00025"},{"key":"15330_CR2","doi-asserted-by":"publisher","first-page":"31401","DOI":"10.1007\/s11042-020-10486-4","volume":"80","author":"A Ali","year":"2021","unstructured":"Ali A, Zhu Y, Zakarya M (2021) A data aggregation based approach to exploit dynamic spatio-temporal correlations for citywide crowd flows prediction in fog computing. Multimed Tools Applic 80:31401\u201331433","journal-title":"Multimed Tools Applic"},{"key":"15330_CR3","doi-asserted-by":"publisher","first-page":"852","DOI":"10.1016\/j.ins.2021.08.042","volume":"577","author":"A Ali","year":"2021","unstructured":"Ali A, Zhu Y, Zakarya M (2021) Exploiting dynamic spatio-temporal correlations for citywide traffic flow prediction using attention based neural networks. Inf Sci 577:852\u2013870. https:\/\/doi.org\/10.1016\/j.ins.2021.08.042, https:\/\/www.sciencedirect.com\/science\/article\/pii\/S0020025521008483","journal-title":"Inf Sci"},{"key":"15330_CR4","doi-asserted-by":"publisher","first-page":"233","DOI":"10.1016\/j.neunet.2021.10.021","volume":"145","author":"A Ali","year":"2022","unstructured":"Ali A, Zhu Y, Zakarya M (2022) Exploiting dynamic spatio-temporal graph convolutional neural networks for citywide traffic flows prediction. Neural Netw 145:233\u2013247. https:\/\/doi.org\/10.1016\/j.neunet.2021.10.021, https:\/\/www.sciencedirect.com\/science\/article\/pii\/S0893608021004123","journal-title":"Neural Netw"},{"key":"15330_CR5","unstructured":"Anwar S, Tahir M, Li C, Mian A, Khan F S, Muzaffar A W (2020) Image colorization: a survey and dataset. arXiv:http:\/\/arxiv.org\/abs\/2008.10774"},{"key":"15330_CR6","doi-asserted-by":"crossref","unstructured":"Bahng H, Yoo S, Cho W, Park D K, Wu Z, Ma X, Choo J (2018) Coloring with words: guiding image colorization through text-based palette generation. In: ECCV","DOI":"10.1007\/978-3-030-01258-8_27"},{"key":"15330_CR7","unstructured":"Bastos R, Wynn W C, Lastra A (2013) Run-time glossy surface self-transfer processing"},{"key":"15330_CR8","doi-asserted-by":"crossref","unstructured":"Caesar H, Uijlings J R R, Ferrari V (2018) Coco-stuff: thing and stuff classes in context. In: 2018 IEEE\/CVF Conference on computer vision and pattern recognition, pp 1209\u20131218","DOI":"10.1109\/CVPR.2018.00132"},{"key":"15330_CR9","doi-asserted-by":"crossref","unstructured":"Carlucci F M, Russo P, Caputo B (2018) (de)2co: deep depth colorization. IEEE Robotics and Automation Letters","DOI":"10.1109\/LRA.2018.2812225"},{"key":"15330_CR10","doi-asserted-by":"crossref","unstructured":"Cheng Z, Yang Q, Sheng B (2015) Deep colorization. In: 2015 IEEE International Conference on Computer Vision (ICCV), pp 415\u2013423","DOI":"10.1109\/ICCV.2015.55"},{"key":"15330_CR11","doi-asserted-by":"crossref","unstructured":"Deng J, Dong W, Socher R, Li L-J, Li K, Fei-Fei L (2009) Imagenet: a large-scale hierarchical image database. In: 2009 IEEE Conference on computer vision and pattern recognition, pp 248\u2013255","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"15330_CR12","doi-asserted-by":"crossref","unstructured":"Deora P, Vasudeva B, Bhattacharya S, Pradhan P M (2020) Structure preserving compressive sensing mri reconstruction using generative adversarial networks. In: 2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), pp 2211\u20132219","DOI":"10.1109\/CVPRW50498.2020.00269"},{"key":"15330_CR13","unstructured":"Goodfellow I, Pouget-Abadie J, Mirza M, Xu B, Warde-Farley D, Ozair S, Courville A, Bengio Y (2014) Generative adversarial nets. In: The Conference on Neural Information Processing Systems (NIPS)"},{"key":"15330_CR14","doi-asserted-by":"crossref","unstructured":"Huang Y-C, Tung Y-S, Chen J-C, Wang S-W, Wu J-L (2005) An adaptive edge detection based colorization algorithm and its applications. In: MULTIMEDIA \u201905","DOI":"10.1145\/1101149.1101223"},{"key":"15330_CR15","doi-asserted-by":"publisher","first-page":"105006","DOI":"10.1016\/j.engappai.2022.105006","volume":"114","author":"S Huang","year":"2022","unstructured":"Huang S, Jin X, Jiang Q, Liu L (2022) Deep learning for image colorization: current and future prospects. Eng Appl Artif Intell 114:105006","journal-title":"Eng Appl Artif Intell"},{"key":"15330_CR16","unstructured":"Iandola F N, Han S, Moskewicz M W, Ashraf K, Dally W J, Keutzer K (2016) SqueezeNet: AlexNet-level accuracy with 50x fewer parameters and <\u20090.5mb model size. arXiv:http:\/\/arxiv.org\/abs\/1602.07360"},{"key":"15330_CR17","unstructured":"Kingma D P, Ba J (2014) Adam: a method for stochastic optimization. arXiv:http:\/\/arxiv.org\/abs\/1412.6980"},{"key":"15330_CR18","unstructured":"Kumar M, Weissenborn D, Kalchbrenner N (2021) Colorization transformer. arXiv:http:\/\/arxiv.org\/abs\/2102.04432"},{"key":"15330_CR19","doi-asserted-by":"crossref","unstructured":"Levin A, Lischinski D, Weiss Y (2004) Colorization using optimization. In: SIGGRAPH 2004","DOI":"10.1145\/1186562.1015780"},{"key":"15330_CR20","doi-asserted-by":"crossref","unstructured":"Lin T-Y, Maire M, Belongie S J, Hays J, Perona P, Ramanan D, Doll\u00e1r P, Zitnick C L (2014) Microsoft coco: common objects in context. In: ECCV","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"15330_CR21","doi-asserted-by":"publisher","first-page":"15808","DOI":"10.1109\/TITS.2022.3145476","volume":"23","author":"F Luo","year":"2022","unstructured":"Luo F, Li Y, Zeng G, Peng P, Wang G, Li Y (2022) Thermal infrared image colorization for nighttime driving scenes with top-down guided attention. IEEE Trans Intell Transp Syst 23:15808\u201315823","journal-title":"IEEE Trans Intell Transp Syst"},{"key":"15330_CR22","unstructured":"Maas A L (2013) Rectifier nonlinearities improve neural network acoustic models"},{"key":"15330_CR23","unstructured":"Mikolov T, Chen K, Corrado G S, Dean J (2013) Efficient estimation of word representations in vector space. In: ICLR"},{"key":"15330_CR24","doi-asserted-by":"crossref","unstructured":"Perazzi F, Pont-Tuset J, McWilliams B, Gool L V, Gross M H, Sorkine-Hornung A (2016) A benchmark dataset and evaluation methodology for video object segmentation. In: 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp 724\u2013732","DOI":"10.1109\/CVPR.2016.85"},{"key":"15330_CR25","doi-asserted-by":"crossref","unstructured":"Simonyan K, Zisserman A (2015) Very deep convolutional networks for large-scale image recognition. In: The International Conference on Learning Representations (ICLR)","DOI":"10.1109\/ICCV.2015.314"},{"key":"15330_CR26","doi-asserted-by":"crossref","unstructured":"Tola E, Lepetit V, Fua P V (2008) A fast local descriptor for dense matching. In: 2008 IEEE Conference on computer vision and pattern recognition, pp 1\u20138","DOI":"10.1109\/CVPR.2008.4587673"},{"key":"15330_CR27","doi-asserted-by":"crossref","unstructured":"Treneska S, Zdravevski E, Pires I, Lameski P, Gievska S (2022) Gan-based image colorization for self-supervised visual feature learning. Sensors (Basel, Switzerland), 22","DOI":"10.3390\/s22041599"},{"key":"15330_CR28","doi-asserted-by":"crossref","unstructured":"Wang P, Patel V M (2018) Generating high quality visible images from sar images using cnns. In: 2018 IEEE Radar Conference (RadarConf18), pp 0570\u20130575","DOI":"10.1109\/RADAR.2018.8378622"},{"key":"15330_CR29","doi-asserted-by":"crossref","unstructured":"Wang Z, Bovik A C, Sheikh H R, Simoncelli E P (2004) Image quality assessment: from error visibility to structural similarity. In: IEEE Transactions on Image Processing (TIP)","DOI":"10.1109\/TIP.2003.819861"},{"key":"15330_CR30","doi-asserted-by":"crossref","unstructured":"Wang X, Yu K, Wu S, Gu J, Liu Y, Dong C, Qiao Y, Loy C C (2018) Esrgan: enhanced super-resolution generative adversarial networks. In: The European Conference on Computer Vision Workshops (ECCVW)","DOI":"10.1007\/978-3-030-11021-5_5"},{"key":"15330_CR31","unstructured":"Welinder P, Branson S, Mita T, Wah C, Schroff F, Belongie S, Perona P (2010) Caltech-UCSD Birds 200. Technical Report CNS-TR-2010-001, California Institute of Technology"},{"key":"15330_CR32","doi-asserted-by":"crossref","unstructured":"Wu Y, Wang X, Li Y, Zhang H, Zhao X, Shan Y (2021) Towards vivid and diverse image colorization with generative color prior. In: 2021 IEEE\/CVF International Conference on Computer Vision (ICCV), pp 14357\u201314366","DOI":"10.1109\/ICCV48922.2021.01411"},{"key":"15330_CR33","doi-asserted-by":"publisher","first-page":"2952","DOI":"10.1002\/int.22726","volume":"37","author":"DNBH Wu","year":"2022","unstructured":"Wu D N B H, Gan J, Zhou J, Wang J, Gao W (2022) Fine-grained semantic ethnic costume high-resolution image colorization with conditional gan. Int J Intell Syst 37:2952\u20132968","journal-title":"Int J Intell Syst"},{"key":"15330_CR34","doi-asserted-by":"publisher","first-page":"1222","DOI":"10.1002\/int.22667","volume":"37","author":"Y Xiao","year":"2022","unstructured":"Xiao Y, Jiang A, Liu C, Wang M (2022) Semantic-aware automatic image colorization via unpaired cycle-consistent self-supervised network. Int J Intell Syst 37:1222\u20131238","journal-title":"Int J Intell Syst"},{"key":"15330_CR35","doi-asserted-by":"crossref","unstructured":"Zhang R, Isola P, Efros A A (2016) Colorful image colorization. In: ECCV","DOI":"10.1007\/978-3-319-46487-9_40"},{"key":"15330_CR36","first-page":"1","volume":"36","author":"R Zhang","year":"2017","unstructured":"Zhang R, Zhu J-Y, Isola P, Geng X, Lin A S, Yu T, Efros A A (2017) Real-time user-guided image colorization with learned deep priors. ACM Transactions on Graphics (TOG) 36:1\u201311","journal-title":"ACM Transactions on Graphics (TOG)"},{"key":"15330_CR37","doi-asserted-by":"crossref","unstructured":"Zhang R, Isola P, Efros A A, Shechtman E, Wang O (2018) The unreasonable effectiveness of deep features as a perceptual metric. In: The IEEE Conference on Computer Vision and Pattern Recognition (CVPR)","DOI":"10.1109\/CVPR.2018.00068"}],"container-title":["Multimedia Tools and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11042-023-15330-z.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11042-023-15330-z\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11042-023-15330-z.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,10,31]],"date-time":"2024-10-31T18:11:18Z","timestamp":1730398278000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11042-023-15330-z"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,10,11]]},"references-count":37,"journal-issue":{"issue":"14","published-online":{"date-parts":[[2024,4]]}},"alternative-id":["15330"],"URL":"https:\/\/doi.org\/10.1007\/s11042-023-15330-z","relation":{},"ISSN":["1573-7721"],"issn-type":[{"value":"1573-7721","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,10,11]]},"assertion":[{"value":"26 July 2022","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"22 September 2022","order":2,"name":"revised","label":"Revised","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"6 April 2023","order":3,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"11 October 2023","order":4,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"Financial interests: The authors declare they have no financial interests.","order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare that they have no conflict of interest","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"<!--Emphasis Type='Bold' removed-->Conflict of Interests"}}]}}