{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,12,18]],"date-time":"2025-12-18T14:29:19Z","timestamp":1766068159245,"version":"3.41.2"},"reference-count":42,"publisher":"Institute of Electronics, Information and Communications Engineers (IEICE)","issue":"8","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["IEICE Trans. Inf. &amp; Syst."],"published-print":{"date-parts":[[2025,8,1]]},"DOI":"10.1587\/transinf.2024edp7201","type":"journal-article","created":{"date-parts":[[2025,1,30]],"date-time":"2025-01-30T17:13:23Z","timestamp":1738257203000},"page":"967-976","source":"Crossref","is-referenced-by-count":1,"title":["Handwritten Character Image Generation for Effective Data Augmentation"],"prefix":"10.1587","volume":"E108.D","author":[{"given":"Chee Siang","family":"LEOW","sequence":"first","affiliation":[{"name":"University of Yamanashi"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Tomoki","family":"KITAGAWA","sequence":"additional","affiliation":[{"name":"SRE Holdings Corporation"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hideaki","family":"YAJIMA","sequence":"additional","affiliation":[{"name":"University of Yamanashi"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hiromitsu","family":"NISHIZAKI","sequence":"additional","affiliation":[{"name":"University of Yamanashi"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"532","reference":[{"key":"1","doi-asserted-by":"publisher","unstructured":"[1] U.-V. Marti and H. Bunke, \u201cThe IAM-database: An english sentence database for off-line handwriting recognition,\u201d International Journal on Document Analysis and Recognition, vol.5, pp.39-46, 2002. 10.1007\/s100320200071","DOI":"10.1007\/s100320200071"},{"key":"2","doi-asserted-by":"crossref","unstructured":"[2] C.-L. Liu, F. Yin, D.-H. Wang, and Q.-F. Wang, \u201cCASIA online and offline Chinese handwriting databases,\u201d 2011 International Conference on Document Analysis and Recognition, pp.37-41, 2011. 10.1109\/icdar.2011.17","DOI":"10.1109\/ICDAR.2011.17"},{"key":"3","doi-asserted-by":"crossref","unstructured":"[3] Y. Xu, T. Lv, L. Cui, G. Wang, Y. Lu, D. Florencio, C. Zhang, and F. Wei, \u201cXFUND: A benchmark dataset for multilingual visually rich form understanding,\u201d Findings of the Association for Computational Linguistics: ACL 2022, ed. S. Muresan, P. Nakov, and A. Villavicencio, pp.3214-3224, 2022. 10.18653\/v1\/2022.findings-acl.253","DOI":"10.18653\/v1\/2022.findings-acl.253"},{"key":"4","unstructured":"[4] \u201cETL character database,\u201d http:\/\/etlcdb.db.aist.go.jp\/?lang=ja, Referred on 2\/8\/2024."},{"key":"5","unstructured":"[5] T. Wang, D.J. Wu, A. Coates, and A.Y. Ng, \u201cEnd-to-end text recognition with convolutional neural networks,\u201d Proc. 21st International Conference on Pattern Recognition (ICPR2012), pp.3304-3308, 2012."},{"key":"6","doi-asserted-by":"crossref","unstructured":"[6] S.F. Rashid, M.-P. Schambach, J. Rottland, and S. von der N\u00fcll, \u201cLow resolution arabic recognition with multidimensional recurrent neural networks,\u201d Proc. 4th International Workshop on Multilingual OCR, Article No.6, 2013. 10.1145\/2505377.2505385","DOI":"10.1145\/2505377.2505385"},{"key":"7","unstructured":"[7] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A.N. Gomez, \u0141. Kaiser, and I. Polosukhin, \u201cAttention is all you need,\u201d Proc. Advances in Neural Information Processing Systems (NeurIPS), pp.5998-6008, 2017."},{"key":"8","unstructured":"[8] A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, \u201cAn image is worth 16x16 words: Transformers for image recognition at scale,\u201d arXiv preprint, arXiv:2010.11929, 2020. 10.48550\/arXiv.2010.11929"},{"key":"9","doi-asserted-by":"publisher","unstructured":"[9] M. Li, T. Lv, J. Chen, L. Cui, Y. Lu, D. Florencio, C. Zhang, Z. Li, and F. Wei, \u201cTrOCR: Transformer-based optical character recognition with pre-trained models,\u201d Proc. AAAI Conference on Artificial Intelligence, vol.37, no.11, pp.13094-13102, 2023. 10.1609\/aaai.v37i11.26538","DOI":"10.1609\/aaai.v37i11.26538"},{"key":"10","unstructured":"[10] J. Ho, A. Jain, and P. Abbeel, \u201cDenoising diffusion probabilistic models,\u201d Proc. Advances in Neural Information Processing Systems (NeurIPS), vol.33, pp.6840-6851, 2020."},{"key":"11","unstructured":"[11] P. Dhariwal and A. Nichol, \u201cDiffusion models beat GANs on image synthesis,\u201d Proc. Advances in Neural Information Processing Systems (NeurIPS), vol.34, pp.8780-8794, 2021."},{"key":"12","doi-asserted-by":"crossref","unstructured":"[12] X. Huang and S. Belongie, \u201cArbitrary style transfer in real-time with adaptive instance normalization,\u201d Proc. IEEE International Conference on Computer Vision (ICCV), pp.1510-1519, 2017. 10.1109\/iccv.2017.167","DOI":"10.1109\/ICCV.2017.167"},{"key":"13","unstructured":"[13] A. Radford, J.W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever, \u201cLearning transferable visual models from natural language supervision,\u201d International Conference on Machine Learning, pp.8748-8763, 2021."},{"key":"14","doi-asserted-by":"crossref","unstructured":"[14] Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner, \u201cGradient-based learning applied to document recognition,\u201d Proc. IEEE, vol.86, no.11, pp.2278-2324, 1998. 10.1109\/5.726791","DOI":"10.1109\/5.726791"},{"key":"15","doi-asserted-by":"crossref","unstructured":"[15] G. Cohen, S. Afshar, J. Tapson, and A. Van Schaik, \u201cEMNIST: Extending MNIST to handwritten letters,\u201d Proc. 2017 International Joint Conference on Neural Networks (IJCNN), pp.2921-2926, 2017. 10.1109\/ijcnn.2017.7966217","DOI":"10.1109\/IJCNN.2017.7966217"},{"key":"16","unstructured":"[16] D.E. Rumelhart, G.E. Hinton, and R.J. Williams, \u201cLearning internal representations by error propagation,\u201d Parallel Distributed Processing: Explorations in the Microstructure of Cognition, vol.1: Foundations of Research, pp.318-362, 1986. 10.7551\/mitpress\/4943.003.0128"},{"key":"17","unstructured":"[17] D.P. Kingma and M. Welling, \u201cAuto-encoding variational Bayes,\u201d Proc. 2nd International Conference on Learning Representations (ICLR), pp.1-14, 2014."},{"key":"18","unstructured":"[18] I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, \u201cGenerative adversarial nets,\u201d Proc. 27th International Conference on Neural Information Processing Systems, vol.2, pp.2672-2680, 2014."},{"key":"19","doi-asserted-by":"crossref","unstructured":"[19] R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, \u201cHigh-resolution image synthesis with latent diffusion models,\u201d Proc. IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pp.10674-10685, 2022. 10.1109\/cvpr52688.2022.01042","DOI":"10.1109\/CVPR52688.2022.01042"},{"key":"20","doi-asserted-by":"crossref","unstructured":"[20] L. Zhang, A. Rao, and M. Agrawala, \u201cAdding conditional control to text-to-image diffusion models,\u201d Proc. IEEE\/CVF International Conference on Computer Vision, pp.3813-3824, 2023. 10.1109\/iccv51070.2023.00355","DOI":"10.1109\/ICCV51070.2023.00355"},{"key":"21","doi-asserted-by":"crossref","unstructured":"[21] B. Chang, Q. Zhang, S. Pan, and L. Meng, \u201cGenerating handwritten Chinese characters using CycleGAN,\u201d Proc. 2018 IEEE Winter Conference on Applications of Computer Vision (WACV), pp.199-207, 2018. 10.1109\/wacv.2018.00028","DOI":"10.1109\/WACV.2018.00028"},{"key":"22","doi-asserted-by":"crossref","unstructured":"[22] J.-Y. Zhu, T. Park, P. Isola, and A.A. Efros, \u201cUnpaired image-to-image translation using cycle-consistent adversarial networks,\u201d Proc. 2017 IEEE International Conference on Computer Vision (ICCV), pp.2242-2251, 2017. 10.1109\/iccv.2017.244","DOI":"10.1109\/ICCV.2017.244"},{"key":"23","unstructured":"[23] W. Kong and B. Xu, \u201cHandwritten Chinese character generation via conditional neural generative models,\u201d Proc. 31st Conference on Neural Information Processing Systems (NeurIPS), pp.4-7, 2017."},{"key":"24","doi-asserted-by":"crossref","unstructured":"[24] S. Kullback and R.A. Leibler, \u201cOn information and sufficiency,\u201d The Annals of Mathematical Statistics, vol.22, no.1, pp.79-86, 1951. 10.1214\/aoms\/1177729694","DOI":"10.1214\/aoms\/1177729694"},{"key":"25","unstructured":"[25] J. Ma, M. Zhao, C. Chen, R. Wang, D. Niu, H. Lu, and X. Lin, \u201cGlyphDraw: Learning to draw Chinese characters in image synthesis models coherently,\u201d arXiv preprint, arXiv:2303.17870, 2023. %2010.48550\/arXiv.2303.17870"},{"key":"26","unstructured":"[26] J. Chen, Y. Huang, T. Lv, L. Cui, Q. Chen, and F. Wei, \u201cTextDiffuser: Diffusion models as text painters,\u201d arXiv preprint, arXiv:2305.10855, 2023. 10.48550\/arXiv.2305.10855"},{"key":"27","unstructured":"[27] Y. Yang, D. Gui, Y. Yuan, H. Ding, H. Hu, and K. Chen, \u201cGlyphControl: Glyph conditional control for visual text generation,\u201d arXiv preprint, arXiv:2305.18259, 2023. 10.48550\/arXiv.2305.18259"},{"key":"28","unstructured":"[28] Y. Tuo, W. Xiang, J.Y. He, Y. Geng, and X. Xie, \u201cAnyText: Multilingual visual text generation and editing,\u201d arXiv preprint, arXiv:2311.03054, 2023. 10.48550\/arXiv.2311.03054"},{"key":"29","doi-asserted-by":"crossref","unstructured":"[29] Y. Zhu, Z. Li, T. Wang, M. He, and C. Yao, \u201cConditional text image generation with diffusion models,\u201d Proc. IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp.14235-14244, 2023. 10.1109\/cvpr52729.2023.01368","DOI":"10.1109\/CVPR52729.2023.01368"},{"key":"30","doi-asserted-by":"crossref","unstructured":"[30] C.S. Leow, T. Kitagawa, H. Yajima, and H. Nishizaki, \u201cData augmentation with automatically generated images for character classifier model training,\u201d Proc. IEEE 12th Global Conference on Consumer Electronics (GCCE), pp.845-849, 2023. 10.1109\/gcce59613.2023.10315447","DOI":"10.1109\/GCCE59613.2023.10315447"},{"key":"31","unstructured":"[31] T. Kitagawa, C.S. Leow, and H. Nishizaki, \u201cHandwritten character generation using Y-autoencoder for character recognition model training,\u201d Proc. 13th Language Resources and Evaluation Conference, pp.7344-7351, 2022."},{"key":"32","doi-asserted-by":"publisher","unstructured":"[32] M. Patacchiola, P. Fox-Roberts, and E. Rosten, \u201cY-Autoencoders: Disentangling latent representations via sequential encoding,\u201d Pattern Recognition Letters, vol.140, pp.59-65, 2020. 10.1016\/j.patrec.2020.09.025","DOI":"10.1016\/j.patrec.2020.09.025"},{"key":"33","doi-asserted-by":"crossref","unstructured":"[33] R. Zhang, P. Isola, A.A. Efros, E. Shechtman, and O. Wang, \u201cThe unreasonable effectiveness of deep features as a perceptual metric,\u201d Proc. IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp.586-595, 2018. 10.1109\/cvpr.2018.00068","DOI":"10.1109\/CVPR.2018.00068"},{"key":"34","unstructured":"[34] K. Simonyan and A. Zisserman, \u201cVery deep convolutional networks for large-scale image recognition,\u201d Proc. International Conference on Learning Representations (ICLR), pp.1-14, 2015."},{"key":"35","doi-asserted-by":"crossref","unstructured":"[35] O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A.C. Berg, and L. Fei-Fei, \u201cImageNet large scale visual recognition challenge,\u201d International Journal of Computer Vision, vol.115, pp.211-252, 2015. 10.1007\/s11263-015-0816-y","DOI":"10.1007\/s11263-015-0816-y"},{"key":"36","doi-asserted-by":"publisher","unstructured":"[36] N. Otsu, \u201cA threshold selection method from gray-level histograms,\u201d IEEE Trans. Syst. Man Cybern., vol.9, no.1, pp.62-66, 1979. 10.1109\/tsmc.1979.4310076","DOI":"10.1109\/TSMC.1979.4310076"},{"key":"37","doi-asserted-by":"publisher","unstructured":"[37] A. Buslaev, V.I. Iglovikov, E. Khvedchenya, A. Parinov, M. Druzhinin, and A.A. Kalinin, \u201cAlbumentations: Fast and flexible image augmentations,\u201d Information, vol.11, no.2, 125, 2020. 10.3390\/info11020125","DOI":"10.3390\/info11020125"},{"key":"38","doi-asserted-by":"crossref","unstructured":"[38] K. He, X. Zhang, S. Ren, and J. Sun, \u201cDeep residual learning for image recognition,\u201d Proc. IEEE Conference on Computer Vision and Pattern Recognition, pp.770-778, 2016. 10.1109\/cvpr.2016.90","DOI":"10.1109\/CVPR.2016.90"},{"key":"39","unstructured":"[39] J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, \u201cBert: Pre-training of deep bidirectional transformers for language understanding,\u201d Proc. 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, vol.1 (Long and Short Papers), pp.4171-4186, 2019. 10.18653\/v1\/n19-1423"},{"key":"40","doi-asserted-by":"crossref","unstructured":"[40] O. Ronneberger, P. Fischer, and T. Brox, \u201cU-Net: Convolutional networks for biomedical image segmentation,\u201d Medical Image Computing and Computer-Assisted Intervention - MICCAI 2015, Lecture Notes in Computer Science, vol.9351, pp.234-241, Springer, Cham, 2015. 10.1007\/978-3-319-24574-4_28","DOI":"10.1007\/978-3-319-24574-4_28"},{"key":"41","doi-asserted-by":"crossref","unstructured":"[41] P. Bergmann, S. L\u00f6we, M. Fauser, D. Sattlegger, and C. Steger, \u201cImproving unsupervised defect segmentation by applying structural similarity to autoencoders,\u201d Proc. 14th International Joint Conference on Computer Vision, Imaging and Computer Graphics Theory and Applications, pp.372-380, 2019. 10.5220\/0007364503720380","DOI":"10.5220\/0007364503720380"},{"key":"42","unstructured":"[42] B. Fuglede and F. Topsoe, \u201cJensen-Shannon divergence and Hilbert space embedding,\u201d Proc. International Symposium on Information Theory 2004, 31, 2004. 10.1109\/isit.2004.1365067"}],"container-title":["IEICE Transactions on Information and Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.jstage.jst.go.jp\/article\/transinf\/E108.D\/8\/E108.D_2024EDP7201\/_pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,8,2]],"date-time":"2025-08-02T03:29:19Z","timestamp":1754105359000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.jstage.jst.go.jp\/article\/transinf\/E108.D\/8\/E108.D_2024EDP7201\/_article"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,8,1]]},"references-count":42,"journal-issue":{"issue":"8","published-print":{"date-parts":[[2025]]}},"URL":"https:\/\/doi.org\/10.1587\/transinf.2024edp7201","relation":{},"ISSN":["0916-8532","1745-1361"],"issn-type":[{"type":"print","value":"0916-8532"},{"type":"electronic","value":"1745-1361"}],"subject":[],"published":{"date-parts":[[2025,8,1]]},"article-number":"2024EDP7201"}}