{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,21]],"date-time":"2026-05-21T14:11:19Z","timestamp":1779372679967,"version":"3.53.1"},"reference-count":62,"publisher":"Association for Computing Machinery (ACM)","issue":"6","license":[{"start":{"date-parts":[[2026,5,21]],"date-time":"2026-05-21T00:00:00Z","timestamp":1779321600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["92370119, 62376113, and 62276258"],"award-info":[{"award-number":["92370119, 62376113, and 62276258"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Jiangsu Science and Technology Programme","award":["BK20251812"],"award-info":[{"award-number":["BK20251812"]}]},{"name":"Open Program of Henan Key Laboratory of Oracle Bone Inscription Information Processing","award":["OIP2024H004"],"award-info":[{"award-number":["OIP2024H004"]}]},{"name":"Foundation of Fujian Key Laboratory of Pattern Recognition and Image Understanding","award":["PRIU25-04"],"award-info":[{"award-number":["PRIU25-04"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2026,6,30]]},"abstract":"<jats:p>\n                    Recognizing oracle bone scripts plays an important role in Chinese archaeology and philology. However, a significant challenge remains because of the scarcity of oracle character images. To overcome this issue, we propose Diff-Oracle, a novel multi-modal conditional diffusion model that generates a diverse range of controllable oracle characters by inputting random combinations of references. Given the challenge of accurately describing oracle character styles using natural language, Diff-Oracle departs from traditional diffusion models that rely primarily on text prompts by introducing a style encoder. This encoder extracts style prompts from existing oracle character images, where style details are converted into a text embedding format via a pre-trained language-vision model. Additionally, given the lack of explicit content information for oracle characters, ensuring that generated characters accurately represent the intended glyphs is challenging. Therefore, we pre-generate pixel-level paired oracle character images (i.e., style and content images) by an image-to-image translation model, providing content information for the generation process. Meanwhile, Diff-Oracle integrates a content encoder designed to capture specific content details from content reference images. Extensive experiments on Oracle-241 and OBC306 datasets demonstrate that Diff-Oracle significantly outperforms existing generative methods in image quality and diversity. Moreover, Diff-Oracle substantially benefits downstream recognition tasks, outperforming all existing state-of-the-art methods by a large margin. In particular, on the challenging OBC306 dataset, Diff-Oracle achieves a 7.70% accuracy gain in the zero-shot setting and reaches 84.62% accuracy for unseen oracle characters, setting a new benchmark for oracle character recognition. The code is available at\n                    <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" ext-link-type=\"uri\" xlink:href=\"https:\/\/github.com\/JJJingLi\/Diff-Oracle\">https:\/\/github.com\/JJJingLi\/Diff-Oracle<\/jats:ext-link>\n                    .\n                  <\/jats:p>","DOI":"10.1145\/3806389","type":"journal-article","created":{"date-parts":[[2026,4,28]],"date-time":"2026-04-28T15:40:47Z","timestamp":1777390847000},"page":"1-24","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Diff-Oracle: Learning Styles and Contents to Augment Realistic Oracle Characters in Diffusion Model"],"prefix":"10.1145","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0009-0000-3411-3140","authenticated-orcid":false,"given":"Jing","family":"Li","sequence":"first","affiliation":[{"name":"College of Science and Technology, Ningbo University, Ningbo, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0918-4606","authenticated-orcid":false,"given":"Qiufeng","family":"Wang","sequence":"additional","affiliation":[{"name":"School of Advanced Technology, Xi\u2019an Jiaotong-Liverpool University, Suzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-3580-9885","authenticated-orcid":false,"given":"Siyuan","family":"Wang","sequence":"additional","affiliation":[{"name":"School of Mathematics and Physics, Xi\u2019an Jiaotong-Liverpool University, Suzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8104-5432","authenticated-orcid":false,"given":"Rui","family":"Zhang","sequence":"additional","affiliation":[{"name":"School of Mathematics and Physics, Xi\u2019an Jiaotong-Liverpool University, Suzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3034-9639","authenticated-orcid":false,"given":"Kaizhu","family":"Huang","sequence":"additional","affiliation":[{"name":"Digital Innovation Research Center, Duke Kunshan University, Suzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3030-1280","authenticated-orcid":false,"given":"Erik","family":"Cambria","sequence":"additional","affiliation":[{"name":"College of Computing and Data Science, Nanyang Technological University, Singapore, Singapore"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,5,21]]},"reference":[{"key":"e_1_3_2_2_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58529-7_43"},{"key":"e_1_3_2_3_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-65414-6_9"},{"key":"e_1_3_2_4_2","first-page":"8780","volume-title":"Proceedings of the Advances in Neural Information Processing Systems (NeurIPS)","author":"Dhariwal Prafulla","year":"2021","unstructured":"Prafulla Dhariwal and Alexander Quinn Nichol. 2021. Diffusion models beat GANs on image synthesis. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), 8780\u20138794."},{"key":"e_1_3_2_5_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v36i1.19938"},{"key":"e_1_3_2_6_2","volume-title":"Proceedings of the 11th International Conference on Learning Representations (ICLR)","author":"Gal Rinon","year":"2023","unstructured":"Rinon Gal, Yuval Alaluf, Yuval Atzmon, Or Patashnik, Amit Haim Bermano, Gal Chechik, and Daniel Cohen-Or. 2023. An image is worth one word: Personalizing text-to-image generation using textual inversion. In Proceedings of the 11th International Conference on Learning Representations (ICLR)."},{"key":"e_1_3_2_7_2","doi-asserted-by":"crossref","unstructured":"Vidit Goel Elia Peruzzo Yifan Jiang Dejia Xu Xingqian Xu Nicu Sebe Trevor Darrell Zhangyang Wang and Humphrey Shi. 2024. PAIR-Diffusion: A Comprehensive Multimodal Object-Level Image Editor. In Proceedings of the 2024 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 8609\u20138618.","DOI":"10.1109\/CVPR52733.2024.00822"},{"issue":"010","key":"e_1_3_2_8_2","first-page":"2001","article-title":"Identification of oracle-bone script fonts based on topological registration (in Chinese)","volume":"44","author":"Gu Shaotong","year":"2016","unstructured":"Shaotong Gu. 2016. Identification of oracle-bone script fonts based on topological registration (in Chinese). Computer & Digital Engineering 44, 010 (2016), 2001\u20132006.","journal-title":"Computer & Digital Engineering"},{"key":"e_1_3_2_9_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.acl-long.831"},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-41679-8_20"},{"key":"e_1_3_2_11_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2015.2500019"},{"key":"e_1_3_2_12_2","first-page":"652","volume-title":"Proceedings of the 15th Asian Conference on Computer Vision (ACCV)","volume":"12627","author":"Han Wenhui","year":"2020","unstructured":"Wenhui Han, Xinlin Ren, Hangyu Lin, Yanwei Fu, and Xiangyang Xue. 2020. Self-supervised learning of Orc-BERT augmentor for recognizing few-shot oracle characters. In Proceedings of the 15th Asian Conference on Computer Vision (ACCV), Vol. 12627, 652\u2013668."},{"key":"e_1_3_2_13_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-024-02137-0"},{"key":"e_1_3_2_14_2","first-page":"6626","volume-title":"Proceedings of the 31st International Conference on Neural Information Processing Systems (NeurIPS)","author":"Heusel Martin","year":"2017","unstructured":"Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. 2017. GANs trained by a two time-scale update rule converge to a local Nash equilibrium. In Proceedings of the 31st International Conference on Neural Information Processing Systems (NeurIPS), 6626\u20136637."},{"key":"e_1_3_2_15_2","doi-asserted-by":"publisher","DOI":"10.5555\/3495724.3496298"},{"key":"e_1_3_2_16_2","unstructured":"Jonathan Ho and Tim Salimans. 2022. Classifier-free diffusion guidance. arXiv:2207.12598. Retrieved from https:\/\/arxiv.org\/abs\/2207.12598"},{"key":"e_1_3_2_17_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.243"},{"key":"e_1_3_2_18_2","doi-asserted-by":"publisher","DOI":"10.1145\/3503161.3548338"},{"key":"e_1_3_2_19_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-06073-2"},{"key":"e_1_3_2_20_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICDAR.2019.00114"},{"key":"e_1_3_2_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.632"},{"key":"e_1_3_2_22_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v33i01.33014015"},{"key":"e_1_3_2_23_2","doi-asserted-by":"publisher","DOI":"10.5555\/3495724.3496739"},{"key":"e_1_3_2_24_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-41685-9_11"},{"key":"e_1_3_2_25_2","doi-asserted-by":"crossref","unstructured":"Jing Li Qiu-Feng Wang Kaizhu Huang Xi Yang Rui Zhang and John Yannis Goulermas. 2023. Towards better long-tailed oracle character recognition with adversarial data augmentation. Pattern Recognition 140 (2023) 109534.","DOI":"10.1016\/j.patcog.2023.109534"},{"key":"e_1_3_2_26_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-86549-8_16"},{"issue":"8","key":"e_1_3_2_27_2","first-page":"112","article-title":"Recognition of inscriptions on bones or tortoise shells based on graph isomorphism (in Chinese)","volume":"47","author":"Li Qingsheng","year":"2011","unstructured":"Qingsheng Li, Xingyu Yang, and Aimin Wang. 2011. Recognition of inscriptions on bones or tortoise shells based on graph isomorphism (in Chinese). Computer Engineering and Applications 47, 8 (2011), 112\u2013114.","journal-title":"Computer Engineering and Applications"},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.02156"},{"key":"e_1_3_2_29_2","doi-asserted-by":"publisher","DOI":"10.5555\/3367243.3367453"},{"key":"e_1_3_2_30_2","doi-asserted-by":"publisher","DOI":"10.1145\/3503161.3547802"},{"key":"e_1_3_2_31_2","first-page":"54","article-title":"Oracle bone inscription recognition based on SVM (in Chinese)","volume":"2","author":"Liu Yongge","year":"2017","unstructured":"Yongge Liu and Guoying Liu. 2017. Oracle bone inscription recognition based on SVM (in Chinese). Journal of Anyang Normal University 2 (2017), 54\u201356.","journal-title":"Journal of Anyang Normal University"},{"key":"e_1_3_2_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICDAR.2017.181"},{"key":"e_1_3_2_33_2","unstructured":"Mehdi Mirza and Simon Osindero. 2014. Conditional generative adversarial nets. arXiv:1411.1784. Retrieved from https:\/\/arxiv.org\/abs\/1411.1784"},{"key":"e_1_3_2_34_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00178"},{"key":"e_1_3_2_35_2","doi-asserted-by":"publisher","DOI":"10.1145\/3686155"},{"key":"e_1_3_2_36_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58545-7_19"},{"key":"e_1_3_2_37_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.00830"},{"key":"e_1_3_2_38_2","first-page":"8748","volume-title":"Proceedings of the International Conference on Machine Learning (ICML)","volume":"139","author":"Radford Alec","year":"2021","unstructured":"Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. In Proceedings of the International Conference on Machine Learning (ICML), Vol. 139, 8748\u20138763."},{"key":"e_1_3_2_39_2","volume-title":"Proceedings of the International Conference on Learning Representations (ICLR Workshop)","author":"Ramachandran Prajit","year":"2018","unstructured":"Prajit Ramachandran, Barret Zoph, and Quoc V. Le. 2018. Searching for activation functions. In Proceedings of the International Conference on Learning Representations (ICLR Workshop)."},{"key":"e_1_3_2_40_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-19784-0_25"},{"key":"e_1_3_2_41_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01042"},{"key":"e_1_3_2_42_2","doi-asserted-by":"publisher","DOI":"10.1145\/3528233.3530757"},{"key":"e_1_3_2_43_2","unstructured":"Joonghyuk Shin Minguk Kang and Jaesik Park. 2023. Fill-Up: Balancing long-tailed data with generative models. arXiv:2306.07200. Retrieved from https:\/\/arxiv.org\/abs\/2306.07200"},{"key":"e_1_3_2_44_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.00852"},{"key":"e_1_3_2_45_2","first-page":"2256","volume-title":"Proceedings of the 32nd International Conference on Machine Learning (ICML)","volume":"37","author":"Sohl-Dickstein Jascha","year":"2015","unstructured":"Jascha Sohl-Dickstein, Eric A. Weiss, Niru Maheswaranathan, and Surya Ganguli. 2015. Deep unsupervised learning using nonequilibrium thermodynamics. In Proceedings of the 32nd International Conference on Machine Learning (ICML), Vol. 37, 2256\u20132265."},{"key":"e_1_3_2_46_2","volume-title":"Proceedings of the 9th International Conference on Learning Representations (ICLR)","author":"Song Jiaming","year":"2021","unstructured":"Jiaming Song, Chenlin Meng, and Stefano Ermon. 2021. Denoising diffusion implicit models. In Proceedings of the 9th International Conference on Learning Representations (ICLR)."},{"key":"e_1_3_2_47_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v31i1.11231"},{"key":"e_1_3_2_48_2","unstructured":"Yuchen Tian. 2017. zi2zi: Master Chinese calligraphy with conditional adversarial networks. Retrieved from https:\/\/github.com\/kaonashi-tyc\/zi2zi"},{"key":"e_1_3_2_49_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.00185"},{"key":"e_1_3_2_50_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2022.3165989"},{"key":"e_1_3_2_51_2","doi-asserted-by":"publisher","DOI":"10.1145\/3576858"},{"key":"e_1_3_2_52_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00509"},{"key":"e_1_3_2_53_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v38i7.28482"},{"key":"e_1_3_2_54_2","volume-title":"Proceedings of the International Conference on Learning Representations (ICLR)","author":"Zhang Hongyi","year":"2018","unstructured":"Hongyi Zhang, Moustapha Ciss\u00e9, Yann N. Dauphin, and David Lopez-Paz. 2018. Mixup: Beyond empirical risk minimization. In Proceedings of the International Conference on Learning Representations (ICLR)."},{"key":"e_1_3_2_55_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.00355"},{"key":"e_1_3_2_56_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00068"},{"key":"e_1_3_2_57_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.00978"},{"issue":"11","key":"e_1_3_2_58_2","first-page":"320","article-title":"SAMControl: Controlling pose and object for image editing with soft attention mask","volume":"21","author":"Zhang Yue","year":"2024","unstructured":"Yue Zhang, Chao Wang, Feifei Fang, Yunzhi Zhuge, Hehe Fan, Xiaojun Chang, Cheng Deng, and Yi Yang. 2024. SAMControl: Controlling pose and object for image editing with soft attention mask. ACM Transactions on Multimedia Computing, Communications, and Applications 21, 11 (2024), 320.","journal-title":"ACM Transactions on Multimedia Computing, Communications, and Applications"},{"key":"e_1_3_2_59_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICDAR.2019.00057"},{"key":"e_1_3_2_60_2","volume-title":"Proceedings of the 36th International Conference on Neural Information Processing Systems (NeurIPS)","author":"Zhao Min","year":"2022","unstructured":"Min Zhao, Fan Bao, Chongxuan Li, and Jun Zhu. 2022. EGSDE: Unpaired image-to-image translation via energy-guided stochastic differential equations. In Proceedings of the 36th International Conference on Neural Information Processing Systems (NeurIPS)."},{"key":"e_1_3_2_61_2","first-page":"37","volume-title":"Proceedings of the 16th Asian Conference on Computer Vision (ACCV)","volume":"13845","author":"Zhao Xinyi","year":"2022","unstructured":"Xinyi Zhao, Siyuan Liu, Yikai Wang, and Yanwei Fu. 2022. FFD augmentor: Towards few-shot oracle character recognition from scratch. In Proceedings of the 16th Asian Conference on Computer Vision (ACCV), Vol. 13845, 37\u201353."},{"key":"e_1_3_2_62_2","first-page":"800","volume-title":"Proceedings of the 16th European Conference on Computer Vision (ECCV)","volume":"12354","author":"Zhao Yihao","year":"2020","unstructured":"Yihao Zhao, Ruihai Wu, and Hao Dong. 2020. Unpaired image-to-image translation using adversarial consistency loss. In Proceedings of the 16th European Conference on Computer Vision (ECCV), Vol. 12354, 800\u2013815."},{"key":"e_1_3_2_63_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.244"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3806389","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,5,21]],"date-time":"2026-05-21T13:32:58Z","timestamp":1779370378000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3806389"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,5,21]]},"references-count":62,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2026,6,30]]}},"alternative-id":["10.1145\/3806389"],"URL":"https:\/\/doi.org\/10.1145\/3806389","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,5,21]]},"assertion":[{"value":"2024-12-18","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-03-22","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-05-21","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}