{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,21]],"date-time":"2026-07-21T05:06:23Z","timestamp":1784610383148,"version":"3.55.0"},"publisher-location":"New York, NY, USA","reference-count":36,"publisher":"ACM","license":[{"start":{"date-parts":[[2026,6,16]],"date-time":"2026-06-16T00:00:00Z","timestamp":1781568000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2026,6,16]]},"DOI":"10.1145\/3810987.3815538","type":"proceedings-article","created":{"date-parts":[[2026,7,21]],"date-time":"2026-07-21T04:55:12Z","timestamp":1784609712000},"page":"36-40","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["VIVID: Vision-and-Text Integrated Vectorization for Image Documents"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0009-0007-0105-953X","authenticated-orcid":false,"given":"Zhixin","family":"Yuan","sequence":"first","affiliation":[{"name":"School of Intelligence Science and Technology, Peking University, Beijing, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0004-5236-7469","authenticated-orcid":false,"given":"Xihong","family":"Wu","sequence":"additional","affiliation":[{"name":"School of Intelligence Science and Technology, Peking University, Beijing, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,7,20]]},"reference":[{"key":"e_1_3_3_1_2_2","doi-asserted-by":"publisher","unstructured":"Abdelrahman Abdallah Daniel Eberharter Zoe Pfister and Adam Jatowt. 2024. A survey of recent approaches to form understanding in scanned documents. Artificial Intelligence Review 57 12 (21 Oct 2024) 342. 10.1007\/s10462-024-11000-0","DOI":"10.1007\/s10462-024-11000-0"},{"key":"e_1_3_3_1_3_2","unstructured":"Shuai Bai Yuxuan Cai Ruizhe Chen Keqin Chen Xionghui Chen Zesen Cheng Lianghao Deng Wei Ding Chang Gao Chunjiang Ge Wenbin Ge Zhifang Guo Qidong Huang Jie Huang Fei Huang Binyuan Hui Shutong Jiang Zhaohai Li Mingsheng Li Mei Li Kaixin Li Zicheng Lin Junyang Lin Xuejing Liu Jiawei Liu Chenglong Liu Yang Liu Dayiheng Liu Shixuan Liu Dunjie Lu Ruilin Luo Chenxu Lv Rui Men Lingchen Meng Xuancheng Ren Xingzhang Ren Sibo Song Yuchong Sun Jun Tang Jianhong Tu Jianqiang Wan Peng Wang Pengfei Wang Qiuyue Wang Yuxuan Wang Tianbao Xie Yiheng Xu Haiyang Xu Jin Xu Zhibo Yang Mingkun Yang Jianxin Yang An Yang Bowen Yu Fei Zhang Hang Zhang Xi Zhang Bo Zheng Humen Zhong Jingren Zhou Fan Zhou Jing Zhou Yuanzhi Zhu and Ke Zhu. 2025. Qwen3-VL Technical Report. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2511.21631 (2025)."},{"key":"e_1_3_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.00973"},{"key":"e_1_3_3_1_5_2","unstructured":"Alexandre Carlier Martin Danelljan Alexandre Alahi and Radu Timofte. 2020. Deepsvg: A hierarchical generative network for vector graphics animation. Advances in Neural Information Processing Systems 33 (2020) 16351\u201316361."},{"key":"e_1_3_3_1_6_2","unstructured":"Cheng Cui Ting Sun Suyin Liang Tingquan Gao Zelun Zhang Jiaxuan Liu Xueqing Wang Changda Zhou Hongen Liu Manhui Lin Yue Zhang Yubo Zhang Handong Zheng Jing Zhang Jun Zhang Yi Liu Dianhai Yu and Yanjun Ma. 2025. PaddleOCR-VL: Boosting Multilingual Document Parsing via a 0.9B Ultra-Compact Vision-Language Model. arxiv:https:\/\/arXiv.org\/abs\/2510.14528\u00a0[cs.CV] https:\/\/arxiv.org\/abs\/2510.14528"},{"key":"e_1_3_3_1_7_2","unstructured":"Cheng Cui Ting Sun Manhui Lin Tingquan Gao Yubo Zhang Jiaxuan Liu Xueqing Wang Zelun Zhang Changda Zhou Hongen Liu Yue Zhang Wenyu Lv Kui Huang Yichao Zhang Jing Zhang Jun Zhang Yi Liu Dianhai Yu and Yanjun Ma. 2025. PaddleOCR 3.0 Technical Report. arxiv:https:\/\/arXiv.org\/abs\/2507.05595\u00a0[cs.CV] https:\/\/arxiv.org\/abs\/2507.05595"},{"key":"e_1_3_3_1_8_2","first-page":"4171","volume-title":"Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, volume 1 (long and short papers)","author":"Devlin Jacob","year":"2019","unstructured":"Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, volume 1 (long and short papers). 4171\u20134186."},{"key":"e_1_3_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.emnlp-main.64"},{"key":"e_1_3_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.5555\/2765856"},{"key":"e_1_3_3_1_11_2","doi-asserted-by":"crossref","unstructured":"Martin\u00a0A Fischler and Robert\u00a0C Bolles. 1981. Random sample consensus: a paradigm for model fitting with applications to image analysis and automated cartography. Commun. ACM 24 6 (1981) 381\u2013395.","DOI":"10.1145\/358669.358692"},{"key":"e_1_3_3_1_12_2","unstructured":"Levon Haroutunian Zhuang Li Lucian Galescu Philip Cohen Raj Tumuluri and Gholamreza Haffari. 2023. Reranking for Natural Language Generation from Logical Forms: A Study based on Large Language Models. arxiv:https:\/\/arXiv.org\/abs\/2309.12294\u00a0[cs.CL] https:\/\/arxiv.org\/abs\/2309.12294"},{"key":"e_1_3_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_3_1_14_2","unstructured":"Dan Hendrycks and Kevin Gimpel. 2023. Gaussian Error Linear Units (GELUs). arxiv:https:\/\/arXiv.org\/abs\/1606.08415\u00a0[cs.LG] https:\/\/arxiv.org\/abs\/1606.08415"},{"key":"e_1_3_3_1_15_2","series-title":"Proceedings of Machine Learning Research","first-page":"2790","volume-title":"Proceedings of the 36th International Conference on Machine Learning","volume":"97","author":"Houlsby Neil","year":"2019","unstructured":"Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De\u00a0Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. 2019. Parameter-Efficient Transfer Learning for NLP. In Proceedings of the 36th International Conference on Machine Learning(Proceedings of Machine Learning Research, Vol.\u00a097), Kamalika Chaudhuri and Ruslan Salakhutdinov (Eds.). PMLR, 2790\u20132799. https:\/\/proceedings.mlr.press\/v97\/houlsby19a.html"},{"key":"e_1_3_3_1_16_2","unstructured":"Edward\u00a0J. Hu Yelong Shen Phillip Wallis Zeyuan Allen-Zhu Yuanzhi Li Shean Wang Lu Wang and Weizhu Chen. 2021. LoRA: Low-Rank Adaptation of Large Language Models. arxiv:https:\/\/arXiv.org\/abs\/2106.09685\u00a0[cs.CL] https:\/\/arxiv.org\/abs\/2106.09685"},{"key":"e_1_3_3_1_17_2","unstructured":"Siddhartha Jain Xiaofei Ma Anoop Deoras and Bing Xiang. 2024. Lightweight reranking for language model generations. arxiv:https:\/\/arXiv.org\/abs\/2307.06857\u00a0[cs.AI] https:\/\/arxiv.org\/abs\/2307.06857"},{"key":"e_1_3_3_1_18_2","series-title":"Proceedings of Machine Learning Research","first-page":"18893","volume-title":"Proceedings of the 40th International Conference on Machine Learning","volume":"202","author":"Lee Kenton","year":"2023","unstructured":"Kenton Lee, Mandar Joshi, Iulia\u00a0Raluca Turc, Hexiang Hu, Fangyu Liu, Julian\u00a0Martin Eisenschlos, Urvashi Khandelwal, Peter Shaw, Ming-Wei Chang, and Kristina Toutanova. 2023. Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding. In Proceedings of the 40th International Conference on Machine Learning(Proceedings of Machine Learning Research, Vol.\u00a0202), Andreas Krause, Emma Brunskill, Kyunghyun Cho, Barbara Engelhardt, Sivan Sabato, and Jonathan Scarlett (Eds.). PMLR, 18893\u201318912. https:\/\/proceedings.mlr.press\/v202\/lee23g.html"},{"key":"e_1_3_3_1_19_2","doi-asserted-by":"crossref","unstructured":"Tzu-Mao Li Michal Luk\u00e1\u010d Gharbi Micha\u00ebl and Jonathan Ragan-Kelley. 2020. Differentiable Vector Graphics Rasterization for Editing and Learning. ACM Trans. Graph. (Proc. SIGGRAPH Asia) 39 6 (2020) 193:1\u2013193:15.","DOI":"10.1145\/3414685.3417871"},{"key":"e_1_3_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00802"},{"key":"e_1_3_3_1_21_2","doi-asserted-by":"publisher","unstructured":"Gioacchino Noris Alexander Hornung Robert\u00a0W. Sumner Maryann Simmons and Markus Gross. 2013. Topology-driven vectorization of clean line drawings. ACM Trans. Graph. 32 1 Article 4 (Feb. 2013) 11\u00a0pages. 10.1145\/2421636.2421640","DOI":"10.1145\/2421636.2421640"},{"key":"e_1_3_3_1_22_2","first-page":"7342","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Reddy Pradyumna","year":"2021","unstructured":"Pradyumna Reddy, Michael Gharbi, Michal Lukac, and Niloy\u00a0J Mitra. 2021. Im2vec: Synthesizing vector graphics without vector supervision. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 7342\u20137351."},{"key":"e_1_3_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52734.2025.01508"},{"key":"e_1_3_3_1_24_2","unstructured":"Juan\u00a0A. Rodriguez Haotian Zhang Abhay Puri Aarash Feizi Rishav Pramanik Pascal Wichmann Arnab Mondal Mohammad\u00a0Reza Samsami Rabiul Awal Perouz Taslakian Spandana Gella Sai Rajeswar David Vazquez Christopher Pal and Marco Pedersoli. 2025. Rendering-Aware Reinforcement Learning for Vector Graphics Generation. arxiv:https:\/\/arXiv.org\/abs\/2505.20793\u00a0[cs.CV] https:\/\/arxiv.org\/abs\/2505.20793"},{"key":"e_1_3_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2011.6126544"},{"key":"e_1_3_3_1_26_2","unstructured":"Chris\u00a0Tsang Sanford\u00a0Pun. 2024. VTracer. https:\/\/github.com\/visioncortex\/vtracer."},{"key":"e_1_3_3_1_27_2","unstructured":"Peter Selinger. 2003. Potrace: a polygon-based tracing algorithm."},{"key":"e_1_3_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v37i2.25326"},{"key":"e_1_3_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.00516"},{"key":"e_1_3_3_1_30_2","doi-asserted-by":"publisher","unstructured":"Xingze Tian and Tobias G\u00fcnther. 2024. A Survey of Smooth Vector Graphics: Recent Advances in Repr esentation Creation Rasterization and Image Vectorization. IEEE Transactions on Visualization and Computer Graphics 30 3 (2024) 1652\u20131671. 10.1109\/TVCG.2022.3220575","DOI":"10.1109\/TVCG.2022.3220575"},{"key":"e_1_3_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICDAR.2003.1227650"},{"key":"e_1_3_3_1_32_2","unstructured":"Ioannis Tsochantaridis Thorsten Joachims Thomas Hofmann and Yasemin Altun. 2005. Large Margin Methods for Structured and Interdependent Output Variables. Journal of Machine Learning Research 6 50 (2005) 1453\u20131484. http:\/\/jmlr.org\/papers\/v6\/tsochantaridis05a.html"},{"key":"e_1_3_3_1_33_2","doi-asserted-by":"crossref","unstructured":"Yael Vinker Ehsan Pajouheshgar Jessica\u00a0Y Bo Roman\u00a0Christian Bachmann Amit\u00a0Haim Bermano Daniel Cohen-Or Amir Zamir and Ariel Shamir. 2022. Clipasso: Semantically-aware object sketching. ACM Transactions on Graphics (TOG) 41 4 (2022) 1\u201311.","DOI":"10.1145\/3528223.3530068"},{"key":"e_1_3_3_1_34_2","doi-asserted-by":"crossref","unstructured":"Zhou Wang Alan\u00a0C Bovik Hamid\u00a0R Sheikh and Eero\u00a0P Simoncelli. 2004. Image quality assessment: from error visibility to structural similarity. IEEE transactions on image processing 13 4 (2004) 600\u2013612.","DOI":"10.1109\/TIP.2003.819861"},{"key":"e_1_3_3_1_35_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACSSC.2003.1292216"},{"key":"e_1_3_3_1_36_2","unstructured":"Martin Weber. 2024. AutoTrace. https:\/\/github.com\/autotrace\/autotrace."},{"key":"e_1_3_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.1145\/3394486.3403172"}],"event":{"name":"ICDAR '26: International Conference on Multimedia Retrieval","location":"Amsterdam Netherlands","acronym":"ICDAR '26","sponsor":["SIGMM ACM Special Interest Group on Multimedia"]},"container-title":["Proceedings of the 7th International Workshop on Intelligent Cross-Data Analysis  and Retrieval"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3810987.3815538","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,7,21]],"date-time":"2026-07-21T04:55:57Z","timestamp":1784609757000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3810987.3815538"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6,16]]},"references-count":36,"alternative-id":["10.1145\/3810987.3815538","10.1145\/3810987"],"URL":"https:\/\/doi.org\/10.1145\/3810987.3815538","relation":{},"subject":[],"published":{"date-parts":[[2026,6,16]]},"assertion":[{"value":"2026-07-20","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}