{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,17]],"date-time":"2026-08-17T15:27:38Z","timestamp":1786980458729,"version":"build-2736575974"},"reference-count":48,"publisher":"Springer Science and Business Media LLC","issue":"2","license":[{"start":{"date-parts":[[2025,12,6]],"date-time":"2025-12-06T00:00:00Z","timestamp":1764979200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2025,12,6]],"date-time":"2025-12-06T00:00:00Z","timestamp":1764979200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100004434","name":"Universit\u00e0 degli Studi di Firenze","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100004434","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["IJDAR"],"published-print":{"date-parts":[[2026,6]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>We present a unified dataset for document Question-Answering (QA), which is obtained combining several public datasets related to Document AI and visually rich document understanding (VRDU). Our main contribution is twofold: on the one hand we reformulate existing Document AI tasks, such as Information Extraction (IE), into a Question-Answering task, making it a suitable resource for training and evaluating Large Language Models; on the other hand, we release the OCR of all the documents and include the exact position of the answer to be found in the document image as a bounding box. Using this dataset, we explore the impact of different prompting techniques (that might include bounding box information) on the performance of open-weight models, identifying the most effective approaches for document comprehension.<\/jats:p>","DOI":"10.1007\/s10032-025-00563-5","type":"journal-article","created":{"date-parts":[[2025,12,6]],"date-time":"2025-12-06T13:24:54Z","timestamp":1765027494000},"page":"447-462","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":1,"title":["BoundingDocs: a Unified Dataset for Document Question Answering with Spatial Annotations"],"prefix":"10.1007","volume":"29","author":[{"given":"Simone","family":"Giovannini","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Fabio","family":"Coppini","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Andrea","family":"Gemelli","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Simone","family":"Marinai","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2025,12,6]]},"reference":[{"key":"563_CR1","unstructured":"Mihalcea, R., Tarau, P.: TextRank: Bringing order into text. In: Lin, D., Wu, D. (eds.) Proc. Conf. Empirical Methods in Natural Language Processing, pp. 404\u2013411. ACL, Barcelona, Spain (2004). https:\/\/aclanthology.org\/W04-3252"},{"key":"563_CR2","doi-asserted-by":"publisher","unstructured":"Devlin, J., et\u00a0al.: BERT: pre-training of deep bidirectional transformers for language understanding. In: Proceedings of the 2019 NAACL-HLT, Volume 1 (Long and Short Papers), pp. 4171\u20134186. Association for Computational Linguistics, (2019). https:\/\/doi.org\/10.18653\/V1\/N19-1423","DOI":"10.18653\/V1\/N19-1423"},{"key":"563_CR3","doi-asserted-by":"publisher","unstructured":"OpenAI: GPT-4 technical report. CoRR arXiv: abs\/2303.08774 (2023) https:\/\/doi.org\/10.48550\/ARXIV.2303.08774","DOI":"10.48550\/ARXIV.2303.08774"},{"key":"563_CR4","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1007\/s10032-006-0020-2","volume":"10","author":"C Nawei","year":"2007","unstructured":"Nawei, C., Blostein, D.: A survey of document image classification classifier: problem statement, architecture and performance evaluation. IJDAR 10, 1\u201316 (2007). https:\/\/doi.org\/10.1007\/s10032-006-0020-2","journal-title":"IJDAR"},{"key":"563_CR5","doi-asserted-by":"publisher","unstructured":"Wang, J., et\u00a0al.: Towards robust visual information extraction in real world: New dataset and novel solution. In: AAAI 2021, pp. 2738\u20132745. AAAI Press, (2021). https:\/\/doi.org\/10.1609\/AAAI.V35I4.16378","DOI":"10.1609\/AAAI.V35I4.16378"},{"key":"563_CR6","doi-asserted-by":"publisher","unstructured":"Mathew, M., Karatzas, D., Jawahar, C.V.: Docvqa: A dataset for VQA on document images. In: WACV, pp. 2199\u20132208. IEEE, (2021). https:\/\/doi.org\/10.1109\/WACV48630.2021.00225","DOI":"10.1109\/WACV48630.2021.00225"},{"key":"563_CR7","doi-asserted-by":"publisher","unstructured":"Wang, W., et\u00a0al.: Layout and task aware instruction prompt for zero-shot document image question answering. CoRR (2023) https:\/\/doi.org\/10.48550\/ARXIV.2306.00526","DOI":"10.48550\/ARXIV.2306.00526"},{"key":"563_CR8","doi-asserted-by":"publisher","unstructured":"Lamott, M., et\u00a0al.: Lapdoc: Layout-aware prompting for documents. In: ICDAR 2024. LNCS, vol. 14807, pp. 142\u2013159. Springer, (2024). https:\/\/doi.org\/10.1007\/978-3-031-70546-5_9","DOI":"10.1007\/978-3-031-70546-5_9"},{"key":"563_CR9","doi-asserted-by":"publisher","first-page":"683","DOI":"10.1007\/s10032-024-00461-2","volume":"27","author":"A Gemelli","year":"2024","unstructured":"Gemelli, A., Marinai, S., Pisaneschi, L., Santoni, F.: Datasets and annotations for layout analysis of scientific articles. IJDAR 27, 683\u2013705 (2024). https:\/\/doi.org\/10.1007\/s10032-024-00461-2","journal-title":"IJDAR"},{"key":"563_CR10","doi-asserted-by":"publisher","unstructured":"Wang, Z., et\u00a0al.: VRDU: A benchmark for visually-rich document understanding. In: Proc. 29th ACM SIGKDD. KDD \u201923. ACM, (2023). https:\/\/doi.org\/10.1145\/3580305.3599929","DOI":"10.1145\/3580305.3599929"},{"key":"563_CR11","unstructured":"Project DeepForm: DeepForm. https:\/\/github.com\/project-deepform\/deepform"},{"key":"563_CR12","doi-asserted-by":"publisher","unstructured":"Landeghem, J.V. et\u00a0al.: Document understanding dataset and evaluation (dude). In: Proc. ICCV, pp. 19471\u201319483. IEEE, (2023). https:\/\/doi.org\/10.1109\/ICCV51070.2023.01789","DOI":"10.1109\/ICCV51070.2023.01789"},{"key":"563_CR13","doi-asserted-by":"publisher","unstructured":"Limam, M., et\u00a0al.: FATURA: A multi-layout invoice image dataset for document analysis and understanding, vol. (2023). https:\/\/doi.org\/10.48550\/ARXIV.2311.11856","DOI":"10.48550\/ARXIV.2311.11856"},{"key":"563_CR14","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2023.109834","volume":"144","author":"R Tito","year":"2023","unstructured":"Tito, R., Karatzas, D., Valveny, E.: Hierarchical multimodal transformers for multipage DocVQA. Pattern Recogn. 144, 109834 (2023). https:\/\/doi.org\/10.1016\/j.patcog.2023.109834","journal-title":"Pattern Recogn."},{"key":"563_CR15","doi-asserted-by":"publisher","unstructured":"Jaume, G., et\u00a0al.: Funsd: A dataset for form understanding in noisy scanned documents. In: 2nd Int. Workshop OST@ICDAR, pp. 22\u201325. IEEE, (2019). https:\/\/doi.org\/10.1109\/ICDARW.2019.10029","DOI":"10.1109\/ICDARW.2019.10029"},{"key":"563_CR16","doi-asserted-by":"crossref","unstructured":"Stanis\u0142awek, T., et\u00a0al.: Kleister: Key information extraction datasets involving long documents with complex layouts. In: Proc. ICDAR, vol. 12821, pp. 564\u2013579. Springer, (2021). http:\/\/dx.doi.org\/10.1007\/978-3-030-86549-8_36","DOI":"10.1007\/978-3-030-86549-8_36"},{"key":"563_CR17","doi-asserted-by":"publisher","unstructured":"Huang, Z., et\u00a0al.: ICDAR2019 competition on scanned receipt OCR and information extraction. In: Proc. ICDAR, pp. 1516\u20131520. IEEE, (2019). https:\/\/doi.org\/10.1109\/ICDAR.2019.00244","DOI":"10.1109\/ICDAR.2019.00244"},{"key":"563_CR18","doi-asserted-by":"crossref","unstructured":"Xu, Y., et\u00a0al.: XFUND: A benchmark dataset for multilingual visually rich form understanding. In: ACL (Findings), pp. 3214\u20133224. ACL, (2022). https:\/\/aclanthology.org\/2022.findings-acl.253","DOI":"10.18653\/v1\/2022.findings-acl.253"},{"key":"563_CR19","doi-asserted-by":"publisher","unstructured":"Nassar, A., et\u00a0al.: Tableformer: Table structure understanding with transformers. In: CVPR, pp. 4604\u20134613. IEEE, (2022). https:\/\/doi.org\/10.1109\/CVPR52688.2022.00457","DOI":"10.1109\/CVPR52688.2022.00457"},{"key":"563_CR20","unstructured":"Park, S., et\u00a0al.: Cord: A consolidated receipt dataset for post-ocr parsing. In: Document Intelligence Workshop at Neural Information Processing Systems (2019). https:\/\/github.com\/clovaai\/cord"},{"key":"563_CR21","unstructured":"Universit\u00e1 degli Studi di Trieste: Ghega Dataset. https:\/\/machinelearning.inginf.units.it\/data-and-tools\/ghega-dataset"},{"key":"563_CR22","unstructured":"HuggingFace: Docmatix - A huge dataset for Document Visual Question Answering. https:\/\/huggingface.co\/blog\/docmatix (2024)"},{"key":"563_CR23","doi-asserted-by":"crossref","unstructured":"Zmigrod, R., et al.: \"what is the value of templates?\" rethinking document information extraction datasets for llms. In: Al-Onaizan, Y., Bansal, M., Chen, Y. (eds.) EMNLP 2024, USA, pp. 13162\u201313185. Association for Computational Linguistics, (2024). https:\/\/aclanthology.org\/2024.findings-emnlp.770","DOI":"10.18653\/v1\/2024.findings-emnlp.770"},{"key":"563_CR24","doi-asserted-by":"publisher","unstructured":"Rodriguez, J.A., et\u00a0al.: Bigdocs: An open and permissively-licensed dataset for training multimodal models on document and code tasks. CoRR (2024) https:\/\/doi.org\/10.48550\/ARXIV.2412.04626","DOI":"10.48550\/ARXIV.2412.04626"},{"key":"563_CR25","doi-asserted-by":"publisher","unstructured":"Hu, A., et\u00a0al.: mPLUG-DocOwl 1.5: Unified structure learning for ocr-free document understanding, vol. (2024). https:\/\/doi.org\/10.48550\/ARXIV.2403.12895","DOI":"10.48550\/ARXIV.2403.12895"},{"key":"563_CR26","doi-asserted-by":"publisher","unstructured":"Garncarek, \u0141., et\u00a0al.: LAMBERT: Layout-Aware Language Modeling for Information Extraction, pp. 532\u2013547. Springer, (2021). https:\/\/doi.org\/10.1007\/978-3-030-86549-8_34","DOI":"10.1007\/978-3-030-86549-8_34"},{"key":"563_CR27","unstructured":"Liu, Y., et\u00a0al.: Roberta: A robustly optimized bert pretraining approach, vol. (2019). https:\/\/arxiv.org\/abs\/1907.11692"},{"key":"563_CR28","doi-asserted-by":"publisher","unstructured":"Perot, V., et\u00a0al.: Lmdx: Language model-based document information extraction and localization. In: Proceedings of ACL (Findings), pp. 15140\u201315168. Association for Computational Linguistics, (2024).https:\/\/doi.org\/10.18653\/V1\/2024.FINDINGS-ACL.899","DOI":"10.18653\/V1\/2024.FINDINGS-ACL.899"},{"key":"563_CR29","doi-asserted-by":"publisher","unstructured":"Wang, D., et\u00a0al.: Docllm: A layout-aware generative language model for multimodal document understanding. In: Proceedings of the 62nd ACL, pp. 8529\u20138548. Association for Computational Linguistics, (2024). https:\/\/doi.org\/10.18653\/V1\/2024.ACL-LONG.463","DOI":"10.18653\/V1\/2024.ACL-LONG.463"},{"key":"563_CR30","unstructured":"Numind: NuExtract 1.5 - Multilingual, Infinite Context, Still Small, and Better than GPT-4o! https:\/\/numind.ai\/blog\/nuextract-1-5---multilingual-infinite-context-still-small-and-better-than-gpt-4o (2024)"},{"key":"563_CR31","doi-asserted-by":"publisher","unstructured":"Dodge, J., et\u00a0al.: Documenting large webtext corpora: A case study on the colossal clean crawled corpus. In: Proceedings of the 2021 EMNLP, pp. 1286\u20131305. Association for Computational Linguistics, (2021). https:\/\/doi.org\/10.18653\/V1\/2021.EMNLP-MAIN.98","DOI":"10.18653\/V1\/2021.EMNLP-MAIN.98"},{"key":"563_CR32","unstructured":"Dosovitskiy, A., et\u00a0al.: An image is worth 16x16 words: Transformers for image recognition at scale. In: ICLR. OpenReview.net, (2021). https:\/\/openreview.net\/pdf?id=YicbFdNTTy"},{"key":"563_CR33","doi-asserted-by":"publisher","unstructured":"Kim, G., et\u00a0al.: Ocr-free document understanding transformer. In: ECCV. LNCS, vol. 13688, pp. 498\u2013517. Springer, (2022). https:\/\/doi.org\/10.1007\/978-3-031-19815-1_29","DOI":"10.1007\/978-3-031-19815-1_29"},{"key":"563_CR34","doi-asserted-by":"publisher","unstructured":"Davis, B., et\u00a0al.: End-to-end document recognition and understanding with dessurt. In: ECCV 2022 Workshops, Proceedings, Part IV. LNCS, vol. 13804, pp. 280\u2013296. Springer, (2022).https:\/\/doi.org\/10.1007\/978-3-031-25069-9_19","DOI":"10.1007\/978-3-031-25069-9_19"},{"key":"563_CR35","doi-asserted-by":"publisher","unstructured":"Huang, Y., et\u00a0al.: Layoutlmv3: Pre-training for document ai with unified text and image masking. In: MM \u201922: The 30th ACM International Conference on Multimedia, Lisboa, Portugal, October 10 - 14, 2022, pp. 4083\u20134091. ACM, (2022).https:\/\/doi.org\/10.1145\/3503161.3548112","DOI":"10.1145\/3503161.3548112"},{"key":"563_CR36","doi-asserted-by":"publisher","unstructured":"Mathew, M., et\u00a0al.: InfographicVQA. In: WACV, pp. 2582\u20132591. IEEE, (2022). https:\/\/doi.org\/10.1109\/WACV51458.2022.00264","DOI":"10.1109\/WACV51458.2022.00264"},{"key":"563_CR37","doi-asserted-by":"publisher","unstructured":"Tanaka, R., et\u00a0al.: Slidevqa: A dataset for document visual question answering on multiple images. In: Williams, B., Chen, Y., Neville, J. (eds.) AAAI 2023, pp. 13636\u201313645. AAAI Press, (2023). https:\/\/doi.org\/10.1609\/AAAI.V37I11.26598","DOI":"10.1609\/AAAI.V37I11.26598"},{"key":"563_CR38","doi-asserted-by":"publisher","unstructured":"Tanaka, R., et\u00a0al.: Visualmrc: Machine reading comprehension on documen images. In: AAAI 2021, pp. 13878\u201313888. AAAI Press, (2021). https:\/\/doi.org\/10.1609\/AAAI.V35I15.17635","DOI":"10.1609\/AAAI.V35I15.17635"},{"key":"563_CR39","unstructured":"Amazon Web Services: Amazon Textract. https:\/\/aws.amazon.com\/it\/textract\/"},{"key":"563_CR40","doi-asserted-by":"publisher","unstructured":"Jiang, A.Q., et\u00a0al.: Mistral 7b, vol. (2023). https:\/\/doi.org\/10.48550\/ARXIV.2310.06825","DOI":"10.48550\/ARXIV.2310.06825"},{"key":"563_CR41","unstructured":"Mistral AI: Mistral Large: Our Flagship Model (2024). https:\/\/mistral.ai\/news\/mistral-large\/"},{"key":"563_CR42","unstructured":"Mistral AI: Mixtral of Experts: A High-Quality Sparse Mixture-of-Experts Model (2023). https:\/\/mistral.ai\/news\/mixtral-of-experts\/"},{"key":"563_CR43","doi-asserted-by":"publisher","unstructured":"Peer, D., et\u00a0al.: Anls* \u2013 a universal document processing metric for generative large language models. CoRR (2024) https:\/\/doi.org\/10.48550\/ARXIV.2402.03848","DOI":"10.48550\/ARXIV.2402.03848"},{"key":"563_CR44","doi-asserted-by":"publisher","unstructured":"Touvron, H., et\u00a0al.: Llama: Open and efficient foundation language models, vol. (2023). https:\/\/doi.org\/10.48550\/ARXIV.2302.13971","DOI":"10.48550\/ARXIV.2302.13971"},{"key":"563_CR45","doi-asserted-by":"publisher","unstructured":"Abdin, M., et\u00a0al.: Phi-3 technical report: A highly capable language model locally on your phone, vol. (2024). https:\/\/doi.org\/10.48550\/ARXIV.2404.14219","DOI":"10.48550\/ARXIV.2404.14219"},{"key":"563_CR46","unstructured":"Anthropic: Claude 3.7 Sonnet (2025). https:\/\/www.anthropic.com\/claude\/sonnet"},{"key":"563_CR47","doi-asserted-by":"publisher","unstructured":"Wang, P., et\u00a0al.: Qwen2-vl: Enhancing vision-language model\u2019s perception of the world at any resolution. CoRR arXv: (2024) https:\/\/doi.org\/10.48550\/ARXIV.2409.12191","DOI":"10.48550\/ARXIV.2409.12191"},{"key":"563_CR48","unstructured":"Chen, A., Giovannini, S., Gemelli, A., Coppini, F., Marinai, S.: Towards Reliable and Interpretable Document Question Answering via VLMs (2025). https:\/\/arxiv.org\/abs\/2509.10129"}],"container-title":["International Journal on Document Analysis and Recognition (IJDAR)"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10032-025-00563-5.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10032-025-00563-5","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10032-025-00563-5.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,20]],"date-time":"2026-06-20T07:07:54Z","timestamp":1781939274000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10032-025-00563-5"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,12,6]]},"references-count":48,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2026,6]]}},"alternative-id":["563"],"URL":"https:\/\/doi.org\/10.1007\/s10032-025-00563-5","relation":{},"ISSN":["1433-2833","1433-2825"],"issn-type":[{"value":"1433-2833","type":"print"},{"value":"1433-2825","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,12,6]]},"assertion":[{"value":"15 November 2024","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"23 October 2025","order":2,"name":"revised","label":"Revised","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"25 November 2025","order":3,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"6 December 2025","order":4,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare no competing interests.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}]}}