{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,20]],"date-time":"2026-06-20T07:50:29Z","timestamp":1781941829072,"version":"3.54.5"},"reference-count":23,"publisher":"Springer Science and Business Media LLC","issue":"2","license":[{"start":{"date-parts":[[2026,3,10]],"date-time":"2026-03-10T00:00:00Z","timestamp":1773100800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2026,3,10]],"date-time":"2026-03-10T00:00:00Z","timestamp":1773100800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100002347","name":"Federal Ministry of Education and Research","doi-asserted-by":"crossref","award":["01IS23064"],"award-info":[{"award-number":["01IS23064"]}],"id":[{"id":"10.13039\/501100002347","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["IJDAR"],"published-print":{"date-parts":[[2026,6]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>\n                    Key Information Extraction (KIE) systems based on Deep Learning achieve strong token-level performance but offer no formal guarantees on prediction reliability, limiting their adoption in business-critical document workflows. In this work, we introduce a post hoc Uncertainty Quantification framework for KIE using Split Conformal Prediction (CP). After fine-tuning multimodal transformer models on a challenging receipt dataset, we reserve a held-out calibration set to derive nonconformity scores and construct entity-level prediction sets that satisfy a user-specified error rate. On unseen receipts, CP achieves tight marginal coverage (98.3% for\n                    <jats:inline-formula>\n                      <jats:alternatives>\n                        <jats:tex-math>$$\\alpha =0.02$$<\/jats:tex-math>\n                        <mml:math xmlns:mml=\"http:\/\/www.w3.org\/1998\/Math\/MathML\">\n                          <mml:mrow>\n                            <mml:mi>\u03b1<\/mml:mi>\n                            <mml:mo>=<\/mml:mo>\n                            <mml:mn>0.02<\/mml:mn>\n                          <\/mml:mrow>\n                        <\/mml:math>\n                      <\/jats:alternatives>\n                    <\/jats:inline-formula>\n                    ), with 70% of predictions being high-confidence singletons. A detailed analysis shows that highly structured fields such as dates and prices yield small, singleton sets with near\u2013perfect reliability, whereas rare or semantically ambiguous fields such as tips or generic keywords produce larger sets and lower coverage. By exposing positional biases and common label confusions that standard F1-scores and document-accuracy metrics overlook, CP reveals critical risk areas for downstream automation. Finally, we demonstrate how calibrated prediction-set sizes can drive risk-aware workflows by automatically processing high-confidence extractions and flagging uncertain cases for human review, thereby enhancing the efficiency, trustworthiness and operational feasibility of real-world document-processing systems.\n                  <\/jats:p>","DOI":"10.1007\/s10032-026-00572-y","type":"journal-article","created":{"date-parts":[[2026,3,10]],"date-time":"2026-03-10T06:09:13Z","timestamp":1773122953000},"page":"551-563","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Beyond Accuracy: Understanding Model Confidence in Key Information Extraction with Conformal Prediction"],"prefix":"10.1007","volume":"29","author":[{"given":"Alexander","family":"Rombach","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Nijat","family":"Mehdiyev","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2026,3,10]]},"reference":[{"key":"572_CR1","doi-asserted-by":"crossref","unstructured":"Angelopoulos, A.N., Bates, S.: A Gentle Introduction to Conformal Prediction and Distribution-Free Uncertainty Quantification. (2022). arXiv: 2107.07511","DOI":"10.1561\/9781638281597"},{"key":"572_CR2","unstructured":"Denk, T.I, Reisswig, C.: BERTgrid: Contextualized Embedding for 2D Document Representation and Understanding. Workshop on Document Intelligence at NeurIPS (2019) arXiv: 1909.04948"},{"issue":"1","key":"572_CR3","doi-asserted-by":"publisher","first-page":"69","DOI":"10.51387\/22-NEJSDS8","volume":"1","author":"N Dey","year":"2023","unstructured":"Dey, N., Ding, J., Ferrell, J., et al.: Conformal prediction for text infilling and part-of-speech prediction. The New England J. Stat. Data Sci. 1(1), 69\u201383 (2023). https:\/\/doi.org\/10.51387\/22-NEJSDS8","journal-title":"The New England J. Stat. Data Sci."},{"key":"572_CR4","unstructured":"Gibbs, I., Candes, E.: Adaptive conformal inference under distribution shift. In: Ranzato, M., Beygelzimer, A., Dauphin, Y., et al. (eds) Advances in Neural Information Processing Systems, vol 34. Curran Associates Inc, pp 1660\u20131672. (2021). https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2021\/file\/0d441de75945e5acbc865406fc9a2559-Paper.pdf"},{"key":"572_CR5","unstructured":"Gibbs, I\/, Cand\u00e8s, E.: Conformal inference for online prediction with arbitrary distribution shifts. J. Mach. Learn. Res. 25(1) (2024)"},{"key":"572_CR6","unstructured":"Giovannotti, P., Gammerman, A.: Transformer-based conformal predictors for paraphrase detection. In: Carlsson, L., Luo, Z., Cherubin, G., et al. (eds) Proceedings of the Tenth Symposium on Conformal and Probabilistic Prediction and Applications, Proceedings of Machine Learning Research, vol 152. PMLR, pp 243\u2013265 (2021), https:\/\/proceedings.mlr.press\/v152\/giovannotti21a.html"},{"key":"572_CR7","first-page":"62","volume-title":"Digitalisierung Von Staat Und Verwaltung","author":"C Houy","year":"2019","unstructured":"Houy, C., Hamberg, M., Fettke, P.: Robotic process automation in public administrations. In: R\u00e4ckers, M., Halsbenning, S., R\u00e4tz, D., et al. (eds.) Digitalisierung Von Staat Und Verwaltung, pp. 62\u201374. Gesellschaft f\u00fcr Informatik e.V, Bonn (2019)"},{"key":"572_CR8","doi-asserted-by":"publisher","unstructured":"Huang, Z., Chen, K., He, J., et al.: ICDAR2019 Competition on Scanned Receipt OCR and Information Extraction. In: 2019 International Conference on Document Analysis and Recognition (ICDAR). IEEE, pp 1516\u20131520 (2019) https:\/\/doi.org\/10.1109\/ICDAR.2019.00244","DOI":"10.1109\/ICDAR.2019.00244"},{"key":"572_CR9","doi-asserted-by":"publisher","unstructured":"Katti, A.R., Reisswig, C., Guder, C., et al.: Chargrid: Towards understanding 2D documents. In: Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, EMNLP 2018. Association for Computational Linguistics, Stroudsburg, PA, USA, pp 4459\u20134469 (2018). https:\/\/doi.org\/10.18653\/v1\/d18-1476","DOI":"10.18653\/v1\/d18-1476"},{"key":"572_CR10","doi-asserted-by":"publisher","unstructured":"Kivim\u00e4ki, J., Lebedev, A., Nurminen, J.K.: Failure Prediction in 2D Document Information Extraction with Calibrated Confidence Scores. In: Proceedings - International Computer Software and Applications Conference, vol 2023-June. IEEE, pp 193\u2013202 (2023). https:\/\/doi.org\/10.1109\/COMPSAC57700.2023.00033","DOI":"10.1109\/COMPSAC57700.2023.00033"},{"key":"572_CR11","doi-asserted-by":"publisher","unstructured":"Liu, X., Gao, F., Zhang, Q., et al.: Graph convolution for multimodal information extraction from visually rich documents. In: NAACL HLT 2019\u20132019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies - Proceedings of the Conference, vol 2. Association for Computational Linguistics, Stroudsburg, PA, USA, pp 32\u201339 (2019). https:\/\/doi.org\/10.18653\/v1\/n19-2005","DOI":"10.18653\/v1\/n19-2005"},{"key":"572_CR12","doi-asserted-by":"publisher","unstructured":"Majumder, B.P., Potti, N., Tata, S., et al.: Representation Learning for Information Extraction from Form-like Documents. In: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. Association for Computational Linguistics, Stroudsburg, PA, USA, pp 6495\u20136504 (2020). https:\/\/doi.org\/10.18653\/v1\/2020.acl-main.580","DOI":"10.18653\/v1\/2020.acl-main.580"},{"key":"572_CR13","unstructured":"Maltoudoglou, L., Paisios, A., Papadopoulos, H.: BERT-based conformal predictor for sentiment analysis. In: Gammerman, A., Vovk, V., Luo, Z., et al. (eds) Proceedings of the Ninth Symposium on Conformal and Probabilistic Prediction and Applications, Proceedings of Machine Learning Research, vol 128. PMLR, pp 269\u2013284 (2020)"},{"key":"572_CR14","doi-asserted-by":"publisher","unstructured":"Nourbakhsh, A., Shah, S., Rose, C.: Towards a new research agenda for multimodal enterprise document understanding: What are we missing? In: Ku LW, Martins A, Srikumar V (eds) Findings of the Association for Computational Linguistics ACL 2024. Association for Computational Linguistics, Stroudsburg, PA, USA, pp 14610\u201314622 (2024). https:\/\/doi.org\/10.18653\/v1\/2024.findings-acl.870","DOI":"10.18653\/v1\/2024.findings-acl.870"},{"key":"572_CR15","doi-asserted-by":"publisher","unstructured":"Peng, Q., Pan, Y., Wang, W., et al.: ERNIE-Layout: Layout Knowledge Enhanced Pre-training for Visually-rich Document Understanding. In: Findings of the Association for Computational Linguistics: EMNLP 2022. Association for Computational Linguistics, Stroudsburg, PA, USA, pp 3744\u20133756 (2022). https:\/\/doi.org\/10.18653\/v1\/2022.findings-emnlp.274","DOI":"10.18653\/v1\/2022.findings-emnlp.274"},{"key":"572_CR16","doi-asserted-by":"publisher","unstructured":"Qian, Y., Santus, E., Jin, Z., et al.: GraphIE: A Graph-Based Framework for Information Extraction. In: NAACL HLT 2019\u20132019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies - Proceedings of the Conference, vol 1. Association for Computational Linguistics, Stroudsburg, PA, USA, pp 751\u201376 (2019) https:\/\/doi.org\/10.18653\/v1\/N19-1082, arXiv:1810.13083","DOI":"10.18653\/v1\/N19-1082"},{"key":"572_CR17","unstructured":"Quach, V., Fisch, A., Schuster, T., et al.: Conformal Language Modeling. In: The Twelfth International Conference on Learning Representations, (2024). https:\/\/openreview.net\/forum?id=pzUhfQ74c5"},{"issue":"2","key":"572_CR18","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3749369","volume":"58","author":"AM Rombach","year":"2025","unstructured":"Rombach, A.M., Fettke, P.: Deep learning based key information extraction from business documents: systematic literature review. ACM Comput. Surv. 58(2), 1\u201337 (2025). https:\/\/doi.org\/10.1145\/3749369","journal-title":"ACM Comput. Surv."},{"key":"572_CR19","doi-asserted-by":"publisher","unstructured":"Rombach, A.M., Lahann, J., Niesen, T., et al.: Utilizing deep learning for field-level information extraction from german real estate tax notices. J. Emerg. Tech. Accounting 22(1), 101\u2013118 (2025). https:\/\/doi.org\/10.2308\/JETA-2023-028","DOI":"10.2308\/JETA-2023-028"},{"key":"572_CR20","doi-asserted-by":"publisher","unstructured":"Sassioui, A., Benouini, R., El Ouargui, Y., et al.: Visually-Rich Document Understanding: Concepts, Taxonomy and Challenges. In: Proceedings - 10th International Conference on Wireless Networks and Mobile Communications, WINCOM 2023. IEEE, pp 1\u20137 (2023). https:\/\/doi.org\/10.1109\/WINCOM59760.2023.10322990","DOI":"10.1109\/WINCOM59760.2023.10322990"},{"key":"572_CR21","unstructured":"Sun, H., Kuang, Z., Yue, X., et al.: Spatial Dual-Modality Graph Reasoning for Key Information Extraction. (2021). arXiv: 2103.14470"},{"key":"572_CR22","doi-asserted-by":"publisher","unstructured":"Xu, Y., Li, M., Cui, L., et al.: LayoutLM: Pre-training of Text and Layout for Document Image Understanding. In: Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. ACM, New York, NY, USA, pp 1192\u20131200 (2020). https:\/\/doi.org\/10.1145\/3394486.3403172, arXiv: 1912.13318","DOI":"10.1145\/3394486.3403172"},{"key":"572_CR23","doi-asserted-by":"publisher","unstructured":"Xu, Y., Xu, Y., Lv, T., et al.: LayoutLMv2: Multi-modal pre-training for visually-rich document understanding. In: ACL-IJCNLP 2021\u201359th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, Proceedings of the Conference. Association for Computational Linguistics, Stroudsburg, PA, USA, pp 2579\u20132591(2021) https:\/\/doi.org\/10.18653\/v1\/2021.acl-long.201arXiv:2012.14740","DOI":"10.18653\/v1\/2021.acl-long.201"}],"container-title":["International Journal on Document Analysis and Recognition (IJDAR)"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10032-026-00572-y.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10032-026-00572-y","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10032-026-00572-y.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,20]],"date-time":"2026-06-20T07:06:48Z","timestamp":1781939208000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10032-026-00572-y"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,3,10]]},"references-count":23,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2026,6]]}},"alternative-id":["572"],"URL":"https:\/\/doi.org\/10.1007\/s10032-026-00572-y","relation":{},"ISSN":["1433-2833","1433-2825"],"issn-type":[{"value":"1433-2833","type":"print"},{"value":"1433-2825","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,3,10]]},"assertion":[{"value":"28 May 2025","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"22 December 2025","order":2,"name":"revised","label":"Revised","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"19 February 2026","order":3,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"10 March 2026","order":4,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare no competing interests.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}]}}