{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,5]],"date-time":"2026-08-05T17:59:07Z","timestamp":1785952747304,"version":"3.56.0"},"reference-count":50,"publisher":"Springer Science and Business Media LLC","issue":"3","license":[{"start":{"date-parts":[[2023,4,29]],"date-time":"2023-04-29T00:00:00Z","timestamp":1682726400000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2023,4,29]],"date-time":"2023-04-29T00:00:00Z","timestamp":1682726400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"name":"FZI Forschungszentrum Informatik"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["IJDAR"],"published-print":{"date-parts":[[2023,9]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Numerous business workflows involve printed forms, such as invoices or receipts, which are often manually digitalized to persistently search or store the data. As hardware scanners are costly and inflexible, smartphones are increasingly used for digitalization. Here, processing algorithms need to deal with prevailing environmental factors, such as shadows or crumples. Current state-of-the-art approaches learn supervised image dewarping models based on pairs of raw images and rectification meshes. The available results show promising predictive accuracies for dewarping, but generated errors still lead to sub-optimal information retrieval. In this paper, we explore the potential of improving dewarping models using additional, structured information in the form of invoice templates. We provide two core contributions: (1) a novel dataset, referred to as Inv3D, comprising synthetic and real-world high-resolution invoice images with structural templates, rectification meshes, and a multiplicity of per-pixel supervision signals and (2) a novel image dewarping algorithm, which extends the state-of-the-art approach GeoTr to leverage structural templates using attention. Our extensive evaluation includes an implementation of DewarpNet and shows that exploiting structured templates can improve the performance for image dewarping. We report superior performance for the proposed algorithm on our new benchmark for all metrics, including an improved local distortion of 26.1 %. We made our new dataset and all code publicly available at<jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" ext-link-type=\"uri\" xlink:href=\"https:\/\/felixhertlein.github.io\/inv3d\">https:\/\/felixhertlein.github.io\/inv3d<\/jats:ext-link>.<\/jats:p>","DOI":"10.1007\/s10032-023-00434-x","type":"journal-article","created":{"date-parts":[[2023,4,29]],"date-time":"2023-04-29T07:02:16Z","timestamp":1682751736000},"page":"175-186","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":15,"title":["Inv3D: a high-resolution 3D invoice dataset for template-guided single-image document unwarping"],"prefix":"10.1007","volume":"26","author":[{"given":"Felix","family":"Hertlein","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Alexander","family":"Naumann","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Patrick","family":"Philipp","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2023,4,29]]},"reference":[{"key":"434_CR1","doi-asserted-by":"crossref","unstructured":"Bandyopadhyay, H., Dasgupta, T., Das, N., et al.: A gated and bifurcated stacked u-net module for document image dewarping. In: 2020 25th International Conference on Pattern Recognition (ICPR), IEEE, pp 10,548\u201310,554 (2021)","DOI":"10.1109\/ICPR48806.2021.9413001"},{"key":"434_CR2","doi-asserted-by":"crossref","unstructured":"Cao, H., Ding, X., Liu, C.: A cylindrical surface model to rectify the bound document image. In: Proceedings Ninth IEEE international conference on computer vision, IEEE, pp 228\u2013233 (2003)","DOI":"10.1109\/ICCV.2003.1238346"},{"key":"434_CR3","unstructured":"Chen, D.: E-commerce data. https:\/\/www.kaggle.com\/carrie1\/ecommerce-data, last retrieved 2022-04-11 (2017)"},{"key":"434_CR4","doi-asserted-by":"crossref","unstructured":"Chua, KB., Zhang, L., Zhang, Y., et al.: A fast and stable approach for restoration of warped document images. In: Eighth International Conference on Document Analysis and Recognition (ICDAR\u201905), IEEE, pp 384\u2013388 (2005)","DOI":"10.1109\/ICDAR.2005.8"},{"key":"434_CR5","doi-asserted-by":"crossref","unstructured":"Cimpoi, M., Maji, S., Kokkinos, I., et al.: Describing textures in the wild. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp 3606\u20133613 (2014)","DOI":"10.1109\/CVPR.2014.461"},{"key":"434_CR6","doi-asserted-by":"crossref","unstructured":"Das, S., Ma, K., Shu, Z., et al.: Dewarpnet: Single-image document unwarping with stacked 3d and 2d regression networks. In: Proceedings of the IEEE\/CVF International Conference on Computer Vision, pp 131\u2013140 (2019)","DOI":"10.1109\/ICCV.2019.00022"},{"key":"434_CR7","unstructured":"Das, S., Sial, HM., Baldrich, R., et al.: Intrinsic decomposition of document images in-the-wild. In: British Machine Vision Conference (BMVC) (2020)"},{"key":"434_CR8","doi-asserted-by":"crossref","unstructured":"Das, S., Singh, KY., Wu, J., et al.: End-to-end piece-wise unwarping of document images. In: Proceedings of the IEEE\/CVF International Conference on Computer Vision, pp 4268\u20134277 (2021)","DOI":"10.1109\/ICCV48922.2021.00423"},{"key":"434_CR9","doi-asserted-by":"crossref","unstructured":"Feng, H., Wang, Y., Zhou, W., et al.: Doctr: Document image transformer for geometric unwarping and illumination correction. In: Proceedings of the 29th ACM International Conference on Multimedia, pp 273\u2013281 (2021a)","DOI":"10.1145\/3474085.3475388"},{"key":"434_CR10","unstructured":"Feng, H., Zhou, W., Deng, J., et al.: Docscanner: Robust document image rectification with progressive learning. arXiv preprint arXiv:2110.14968 (2021b)"},{"key":"434_CR11","doi-asserted-by":"crossref","unstructured":"Feng, H., Zhou, W., Deng, J., et al.: Geometric representation learning for document image rectification. In: European Conference on Computer Vision, Springer, pp 475\u2013492 (2022)","DOI":"10.1007\/978-3-031-19836-6_27"},{"issue":"28","key":"434_CR12","doi-asserted-by":"publisher","first-page":"36009","DOI":"10.1007\/s11042-021-10507-w","volume":"80","author":"A Garai","year":"2021","unstructured":"Garai, A., Biswas, S., Mandal, S., et al.: Dewarping of document images: a semi-cnn based approach. Multimed. Tools Appl. 80(28), 36009\u201336032 (2021)","journal-title":"Multimed. Tools Appl."},{"issue":"6","key":"434_CR13","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3130800.3130891","volume":"36","author":"MA Gardner","year":"2017","unstructured":"Gardner, M.A., Sunkavalli, K., Yumer, E., et al.: Learning to predict indoor illumination from a single image. ACM Trans. Graph. (TOG) 36(6), 1\u201314 (2017)","journal-title":"ACM Trans. Graph. (TOG)"},{"key":"434_CR14","doi-asserted-by":"crossref","unstructured":"Huang, Z., Gu, J., Meng, G., et al.: Text line extraction of curved document images using hybrid metric. In: 2015 3rd IAPR Asian Conference on Pattern Recognition (ACPR), IEEE, pp 251\u2013255 (2015)","DOI":"10.1109\/ACPR.2015.7486504"},{"key":"434_CR15","doi-asserted-by":"crossref","unstructured":"Jiang, X., Long, R., Xue, N., et al.: Revisiting document image dewarping by grid regularization. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pp 4543\u20134552 (2022)","DOI":"10.1109\/CVPR52688.2022.00450"},{"key":"434_CR16","doi-asserted-by":"crossref","unstructured":"Jung, ES., Son, H., Oh, K., et al.: Duet: Detection utilizing enhancement for text in scanned or captured documents. In: 2020 25th International Conference on Pattern Recognition (ICPR), IEEE, pp 5466\u20135473 (2021)","DOI":"10.1109\/ICPR48806.2021.9412928"},{"key":"434_CR17","doi-asserted-by":"crossref","unstructured":"Kil, T., Seo, W., Koo, HI., et al.: Robust document image dewarping method using text-lines and line segments. In: 2017 14Th IAPR international conference on document analysis and recognition (ICDAR), IEEE, pp 865\u2013870 (2017)","DOI":"10.1109\/ICDAR.2017.146"},{"issue":"11","key":"434_CR18","doi-asserted-by":"publisher","first-page":"3600","DOI":"10.1016\/j.patcog.2015.04.026","volume":"48","author":"BS Kim","year":"2015","unstructured":"Kim, B.S., Koo, H.I., Cho, N.I.: Document dewarping via text-line based optimization. Patt. Recogn. 48(11), 3600\u20133614 (2015)","journal-title":"Patt. Recogn."},{"key":"434_CR19","first-page":"84","volume":"25","author":"A Krizhevsky","year":"2012","unstructured":"Krizhevsky, A., Sutskever, I., Hinton, G.E.: Imagenet classification with deep convolutional neural networks. Adv. Neural Inform. Process. Syst. 25, 84 (2012)","journal-title":"Adv. Neural Inform. Process. Syst."},{"key":"434_CR20","unstructured":"Levenshtein, VI., et al.: Binary codes capable of correcting deletions, insertions, and reversals. In: Soviet physics doklady, Soviet Union, pp 707\u2013710 (1966)"},{"issue":"6","key":"434_CR21","first-page":"1","volume":"38","author":"X Li","year":"2019","unstructured":"Li, X., Zhang, B., Liao, J., et al.: Document rectification and illumination correction using a patch-based cnn. ACM Trans. Graph. (TOG) 38(6), 1\u201311 (2019)","journal-title":"ACM Trans. Graph. (TOG)"},{"issue":"4","key":"434_CR22","doi-asserted-by":"publisher","first-page":"591","DOI":"10.1109\/TPAMI.2007.70724","volume":"30","author":"J Liang","year":"2008","unstructured":"Liang, J., DeMenthon, D., Doermann, D.: Geometric rectification of camera-captured document images. IEEE Trans. Patt. Anal. Mach. Intell. 30(4), 591\u2013605 (2008)","journal-title":"IEEE Trans. Patt. Anal. Mach. Intell."},{"key":"434_CR23","doi-asserted-by":"crossref","unstructured":"Lilienblum, E., Michaelis, B.: Book scanner dewarping with weak 3d measurements and a simplified surface model. In: International Conference on Discrete Geometry for Computer Imagery, Springer, pp 529\u2013540 (2008)","DOI":"10.1007\/978-3-540-79126-3_47"},{"issue":"5","key":"434_CR24","doi-asserted-by":"publisher","first-page":"978","DOI":"10.1109\/TPAMI.2010.147","volume":"33","author":"C Liu","year":"2010","unstructured":"Liu, C., Yuen, J., Torralba, A.: Sift flow: dense correspondence across scenes and its applications. IEEE Trans. Patt. Anal. Mach. Intell. 33(5), 978\u2013994 (2010)","journal-title":"IEEE Trans. Patt. Anal. Mach. Intell."},{"key":"434_CR25","unstructured":"Loshchilov, I., Hutter, F.: Fixing weight decay regularization in adam. https:\/\/openreview.net\/forum?id=rk6qdGgCZ, last retrieved 2022-04-11 (2018)"},{"key":"434_CR26","doi-asserted-by":"crossref","unstructured":"Lu, S., Tan, CL .: Document flattening through grid modeling and regularization. In: 18th International Conference on Pattern Recognition (ICPR\u201906), IEEE, pp 971\u2013974 (2006a)","DOI":"10.1109\/ICPR.2006.458"},{"key":"434_CR27","doi-asserted-by":"crossref","unstructured":"Lu, S., Tan, CL.: The restoration of camera documents through image segmentation. In: Document Analysis Systems. p 484\u2013495 (2006b)","DOI":"10.1007\/11669487_43"},{"issue":"8","key":"434_CR28","doi-asserted-by":"publisher","first-page":"837","DOI":"10.1016\/j.imavis.2006.02.008","volume":"24","author":"S Lu","year":"2006","unstructured":"Lu, S., Chen, B.M., Ko, C.C.: A partition approach for the restoration of camera images of planar and curled document. Image Vis. Comput. 24(8), 837\u2013848 (2006)","journal-title":"Image Vis. Comput."},{"key":"434_CR29","doi-asserted-by":"crossref","unstructured":"Ma, K., Shu, Z., Bai, X., et al.: DocUNet: Document Image Unwarping via a Stacked U-Net. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp 4700\u20134709 (2018)","DOI":"10.1109\/CVPR.2018.00494"},{"key":"434_CR30","doi-asserted-by":"crossref","unstructured":"Ma, K., Das, S., Shu, Z., et al.: Learning from documents in the wild to improve document unwarping. In: ACM SIGGRAPH 2022 Conference Proceedings, pp 1\u20139 (2022)","DOI":"10.1145\/3528233.3530756"},{"key":"434_CR31","doi-asserted-by":"crossref","unstructured":"Markovitz, A., Lavi, I., Perel, O., et al.: Can you read me now? content aware rectification using angle supervision. In: European Conference on Computer Vision, Springer, pp 208\u2013223 (2020)","DOI":"10.1007\/978-3-030-58610-2_13"},{"issue":"107","key":"434_CR32","first-page":"404","volume":"106","author":"X Qin","year":"2020","unstructured":"Qin, X., Zhang, Z., Huang, C., et al.: U2-net: going deeper with nested u-structure for salient object detection. Patt. Recognit. 106(107), 404 (2020)","journal-title":"Patt. Recognit."},{"key":"434_CR33","doi-asserted-by":"crossref","unstructured":"Ramanna, VKB., Bukhari, SS., Dengel, A.: Document image dewarping using deep learning. In: ICPRAM, pp 524\u2013531 (2019)","DOI":"10.5220\/0007368405240531"},{"key":"434_CR34","unstructured":"Sage, A., Agustsson, E., Timofte, R., et al.: Lld - large logo dataset - version 0.1. https:\/\/data.vision.ee.ethz.ch\/cvl\/lld, last retrieved 2022-04-11 (2017)"},{"key":"434_CR35","unstructured":"Shafait, F., Breuel, T.M.: Document image dewarping contest. In: 2nd Int. Workshop on Camera-Based Document Analysis and Recognition, Curitiba, Brazil, pp. 181\u2013188 (2007)"},{"key":"434_CR36","doi-asserted-by":"crossref","unstructured":"Simon, G., Tabbone, S.: Generic document image dewarping by probabilistic discretization of vanishing points. In: 2020 25th International Conference on Pattern Recognition (ICPR), IEEE, pp 2344\u20132351 (2021)","DOI":"10.1109\/ICPR48806.2021.9412649"},{"key":"434_CR37","doi-asserted-by":"crossref","unstructured":"Smith, LN., Topin, N.: Super-convergence: Very fast training of neural networks using large learning rates. In: Artificial intelligence and machine learning for multi-domain operations applications, International Society for Optics and Photonics, p 1100612 (2019)","DOI":"10.1117\/12.2520589"},{"key":"434_CR38","doi-asserted-by":"crossref","unstructured":"Smith, R.: An overview of the tesseract ocr engine. In: Ninth international conference on document analysis and recognition (ICDAR 2007), IEEE, pp 629\u2013633 (2007)","DOI":"10.1109\/ICDAR.2007.4376991"},{"key":"434_CR39","doi-asserted-by":"crossref","unstructured":"Tian, Y., Narasimhan, SG.: Rectification and 3d reconstruction of curved document images. In: CVPR 2011, IEEE, pp 377\u2013384 (2011)","DOI":"10.1109\/CVPR.2011.5995540"},{"key":"434_CR40","doi-asserted-by":"crossref","unstructured":"Ulges, A., Lampert, CH., Breuel, T.: Document capture using stereo vision. In: Proceedings of the 2004 ACM symposium on Document engineering, pp 198\u2013200 (2004)","DOI":"10.1145\/1030397.1030434"},{"key":"434_CR41","doi-asserted-by":"crossref","unstructured":"Wang, Y., Zhou, W., Lu, Z., et al.: Udoc-gan: Unpaired document illumination correction with background light prior. In: Proceedings of the 30th ACM International Conference on Multimedia, pp 5074\u20135082 (2022)","DOI":"10.1145\/3503161.3547916"},{"key":"434_CR42","doi-asserted-by":"publisher","unstructured":"Wang, Z., Simoncelli, E., Bovik, A.: Multiscale structural similarity for image quality assessment. In: The Thrity-Seventh Asilomar Conference on Signals, Systems Computers, 2003, pp 1398\u20131402 Vol.2, (2003) https:\/\/doi.org\/10.1109\/ACSSC.2003.1292216","DOI":"10.1109\/ACSSC.2003.1292216"},{"key":"434_CR43","doi-asserted-by":"publisher","first-page":"131","DOI":"10.1007\/978-3-030-57058-3_10","volume-title":"International Workshop on Document Analysis Systems","author":"GW Xie","year":"2020","unstructured":"Xie, G.W., Yin, F., Zhang, X.Y., et al.: Dewarping document image by displacement flow estimation with fully convolutional network. In: International Workshop on Document Analysis Systems, pp. 131\u2013144. Springer, London (2020)"},{"key":"434_CR44","doi-asserted-by":"crossref","unstructured":"Xie, GW., Yin, F., Zhang, XY., et al.: Document dewarping with control points. In: International Conference on Document Analysis and Recognition, Springer, pp 466\u2013480 (2021)","DOI":"10.1007\/978-3-030-86549-8_30"},{"key":"434_CR45","doi-asserted-by":"crossref","unstructured":"Xie, Q., Luong, MT., Hovy, E., et al.: Self-training with noisy student improves imagenet classification. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp 10,687\u201310,698 (2020b)","DOI":"10.1109\/CVPR42600.2020.01070"},{"key":"434_CR46","doi-asserted-by":"crossref","unstructured":"Xue, C., Tian, Z., Zhan, F., et al.: Fourier document restoration for robust document dewarping and recognition. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pp 4573\u20134582 (2022)","DOI":"10.1109\/CVPR52688.2022.00453"},{"key":"434_CR47","doi-asserted-by":"crossref","unstructured":"Yamashita, A., Kawarago, A., Kaneko, T., et al.: Shape reconstruction and image restoration for non-flat surfaces of documents with a stereo vision system. In: Proceedings of the 17th International Conference on Pattern Recognition, 2004. ICPR 2004., IEEE, pp 482\u2013485 (2004)","DOI":"10.1109\/ICPR.2004.1334171"},{"issue":"2","key":"434_CR48","doi-asserted-by":"publisher","first-page":"505","DOI":"10.1109\/TPAMI.2017.2675980","volume":"40","author":"S You","year":"2017","unstructured":"You, S., Matsushita, Y., Sinha, S., et al.: Multiview rectification of folded documents. IEEE Trans. Patt. Anal. Mach. Intell. 40(2), 505\u2013511 (2017)","journal-title":"IEEE Trans. Patt. Anal. Mach. Intell."},{"key":"434_CR49","unstructured":"Zhang, J., Luo, C., Jin, L., et al.: Marior: Margin removal and iterative content rectification for document dewarping in the wild. arXiv preprint arXiv:2207.11515"},{"key":"434_CR50","doi-asserted-by":"crossref","unstructured":"Zhang, R., Isola, P., Efros, AA., et al.: The unreasonable effectiveness of deep features as a perceptual metric. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 586\u2013595 (2018)","DOI":"10.1109\/CVPR.2018.00068"}],"container-title":["International Journal on Document Analysis and Recognition (IJDAR)"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10032-023-00434-x.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10032-023-00434-x\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10032-023-00434-x.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,10,19]],"date-time":"2024-10-19T13:22:01Z","timestamp":1729344121000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10032-023-00434-x"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,4,29]]},"references-count":50,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2023,9]]}},"alternative-id":["434"],"URL":"https:\/\/doi.org\/10.1007\/s10032-023-00434-x","relation":{},"ISSN":["1433-2833","1433-2825"],"issn-type":[{"value":"1433-2833","type":"print"},{"value":"1433-2825","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,4,29]]},"assertion":[{"value":"15 November 2022","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"13 February 2023","order":2,"name":"revised","label":"Revised","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"12 April 2023","order":3,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"29 April 2023","order":4,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare no competing interests.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}}]}}