SyntheticDoc: A Large Synthetic Dataset for Document Unwarping and Illumination Correction

D. Woortmann*, T. Magne*, O. Sorkine-Hornung

* joint first authors

European Conference on Computer Vision (ECCV) 2026 Spotlight

SyntheticDoc Teaser image - Render SyntheticDoc Teaser image - Albedo SyntheticDoc Teaser image - Shading SyntheticDoc Teaser image - Normal SyntheticDoc Teaser image - 3D SyntheticDoc Teaser image - UV Inverse

Rendered image

Albedo

Shading

Normal map

3D coordinates

UV map

A sample from our SyntheticDoc dataset, with all its annotations.

Abstract

Deep learning models have become the standard tool for document rectification and illumination correction, yet their performance is fundamentally bound by their training data. For nearly a decade, the community has heavily relied on Doc3D, a pioneering but increasingly limited document unwarping dataset in terms of scale and quality. To address this bottleneck, we introduce SyntheticDoc, a massive, high-quality dataset designed to push the boundaries of document unwarping. SyntheticDoc is composed of 1,000,000 high-resolution procedurally generated training samples, alongside extensive validation and test sets. Each sample is paired with rich, pixel-perfect annotations, including UV maps, normal maps, albedo and shading. To ensure physical accuracy and photorealism, the paper geometries are generated via a physics-based simulator and rendered using a path tracer. To demonstrate the benefit of our dataset, we train a simple baseline model on SyntheticDoc and report on its performance in comparison to state-of-the-art methods on both document unwarping and illumination correction tasks.

SyntheticDoc in number

300,000

Simulated meshes

>500,000

Documents

1,000,000

Training samples

1024×1440

High resolution

Simulating the meshes

We generate the meshes representing the paper with a physics-based simulator, and design 6 scenarios that mimic how paper deforms in real life, including bending, folding and crumpling.

Rendering the samples

The samples are rendered using a path tracer under various lighting conditions, with several backgrounds, document textures and paper materials used to ensure a diverse dataset.

A sample with all its annotations is presented at the top of this page.

Video overview

Acknowledgments

We thank the anonymous reviewers for their insightful feedback and constructive suggestions. We are also grateful to Danielle Luterbacher for her help in managing the hardware required to create and store a dataset of this size.

Citation

@article{Woortmann:SyntheticDoc:2026,
    title = {{SyntheticDoc}: A Large Synthetic Dataset for Document Unwarping and Illumination Correction},
    author = {Woortmann, Daniel and Magne, Tanguy and Sorkine-Hornung, Olga},
    booktitle={Computer Vision -- ECCV 2026},
    year = {2026},
}