DARE: A large-scale handwritten DAte REcognition system
Published in International Journal on Document Analysis and Recognition (IJDAR), 2026
Recommended citation: Dahl, Christian M., Torben S. D. Johansen, Emil N. Sørensen, Christian E. Westermann, and Simon F. Wittrock (2026). “DARE: A large-scale handwritten DAte REcognition system”. In: International Journal on Document Analysis and Recognition (IJDAR). doi: 10.1007/s10032-026-00587-5. https://link.springer.com/article/10.1007/s10032-026-00587-5
Authors: The paper is written by Christian M. Dahl, Torben S. D. Johansen, Emil N. Sørensen, Christian E. Westermann, and Simon F. Wittrock.
Download: You can access the paper here or the earlier arXiv version here.
Code and Data: You can find the code for the project here and the database here.
Abstract: Handwritten text recognition for historical documents is an important task, but it remains challenging due to insufficient training data combined with wide variability in writing styles and degradation of historical documents. In the context of recognizing and extracting handwritten dates, we propose a model based on the EfficientNetV2 architecture. The model is characterized by its fast training speed, robustness to parameter choices, and accurate extraction of handwritten dates from various sources. For our training process, we build and introduce a database containing nearly 10 million tokens derived from over 2.2 million images of handwritten dates, extracted and segmented from diverse historical documents. Considering that dates are among the most prevalent pieces of information in historical documents, and given the existence of millions of such documents in historical archives, achieving efficient and automated extraction of dates holds the potential for substantial cost savings compared to manual transcription efforts. We demonstrate that training on handwritten text that exhibits substantial variability in writing styles yields robust models for recognizing general handwritten text and that transfer learning from the DARE system increases transcription accuracy substantially, allowing one to obtain high accuracy even when using relatively small training samples on entirely new types of documents.
Citing
If you would like to cite our paper, please use
@article{dahl2026dare,
title={DARE: A large-scale handwritten {DA}te {RE}cognition system},
author={Dahl, Christian M. and Johansen, Torben S. D. and S{\o}rensen, Emil N. and Westermann, Christian E. and Wittrock, Simon F.},
journal={International Journal on Document Analysis and Recognition (IJDAR)},
year={2026},
publisher={Springer},
doi={10.1007/s10032-026-00587-5}
}
or (as a reference to the earlier arXiv version)
@article{dahl2022dare,
title={DARE: A large-scale handwritten {DA}te {Re}cognition system},
author={Dahl, Christian M. and Johansen, Torben S. D. and S{\o}rensen, Emil N. and Westermann, Christian E. and Wittrock, Simon F.},
journal={arXiv preprint arXiv:2210.00503},
year={2022}
}
