Swedish Fraktur — Kraken OCR/HTR model

A Kraken recognition model (.mlmodel) for OCR of older Swedish text printed in Fraktur (blackletter). It produces a diplomatic transcription that keeps historical orthography, the long s (ſ) and period spelling, and works on scanned page images.

  • File: svensk_fraktur.mlmodel
  • Type: Kraken recognition model (baseline/HTR pipeline)
  • DOI: 10.5281/zenodo.20702142
  • Best validation accuracy: 0.9880 (≈ 1.2 % CER) on a held-out split
  • Base model: german_print (fine-tuned from it)
  • Training data: Språkbanken, Svensk fraktur 1626–1816

Usage

Fetch the model directly from Zenodo: kraken get 10.5281/zenodo.20702142

Kraken (CLI):

kraken -i page.jpg out.txt segment -bl ocr -m svensk_fraktur.mlmodel

For multi-page PDFs, render pages to images first (e.g. ~300 DPI) and pass them to Kraken. eScriptorium: import the .mlmodel under Models.

How it was trained

Fine-tuned from german_print on the Språkbanken corpus Svensk fraktur 1626–1816 (199 page images with line-level diplomatic transcriptions). The page images were OCR-bootstrapped and the recognized lines aligned to the ground-truth line text (≈98 % coverage), producing PageXML for ketos train. A low learning rate (-r 0.0001) was decisive — it let the model improve steadily past epoch 0 instead of drifting away from the strong starting point. See training/ for the scripts and exact commands.

The validation accuracy is measured on a held-out split of the same corpus and is therefore optimistic relative to entirely new documents; on a real volume (1600s Swedish Fraktur) it produced a near-flawless body-text transcription with preserved long-s and period spelling and no systematic substitution errors.

Limitations

  • Trained on Swedish Fraktur print; not intended for handwriting or modern (antiqua/roman) type.
  • Ornate/decorated title-page initials are read less reliably than body text.
  • Output is diplomatic (verbatim): historical spelling, long-s and printed line-break hyphens are preserved. Modernisation/normalisation should happen in a downstream step, not here.

Provenance, rights and attribution

This is a derivative model. Full chain:

Layer Resource By DOI License
Base model german_print (OCR model for German prints) S. Weil, J. Kamlah, T. Schmidt (2023) 10.5281/zenodo.10519596 CC0-1.0
Training data Svensk fraktur 1626–1816 Språkbanken Text, University of Gothenburg 10.23695/5sme-7437 CC-BY-4.0

The base model is CC0 (no attribution legally required; cited as courtesy). The training data is CC-BY-4.0, which requires attribution. When you use or redistribute this model, please keep the following attribution:

Trained on Svensk fraktur 1626–1816, Språkbanken Text, University of Gothenburg (digitisation: Gothenburg University Library; transcription: GREPECT), doi.org/10.23695/5sme-7437, licensed CC-BY-4.0.

This model is released under CC-BY-4.0 (see LICENSE).

Citation

See CITATION.cff.

Downloads last month
8
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support