MuseumSCAT TrOCR glyph scorer

This is the final auxiliary handwriting checkpoint used in the fourth-place solution to the MuseumSCAT Specimen Collection Annotation Task.

The model was not used to generate the final locality transcription freely. Its intended use is candidate-conditioned scoring: given a locality crop and several strings proposed by vision-language models, it estimates which candidate is most strongly supported by the visible glyphs. The complete solution code is available in octavigrau/kaggle-museumSCAT.

Base model

The checkpoint was initialized from Riksarkivet/trocr-base-handwritten-hist-swe-2, an Apache-2.0 TrOCR checkpoint for historical Swedish handwriting.

Training data

The final training mixture contained:

  • 148 labeled locality crops from the MuseumSCAT training set;
  • approximately 3,500 synthetic Danish place-name lines rendered with handwriting fonts and image degradation;
  • 1,149 conservatively pseudo-labeled crops from the unlabeled MuseumSCAT test set.

The test pseudo-labels were produced automatically from high-confidence model agreement; they were not manually transcribed. No MuseumSCAT images, crops, labels, or pseudo-label manifests are distributed in this model repository.

The MuseumSCAT competition data is provided under CC BY-NC 4.0. Users should also acknowledge the competition and the Natural History Museum of Denmark as required by the competition rules.

Training

  • encoder frozen;
  • published weights selected at epoch 3 of a restarted run using candidate-selection accuracy; the run continued through epoch 7 before ending;
  • AdamW, learning rate 2e-5, weight decay 0.01;
  • batch size 16;
  • rotation, contrast, brightness and sharpness augmentation;
  • Transformers 4.46.3.

The checkpoint was selected using candidate-selection accuracy. Adding the pseudo-labeled crops moved the held-out result from 33/40 to 34/40 selections, a noise-sized improvement. In the competition pipeline, replacing the VLM medoid broadly with the TrOCR choice was harmful; the model was useful as a confidence signal and inside a strict correction gate.

Candidate scoring

For an image crop $x$ and candidate string $y$, the score is mean next-token log-likelihood:

s(yx)=1ytlogp(yty<t,x). s(y\mid x)=\frac{1}{|y|}\sum_t \log p(y_t\mid y_{<t},x).

import torch
from PIL import Image
from transformers import TrOCRProcessor, VisionEncoderDecoderModel

checkpoint = "octavigrau/museumscat-trocr-glyph-scorer"
processor = TrOCRProcessor.from_pretrained(checkpoint)
model = VisionEncoderDecoderModel.from_pretrained(checkpoint).eval()

@torch.inference_mode()
def score_candidate(image: Image.Image, candidate: str) -> float:
    pixels = processor(images=image.convert("RGB"), return_tensors="pt").pixel_values
    target = processor.tokenizer(candidate, return_tensors="pt").input_ids
    encoded = model.encoder(pixel_values=pixels).last_hidden_state
    logits = model.decoder(input_ids=target, encoder_hidden_states=encoded).logits[:, :-1, :]
    next_tokens = target[:, 1:]
    token_scores = logits.log_softmax(-1).gather(-1, next_tokens.unsqueeze(-1)).squeeze(-1)
    return float(token_scores.mean())

scores = {candidate: score_candidate(Image.open("locality_crop.png"), candidate)
          for candidate in ["Ørslev", "Ørholm"]}
print(max(scores, key=scores.get), scores)

Higher scores are preferred. Scores are meaningful for comparing candidates for the same crop; they should not be interpreted as calibrated probabilities across images.

Limitations

  • The labeled in-domain set contained only 148 crops.
  • The evaluation set was small and designed around candidate-selection failures rather than general handwriting recognition.
  • High-confidence pseudo-labels were mostly easy cases and added little.
  • The checkpoint inherits the biases and vocabulary of a Swedish historical-handwriting base model.
  • It expects an already-localized text crop and is not a detector.
  • It should not be presented as a general Danish HTR benchmark result.

Attribution

  • MuseumSCAT organizers and the Natural History Museum of Denmark for the competition data.
  • The Swedish National Archives and collaborators for the Riksarkivet TrOCR checkpoint.
  • GeoNames supplied a separate CC BY 4.0 gazetteer signal in the full competition solution; GeoNames data was not used to train this checkpoint.
Downloads last month
35
Safetensors
Model size
0.4B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for octavigrau/museumscat-trocr-glyph-scorer