Document-level language ID gives one language per
page. Real pages are not one language: the navigation is English, the body
might be Ojibwe, the comments might be three other things. Filtering a corpus on the
document label throws that text away, and low-resource languages are the ones that
cannot spare it.
On the benchmark corpus, document-level LID
recovers 0.0% of the low-resource spans it would otherwise discard
or misfile. Tessera recovers 86% (90% on the rarest tier).
Tessera runs behind GlotLID / OpenLID, not instead of them.
Loading the model (17 MB, runs entirely in your browser)…
Analyse some text to see what a single-label filter would
cost.