--- license: mit language: - ar - aeb pipeline_tag: image-to-text base_model: microsoft/trocr-base-handwritten tags: - image-to-text - ocr - handwriting-recognition - htr - arabic - tunisian - trocr - vision-encoder-decoder - offline --- # trocr-tunisian-arabic — handwritten Tunisian Arabic recognition A TrOCR vision-encoder/decoder fine-tuned for **handwritten Tunisian Arabic**, with a 30k-token Arabic-native vocabulary. ## Why this exists Cloud OCR services handle printed Arabic reasonably and handwritten dialectal Arabic poorly. For the use case this was built for, they were not an option at all: Tunisian legal documents contain privileged client data that cannot leave the premises under the country's data-protection regime. That ruled out every managed service and required a model that runs fully offline. ## Intended use Offline extraction of handwritten Arabic from structured forms. Built for and deployed in [Wathiqa](https://github.com/GhwazouaniWala/Wathiqa), which turns handwritten Tunisian legal forms into structured, searchable client records without any cloud dependency. **This is a word-level model.** In Wathiqa it sits behind a segmentation stage: the ink mask is dilated horizontally by a kernel scaled to text height — Arabic joins letters within a word but not between words, so inter-word gaps are the reliable separator — and the resulting components are read right to left. Feeding a full unsegmented line will not work well. ## Usage ## Where it sits in a real pipeline Wathiqa routes each field type to a purpose-built engine rather than asking one model to cover everything: | Field type | Engine | |---|---| | Arabic text, names, notes | **This model** | | CIN, passport, number, phone, date | Dedicated digit recogniser with a constrained alphabet and shape validation | | Printed captions | EasyOCR (Arabic + Latin) | | Checkboxes | Fill-ratio detector | | Signatures | Ink-presence detector | That routing exists because these are genuinely different problems, and the limitations below are the reason. ## Limitations - **Trained on isolated words**, not lines or sentences. Word segmentation is the caller's responsibility. - **Not for digits.** Numeric fields need a separate path — a word model asked for a digit string produces confident nonsense. Wathiqa isolates characters, constrains the alphabet to `0-9` and separators, then validates the assembled value against the field's expected shape. - **Not for printed text.** Printed captions are a different problem and route elsewhere. - **A blank crop returns confident nonsense rather than silence.** Empty-field detection must happen before the model is called; Wathiqa reports fields with no ink as empty and never sends them to a model. - Trained on Tunisian handwriting. Performance on other regional Arabic hands is untested. ## Training Fine-tuned from [`microsoft/trocr-base-handwritten`](https://huggingface.co/microsoft/trocr-base-handwritten) on handwritten Tunisian Arabic, with the decoder vocabulary replaced by a 30k-token Arabic-native vocabulary. ## Related work | Model | Purpose | |---|---| | [`Ghazouaniwala/emotions_speech`](https://huggingface.co/Ghazouaniwala/emotions_speech) | Speaker-independent speech emotion recognition | | [`Ghazouaniwala/silma-tts-derja`](https://huggingface.co/Ghazouaniwala/silma-tts-derja) | Tunisian Derja speech synthesis | ## Citation ```bibtex @misc{ghazouani2025trocrtunisian, author = {Ghazouani, Wala Eddine}, title = {trocr-tunisian-arabic: handwritten Tunisian Arabic recognition}, year = {2025}, url = {https://huggingface.co/Ghazouaniwala/trocr-tunisian-arabic} } ``` --- Built by [Wala Eddine Ghazouani](https://www.walaghazouani.com) · [GitHub](https://github.com/GhwazouaniWala) · [LinkedIn](https://www.linkedin.com/in/ghazouani-wala-eddine)