Yo-ByT5: Byte-Level Diacritic Restoration for Yorùbá

Yo-ByT5 is a fine-tune of google/byt5-small that restores Yorùbá diacritics (tone marks: acute ◌́, grave ◌̀, mid unmarked; underdots: ẹ, ọ, ṣ) to undiacritized text.

  • Developer: Gali Ahmad Samuel (lazymonster)
  • Language: Yorùbá (yo)
  • Task: Automatic Diacritic Restoration (ADR)
  • License: Apache 2.0

Quickstart

import unicodedata
from transformers import AutoTokenizer, AutoModelForSeq2SeqLM

model_id = "lazymonster/yobyt5-restoration"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForSeq2SeqLM.from_pretrained(model_id)

text = "Eko ni kokoro aseyori."
inputs = tokenizer(text, return_tensors="pt", max_length=1024, truncation=True)
outputs = model.generate(inputs["input_ids"], max_length=1024, num_beams=1)

predicted = tokenizer.decode(outputs[0], skip_special_tokens=True)
predicted = unicodedata.normalize("NFC", predicted)
print(predicted)
# Ẹ̀kọ́ ni kọ́kọ́rọ́ àṣeyọrí.

Benchmark Results

Evaluated on the YAD Test Benchmark (3,330 sentences) using base-word Needleman-Wunsch sequence alignment and Unicode NFC normalization.

The Corrected Reference repairs 1,415 non-standard underdot codepoints (U+0329 to standard U+0323) in the raw YAD release.

Metric Decoding Official Reference Corrected Reference
DER (Total) Beam (num_beams=5) 10.58% 10.14%
DER (Total) Greedy (num_beams=1) 10.45% 10.00%
CER Beam (num_beams=5) 3.75% 3.48%
CER Greedy (num_beams=1) 3.82% 3.55%
WER Beam (num_beams=5) 16.03% 14.86%
WER Greedy (num_beams=1) 16.11% 14.93%
DER-tone Beam (num_beams=5) 8.93% 8.93%
DER-tone Greedy (num_beams=1) 8.76% 8.76%
DER-underdot Beam (num_beams=5) 6.12% 6.15%
DER-underdot Greedy (num_beams=1) 5.66% 5.69%
WDER Beam (num_beams=5) 16.88% 15.64%
WDER Greedy (num_beams=1) 16.81% 15.58%
BLEU Beam (num_beams=5) 0.6841 0.6841
BLEU Greedy (num_beams=1) 0.6837 0.6837
ChrF Greedy / Beam 0.8431 0.8431

Training Data

Dataset Sentences License
MENYO-20k (Yorùbá train split) 9,942 CC BY-NC 4.0
Biblica® Open Yorùbá Contemporary Bible 2017 36,371 CC BY-SA
Total Train 46,313
Validation Set 5,305

Training Procedure

  • Architecture: google/byt5-small (299.6M parameters)
  • Compute: Google Cloud TPU v6e-8
  • Optimizer: AdamW, linear decay, weight decay 0.01, gradient clip norm 0.5
  • Batch Size: Per-device batch 4, gradient accumulation 2 (effective global batch size 64)
  • Phase 1: Learning rate 2e-4, 300 warmup steps, 4 epochs
  • Phase 2: Learning rate 1e-4, 0 warmup steps, 4 epochs (released checkpoint)

Recommendations

  1. Unicode normalization: Always normalize input and generated text to NFC before scoring or processing.
  2. Chunking: Chunk inputs longer than 1024 bytes at sentence or clause boundaries.
  3. Decoding: Greedy decoding is fast and accurate. Beam search (num_beams=5) offers slight quality gains on edit distance metrics.

Citation & License

License

  • Model weights: Apache 2.0
  • Training data: Subject to original licenses (MENYO-20k: CC BY-NC 4.0; Biblica: CC BY-SA)

Citation

@misc{gali2025yobyt5,
  author       = {Gali Ahmad, Samuel},
  title        = {Yo-ByT5: Byte-Level Diacritic Restoration for Yor\`ub\'a},
  year         = {2025},
  publisher    = {Hugging Face},
  howpublished = {\url{https://huggingface.co/lazymonster/yobyt5-restoration}}
}

Acknowledgments

Trained as a member of the HausaNLP Research Group with compute support from the Google TPU Research Cloud (TRC) program.

Downloads last month
384
Safetensors
Model size
0.3B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for lazymonster/yobyt5-restoration

Finetuned
(334)
this model

Dataset used to train lazymonster/yobyt5-restoration

Space using lazymonster/yobyt5-restoration 1