You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

zipformer-quran-ternary

Ternary Zipformer Quran phoneme ASR model from Quran Lab.

Model

Repository:

raufazhda/zipformer-quran-ternary

Model name:

zipformer-quran-ternary

Final checkpoint:

zipformer-quran-ternary.pt

SHA256:

8d11f69b7190b2e4d4121c208d946200ca9d58eb410e70e874f470f13cac85c4

The model uses 250 non-blank phoneme units with CTC blank ID 250, for 251 total CTC output classes.

Ternary quantization

The model was trained with ternary quantization-aware training.

Ternary candidate weights use three states:

{-1, 0, +1}

with FP32 per-output-channel scaling.

Quantization configuration:

  • tau: 0.70
  • candidate matrices: 281
  • candidate parameters: 63,908,352
  • candidate coverage: 97.162%

The optimized packed runtime artifact is still under development. It is intentionally not published in this repository revision.

DEV-600

The locked DEV-600 benchmark contains:

  • 600 evaluation rows
  • 599 physical audio files
  • 25,989 reference phoneme tokens

Final zipformer-quran-ternary result:

  • PER: 3.3360267805610064%
  • substitutions: 350
  • deletions: 182
  • insertions: 335
  • total errors: 867

Comparison:

Model PER
zipformer-quran-ternary 3.336%
Original FP32 4.129%
Direct hard ternary PTQ 13.648%

The final QAT model received additional continued training relative to the original FP32 checkpoint. The comparison therefore should not be interpreted as isolating ternary regularization as the sole cause of the difference.

QuranTTS unseen benchmark

The additional unseen QuranTTS benchmark contains 16 qari:

  • 3,200 clean clips
  • 3,200 RIR/reverberant clips

Results:

Model Clean PER RIR PER
zipformer-quran-ternary 2.217% 3.043%
Original FP32 3.963% 4.476%
Direct hard ternary PTQ 8.154% 11.217%

The final model had lower PER than the original FP32 model for 16/16 qari on clean evaluation and 16/16 qari on RIR evaluation.

This does not by itself establish intrinsically greater RIR robustness because the models differ in continued-training history.

Benchmark packaging

To avoid thousands of small repository files, QuranTTS benchmark audio is stored in two uncompressed TAR archives:

  • benchmark/qurantts/qurantts_clean.tar
  • benchmark/qurantts/qurantts_rir.tar

benchmark/qurantts/audio_index.jsonl maps each benchmark row to its archive member.

Repository layout

zipformer-quran-ternary.pt
README.md
LICENSE
config.json
SHA256SUMS

tokenizer/
  phoneme_units.json

benchmark/
  dev600/
    manifest.jsonl
    results.json

  qurantts/
    manifest.jsonl
    audio_index.jsonl
    results.json
    qurantts_clean.tar
    qurantts_rir.tar

License

See LICENSE.

This repository uses the Quran-Lab No-Profit License 1.2.

Downloads last month
24
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support