πŸ“– Quranic ASR Leaderboard

A leakage-free, held-out benchmark of Arabic ASR systems on Quranic recitation.

⭐ Head-to-head vs Tarteel's official realtime model: all 600 clips through Tarteel's production ASR (voice-v2.tarteel.io), same scorer: 9.56 overall WER (benchmark v1.1). The overall leader is zipformer_p-arabic-v3, a streaming phoneme model, at 3.46* WER; the best open-vocabulary system is our offline FastConformer at 4.13. Streaming field: v3 (3.46*), v2 (5.64*), zipformer_p-quran (5.83*), Tarteel (9.56), streaming FastConformer (10.53). *Closed-vocab nearest-ayah retrieval β€” the PER column and the "(PER)" subset cells give the phoneme models' native accuracy; v2/p-quran WER still scored against v1.0 references, re-run pending. Β πŸ€— Benchmark dataset Β· πŸš€ Submit your model

Every model is evaluated on the same 600 held-out clips (200 per source) with the same scorer; every clip is verified absent from our training data. Rankings use overall WER on all 600 clips (lower is better).

Metric
Mode
Sort by
Order
Rank
Model
Mode
Size (B)
Overall
PER
Phone (real-world)
EveryAyah
QUL (unseen reciter)
WER alef-ins.
πŸ₯‡ 1
🟒 Streaming
0.0655
10.32
11.54
22.75
12.39
11.91
10.14

Overall spans all 600 clips; per-source columns are 200 clips each. Lower is better.

@misc{quranlab2026benchmark,
  title        = {Quranic ASR Benchmark: a leakage-free held-out evaluation of Arabic ASR on Quranic recitation},
  author       = {{Quran-Lab}},
  year         = {2026},
  howpublished = {\url{https://huggingface.co/datasets/Quran-Lab/quranic-asr-benchmark}}
}