Tilawi FastConformer (Quran recitation, on-device)

The speech model behind the Tilawi Quran app: it runs on the phone, so recitations never leave the device. Given a recitation, it produces a CTC transcript that @tilawi/quran-asr matches against the fixed Quran text to find the recited verse, or to give word-by-word memorization feedback.

  • Fine-tuned from NVIDIA's stt_ar_fastconformer_hybrid_large_pcd_v1.0 (CTC branch) on the author's own recitation recordings.
  • Exported to a single ONNX graph with the preprocessor inside, so the input is raw audio. Mixed quantization (int4 MatMul, int8 Conv/LayerNorm): 88 MB.
  • Runs with onnxruntime-react-native (iOS/Android) or onnxruntime-node.

Links: tilawi.ai (the app) · github.com/Tilawi (all open source) · npm @tilawi/quran-asr, @tilawi/react-native-quran-asr, @tilawi/expo-pcm-recorder

The model never outputs or alters Quran text. It transcribes the user's recitation; the matcher compares that transcript to the canonical text.

Files

File What it is
fastconformer_full_mixed.onnx The model. Input audio_signal float32 [1, N] (16 kHz mono PCM) and length int64 [1]; output log-probabilities [1, T, 1025].
vocab.json CTC token id to SentencePiece piece (1,025 entries; blank id 1024).
quran_ctc_tokens.json Each verse's text as model token ids, for CTC re-ranking.
quran.json 6,236 verse records (surah, ayah, text_uthmani, text_clean, surah names). Tanzil text; see TANZIL-NOTICE.txt.
SHA256SUMS Checksums of the four files above.

Accuracy

Verse identification through @tilawi/quran-asr. The numbers come from 600 verses (200 per reciter, sampled across the whole Quran, seed 1) from EveryAyah; real live tests, with people reciting into phones, were done as well:

Reciter Correct verse No match Wrong verse
Mohamed Siddiq al-Minshawi (murattal) 200 / 200 0 0
Mahmoud Khalil al-Husary 200 / 200 0 0
Mishary Rashid Alafasy 199 / 200 1 0

"Correct" includes 12 verses whose text is word-for-word identical to another verse (e.g. 55:13), which audio cannot tell apart. The harness and per-clip results are in the quran-asr repo.

Limitations. The table uses studio recordings because they can be scored automatically; the model has also been tested with live recitation recorded on phones. It isn't perfect: noise, very short verses or unclear recitation can lead to a missed word or no match, and the matcher prefers no answer over a wrong one. It is trained for Quran recitation (Hafs), not general Arabic speech. It's free and open source in the hope that it helps.

Use

In your own project: npm install @tilawi/quran-asr onnxruntime-node, then see the usage example. To try it from the repo:

git clone https://github.com/Tilawi/quran-asr && cd quran-asr
npm install && npm run build && npm run fetch-assets
node examples/transcribe.mjs recitation.mp3

Credits and license

  • Model weights: CC-BY-4.0. Base model: NVIDIA, stt_ar_fastconformer_hybrid_large_pcd_v1.0, CC-BY-4.0. Fine-tuning and export: Muhammed Durakovic.
  • Quran text in quran.json: Tanzil Project, CC-BY-3.0, redistributed verbatim; the full notice is in TANZIL-NOTICE.txt. The token table is derived from that text.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for muhdur/tilawi-fastconformer-quran

Quantized
(8)
this model