Automatic Speech Recognition
Transformers
Safetensors
Mongolian
whisper
mongolian
cyrillic
speech
Eval Results (legacy)
Instructions to use Blgn94/whisper-small-mn-v3 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Blgn94/whisper-small-mn-v3 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("automatic-speech-recognition", model="Blgn94/whisper-small-mn-v3")# Load model directly from transformers import AutoProcessor, AutoModelForSpeechSeq2Seq processor = AutoProcessor.from_pretrained("Blgn94/whisper-small-mn-v3") model = AutoModelForSpeechSeq2Seq.from_pretrained("Blgn94/whisper-small-mn-v3", device_map="auto") - Notebooks
- Google Colab
- Kaggle
| language: | |
| - mn | |
| license: apache-2.0 | |
| base_model: openai/whisper-small | |
| tags: | |
| - whisper | |
| - automatic-speech-recognition | |
| - mongolian | |
| - cyrillic | |
| - speech | |
| datasets: | |
| - Blgn94/mongolian-stt-dataset-v24 | |
| metrics: | |
| - wer | |
| - cer | |
| library_name: transformers | |
| pipeline_tag: automatic-speech-recognition | |
| model-index: | |
| - name: whisper-small-mn-v25 | |
| results: | |
| - task: | |
| type: automatic-speech-recognition | |
| name: Automatic Speech Recognition | |
| dataset: | |
| name: mongolian-stt-dataset-v24 (text-disjoint clean split) | |
| type: Blgn94/mongolian-stt-dataset-v24 | |
| split: validation (clean) | |
| metrics: | |
| - type: wer | |
| value: 21.33 | |
| name: WER | |
| - type: wer | |
| value: 19.67 | |
| name: WER (normalized) | |
| - type: cer | |
| value: 9.13 | |
| name: CER (normalized) | |
| # whisper-small-mn-v25 | |
| Fine-tune of [`openai/whisper-small`](https://huggingface.co/openai/whisper-small) for | |
| **Mongolian (Cyrillic Khalkha) speech-to-text**. This is the **v25** run β the winning | |
| configuration of a controlled A/B against a plain fine-tuning baseline, selected on a | |
| **text-disjoint** evaluation set (transcripts that never appear in training) rather than the | |
| optimistic in-domain split. | |
| - **Base model:** `openai/whisper-small` (244M params) | |
| - **Language:** Mongolian, Cyrillic script (mn) | |
| - **Dataset:** [`Blgn94/mongolian-stt-dataset-v24`](https://huggingface.co/datasets/Blgn94/mongolian-stt-dataset-v24) (~148 h of audio) | |
| - **Honest clean WER:** **21.33** (raw) / **19.67** (normalized) | |
| ## Results | |
| All numbers are Word/Character Error Rate (lower is better). Normalized = lowercased, | |
| punctuation stripped, NFC-normalized Cyrillic (Ρ/ΠΉ) β this is what published benchmarks use, | |
| so compare against the `_norm` columns. | |
| Three evaluation sets are reported: | |
| | eval set | what it is | trust | | |
| |---|---|---| | |
| | **clean** | dev clips whose transcripts **never** appear in training (241 clips) | the honest metric β **model was selected on this** | | |
| | **indomain** | old random split; ~89% of its transcripts also appear in training (300 clips) | optimistic (memorization), kept for continuity | | |
| | **mic** | 10 clips recorded on a laptop mic β real, noisy, out-of-distribution | reported only, too small to trust | | |
| ### v25 final metrics | |
| | eval set | WER | CER | WER (norm) | CER (norm) | | |
| |---|---|---|---|---| | |
| | **clean** | 21.33 | 9.78 | **19.67** | 9.13 | | |
| | indomain | 23.01 | 8.40 | 21.04 | 7.92 | | |
| | mic | 84.36 | 48.66 | 72.51 | 44.87 | | |
| Per-source normalized WER on the clean set: **mbspeech 6.03**, **fleurs 21.67**, | |
| **youtube 57.65** (spontaneous speech is the weak spot β see Limitations). | |
| ### A/B vs baseline (the reason this recipe was chosen) | |
| Both runs share the identical text-disjoint `clean` eval set, so the comparison is fair. | |
| | metric (normalized) | baseline | **v25** | Ξ | | |
| |---|---|---|---| | |
| | **clean WER** (selection metric) | 22.85 | **19.67** | **β3.18** | | |
| | clean CER | 9.82 | 9.13 | β0.69 | | |
| | clean β mbspeech | 8.18 | 6.03 | β2.14 | | |
| | clean β youtube (spontaneous) | 66.15 | 57.65 | **β8.50** | | |
| | mic (real-world) | 87.68 | 72.51 | **β15.17** | | |
| | indomain (optimistic) | 17.07 | 21.04 | +3.97 | | |
| The baseline is a plain fine-tune on the full corpus with no de-duplication or augmentation. | |
| v25 improves the **honest** clean WER by 3.2 points and generalizes markedly better on | |
| spontaneous YouTube speech and on real microphone recordings. The in-domain score *rises* β | |
| this is expected and desirable: capping duplicate transcripts removes the memorization crutch | |
| (the corpus repeats each Common Voice sentence ~14.5Γ), so the optimistic in-domain number | |
| falls while genuine transcription of unseen speech improves. | |
| ## Training data | |
| [`Blgn94/mongolian-stt-dataset-v24`](https://huggingface.co/datasets/Blgn94/mongolian-stt-dataset-v24) | |
| β ~148 h of 16 kHz Mongolian Cyrillic speech across four sources (measured by decoding all clips): | |
| | source | ~hours | notes | | |
| |---|---|---| | |
| | Common Voice | ~126 | read speech; only ~6,000 distinct sentences (~14.5Γ text reuse) | | |
| | FLEURS | ~14 | read speech | | |
| | MBSpeech | ~7 | read speech | | |
| | YouTube | ~1 | manual-subtitle clips β the only spontaneous speech | | |
| **Known weakness:** the corpus is ~85% Common Voice read speech and very light on spontaneous | |
| speech (~1 h). Read-speech sources score in the single-to-low-twenties WER; spontaneous | |
| YouTube sits near 58%. Expect the model to be strongest on clear, read-style Mongolian and | |
| weaker on conversational/noisy audio. | |
| ### v25 corpus shaping (what makes this run different) | |
| - **Per-sentence cap = 4:** at most 4 recordings per distinct transcript, dropping 66.5% of | |
| rows (95,946 β 32,134) and cutting Common Voice text reuse from 14.5Γ to 4Γ. Breaks the | |
| transcript-memorization shortcut. | |
| - **Source-balanced sampling (Ξ± = 0.5):** re-weights so Common Voice does not drown out FLEURS, | |
| MBSpeech, and YouTube. | |
| - **Audio augmentation (train only; eval always clean):** | |
| - SpecAugment (time masking, p = 0.05) | |
| - Waveform augmentation (p = 0.5, up to 2 ops): speed 0.9β1.1Γ, gain, additive noise at | |
| 10β30 dB SNR, light reverb. | |
| ## Training procedure | |
| | | | | |
| |---|---| | |
| | Base model | `openai/whisper-small` | | |
| | Precision | bf16 | | |
| | Effective batch | 16 (batch 4 Γ grad-accum 4) | | |
| | Learning rate | 1e-5, linear schedule, warmup ratio 0.05 | | |
| | Weight decay | 0.01 | | |
| | Label smoothing | 0.0 | | |
| | Epochs | 4 (8,036 steps) | | |
| | Max label tokens | 440 | | |
| | Seed | 42 | | |
| | Hardware | 1Γ NVIDIA RTX 5060 (8 GB, Blackwell) | | |
| | Wall-clock | ~2 h 53 m | | |
| | transformers / torch | 5.14.1 / 2.13.0+cu130 | | |
| Model selection: best checkpoint by `eval_clean_wer_norm` (`load_best_model_at_end`). | |
| Labels drop the leading `<|startoftranscript|>` token (the model re-prepends | |
| `decoder_start_token_id`). | |
| ### Reproduce | |
| ```bash | |
| MAX_PER_SENTENCE=4 SOURCE_ALPHA=0.5 AUG_PROB=0.5 EPOCHS=4 OUT=out_v25 \ | |
| python scripts/train.py | |
| ``` | |
| Full config and final metrics are in `run_manifest.json` in this repo. | |
| ## Usage | |
| ```python | |
| import torch | |
| from transformers import pipeline | |
| pipe = pipeline( | |
| "automatic-speech-recognition", | |
| model="Blgn94/whisper-small-mn-v25", | |
| device=0 if torch.cuda.is_available() else -1, | |
| ) | |
| result = pipe( | |
| "audio.mp3", | |
| generate_kwargs={"language": "mn", "task": "transcribe"}, | |
| ) | |
| print(result["text"]) | |
| ``` | |
| For long audio, add `chunk_length_s=30`. For raw model + processor access: | |
| ```python | |
| from transformers import WhisperForConditionalGeneration, WhisperProcessor | |
| processor = WhisperProcessor.from_pretrained("Blgn94/whisper-small-mn-v25") | |
| model = WhisperForConditionalGeneration.from_pretrained("Blgn94/whisper-small-mn-v25") | |
| ``` | |
| ## Limitations | |
| - **Spontaneous / conversational speech:** weak (~58% WER on YouTube), because the corpus has | |
| only ~1 h of it. Real-world microphone input is also harder (mic set ~72% WER, but only | |
| 10 clips β treat as directional, not a benchmark). | |
| - **Not evaluated on the official Common Voice test split.** The `clean` set is a custom | |
| text-disjoint split of v24; it is *not* speaker-disjoint from training, so real-world WER on | |
| entirely new speakers may be somewhat higher. For a speaker-disjoint reference point, | |
| `bayartsogt/whisper-small-mn-8` reports ~26.5% WER on the official Common Voice mn test. | |
| - **Cyrillic Khalkha only.** No traditional Mongolian script; not tuned for other dialects. | |
| ## License | |
| Apache-2.0, inheriting from the `openai/whisper-small` base model. Note the underlying | |
| training data combines sources under their own licenses (Common Voice, FLEURS, MBSpeech, | |
| YouTube manual subtitles); review those before commercial use. | |