🇬🇪 Georgian ASR — End-to-End Speech Recognition System
Automatic Speech Recognition system for the Georgian language, built from the ground up using state-of-the-art deep learning techniques, custom data pipelines, and domain-adapted language models.
Overview
Georgian (ქართული) is a South Caucasian language spoken by approximately 3.7 million people. Despite its rich linguistic heritage and unique script (მხედრული), Georgian remains severely underserved in the NLP and speech recognition space. This project addresses that gap by building a robust, general-purpose Georgian ASR system capable of handling diverse acoustic conditions — from clean studio recordings to noisy phone call audio.
The system achieves competitive Word Error Rate (WER) on Georgian conversational speech and is released for research and educational purposes only.
Architecture
Core Model — NeMo FastConformer Hybrid (RNNT + CTC)
The acoustic model is built on NVIDIA's FastConformer-Hybrid-Large architecture, fine-tuned from the pretrained checkpoint nvidia/stt_ka_fastconformer_hybrid_large_pc.
What is FastConformer?
FastConformer is an evolution of the Conformer architecture introduced in "Conformer: Convolution-augmented Transformer for Speech Recognition" (arxiv.org/abs/2005.08100). It combines:
- Multi-Head Self-Attention — captures long-range temporal dependencies in speech, allowing the model to understand context across the entire utterance
- Depthwise Separable Convolution — captures local acoustic patterns (phoneme-level features) efficiently
- Feed-Forward Layers — non-linear transformations between attention and convolution blocks
- Subsampling — reduces sequence length by 8x using strided convolutions, making attention computationally feasible on long sequences
FastConformer specifically improves over standard Conformer by using 8x subsampling instead of 4x and global attention with local windows, reducing compute by approximately 2.5x while maintaining accuracy.
Hybrid RNNT + CTC Decoder
The model uses a hybrid decoding architecture:
RNNT (Recurrent Neural Network Transducer) — the primary decoder used at inference time. RNNT jointly models acoustic and language probability, producing output tokens conditioned on both the audio encoding and previously predicted tokens. It handles variable-length input-output alignment natively through a special blank token mechanism, making it significantly more accurate than CTC for conversational speech.
CTC (Connectionist Temporal Classification) — used as an auxiliary loss during training only (weight 0.3). CTC enforces alignment structure on the encoder, stabilizing early training and improving convergence speed. At inference, CTC is discarded.
The joint loss during training is:
Total Loss = 0.7 × RNNT_Loss + 0.3 × CTC_Loss
Language Model — KenLM n-gram
KenLM (Heafield, 2011 — "KenLM: Faster and Smaller Language Model Queries", arxiv.org/abs/1106.1234) is used for shallow fusion rescoring at inference time. An n-gram language model trained on Georgian text corpora assigns probability to word sequences, boosting acoustically ambiguous but linguistically plausible outputs.
The KenLM model was trained on public Georgian text from multiple domains including financial documents, legal texts, news articles.
Shallow fusion combines the acoustic model score with the language model score at decode time:
Score(y|x) = log P_acoustic(y|x) + λ × log P_LM(y)
Where λ is a tunable interpolation weight.
Voice Activity Detection — Silero VAD
Silero VAD (github.com/snakers4/silero-vad) is used for audio segmentation before ASR inference. It is a lightweight temporal convolutional network (~1MB) that classifies 32ms audio frames as speech or silence with high accuracy across languages.
The VAD pipeline:
- Loads full-length audio
- Detects speech segment boundaries
- Splits into chunks at natural silence boundaries
- Passes cleaned segments to the ASR model
Silero VAD is language-agnostic — speech vs silence detection relies on universal acoustic properties (harmonic structure, energy distribution) that generalize across all human languages without fine-tuning.
Noise Reduction — Spectral Gating
Pre-processing applies spectral noise reduction (Sainburg et al., 2020 — based on Boll 1979's spectral subtraction) via the noisereduce library:
- Estimate noise spectral profile from silent segments
- Compute per-frequency noise threshold
- Apply sigmoid spectral gate — attenuate bins where signal ≈ noise floor
- Reconstruct via inverse STFT
This improves transcription accuracy on noisy phone call audio without any model retraining.
Data Pipeline
Data Collection
Training data was collected from diverse Georgian audio sources to maximize acoustic and linguistic coverage:
| Source | Type | Approximate Size |
|---|---|---|
| Mozilla Common Voice (Georgian) | Read speech, multiple speakers | ~4,200 utterances |
| Public Georgian Data | ~7,300 utterances | |
| Phone call recordings | Telephony, noisy, short utterances | ~7,700 utterances |
Total: ~20,000 unique original utterances, ~108 hours with augmentation
Data Labeling
Manual labeling was performed for original recordings. Gemini 2.5 Flash API was used for pseudo-labeling of new data, with a confidence filtering pipeline:
- Run current ASR model on unlabeled audio → model transcription
- Run Gemini API on same audio → Gemini transcription
- Compute symmetric Character Error Rate between the two
- Auto-accept entries with symmetric CER ≤ 10% (using Gemini text as label)
- Manual review for entries with CER 10-30%
- Discard entries with CER > 30%
This approach reduced manual labeling effort by approximately 60-70% while maintaining label quality.
KenLM vs RNNT — empirical comparison: A domain-specific KenLM n-gram language model was trained on Georgian text from financial documents, legal texts, and news articles, and applied via shallow fusion with the CTC decoder head as an alternative decoding strategy. Despite KenLM providing meaningful language model support to CTC, RNNT greedy decoding still outperformed CTC+KenLM on the test set. This confirms that RNNT's joint acoustic-language modeling — where the decoder conditions on both audio representations and previously predicted tokens simultaneously — is more powerful than post-hoc language model rescoring for Georgian conversational speech. KenLM remains available as a lightweight alternative for latency-constrained deployments where RNNT beam search is too slow.
Audio Preprocessing
Raw audio goes through the following pipeline before training:
Raw WAV
→ Resample to 16kHz mono (librosa)
→ VAD-based chunking (Silero VAD)
- Min chunk: 0.3s
- Max chunk: 15.0s
- Split on silence ≥ 200ms
→ Duration filtering
- Remove: duration < 0.3s or > 15.0s
→ Add to manifest
Augmentation Strategy
A hybrid static + dynamic augmentation approach was used:
Static augmentations (pre-generated, saved as separate files):
| Augmentation | Purpose | Ratio |
|---|---|---|
| Telephone bandwidth filter (300-3400Hz) | Simulate phone call audio | 1:1 |
| Codec noise (MP3/GSM compression) | Simulate compressed audio | 1:1 |
| Room Impulse Response convolution | Simulate reverberant environments | 1:1 |
| SNR noise mixing (MUSAN/DNS) | Simulate background noise | 1:1 |
Dynamic augmentations (applied on-the-fly during training via NeMo preprocessor):
| Augmentation | Parameters | Probability |
|---|---|---|
| Gain perturbation | -10 to +5 dBFS | 50% |
| Speed perturbation | 0.9x to 1.1x | 50% |
SpecAugment (applied on mel spectrogram during training):
- Frequency masking: 2 masks, max width 27 bins
- Time masking: 10 masks, max width 5% of utterance
The philosophy: static augmentations teach domain-specific acoustic robustness. Dynamic augmentations introduce infinite variety so the model never sees the exact same file twice across epochs.
Training Configuration
Model: nvidia/stt_ka_fastconformer_hybrid_large_pc (fine-tuned)
Optimizer: AdamW (weight_decay=1e-3)
LR: 1e-4 with CosineAnnealing (warmup=1000 steps, min_lr=1e-6)
Batch size: 32 (effective 64 with gradient accumulation × 2)
Precision: BF16 mixed precision
GPU: NVIDIA L4 (24GB)
Max epochs: 50 (early stopping patience=8 on val_wer)
CTC weight: 0.3 (auxiliary loss during training only)
Validation Strategy
Validation set: 400 unique original utterances (no augmented variants), distributed to mirror training domain proportions:
Zero leakage verified — no training stem appears in validation in any augmented form.
Checkpoint Averaging
After training, top-5 checkpoints (by val_wer) are averaged using uniform weight averaging:
averaged_weights[key] = sum(ckpt[key] for ckpt in top5) / 5
This smooths the loss landscape and consistently gives 0.1-0.3% absolute WER improvement over the single best checkpoint.
Results
| Model | WER | CER |
|---|---|---|
| This model (fine-tuned) | 12.32% | 3.89% |
| NVIDIA Base (starting point) | 26.70% | 10.38% |
| Azure Speech Service | 49.29% | 40.93% |
Evaluated on 575 clips from publicly available Georgian government and broadcaster content — National Bank press conferences, NCDC health briefings, etc. 54% relative WER improvement over base. 75% over Azure.
Fleurs KA Benchmark (google/fleurs, ka_ge test set)
| Model | WER | CER |
|---|---|---|
| This model (Last_Average.ckpt) | 11.45% | 4.06% |
| NVIDIA Base (nvidia/stt_ka_fastconformer_hybrid_large_pc) | 13.63% | 4.39% |
| Whisper Large-v3 (forced language=ka) | 76.33% | 37.05% |
Notably, the base NVIDIA model was trained directly on Fleurs Georgian train/dev splits — yet our fine-tuned model still outperforms it on the Fleurs test set, demonstrating genuine generalization rather than benchmark overfitting.
Key Technical Decisions
Why FastConformer over Whisper? FastConformer's hybrid RNNT+CTC architecture trains more stably on limited data than Whisper's encoder-decoder approach. RNNT handles streaming inference naturally and has lower latency for real-time applications. For a low-resource language like Georgian, fine-tuning a pretrained NVIDIA model proved more effective than fine-tuning Whisper.
Why static + dynamic augmentation? Static augmentations (telephone, codec, RIR, noise) teach the model specific acoustic degradation patterns. Dynamic augmentations (gain, speed) prevent memorization — the model never hears the exact same file twice across epochs. Separating these concerns gives cleaner control over what the model learns.
Why KenLM over neural LM rescoring? KenLM is fast, interpretable, and requires minimal Georgian text to train effectively. Neural LM rescoring (e.g. with GPT-style models) would give marginally better results but adds significant inference latency. For a production system targeting call center transcription, latency matters.
Why confidence filtering for pseudo-labels? Fully manual labeling of 20,000+ utterances is not feasible for a solo project. Using the ASR model and Gemini as dual references and keeping only high-agreement entries (symmetric CER ≤ 15%) gives near-manual label quality at a fraction of the cost. The two-reference approach catches cases where both systems hallucinate on the same audio (rare but possible) better than single-model filtering.
References
Gulati et al. (2020). Conformer: Convolution-augmented Transformer for Speech Recognition. arxiv.org/abs/2005.08100
Graves et al. (2012). Sequence Transduction with Recurrent Neural Networks. arxiv.org/abs/1211.3711 — Original RNNT paper.
Graves et al. (2006). Connectionist Temporal Classification. — Original CTC paper, ICML 2006.
Heafield (2011). KenLM: Faster and Smaller Language Model Queries. Proceedings of the Sixth Workshop on Statistical Machine Translation.
Park et al. (2019). SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition. arxiv.org/abs/1904.08779
Boll (1979). Suppression of acoustic noise in speech using spectral subtraction. IEEE Transactions on Acoustics, Speech, and Signal Processing.
Sainburg et al. (2020). Finding, visualizing, and quantifying latent structure across diverse animal acoustic repertoires. PLOS Computational Biology. — Spectral gating formalization used in noisereduce.
Ko et al. (2015). Audio augmentation for speech recognition. Interspeech 2015. — Speed perturbation augmentation.
Snyder et al. (2015). MUSAN: A Music, Speech, and Noise Corpus. arxiv.org/abs/1510.08484 — Noise corpus used for SNR augmentation.
Author
Built by — Senior AI/ML Engineer
This project demonstrates end-to-end ML engineering for a low-resource language: data collection and pseudo-labeling at scale, augmentation pipeline design, model fine-tuning, evaluation framework, and production inference pipeline.
For research and demonstration purposes.