🇬🇪 Georgian ASR — End-to-End Speech Recognition System

Automatic Speech Recognition system for the Georgian language, built from the ground up using state-of-the-art deep learning techniques, custom data pipelines, and domain-adapted language models.


Overview

Georgian (ქართული) is a South Caucasian language spoken by approximately 3.7 million people. Despite its rich linguistic heritage and unique script (მხედრული), Georgian remains severely underserved in the NLP and speech recognition space. This project addresses that gap by building a robust, general-purpose Georgian ASR system capable of handling diverse acoustic conditions — from clean studio recordings to noisy phone call audio.

The system achieves competitive Word Error Rate (WER) on Georgian conversational speech and is released for research and educational purposes only.


Architecture

Core Model — NeMo FastConformer Hybrid (RNNT + CTC)

The acoustic model is built on NVIDIA's FastConformer-Hybrid-Large architecture, fine-tuned from the pretrained checkpoint nvidia/stt_ka_fastconformer_hybrid_large_pc.

What is FastConformer?

FastConformer is an evolution of the Conformer architecture introduced in "Conformer: Convolution-augmented Transformer for Speech Recognition" (arxiv.org/abs/2005.08100). It combines:

  • Multi-Head Self-Attention — captures long-range temporal dependencies in speech, allowing the model to understand context across the entire utterance
  • Depthwise Separable Convolution — captures local acoustic patterns (phoneme-level features) efficiently
  • Feed-Forward Layers — non-linear transformations between attention and convolution blocks
  • Subsampling — reduces sequence length by 8x using strided convolutions, making attention computationally feasible on long sequences

FastConformer specifically improves over standard Conformer by using 8x subsampling instead of 4x and global attention with local windows, reducing compute by approximately 2.5x while maintaining accuracy.

Hybrid RNNT + CTC Decoder

The model uses a hybrid decoding architecture:

RNNT (Recurrent Neural Network Transducer) — the primary decoder used at inference time. RNNT jointly models acoustic and language probability, producing output tokens conditioned on both the audio encoding and previously predicted tokens. It handles variable-length input-output alignment natively through a special blank token mechanism, making it significantly more accurate than CTC for conversational speech.

CTC (Connectionist Temporal Classification) — used as an auxiliary loss during training only (weight 0.3). CTC enforces alignment structure on the encoder, stabilizing early training and improving convergence speed. At inference, CTC is discarded.

The joint loss during training is:

Total Loss = 0.7 × RNNT_Loss + 0.3 × CTC_Loss

Language Model — KenLM n-gram

KenLM (Heafield, 2011 — "KenLM: Faster and Smaller Language Model Queries", arxiv.org/abs/1106.1234) is used for shallow fusion rescoring at inference time. An n-gram language model trained on Georgian text corpora assigns probability to word sequences, boosting acoustically ambiguous but linguistically plausible outputs.

The KenLM model was trained on public Georgian text from multiple domains including financial documents, legal texts, news articles.

Shallow fusion combines the acoustic model score with the language model score at decode time:

Score(y|x) = log P_acoustic(y|x) + λ × log P_LM(y)

Where λ is a tunable interpolation weight.


Voice Activity Detection — Silero VAD

Silero VAD (github.com/snakers4/silero-vad) is used for audio segmentation before ASR inference. It is a lightweight temporal convolutional network (~1MB) that classifies 32ms audio frames as speech or silence with high accuracy across languages.

The VAD pipeline:

  1. Loads full-length audio
  2. Detects speech segment boundaries
  3. Splits into chunks at natural silence boundaries
  4. Passes cleaned segments to the ASR model

Silero VAD is language-agnostic — speech vs silence detection relies on universal acoustic properties (harmonic structure, energy distribution) that generalize across all human languages without fine-tuning.


Noise Reduction — Spectral Gating

Pre-processing applies spectral noise reduction (Sainburg et al., 2020 — based on Boll 1979's spectral subtraction) via the noisereduce library:

  1. Estimate noise spectral profile from silent segments
  2. Compute per-frequency noise threshold
  3. Apply sigmoid spectral gate — attenuate bins where signal ≈ noise floor
  4. Reconstruct via inverse STFT

This improves transcription accuracy on noisy phone call audio without any model retraining.


Data Pipeline

Data Collection

Training data was collected from diverse Georgian audio sources to maximize acoustic and linguistic coverage:

Source Type Approximate Size
Mozilla Common Voice (Georgian) Read speech, multiple speakers ~4,200 utterances
Public Georgian Data ~7,300 utterances
Phone call recordings Telephony, noisy, short utterances ~7,700 utterances

Total: ~20,000 unique original utterances, ~108 hours with augmentation

Data Labeling

Manual labeling was performed for original recordings. Gemini 2.5 Flash API was used for pseudo-labeling of new data, with a confidence filtering pipeline:

  1. Run current ASR model on unlabeled audio → model transcription
  2. Run Gemini API on same audio → Gemini transcription
  3. Compute symmetric Character Error Rate between the two
  4. Auto-accept entries with symmetric CER ≤ 10% (using Gemini text as label)
  5. Manual review for entries with CER 10-30%
  6. Discard entries with CER > 30%

This approach reduced manual labeling effort by approximately 60-70% while maintaining label quality.

KenLM vs RNNT — empirical comparison: A domain-specific KenLM n-gram language model was trained on Georgian text from financial documents, legal texts, and news articles, and applied via shallow fusion with the CTC decoder head as an alternative decoding strategy. Despite KenLM providing meaningful language model support to CTC, RNNT greedy decoding still outperformed CTC+KenLM on the test set. This confirms that RNNT's joint acoustic-language modeling — where the decoder conditions on both audio representations and previously predicted tokens simultaneously — is more powerful than post-hoc language model rescoring for Georgian conversational speech. KenLM remains available as a lightweight alternative for latency-constrained deployments where RNNT beam search is too slow.

Audio Preprocessing

Raw audio goes through the following pipeline before training:

Raw WAV
  → Resample to 16kHz mono (librosa)
  → VAD-based chunking (Silero VAD)
     - Min chunk: 0.3s
     - Max chunk: 15.0s
     - Split on silence ≥ 200ms
  → Duration filtering
     - Remove: duration < 0.3s or > 15.0s
  → Add to manifest

Augmentation Strategy

A hybrid static + dynamic augmentation approach was used:

Static augmentations (pre-generated, saved as separate files):

Augmentation Purpose Ratio
Telephone bandwidth filter (300-3400Hz) Simulate phone call audio 1:1
Codec noise (MP3/GSM compression) Simulate compressed audio 1:1
Room Impulse Response convolution Simulate reverberant environments 1:1
SNR noise mixing (MUSAN/DNS) Simulate background noise 1:1

Dynamic augmentations (applied on-the-fly during training via NeMo preprocessor):

Augmentation Parameters Probability
Gain perturbation -10 to +5 dBFS 50%
Speed perturbation 0.9x to 1.1x 50%

SpecAugment (applied on mel spectrogram during training):

  • Frequency masking: 2 masks, max width 27 bins
  • Time masking: 10 masks, max width 5% of utterance

The philosophy: static augmentations teach domain-specific acoustic robustness. Dynamic augmentations introduce infinite variety so the model never sees the exact same file twice across epochs.


Training Configuration

Model:        nvidia/stt_ka_fastconformer_hybrid_large_pc (fine-tuned)
Optimizer:    AdamW (weight_decay=1e-3)
LR:           1e-4 with CosineAnnealing (warmup=1000 steps, min_lr=1e-6)
Batch size:   32 (effective 64 with gradient accumulation × 2)
Precision:    BF16 mixed precision
GPU:          NVIDIA L4 (24GB)
Max epochs:   50 (early stopping patience=8 on val_wer)
CTC weight:   0.3 (auxiliary loss during training only)

Validation Strategy

Validation set: 400 unique original utterances (no augmented variants), distributed to mirror training domain proportions:

Zero leakage verified — no training stem appears in validation in any augmented form.

Checkpoint Averaging

After training, top-5 checkpoints (by val_wer) are averaged using uniform weight averaging:

averaged_weights[key] = sum(ckpt[key] for ckpt in top5) / 5

This smooths the loss landscape and consistently gives 0.1-0.3% absolute WER improvement over the single best checkpoint.


Results

Model WER CER
This model (fine-tuned) 12.32% 3.89%
NVIDIA Base (starting point) 26.70% 10.38%
Azure Speech Service 49.29% 40.93%

Evaluated on 575 clips from publicly available Georgian government and broadcaster content — National Bank press conferences, NCDC health briefings, etc. 54% relative WER improvement over base. 75% over Azure.


Fleurs KA Benchmark (google/fleurs, ka_ge test set)

Model WER CER
This model (Last_Average.ckpt) 11.45% 4.06%
NVIDIA Base (nvidia/stt_ka_fastconformer_hybrid_large_pc) 13.63% 4.39%
Whisper Large-v3 (forced language=ka) 76.33% 37.05%

Notably, the base NVIDIA model was trained directly on Fleurs Georgian train/dev splits — yet our fine-tuned model still outperforms it on the Fleurs test set, demonstrating genuine generalization rather than benchmark overfitting.

Key Technical Decisions

Why FastConformer over Whisper? FastConformer's hybrid RNNT+CTC architecture trains more stably on limited data than Whisper's encoder-decoder approach. RNNT handles streaming inference naturally and has lower latency for real-time applications. For a low-resource language like Georgian, fine-tuning a pretrained NVIDIA model proved more effective than fine-tuning Whisper.

Why static + dynamic augmentation? Static augmentations (telephone, codec, RIR, noise) teach the model specific acoustic degradation patterns. Dynamic augmentations (gain, speed) prevent memorization — the model never hears the exact same file twice across epochs. Separating these concerns gives cleaner control over what the model learns.

Why KenLM over neural LM rescoring? KenLM is fast, interpretable, and requires minimal Georgian text to train effectively. Neural LM rescoring (e.g. with GPT-style models) would give marginally better results but adds significant inference latency. For a production system targeting call center transcription, latency matters.

Why confidence filtering for pseudo-labels? Fully manual labeling of 20,000+ utterances is not feasible for a solo project. Using the ASR model and Gemini as dual references and keeping only high-agreement entries (symmetric CER ≤ 15%) gives near-manual label quality at a fraction of the cost. The two-reference approach catches cases where both systems hallucinate on the same audio (rare but possible) better than single-model filtering.


References

  1. Gulati et al. (2020). Conformer: Convolution-augmented Transformer for Speech Recognition. arxiv.org/abs/2005.08100

  2. Graves et al. (2012). Sequence Transduction with Recurrent Neural Networks. arxiv.org/abs/1211.3711 — Original RNNT paper.

  3. Graves et al. (2006). Connectionist Temporal Classification. — Original CTC paper, ICML 2006.

  4. Heafield (2011). KenLM: Faster and Smaller Language Model Queries. Proceedings of the Sixth Workshop on Statistical Machine Translation.

  5. Park et al. (2019). SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition. arxiv.org/abs/1904.08779

  6. Boll (1979). Suppression of acoustic noise in speech using spectral subtraction. IEEE Transactions on Acoustics, Speech, and Signal Processing.

  7. Sainburg et al. (2020). Finding, visualizing, and quantifying latent structure across diverse animal acoustic repertoires. PLOS Computational Biology. — Spectral gating formalization used in noisereduce.

  8. Ko et al. (2015). Audio augmentation for speech recognition. Interspeech 2015. — Speed perturbation augmentation.

  9. Snyder et al. (2015). MUSAN: A Music, Speech, and Noise Corpus. arxiv.org/abs/1510.08484 — Noise corpus used for SNR augmentation.


Author

Built by — Senior AI/ML Engineer

This project demonstrates end-to-end ML engineering for a low-resource language: data collection and pseudo-labeling at scale, augmentation pipeline design, model fine-tuning, evaluation framework, and production inference pipeline.


For research and demonstration purposes.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Papers for EchoesML/Georgian-NeMo