Paza
Speech Recognition for Low-Resource Languages
About Paza
Paza is a speech research initiative from Microsoft Research Africa, Nairobi — part of Project Gecko — focused on automatic speech recognition for low-resource languages that have largely been left behind by recent AI progress. It pairs a benchmark, a set of fine-tuned ASR models, and a practical playbook, and centers on identifying which fine-tuning and data-adaptation strategies most effectively adapt modern speech models to these languages. PazaBench, the first ASR benchmark for low-resource languages, evaluates 52 state-of-the-art models across 39 African languages, while the released models cover six Kenyan languages — Swahili, Dholuo, Kalenjin, Kikuyu, Maasai, and Somali.
Building speech systems for these languages is constrained less by model architecture than by data and adaptation strategy. Paza’s evaluation shows that encoder-only models run roughly 28× faster than multitask LLMs with competitive accuracy, that performance saturates around 150 to 300 hours of labelled data, and that adding just 20 percent in-domain data yields up to an 83 percent character-error-rate improvement for LLM-based models. Developed in partnership with the communities who use the technology, Paza ships its benchmark, models, and playbook openly so practitioners can reproduce and extend the work.
Key capabilities
- PazaBench — first ASR benchmark for low-resource languages, spanning 52 models and 39 African languages
- Fine-tuned ASR models covering six Kenyan languages, outperforming prior state-of-the-art systems
- Encoder-only models run roughly 28× faster than multitask LLMs with competitive accuracy
- Up to 83% character-error-rate improvement from adding 20% in-domain data
- Community playbook for collecting data, selecting models, fine-tuning, and evaluation
Ready to Explore?
Dive into platform integrations, source code, research papers, and announcements.