Swahili speech dataset
Swahili speech spanning standard Kiunguja through urban Sheng and Congolese Swahili, with mobile-band captures matching the channels most East African deployments run on.
| Language | Swahili |
|---|---|
| Endonym | Kiswahili |
| ISO 639-3 | swh |
| Speaker population | ~200 million (L1 + L2) |
| Primary regions | Tanzania, Kenya, Uganda, DRC, Rwanda, Mozambique |
| Dialects covered | Kiunguja (standard), Kimvita, Sheng (urban Nairobi), Congolese Swahili |
| Hours available / collectable | 1,300 hours available · up to 3,000 hours/quarter collectable |
| Recording spec | 48 kHz / 24-bit WAV plus mobile and telephony bands, mono |
| Transcript availability | Standard Latin orthography with Sheng lexicon annotation |
| Consent basis | Written informed consent, translated into Swahili |
| License terms | Perpetual, worldwide, non-exclusive |
| Related languages | Hausa, Yoruba, Zulu, Gulf Arabic, Haitian Creole |
Why Swahili data is hard to find
Congolese Swahili and Nairobi Sheng diverge far enough from standard Kiunguja that a single model underperforms on both. Public corpora are dominated by broadcast and religious material. Urban slang shifts fast, so useful lexicons need recent, dated recordings rather than archives.
What we can collect
- —Sheng and urban conversational speech with dated lexicon annotation
- —Congolese Swahili field collection
- —Mobile-money, agriculture and health domain audio
- —Swahili–English translation pairs
Licensing and consent
Written informed consent, translated into Swahili. Delivery is under perpetual, worldwide, non-exclusive. Call and conversational audio is redacted for personal information before hand-off, and every recording carries a documented consent chain you can audit.
Need Swahili data at scale?
Send us hours, dialect mix and domain — we'll come back with a collection plan.
Talk to our sourcing team