← All languages

Swahili speech dataset

Swahili speech spanning standard Kiunguja through urban Sheng and Congolese Swahili, with mobile-band captures matching the channels most East African deployments run on.

LanguageSwahili
EndonymKiswahili
ISO 639-3swh
Speaker population~200 million (L1 + L2)
Primary regionsTanzania, Kenya, Uganda, DRC, Rwanda, Mozambique
Dialects coveredKiunguja (standard), Kimvita, Sheng (urban Nairobi), Congolese Swahili
Hours available / collectable1,300 hours available · up to 3,000 hours/quarter collectable
Recording spec48 kHz / 24-bit WAV plus mobile and telephony bands, mono
Transcript availabilityStandard Latin orthography with Sheng lexicon annotation
Consent basisWritten informed consent, translated into Swahili
License termsPerpetual, worldwide, non-exclusive
Related languagesHausa, Yoruba, Zulu, Gulf Arabic, Haitian Creole

Why Swahili data is hard to find

Congolese Swahili and Nairobi Sheng diverge far enough from standard Kiunguja that a single model underperforms on both. Public corpora are dominated by broadcast and religious material. Urban slang shifts fast, so useful lexicons need recent, dated recordings rather than archives.

What we can collect

  • Sheng and urban conversational speech with dated lexicon annotation
  • Congolese Swahili field collection
  • Mobile-money, agriculture and health domain audio
  • Swahili–English translation pairs

Licensing and consent

Written informed consent, translated into Swahili. Delivery is under perpetual, worldwide, non-exclusive. Call and conversational audio is redacted for personal information before hand-off, and every recording carries a documented consent chain you can audit.

Need Swahili data at scale?

Send us hours, dialect mix and domain — we'll come back with a collection plan.

Talk to our sourcing team