Yoruba speech dataset
Yoruba speech with fully diacritised, tone-verified transcripts across the major dialect groups — the marking that scraped Yoruba text almost never has.
| Language | Yoruba |
|---|---|
| Endonym | Èdè Yorùbá |
| ISO 639-3 | yor |
| Speaker population | ~46 million |
| Primary regions | Southwestern Nigeria, Benin, Togo; UK and US diaspora |
| Dialects covered | Ọ̀yọ́ (standard), Ìjẹ̀bú, Ẹ̀gbá, Èkìtì, Lagos urban |
| Hours available / collectable | 500 hours available · up to 1,500 hours/quarter collectable |
| Recording spec | 48 kHz / 24-bit WAV, mono, studio and in-field |
| Transcript availability | Full diacritic marking of tone and sub-dot vowels, human-verified |
| Consent basis | Written informed consent, translated into Yoruba |
| License terms | Perpetual, worldwide, non-exclusive |
| Related languages | Hausa, Swahili, Zulu, Haitian Creole, Burmese |
Why Yoruba data is hard to find
Yoruba is three-tone and its written form depends on diacritics that are routinely omitted online, making public text ambiguous at the word level. Dialects differ enough in tone patterns to affect recognition. Accurate transcription requires trained native annotators, not crowdworkers.
What we can collect
- —Tone-verified read speech across Ọ̀yọ́, Ìjẹ̀bú, Ẹ̀gbá and Èkìtì
- —Lagos urban conversational and code-switched speech
- —Market, health and civic domain audio
- —Yoruba–English translation pairs
Licensing and consent
Written informed consent, translated into Yoruba. Delivery is under perpetual, worldwide, non-exclusive. Call and conversational audio is redacted for personal information before hand-off, and every recording carries a documented consent chain you can audit.
Need Yoruba data at scale?
Send us hours, dialect mix and domain — we'll come back with a collection plan.
Talk to our sourcing team