← All languages

Yoruba speech dataset

Yoruba speech with fully diacritised, tone-verified transcripts across the major dialect groups — the marking that scraped Yoruba text almost never has.

LanguageYoruba
EndonymÈdè Yorùbá
ISO 639-3yor
Speaker population~46 million
Primary regionsSouthwestern Nigeria, Benin, Togo; UK and US diaspora
Dialects coveredỌ̀yọ́ (standard), Ìjẹ̀bú, Ẹ̀gbá, Èkìtì, Lagos urban
Hours available / collectable500 hours available · up to 1,500 hours/quarter collectable
Recording spec48 kHz / 24-bit WAV, mono, studio and in-field
Transcript availabilityFull diacritic marking of tone and sub-dot vowels, human-verified
Consent basisWritten informed consent, translated into Yoruba
License termsPerpetual, worldwide, non-exclusive
Related languagesHausa, Swahili, Zulu, Haitian Creole, Burmese

Why Yoruba data is hard to find

Yoruba is three-tone and its written form depends on diacritics that are routinely omitted online, making public text ambiguous at the word level. Dialects differ enough in tone patterns to affect recognition. Accurate transcription requires trained native annotators, not crowdworkers.

What we can collect

  • Tone-verified read speech across Ọ̀yọ́, Ìjẹ̀bú, Ẹ̀gbá and Èkìtì
  • Lagos urban conversational and code-switched speech
  • Market, health and civic domain audio
  • Yoruba–English translation pairs

Licensing and consent

Written informed consent, translated into Yoruba. Delivery is under perpetual, worldwide, non-exclusive. Call and conversational audio is redacted for personal information before hand-off, and every recording carries a documented consent chain you can audit.

Need Yoruba data at scale?

Send us hours, dialect mix and domain — we'll come back with a collection plan.

Talk to our sourcing team