← All languages

Zulu speech dataset

Consented isiZulu speech with conjunctively written transcripts, including the heavy English code-switching typical of urban speakers — available now and collectable to spec across KwaZulu-Natal and Gauteng.

LanguageZulu
EndonymisiZulu
ISO 639-3zul
Speaker population~28 million (12M native)
Primary regionsKwaZulu-Natal and Gauteng, South Africa; Eswatini border regions
Dialects coveredStandard isiZulu, urban Johannesburg Zulu, rural KZN varieties
Hours available / collectable600 hours available · up to 1,800 hours/quarter collectable
Recording spec48 kHz / 24-bit WAV, mono, studio and in-home capture
Transcript availabilityStandard isiZulu orthography, conjunctive writing, code-switch tagged
Consent basisWritten informed consent, POPIA-aligned processing notice
License termsPerpetual, worldwide, non-exclusive; POPIA-compliant transfer terms
Related languagesHausa, Swahili, Yoruba, Gulf Arabic, Tagalog

Why Zulu data is hard to find

isiZulu is written conjunctively, so words are long agglutinative strings that inflate vocabulary and defeat models trained on disjunctive Nguni text. Real urban speech mixes English constantly, but public corpora are clean monolingual read text. Click consonants and tone are poorly represented in scraped material.

What we can collect

  • Urban code-switched conversational speech with switch-point tagging
  • Rural KZN read speech for dialect balance
  • Clinic and public-service domain audio
  • isiZulu–English translation pairs

Licensing and consent

Written informed consent, POPIA-aligned processing notice. Delivery is under perpetual, worldwide, non-exclusive; popia-compliant transfer terms. Call and conversational audio is redacted for personal information before hand-off, and every recording carries a documented consent chain you can audit.

Need Zulu data at scale?

Send us hours, dialect mix and domain — we'll come back with a collection plan.

Talk to our sourcing team