Zulu speech dataset
Consented isiZulu speech with conjunctively written transcripts, including the heavy English code-switching typical of urban speakers — available now and collectable to spec across KwaZulu-Natal and Gauteng.
| Language | Zulu |
|---|---|
| Endonym | isiZulu |
| ISO 639-3 | zul |
| Speaker population | ~28 million (12M native) |
| Primary regions | KwaZulu-Natal and Gauteng, South Africa; Eswatini border regions |
| Dialects covered | Standard isiZulu, urban Johannesburg Zulu, rural KZN varieties |
| Hours available / collectable | 600 hours available · up to 1,800 hours/quarter collectable |
| Recording spec | 48 kHz / 24-bit WAV, mono, studio and in-home capture |
| Transcript availability | Standard isiZulu orthography, conjunctive writing, code-switch tagged |
| Consent basis | Written informed consent, POPIA-aligned processing notice |
| License terms | Perpetual, worldwide, non-exclusive; POPIA-compliant transfer terms |
| Related languages | Hausa, Swahili, Yoruba, Gulf Arabic, Tagalog |
Why Zulu data is hard to find
isiZulu is written conjunctively, so words are long agglutinative strings that inflate vocabulary and defeat models trained on disjunctive Nguni text. Real urban speech mixes English constantly, but public corpora are clean monolingual read text. Click consonants and tone are poorly represented in scraped material.
What we can collect
- —Urban code-switched conversational speech with switch-point tagging
- —Rural KZN read speech for dialect balance
- —Clinic and public-service domain audio
- —isiZulu–English translation pairs
Licensing and consent
Written informed consent, POPIA-aligned processing notice. Delivery is under perpetual, worldwide, non-exclusive; popia-compliant transfer terms. Call and conversational audio is redacted for personal information before hand-off, and every recording carries a documented consent chain you can audit.
Need Zulu data at scale?
Send us hours, dialect mix and domain — we'll come back with a collection plan.
Talk to our sourcing team