Low-resource language speech and translation data.
Speech, transcription, and translation data collected from real speakers — in languages where public corpora barely exist.
Language coverage
Burmese
myaမြန်မာဘာသာ
~43 million (33M native)
View coverage →Hausa
hauHarshen Hausa
~80 million (L1 + L2)
View coverage →Zulu
zulisiZulu
~28 million (12M native)
View coverage →Tibetan
bodབོད་སྐད་
~6 million
View coverage →Haitian Creole
hatKreyòl Ayisyen
~12 million
View coverage →Gulf Arabic
afbخليجي
~36 million
View coverage →Vietnamese
vieTiếng Việt
~85 million
View coverage →Khmer
khmភាសាខ្មែរ
~17 million
View coverage →Thai
thaภาษาไทย
~61 million
View coverage →Malaysian Malay
zsmBahasa Melayu
~33 million
View coverage →Tagalog
tglWikang Tagalog
~82 million (incl. Filipino L2)
View coverage →Swahili
swhKiswahili
~200 million (L1 + L2)
View coverage →Yoruba
yorÈdè Yorùbá
~46 million
View coverage →Latin American Spanish
spaEspañol Latinoamericano
~450 million
View coverage →Data types
Highlights
translation pairs across rare-language coverage
Burmese–English pairs
Gulf Arabic, Spanish, Malaysian call-centre recordings
English and German from native Vietnamese speakers
Consent and licensing
Every speaker gives informed consent in their own language, with commercial AI training named explicitly. Call audio is collected on a two-party consent basis and delivered with personal information redacted. Data ships under a perpetual, worldwide, non-exclusive licence by default; exclusivity and per-domain terms are negotiable before collection begins.
Need a language not listed? We recruit and record.
Tell us the language, dialect, domain and hours you need. We build the speaker panel and collect to your spec.
Talk to our sourcing team