Voice & Language

Low-resource language speech and translation data.

Speech, transcription, and translation data collected from real speakers — in languages where public corpora barely exist.

Data types

Read speechConversational / call audioInterview audioAccented speechTranscriptionTranslation pairs

Highlights

7.4B

translation pairs across rare-language coverage

4M

Burmese–English pairs

Call audio

Gulf Arabic, Spanish, Malaysian call-centre recordings

Accented

English and German from native Vietnamese speakers

Consent and licensing

Every speaker gives informed consent in their own language, with commercial AI training named explicitly. Call audio is collected on a two-party consent basis and delivered with personal information redacted. Data ships under a perpetual, worldwide, non-exclusive licence by default; exclusivity and per-domain terms are negotiable before collection begins.

Need a language not listed? We recruit and record.

Tell us the language, dialect, domain and hours you need. We build the speaker panel and collect to your spec.

Talk to our sourcing team