← All languages

Vietnamese speech dataset

Vietnamese speech across all three regional varieties with fully diacritised transcripts — plus accented English and German recorded from native Vietnamese speakers for accent-robust ASR.

LanguageVietnamese
EndonymTiếng Việt
ISO 639-3vie
Speaker population~85 million
Primary regionsVietnam; diaspora in the US, Germany, Czechia, Japan
Dialects coveredNorthern (Hanoi), Central (Huế), Southern (Ho Chi Minh City)
Hours available / collectable2,000 hours available · up to 4,000 hours/quarter collectable
Recording spec48 kHz / 24-bit WAV, mono, plus accented-English and German sessions
Transcript availabilityQuốc ngữ with full diacritics, tone-verified
Consent basisWritten informed consent, commercial AI training explicitly named
License termsPerpetual, worldwide, non-exclusive; exclusivity negotiable
Related languagesThai, Khmer, Burmese, Malaysian Malay, Tagalog

Why Vietnamese data is hard to find

Vietnamese is six-tone in the north and five in the south, and public corpora rarely tag which variety a speaker uses. Diacritics are routinely dropped in online text, so scraped transcripts are tonally ambiguous. Accented-English data from Vietnamese speakers barely exists publicly despite being what many deployments actually encounter.

What we can collect

  • Region-tagged read and conversational speech (North, Central, South)
  • Accented English and German from native Vietnamese speakers
  • Call-centre and customer-service audio
  • Vietnamese–English translation pairs

Licensing and consent

Written informed consent, commercial AI training explicitly named. Delivery is under perpetual, worldwide, non-exclusive; exclusivity negotiable. Call and conversational audio is redacted for personal information before hand-off, and every recording carries a documented consent chain you can audit.

Need Vietnamese data at scale?

Send us hours, dialect mix and domain — we'll come back with a collection plan.

Talk to our sourcing team