← All languages

Thai speech dataset

Thai speech including the Isan and southern varieties that public corpora ignore, transcribed in Thai script with explicit word boundaries and tone verification.

LanguageThai
Endonymภาษาไทย
ISO 639-3tha
Speaker population~61 million
Primary regionsThailand; northeastern Isan region, southern provinces
Dialects coveredCentral (standard), Isan, Northern (Lanna), Southern
Hours available / collectable1,100 hours available · up to 2,500 hours/quarter collectable
Recording spec48 kHz / 24-bit WAV, mono, plus telephony-band variants
Transcript availabilityThai script with word-boundary annotation and tone verification
Consent basisWritten informed consent under PDPA-aligned notice
License termsPerpetual, worldwide, non-exclusive; PDPA-compliant transfer terms
Related languagesKhmer, Vietnamese, Burmese, Malaysian Malay, Tibetan

Why Thai data is hard to find

Thai is written without spaces and is tonal, so both segmentation and tone need human annotation to be trustworthy. Nearly all public data is Central Thai, while a third of the country speaks Isan at home. Telephony-band Thai data with consent is scarce.

What we can collect

  • Dialect-balanced speech across Central, Isan, Northern and Southern varieties
  • Call-centre and telephony-band audio
  • Retail and logistics domain conversational speech
  • Thai–English translation pairs

Licensing and consent

Written informed consent under PDPA-aligned notice. Delivery is under perpetual, worldwide, non-exclusive; pdpa-compliant transfer terms. Call and conversational audio is redacted for personal information before hand-off, and every recording carries a documented consent chain you can audit.

Need Thai data at scale?

Send us hours, dialect mix and domain — we'll come back with a collection plan.

Talk to our sourcing team