Thai speech dataset
Thai speech including the Isan and southern varieties that public corpora ignore, transcribed in Thai script with explicit word boundaries and tone verification.
| Language | Thai |
|---|---|
| Endonym | ภาษาไทย |
| ISO 639-3 | tha |
| Speaker population | ~61 million |
| Primary regions | Thailand; northeastern Isan region, southern provinces |
| Dialects covered | Central (standard), Isan, Northern (Lanna), Southern |
| Hours available / collectable | 1,100 hours available · up to 2,500 hours/quarter collectable |
| Recording spec | 48 kHz / 24-bit WAV, mono, plus telephony-band variants |
| Transcript availability | Thai script with word-boundary annotation and tone verification |
| Consent basis | Written informed consent under PDPA-aligned notice |
| License terms | Perpetual, worldwide, non-exclusive; PDPA-compliant transfer terms |
| Related languages | Khmer, Vietnamese, Burmese, Malaysian Malay, Tibetan |
Why Thai data is hard to find
Thai is written without spaces and is tonal, so both segmentation and tone need human annotation to be trustworthy. Nearly all public data is Central Thai, while a third of the country speaks Isan at home. Telephony-band Thai data with consent is scarce.
What we can collect
- —Dialect-balanced speech across Central, Isan, Northern and Southern varieties
- —Call-centre and telephony-band audio
- —Retail and logistics domain conversational speech
- —Thai–English translation pairs
Licensing and consent
Written informed consent under PDPA-aligned notice. Delivery is under perpetual, worldwide, non-exclusive; pdpa-compliant transfer terms. Call and conversational audio is redacted for personal information before hand-off, and every recording carries a documented consent chain you can audit.
Need Thai data at scale?
Send us hours, dialect mix and domain — we'll come back with a collection plan.
Talk to our sourcing team