Khmer speech dataset
Khmer speech with segmentation-marked Khmer-script transcripts, available now and collectable at scale across Phnom Penh and provincial varieties.
| Language | Khmer |
|---|---|
| Endonym | ភាសាខ្មែរ |
| ISO 639-3 | khm |
| Speaker population | ~17 million |
| Primary regions | Cambodia; Thailand border provinces, Vietnam delta, diaspora |
| Dialects covered | Central (Phnom Penh), Northern (Surin), Cardamom, Khmer Krom |
| Hours available / collectable | 350 hours available · up to 1,000 hours/quarter collectable |
| Recording spec | 48 kHz / 24-bit WAV, mono, studio and in-field |
| Transcript availability | Khmer script with explicit word segmentation markers |
| Consent basis | Written informed consent, translated into Khmer |
| License terms | Perpetual, worldwide, non-exclusive |
| Related languages | Thai, Vietnamese, Burmese, Tibetan, Malaysian Malay |
Why Khmer data is hard to find
Khmer script has no spaces between words and a complex stacked-consonant system, so segmentation must be annotated rather than inferred. The public corpus is tiny and dominated by news readers. Provincial and Khmer Krom varieties are effectively absent from anything scrapeable.
What we can collect
- —Read speech with prompt sets you define
- —Spontaneous provincial-dialect recordings
- —Agricultural, microfinance and health-domain speech
- —Khmer–English translation pairs
Licensing and consent
Written informed consent, translated into Khmer. Delivery is under perpetual, worldwide, non-exclusive. Call and conversational audio is redacted for personal information before hand-off, and every recording carries a documented consent chain you can audit.
Need Khmer data at scale?
Send us hours, dialect mix and domain — we'll come back with a collection plan.
Talk to our sourcing team