← All languages

Khmer speech dataset

Khmer speech with segmentation-marked Khmer-script transcripts, available now and collectable at scale across Phnom Penh and provincial varieties.

LanguageKhmer
Endonymភាសាខ្មែរ
ISO 639-3khm
Speaker population~17 million
Primary regionsCambodia; Thailand border provinces, Vietnam delta, diaspora
Dialects coveredCentral (Phnom Penh), Northern (Surin), Cardamom, Khmer Krom
Hours available / collectable350 hours available · up to 1,000 hours/quarter collectable
Recording spec48 kHz / 24-bit WAV, mono, studio and in-field
Transcript availabilityKhmer script with explicit word segmentation markers
Consent basisWritten informed consent, translated into Khmer
License termsPerpetual, worldwide, non-exclusive
Related languagesThai, Vietnamese, Burmese, Tibetan, Malaysian Malay

Why Khmer data is hard to find

Khmer script has no spaces between words and a complex stacked-consonant system, so segmentation must be annotated rather than inferred. The public corpus is tiny and dominated by news readers. Provincial and Khmer Krom varieties are effectively absent from anything scrapeable.

What we can collect

  • Read speech with prompt sets you define
  • Spontaneous provincial-dialect recordings
  • Agricultural, microfinance and health-domain speech
  • Khmer–English translation pairs

Licensing and consent

Written informed consent, translated into Khmer. Delivery is under perpetual, worldwide, non-exclusive. Call and conversational audio is redacted for personal information before hand-off, and every recording carries a documented consent chain you can audit.

Need Khmer data at scale?

Send us hours, dialect mix and domain — we'll come back with a collection plan.

Talk to our sourcing team