Burmese speech dataset
We hold transcribed Burmese speech from consented speakers and can record substantially more to order — read speech, conversational audio, and interview recordings with aligned Burmese-script transcripts, plus 4M curated Burmese–English translation pairs.
| Language | Burmese |
|---|---|
| Endonym | မြန်မာဘာသာ |
| ISO 639-3 | mya |
| Speaker population | ~43 million (33M native) |
| Primary regions | Myanmar; diaspora in Thailand, Malaysia, Singapore |
| Dialects covered | Standard (Yangon/Mandalay), Rakhine, Danu, Intha, Taungyo |
| Hours available / collectable | 1,200 hours available · up to 3,000 hours/quarter collectable |
| Recording spec | 48 kHz / 24-bit WAV, mono, close-mic and far-field pairs, <30 dB(A) noise floor |
| Transcript availability | Human transcription in Burmese script, utterance-aligned, 98%+ WER-checked |
| Consent basis | Written informed consent per speaker, commercial AI training explicitly named |
| License terms | Perpetual, worldwide, non-exclusive by default; exclusivity negotiable |
| Related languages | Thai, Vietnamese, Tibetan, Khmer, Hausa |
Why Burmese data is hard to find
Burmese uses a script with no whitespace word boundaries, which breaks most off-the-shelf transcription pipelines and makes scraped corpora nearly unusable. Public audio is dominated by broadcast news read by a handful of voices, so speaker diversity is thin. Political conditions in Myanmar have also pushed most usable material offline or behind closed community channels.
What we can collect
- —Read speech from prompt sets you supply, balanced across age, gender and region
- —Conversational and call-centre audio between consented participants
- —Interview and long-form spontaneous speech
- —Burmese–English translation pairs and parallel sentence data
Licensing and consent
Written informed consent per speaker, commercial AI training explicitly named. Delivery is under perpetual, worldwide, non-exclusive by default; exclusivity negotiable. Call and conversational audio is redacted for personal information before hand-off, and every recording carries a documented consent chain you can audit.
Need Burmese data at scale?
Send us hours, dialect mix and domain — we'll come back with a collection plan.
Talk to our sourcing team