Tibetan speech dataset
Tibetan speech across Ü-Tsang, Amdo and Kham with dual Tibetan-script and Wylie transcripts — a small standing corpus plus recruited-speaker collection through diaspora communities.
| Language | Tibetan |
|---|---|
| Endonym | བོད་སྐད་ |
| ISO 639-3 | bod |
| Speaker population | ~6 million |
| Primary regions | Tibetan Plateau; Nepal, Bhutan, northern India diaspora |
| Dialects covered | Ü-Tsang (Lhasa), Amdo, Kham |
| Hours available / collectable | 250 hours available · up to 700 hours/quarter collectable |
| Recording spec | 48 kHz / 24-bit WAV, mono, field and studio capture |
| Transcript availability | Tibetan script with Wylie transliteration alongside |
| Consent basis | Written informed consent with diaspora-community review |
| License terms | Perpetual, worldwide, non-exclusive; attribution optional |
| Related languages | Burmese, Khmer, Thai, Vietnamese, Hausa |
Why Tibetan data is hard to find
The three main Tibetan varieties are barely mutually intelligible, so a single pooled corpus trains a model that fits none of them. Tibetan script has a large syllable inventory and no reliable public transcription convention. Access to speakers inside the plateau is limited, which makes diaspora recruitment the only ethical route at scale.
What we can collect
- —Dialect-separated read speech for Ü-Tsang, Amdo and Kham
- —Interview and oral-history long-form audio
- —Religious and literary text read-aloud sets
- —Tibetan–English and Tibetan–Chinese translation pairs
Licensing and consent
Written informed consent with diaspora-community review. Delivery is under perpetual, worldwide, non-exclusive; attribution optional. Call and conversational audio is redacted for personal information before hand-off, and every recording carries a documented consent chain you can audit.
Need Tibetan data at scale?
Send us hours, dialect mix and domain — we'll come back with a collection plan.
Talk to our sourcing team