← All languages

Tibetan speech dataset

Tibetan speech across Ü-Tsang, Amdo and Kham with dual Tibetan-script and Wylie transcripts — a small standing corpus plus recruited-speaker collection through diaspora communities.

LanguageTibetan
Endonymབོད་སྐད་
ISO 639-3bod
Speaker population~6 million
Primary regionsTibetan Plateau; Nepal, Bhutan, northern India diaspora
Dialects coveredÜ-Tsang (Lhasa), Amdo, Kham
Hours available / collectable250 hours available · up to 700 hours/quarter collectable
Recording spec48 kHz / 24-bit WAV, mono, field and studio capture
Transcript availabilityTibetan script with Wylie transliteration alongside
Consent basisWritten informed consent with diaspora-community review
License termsPerpetual, worldwide, non-exclusive; attribution optional
Related languagesBurmese, Khmer, Thai, Vietnamese, Hausa

Why Tibetan data is hard to find

The three main Tibetan varieties are barely mutually intelligible, so a single pooled corpus trains a model that fits none of them. Tibetan script has a large syllable inventory and no reliable public transcription convention. Access to speakers inside the plateau is limited, which makes diaspora recruitment the only ethical route at scale.

What we can collect

  • Dialect-separated read speech for Ü-Tsang, Amdo and Kham
  • Interview and oral-history long-form audio
  • Religious and literary text read-aloud sets
  • Tibetan–English and Tibetan–Chinese translation pairs

Licensing and consent

Written informed consent with diaspora-community review. Delivery is under perpetual, worldwide, non-exclusive; attribution optional. Call and conversational audio is redacted for personal information before hand-off, and every recording carries a documented consent chain you can audit.

Need Tibetan data at scale?

Send us hours, dialect mix and domain — we'll come back with a collection plan.

Talk to our sourcing team