← All languages

Malaysian Malay speech dataset

Malaysian Malay speech including real call-centre audio, transcribed in Rumi with tagged code-switching into English, Mandarin and Tamil — the way the language is actually spoken.

LanguageMalaysian Malay
EndonymBahasa Melayu
ISO 639-3zsm
Speaker population~33 million
Primary regionsMalaysia, Singapore, Brunei
Dialects coveredStandard Malay, Kelantanese, Sarawak Malay, Manglish register
Hours available / collectable900 hours available (incl. call audio) · 2,000 hours/quarter collectable
Recording spec8 kHz telephony and 48 kHz studio, mono and dual-channel
Transcript availabilityRumi (Latin) orthography with code-switch tagging (Malay/English/Chinese/Tamil)
Consent basisTwo-party consent for call audio; written consent for recorded sessions
License termsPerpetual, worldwide, non-exclusive; PII-redacted delivery standard
Related languagesTagalog, Thai, Vietnamese, Gulf Arabic, Khmer

Why Malaysian Malay data is hard to find

Everyday Malaysian speech switches languages mid-sentence, so monolingual Malay corpora badly mismatch deployment. Regional varieties like Kelantanese differ enough to break standard-trained ASR. Consented, redacted call audio in this market is almost never available off the shelf.

What we can collect

  • Dual-channel call-centre audio with consent chain and PII redaction
  • Code-switched conversational speech with switch-point tags
  • Kelantanese and Sarawak dialect collection
  • Malay–English translation pairs

Licensing and consent

Two-party consent for call audio; written consent for recorded sessions. Delivery is under perpetual, worldwide, non-exclusive; pii-redacted delivery standard. Call and conversational audio is redacted for personal information before hand-off, and every recording carries a documented consent chain you can audit.

Need Malaysian Malay data at scale?

Send us hours, dialect mix and domain — we'll come back with a collection plan.

Talk to our sourcing team