Tagalog speech dataset
Tagalog and Taglish speech including BPO-style call audio, transcribed with explicit code-switch tagging and available in both studio and telephony bands.
| Language | Tagalog |
|---|---|
| Endonym | Wikang Tagalog |
| ISO 639-3 | tgl |
| Speaker population | ~82 million (incl. Filipino L2) |
| Primary regions | Philippines (Luzon, Metro Manila); global diaspora |
| Dialects covered | Manila standard, Batangas, Bulacan, Taglish register |
| Hours available / collectable | 1,000 hours available · up to 2,500 hours/quarter collectable |
| Recording spec | 48 kHz / 24-bit WAV and 8 kHz telephony, mono and dual-channel |
| Transcript availability | Latin orthography with Taglish code-switch tagging |
| Consent basis | Written informed consent; two-party consent for call audio |
| License terms | Perpetual, worldwide, non-exclusive; PII-redacted delivery standard |
| Related languages | Malaysian Malay, Vietnamese, Thai, Zulu, Haitian Creole |
Why Tagalog data is hard to find
Real Tagalog is Taglish, and corpora that strip English out no longer resemble anything a model will hear. Provincial varieties such as Batangas differ sharply from Manila standard. The Philippines' large BPO sector produces enormous call audio, almost none of it licensable with a clean consent chain.
What we can collect
- —Taglish conversational and call-centre audio with consent and redaction
- —Provincial dialect read speech
- —Accented English from native Tagalog speakers
- —Tagalog–English translation pairs
Licensing and consent
Written informed consent; two-party consent for call audio. Delivery is under perpetual, worldwide, non-exclusive; pii-redacted delivery standard. Call and conversational audio is redacted for personal information before hand-off, and every recording carries a documented consent chain you can audit.
Need Tagalog data at scale?
Send us hours, dialect mix and domain — we'll come back with a collection plan.
Talk to our sourcing team