← All languages

Latin American Spanish speech dataset

Latin American Spanish at volume, weighted toward genuine call-centre audio and regional varieties rather than the neutral broadcast Spanish that dominates public corpora.

LanguageLatin American Spanish
EndonymEspañol Latinoamericano
ISO 639-3spa
Speaker population~450 million
Primary regionsMexico, Colombia, Argentina, Peru, Chile, Central America
Dialects coveredMexican, Rioplatense, Andean, Caribbean, Chilean
Hours available / collectable3,000+ hours available (incl. call audio) · 5,000 hours/quarter collectable
Recording spec8 kHz telephony and 48 kHz studio, mono and dual-channel
Transcript availabilityStandard orthography with regionalism and code-switch annotation
Consent basisTwo-party consent for call audio; written consent for recorded sessions
License termsPerpetual, worldwide, non-exclusive; PII-redacted delivery standard
Related languagesHaitian Creole, Gulf Arabic, Tagalog, Swahili, Malaysian Malay

Why Latin American Spanish data is hard to find

"Spanish" corpora are usually neutral or Peninsular, and models trained on them degrade badly on Caribbean and Chilean speech. Consented, redacted call audio is the scarce part, not the language itself. Indigenous-contact varieties in the Andes are effectively unrepresented online.

What we can collect

  • Dual-channel call-centre audio across multiple countries, PII-redacted
  • Regional dialect read and conversational speech
  • Spanish–English code-switched border speech
  • Spanish–English translation pairs

Licensing and consent

Two-party consent for call audio; written consent for recorded sessions. Delivery is under perpetual, worldwide, non-exclusive; pii-redacted delivery standard. Call and conversational audio is redacted for personal information before hand-off, and every recording carries a documented consent chain you can audit.

Need Latin American Spanish data at scale?

Send us hours, dialect mix and domain — we'll come back with a collection plan.

Talk to our sourcing team