Speech data for the next generation of voice AI
Spontaneous, transcribed, and consented recordings from vetted native speakers. Ready to license today.
Speech data your models are missing
Every configuration is spontaneous pair conversation, recorded as audio plus video. Speaker profile, format, and rights are confirmed in your written proposal.
Spanish (LatAm)
Mexico · ConversationalGerman
Germany · ConversationalHindi
India · ConversationalJapanese
Japan · ConversationalArabic MSA (Modern)
Saudi Arabia · ConversationalBuilt for training, cleared for commercial use
Spontaneous, not scripted
Real two-person conversations with natural disfluencies, turn-taking, and code-switching. These are the acoustics your ASR model meets in production, not a reading booth.
Consent chain, documented
Every speaker signed an explicit commercial-use consent. We hand you the documentation, so GDPR and EU AI Act reviews don’t stall your procurement.
Transcripts & metadata included
Time-aligned, QA’d transcripts plus speaker age, gender, region, and recording-device metadata, ready for filtering, balancing, and evaluation splits.
Samples before you commit
Request free samples from any dataset and run your own audit: audio quality, transcript accuracy, speaker diversity. Judge the data, not the deck.
Transparent per-hour pricing
$60-$95 per hour depending on language, published on our catalogue dataset pages. Volume and multi-language discounts on request.
Custom collection on tap
Need a language, dialect, or domain we don’t stock? Our custom collection network records to your exact spec.
From high-demand markets to underserved languages
Tell us the speech data you need
Send your speaker profiles, languages, and recording conditions. We review the brief and reply with feasibility questions and next steps.
Request a quoteOr write to us at hello@spirelight.net