Conversational speech data · 60 languages

Speech data for the next generation of voice AI

Spontaneous, transcribed, and consented recordings from vetted native speakers. Ready to license today.

5.0 on Datarade 4.4 on Trustpilot SKI Verified Supplier Backed by
60 Languages & dialects
30+ Contributor markets
61 Dataset configurations
2,000h Largest configuration
Featured datasets

Speech data your models are missing

Every configuration is spontaneous pair conversation, recorded as audio plus video. Speaker profile, format, and rights are confirmed in your written proposal.

Why teams buy from us

Built for training, cleared for commercial use

01

Spontaneous, not scripted

Real two-person conversations with natural disfluencies, turn-taking, and code-switching. These are the acoustics your ASR model meets in production, not a reading booth.

02

Consent chain, documented

Every speaker signed an explicit commercial-use consent. We hand you the documentation, so GDPR and EU AI Act reviews don’t stall your procurement.

03

Transcripts & metadata included

Time-aligned, QA’d transcripts plus speaker age, gender, region, and recording-device metadata, ready for filtering, balancing, and evaluation splits.

04

Samples before you commit

Request free samples from any dataset and run your own audit: audio quality, transcript accuracy, speaker diversity. Judge the data, not the deck.

05

Transparent per-hour pricing

$60-$95 per hour depending on language, published on our catalogue dataset pages. Volume and multi-language discounts on request.

06

Custom collection on tap

Need a language, dialect, or domain we don’t stock? Our custom collection network records to your exact spec.

Tell us the speech data you need

Send your speaker profiles, languages, and recording conditions. We review the brief and reply with feasibility questions and next steps.

Request a quote