Data Menu
AmplLab records, cleans and annotates conversational speech for speech-to-speech, full-duplex and expressive TTS models. Voice actors perform every conversation on isolated, synchronized channels; trained annotators then transcribe each turn and label emotion, vocal bursts, overlaps and background noise, with automated QA/QC on top.
- Accepted audio
- 179h
- Fully annotated
- 142h
- Time-aligned turns
- 148K
- Labelled noise events
- 49K
Included Metadata
Datasets ship with file, annotation and speaker metadata, produced by AmplLab's recording, cleanup and QA/QC pipeline.
File Metadata
11 fields- Raw Audio (FLAC, lossless)
- Per-speaker Channel Tracks
- Merged Mix
merged-<take>.flac - Sample Rate
sampleRate - Bit Depth
bitDepth - Codec
codec - Channels
channels - Duration
take.duration - Session / Scene / Take IDs
scene.id, take.uid - Domain & Scenario
script.category - Speaker Roles
speakers.*.role
Annotation Metadata
11 fields- Speaker-attributed Transcript
transcript - Turn Timestamps
start_time, end_time - Emotion per Turn
emotion - Language per Turn
languages - Vocal Bursts & Fillers
[laughing] [um] - Overlap Spans
<overlap> - Speaking-style Spans
<whisper> - Voice Activity Segments
silero_vad - Noise Events
label, start, end - Speech vs Noise-floor Level
speakingRMS - QA/QC Checks
wpm, coverage
Speaker Metadata
7 fields- Gender
gender - Birth Year
birth_year - Native & Secondary Accent
native_accent - Primary & Other Languages
primary_language - Voice Description
voice - Recording Setup
home_studio - Voice-acting Experience
voice_actor_experience
Datasets
Quantities are a snapshot of accepted, annotated data as of 2026-10-01, and collection is ongoing. We also take bespoke requests for data that isn't listed.
Conversational Audio
Dual-Channel Conversations
Two-speaker conversations with each voice on its own isolated, time-synchronized channel, plus a merged mix of the take.
179 hours · 369 sessions
Domain Role-Play
Scripted agent-customer conversations across banking, insurance, travel, support, sales and negotiation, performed with natural pacing.
9 domains · 44 scenarios
Overlapping Speech
Interruptions, back-channels and cross-talk, with every overlap span marked in the transcript. Signal for full-duplex models.
125K overlap spans
Expressivity / Emotional Intelligence
Emotion-Annotated Transcripts
Speaker-attributed, time-aligned transcripts with a per-turn emotion label and language tag, run through automated QA/QC.
148K turns · 12.8K emotion-labelled
Vocal Bursts & Disfluencies
Laughs, sighs, gasps, coughs and fillers tagged inline exactly where they occur, plus whisper and elongation spans.
11K bursts · 69K fillers
Emotion-Directed Statements
Single-speaker statements performed under directed emotions (excitement, happiness, sadness and fear) for controllable TTS.
4 directed emotions · Pilot
Reliability / Voice & Acoustics
Voice & Speaker Profiles
Per-speaker accent, languages, age, gender, a self-described voice profile and recording setup, for voice design and balanced sampling.
49 voice actors
Noise-Event Annotations
Time-stamped labels for 15 classes of real-world noise, from mouth clicks and mic handling to traffic, on each speaker's track.
49K events · 3.2K tracks
Raw / Clean Pairs
Takes as recorded and after noise reduction, with the noise-only segments that were profiled. Paired data for speech enhancement.
5.5K noise profiles
Need data that isn't listed?
Tell us the domains, speakers, emotions and labels you need. We script, record, clean and annotate it through the same pipeline, with the same metadata.