Expressivity / Emotional Intelligence
Vocal Bursts & Disfluencies
Non-verbal sounds stay in the transcript instead of being stripped out. Vocal bursts like [laughing], [sigh] and [gasp], fillers like [um] and [hmm], and delivery spans like <whisper> and <elongated> are tagged at the exact position they occur, so models learn when they happen, not just how they sound.
Specification
- Bursts
- [laughing] [sigh] [gasp] [cough] [crying] [groan]
- Fillers
- [um] [uh] [hmm] [mm] [ah] [ohh]
- Style spans
- <whisper> <elongated>
- Placement
- Inline, at the exact word position
Delivery layout
- ampllab-speech-to-speech-batch-<timestamp>/
- session_<id>/
- scene_<n>/
- take-<n>/
- <speaker-a>.flac
- <speaker-b>.flac
- merged-<take>.flac
- annotation-<id>.json
Vocal bursts
- [laughing]
- [sigh]
- [gasp]
- [cough]
- [crying]
- [groan]
- [singing]
Fillers
- [um]
- [uh]
- [hmm]
- [mm]
- [ah]
- [ohh]
- [haan]
Style spans
- <whisper>...</whisper>
- <elongated>...</elongated>
take-<n>/annotation-<id>.jsonIllustrative values
{ "script": { "category": "Banking", "sub_category": "Card Servicing" }, "scene": { "id": "scene_2", "name": "Declined card at checkout" }, "take": { "uid": "7f3c9e21-...", "seq": 4, "duration": 287.42 }, "speakers": { "a41d07c2-...": { "role": "Priya (Card Services Agent)" }, "e9b2f6a8-...": { "role": "Daniel (Customer)" } }, "transcription": [ { "speaker": "e9b2f6a8-...", "start_time": 3.05, "end_time": 7.48, "transcript": "[sigh] Hi, my card got declined twice <overlap>this morning</overlap>.", "emotion": "Frustrated", "languages": "English" }, { "speaker": "a41d07c2-...", "start_time": 7.1, "end_time": 11.86, "transcript": "<overlap>Oh no,</overlap> let me take a look [mm]. Can you confirm the last four digits?", "emotion": "Calm", "languages": "English" } ]}Sample pack
A curated set of takes from this dataset, with audio and annotation files in the delivery format.
More in Expressivity / Emotional Intelligence
Emotion-Annotated Transcripts
Speaker-attributed, time-aligned transcripts with a per-turn emotion label and language tag, run through automated QA/QC.
148K turns · 12.8K emotion-labelled
Emotion-Directed Statements
Single-speaker statements performed under directed emotions (excitement, happiness, sadness and fear) for controllable TTS.
4 directed emotions · Pilot