Expressivity / Emotional Intelligence
Emotion-Annotated Transcripts
Every turn carries its speaker, start and end time, verbatim transcript, an emotion label and its language. Labels come from trained human annotators, and every file then goes through automated QA/QC: spelling, speech-rate outliers and transcript-versus-voice-activity coverage.
Specification
- Granularity
- Per turn
- Fields
- speaker, start_time, end_time, transcript, emotion, languages
- Emotion classes
- 19 + neutral
- QA/QC
- Spelling, speech rate, VAD coverage
- Annotated audio
- 142 hours
Delivery layout
- ampllab-speech-to-speech-batch-<timestamp>/
- session_<id>/
- scene_<n>/
- take-<n>/
- <speaker-a>.flac
- <speaker-b>.flac
- merged-<take>.flac
- annotation-<id>.json
Emotion labels
- Happy
- Angry
- Sad
- Fear
- Excited
- Frustrated
- Surprise
- Worried
- Sarcastic
- Curious
- Anxious
- Enthusiastic
- Calm
- Amazed
- Nervous
- Disgust
- Annoyed
- Irritated
- Exhausted
take-<n>/annotation-<id>.jsonIllustrative values
{ "script": { "category": "Banking", "sub_category": "Card Servicing" }, "scene": { "id": "scene_2", "name": "Declined card at checkout" }, "take": { "uid": "7f3c9e21-...", "seq": 4, "duration": 287.42 }, "speakers": { "a41d07c2-...": { "role": "Priya (Card Services Agent)" }, "e9b2f6a8-...": { "role": "Daniel (Customer)" } }, "transcription": [ { "speaker": "e9b2f6a8-...", "start_time": 3.05, "end_time": 7.48, "transcript": "[sigh] Hi, my card got declined twice <overlap>this morning</overlap>.", "emotion": "Frustrated", "languages": "English" }, { "speaker": "a41d07c2-...", "start_time": 7.1, "end_time": 11.86, "transcript": "<overlap>Oh no,</overlap> let me take a look [mm]. Can you confirm the last four digits?", "emotion": "Calm", "languages": "English" } ]}Sample pack
A curated set of takes from this dataset, with audio and annotation files in the delivery format.
More in Expressivity / Emotional Intelligence
Vocal Bursts & Disfluencies
Laughs, sighs, gasps, coughs and fillers tagged inline exactly where they occur, plus whisper and elongation spans.
11K bursts · 69K fillers
Emotion-Directed Statements
Single-speaker statements performed under directed emotions (excitement, happiness, sadness and fear) for controllable TTS.
4 directed emotions · Pilot