AmplLab/samples
Expressivity / Emotional Intelligence

Vocal Bursts & Disfluencies

Non-verbal sounds stay in the transcript instead of being stripped out. Vocal bursts like [laughing], [sigh] and [gasp], fillers like [um] and [hmm], and delivery spans like <whisper> and <elongated> are tagged at the exact position they occur, so models learn when they happen, not just how they sound.

11Kbursts
69Kfillers
Request samples

Specification

Bursts
[laughing] [sigh] [gasp] [cough] [crying] [groan]
Fillers
[um] [uh] [hmm] [mm] [ah] [ohh]
Style spans
<whisper> <elongated>
Placement
Inline, at the exact word position

Delivery layout

  • ampllab-speech-to-speech-batch-<timestamp>/
  • session_<id>/
  • scene_<n>/
  • take-<n>/
  • <speaker-a>.flac
  • <speaker-b>.flac
  • merged-<take>.flac
  • annotation-<id>.json

Vocal bursts

  • [laughing]
  • [sigh]
  • [gasp]
  • [cough]
  • [crying]
  • [groan]
  • [singing]

Fillers

  • [um]
  • [uh]
  • [hmm]
  • [mm]
  • [ah]
  • [ohh]
  • [haan]

Style spans

  • <whisper>...</whisper>
  • <elongated>...</elongated>
take-<n>/annotation-<id>.jsonIllustrative values
{  "script": {    "category": "Banking",    "sub_category": "Card Servicing"  },  "scene": {    "id": "scene_2",    "name": "Declined card at checkout"  },  "take": {    "uid": "7f3c9e21-...",    "seq": 4,    "duration": 287.42  },  "speakers": {    "a41d07c2-...": {      "role": "Priya (Card Services Agent)"    },    "e9b2f6a8-...": {      "role": "Daniel (Customer)"    }  },  "transcription": [    {      "speaker": "e9b2f6a8-...",      "start_time": 3.05,      "end_time": 7.48,      "transcript": "[sigh] Hi, my card got declined twice <overlap>this morning</overlap>.",      "emotion": "Frustrated",      "languages": "English"    },    {      "speaker": "a41d07c2-...",      "start_time": 7.1,      "end_time": 11.86,      "transcript": "<overlap>Oh no,</overlap> let me take a look [mm]. Can you confirm the last four digits?",      "emotion": "Calm",      "languages": "English"    }  ]}

Sample pack

A curated set of takes from this dataset, with audio and annotation files in the delivery format.

Request samples

More in Expressivity / Emotional Intelligence