AmplLab/samples
Conversational Audio

Dual-Channel Conversations

Every take is captured as one isolated mono track per speaker, recorded simultaneously in separate rooms, with a merged mix alongside. There is no bleed between speakers and no diarization guesswork, so each channel can be used as-is for speech-to-speech, turn-taking and full-duplex training.

179hours
369sessions
Request samples

Specification

Speakers
2 per conversation
Channels
1 isolated mono track per speaker + merged mix
Audio
FLAC, 24-bit, 48 kHz typical (44.1-96 kHz)
Language
English
Recording
Remote sessions, one microphone per speaker
Structure
Session / scene / take

Delivery layout

  • ampllab-speech-to-speech-batch-<timestamp>/
  • session_<id>/
  • scene_<n>/
  • take-<n>/
  • <speaker-a>.flac
  • <speaker-b>.flac
  • merged-<take>.flac
  • annotation-<id>.json
take-<n>/annotation-<id>.jsonIllustrative values
{  "script": {    "category": "Banking",    "sub_category": "Card Servicing"  },  "scene": {    "id": "scene_2",    "name": "Declined card at checkout"  },  "take": {    "uid": "7f3c9e21-...",    "seq": 4,    "duration": 287.42  },  "speakers": {    "a41d07c2-...": {      "role": "Priya (Card Services Agent)"    },    "e9b2f6a8-...": {      "role": "Daniel (Customer)"    }  },  "transcription": [    {      "speaker": "e9b2f6a8-...",      "start_time": 3.05,      "end_time": 7.48,      "transcript": "[sigh] Hi, my card got declined twice <overlap>this morning</overlap>.",      "emotion": "Frustrated",      "languages": "English"    },    {      "speaker": "a41d07c2-...",      "start_time": 7.1,      "end_time": 11.86,      "transcript": "<overlap>Oh no,</overlap> let me take a look [mm]. Can you confirm the last four digits?",      "emotion": "Calm",      "languages": "English"    }  ]}

Sample pack

A curated set of takes from this dataset, with audio and annotation files in the delivery format.

Request samples

More in Conversational Audio