Kenpath Labs
Get started

অসমীয়াবাংলাEnglishગુજરાતીहिन्दीಕನ್ನಡമലയാളംमराठीଓଡ଼ିଆਪੰਜਾਬੀதமிழ்తెలుగుاردو

NewSample release · recorded to order

Indian-language calls between people and an AI voice agent

Phone calls between a person and an AI voice agent in thirteen Indian languages, on recruitment, telecom, healthcare, e-commerce and insurance, each voice on its own channel at 48 kHz. A sample release: 27 calls, 100 minutes, with verbatim transcripts of both sides. Recorded to order in your languages, domains and scenarios, at any volume.

27 calls · 100 min
13 languages
Two channels
Transcribed
48 kHz · 16-bit FLAC
Transcribed in full
Fit for
  • ASR
  • Full duplex
  • Turn-taking
  • Voice agents
  • Diarisation
  • TTS

The sample, measured.

Figures measured from every call in the sample release. A collection recorded to order is measured the same way before delivery.

of audio
100 min
27 conversations
audio files
27
FLAC · 48 kHz · 16-bit
transcribed words
12,098
683 turns
two channels
27
bandwidth
wideband
15 of 27 files measured
median SNR
38.1 dB
clean

Personal data. Numbers of 4 or more digits, digit strings spoken as words, and email addresses are masked as [pii] in text and replaced by a tone in audio. First names are not masked.

Built to order.

A sample of a collection recorded to order. The full set is built to your brief. What you hear on this page is the spec: the same channels, sample rate and transcripts, in the languages, domains and scenarios you name, at the volume you need.

In the sample
13 languages: Assamese, Bengali, Indian English, Gujarati, Hindi, Kannada, Malayalam, Marathi, Odia, Punjabi, Tamil, Telugu, Urdu
Languages next
Yours on request
Layout
Each speaker on a separate channel
Audio
48 kHz · 16-bit
  • Recorded to order in the languages, domains and scenarios you need, at the same spec, and scales to large volumes
  • Each call is recorded on two channels, the person on channel 1 and the AI voice agent on channel 2
  • 48 kHz, 16-bit FLAC master
  • Verbatim transcripts of both sides, in each language's own script, with English words as spoken

Sample 1 of 26 · Odia

Customer service: mobile network complaint

1:29 from a 4-minute conversation, at 1:01

Caller

Female

AI voice agent

Female · synthetic voice

0:00 / 1:29
Channel 1 · CallerChannel 2 · AI voice agent

Speakers

Across the whole dataset

21 · 5 M · 16 F
distinct callers
1 · F
AI agent voice

Callers are people from our contributor network, counted by contributor as the collection team recorded them. The agent is one synthetic voice on every call.

Voices in the sample

In the transcripts

12,098 words in 683 turns

  • Bengali-Assamese and Bengali and Latin and Gujarati and Devanagari and Kannada and Malayalam and Odia and Gurmukhi and Tamil and Telugu and Perso-Arabic script, English as spoken

    Words are written in the script the speaker would use; English words stay in Latin script where they were said, so code-mixing is preserved as spoken.

  • Aligned per segment, speakers labelled

    One segment per turn with its start and end time and the speaker, no word-level timestamps, on the channel that speaker was recorded to. Overlaps are kept as overlapping segments.

  • One tag vocabulary, in square brackets

    Anything that is not a spoken word is a bracket tag from a single list: [pii] for masked personal data, [filler] for hesitations, [overlap] where both speak at once. Plain-text fields carry no tags.

  • Personal data masked, in text and audio

    Numbers of 4 or more digits, digit strings spoken as words, and email addresses are masked as [pii] in text and replaced by a tone in audio. First names are not masked.

Delivery layouts are on formats.

Conversation profile

Across every conversation in the dataset

Talk time60% / 40%
  • Caller 60%
  • AI voice agent 40%
Speech, overlap and silence74% / 2% / 24%
  • Speech 74%
  • Overlap 2%
  • Silence 24%
Transcript words in Latin script16%
  • Latin script 16%
  • Native script 84%
Gap at each change of speakermedian +1.36 s

Across 27 two-channel calls. The middle 80% of changes fall between -0.16 s and +2.25 s; a negative gap means the next speaker came in early.

Turn-taking events a minute across the set, beside the Fisher corpus

Turn-taking events per minute, this dataset beside the Fisher corpus
EventThis datasetFisher
Inter-pausal units18.521.6
Pauses9.37.0
Gaps6.17.5
Overlaps0.96.5
Backchannels1.8not reported
  • This dataset
  • Fisher corpus

Channel isolation: 60 dB. How much of the other speaker bleeds into each channel. Lower is cleaner.

3.7 min
median conversation
27
conversations
6.4
turns a minute
120
words a minute

Audio quality

Across every call in the sample release.

Across all 27 files38.1 dB median SNR
noisysome backgroundclean

Clean across the set. Measured on the caller's channel; the agent's channel is synthetic and silent between turns. Bandwidth wideband (8 kHz).

Effective bandwidth8.0 kHz
0telephone band8 kHz

Wideband or better. This reading is taken on a 16 kHz copy, so it stops at 8 kHz; measured file by file, the set holds 15 wideband, 11 super-wideband, 1 fullband.

Listener-rated quality, estimatedDNSMOS P.835
  • Background4.01
  • Speech3.54
  • Overall3.24

Clipping in 0.00% of the audio.

Start with a sample. License the full set when it fits.

Recorded to order

Request a collection

Name the languages, domains and scenarios, and the hours you need. The collection is recorded to the spec you hear in the sample, at any volume, and delivered in stages.

  • Same spec as the sample: two channels, 48 kHz, verbatim transcripts of both sides
  • Your languages, domains and call scenarios
  • Licensed per use, priced per collection, delivered in stages

Sample first

Get a sample by email

Ten to thirty minutes of this dataset, with transcripts, in the same files and fields as the full delivery. The link works for 24 hours.

  • Real recordings from this dataset
  • Same layout, naming and fields as the full set
  • For evaluation only

What it is good for.

  • Speech recognitionFitsTime-aligned transcripts in native script, English kept as spoken. 38.1 dB median SNR across the full set, wideband (44.3 dB in the sample conversations).
  • Full-duplex speech to speechFitsEach speaker on a separate channel, so overlap, backchannels and turn timing survive. Moshi and PersonaPlex train on exactly this layout.
  • Turn-taking and endpointingFitsGaps and overlaps at every change of speaker are measured from the two channels; see the profile.
  • Voice agents for supportFitsCaller and agent turns across support and recruitment scenarios, with the agent on its own channel.
  • Speaker diarisationFitsSpeaker-attributed segments across full conversations.
  • Text to speechPartly38.1 dB median SNR across the full set, wideband (44.3 dB in the sample conversations): clean and wideband enough for conversational prosody data, though not a studio voice.

Specification

Figures marked measured were read from the audio. Anything we cannot confirm is listed under Ask us about.

id
LK-SP-MUL-001
type
speech · conversational · call centre
language
অসমীয়া · Assamese · as-IN / বাংলা · Bengali · bn-IN / English · Indian English · en-IN / ગુજરાતી · Gujarati · gu-IN / हिन्दी · Hindi · hi-IN / ಕನ್ನಡ · Kannada · kn-IN / മലയാളം · Malayalam · ml-IN / मराठी · Marathi · mr-IN / ଓଡ଼ିଆ · Odia · or-IN / ਪੰਜਾਬੀ · Punjabi · pa-IN / தமிழ் · Tamil · ta-IN / తెలుగు · Telugu · te-IN / اردو · Urdu · ur-IN
hours
2 h
channels
Two channels, one per speaker
files
27 files · 27 conversationscounted across the sample release
layouts
27 two-channelcounted across the sample release
audio
FLAC · 48 kHz · 16-bitmeasured across the sample release
bandwidth
wideband (8 kHz)
snr
38.1 dB medianacross the sample release, on the caller's channel
release
sample releasethe full collection is recorded to order
transcript
time-aligned by segment · Bengali-Assamese and Bengali and Latin and Gujarati and Devanagari and Kannada and Malayalam and Odia and Gurmukhi and Tamil and Telugu and Perso-Arabic script · delivered as JSON
speakers
21 callers and one AI voice agent (female synthetic voice)
pii
Numbers of 4 or more digits, digit strings spoken as words, and email addresses are masked as [pii] in text and replaced by a tone in audio. First names are not masked.
source
Recorded to order with an AI voice agent
review
Every recording has a complete, segment-level transcript
personal data
Redacted
licence
Customquoted per use

Not exactly what you need?

A different domain, more hours, another channel layout or speaker mix. Tell us, and it becomes a collection built to the same specification.

Scope a collection

Related datasets

Talk to us

Start a collection.

Tell us what you want to collect and from whom. We reply with how we would run it.

  • Speech, images, documents, feedback or annotation
  • A new collection, or a dataset from the catalogue
  • Licensing and data handling

Prefer to reach out directly? Write to hello@kenpathlabs.com.