---
pretty_name: "Hindi and Tamil full-duplex call-centre conversations"
language:
- hi
- ta
language_details: hi-IN, ta-IN
license: other
license_name: lokah-custom
license_link: https://kenpathlabs.com/lokah/licensing
task_categories:
- audio-to-audio
- automatic-speech-recognition
size_categories:
- 1K<n<10K
tags:
- Hindi speech dataset
- Tamil speech dataset
- conversational speech
- full-duplex dialogue data
- call centre audio
- dual-channel audio
---

# Hindi and Tamil full-duplex call-centre conversations

> Hindi and Tamil full-duplex conversation data: 1,265 two-channel call-centre calls with each speaker on a separate channel, overlaps, backchannels and turn timing preserved, transcripts time-aligned per channel. 264 hours.

`LK-SP-MUL-002` · [Get a quote](https://kenpathlabs.com/lokah/datasets/hindi-tamil-full-duplex-call-centre-conversations#contact) · [Record as JSON](https://kenpathlabs.com/api/lokah/datasets/hindi-tamil-full-duplex-call-centre-conversations) · [Croissant](https://kenpathlabs.com/lokah/datasets/hindi-tamil-full-duplex-call-centre-conversations/croissant.json)

Hindi and Tamil full-duplex conversation data: 1,265 two-channel call-centre calls with each speaker on a separate channel, overlaps, backchannels and turn timing preserved, transcripts time-aligned per channel. 264 hours. Measured from 6 sample conversations (25 minutes): 16 kHz, 16-bit FLAC, one two-channel file per conversation. Across the whole set 20% of transcript words are English written in Latin script and 20% of the time is silence; the sample conversations hold 278 turns. 11 distinct voices in the sample (6 male, 5 female). Every recording has a complete, segment-level transcript. Personal data: redacted. Licence: custom, quoted per use.

## Specification

| Field | Value | Note |
| --- | --- | --- |
| id | LK-SP-MUL-002 |  |
| type | speech · full-duplex calls · call centre |  |
| language | हिन्दी · Hindi · hi-IN / தமிழ் · Tamil · ta-IN |  |
| hours | 264 h |  |
| channels | Two channels, one per speaker |  |
| files | 1,265 files · 1,265 calls | counted across the full set |
| layouts | 1,265 mono conversation | counted across the full set |
| audio | FLAC · 16 kHz · 16-bit | measured across the full set |
| bandwidth | wideband (8 kHz) |  |
| snr | 29.2 dB median | across the full set |
| release | v1.0 |  |
| transcript | time-aligned by segment · Devanagari and Tamil script · delivered as JSON |  |
| speakers | 529 across the full set · id and gender per speaker |  |
| pii | Numbers of 4 or more digits, digit strings spoken as words, and email addresses are masked as [pii] in text and replaced by a tone in audio. First names are not masked. |  |
| source | Recorded for the dataset |  |
| review | Every recording has a complete, segment-level transcript |  |
| personal data | Redacted |  |
| licence | Custom | quoted per use |


## What it is good for

| Task | Fit | Why |
| --- | --- | --- |
| Speech recognition | partly | Per-channel transcripts are included; built for dialogue timing, not for utterance-level training. |
| Full-duplex speech to speech | yes | Built for it: 1,265 two-channel calls, each speaker on a separate channel, overlaps and timing intact. |
| Turn-taking | yes | Floor transfers, overlaps and backchannels are measurable from the two channels. |
| Voice agents | yes | Real agent and customer turns in a support flow, the condition a deployed agent hears. |
| Diarisation | partly | Two speakers, already separated by channel; useful as ground truth, not as a hard case. |
| Text to speech | no | Conversational call audio, not studio voice. |

## Conversation profile

Across every conversation in the dataset (1,265 conversations, median 12.45 minutes), as measured by the delivery.

| Measure | Value |
| --- | --- |
| Talk time, most-talking speaker / the other | 72% / 28% |
| Silence | 20% |
| Overlapping speech | 8% |
| English words, written in Latin script | 20% |
| Turns a minute | 7.38 |
| Speaking rate | 150.5 words a minute |

### Turn-taking, measured from the two channels across the set

| Per minute | This dataset | Fisher corpus |
| --- | --- | --- |
| Inter-pausal units | 27.36 | 21.6 |
| Pauses | 10.23 | 7 |
| Gaps | 7.09 | 7.5 |
| Overlaps | 3.8 | 6.5 |
| Backchannels | 5.86 | not reported |

Floor-transfer offset: median +0.19 s, 10th to 90th percentile -0.66 s to +1.38 s. Channel isolation: 45.2 dB.


## Audio quality

Frame RMS at 20 ms on the channels mixed to mono. Noise floor is the 10th percentile, speech level the 90th; SNR is their difference. Bandwidth is the highest frequency at which speech still rises 6 dB above the recording's own noise spectrum.

| Measure | Value | Reading |
| --- | --- | --- |
| SNR across every file of the set | 29.2 dB median |  |
| Speech above noise floor (SNR), sample conversations | 35 dB (29.8 to 48.3) | clean |
| Noise floor | -68.1 dBFS |  |
| Effective bandwidth | 8.0 kHz | wideband (8 kHz) across the full set |
| Clipping | 0.000% of samples |  |
| DNSMOS P.835 (1 to 5) | background 3.55, speech 3.17, overall 2.65 | listener-rated quality, estimated |

## Samples

4 public excerpts, 39 to 43 seconds each, from different conversations in the dataset.

- Sample 1, Two-channel call (Hindi), from a 4-minute conversation (Speaker 1: male; Speaker 2: male): [audio](https://kenpathlabs.com/lokah-samples/hindi-tamil-full-duplex-call-centre-conversations--1.m4a) · [segments](https://kenpathlabs.com/lokah-samples/hindi-tamil-full-duplex-call-centre-conversations--1.segments.json)
- Sample 2, Two-channel call (Hindi), from a 4-minute conversation (Speaker 1: male; Speaker 2: male): [audio](https://kenpathlabs.com/lokah-samples/hindi-tamil-full-duplex-call-centre-conversations--2.m4a) · [segments](https://kenpathlabs.com/lokah-samples/hindi-tamil-full-duplex-call-centre-conversations--2.segments.json)
- Sample 3, Two-channel call (Hindi), from a 4-minute conversation (Speaker 1: male; Speaker 2: female): [audio](https://kenpathlabs.com/lokah-samples/hindi-tamil-full-duplex-call-centre-conversations--3.m4a) · [segments](https://kenpathlabs.com/lokah-samples/hindi-tamil-full-duplex-call-centre-conversations--3.segments.json)
- Sample 4, Two-channel call (Hindi), from a 4-minute conversation (Speaker 1: female; Speaker 2: male): [audio](https://kenpathlabs.com/lokah-samples/hindi-tamil-full-duplex-call-centre-conversations--4.m4a) · [segments](https://kenpathlabs.com/lokah-samples/hindi-tamil-full-duplex-call-centre-conversations--4.segments.json)

### Sample 1 transcript

Two-channel call (Hindi), starting at 0:03. Stereo: left channel is speaker 1, right channel is speaker 2. Audio sha256 `0be133fcc76d86e7c37e671583aba4c7af90ce1d9f031790dc89f4e4209f37c4`.

| Start | Speaker | Text |
| --- | --- | --- |
| 0:00 | Speaker 1 | hello |
| 0:01 | Speaker 2 | hello |
| 0:03 | Speaker 1 | नमस्कार sir |
| 0:04 | Speaker 2 | नमश्कार. |
| 0:06 | Speaker 1 | sir आप company |
| 0:08 | Speaker 2 | जी बताइए आप कौन? |
| 0:11 | Speaker 1 | sir मैं अरविंदर सिंह बात कर रहा हूँ. |
| 0:15 | Speaker 2 | जी |
| 0:17 | Speaker 1 | OK |
| 0:20 | Speaker 2 | जी जरूर आप की ही सेवा में बैठे हैं. |
| 0:23 | Speaker 1 | actually मैंने अभी तक कोई policy ली नहीं है और कोई बीमा भी नहीं करवाया है. तो मुझे उसके बारे में जानकारी नहीं है तो sir मुझे थोड़ा आप विस्तार से उसके बारे में बता सके बीमा के. |
| 0:27 | Speaker 2 | जी |
| 0:34 | Speaker 2 | अच्छा जी जी जी. |
| 0:36 | Speaker 1 | जी sir actually क्या है कि बीमा मुझे तो दो तीन बीमे करवाने है. |
| 0:41 | Speaker 2 | अच्छा. |

## Licence

Licence: custom, quoted per use (training, evaluation or both; internal or commercial; exclusive or not). Delivered in the layout your training code reads: https://kenpathlabs.com/lokah/formats.

---

Machine-readable: [llms.txt](https://kenpathlabs.com/llms.txt) · [JSON API](https://kenpathlabs.com/api/lokah/datasets) · [OpenAPI](https://kenpathlabs.com/lokah/openapi.json) · [MCP](https://kenpathlabs.com/api/lokah/mcp) · [for agents](https://kenpathlabs.com/lokah/for-agents)
