Speech AI Engineer
Kenpath Labs is a frontier AI and data company in Bengaluru. We build speech models, and we collect and license human data for AI, in the languages most systems skip.
Our open model, Svara TTS v1, has passed a million downloads. Our current model, Svara TTS Turbo, speaks 82 languages in hundreds of voices, switches language mid-sentence, and returns first audio in about 80 ms. It runs today behind a self-serve API at platform.kenpathlabs.com.
Lokah, our data platform, licenses conversational speech datasets in Indian languages and collects human data to a brief: speech, images, documents, human feedback and expert annotation.
We are a small team. You would work on a model that people already use in production, and you would see your work ship.
Location & work mode
This role is based in Bengaluru, India, and works from our office. We care about focused time and short feedback loops, which is easier in a room together while the team is this size.
What you'll do
- Train and fine-tune text-to-speech and speech-to-speech models, from data through to a checkpoint that ships.
- Own the quality bar for a language or a voice family: what good sounds like, how it is measured, and what has to change to get there.
- Work on the parts that make speech feel live, streaming inference, time to first audio, and the tradeoffs between latency and quality.
- Build the evaluation that tells us we improved: objective metrics, listening tests, and the tooling around both.
- Take on the hard cases directly, code-switching, low-resource languages, expressive delivery, and voice cloning.
What we're looking for
- Experience training speech or audio models end to end, in PyTorch, and getting them to a state someone else can run.
- A working understanding of modern TTS: neural codecs, autoregressive and diffusion approaches, vocoders, and where each breaks.
- Comfort with the unglamorous half, data collection, cleaning, alignment, and normalization, which is usually where the quality actually comes from.
- The judgement to tell a real improvement from a metric that moved.
- Clear writing. We work from written arguments more than meetings.
Bonus
- You have worked on Indic or African languages, or any low-resource setting.
- You have shipped speech-to-speech or real-time voice systems.
- You have published, or contributed to open-source speech work.
- You speak more than one of the languages we support, and can hear when a voice is wrong.
What we offer
- Competitive salary, discussed early in the process.
- Direct ownership of a system in production, not a slice of a backlog.
- A small team and a short path from idea to shipped.
- The hardware and compute you need to do the work.
- Based in our Bengaluru office.
Apply
This goes straight to a person, not a queue. A short note about what you have built beats a long one about what you are looking for.