Never miss an opening. Get the daily email.
Daily Tech Jobs India

Machine Learning Engineer, Dubbing

Sarvam · Bengaluru

Verified live on July 29, 2026
Get every day's jobs where you already are
Sarvam
BengaluruFull-timeNot stated by the employer

About this role

Sarvam is building what it calls sovereign AI for India, backed by Lightspeed, Peak XV and Khosla Ventures, and working with Tata Capital, SBI Life, CRED, IDFC and LIC. This role owns the machine learning integration layer for its dubbing and live translation products, stitching speech recognition, translation, text to speech and voice cloning into end to end systems. Two very different workloads sit in scope: batch video dubbing across 12 or more Indian languages, and real time speech to speech translation in multi participant settings where the latency budget is measured in hundreds of milliseconds. That real time half is the hard part, and it brings streaming architecture, voice cloning that holds speaker identity across a session, and automated quality control into the job.

Who this is for

Sarvam publishes no years of experience figure on this posting, so none has been recorded here rather than guessing a number.

Required: strong Python and PyTorch, described as being comfortable reading model internals, profiling inference and debugging production failures; hands on experience integrating and optimising speech models (ASR or TTS) in production; experience with real time and streaming systems including WebSocket pipelines, chunked audio processing and latency sensitive async architectures; a solid understanding of modern speech system architectures covering sequence to sequence models, attention mechanisms, flow matching or diffusion based TTS and streaming ASR; familiarity with model serving infrastructure such as Triton, TorchServe or ONNX Runtime; audio signal processing fundamentals including sample rates, PCM formats, spectrograms, vocoding and time stretching; and strong async Python.

The day to day: building and optimising the real time speech to speech translation pipeline with streaming ASR, server side voice activity detection, low latency translation and live TTS audio streams; designing fan out architectures where one ASR stream serves many concurrent listeners with personalised translated audio; implementing voice cloning in both streaming and batch contexts; optimising end to end latency across the ASR, translation and TTS chain; integrating with real time media infrastructure such as WebRTC, RTMP and SRT; and building evaluation harnesses tracking WER and CER, tempo and pronunciation.

Location is Bengaluru. Good fit if you have shipped speech models in production. Not a fit if your ML work has been offline and batch only, because latency is the whole problem here.
Apply on company site Opens jobs.ashbyhq.com, the employer's own application page. Applying is always free.
More roles like this, every day

Every link is checked live before we post it. Get the day's list in your inbox.

Free · one email a day · unsubscribe anytime

More from July 29, 2026

← Back to all jobs