Sarvam
About this role
This is the second Sarvam field engineering role today, on its AI dubbing platform, which dubs video into more than 12 Indian languages while preserving speaker voice, tone and timing. You would be the senior technical partner to media companies, OTT platforms, content studios and enterprise localisation teams, owning integration architecture and pipeline tuning so clients get production quality dubbed content at scale. The debugging surface is specific and interesting: audio separation, speech recognition, translation, text to speech and video stitching, plus tuning parameters like voice activity detection thresholds, translation glossaries, TTS voice profiles and audio mixing for each client's content type. Unlike many customer facing roles this one expects you to contribute code back, fixing platform bugs that client deployments surface. Two requirements are marked non negotiable: strong Python, and audio or video processing experience.
Who this is for
What the posting requires:
- 5 to 8 years in field engineering, solutions engineering, technical account management or senior client facing engineering roles.
- Strong Python proficiency, with the ability to read, debug and contribute to production FastAPI services and ML pipelines. Marked **non negotiable**.
- Experience with audio and video processing workflows: FFmpeg, codec pipelines, media formats or streaming infrastructure. Also marked **non negotiable**.
- A proven track record working with enterprise media, OTT or content localisation clients.
- Comfort operating across the stack: REST APIs, async job queues such as Celery or Redis, PostgreSQL, cloud storage on Azure, GCP or AWS, and Kubernetes.
- Strong debugging instincts, tracing failures across distributed systems from API through queue and worker to ML inference and storage.
- Experience owning SLA management and escalation governance across multiple enterprise accounts.
- Excellent communication, comfortable engaging CXO and VP level stakeholders.
What the work actually looks like:
- Lead end to end integration of the dubbing platform into enterprise content workflows across OTT, media houses, ed-tech and enterprise learning and development.
- Own the technical relationship with strategic accounts: scoping requirements, designing integration architecture and ensuring production readiness.
- Debug and resolve complex pipeline issues across the full dubbing stack: audio separation, speech recognition, translation, text to speech and video stitching.
- Tune pipeline parameters for client specific content: voice activity detection thresholds, translation glossaries, TTS voice profiles and audio mixing.
- Drive presales engagements including technical discovery and proof of concept scoping.
- Build and maintain integration playbooks, API guides and troubleshooting runbooks.
- Define SLA governance across enterprise accounts covering turnaround time, quality benchmarks and escalation resolution.
- Act as primary technical liaison between enterprise clients and Sarvam's dubbing product and ML engineering teams.
- Mentor other field engineers, and contribute fixes back to internal platform codebases when client deployments surface bugs.
Location and working pattern: Bengaluru, with significant customer facing responsibility.
Honest fit guidance: the media pipeline experience is the differentiator. Plenty of engineers can debug a distributed system, far fewer have worked with FFmpeg and codec pipelines, and that is what the posting marks non negotiable alongside Python. If you have both, the customer facing element is learnable. If you have neither, the title being FDSE will not carry the application.
- 5 to 8 years in field engineering, solutions engineering, technical account management or senior client facing engineering roles.
- Strong Python proficiency, with the ability to read, debug and contribute to production FastAPI services and ML pipelines. Marked **non negotiable**.
- Experience with audio and video processing workflows: FFmpeg, codec pipelines, media formats or streaming infrastructure. Also marked **non negotiable**.
- A proven track record working with enterprise media, OTT or content localisation clients.
- Comfort operating across the stack: REST APIs, async job queues such as Celery or Redis, PostgreSQL, cloud storage on Azure, GCP or AWS, and Kubernetes.
- Strong debugging instincts, tracing failures across distributed systems from API through queue and worker to ML inference and storage.
- Experience owning SLA management and escalation governance across multiple enterprise accounts.
- Excellent communication, comfortable engaging CXO and VP level stakeholders.
What the work actually looks like:
- Lead end to end integration of the dubbing platform into enterprise content workflows across OTT, media houses, ed-tech and enterprise learning and development.
- Own the technical relationship with strategic accounts: scoping requirements, designing integration architecture and ensuring production readiness.
- Debug and resolve complex pipeline issues across the full dubbing stack: audio separation, speech recognition, translation, text to speech and video stitching.
- Tune pipeline parameters for client specific content: voice activity detection thresholds, translation glossaries, TTS voice profiles and audio mixing.
- Drive presales engagements including technical discovery and proof of concept scoping.
- Build and maintain integration playbooks, API guides and troubleshooting runbooks.
- Define SLA governance across enterprise accounts covering turnaround time, quality benchmarks and escalation resolution.
- Act as primary technical liaison between enterprise clients and Sarvam's dubbing product and ML engineering teams.
- Mentor other field engineers, and contribute fixes back to internal platform codebases when client deployments surface bugs.
Location and working pattern: Bengaluru, with significant customer facing responsibility.
Honest fit guidance: the media pipeline experience is the differentiator. Plenty of engineers can debug a distributed system, far fewer have worked with FFmpeg and codec pipelines, and that is what the posting marks non negotiable alongside Python. If you have both, the customer facing element is learnable. If you have neither, the title being FDSE will not carry the application.
Apply on company site
Opens jobs.ashbyhq.com, the employer's own application page. Applying is always free.