Sarvam
About this role
Sarvam is building a full stack sovereign AI platform for India, backed by Lightspeed, Peak XV and Khosla Ventures, and working with Tata Capital, SBI Life, CRED, IDFC and LIC. Studio is its creative media platform, doing AI dubbing, live translation and voice cloning across more than 12 Indian languages. This role keeps that platform running and shipping. You own production Kubernetes clusters with GPU inference servers, async task queues and multi stage ML pipelines across a multi cloud setup, run Helm based deployments and CI/CD, and build the observability that catches failures before customers see them. There is real depth here on secrets management, cost optimisation and incident response for a system where a single job spans minutes and one failure cascades through pipeline stages. The stack is Kubernetes, Helm, Docker, Prometheus and Grafana, Python and Bash, across Azure, GCP or AWS.
Who this is for
What the posting asks for: 3 or more years in a DevOps, SRE or Infrastructure Engineering role. Sarvam marks two requirements non negotiable: strong hands on production Kubernetes covering deployments, Helm charts, HPAs, CronJobs, node affinity, resource management and debugging pod failures, and proficiency with at least one major cloud among Azure, GCP or AWS covering managed Kubernetes, container registries, blob storage, secrets vaults and IAM. Beyond those it asks for experience building and maintaining CI/CD pipelines with container builds, automated testing gates, multi environment promotion and deployment automation, solid containerisation including production Dockerfiles, multi stage builds and image optimisation, experience with monitoring and observability stacks such as Prometheus, Grafana, OpenTelemetry or Sentry, strong Linux systems knowledge across networking, process management, storage and system level debugging, scripting in Python or Bash, a good understanding of networking including DNS, load balancing, ingress controllers, TLS termination and CDN configuration, experience with secrets management patterns such as External Secrets Operator or Sealed Secrets, and familiarity with Redis or similar in memory stores for queues and pub sub.
The real day to day: operating production Kubernetes with multi role deployments across API servers, task schedulers, per stage workers and WebSocket servers. Building and optimising CI/CD with staged rollouts across QA, staging and production. Maintaining Helm charts, shared module dependencies and environment specific value overlays. Implementing metrics, dashboards, error tracking and distributed tracing. Operating blob storage with CDN for media delivery, ingress, secrets and IAM. Coordinating multi cloud deployments and artifact management for Docker images and internal Python packages. Running vault to cluster secret sync. Optimising async task queues including dead letter handling and stuck message recovery. Monitoring database query health, connection pooling and backup and recovery. Building developer productivity tooling and local dev environments. Owning incident response including runbooks, alerting rules and post mortems. Driving cost optimisation through right sizing, autoscaling policies and storage lifecycle management.
Location and setup: Bengaluru. The posting does not state office days.
Honest fit guidance: the workload is ML heavy and GPU backed, which is meaningfully harder to operate than a standard web stack, and the posting is honest that failures cascade across pipeline stages. Good fit if you want infrastructure work at a company where the infrastructure is the hard part. Sarvam also has a Studio backend role open today at 4 to 6 years, which is the same platform from the application side.
The real day to day: operating production Kubernetes with multi role deployments across API servers, task schedulers, per stage workers and WebSocket servers. Building and optimising CI/CD with staged rollouts across QA, staging and production. Maintaining Helm charts, shared module dependencies and environment specific value overlays. Implementing metrics, dashboards, error tracking and distributed tracing. Operating blob storage with CDN for media delivery, ingress, secrets and IAM. Coordinating multi cloud deployments and artifact management for Docker images and internal Python packages. Running vault to cluster secret sync. Optimising async task queues including dead letter handling and stuck message recovery. Monitoring database query health, connection pooling and backup and recovery. Building developer productivity tooling and local dev environments. Owning incident response including runbooks, alerting rules and post mortems. Driving cost optimisation through right sizing, autoscaling policies and storage lifecycle management.
Location and setup: Bengaluru. The posting does not state office days.
Honest fit guidance: the workload is ML heavy and GPU backed, which is meaningfully harder to operate than a standard web stack, and the posting is honest that failures cascade across pipeline stages. Good fit if you want infrastructure work at a company where the infrastructure is the hard part. Sarvam also has a Studio backend role open today at 4 to 6 years, which is the same platform from the application side.
Apply on company site
Opens jobs.ashbyhq.com, the employer's own application page. Applying is always free.