Never miss an opening. Get the daily email.
Daily Tech Jobs India

Platform Engineer, AI Infrastructure

Sarvam · Bengaluru

Verified live on July 24, 2026
Get every day's jobs where you already are
Sarvam
BengaluruFull-time5+ years

About this role

Sarvam runs a large multi vendor GPU fleet carrying two workloads that fight each other: training jobs spanning hundreds of GPUs, and inference services that have to hold a flat p99 under production load. This role builds the platform on top of that fleet, the scheduling, scaling, multi tenancy and self service layers that let ML teams use thousands of GPUs without a human approving every job. The posting draws a clear line: this is the build side of infrastructure, not the operate side, and the SREs carry the pager. The work is heavy software engineering in Go or Python, with Kubernetes at controller and internals level and GPU specific constraints such as MIG, gang scheduling, topology aware placement and RDMA. It asks 5+ years.

Who this is for

What the posting asks for:
- 5+ years building infrastructure or platform software, with a track record of systems you designed and shipped. The posting is pointed about this: services and control planes others built on, not scripts and glue.
- Strong software engineering in Go or Python, with the judgement to build maintainable systems and the depth to debug them in production.
- Kubernetes at the controller and internals level. You should have written operators or controllers, understand the scheduler and API machinery, and know where the abstractions leak.
- Working literacy in GPU specific platform constraints: MIG and GPU sharing, gang scheduling, topology and fabric aware placement, and why training and serving contend for the same hardware.
- A product mindset toward internal users, measuring the platform by adoption and self service rather than tickets closed.
- The range to own a capability end to end, from design through rollout to the documentation that makes it self serve.

Counts as a bonus:
- Having built a serving, inference or training platform: routing, autoscaling and rollout for model endpoints at scale.
- GPU scheduling systems in multi tenant production (Kueue, Volcano, Slurm, Run:ai or custom).
- Multi tenant isolation (MIG, MPS, time slicing) shipped as a self service capability.
- Deep Kubernetes networking: CNI internals, custom network components, or RDMA and SR-IOV in pods.
- On premise GPU platform work, including multi vendor or Indian NCP environments.
- Open source contributions to Kubernetes, scheduling or GPU platform projects.

The surface area, from the posting: the serving platform control plane with routing, canary and blue green rollouts and traffic splitting; the scaling and elasticity layer for both training and serving; scheduling and orchestration with gang scheduling, priority and preemption and quota enforcement; multi tenancy, RBAC and isolation; platform networking; observability and cost tooling; the storage and data path; developer experience through CLI, SDK and APIs; and provisioning through Terraform, Crossplane or operators. You would take one capability and build it end to end rather than own all of it at once.

Experience: the posting states 5+ years.

Location: Bengaluru.

Who should apply: a platform engineer who has built control planes rather than maintained them, and who wants GPU scale problems. If your Kubernetes experience is writing manifests rather than controllers, the bar here is above you.
Apply on company site Opens jobs.ashbyhq.com, the employer's own application page. Applying is always free.
More roles like this, every day

Every link is checked live before we post it. Get the day's list in your inbox.

Free · one email a day · unsubscribe anytime

More from July 24, 2026

← Back to all jobs