Sarvam
About this role
Sarvam is building India's sovereign AI stack, backed by Lightspeed, Peak XV and Khosla Ventures, and working with Tata Capital, SBI Life, CRED, IDFC and LIC. This role owns the infrastructure that Sarvam's next family of foundation models trains on, which is one of the harder systems problems available in India right now. Expect distributed training across large GPU clusters, parallelism strategies spanning data, tensor, pipeline, sequence and expert, custom CUDA and Triton kernels where the off the shelf option leaves performance behind, and the reliability engineering that decides whether a training run of weeks actually finishes. The posting is refreshingly blunt that a bad MFU number or a slow data loader costs weeks and millions. Sarvam says exceptional early career candidates with a strong systems background will be considered.
Who this is for
What the posting asks for:
- A BS or MS in computer science or a closely related technical field, or equivalent demonstrated experience.
- 3+ years building ML training infrastructure or large scale distributed systems. The posting adds that exceptional early career candidates with a strong systems background will be considered, which is worth noting if you are under three years.
- Hands on experience training large models with distributed training frameworks: Megatron LM, DeepSpeed, FSDP, NeMo or equivalent. The posting says you should have been on call for a real pretraining run.
- Deep working knowledge of GPU architecture, the CUDA programming model, and profiling tools such as Nsight and the PyTorch profiler.
- Strong PyTorch internals knowledge, comfortable reading and modifying low level training code rather than only using high level APIs.
- Meaningful open source contributions in the training infrastructure ecosystem: Megatron, DeepSpeed, PyTorch, vLLM, Triton, NCCL or similar.
Counts as a bonus:
- Custom CUDA or Triton kernel development with measurable wins on real workloads.
- Direct experience training models at 10B+ parameters or on 1000+ GPU clusters.
- Cluster orchestration and job scheduling at scale (Slurm, Kubernetes).
- Mixed precision (BF16, FP8), quantization aware training or similar numerical work in production runs.
- First author papers or technical reports on training systems, scaling or model efficiency.
The day to day: building and pushing the limits of the distributed training stack across large GPU clusters, designing parallelism strategies and reasoning about which combinations suit which architectures and scales, profiling and optimising end to end throughput including kernel performance, communication overlap, memory layout, checkpointing and data loading, writing and tuning custom GPU kernels, owning the reliability of long running jobs through fault tolerance and checkpoint integrity and deterministic restarts, and partnering with researchers and the data team.
Experience: the posting states 3+ years, with an explicit exception for exceptional early career systems engineers.
Location: Bengaluru.
Who should apply: a systems engineer who wants frontier model training work without leaving India. Sarvam is direct that the bar is high and the cost of mistakes is measured in weeks. If you have real distributed training experience, very few Indian employers can offer this problem.
- A BS or MS in computer science or a closely related technical field, or equivalent demonstrated experience.
- 3+ years building ML training infrastructure or large scale distributed systems. The posting adds that exceptional early career candidates with a strong systems background will be considered, which is worth noting if you are under three years.
- Hands on experience training large models with distributed training frameworks: Megatron LM, DeepSpeed, FSDP, NeMo or equivalent. The posting says you should have been on call for a real pretraining run.
- Deep working knowledge of GPU architecture, the CUDA programming model, and profiling tools such as Nsight and the PyTorch profiler.
- Strong PyTorch internals knowledge, comfortable reading and modifying low level training code rather than only using high level APIs.
- Meaningful open source contributions in the training infrastructure ecosystem: Megatron, DeepSpeed, PyTorch, vLLM, Triton, NCCL or similar.
Counts as a bonus:
- Custom CUDA or Triton kernel development with measurable wins on real workloads.
- Direct experience training models at 10B+ parameters or on 1000+ GPU clusters.
- Cluster orchestration and job scheduling at scale (Slurm, Kubernetes).
- Mixed precision (BF16, FP8), quantization aware training or similar numerical work in production runs.
- First author papers or technical reports on training systems, scaling or model efficiency.
The day to day: building and pushing the limits of the distributed training stack across large GPU clusters, designing parallelism strategies and reasoning about which combinations suit which architectures and scales, profiling and optimising end to end throughput including kernel performance, communication overlap, memory layout, checkpointing and data loading, writing and tuning custom GPU kernels, owning the reliability of long running jobs through fault tolerance and checkpoint integrity and deterministic restarts, and partnering with researchers and the data team.
Experience: the posting states 3+ years, with an explicit exception for exceptional early career systems engineers.
Location: Bengaluru.
Who should apply: a systems engineer who wants frontier model training work without leaving India. Sarvam is direct that the bar is high and the cost of mistakes is measured in weeks. If you have real distributed training experience, very few Indian employers can offer this problem.
Apply on company site
Opens jobs.ashbyhq.com, the employer's own application page. Applying is always free.