Sarvam
About this role
This is a full lifecycle vision language model role at Sarvam, covering data, training, evaluation and production rather than one slice of it. On the research side you run training and fine tuning pipelines for large vision language models on GPU clusters, build multimodal data pipelines including synthetic generation and deduplication, and implement architectures and training techniques from recent papers. On the applied side you work directly with clients on document processing, visual search and form extraction, and own those solutions end to end including latency, accuracy and edge cases after deployment. Sarvam says openly that the team's scope will change as the field does and that it wants people comfortable with an unfixed roadmap, which is a fair warning and worth taking at face value.
Who this is for
Sarvam publishes no years of experience figure on this posting, so none has been recorded here rather than guessing.
Required: strong Python and PyTorch, described as comfort reading and modifying model internals; hands on experience training or fine tuning large models including debugging broken runs; experience building data pipelines at scale; a solid grounding in transformer architectures and modern training techniques; comfort with ambiguity, since the roadmap is not fully pre specified; a strong focus on secure coding, code quality and system reliability; and an undergraduate degree in a technical discipline such as computer science, statistics or physics.
Bonus points listed: experience with vision language or multimodal systems, distributed training using FSDP, DeepSpeed or Megatron-LM, post training methods such as RLHF, DPO or other alignment techniques, and inference optimisation covering quantisation, distillation and serving.
The day to day: designing and running training and fine tuning pipelines on GPU clusters; building multimodal data pipelines covering ingestion, filtering, deduplication, synthetic generation and quality assurance; implementing new architectures from research; building evaluation harnesses, benchmarks and automated regression tracking; optimising models for inference; working directly with clients on document processing, visual search and form extraction; and debugging deployed solutions for latency, accuracy and edge cases.
Location is Bengaluru. Good fit if you have trained or fine tuned large models yourself and want both research and client facing delivery. Not a fit if you only consume model APIs without touching training.
Required: strong Python and PyTorch, described as comfort reading and modifying model internals; hands on experience training or fine tuning large models including debugging broken runs; experience building data pipelines at scale; a solid grounding in transformer architectures and modern training techniques; comfort with ambiguity, since the roadmap is not fully pre specified; a strong focus on secure coding, code quality and system reliability; and an undergraduate degree in a technical discipline such as computer science, statistics or physics.
Bonus points listed: experience with vision language or multimodal systems, distributed training using FSDP, DeepSpeed or Megatron-LM, post training methods such as RLHF, DPO or other alignment techniques, and inference optimisation covering quantisation, distillation and serving.
The day to day: designing and running training and fine tuning pipelines on GPU clusters; building multimodal data pipelines covering ingestion, filtering, deduplication, synthetic generation and quality assurance; implementing new architectures from research; building evaluation harnesses, benchmarks and automated regression tracking; optimising models for inference; working directly with clients on document processing, visual search and form extraction; and debugging deployed solutions for latency, accuracy and edge cases.
Location is Bengaluru. Good fit if you have trained or fine tuned large models yourself and want both research and client facing delivery. Not a fit if you only consume model APIs without touching training.
Apply on company site
Opens jobs.ashbyhq.com, the employer's own application page. Applying is always free.