Sarvam
About this role
Sarvam is building India's own full stack AI platform, and Chanakya is the team handling its defence and strategic sector work. This role owns the model lifecycle across those deployments, which covers serving infrastructure, monitoring, evaluation pipelines and environment management. The posting sets the standard directly: the system has to be always on, always accurate and always auditable, and it notes that a model failure here is not a user experience problem but an operational risk. The work spans two layers, supporting the deployment engineers who are physically in the field and owning model deployment infrastructure for new products being built by the product team. A lot of it happens in unusual conditions, including on premise, air gapped and edge environments where the normal managed cloud tooling is unavailable. You would also write the runbooks that field engineers rely on and own incident response when a model layer fails anywhere across the active deployments.
Who this is for
What the posting requires
- 3 to 5 years in machine learning engineering or MLOps, with at least one production language model or ML system in continuous operation
- Deep model serving expertise using vLLM, TGI, Triton Inference Server or equivalent, including quantised formats such as GGUF, AWQ or GPTQ
- Experience fine tuning and adapting models in constrained, on premise or air gapped environments, managing the data pipeline and compute limits those bring
- Containerisation with Docker and Kubernetes, or lightweight alternatives such as K3s or K0s for constrained and edge deployments, across varied hardware
- Monitoring and observability with Prometheus, Grafana or equivalent, including building custom evaluation dashboards
- Python fluency, with familiarity with fine tuning workflows and model evaluation frameworks
- CI/CD tooling for ML pipelines such as GitHub Actions, ArgoCD or DVC
The signals they say they look for
- You have kept a production ML system running under load, and debugged it when it broke
- You do not wait for failures, you build systems that warn you before they happen
- You write documentation that people other than you actually use
What you would actually be doing
- Designing and operating model serving infrastructure across on premise and cloud deployments
- Building CI/CD pipelines for model updates, rollbacks and evaluation gated releases
- Monitoring latency, accuracy drift, throughput and failure modes, and building systems that surface problems before clients notice
- Building evaluation infrastructure: harnesses, A/B testing and model comparison tooling for both lab and field use
- Managing containerised serving in constrained, air gapped and edge environments
- Writing runbooks and operational playbooks for deployment engineers in the field
- Owning incident response for model layer failures across every active deployment
Location and working pattern
Based in Delhi rather than Bengaluru, which is worth noting because most Indian AI roles cluster in Bangalore.
A good fit if
You would rather keep models running reliably than train them, and you treat uptime and correctness as equally non negotiable.
Think twice if
You want research or modelling work. This is squarely infrastructure, reliability and operations, and the air gapped and edge constraints mean the usual managed cloud conveniences often are not available.
- 3 to 5 years in machine learning engineering or MLOps, with at least one production language model or ML system in continuous operation
- Deep model serving expertise using vLLM, TGI, Triton Inference Server or equivalent, including quantised formats such as GGUF, AWQ or GPTQ
- Experience fine tuning and adapting models in constrained, on premise or air gapped environments, managing the data pipeline and compute limits those bring
- Containerisation with Docker and Kubernetes, or lightweight alternatives such as K3s or K0s for constrained and edge deployments, across varied hardware
- Monitoring and observability with Prometheus, Grafana or equivalent, including building custom evaluation dashboards
- Python fluency, with familiarity with fine tuning workflows and model evaluation frameworks
- CI/CD tooling for ML pipelines such as GitHub Actions, ArgoCD or DVC
The signals they say they look for
- You have kept a production ML system running under load, and debugged it when it broke
- You do not wait for failures, you build systems that warn you before they happen
- You write documentation that people other than you actually use
What you would actually be doing
- Designing and operating model serving infrastructure across on premise and cloud deployments
- Building CI/CD pipelines for model updates, rollbacks and evaluation gated releases
- Monitoring latency, accuracy drift, throughput and failure modes, and building systems that surface problems before clients notice
- Building evaluation infrastructure: harnesses, A/B testing and model comparison tooling for both lab and field use
- Managing containerised serving in constrained, air gapped and edge environments
- Writing runbooks and operational playbooks for deployment engineers in the field
- Owning incident response for model layer failures across every active deployment
Location and working pattern
Based in Delhi rather than Bengaluru, which is worth noting because most Indian AI roles cluster in Bangalore.
A good fit if
You would rather keep models running reliably than train them, and you treat uptime and correctness as equally non negotiable.
Think twice if
You want research or modelling work. This is squarely infrastructure, reliability and operations, and the air gapped and edge constraints mean the usual managed cloud conveniences often are not available.
Apply on company site
Opens jobs.ashbyhq.com, the employer's own application page. Applying is always free.