Sarvam
About this role
Sarvam is building India's sovereign AI stack across research, models, infrastructure and applications, backed by Lightspeed, Peak XV and Khosla Ventures, and working with Tata Capital, SBI Life, CRED, IDFC and LIC. This role takes Sarvam's models from the state researchers hand them over in to production ready artifacts on at least two target chipsets, chosen from Intel, ARM and Apple processors and Nvidia or AMD GPUs. You would own one or two model and chipset pairs end to end: quantise the model, validate accuracy, benchmark it, document it, and write the deployment workbook for each pair you own. You also embed part time with the app team consuming your work and debug performance and accuracy problems alongside them, and you maintain the team's benchmark harness. The ask is 3+ years on ML systems with real production quantisation experience.
Who this is for
What the posting requires:
- 3+ years working on ML systems.
- Solid PyTorch and ONNX export experience, including the awkward parts: dynamic shapes, control flow and custom ops. The posting names these specifically.
- Quantisation in production on at least one real model.
- Comfort with at least two of ONNX Runtime, TensorRT, CoreML, OpenVINO, QNN and LiteRT.
- Profiling fluency on at least one platform.
Nice to have:
- Custom op authoring in any runtime.
What the work actually looks like:
- Own one or two model and chipset pairs end to end, taking each from research handoff to a production ready artifact.
- Quantise, validate accuracy, benchmark and document.
- Author the deployment workbook for each pair you own.
- Embed part time with the consuming app team during integration, debugging performance and accuracy issues with them.
- Maintain and extend the team's benchmark harness.
Target hardware named in the posting: Intel xPU, ARM xPU, Apple xPU, and Nvidia or AMD GPUs. You are expected to deliver on at least two of these.
Location and working pattern: Bengaluru. The posting does not state a hybrid or remote arrangement, so assume in office.
What the employer says about the team: a high talent density team, AI first in how it builds and ships, with high ownership from day one and population scale impact as the stated draw.
Honest fit guidance: this is a narrow, deep specialism. Three years is enough if those years included real quantisation and profiling work on shipped models. General ML engineering experience without on device deployment will not clear the bar, because the entire role is about the gap between a research checkpoint and a model that runs fast on specific silicon. No interview process is published.
- 3+ years working on ML systems.
- Solid PyTorch and ONNX export experience, including the awkward parts: dynamic shapes, control flow and custom ops. The posting names these specifically.
- Quantisation in production on at least one real model.
- Comfort with at least two of ONNX Runtime, TensorRT, CoreML, OpenVINO, QNN and LiteRT.
- Profiling fluency on at least one platform.
Nice to have:
- Custom op authoring in any runtime.
What the work actually looks like:
- Own one or two model and chipset pairs end to end, taking each from research handoff to a production ready artifact.
- Quantise, validate accuracy, benchmark and document.
- Author the deployment workbook for each pair you own.
- Embed part time with the consuming app team during integration, debugging performance and accuracy issues with them.
- Maintain and extend the team's benchmark harness.
Target hardware named in the posting: Intel xPU, ARM xPU, Apple xPU, and Nvidia or AMD GPUs. You are expected to deliver on at least two of these.
Location and working pattern: Bengaluru. The posting does not state a hybrid or remote arrangement, so assume in office.
What the employer says about the team: a high talent density team, AI first in how it builds and ships, with high ownership from day one and population scale impact as the stated draw.
Honest fit guidance: this is a narrow, deep specialism. Three years is enough if those years included real quantisation and profiling work on shipped models. General ML engineering experience without on device deployment will not clear the bar, because the entire role is about the gap between a research checkpoint and a model that runs fast on specific silicon. No interview process is published.
Apply on company site
Opens jobs.ashbyhq.com, the employer's own application page. Applying is always free.