Sarvam
About this role
Sarvam is building what it calls sovereign AI for India, a full stack platform spanning research, models, infrastructure and applications, backed by Lightspeed, Peak XV and Khosla Ventures and deployed with Tata Capital, SBI Life, CRED, IDFC and LIC. The API team is the layer that exposes its model family, speech recognition, text to speech, language models and vision, over high performance APIs. That makes this a real low latency backend problem: serving millions of requests reliably across Azure, AWS, GCP and on premises deployments, in Python with FastAPI, and using HTTP, WebSockets and gRPC depending on the workload. Readers of this list have repeatedly asked for backend roles specifically, and this is the cleanest backend listing of today's thirty.
Who this is for
What the posting requires: strong Python with hands on FastAPI, Django, Flask or similar. Deep understanding of HTTP, WebSockets and gRPC. Proven experience building low latency distributed backend systems. Hands on experience with PostgreSQL, Redis, ClickHouse or related data systems. Familiarity with Kafka or Redis Streams. Solid understanding of API authentication, authorisation and security practice. Experience with Docker, Kubernetes and CI/CD. Hands on experience with at least one major cloud platform, with Azure preferred.
Nice to have: canary deployments, progressive rollouts or feature flag systems; prior work on ML inference or model serving infrastructure; observability tooling such as Prometheus, Grafana and OpenTelemetry; and open source contributions or a solid GitHub portfolio.
The day to day: designing and optimising Python APIs for serving ML models at scale; building communication layers over HTTP, WebSockets and gRPC; architecting low latency, fault tolerant and secure backends for real time inference; implementing authentication, rate limiting and request prioritisation; integrating voice agent and LLM SDKs; working with PostgreSQL, Redis and ClickHouse; building event driven and streaming architectures on Kafka and Redis Streams; and working on canary deployments and CI/CD.
Location and office reality: Bengaluru.
⚠️ Years of experience: Sarvam publishes no experience figure anywhere in this posting. We write no number rather than guessing one, which means this role will not appear if you filter the board by years. Judged by the requirement list it reads mid to senior, but that is our reading and not the employer's statement.
Who this is for: a Python backend engineer who wants latency and reliability problems rather than CRUD, and who is interested in AI infrastructure without needing to be a researcher. The posting describes a high talent density team with high ownership from day one, which usually means a demanding pace.
Nice to have: canary deployments, progressive rollouts or feature flag systems; prior work on ML inference or model serving infrastructure; observability tooling such as Prometheus, Grafana and OpenTelemetry; and open source contributions or a solid GitHub portfolio.
The day to day: designing and optimising Python APIs for serving ML models at scale; building communication layers over HTTP, WebSockets and gRPC; architecting low latency, fault tolerant and secure backends for real time inference; implementing authentication, rate limiting and request prioritisation; integrating voice agent and LLM SDKs; working with PostgreSQL, Redis and ClickHouse; building event driven and streaming architectures on Kafka and Redis Streams; and working on canary deployments and CI/CD.
Location and office reality: Bengaluru.
⚠️ Years of experience: Sarvam publishes no experience figure anywhere in this posting. We write no number rather than guessing one, which means this role will not appear if you filter the board by years. Judged by the requirement list it reads mid to senior, but that is our reading and not the employer's statement.
Who this is for: a Python backend engineer who wants latency and reliability problems rather than CRUD, and who is interested in AI infrastructure without needing to be a researcher. The posting describes a high talent density team with high ownership from day one, which usually means a demanding pace.
Apply on company site
Opens jobs.ashbyhq.com, the employer's own application page. Applying is always free.