Sarvam
About this role
Sarvam's model APIs for speech recognition, text to speech, vision, LLM and translation are how developers and enterprises build on its foundation models, and the platform handles tens of millions of API calls a day over HTTP and WebSockets with metering, billing, prepaid wallets and rate limiting in the same stack. This role owns that platform end to end: architecture, reliability, performance and standards. The concrete headline is a rewrite, since the platform is FastAPI and Python today and is being moved to Go, and this engineer is expected to lead that architecture and set the patterns without compromising reliability. Beyond the rewrite there is audio pipeline work, vision pipelines orchestrating OCR and layout detection, streaming infrastructure with backpressure, and the commercial metering and wallet layer held to the same engineering bar as the model APIs. The posting states clearly that this is an individual contributor role.
Who this is for
What the posting requires:
- 7+ years building production backend systems at scale.
- Strong Go in production: designed, built and operated Go services under real load.
- Comfort in Python, since you work in both languages through the rewrite and FastAPI is the current foundation.
- A track record designing and operating high scale, low latency, multi tenant distributed systems.
- Hands on experience with real time and streaming systems: WebSockets, long lived connections and backpressure.
- Strong PostgreSQL and Redis fundamentals covering schema design, query performance and caching strategy.
- Comfort running services on Kubernetes in production.
- Senior individual contributor judgement: when to build versus buy, optimise versus ship, abstract versus inline, and how to bring others with you.
Nice to have:
- Experience serving LLM, ASR, TTS or vision models in production.
- Background in audio processing with pyav, FFmpeg, codec and sample rate work, or VAD.
- Experience building metering, billing or wallet and payments infrastructure.
- Time at an early or growth stage startup.
What the work actually looks like:
- Own the end to end design and evolution of the platform, from the moment a request hits the edge to the response going back out.
- Lead the Python to Go rewrite: architecture, patterns and migration without losing reliability.
- Audio pipeline engineering: TTS chunking around model context limits, sample rate adjustment, format encoding, VAD based silence detection for ASR.
- Vision pipelines: orchestrating OCR, layout detection and VLM harnesses for structured data extraction.
- Streaming infrastructure: WebSocket connections, queue based batch processing, backpressure and low latency model invocation.
- The commercial layer: metering, billing, prepaid wallet management and rate limiting.
- Observability across logging, metrics and tracing, and the integration test harness that lets the team ship without breaking customer APIs.
- Partner with the Inference and MLOps teams and mentor them on production engineering.
How the team works: design first and documentation first, with RFCs before code, decisions written down, and an explicit goal of no tribal knowledge.
Location and working pattern: Bengaluru. No remote or hybrid arrangement is stated.
Honest fit guidance: this is a senior IC ownership role, and the Go requirement is not negotiable given the rewrite sits at the centre of the job. Python only backend engineers would be applying against the primary ask.
- 7+ years building production backend systems at scale.
- Strong Go in production: designed, built and operated Go services under real load.
- Comfort in Python, since you work in both languages through the rewrite and FastAPI is the current foundation.
- A track record designing and operating high scale, low latency, multi tenant distributed systems.
- Hands on experience with real time and streaming systems: WebSockets, long lived connections and backpressure.
- Strong PostgreSQL and Redis fundamentals covering schema design, query performance and caching strategy.
- Comfort running services on Kubernetes in production.
- Senior individual contributor judgement: when to build versus buy, optimise versus ship, abstract versus inline, and how to bring others with you.
Nice to have:
- Experience serving LLM, ASR, TTS or vision models in production.
- Background in audio processing with pyav, FFmpeg, codec and sample rate work, or VAD.
- Experience building metering, billing or wallet and payments infrastructure.
- Time at an early or growth stage startup.
What the work actually looks like:
- Own the end to end design and evolution of the platform, from the moment a request hits the edge to the response going back out.
- Lead the Python to Go rewrite: architecture, patterns and migration without losing reliability.
- Audio pipeline engineering: TTS chunking around model context limits, sample rate adjustment, format encoding, VAD based silence detection for ASR.
- Vision pipelines: orchestrating OCR, layout detection and VLM harnesses for structured data extraction.
- Streaming infrastructure: WebSocket connections, queue based batch processing, backpressure and low latency model invocation.
- The commercial layer: metering, billing, prepaid wallet management and rate limiting.
- Observability across logging, metrics and tracing, and the integration test harness that lets the team ship without breaking customer APIs.
- Partner with the Inference and MLOps teams and mentor them on production engineering.
How the team works: design first and documentation first, with RFCs before code, decisions written down, and an explicit goal of no tribal knowledge.
Location and working pattern: Bengaluru. No remote or hybrid arrangement is stated.
Honest fit guidance: this is a senior IC ownership role, and the Go requirement is not negotiable given the rewrite sits at the centre of the job. Python only backend engineers would be applying against the primary ask.
Apply on company site
Opens jobs.ashbyhq.com, the employer's own application page. Applying is always free.