Never miss an opening. Get the daily email.
Daily Tech Jobs India

ML Engineer (Data), Foundational Models

Sarvam · Bengaluru, India

Verified live on July 23, 2026
Get every day's jobs where you already are
Sarvam
Bengaluru, IndiaFull-time3+ years

About this role

Sarvam is building sovereign, full stack AI for India, backed by Lightspeed, Peak XV and Khosla Ventures, and this seat owns the data infrastructure that feeds its next family of foundational models. It is explicitly not a glue code job. The work is engineering and research heavy: petabyte scale curation and filtering pipelines, deduplication at scale, model based quality classifiers, contamination detection, and the mixture, curriculum and annealing decisions that determine what a model actually learns and in what order. The stated bar is 3+ years, but Sarvam says strong early career candidates with a real systems background will be considered, so it is more accessible than the title suggests. Based in Bengaluru. This is for a data or ML systems engineer who wants to be able to walk through a pretraining corpus end to end and defend every choice that went into it.

Who this is for

What the posting asks for:
- 3+ years building large scale data systems (petabyte scale processing, distributed data pipelines or comparable). Exceptional early career candidates with a strong systems background will also be considered.
- A BS or MS in Computer Science or a closely related field, or equivalent demonstrated experience.
- Hands on experience with data curation and filtering for LLM training, and the ability to defend the choices in a corpus you helped build.
- Depth in distributed data frameworks (Spark, Ray, Beam, Dask or equivalent) and the storage underneath, plus strong Python and comfort with tokenization, sharding, packing and IO tradeoffs.

Nice to have: experience with large open pretraining corpora, model based quality classifiers, contamination detection, or data attribution research.

The day to day: you build ingestion, parsing, filtering, dedup, tokenization and packing pipelines at petabyte scale, improve quality filtering, own mixture and curriculum design with the research team, and build tooling so researchers can slice and debug the data.

Location: Bengaluru.

Honest read: a serious data engineering role at a pretraining lab. The 3+ years is a soft floor, so strong systems focused early career engineers should still consider it.
Apply on company site Opens jobs.ashbyhq.com, the employer's own application page. Applying is always free.
More roles like this, every day

Every link is checked live before we post it. Get the day's list in your inbox.

Free · one email a day · unsubscribe anytime

More from July 23, 2026

← Back to all jobs