munotes®
Never miss an opening. Get the daily email.
Daily Tech Jobs India

Researcher | Vision Language Models, Training and Evaluation in PyTorch

Sarvam · Bengaluru, India

Verified live on August 7, 2026
Sarvam
Apply
Bengaluru, IndiaFull-timeNot stated

About this role

Sarvam builds full stack AI for India, and this is a research seat on vision language models rather than an applied engineering one. The scope runs the whole lifecycle: data, training, evaluation and production, with the team's remit expected to change as the field does, so the posting asks for researchers comfortable with that and able to lead. The named work is researching vision language architectures, encoders and the tradeoffs between them, with rigorous experimental design as an explicit requirement, meaning the ability to isolate variables and draw defensible conclusions. PyTorch is the stated tool and the bar is running experiments end to end yourself. What makes this genuinely interesting is the Indian language angle: multilingual and low resource language modelling, document understanding, OCR and structured visual prediction all appear as valued experience, which is population scale work that very few labs anywhere are doing.

Who this is for

Sarvam publishes no years of experience figure on this posting, so we have left the experience field blank rather than guess. It asks for capabilities and evidence instead:

Deep understanding of vision language models covering training dynamics, architecture tradeoffs and failure modes. A track record of good research, demonstrated through publications, technical reports or impactful shipped work. Rigorous experimental design, specifically the ability to isolate variables and draw defensible conclusions. Strong PyTorch skills, with the ability to run experiments end to end. Intellectual range, meaning a willingness to work across data, training and evaluation problems.

Bonus points: a PhD or Master's with relevant research experience in ML, computer vision, NLP or a related field. Research papers published at A or A star venues. Experience with multilingual or low resource language modelling. Familiarity with document understanding, OCR or structured visual prediction. Experience with large scale data curation and its effect on model quality.

What you will do: work across the full lifecycle of vision language model development, from data through training and evaluation to production. Research vision language architectures including encoders and their tradeoffs.

Location: Bengaluru.

Honest fit guidance: note that the advanced degree and the publications sit under bonus points, not requirements, and "impactful shipped work" is offered as an alternative to papers. That is a deliberately wider door than most research postings, and it means a strong engineer with real model training experience and no PhD should not self reject. What is not optional is the PyTorch depth and the experimental rigour: Sarvam wants someone who designs an experiment properly rather than someone who runs training scripts. If your ML work has been fine tuning with off the shelf recipes, this is a step up rather than a lateral move.
Apply on company site Opens jobs.ashbyhq.com, the employer's own application page. Applying is always free.

More from August 7, 2026

More roles like this, every day

Every link is checked live before we post it. Get the day's list in your inbox.

Free · one email a day · unsubscribe anytime
Or where you already are: WhatsApp Telegram

← Back to all jobs