HackerRank
About this role
HackerRank has assessed developers for over 3,000 companies, and Chakra is its bet on what comes next: an AI interviewer that holds a real conversation with a candidate, asks follow up questions, evaluates how they think and produces a report a hiring manager can act on. The posting is refreshingly honest about why this is hard. Getting a model to hold a conversation is solved. Getting it to exercise judgment, knowing what to probe, knowing when an answer is shallow versus when it merely sounds shallow, and doing that 200,000 times without drifting or being gamed, is not. You would own the whole thing end to end: agent design, conversation management, real time evaluation, scoring methodology and report generation, plus the benchmarking that proves it works.
Who this is for
What the posting requires: experience building and shipping agentic or conversational AI systems in production rather than prototypes. Strong intuition for where LLM behaviour breaks down under real world conditions and how to fix it systematically. Systems thinking that treats conversation architecture, evaluation, serving infrastructure and candidate experience as one problem. Care about the quality bar at the level of a user who depends on the output rather than a researcher measuring aggregate metrics.
The day to day: architecting and developing Chakra end to end, covering agent design, conversation management, real time response evaluation, scoring methodology and report generation; building the infrastructure that keeps the 200,000th interview as coherent as the first; designing evaluation and benchmarking pipelines for interview quality, candidate experience consistency and report defensibility; building fine tuning and RLHF workflows to push model judgment past off the shelf behaviour; defining and instrumenting the quality bar; and working across data pipelines, model serving, latency constraints and the product experience itself.
Location and office reality: hybrid in Bengaluru, India.
⚠️ Years of experience: HackerRank publishes no figure on this posting. We write no number rather than guessing, so it will not appear under an experience filter. The requirement to have shipped production agentic systems is the real gate, and that tends to imply several years.
Who this is for: an ML engineer who has taken LLM systems past the demo stage and has opinions about evaluation. If your experience is model training rather than production behaviour and guardrails, the emphasis here is on the latter. There is a second HackerRank ML role open on the evaluation side if this specific product does not appeal.
The day to day: architecting and developing Chakra end to end, covering agent design, conversation management, real time response evaluation, scoring methodology and report generation; building the infrastructure that keeps the 200,000th interview as coherent as the first; designing evaluation and benchmarking pipelines for interview quality, candidate experience consistency and report defensibility; building fine tuning and RLHF workflows to push model judgment past off the shelf behaviour; defining and instrumenting the quality bar; and working across data pipelines, model serving, latency constraints and the product experience itself.
Location and office reality: hybrid in Bengaluru, India.
⚠️ Years of experience: HackerRank publishes no figure on this posting. We write no number rather than guessing, so it will not appear under an experience filter. The requirement to have shipped production agentic systems is the real gate, and that tends to imply several years.
Who this is for: an ML engineer who has taken LLM systems past the demo stage and has opinions about evaluation. If your experience is model training rather than production behaviour and guardrails, the emphasis here is on the latter. There is a second HackerRank ML role open on the evaluation side if this specific product does not appeal.
Apply on company site
Opens job-boards.greenhouse.io, the employer's own application page. Applying is always free.