Databricks
About this role
A customer facing role with real engineering depth behind it. When a Databricks customer's generative AI or machine learning workload misbehaves in production, this is who works out why. The posting is specific that you analyse and troubleshoot production workloads at the code level and optimise for performance, reliability, latency and cost, across products like Agent Bricks, Vector Search and Model Serving. That means diagnosing real time and batch inference, autoscaling, monitoring and alerting, and advising on experiment tracking, model registry, versioning, evaluation and lifecycle observability. You also feed what you learn back into the roadmap and into internal documentation. It reports into a technical solutions manager within the global support engineering organisation.
Who this is for
What the posting requires: 8 or more years designing, building and scaling data, machine learning and AI systems on premises and in the cloud using Python, Scala and Java in production, with expertise in machine learning or generative AI. Experience with AWS, Azure or GCP, with Databricks familiarity a plus. Proficiency in the data engineering needed to orchestrate end to end ML training pipelines, ideally processing large datasets with Apache Spark. Subject matter expertise in feature engineering, ML frameworks, model training, monitoring, drift detection and retraining. Proficiency with algorithms, deep learning and NLP techniques. Prior experience building, designing or troubleshooting LLM based generative AI applications. Familiarity with agentic frameworks such as LangChain and LangGraph. Expertise in context orchestration including prompt design, memory management, retrieval systems, vector embeddings, semantic search and tool integrations. Comprehensive MLOps and LLMOps knowledge. Prior support or customer facing experience.
The day to day: acting as senior technical expert on complex issues spanning data pipelines, ML pipelines and AI applications; troubleshooting production workloads at code level; diagnosing ML and LLM deployments including inference, autoscaling, monitoring and alerting; guiding customers on model lifecycle and observability; supporting generative AI use cases across RAG, agents, vector search and prompt engineering; and collaborating internally to influence the roadmap and contribute documentation.
Location and office reality: Bengaluru, India.
Who this is for: an ML engineer who genuinely enjoys debugging other people's systems and talking to customers while doing it. This is a support engineering seat, not a platform building one, but the technical bar is the same. Readers asked for customer facing technical roles by name, and this is one of the more hands on examples on today's list.
The day to day: acting as senior technical expert on complex issues spanning data pipelines, ML pipelines and AI applications; troubleshooting production workloads at code level; diagnosing ML and LLM deployments including inference, autoscaling, monitoring and alerting; guiding customers on model lifecycle and observability; supporting generative AI use cases across RAG, agents, vector search and prompt engineering; and collaborating internally to influence the roadmap and contribute documentation.
Location and office reality: Bengaluru, India.
Who this is for: an ML engineer who genuinely enjoys debugging other people's systems and talking to customers while doing it. This is a support engineering seat, not a platform building one, but the technical bar is the same. Readers asked for customer facing technical roles by name, and this is one of the more hands on examples on today's list.
Apply on company site
Opens databricks.com, the employer's own application page. Applying is always free.