Databricks
About this role
This is the second Databricks role on today's board and it sits on the Observability team, which builds the tooling that tells the company how its own platform is behaving. The scale described in the posting is the point: a fleet of millions of virtual machines generating terabytes of logs and processing exabytes of data per day, where cloud hardware, network and operating system faults are routine and the software has to shield customers from them. You would set the standards for logging, metrics and tracing, work with other teams to identify the metrics that matter, and build the infrastructure that lets components emit, aggregate and store those metrics for dashboards and alerting. There is a cost angle too: analysing system expense, enforcing retention, streamlining queries and right sizing resources to cut observability spend. On call is part of the role.
Who this is for
What the posting requires:
- A BS or higher in Computer Science or a related field.
- 5+ years of production level experience in one of Python, Java, Scala, C++ or similar languages.
- Experience in software development on large scale distributed systems.
- Familiarity with metrics collection, health monitoring and observability tools.
What the work actually looks like:
- Establish standards for logging, metrics and tracing across the platform.
- Work with different teams to identify the metrics engineers need to see how the system and its subcomponents are performing.
- Build tooling and infrastructure so components can efficiently emit, aggregate and store metrics for dashboards and alerting.
- Contribute to and execute the technical roadmap for scalability, performance and reliability.
- Participate in on call rotations and reduce incident response times.
- Optimise platform and infrastructure cost by analysing system expenses, improving visibility, enforcing retention policies, streamlining queries and right sizing resources.
The scale you would be working at, as stated: millions of virtual machines, terabytes of logs and exabytes of data processed per day, with hardware, network and OS faults treated as expected conditions.
Location and working pattern: Bengaluru. The posting mentions Databricks is standing up ten new teams from scratch at this site, which suggests early team formation and the ambiguity that comes with it.
Compliance note in the posting: the same export controlled technology clause appears here as on other Databricks roles. Where access to export controlled technology or source code is needed, applying for a US government licence is at the employer's discretion.
Honest fit guidance: this is infrastructure engineering, not product. If your background is observability, metrics pipelines or large scale operations, the five year bar is met easily by the right experience. If you have never carried a pager, the on call requirement is stated plainly and is a real part of the job.
- A BS or higher in Computer Science or a related field.
- 5+ years of production level experience in one of Python, Java, Scala, C++ or similar languages.
- Experience in software development on large scale distributed systems.
- Familiarity with metrics collection, health monitoring and observability tools.
What the work actually looks like:
- Establish standards for logging, metrics and tracing across the platform.
- Work with different teams to identify the metrics engineers need to see how the system and its subcomponents are performing.
- Build tooling and infrastructure so components can efficiently emit, aggregate and store metrics for dashboards and alerting.
- Contribute to and execute the technical roadmap for scalability, performance and reliability.
- Participate in on call rotations and reduce incident response times.
- Optimise platform and infrastructure cost by analysing system expenses, improving visibility, enforcing retention policies, streamlining queries and right sizing resources.
The scale you would be working at, as stated: millions of virtual machines, terabytes of logs and exabytes of data processed per day, with hardware, network and OS faults treated as expected conditions.
Location and working pattern: Bengaluru. The posting mentions Databricks is standing up ten new teams from scratch at this site, which suggests early team formation and the ambiguity that comes with it.
Compliance note in the posting: the same export controlled technology clause appears here as on other Databricks roles. Where access to export controlled technology or source code is needed, applying for a US government licence is at the employer's discretion.
Honest fit guidance: this is infrastructure engineering, not product. If your background is observability, metrics pipelines or large scale operations, the five year bar is met easily by the right experience. If you have never carried a pager, the on call requirement is stated plainly and is a real part of the job.
Apply on company site
Opens databricks.com, the employer's own application page. Applying is always free.