Sumo Logic
About this role
Sumo Logic runs a cloud native observability and security analytics platform, and this is a site reliability role embedded with a product area rather than a central operations one. The posting is clear about that: you maintain and execute a reliability roadmap for your product area, define and manage SLOs for several teams, and join the on call rotation specifically so you understand the operational load well enough to reduce it. There is a strong emphasis on writing code and automation to eliminate toil, plus facilitating blame free root cause analysis after incidents. The stack is AWS heavy with Kubernetes, Terraform, Ansible and Jenkins, and you need to author production ready code in Java, Scala or Go. At 4 to 6 years it is a good mid level SRE seat. Noida, and the posting states hybrid.
Who this is for
Required by the posting: 4 to 6 years of industry experience. A Bachelor's or Master's degree in Computer Science, Electrical Engineering or another scientific or technical discipline. Cloud native application development experience using best practices and design patterns. Strong debugging and troubleshooting across the entire technology stack. Deep understanding of AWS networking, compute, storage and managed services. Competency with modern CI/CD tooling including Kubernetes, Terraform, Ansible and Jenkins. Experience with full lifecycle support of services from creation to production support. Infrastructure as code practice using Terraform or CloudFormation. Ability to author production ready code in at least one of Java, Scala or Go. Experience with Linux systems and comfort on the command line. Understanding of modern cloud native software security. Experience with agile frameworks such as Scrum and Kanban. Flexibility to step into new roles and responsibilities, and willingness to learn and use Sumo Logic's own products for solving reliability and security issues.
Desirable skills: experience using Sumo Logic or other observability products; planet scale product development; running and operating SaaS products on AWS at expert level; streaming technologies such as Kafka, Kafka Streams or KSQL; expert level experience in one or more of Java, Go, Scala or Python; and expert level experience in one or more of Terraform, Jenkins or Kubernetes.
The real day to day: maintaining and executing a reliability roadmap for your product area covering reliability, maintainability, security, efficiency and velocity; collaborating with development infrastructure, Global SRE and product area engineering teams to refine that roadmap; defining, evolving and managing SLOs for several teams; participating in on call rotations specifically to understand the operational workload well enough to improve it; improving the lifecycle of microservices from design through deployment and refinement; writing code and automation to reduce operational workload, improve security posture and eliminate toil; working with developer infrastructure teams and contributing features and bug fixes back; scaling systems sustainably through automation; and facilitating blame free root cause analysis meetings after incidents.
Location: the posting body states Noida (Hybrid), while the Greenhouse metadata field for this requisition carries an office requirement value of Remote. We have published the body's statement because it is the more specific one, but this is a genuine ambiguity worth resolving with the recruiter on the first call.
Honest fit guidance: the reliability roadmap and SLO ownership make this more strategic than a firefighting SRE seat, and the requirement to author production code in Java, Scala or Go means it is an engineering job rather than an operations one. If you have four to six years and can write real code, this is a strong mid level move.
Desirable skills: experience using Sumo Logic or other observability products; planet scale product development; running and operating SaaS products on AWS at expert level; streaming technologies such as Kafka, Kafka Streams or KSQL; expert level experience in one or more of Java, Go, Scala or Python; and expert level experience in one or more of Terraform, Jenkins or Kubernetes.
The real day to day: maintaining and executing a reliability roadmap for your product area covering reliability, maintainability, security, efficiency and velocity; collaborating with development infrastructure, Global SRE and product area engineering teams to refine that roadmap; defining, evolving and managing SLOs for several teams; participating in on call rotations specifically to understand the operational workload well enough to improve it; improving the lifecycle of microservices from design through deployment and refinement; writing code and automation to reduce operational workload, improve security posture and eliminate toil; working with developer infrastructure teams and contributing features and bug fixes back; scaling systems sustainably through automation; and facilitating blame free root cause analysis meetings after incidents.
Location: the posting body states Noida (Hybrid), while the Greenhouse metadata field for this requisition carries an office requirement value of Remote. We have published the body's statement because it is the more specific one, but this is a genuine ambiguity worth resolving with the recruiter on the first call.
Honest fit guidance: the reliability roadmap and SLO ownership make this more strategic than a firefighting SRE seat, and the requirement to author production code in Java, Scala or Go means it is an engineering job rather than an operations one. If you have four to six years and can write real code, this is a strong mid level move.
Apply on company site
Opens job-boards.greenhouse.io, the employer's own application page. Applying is always free.