PhonePe
About this role
PhonePe processes over 330 million transactions a day, and this role keeps the big data infrastructure behind that running. It is SRE work on distributed Hadoop ecosystems: HDFS, HBase, Airflow, YARN, Ranger, Kafka and Pinot, with the automation around provisioning, scaling, upgrades and cluster patching built by you. You lead on call rotations and incident response, run root cause analysis and postmortems, and design system architectures reviewed for scale and reliability. The posting is unusually direct that this is a senior independent seat: it asks you to set technical direction, drive standardisation and operate independently across multiple business verticals. Rare among today's listings, PhonePe puts the experience band in the job title itself, so there is no guessing what level it is aimed at.
Who this is for
What the posting requires: over 7 years managing and maintaining distributed big data ecosystems, with the title stating a 7 to 11 year band. Strong Linux expertise including IP, iptables and IPsec. Scripting or programming in Perl, Golang or Python. Hands on experience with the Hadoop stack: HDFS, HBase, Airflow, YARN, Ranger, Kafka and Pinot. Familiarity with open source configuration management and deployment tooling such as Puppet, Salt, Chef or Ansible. Solid networking and open source fundamentals. DevOps tooling: SaltStack, Ansible, Docker and Git. Logging and monitoring: the ELK stack, Grafana, Prometheus, OpenTSDB and OpenTelemetry. Strong communication and collaboration.
Nice to have: managing infrastructure on AWS, Azure or GCP; designing and reviewing system architectures for scale; and observability tooling for visualisation and alerting.
The day to day: supporting incremental change across Linux and Unix environments; leading on call rotations, incident response, root cause analysis and postmortems; building automation for provisioning, scaling, upgrades and patching; capacity planning and performance tuning; enforcing security standards; and folding reliability practice into how product teams build.
Benefits worth knowing: PhonePe publishes its full time benefits on the posting, including medical, critical illness, accident and life insurance, an employee assistance programme and on site medical centre, maternity and paternity benefits, adoption and day care support, relocation and transfer support, PF, gratuity, NPS, leave encashment and higher education support. Very few postings on this board list any of this.
Who this is for: an SRE or infrastructure engineer who has actually operated Hadoop at scale, not just used it. If your big data experience is analytics rather than keeping clusters alive at three in the morning, this is not the fit.
Nice to have: managing infrastructure on AWS, Azure or GCP; designing and reviewing system architectures for scale; and observability tooling for visualisation and alerting.
The day to day: supporting incremental change across Linux and Unix environments; leading on call rotations, incident response, root cause analysis and postmortems; building automation for provisioning, scaling, upgrades and patching; capacity planning and performance tuning; enforcing security standards; and folding reliability practice into how product teams build.
Benefits worth knowing: PhonePe publishes its full time benefits on the posting, including medical, critical illness, accident and life insurance, an employee assistance programme and on site medical centre, maternity and paternity benefits, adoption and day care support, relocation and transfer support, PF, gratuity, NPS, leave encashment and higher education support. Very few postings on this board list any of this.
Who this is for: an SRE or infrastructure engineer who has actually operated Hadoop at scale, not just used it. If your big data experience is analytics rather than keeping clusters alive at three in the morning, this is not the fit.
Apply on company site
Opens job-boards.greenhouse.io, the employer's own application page. Applying is always free.