Airbnb
About this role
Airbnb's BizTech group builds and runs the internal corporate technology the company itself works on, and the Global Operations team inside it keeps those production services healthy through observability, incident management and automation. This role is unusual in how directly it states its premise: AI fluency is core to the job, not an add on. You would use LLM assisted triage to resolve tickets, prototype agentic workflows, embed AI into runbooks and diagnostics, and build self healing observability that catches problems before they escalate. The rest is recognisable operations engineering: owning a ticket queue, a rotating on call that includes weekends, dashboards in Tableau, Superset and Grafana, and automation code in Python, Go, JavaScript, TypeScript and Bash. Success is measured as a shrinking backlog of recurring ticket categories and faster mean time to resolution. The posting is remote and eligible anywhere within India.
Who this is for
What the posting requires:
- 3+ years with observability and metrics tooling such as Prometheus, Grafana, Datadog or ElasticSearch.
- 3+ years working with data querying and pipelines such as SQL, Airflow, Trino or SQS.
- Working knowledge of network fundamentals and hardware, for example Cisco and Palo Alto.
- Experience with Infrastructure as Code.
- Hands on experience with CI/CD and automation tooling such as Jenkins, ArgoCD or GitHub Actions, across AWS, GCP or Oracle Cloud.
- Comfort working in ticket and workflow driven environments using Jira and Confluence.
Nice to have:
- Experience supporting enterprise SaaS tools such as Salesforce and Workday, and SSO or identity integration such as Okta.
What the work actually looks like:
- Manage the ticket queue, prioritise and resolve requests, and identify recurring categories to automate or deflect.
- Take part in a rotating on call and incident response schedule, including weekends, troubleshooting and documenting in real time.
- Build and maintain monitoring dashboards in Tableau, Superset and Grafana covering service health, availability and data quality.
- Use AI assisted tools to speed up triage, root cause analysis, scripting and documentation, while knowing when a problem needs hands on judgement instead.
- Write and maintain automation code in Python, Go, JavaScript, TypeScript and Bash.
- Partner with global stakeholder teams to drive issues to resolution.
Location and working pattern: India remote, eligible anywhere within India. The posting says the role may include occasional work at an Airbnb office or attendance at offsites, as agreed with your manager. That is a genuinely remote posting rather than a hybrid one described loosely, which matters if relocation is the thing standing between you and a job.
Honest fit guidance: this is an operations and reliability seat, not a product engineering one, and the weekend on call is stated up front. If you have three years in observability and pipelines and you want remote work in India, this is one of the few postings today that qualifies on both counts. No interview process is published.
- 3+ years with observability and metrics tooling such as Prometheus, Grafana, Datadog or ElasticSearch.
- 3+ years working with data querying and pipelines such as SQL, Airflow, Trino or SQS.
- Working knowledge of network fundamentals and hardware, for example Cisco and Palo Alto.
- Experience with Infrastructure as Code.
- Hands on experience with CI/CD and automation tooling such as Jenkins, ArgoCD or GitHub Actions, across AWS, GCP or Oracle Cloud.
- Comfort working in ticket and workflow driven environments using Jira and Confluence.
Nice to have:
- Experience supporting enterprise SaaS tools such as Salesforce and Workday, and SSO or identity integration such as Okta.
What the work actually looks like:
- Manage the ticket queue, prioritise and resolve requests, and identify recurring categories to automate or deflect.
- Take part in a rotating on call and incident response schedule, including weekends, troubleshooting and documenting in real time.
- Build and maintain monitoring dashboards in Tableau, Superset and Grafana covering service health, availability and data quality.
- Use AI assisted tools to speed up triage, root cause analysis, scripting and documentation, while knowing when a problem needs hands on judgement instead.
- Write and maintain automation code in Python, Go, JavaScript, TypeScript and Bash.
- Partner with global stakeholder teams to drive issues to resolution.
Location and working pattern: India remote, eligible anywhere within India. The posting says the role may include occasional work at an Airbnb office or attendance at offsites, as agreed with your manager. That is a genuinely remote posting rather than a hybrid one described loosely, which matters if relocation is the thing standing between you and a job.
Honest fit guidance: this is an operations and reliability seat, not a product engineering one, and the weekend on call is stated up front. If you have three years in observability and pipelines and you want remote work in India, this is one of the few postings today that qualifies on both counts. No interview process is published.
Apply on company site
Opens careers.airbnb.com, the employer's own application page. Applying is always free.