Sr Staff Site Reliability Engineer at Palo Alto Networks

on-site · full time · Visa sponsorship

Apply for this role at Palo Alto Networks

Responsibilities:
- Operate and own global, multi-cloud production environments (GCP, AWS, Azure).
- Monitor, investigate and resolve incidents; lead post-incident analysis and remediation.
- Troubleshoot complex distributed systems and drive end-to-end fault isolation and fixes.
- Design, deploy and improve observability platforms (Prometheus, Grafana) and alerting.
- Build automation, tooling and infrastructure-as-code (Terraform) to reduce toil (Python preferred).
- Collaborate with CX, CS and engineering teams; champion documentation, async communication and best practices.
- Participate in on-call rotations (daytime coverage and occasional weekends/holidays).

Requirements:
- 5+ years in SRE or similar production engineering roles supporting large-scale systems.
- Hands-on experience with Kubernetes, Terraform and cloud provider operations (GCP/AWS/Azure).
- Strong skills in monitoring/observability tooling (Prometheus, Grafana) and incident response workflows.
- Experience developing automation and tools, ideally in Python.
- Familiarity with GitOps, CI/CD pipelines and modern deployment practices.
- Excellent troubleshooting, communication and cross-team collaboration skills for a distributed remote environment.