Senior Site Reliability Engineer (SRE & Platform Reliability) at Affirm

on-site · full-time · Visa sponsorship

Apply for this role at Affirm

Responsibilities:
- Own and deliver quarterly team goals; lead engineers through ambiguity to solve open-ended reliability problems.
- Define and drive SLOs, observability, alerting, incident response, post-incident analysis and runbooks.
- Lead incident management and change/deployment practices; recommend deployment and configuration controls.
- Design and run capacity planning, load and chaos testing to improve resilience and scalability.
- Build automation and tooling to reduce toil and improve operational readiness; support on-call and "keep the lights on" activities.
- Partner with infrastructure, product, developer experience and analytics on architecture, trade-offs and operational risk.
- Mentor engineers; raise code-review and design standards; deliver documentation and tech talks.

Requirements:
- 4+ years building and launching backend systems at scale using scripting/development languages (Bash, Python, Kotlin).
- Proven experience with distributed systems and technologies such as AWS, MySQL and Kubernetes.
- Strong background in SRE practices: SLOs, incident lifecycle, observability, capacity management, and automation.
- Experience with load/chaos testing, configuration/change management and deployment pipelines.
- Comfortable with participating in an on-call rotation and supporting production incidents.