Senior Software Engineer, Reliability at Klaviyo

Dublin, IE

on-site · full-time · Visa sponsorship

Apply for this role at Klaviyo

Responsibilities:
- Design, build and operate foundational, security-critical services with emphasis on availability, scalability, latency and fault tolerance.
- Apply software engineering to automate infrastructure, reduce operational toil and improve system reliability at scale.
- Define and refine SLIs, SLOs and error budgets; improve observability, alerting and incident response workflows.
- Perform quantitative analysis of system behavior, capacity and scaling limits; drive long-term preventative fixes.
- Participate in on-call rotations focused on sustainable operations and automatic remediation.
- Mentor peers, influence architecture early, and raise the bar for reliability and operational maturity across teams.

Requirements:
- Proven cloud-native SRE experience operating production systems at scale and deep understanding of distributed systems.
- Production-quality coding skills in Python, Go or similar for automation and platform development.
- Experience with Kubernetes, Terraform/AWS (or equivalent IaC), observability tooling, capacity planning and chaos testing.
- Track record of diagnosing complex failures, designing preventative solutions and leading reliability initiatives.
- Strong communication, mentoring and cross-team collaboration skills.