Senior Site Reliability Engineer I at Etsy

on-site · full_time · Visa sponsorship

Apply for this role at Etsy

Responsibilities:
- Own and operate high-impact platform systems that power Etsy Search: Kubernetes runtimes, CI/CD pipelines, deployment platforms and model serving.
- Build developer-facing tooling (self-service deploys, observability, release automation) to increase velocity and reduce operational toil.
- Collaborate with ML, platform and product teams to define standards, influence architecture and scale infrastructure reliably.
- Implement automation, monitoring, incident response and capacity planning for distributed, high-availability services.
- Improve developer experience, document runbooks and mentor engineers on platform best practices.

Requirements:
- 5+ years software engineering experience with at least ~2 years in SRE/DevOps or infrastructure in cloud environments.
- Hands-on experience with Kubernetes, Terraform (or IaC), Linux and public cloud (GCP/AWS preferred).
- Strong understanding of distributed systems, CI/CD, containerisation and observability tools (e.g., Grafana, Prometheus).
- Proficiency in at least one programming language (Golang preferred) and experience automating operational workflows.
- Excellent collaboration and communication skills, and experience operating production services at scale.