Staff Infrastructure Engineer, Cluster Infrastructure at Anthropic

London, UK

on-site · full-time · Visa sponsorship

Apply for this role at Anthropic

Responsibilities:
- Lead the technical strategy and roadmap for agent-driven cluster lifecycle management: provisioning, updates, scaling, and decommissioning.
- Build and maintain automation that provisions secure-by-default clusters across multiple cloud providers and datacenters.
- Coordinate with partner teams to ensure new compute capacity is ingested on schedule and meets bandwidth/connectivity requirements.
- Align physical build-out and cloud solutions to deliver high-bandwidth inter-cluster connectivity and fault-tolerant topology.
- Collaborate with security teams to bake in secure configurations, RBAC, networking and compliance controls at provisioning time.
- Define strategy for cluster scalability, homogeneity, and fault tolerance; implement mechanisms for automated draining, recovery, and lifecycle updates.
- Establish operational excellence: incident response, postmortems, runbooks, and healthy on-call practices; mentor engineers.

Requirements:
- Deep expertise in distributed systems, reliability, and cloud platforms (Kubernetes, IaC, AWS/GCP/Azure).
- Strong proficiency in at least one systems language (Rust, Go, or Python) and IaC tooling experience (Terraform).
- Proven track record leading complex, multi-quarter technical initiatives that span teams and systems.
- Experience aligning multiple stakeholders and communicating technical direction clearly at all levels.
- Preferred: 8+ years engineering experience; operating hyperscale infrastructure (100+ clusters, 10k+ nodes); deep knowledge of Kubernetes internals and cluster provisioning systems.