Staff Software Engineer, Node Infra at Anthropic

London, UK

on-site · full-time · Visa sponsorship

Apply for this role at Anthropic

Responsibilities:
- Own the technical strategy and roadmap for node lifecycle management: ingestion, bring-up, health checking, repair automation, and decommissioning.
- Drive cross-team projects to provision and scale accelerator fleets across multiple cloud providers and datacenters.
- Design and operate systems to detect, isolate, and automatically remediate unhealthy hardware to increase fleet MTBI and reduce stranded capacity.
- Define infrastructure architecture and lead solutions for the hardest systems problems, either hands-on or through other engineers.
- Collaborate with cloud providers and internal research/inference/product teams to shape long-term compute and data strategy.
- Establish and evolve operational practices: incident response, postmortems, runbooks, on-call rotation, and reliability metrics.
- Mentor and coach engineers to grow technical capability across the team.

Requirements:
- Deep expertise in distributed systems, reliability engineering, and cloud platforms (Kubernetes, IaC, AWS/GCP/Azure).
- Strong proficiency in at least one systems language (Rust, Go, or Python) and IaC experience with Terraform.
- Hands-on experience with ML accelerators (GPUs, TPUs, Trainium) and managing node fleets.
- Proven track record leading complex, multi-quarter technical initiatives spanning multiple teams.
- Strong communication skills and ability to align senior stakeholders and cross-functional partners.
- Preferred: 8+ years engineering experience; experience operating hyperscale compute (10k+ nodes); deep knowledge of Kubernetes internals (scheduler, autoscaler, kubelet, Karpenter) or large-scale node provisioning systems.