Senior Staff+ Software Engineer, Kubernetes Platform at Anthropic
London, UK
on-site · full-time · Visa sponsorship
Apply for this role at Anthropic
Responsibilities:
- Own and extend the Kubernetes scheduler: implement custom plugins, gang scheduling, topology-aware placement, and preemption policies.
- Scale control plane components (apiserver, etcd, controllers) to support extremely large clusters and high object counts.
- Design, build, and operate core cluster services (service discovery, networking, etc.) that must sustain heavy ML workloads.
- Implement and maintain custom controllers, operators, and CRDs to express platform capabilities.
- Collaborate with research, training, and inference teams to translate workload shapes into platform features.
- Lead incident response, on-call rotations, runbooks, postmortems, and SLO-driven reliability improvements.
Requirements:
- Significant production experience building and operating distributed systems at scale.
- Proficiency in systems languages (Go, Python, Rust, or C++) and strong engineering fundamentals.
- Deep, hands-on knowledge of Kubernetes internals—scheduler, apiserver, controllers, and etcd—beyond typical user-level experience.
- Proven ability to debug complex cross-stack issues (API, network, node-level) and design for reliability and correct failure semantics.
- Strong written and verbal communication; ability to build consensus with internal stakeholders.
Preferred:
- Experience operating very large multi-tenant clusters, ML/accelerator workloads, cloud provider collaboration, and kernel-level tuning (e.g., cgroups, eBPF).