Senior Staff+ Software Engineer, Kubernetes Platform at Anthropic

London, UK

on-site · full-time · Visa sponsorship

Apply for this role at Anthropic

Responsibilities:
- Own and extend the Kubernetes scheduler: implement custom plugins, gang scheduling, topology-aware placement, and preemption policies.
- Scale control plane components (apiserver, etcd, controllers) to support extremely large clusters and high object counts.
- Design, build, and operate core cluster services (service discovery, networking, etc.) that must sustain heavy ML workloads.
- Implement and maintain custom controllers, operators, and CRDs to express platform capabilities.
- Collaborate with research, training, and inference teams to translate workload shapes into platform features.
- Lead incident response, on-call rotations, runbooks, postmortems, and SLO-driven reliability improvements.

Requirements:
- Significant production experience building and operating distributed systems at scale.
- Proficiency in systems languages (Go, Python, Rust, or C++) and strong engineering fundamentals.
- Deep, hands-on knowledge of Kubernetes internals—scheduler, apiserver, controllers, and etcd—beyond typical user-level experience.
- Proven ability to debug complex cross-stack issues (API, network, node-level) and design for reliability and correct failure semantics.
- Strong written and verbal communication; ability to build consensus with internal stakeholders.

Preferred:
- Experience operating very large multi-tenant clusters, ML/accelerator workloads, cloud provider collaboration, and kernel-level tuning (e.g., cgroups, eBPF).