Staff Software Engineer, Observability & Profiling at Anthropic

on-site · full-time · Visa sponsorship

Apply for this role at Anthropic

Responsibilities:
- Design and build scalable telemetry ingest and storage pipelines for metrics, logs, traces, and error data across multi-cluster infrastructure.
- Build low-overhead observability solutions (continuous profiling, eBPF tracing, network visibility) that surface root causes from kernel/accelerator to application.
- Own and evolve core observability platforms, driving migrations, cost/reliability improvements, and platform-wide instrumentation.
- Create instrumentation libraries, SDKs, and auto-instrumentation (including eBPF) to emit high-quality telemetry without heavy code changes.
- Implement cross-signal correlation and unified query interfaces to reduce mean time to detection and resolution.
- Partner with Research, Inference, Product, and Infrastructure teams to tailor observability for GPU/TPU/Trainium workloads and fleet efficiency.

Requirements:
- Hands-on experience building and operating large-scale observability or monitoring infrastructure.
- Deep familiarity with telemetry signals end-to-end: instrumentation, ingest, storage, query, and analysis (metrics, logs, traces, profiles).
- Experience with high-throughput telemetry pipelines, sampling strategies, and trade-offs for storage and query cost vs. fidelity.
- Practical knowledge of eBPF, kernel-level debugging, continuous profiling, and accelerator telemetry is highly desirable.
- Strong system-design and performance-engineering skills, plus experience collaborating across research and infrastructure teams.
- Ability to deliver production-grade tooling that scales across large multi-cluster deployments and reduces operational toil.