Site Reliability Engineer - Observability at N26

on-site · full-time · Visa sponsorship

Apply for this role at N26

Responsibilities:
- Design, build and maintain end-to-end observability pipelines (metrics, logs, traces) for cloud-native infrastructure and microservices.
- Implement and operate extraction, transformation and ingestion tooling (Prometheus, StatsD, OpenTelemetry, Vector, Beats, FluentBit) and storage/visualization (OpenSearch, Grafana, Datadog as needed).
- Deliver developer-focused dashboards, alerts and runbooks to improve mean-time-to-detect and mean-time-to-recover.
- Automate repetitive tasks, alerting, incident playbooks and reliability checks using code and CI/CD.
- Apply infrastructure-as-code to provision observability resources and enforce secure, compliant defaults.
- Collaborate with platform teams and service owners to instrument applications, define SLOs and run post-incident reviews.

Requirements:
- Solid understanding of observability foundations: metrics, logs and distributed traces.
- Hands-on experience with Prometheus/StatsD/OpenTelemetry and log pipeline tooling (Vector, Beats, FluentBit) and dashboards/storage (OpenSearch, Grafana, Datadog).
- Proficiency in a scripting/compiled glue language (Go or Python) and comfortable working on Linux systems.
- Experience operating in cloud environments and managing infra as code (Terraform, AWS CDK or similar).
- Strong troubleshooting, automation-first mindset, and ability to communicate observability patterns to engineers and stakeholders.
- Willingness to learn, adapt quickly, and work in a fast-paced, security-conscious, cloud-hosted banking environment.