(Senior) Cloud Site Reliability Engineer (Scalability) (m/f/x) at Scalable GmbH
on-site · full-time · Visa sponsorship
Apply for this role at Scalable GmbH
Responsibilities:
- Define and implement scalability, reliability, and cost-optimization strategies for microservices and storage across AWS (ECS, Fargate, Lambda) and multi-account setups.
- Design and roll out monitoring and observability best practices in Datadog, including SLIs, SLOs and alerts.
- Build and maintain self-service tools for autoscaling, load testing, chaos engineering, and developer discovery/portals.
- Drive performance and cost improvements via FinOps practices and service/storage optimizations, including serverless evaluations.
- Run and analyze load tests, capacity planning, and chaos experiments to validate resilience and performance.
- Mentor and enable development teams to adopt reusable, unified building blocks and DevOps practices.
- Collaborate with cross-functional teams to translate scalability requirements into technical designs and roadmaps.
Requirements:
- Several years of hands-on experience with AWS (ECS/Fargate/Lambda), microservices, and distributed systems.
- Proficiency with infrastructure-as-code (Terraform or similar) and CI/CD pipelines.
- Strong monitoring/observability experience, ideally with Datadog; familiar with SLI/SLO concepts.
- Practical experience with load testing, chaos engineering and performance tuning.
- Programming skills in Python plus at least one other language (Java, Kotlin, JavaScript/Node.js preferred).
- Familiarity with FinOps, cost analysis and multi-account cloud operations.
- Excellent collaboration and mentoring skills; ability to influence engineering teams and stakeholders.