Senior Site Reliability Engineer at Photoroom

on-site · full_time · Visa sponsorship

Apply for this role at Photoroom

Responsibilities:
- Own and evolve the ML inference infrastructure that serves millions of requests daily across GPU-based systems.
- Design cloud-agnostic, scalable solutions for model serving, load balancing, autoscaling and queuing to meet latency and throughput targets.
- Implement observability and monitoring (metrics, traces, alerts) and use tools like Datadog to maintain platform health.
- Run capacity planning, performance tuning and cost optimisation for GPU workloads and inference clusters.
- Collaborate closely with ML, Product, Web and Mobile teams to identify bottlenecks, improve deployment workflows and enable faster iteration.
- Lead incident response, postmortems and reliability improvements; optionally participate in on-call rotation.

Requirements:
- Significant experience designing and operating large-scale distributed systems with strict availability and latency requirements.
- Hands-on expertise with load balancing, autoscaling, queuing systems, container orchestration (Kubernetes/ECS) and CI/CD pipelines.
- Practical experience running GPU inference workloads and optimising cost/performance for ML deployments.
- Strong monitoring and incident management skills; experience with Datadog or similar observability stacks.
- Familiarity with cloud providers, networking, and security best practices for production services.
- Clear communicator, able to work cross-functionally in a fast-paced startup environment and make pragmatic technical decisions.