Platform Site Reliability Engineer - Member of Technical Staff at Callosum
United Kingdom
on-site · full_time · {"@type":"MonetaryAmount","currency":"GBP","value":{"@type":"QuantitativeValue","minValue":101000,"maxValue":192000,"unitText":"YEAR"}} · Visa sponsorship
Apply for this role at Callosum
Responsibilities:
- Own end-to-end operational health of the production platform: define SLOs, monitoring, alerting and observability.
- Run and improve on-call processes: create runbooks, lead incident response, conduct blameless postmortems and ensure follow-through.
- Lead capacity planning and operational strategy for heterogeneous compute backends to handle large-scale concurrent traffic.
- Define "production-grade" standards and build reliability practice for a fast-growing platform serving real customer traffic.
- Collaborate closely with platform, hardware and orchestration teams to expose diverse compute backends reliably.
Requirements:
- Strong SRE or production-engineering background running customer-facing systems at scale.
- Experience defining and operating SLOs, monitoring pipelines, alerting and incident response playbooks.
- Proven capacity planning and performance tuning experience for high-concurrency services across diverse infrastructure.
- Hands-on familiarity with container orchestration, observability tooling and infrastructure-as-code practices.
- Strong collaboration skills and a bias for building durable operational systems rather than temporary fixes.