Senior Site Reliability Engineer at EMBL
on-site · full time · 3861.91 · Visa sponsorship
Responsibilities:
- Ensure availability, performance and reliability of services used by global researchers; own incident and problem response for platform services.
- Operate and evolve identity & access systems (Active Directory, Entra-ID, Red Hat IDP) and support directory services (389 DS/OpenLDAP).
- Maintain and develop email systems (Postfix, Cyrus, Roundcube, Mailman) and support migrations (e.g., to O365).
- Support data transfer and storage services (Globus, Aspera, FTP/HTTP, FIRE object storage).
- Maintain monitoring and reliability tooling (Check_mk), define observability baselines and contribute to monitoring strategy.
- Automate lifecycle management and orchestration using Foreman, Puppet, Gerrit, RPM repos to manage thousands of servers.
- Produce clear documentation and SOPs to empower Service Desk and improve user experience.
Requirements:
- 5+ years operating Linux production systems with strong troubleshooting skills (tcpdump, strace, large-scale log parsing).
- Hands-on experience with orchestration/automation tools (Foreman, Puppet) and configuration management.
- Experience with directory services (389 DS or OpenLDAP) and identity systems.
- Ability to read/review and contribute code in Python, Bash and Puppet.
- Familiarity with monitoring, backup and lifecycle management for large server estates.
- Strong communication skills, willingness to learn new technologies and collaborate in a hybrid working environment.