Mistral Cloud - Site Reliability Engineer at Mistral AI
Amsterdam
on-site · full-time · Visa sponsorship
Apply for this role at Mistral AI
Responsibilities:
- Design, build and maintain scalable, highly available and fault-tolerant cloud and on-prem infrastructure.
- Operate production systems: incident response, on-call rotations, troubleshooting and root cause analysis.
- Implement and improve monitoring, alerting and observability for customer APIs and large model-training runs.
- Develop and maintain CI/CD pipelines, container orchestration and automation tooling (IaC, Terraform).
- Collaborate with software, ML and security teams to enable reproducible model-training experiments and secure deployments.
- Document processes and contribute to open-source projects, research outputs and internal knowledge sharing.
Requirements:
- Master’s degree preferred and 5+ years in SRE, DevOps or infrastructure engineering roles.
- Strong experience with Kubernetes, Docker, Terraform and infrastructure-as-code practices.
- Proficiency in scripting and automation (Python, Go, Bash) and building deployment workflows.
- Familiarity with observability stacks (Prometheus, Grafana, ELK, Datadog) and incident management.
- Experience with bare-metal/HPC environments, large-scale distributed systems and networking.
- Excellent troubleshooting skills, security awareness and ability to work cross-functionally.