Site Reliability Engineer at Mistral
on-site · full_time · Visa sponsorship
Apply for this role at Mistral
Responsibilities:
- Design, build and maintain scalable, highly available infrastructure to support web services and ML workloads.
- Ensure platform, inference and model-training environments remain highly available and reproducible across HPC clusters.
- Operate production systems, respond to incidents and participate in on-call rotations.
- Implement and improve monitoring, alerting, logging and incident response (Prometheus, Grafana, ELK, Datadog).
- Automate provisioning and workflows using infrastructure-as-code (Terraform) to reduce operational toil.
- Collaborate with research and engineering teams to enable cloud-agnostic deployments and capacity planning.
Requirements:
- Master’s degree and 7+ years’ experience in SRE/DevOps or similar production-infrastructure roles.
- Strong hands-on experience with Kubernetes, Docker, Terraform and infrastructure automation.
- Proficiency in scripting and tooling (Python, Go, Bash) and CI/CD pipelines.
- Familiarity with HPC schedulers (Slurm), GPU infrastructure and cloud provider ecosystems.
- Demonstrated incident management, root-cause analysis and observability best practices.