Mistral Cloud - Site Reliability Engineer at Mistral
on-site · full_time · Visa sponsorship
Apply for this role at Mistral
Responsibilities:
- Design, deploy and operate scalable, highly available infrastructure spanning cloud and bare-metal for customer APIs and model training.
- Run production operations: incident triage, on-call response, user/admin support, capacity planning and scaling.
- Implement and tune monitoring, logging and alerting to meet reliability and performance targets.
- Build and maintain CI/CD, containerization and orchestration tooling to support large training runs and customer workloads.
- Automate infrastructure and deployment tasks using IaC and scripting to reduce toil and improve reproducibility.
- Conduct post-incident RCA and implement mitigations to prevent recurrence.
- Partner with software and security teams to ensure deployments meet performance, security and compliance needs.
Requirements:
- Master’s degree (preferred) and 5+ years in SRE/DevOps with experience on bare-metal and distributed systems.
- Proven expertise with Kubernetes, Docker, Terraform/CloudFormation and infrastructure-as-code workflows.
- Strong scripting/automation skills (Python, Go, Bash) and experience building CI/CD pipelines.
- Familiarity with observability stacks (Prometheus, Grafana, ELK, Datadog) and incident management.
- Knowledge of networking, security best practices, HPC or GPU training environments is a plus.
- Excellent troubleshooting, communication skills and readiness to participate in on-call rotations.