Site Reliability Engineer at Wheely
on-site · full-time · Visa sponsorship
Responsibilities:
- Respond to alerts and incidents; triage and resolve availability, performance and security issues.
- Support development teams during incidents and help debug production problems.
- Automate repetitive operational tasks using IaC (Terraform, Ansible), scripts (Bash, Python) and CI/CD pipelines.
- Build and maintain application deployment pipelines, observability (Prometheus, Grafana, Sentry, Loki) and resilient infrastructure on AWS.
- Design and migrate infrastructure and services to improve scalability, resilience and operational efficiency.
Requirements:
- Proven experience with cloud platforms (AWS/GCP) and container orchestration (Kubernetes, Docker).
- Strong Linux troubleshooting skills (networking, file systems) and database/admin experience (PostgreSQL, MongoDB, Redis).
- Experience with messaging systems (RabbitMQ/Kafka) and monitoring stacks (Prometheus, Grafana, Sentry, Loki).
- Proficiency in infrastructure automation: Terraform, Ansible and scripting (Bash, Python).
- Good written and verbal communication, ability to prioritize, collaborate with engineers and teach best practices.
- Mindset for automation, reliability engineering and continuous improvement; team-player attitude.