Staff Site Reliability Engineer (AI Platform) at Manychat

on-site · full-time · Visa sponsorship

Apply for this role at Manychat

Responsibilities:
- Own reliability, performance and cost of the AI platform: AI Gateway, inference services and provider integrations.
- Design and evolve the AI Gateway (routing, provider failover, rate limiting, caching, guardrails).
- Define and implement observability and SLOs for models and providers (latency, throughput, token metrics, quality signals).
- Drive FinOps for AI workloads: per-feature cost visibility, model sizing, caching and provider mix optimization.
- Lead capacity planning, incident response, runbooks and postmortems for inference services.
- Mentor teams, set platform standards and review designs for LLM-backed features.

Requirements:
- 5+ years in SRE/platform/infrastructure engineering with production ownership at scale.
- Hands-on experience operating LLM-backed systems and provider APIs (Bedrock, OpenAI, Anthropic or similar).
- Strong cloud-native skills: AWS, Kubernetes, Terraform/IaC and CI/CD pipelines.
- Proven observability practice (Prometheus/Grafana, OpenTelemetry) and experience defining SLOs for non-deterministic systems.
- Experience with cost optimization, capacity planning and incident management for inference workloads.
- Strong communication and collaboration skills to drive standards across engineering teams.