Staff Site Reliability Engineer — AI Platform at Manychat, Inc

on-site · full_time · Visa sponsorship

Apply for this role at Manychat, Inc

Responsibilities:
- Own reliability, performance, and cost of the AI platform: AI Gateway, inference services, and provider integrations (Bedrock, Azure OpenAI, Anthropic, etc.).
- Design and evolve the AI Gateway: routing, failover, rate limiting, caching, and guardrails for multi-provider setups.
- Build observability and SLOs for LLM systems: latency/throughput/error metrics, token-level telemetry, quality and drift signals.
- Lead FinOps for AI workloads: per-feature cost visibility, model sizing, caching strategies, and provider mix optimization.
- Run capacity planning, incident response, runbooks, and postmortems for inference infrastructure.
- Evangelize standards, review designs, and coach teams to raise reliability across the org.

Requirements:
- 5+ years in SRE, platform, or infrastructure engineering with production ownership at scale.
- Hands-on experience operating LLM-backed systems and provider APIs (Bedrock, OpenAI, Anthropic, or similar).
- Strong cloud-native skills: AWS, Kubernetes, Terraform/IaC, and CI/CD pipelines.
- Deep observability practice (Prometheus/Grafana, OpenTelemetry) and experience defining SLOs for non-deterministic systems.
- Demonstrated experience with cost optimization for ML/AI workloads and capacity planning for inference.
- Proven ability to lead technical direction, write runbooks, and drive cross-team reliability improvements (on-site in Amsterdam).