Data Scientist at Kaizo

on-site · full_time · Visa sponsorship

Apply for this role at Kaizo

Responsibilities:
- Translate customer QA rubrics into precise, testable instructions for LLM-based AutoQA systems.
- Design and run experiments on large, representative datasets to measure and improve model accuracy (LangSmith, BigQuery).
- Curate and build golden datasets and synthetic examples to cover rare but important cases and imbalanced labels.
- Implement LLM-as-a-judge pipelines to produce repeatable, auditable evaluation metrics and enable model-driven decisions.
- Work with AI engineers to turn experiment results into product changes (prompts, tool usage, retrieval improvements) and validate end-to-end quality.
- Join customer calls to understand quality definitions and communicate results to technical and non-technical stakeholders.

Requirements:
- 0–2 years experience; strong fundamentals in LLMs, applied statistics, and experimental design.
- Proficient in Python and SQL; experience with data tooling (Pandas, NumPy) and BigQuery preferred.
- Experience with evaluation methodologies: classifier metrics, precision/recall trade-offs, sampling for rare events.
- Familiarity with LLM Ops tooling (e.g., LangSmith), RAG components, and synthetic data generation is strongly valued.
- Strong analytical thinking, clear written/verbal communication, and the ability to work closely with customers and engineering teams.