Data Scientist - AI Evaluation & Benchmarking Manager at PwC

on-site · full time · Visa sponsorship

Apply for this role at PwC

Responsibilities:
- Design and run end-to-end AI benchmarking and experimentation workflows tailored to client use cases.
- Build scalable evaluation frameworks, metrics and repeatable pipelines that combine engineering and research best practices.
- Develop, maintain and harden experimentation infrastructure; ensure reproducibility, CI/CD integration and containerised deployments.
- Analyse results using statistical methods and experimental design; translate findings into clear, actionable recommendations and client-ready reports.
- Support technical demos, deep-dive sessions and influence model selection and strategy across engagements.
- Mentor and collaborate with team members to raise engineering standards and maintain rigorous evaluation practices.

Requirements:
- Strong hands-on experience in data science, LLM experimentation or model benchmarking with structured evaluation frameworks.
- Proficiency in Python, including asynchronous programming and multithreading, and writing maintainable production-grade code.
- Experience deploying ML workloads to cloud platforms (Azure, AWS or GCP) and familiarity with CI/CD and containerisation (Docker/Podman).
- Applied knowledge of statistics, experiment design and evaluation metrics; ability to convert results into business insights.
- Excellent communication skills, ability to manage fast-moving workstreams and operate autonomously.
- Emerging leadership capabilities: influencing technical direction, mentoring colleagues and taking ownership of complex projects.