Data Scientist (Moderation) at Mayflower
on-site · full_time · Visa sponsorship
Apply for this role at Mayflower
Responsibilities:
- Build, train and evaluate ML models for classification, regression, clustering and time-series tasks relevant to content moderation.
- Design feature engineering pipelines and data preprocessing workflows for large, high-throughput datasets.
- Deploy and maintain production models with ML-ops partners; implement monitoring for model performance and data drift.
- Run statistical analyses, hypothesis tests and A/B experiments to validate product and model changes.
- Work with streaming and analytical systems (Kafka, ClickHouse); process and transform data with Python and SQL.
- Investigate anomalies, translate product questions into analytical tasks and communicate findings to product and engineering teams.
Requirements:
- 5+ years experience in applied data science or ML, with strong Python and SQL skills and experience in notebook-driven workflows.
- Hands-on experience with pandas and large-scale data processing; familiarity with ClickHouse or similar columnar databases.
- Experience with streaming platforms (Kafka) and working with event-driven data.
- Proficiency with ML libraries (scikit-learn, XGBoost, LightGBM) and evaluation/experiment design methods.
- Strong understanding of statistics, probability and experiment design; ability to monitor, backtest and iterate models in production.
- Clear communicator able to work cross-functionally and contribute to ML best practices.
Nice to have:
- Experience with deep learning (PyTorch/TensorFlow), computer vision, model deployment at scale or distributed processing.