DataOps Engineer (AI Platform Engineer) at Exness jobs for internal candidates

on-site · full-time · Visa sponsorship

Apply for this role at Exness jobs for internal candidates

Responsibilities:
- Design, deploy and operate an on-prem AI platform for multi-node GPU clusters and Kubernetes.
- Collaborate on GPU server selection, high-performance networking and RDMA-enabled cluster configuration.
- Manage GPU partitioning/MIG, scheduling and runtime integration to maximise utilization and reliability.
- Deploy and maintain model serving runtimes (vLLM, ONNX, SGLang, Nvidia Triton, KServe) and benchmarking setups.
- Build CI/CD pipelines and tooling for model packaging, versioning, deployment and reproducible delivery.
- Implement model lifecycle tooling: experiment tracking, registries and support for fine-tuning (e.g., LoRA).
- Create observability and tracing for GPU workloads and inference flows (utilisation, memory, throughput, latency).
- Run performance tests and benchmarks across runtimes and hardware; identify and resolve bottlenecks.
- Ensure platform security and compliance for model serving and data handling.

Requirements:
- Bachelor’s or Master’s in CS, Engineering or related field and 5+ years in infrastructure/platform or distributed systems.
- Strong Kubernetes experience running production workloads and Linux proficiency.
- Programming skills in Python and/or Go; experience building automation and CI/CD for ML workloads.
- Hands-on experience with GPU infrastructure (NVIDIA/AMD), multi-GPU clusters and MIG configurations.
- Familiarity with model serving runtimes (Triton, KServe, vLLM, ONNX) and model lifecycle tools (MLflow).
- Understanding of high-performance networking, RDMA, GPU scheduling and performance benchmarking.
- Good troubleshooting, observability and system optimisation skills.