Senior / Lead Machine Learning Engineer, Serving - Germany at Inworld AI
on-site · full_time · Visa sponsorship
Apply for this role at Inworld AI
Responsibilities:
- Lead efforts to optimize multimodal model serving for low-latency, high-throughput production deployments.
- Convert research models into production-ready services: containerization, profiling, quantization and performance tuning.
- Implement serving improvements: continuous batching, caching, paged attention, speculative decoding and inference framework tuning (vLLM/TRT-LLM).
- Design and operate scalable inference clusters across GPUs and nodes using Kubernetes, Ray or custom orchestration.
- Profile and optimize GPU-accelerated code (C++, CUDA, Rust or optimized Python) and ensure robust monitoring and alerting.
- Collaborate closely with research, product and SRE teams to meet SLA, cost and reliability goals.
Requirements:
- Strong experience with inference optimization and model acceleration techniques (quantization, distillation, batching).
- Proficiency in systems-level programming and GPU profiling (C++, CUDA, Rust or high-performance Python).
- Hands-on experience with distributed inference at scale: Kubernetes, Ray, multi-GPU/multi-node setups and custom load balancing.
- Familiarity with advanced serving strategies (caching, paged attention, speculative decoding) and memory management for large models.
- Proven track record of shipping production inference systems; open-source contributions or publications are advantageous.
- PhD or equivalent practical experience in CS/Math/Physics and excellent English communication skills.