Senior / Lead Machine Learning Engineer, Serving - Serbia at Inworld AI
on-site · full_time · Visa sponsorship
Apply for this role at Inworld AI
Responsibilities:
- Optimize and deploy multimodal model serving to meet strict latency and throughput targets (sub-second where required).
- Take models from research to production: containerize, profile, optimize (quantization, distillation) and validate performance.
- Implement advanced serving techniques: continuous batching, caching, paged attention, speculative decoding, and framework tuning (vLLM, TRT-LLM).
- Build and operate scalable inference infrastructure across multi-GPU/multi-node clusters using Kubernetes, Ray or custom schedulers.
- Profile and optimize GPU code paths (C++, CUDA, Rust or highly optimized Python) and integrate monitoring/observability.
- Collaborate with research and product teams to translate model requirements into reliable, cost-effective production systems.
Requirements:
- Demonstrated experience in inference optimization and model acceleration (vLLM/TRT-LLM, quantization, distillation, batching).
- Strong systems programming and profiling skills in C++, CUDA, Rust or optimized Python with GPU experience.
- Experience running large-scale distributed inference: Kubernetes, Ray, multi-GPU/multi-node setups, and custom load balancing.
- Familiarity with performance techniques: caching strategies, paged attention, speculative decoding, and memory management.
- Track record of deploying models to production and owning reliability at scale; public contributions or technical write-ups are a plus.
- PhD in CS/Math/Physics or equivalent practical experience building backend or ML systems; strong English communication skills.