Audio | Multimodal ML Engineer at ConnectHum
on-site · full_time · Visa sponsorship
Apply for this role at ConnectHum
Responsibilities:
- Train and fine-tune large-scale audio and multimodal models using PyTorch and distributed training stacks.
- Design and run experiments (architectures, data mixtures, training strategies) to improve performance and robustness.
- Build and maintain audio data pipelines and preprocessing for scalable training and evaluation.
- Optimize models for production: quantization, distillation, streaming inference and latency reduction.
- Deploy models end-to-end to low-latency serving environments and define meaningful evaluation metrics beyond benchmarks.
- Collaborate closely with research and engineering to move models from research to production.
Requirements:
- 3+ years training deep learning models in audio/speech domains with strong PyTorch experience.
- Hands-on experience with distributed training frameworks and large-scale training workflows.
- Solid understanding of audio signal processing fundamentals and speech modeling.
- Experience shipping models to production with attention to latency and reliability.
- Strong engineering hygiene: clean code, testing, versioning, and data pipeline construction.
- Bonus: multimodal architecture experience, alignment/fine-tuning techniques, or model optimization/infrastructure background.