About this job
Nimbus Labs builds the monitoring plane for large-scale machine-learning systems. As a Staff ML Engineer on the Platform & ML team, you will own the training and evaluation infrastructure used by every model team, from feature pipelines to online metrics.
What you'll do
- Design and drive the roadmap for training infrastructure used by 40+ engineers.
- Build distributed data and evaluation pipelines in Python and Go.
- Own reliability and cost for GPU training clusters and inference services.
- Mentor senior engineers and run design reviews.
What you'll bring
- 8+ years building machine-learning systems, with 2+ at staff/lead level.
- Deep experience with PyTorch, Kubeflow, and GPU scheduling.
- Strong engineering judgment and a track record of shipping.
Benefits
📱 Fully remote with quarterly in-person offsites
🏆 Generous equity package
💵 401(k) match, health/dental/vision for you and your family
🛠️ $2,000 yearly home office budget
Interview process
- Recruiter screen (30 min)
- Technical screen with a staff engineer (60 min)
- Virtual on-site: system design + coding (3 sessions)
- Team & leadership interviews
Questions about this role? Email recruiting@nimbus.example.