/roles — ROLE_311
ML Infrastructure Engineer
Well-funded robotics startup training embodied AI foundation models for robot arms already working at customer sites
The role
- COMP
- $220K - $350K
- EQUITY
- Competitive equity
- LOCATION
- San Francisco · South Bay Area
- WORKPLACE
- On-site
- EXPERIENCE
- 5+ years
- VISA
- None, Visa transfers
- STACK
- PyTorch, Kubernetes, GCP, AWS, Docker, Python
- INDUSTRY
- AI, Robotics, Hardware
The company
Heavily funded robotics company from repeat founders, building foundation models and low-cost robot arms that automate repetitive tasks for commercial customers, starting with food service and hospitality.
- STAGE
- scale-up
- FUNDING
- $140M+ raised
- TEAM
- 100+ people
- FOUNDED
- 2024
- BACKING
- VC-backed
JD — the work
About the role
You would take full ownership of training infrastructure and make GPU capacity spread across several clouds work as one fast, dependable engine for large multimodal models. The role connects researchers with compute: distributed training, GPU efficiency, and the pipelines that pull in terabytes of robot data. Every research effort here is aimed at real deployment, so your work directly shortens the path from a trained model to a robot on a customer site.
What you'll do
- Architect distributed training for big GPU clusters, applying FSDP, ZeRO, activation checkpointing and sharding to multimodal models
- Give researchers easy tooling and Kubernetes or SLURM scheduling that supports quick iteration, automatic retries and graceful recovery
- Build fast data pipelines that move terabytes of robot video, proprioception and 3D sensor data so GPUs stay fed
- Speed up on-robot inference for live control through distillation, quantization and compiling models with TensorRT and Triton
- Profile the stack in depth, tracking down memory fragmentation, I/O stalls and idle GPU time as the fleet grows
What they're looking for
- 7+ years working on ML infrastructure or high-performance computing
- A history of leading technical projects in that area
- Hands-on distributed training expertise, including memory optimization in the FSDP or ZeRO style
- Comfort with PyTorch, Kubernetes, Docker and at least one major cloud
- Genuine enthusiasm for robotics and physical AI
- Comfort with a demanding pace of roughly 55 hours a week, balanced by unlimited time off and flexibility