/roles — ROLE_311

ML Infrastructure Engineer

Well-funded robotics startup training embodied AI foundation models for robot arms already working at customer sites

The role

COMP
$220K - $350K
EQUITY
Competitive equity
LOCATION
San Francisco · South Bay Area
WORKPLACE
On-site
EXPERIENCE
5+ years
VISA
None, Visa transfers
STACK
PyTorch, Kubernetes, GCP, AWS, Docker, Python
INDUSTRY
AI, Robotics, Hardware

The company

Heavily funded robotics company from repeat founders, building foundation models and low-cost robot arms that automate repetitive tasks for commercial customers, starting with food service and hospitality.

STAGE
scale-up
FUNDING
$140M+ raised
TEAM
100+ people
FOUNDED
2024
BACKING
VC-backed

JD — the work

About the role

You would take full ownership of training infrastructure and make GPU capacity spread across several clouds work as one fast, dependable engine for large multimodal models. The role connects researchers with compute: distributed training, GPU efficiency, and the pipelines that pull in terabytes of robot data. Every research effort here is aimed at real deployment, so your work directly shortens the path from a trained model to a robot on a customer site.

What you'll do

  • Architect distributed training for big GPU clusters, applying FSDP, ZeRO, activation checkpointing and sharding to multimodal models
  • Give researchers easy tooling and Kubernetes or SLURM scheduling that supports quick iteration, automatic retries and graceful recovery
  • Build fast data pipelines that move terabytes of robot video, proprioception and 3D sensor data so GPUs stay fed
  • Speed up on-robot inference for live control through distillation, quantization and compiling models with TensorRT and Triton
  • Profile the stack in depth, tracking down memory fragmentation, I/O stalls and idle GPU time as the fleet grows

What they're looking for

  • 7+ years working on ML infrastructure or high-performance computing
  • A history of leading technical projects in that area
  • Hands-on distributed training expertise, including memory optimization in the FSDP or ZeRO style
  • Comfort with PyTorch, Kubernetes, Docker and at least one major cloud
  • Genuine enthusiasm for robotics and physical AI
  • Comfort with a demanding pace of roughly 55 hours a week, balanced by unlimited time off and flexibility
APPLY FOR THIS ROLE →All open rolesOne application covers up to 3 roles.