/roles — ROLE_317

Member of Technical Staff, ML Systems

Seed-stage team making world-model and video generation run far faster and cheaper on GPUs

The role

COMP
$180K - $230K
EQUITY
Competitive
LOCATION
South Bay Area
WORKPLACE
On-site
EXPERIENCE
1+ years
VISA
None, Visa transfers, New visa sponsorships
STACK
PyTorch
INDUSTRY
AI, Devtools, Robotics, API SDK

The company

Seed-stage startup rebuilding AI compute software for world-model, video, and image workloads, spanning GPU kernels to distributed serving, with customers in generative media and robotics.

STAGE
Series A-stage
FUNDING
$10M+ raised

JD — the work

About the role

You would be responsible for making every layer of a GPU compute stack, built for world-model and video and image generation, fast and cheap to run, from hand-tuned kernels up to multi-node inference and training systems. The company is a small, in-person founding team that recently emerged from stealth with public benchmarks and a production API, and you would work directly with founders whose experience covers research, cloud infrastructure, kernel work, and distributed systems. It suits someone more excited by making a video model run an order of magnitude faster than by training it.

What you'll do

  • Push GPU efficiency for training and serving world-model, video, and image workloads
  • Profile kernels, memory, systems, and whole clusters with Nsight and similar tools to remove bottlenecks
  • Write CUDA and Triton optimizations that ship on production code paths
  • Build distributed inference and training engines for diffusion models spanning many GPUs and nodes
  • Own communication speed across NCCL, RDMA on InfiniBand or RoCE fabrics, and disaggregated serving setups
  • Set up benchmarks and regression tests that keep performance gains from eroding in production

What they're looking for

  • 1+ years of hands-on GPU performance or ML systems engineering
  • Working knowledge of CUDA or Triton and GPU profiling tools
  • Experience with PyTorch and distributed training or inference across multiple nodes
  • Familiarity with GPU communication such as NCCL and RDMA networking
  • Willingness to work in person with a small founding team in the Bay Area
APPLY FOR THIS ROLE →All open rolesOne application covers up to 3 roles.