/roles — ROLE_72

Member of Technical Staff

A company that builds AI agents that work as autonomous performance engineers, optimizing GPU kernels for AI inference

The role

COMP
$200K - $200K
EQUITY
0.5 - 1%
LOCATION
San Francisco
WORKPLACE
On-site
EXPERIENCE
1 - 5 years
VISA
None, Visa transfers
STACK
Python, AWS, C/C++, CUDA, HIP, Triton, vLLM-class serving, Kubernetes
INDUSTRY
AI, Hardware

The company

AI-infrastructure company building the fastest inference for open models in the enterprise.

STAGE
seed-stage
FUNDING
seed funding
TEAM
founding team of <10
FOUNDED
2025

JD — the work

About the role

Maximize intelligence per watt: serve open-source LLM inference at the best performance per dollar through autonomous optimization of heterogeneous hardware — NVIDIA, AMD, TPU, Trainium and beyond. Small SF team, complete autonomy, kernels to customers.

What you'll do

  • Ship day-zero support for new open models, tuned for latency and throughput
  • Optimize serving: batching, KV cache, speculative decoding, quantization
  • Write and tune kernels in CUDA, HIP, and Triton across vendors
  • Design and operate heterogeneous clusters
  • Run production inference at scale: reliability, observability, cost per token

What they're looking for

  • Deep systems/GPU engineering ability
  • First-principles problem solving with unreasonable standards
  • San Francisco, on-site 5 days/week
APPLY FOR THIS ROLE →All open rolesOne application covers up to 3 roles.