The role
- COMP
- $200K - $200K
- EQUITY
- 0.5 - 1%
- LOCATION
- San Francisco
- WORKPLACE
- On-site
- EXPERIENCE
- 1 - 5 years
- VISA
- None, Visa transfers
- STACK
- Python, AWS, C/C++, CUDA, HIP, Triton, vLLM-class serving, Kubernetes
- INDUSTRY
- AI, Hardware
/roles — ROLE_72
A company that builds AI agents that work as autonomous performance engineers, optimizing GPU kernels for AI inference
AI-infrastructure company building the fastest inference for open models in the enterprise.
JD — the work
Maximize intelligence per watt: serve open-source LLM inference at the best performance per dollar through autonomous optimization of heterogeneous hardware — NVIDIA, AMD, TPU, Trainium and beyond. Small SF team, complete autonomy, kernels to customers.