/roles — ROLE_319

Member of Technical Staff, Inference Systems

Stealth inference startup building a new LLM serving engine in Rust from first principles

The role

COMP
$230K - $350K
EQUITY
Competitive
LOCATION
South Bay Area
WORKPLACE
On-site
EXPERIENCE
2 - 10 years
VISA
None, Visa transfers, New visa sponsorships
STACK
Rust, Python, PyTorch, C/C++, Go
INDUSTRY
AI

The company

Stealth-stage company building a high-performance LLM inference platform in Rust, focused on raw speed and serving efficiency, led by founders with deep AI infrastructure experience.

STAGE
Series A-stage
FUNDING
$10M+ raised
FOUNDED
2026

JD — the work

About the role

You would help build an LLM inference platform from scratch, written primarily in Rust and with no legacy code to work around. The role suits someone with a systems background who knows inference internals such as KV caching, attention, batching, and scheduling, and who wants to own the whole serving stack rather than a narrow slice. The company is small, backed by well-known investors, founded by engineers with deep AI infrastructure backgrounds, and still in stealth ahead of announcing its funding and product.

What you'll do

  • Build a new Rust inference runtime, including request routing, the batching and scheduling logic, and serving
  • Design KV cache management and prefix caching that cut latency and per-token cost
  • Extend serving to many GPUs and multiple nodes, working through the distributed problems that come with it
  • Measure and profile the full inference path, then land the speedups the data points to
  • Shape foundational architecture choices alongside the small founding team

What they're looking for

  • 2+ years of systems engineering experience, ideally on inference or model serving
  • Deep knowledge of how inference works under the hood, including KV caching, attention, batching, and scheduling
  • Hands-on work in the internals of SGLang, vLLM, TensorRT-LLM, or comparable engines
  • Strong Rust skills, plus comfort with Python, PyTorch, C/C++, or Go
  • Desire to own the full stack and work on site with a small team
APPLY FOR THIS ROLE →All open rolesOne application covers up to 3 roles.