/roles — ROLE_337

Member of ML Technical Staff

Tiny, venture-backed lab scaling extremely efficient low-bit language models that can run on-device

The role

COMP
$200K - $350K
EQUITY
Competitive equity
LOCATION
San Francisco
WORKPLACE
On-site
EXPERIENCE
1 - 5 years
VISA
None, Visa transfers
STACK
C/C++, PyTorch
INDUSTRY
AI

The company

Very small, venture-backed AI lab training highly efficient low-bit language models that need far less hardware and can run entirely on edge devices.

STAGE
Series A-stage
FUNDING
$9M raised
TEAM
founding team of <10
FOUNDED
2025

JD — the work

About the role

The team is rethinking every layer of how models are built, including data, optimization, infrastructure, pretraining, and post-training, to make heavily compressed language models competitive with the strongest open-weight models from US labs. You would join a very small team where each person has an unusually direct line to model quality, focusing on pretraining, post-training, or infrastructure. The role is on-site in San Francisco.

What you'll do

  • Pretraining: train large models, choosing and combining parallelism strategies to suit the model's shape and the cluster topology
  • Pretraining: co-design software and hardware for throughput through kernels, compute and communication overlap, precision, and sharding
  • Post-training: design RL environments, build training infrastructure, and run large-scale RL jobs for language models
  • Post-training: support both async and sync RL with multi-node training, rollout serving, and weight sync
  • Infrastructure: tune inference engines for speculative decoding, prefill and decode disaggregation, KV-cache pressure, batching, and long context
  • Write low-level kernels in CUDA, Triton, or similar DSLs, and keep long training jobs healthy

What they're looking for

  • Current with papers and technical reports from the past month in your focus area
  • A research mindset backed by written work such as technical blogs or publications
  • Pretraining track: experience training large models, or intense academic work on smaller ones, plus model parallelism knowledge
  • Post-training track: hands-on RL for LLMs and a clear view of PPO, GRPO, reward design, and stability
  • Infrastructure track: deep understanding of inference engines under load, plus Triton, CUDA, CuTe, PTX, TileLang, or C++

Nice to have

  • LLM work at a major AI lab or in an academic group
  • Contributions to open efficient-pretraining efforts or training repos such as OLMo, TorchTitan, or Megatron
  • Contributions to open inference engines such as vLLM or SGLang, or RL frameworks such as veRL or NeMo-RL
  • Adjacent research in scaling behavior, data mixing, architecture, tokenization, tool use, or eval design
APPLY FOR THIS ROLE →All open rolesOne application covers up to 3 roles.