/roles — ROLE_337
Member of ML Technical Staff
Tiny, venture-backed lab scaling extremely efficient low-bit language models that can run on-device
The role
- COMP
- $200K - $350K
- EQUITY
- Competitive equity
- LOCATION
- San Francisco
- WORKPLACE
- On-site
- EXPERIENCE
- 1 - 5 years
- VISA
- None, Visa transfers
- STACK
- C/C++, PyTorch
- INDUSTRY
- AI
The company
Very small, venture-backed AI lab training highly efficient low-bit language models that need far less hardware and can run entirely on edge devices.
- STAGE
- Series A-stage
- FUNDING
- $9M raised
- TEAM
- founding team of <10
- FOUNDED
- 2025
JD — the work
About the role
The team is rethinking every layer of how models are built, including data, optimization, infrastructure, pretraining, and post-training, to make heavily compressed language models competitive with the strongest open-weight models from US labs. You would join a very small team where each person has an unusually direct line to model quality, focusing on pretraining, post-training, or infrastructure. The role is on-site in San Francisco.
What you'll do
- Pretraining: train large models, choosing and combining parallelism strategies to suit the model's shape and the cluster topology
- Pretraining: co-design software and hardware for throughput through kernels, compute and communication overlap, precision, and sharding
- Post-training: design RL environments, build training infrastructure, and run large-scale RL jobs for language models
- Post-training: support both async and sync RL with multi-node training, rollout serving, and weight sync
- Infrastructure: tune inference engines for speculative decoding, prefill and decode disaggregation, KV-cache pressure, batching, and long context
- Write low-level kernels in CUDA, Triton, or similar DSLs, and keep long training jobs healthy
What they're looking for
- Current with papers and technical reports from the past month in your focus area
- A research mindset backed by written work such as technical blogs or publications
- Pretraining track: experience training large models, or intense academic work on smaller ones, plus model parallelism knowledge
- Post-training track: hands-on RL for LLMs and a clear view of PPO, GRPO, reward design, and stability
- Infrastructure track: deep understanding of inference engines under load, plus Triton, CUDA, CuTe, PTX, TileLang, or C++
Nice to have
- LLM work at a major AI lab or in an academic group
- Contributions to open efficient-pretraining efforts or training repos such as OLMo, TorchTitan, or Megatron
- Contributions to open inference engines such as vLLM or SGLang, or RL frameworks such as veRL or NeMo-RL
- Adjacent research in scaling behavior, data mixing, architecture, tokenization, tool use, or eval design