/roles — ROLE_7

Member of Technical Staff — RL Research (New PhD Grad)

Face-to-face AI interaction that feels human

The role

COMP
$250K - $350K
EQUITY
Highly competitive equity
LOCATION
Seattle
WORKPLACE
On-site
EXPERIENCE
0 - 2 years
VISA
None, Visa transfers
STACK
Python, PyTorch, RL/post-training frameworks
INDUSTRY
Software Development, AI

The company

Applied-AI lab building visual conversational AI — real-time, face-to-face interaction that feels human.

STAGE
growth-stage
FUNDING
$60M+ raised
TEAM
~25 people
FOUNDED
2024

JD — the work

About the role

A rare 0→1 post-training role for a finishing or recent PhD: build the RL and post-training stack of a frontier lab whose models are omni from the ground up — audio, video, language, and real-time full-duplex interaction. The problems go beyond text: timing, interruption, emotional response, and audiovisual coherence are all training targets.

What you'll do

  • Build the RL/post-training stack from scratch: rollouts, policy optimization, reward serving, feedback loops, evaluation, observability
  • Develop and scale methods like PPO, GRPO, DPO, rejection sampling, RLHF/RLAIF, and online RL
  • Design the abstractions connecting research ideas to production-scale runs: trainers, rollout workers, reward models, experience buffers
  • Build evaluation loops for interactive behavior: turn-taking, interruption, timing, emotional response
  • Optimize the full loop across rollout throughput, serving latency, and GPU utilization

What they're looking for

  • Completing or recently completed PhD with deep RL/post-training grounding
  • Systems ability to build and debug distributed training infrastructure
  • Drive to turn evolving research ideas into reliable, fast tooling
APPLY FOR THIS ROLE →All open rolesOne application covers up to 3 roles.