/roles — ROLE_286
RL Engineer
Early-stage company producing training data and evaluations that frontier AI labs use to improve models
The role
- COMP
- $260K - $290K
- EQUITY
- Competitive equity
- LOCATION
- San Francisco
- WORKPLACE
- On-site
- EXPERIENCE
- 1 - 8 years
- VISA
- None, Visa transfers
- STACK
- Python, TypeScript, C/C++, AWS
- INDUSTRY
- AI, Data
The company
Frontier-data company capturing how experts reason and turning real professional work into training data for foundation models.
- STAGE
- growth-stage
- FUNDING
- $30M+ raised
- TEAM
- ~80 people
- FOUNDED
- 2024
JD — the work
About the role
You would shape how frontier models learn by designing datasets, reward signals, and grading rubrics, collaborating directly with research groups at leading AI labs. Ideas move quickly from a hunch to a running experiment, and your work flows straight into large-scale training. The team is small and early, so individual engineers have real influence on how models get better. The role is on-site in San Francisco.
What you'll do
- Slice data to expose where models genuinely fail on financial, coding, and enterprise tasks
- Develop reward signals and grading rubrics used in RLHF/RLVR training
- Study how annotators behave and test ways to strengthen particular model skills
- Measure dataset diversity and quality, and how each dataset changes alignment and capability downstream
- Operate pipelines for both synthetic and real-world data
- Translate what lab researchers want from training into specific data and eval requirements
What they're looking for
- 1 to 8 years of engineering experience in ML or software
- Deep backend or full-stack skill in TypeScript, Python, or another comparable language
- A degree in computer science is required
- A bias for doing, including the unglamorous and tedious parts of the job
Nice to have
- AI side projects or research papers
- Hands-on RL environment building, or a researcher's instinct for designing agent tasks
- Frameworks you built to score dataset diversity or quality