The role
- COMP
- $150K - $210K
- EQUITY
- Competitive equity
- LOCATION
- San Francisco
- WORKPLACE
- On-site
- EXPERIENCE
- 2 - 6 years
- VISA
- None, Visa transfers
- STACK
- Python, Django, React, TypeScript, AWS, Git, FastAPI
- INDUSTRY
- AI, Software Development
/roles — ROLE_331
Growing AI evaluation company benchmarking frontier language models on real professional tasks
AI evaluation company that benchmarks leading models on demanding, domain-specific tasks in law, finance, healthcare, software, and more, building most of its benchmarks in-house.
JD — the work
You would join the platform team as a generalist engineer, building the infrastructure that runs large-scale LLM evaluations. The work covers the whole stack, from backend services in Python to React on the front end, in a high-autonomy setting where you ship fast. It is a pure individual-contributor role, hands-on every day with open-ended problems. The team looks for evidence of excellence, for example a standout employer, strong research, an elite school, or a notable side project.