The role
- COMP
- $140K - $185K
- EQUITY
- Competitive equity
- LOCATION
- San Francisco
- WORKPLACE
- On-site
- EXPERIENCE
- 0 - 4 years
- VISA
- None, Visa transfers
- STACK
- Python, Django, React, AWS, Git, OpenAI
- INDUSTRY
- AI, Software Development
/roles — ROLE_442
AI benchmarking company publishing public model leaderboards covering law, tax, finance, coding, and more
AI evaluation company that benchmarks leading models on demanding, domain-specific tasks in law, finance, healthcare, software, and more, building most of its benchmarks in-house.
JD — the work
You would own the public model leaderboards the company publishes, running every new model through its legal, tax, finance, and coding benchmarks soon after release. That means analyzing how models fail, weighing where each model is strong or weak, and partnering with communications to publish findings. Research labs, enterprises, and startups all rely on the results; the company works with the big model labs, hospital systems, and major financial institutions, and its work has been covered by major national press. Expect intense sprints after big model launches and calmer periods in between.