/roles — ROLE_442

Evaluations Engineer

AI benchmarking company publishing public model leaderboards covering law, tax, finance, coding, and more

The role

COMP
$140K - $185K
EQUITY
Competitive equity
LOCATION
San Francisco
WORKPLACE
On-site
EXPERIENCE
0 - 4 years
VISA
None, Visa transfers
STACK
Python, Django, React, AWS, Git, OpenAI
INDUSTRY
AI, Software Development

The company

AI evaluation company that benchmarks leading models on demanding, domain-specific tasks in law, finance, healthcare, software, and more, building most of its benchmarks in-house.

STAGE
Series A-stage
FUNDING
$5M raised
TEAM
~25 people
FOUNDED
2024
BACKING
VC-backed

JD — the work

About the role

You would own the public model leaderboards the company publishes, running every new model through its legal, tax, finance, and coding benchmarks soon after release. That means analyzing how models fail, weighing where each model is strong or weak, and partnering with communications to publish findings. Research labs, enterprises, and startups all rely on the results; the company works with the big model labs, hospital systems, and major financial institutions, and its work has been covered by major national press. Expect intense sprints after big model launches and calmer periods in between.

What you'll do

  • Run each newly released model through the full benchmark suite
  • Coordinate directly with open and closed model providers during evaluations
  • Dig into model transcripts with analysis tooling to spot recurring failures and trends
  • Help the social team turn notable results into public posts
  • Onboard new models and keep existing model integrations working
  • Improve and maintain benchmark infrastructure for both agentic and non-agentic tasks
  • Collaborate with researchers on building new benchmarks

What they're looking for

  • 0 to 4 years of software engineering experience with strong fundamentals
  • Solid Python, with exposure to Django, React, and AWS
  • Curiosity about model behavior and care in analyzing errors
  • Comfort with bursty, launch-driven weeks and intense sprints
APPLY FOR THIS ROLE →All open rolesOne application covers up to 3 roles.