/roles — ROLE_331

Member of Technical Staff - Platform

Growing AI evaluation company benchmarking frontier language models on real professional tasks

The role

COMP
$150K - $210K
EQUITY
Competitive equity
LOCATION
San Francisco
WORKPLACE
On-site
EXPERIENCE
2 - 6 years
VISA
None, Visa transfers
STACK
Python, Django, React, TypeScript, AWS, Git, FastAPI
INDUSTRY
AI, Software Development

The company

AI evaluation company that benchmarks leading models on demanding, domain-specific tasks in law, finance, healthcare, software, and more, building most of its benchmarks in-house.

STAGE
Series A-stage
FUNDING
$5M raised
TEAM
~25 people
FOUNDED
2024
BACKING
VC-backed

JD — the work

About the role

You would join the platform team as a generalist engineer, building the infrastructure that runs large-scale LLM evaluations. The work covers the whole stack, from backend services in Python to React on the front end, in a high-autonomy setting where you ship fast. It is a pure individual-contributor role, hands-on every day with open-ended problems. The team looks for evidence of excellence, for example a standout employer, strong research, an elite school, or a notable side project.

What you'll do

  • Maintain and extend the system that executes benchmark runs, spanning Python libraries, the web app, cloud infra, and tools
  • Work across Django backend services and React/TypeScript frontend features as team needs shift
  • Own vague, open-ended problems from design through shipping without much hand-holding
  • Review teammates' code and architecture designs
  • Partner with researchers so the platform supports the evaluations they need to run

What they're looking for

  • 2 to 6 years of software engineering experience as a generalist
  • Strong Python (Django or FastAPI) and React with TypeScript
  • Experience with AWS and with shipping fast under high autonomy
  • Preference for hands-on individual contributor work
APPLY FOR THIS ROLE →All open rolesOne application covers up to 3 roles.