/roles — ROLE_327

Member of Technical Staff - Research Scientist

Early-stage startup that benchmarks large language models for enterprises and leading AI labs

The role

COMP
$150K - $200K
EQUITY
0.15-0.5% (with flexibility upwards)
LOCATION
San Francisco
WORKPLACE
On-site
EXPERIENCE
0 - 3 years
VISA
None, Visa transfers
STACK
Python, Git, PyTorch, TensorFlow, Django, AWS, React, TypeScript
INDUSTRY
AI, Software Development

The company

AI evaluation company that benchmarks leading models on demanding, domain-specific tasks in law, finance, healthcare, software, and more, building most of its benchmarks in-house.

STAGE
Series A-stage
FUNDING
$5M raised
TEAM
~25 people
FOUNDED
2024
BACKING
VC-backed

JD — the work

About the role

You would help decide how advanced AI systems get tested before enterprises put them into production, building benchmarks and evaluation methods for LLMs. The work ranges from assessing new models as soon as they launch to creating benchmarks from scratch and improving automatic scoring of generated text, in close partnership with enterprise customers, AI labs, and the engineering team. The small team is based in San Francisco and works on-site, with relocation or commuting support available.

What you'll do

  • Assess new models from major AI labs soon after release
  • Build benchmarks from the ground up: recruit labelers, assemble datasets, and write white papers
  • Improve the methods used to automatically judge generated text
  • Partner with engineering to put evaluation methods into production and scale them up
  • Learn what enterprise customers and AI labs need to measure, and design for it

What they're looking for

  • 0 to 3 years in applied AI, ideally focused on benchmarking, evaluation, or language models
  • A CS or ML degree (or related field); master's or PhD preferred, or a bachelor's plus 1 to 3 years
  • Genuine interest in LLM infrastructure and working knowledge of it
  • An applied focus that values shipped work over purely academic publishing
  • Strong communication and comfort giving and receiving feedback
  • Happy to be in the San Francisco office in person

Nice to have

  • Strong, maintainable Python, plus Django, Flask, or another Python web server
  • Team development habits such as sprints, Git hygiene, and code review
  • NLP research or publications, especially on new benchmarks or evaluation methods
  • Experience with TensorFlow or PyTorch, language modeling, or diffusion models
APPLY FOR THIS ROLE →All open rolesOne application covers up to 3 roles.