/roles — ROLE_327
Member of Technical Staff - Research Scientist
Early-stage startup that benchmarks large language models for enterprises and leading AI labs
The role
- COMP
- $150K - $200K
- EQUITY
- 0.15-0.5% (with flexibility upwards)
- LOCATION
- San Francisco
- WORKPLACE
- On-site
- EXPERIENCE
- 0 - 3 years
- VISA
- None, Visa transfers
- STACK
- Python, Git, PyTorch, TensorFlow, Django, AWS, React, TypeScript
- INDUSTRY
- AI, Software Development
The company
AI evaluation company that benchmarks leading models on demanding, domain-specific tasks in law, finance, healthcare, software, and more, building most of its benchmarks in-house.
- STAGE
- Series A-stage
- FUNDING
- $5M raised
- TEAM
- ~25 people
- FOUNDED
- 2024
- BACKING
- VC-backed
JD — the work
About the role
You would help decide how advanced AI systems get tested before enterprises put them into production, building benchmarks and evaluation methods for LLMs. The work ranges from assessing new models as soon as they launch to creating benchmarks from scratch and improving automatic scoring of generated text, in close partnership with enterprise customers, AI labs, and the engineering team. The small team is based in San Francisco and works on-site, with relocation or commuting support available.
What you'll do
- Assess new models from major AI labs soon after release
- Build benchmarks from the ground up: recruit labelers, assemble datasets, and write white papers
- Improve the methods used to automatically judge generated text
- Partner with engineering to put evaluation methods into production and scale them up
- Learn what enterprise customers and AI labs need to measure, and design for it
What they're looking for
- 0 to 3 years in applied AI, ideally focused on benchmarking, evaluation, or language models
- A CS or ML degree (or related field); master's or PhD preferred, or a bachelor's plus 1 to 3 years
- Genuine interest in LLM infrastructure and working knowledge of it
- An applied focus that values shipped work over purely academic publishing
- Strong communication and comfort giving and receiving feedback
- Happy to be in the San Francisco office in person
Nice to have
- Strong, maintainable Python, plus Django, Flask, or another Python web server
- Team development habits such as sprints, Git hygiene, and code review
- NLP research or publications, especially on new benchmarks or evaluation methods
- Experience with TensorFlow or PyTorch, language modeling, or diffusion models