/roles — ROLE_340

Machine Learning Engineer - Evals

Research-driven AI lab building a memory and user-modeling layer for apps and agents

The role

COMP
$220K - $300K
EQUITY
Competitive equity
LOCATION
New York
WORKPLACE
On-site
EXPERIENCE
3+ years
VISA
None
STACK
Python, PyTorch, Kafka
INDUSTRY
AI, Devtools

The company

Research-driven AI lab building identity and memory infrastructure that lets AI apps and agents learn rich, evolving models of their users, powered by its own trained models.

STAGE
Series A-stage
FUNDING
$9M raised
TEAM
~15 people
FOUNDED
2023
BACKING
VC-backed

JD — the work

About the role

You would own every part of evaluating the company's core product, which is the yardstick a research-heavy team relies on to know whether its representations of users and entities are really getting better. The system is a multi-agent setup whose behavior is opaque and hard to inspect, so the core challenge is deciding what improvement even looks like for representations that evolve, with no ground truth to lean on, then building the tooling to measure it. It suits someone who enjoys building measurements from scratch, digging through traces to locate the real problems, and shipping fixes personally.

What you'll do

  • Design the evals, choosing the scores and the definition of progress for entity representations as they evolve
  • Build pipelines and harnesses for ingestion, labeling, versioning, reruns, and judge models, so new questions get quick answers
  • Run the evals, study traces and results, and recommend changes that sharpen representation quality
  • Expand simulated agents so the team can model and score far more entities at the same time
  • Carry each investigation yourself, from the question and pipeline to reruns and the final writeup

What they're looking for

  • 3+ years of ML engineering experience, ideally including evaluation systems
  • Strong Python and PyTorch, with familiarity with streaming data tools such as Kafka
  • Comfort designing metrics where no ground truth exists
  • A habit of reading raw traces and outputs to diagnose failures
  • Preference for owning work end to end, including shipping fixes yourself
APPLY FOR THIS ROLE →All open rolesOne application covers up to 3 roles.