/roles — ROLE_273
Senior AI/ML Engineer
Seed-stage enterprise AI startup turning messy operational data into structured insight on how work gets done
The role
- COMP
- $180K - $275K
- EQUITY
- Competitive equity
- LOCATION
- San Francisco
- WORKPLACE
- On-site
- EXPERIENCE
- 4 - 10 years
- VISA
- None, Visa transfers
- STACK
- Python, Go, PostgreSQL, Redis, React, TypeScript, GCP, Terraform, Docker, Git, OpenAI, C/C++, Java, Scala, CI/CD
- INDUSTRY
- AI, Enterprise, B2B
The company
Seed-stage startup that maps how very large enterprises really get work done and turns that know-how into a structured, evolving blueprint AI agents can act on.
- STAGE
- Series A-stage
- FUNDING
- $5M raised
- TEAM
- ~10 people
- FOUNDED
- 2024
JD — the work
About the role
You would own the AI layer of an enterprise process-intelligence platform: retrieval, context assembly, model and provider choice, structured extraction, evaluation, orchestration, and monitoring that convert disorganized enterprise data into a structured picture of how work happens. The team builds on hosted frontier models instead of training its own, so the focus is on AI that behaves predictably, can be measured, and holds up in production. It is an early, on-site role at a seed-stage company of about 10 people.
What you'll do
- Own retrieval, context assembly, and model and provider selection across the product
- Build extraction pipelines that pull trustworthy, structured information out of long, messy documents
- Define what good looks like, assemble representative datasets, and run evaluations with LLM judges and human review
- Put guardrails, regression suites, launch gates, and fallbacks around every AI feature
- Make multi-step, tool-using agent workflows dependable even when intermediate steps fail
- Track latency, cost, retries, and failures in production, and debug whatever breaks
What they're looking for
- 4 to 6+ years of production engineering across software, ML, or applied AI
- A record of running dependable AI services in production, with good visibility into how they behave
- Hands-on experience shaping product behavior with hosted LLMs via retrieval, tool use, structured outputs, or multi-step flows
- Evaluation know-how: golden datasets, regression testing, hallucination reduction, and output verification
- End-to-end pipeline thinking, from incoming client data to what users finally see
- Skill at converting vague product goals into experiments and measurable, shipped gains
Nice to have
- Retrieval depth: chunking strategies, reranking, embeddings, and hybrid search
- Background in enterprise workflows, including process or task mining
- Multimodal work with transcripts, screenshots, video, or mixed document sets
- Backend services in Go, plus everyday use of AI coding assistants