/roles — ROLE_6
Member of Technical Staff - Model Optimization and Inference (Experienced)
Face-to-face AI interaction that feels human
The role
- COMP
- $250K - $350K
- LOCATION
- Seattle
- WORKPLACE
- On-site
- EXPERIENCE
- 2+ years
- VISA
- None, Visa transfers
- STACK
- Kubernetes, K8s, Terraform, Python, Rust, Go, Airflow, PyTorch, vLLM, SGLang, TensorRT-LLM, CUDA
- INDUSTRY
- Software Development, AI
The company
Applied-AI lab building visual conversational AI — real-time, face-to-face interaction that feels human.
- STAGE
- growth-stage
- FUNDING
- $60M+ raised
- TEAM
- ~25 people
- FOUNDED
- 2024
- NOTE
- research team with PhDs from top programs
JD — the work
About the role
An ML-systems role at a research lab building real-time, photorealistic conversational AI: own inference performance across the entire model stack — LLMs, audio models, and diffusion components — from serving frameworks down to custom kernels, for a product where latency is the product.
What you'll do
- Own end-to-end inference optimization across LLM, audio, and diffusion models
- Design KV-cache strategy for long conversations: eviction, compression, memory-efficient attention
- Deploy and extend serving frameworks (vLLM, SGLang, TensorRT-LLM) for unusual workloads
- Profile latency and throughput; systematically eliminate bottlenecks
- Accelerate diffusion inference: consistency models, step distillation, caching, custom kernels
- Apply quantization (INT8/INT4, GPTQ, AWQ) without meaningful quality loss
What they're looking for
- 2+ years building production ML systems
- Scalable-infrastructure design from scratch, with well-reasoned technology choices
- Track record optimizing for latency, throughput, and cost
- Broad ML-infra fluency: inference, real-time streaming, data engineering
Nice to have
- Video or audio model experience
- CUDA kernel-level optimization