/roles — ROLE_75
Member of Technical Staff - Model Optimization and Inference (Experienced)
Face-to-face AI interaction that feels human
The role
- COMP
- $250K - $350K
- LOCATION
- Seattle
- WORKPLACE
- On-site
- EXPERIENCE
- 2+ years
- VISA
- None, Visa transfers
- STACK
- Kubernetes, K8s, Terraform, Python, Rust, Go, Airflow, PyTorch, vLLM, SGLang, TensorRT-LLM, CUDA
- INDUSTRY
- Software Development, AI
The company
Applied-AI lab building visual conversational AI — real-time, face-to-face interaction that feels human.
- STAGE
- growth-stage
- FUNDING
- $60M+ raised
- TEAM
- ~25 people
- FOUNDED
- 2024
JD — the work
About the role
ML-systems engineering at a research lab building real-time, photorealistic conversational AI: own inference performance across LLMs, audio, and diffusion — from serving frameworks to custom kernels, where latency is the product.
What you'll do
- Own end-to-end inference optimization across the model stack
- Tune KV-cache strategy: eviction, compression, memory-efficient attention
- Extend serving frameworks (vLLM, SGLang, TensorRT-LLM) for unusual workloads
- Profile and systematically eliminate latency/throughput bottlenecks
- Accelerate diffusion inference: distillation, caching, custom kernels
- Quantize (INT8/INT4, GPTQ, AWQ) without losing quality
What they're looking for
- 2+ years building production ML systems
- Infrastructure design from scratch with sound technology choices
- Latency/throughput/cost optimization track record
Nice to have
- Video/audio models
- CUDA kernel work