/roles — ROLE_76

Member of Technical Staff - Model Optimization and Inference (New Grad)

Face-to-face AI interaction that feels human

The role

COMP
$200K - $300K
EQUITY
Competitive Equity
LOCATION
Seattle
WORKPLACE
On-site
EXPERIENCE
0 - 2 years
VISA
None, Visa transfers
STACK
Kubernetes, Terraform, Python, Rust, Go, Airflow, PyTorch
INDUSTRY
Software Development, AI

The company

Applied-AI lab building visual conversational AI — real-time, face-to-face interaction that feels human.

STAGE
growth-stage
FUNDING
$60M+ raised
TEAM
~25 people
FOUNDED
2024

JD — the work

About the role

For early-career engineers who want to squeeze every millisecond out of trained models: full-stack inference work — quantization, KV cache, kernels, batching — at a research lab building real-time conversational AI. BS/MS/PhD all welcome; systems intuition matters more than credentials.

What you'll do

  • Contribute to inference optimization across LLM, audio, and diffusion models
  • Implement KV-cache strategies for long-context conversation
  • Work with and extend vLLM/SGLang/TensorRT-LLM-class frameworks
  • Profile latency and throughput; eliminate bottlenecks
  • Build internal profiling and test tooling
  • Apply quantization techniques throughput-first

What they're looking for

  • Finishing or recently finished BS/MS/PhD
  • Framework exposure through coursework, research, internships, or open source — with opinions on where they fall short
  • Appetite to go deep
APPLY FOR THIS ROLE →All open rolesOne application covers up to 3 roles.