Search

Member of Technical Staff - Model Optimization and Inference

PublishedPublished: 6/14/2022
Technology

Job Description

Job DescriptionMember of Technical Staff - Model Optimization and Inference

Company: Nuance Labs
Location: Seattle, WA (in office 5 days per week; relocation assistance available)
Compensation: $250,000 - $350,000 + equity
Employment Type: Full-time
Visa Sponsorship: Visa sponsorship available

About Nuance Labs

Nuance Labs builds photorealistic, real-time AI avatars with emotional intelligence: a full-duplex audiovisual system that can listen, speak, react and respond like a real person. The research team includes PhDs from MIT, UW, Oxford, CMU and Johns Hopkins.

Founded in 2024, Nuance Labs has raised $60M and has about 25 people.

The Role

Nuance Labs is hiring an experienced ML Infrastructure/Systems Engineer (2+ years) to own end-to-end inference optimization across LLMs, audio models and diffusion components, with a focus on latency, throughput and cost.

What You Will Do

  • Own end-to-end inference optimization across the model stack.
  • Implement and tune KV cache strategies for long-context conversations.
  • Evaluate, deploy and extend serving frameworks such as vLLM, SGLang and TensorRT-LLM.
  • Profile and benchmark latency and throughput and remove bottlenecks.
  • Accelerate diffusion inference and apply quantization techniques (INT8, INT4, GPTQ, AWQ).
  • Build internal tooling that makes optimization work faster and more rigorous.

What You Bring

  • 2+ years building and maintaining production ML systems
  • Designing scalable infrastructure from scratch
  • Track record optimizing latency, throughput and cost
  • Debugging distributed systems

Nice to Have

  • Video or audio model experience
  • CUDA kernels and low-level optimization
  • Real-time video streaming (WebRTC)

Benefits

HSA with about $2,000 annual company contribution, 15 days PTO plus public holidays and a company-wide office closure week.

Tech Stack

Kubernetes, Terraform, Python, Rust, Go, Dagster, Ray, Airflow, WebRTC, vLLM, Triton Inference Server, TensorRT

Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...