Skip to main content

Cloud Inference Engineer

LuminalSoftware
Pay
$150K–$250K
per year
Work mode
On-site
Full Time
Level
Entry

San Francisco, CA at a glance

Rent
#2 of 51
$2,680/mo+46% vs US avg
Weather
#17 of 51
295 mild days0 hot · 0 cold
Income tax
#1 of 51
13.3% top rateCalifornia

What you need

Torch/PyTorchCUDA

Qualifications

  • CUDA + GPU inference optimization
  • vLLM, SGLang, or TensorRT-LLM experience
  • KV caching, paged attention, batching, token streaming, etc.
  • Distributed compute (with GPUs is a super plus)
  • No degree required

Company

Luminal (YC S25) builds an AI compiler and serving stack that makes models 10x faster and production ready with one line.

Role

Founding, on site in downtown SF. Ship low latency, high throughput model serving on Luminal Cloud.

Day to day responsibilities:

  • Deploy and tune models with optimizations like KV caching, paged attention, sequence packing, etc.
  • Conducting model performance reviews
  • Improve scheduler, batcher, autoscaling; profile latency, cost, utilization
  • Sometimes write kernels and, yes, occasional tasteful shitposting

Optimize your resume for this job

Get a match score and the keywords you're missing

Optimize resume

About Luminal

Luminal provides a machine learning compiler and serverless cloud platform that automates PyTorch model optimization and deployment.

Similar Software roles

Cloud Inference Engineer
$150K–$250K · Luminal
Apply