
Cloud Inference Engineer
LuminalSoftware
Pay
$150K–$250K
per year
Work mode
On-site
Full Time
Location
Level
Entry
San Francisco, CA at a glance
- Rent
- #2 of 51$2,680/mo+46% vs US avg
- Weather
- #17 of 51295 mild days0 hot · 0 cold
- Income tax
- #1 of 5113.3% top rateCalifornia
What you need
Torch/PyTorchCUDA
Qualifications
- CUDA + GPU inference optimization
- vLLM, SGLang, or TensorRT-LLM experience
- KV caching, paged attention, batching, token streaming, etc.
- Distributed compute (with GPUs is a super plus)
- No degree required
Company
Luminal (YC S25) builds an AI compiler and serving stack that makes models 10x faster and production ready with one line.
Role
Founding, on site in downtown SF. Ship low latency, high throughput model serving on Luminal Cloud.
Day to day responsibilities:
- Deploy and tune models with optimizations like KV caching, paged attention, sequence packing, etc.
- Conducting model performance reviews
- Improve scheduler, batcher, autoscaling; profile latency, cost, utilization
- Sometimes write kernels and, yes, occasional tasteful shitposting
Optimize your resume for this job
Get a match score and the keywords you're missing
About Luminal
Luminal provides a machine learning compiler and serverless cloud platform that automates PyTorch model optimization and deployment.
Similar Software roles

Thunder Compute
Software
Cloud Infrastructure Engineer
San Francisco, CA Full Time
6 months ago
Software
6 months ago
Character.AI
Software
Machine Learning Infrastructure Engineer
Redwood City, CA$150K - $350K Full Time
6 months ago
Software
6 months ago
Cerebras
Software
Principal Engineer, Inference Cloud
Sunnyvale, CA Full Time
1 month ago
Software
1 month ago