
Compensation
$150,000–$250,000/year
Job Description
Qualifications
- CUDA + GPU inference optimization
- vLLM, SGLang, or TensorRT-LLM experience
- KV caching, paged attention, batching, token streaming, etc.
- Distributed compute (with GPUs is a super plus)
- No degree required
Company
Luminal (YC S25) builds an AI compiler and serving stack that makes models 10x faster and production ready with one line.
Role
Founding, on site in downtown SF. Ship low latency, high throughput model serving on Luminal Cloud.
Day to day responsibilities:
- Deploy and tune models with optimizations like KV caching, paged attention, sequence packing, etc.
- Conducting model performance reviews
- Improve scheduler, batcher, autoscaling; profile latency, cost, utilization
- Sometimes write kernels and, yes, occasional tasteful shitposting
Optimize Your Resume for This Job
Get a match score and see exactly which keywords you're missing
Job Details
- Category
- Software
- Employment Type
- Full Time
- Location
- San Francisco, CA
- Posted
- Last updated
- Aug 16, 2026, 09:40 PM
- Compensation
- $150,000 - $250,000 per year
About Luminal
Luminal provides a machine learning compiler and serverless cloud platform that automates PyTorch model optimization and deployment.
More Roles at Luminal
Similar Software Roles

Haladir
Software
Software Engineer (Full Time)
San Francisco, CA$100K - $170K Full Time
2 hours ago
Software
2 hours ago
Haladir
Software
Software Engineering Intern (Fall 2026)
San Francisco, CA$6K - $8K/month Internship
2 hours ago
Software
2 hours ago
RetailReady
Software
Full Stack Software Engineering Intern
San Francisco, CA$6K - $10K/month Internship
2 hours ago
Software
2 hours agoFound this role interesting?