Skip to main content
Back to companies
San FranciscoFounded 20252+ employeesPrivate
$500K raisedSeed· last: grant (Apr 2026)· 3 rounds
Backed by(2 investors)
Toloka.vc · LeadY Combinator · Lead

Cumulus Labs is a fast multimodal inference provider, purpose-built for AI teams who want faster performance, lower costs, and zero infrastructure work on fine-tuned & open source models. Most teams today are stuck choosing between bad options. Self-hosting inference means wrestling with configurations and babysitting infrastructure that slows/breaks at scale. Big providers like Fireworks are convenient but extremely expensive and idle GPUs. Cumulus ships Ion, a proprietary inference engine that run LLMs, VLMs, and audio/video gen with high performance and lower cost.

Founders

Skip jobs list

Open Positions at Cumulus Labs (1 Jobs)

1 open

Showing 1 jobs