Skip to main content

Staff Infrastructure Engineer — Device Cloud

Revyl
San Francisco, CA
Full Time
Compensation
$150,000–$250,000/year

Job Description

As a Staff Infrastructure Engineer at Revyl, you'll own the technical strategy and systems required to scale the infrastructure that powers our mobile development platform.

Revyl gives developers and AI agents access to cloud iOS and Android environments where they can build, run, inspect, and verify mobile applications. Behind that experience is a distributed compute platform spanning GCP, physical Mac infrastructure, simulators, emulators, build runners, streaming, and test execution.

We're entering a period where our infrastructure needs to support significantly more customers, workloads, and compute. You'll be responsible not only for operating that infrastructure, but for figuring out how it needs to evolve.

You'll dig into how our platform operates today, identify bottlenecks and scaling limits, understand where operational toil and reliability issues are coming from, model future capacity requirements, and develop the technical roadmap that gets us from where we are today to where we need to be.

This isn't a traditional DevOps or SRE role. We're looking for someone who can move fluidly between understanding the current system, designing the next version of it, and getting hands-on to build it.

What you'll work on

  • Define our infrastructure scaling strategy. Understand the architecture and operational characteristics of Revyl's platform today, identify the systems that will break as load increases, and develop a roadmap for scaling them ahead of demand.
  • Capacity planning and modeling. Translate expected customer and partnership growth into compute, storage, network, device, and hardware requirements. Build the models and instrumentation that let us understand when and where we need to add capacity.
  • Scaling our device cloud. Design and build the systems that provision, schedule, manage, and recycle large fleets of iOS simulators, Android emulators, and physical compute.
  • Distributed workload orchestration. Evolve how builds, tests, agent sessions, and interactive workloads are scheduled across heterogeneous pools of compute as concurrency grows.
  • Cloud infrastructure. Scale and evolve the GCP infrastructure behind Revyl's APIs, control plane, workflow execution, storage, networking, and supporting services.
  • Identify and eliminate bottlenecks. Use production data to understand where we're constrained by CPU, memory, disk, network, scheduler throughput, database performance, host capacity, or architecture—and determine the right solution rather than simply adding more machines.
  • Build for step-function growth. Prepare Revyl for customers and partnerships that can introduce significantly more traffic than the platform handles today. Design load tests, failure tests, and capacity plans that give us confidence before traffic arrives.

What we're looking for

We are looking for someone who has taken an infrastructure platform from one stage of scale to the next.

You likely have:

  • Has 5+ years of software engineering, infrastructure, platform engineering, or SRE experience, with significant technical ownership.
  • Has previously helped scale a production infrastructure or compute platform through a meaningful increase in load.
  • Can analyze an existing architecture, identify scaling constraints, and turn those findings into a prioritized technical roadmap.
  • Understands capacity planning and can translate business forecasts and expected workload growth into infrastructure requirements.
  • Is capable of making architectural decisions under uncertainty and knows how to validate assumptions through instrumentation, benchmarking, load testing, and production data.
  • Has built or operated distributed systems at meaningful scale.
  • Is a strong software engineer who can build infrastructure services and control-plane systems, not just configure infrastructure tooling.
  • Understands scheduling, queues, worker pools, concurrency, backpressure, resource allocation, failure recovery, and distributed systems fundamentals.
  • Has deep experience with cloud infrastructure such as GCP, AWS, or Azure.

What success looks like

Within your first year, we'd expect you to have materially changed both Revyl's infrastructure and our understanding of how it scales.

That means:

  • We have a clear model of the capacity and architecture required to support our next stages of growth.
  • We know where our major scaling limits are before customers encounter them.
  • Large new customers and partnerships have concrete capacity and load plans before launch.
  • Revyl can handle an order of magnitude more concurrent compute without requiring an order of magnitude more operational work.
  • Common infrastructure failures are automatically detected and remediated.
  • Device scheduling and capacity management are predictable rather than reactive.
  • We can confidently answer how much additional traffic the platform can support.
  • Infrastructure projects are prioritized based on real scaling constraints rather than whichever fire happened most recently.
  • Our existing SRE and engineering team spend substantially less time firefighting.
  • The Device Cloud becomes a platform we can scale repeatedly rather than an infrastructure stack that needs to be reinvented every time Revyl grows.

Optimize Your Resume for This Job

Get a match score and see exactly which keywords you're missing

Optimize Resume

Job Details

Category
Software
Employment Type
Full Time
Location
San Francisco, CA (Remote)
Posted
Compensation
$150,000 - $250,000 per year

About Revyl

Revyl is an observability platform that helps teams identify and resolve bugs before they reach production. It offers real-time monitoring and insights, allowing developers to catch issues early in the development process. With features such as error detection and performance analytics, Revyl enhances software reliability and improves the overall user experience.

Found this role interesting?

Staff Infrastructure Engineer — Device Cloud
Revyl
Apply