Skip to main content
frontier

Working at Inferact: Culture, Pace and Who Thrives

By Sarah Mitchell

How Work Gets Done

Fifty-plus engineers push code daily to an inference runtime that runs on half a million GPUs around the clock — a feedback loop measured in hours, not sprints. That volume, tracked through the project's own usage telemetry as of January 2026, means every optimization, every memory fix, every scheduler tweak lands in production environments almost immediately.

Inferact grew out of the vLLM open-source project, which began as a PhD prototype at UC Berkeley in 2022 after Meta released the OPT model. The founding team (original maintainers from Berkeley and early contributors from Meta and Red Hat) formed the company to steward and accelerate that ecosystem. The open-source DNA remains the operating system: quarterly vision documents publish publicly, then the community shapes the roadmap. Pull requests from outside contributors receive the same rigor as internal ones. Code quality is enforced through mandatory reviews and what one lead calls "constant refactoring and iteration."

Decision-making follows the artifact, not the hierarchy. When a new model architecture drops (multi-trillion parameter weights, novel attention variants, custom tool-call formats), the team closest to the runtime evaluates the impact and proposes the adaptation. The inference layer must absorb that diversity without forcing model providers to negotiate with every hardware vendor. That M×N compatibility problem (models times chips) is the organizing constraint. It dictates why the headcount leans heavily toward performance engineering: TPU specialization in Singapore, AMD GPU tuning in both San Francisco and Singapore, cluster administration, cloud orchestration, and site reliability. Sixteen salaried roles span these functions, with a compensation band reflecting the depth required to optimize kernels across Nvidia, AMD, Google, AWS, and Intel silicon simultaneously.

Role Grouping Count Salary Band Median
Salaried roles (TPU specialization, AMD GPU tuning, cluster administration, cloud orchestration, site reliability) 16 $200k–$400k $400k
Member of Technical Staff roles (Cloud Orchestration, Cluster Administration, Site Reliability, TPU Performance, AMD GPU Performance (both hubs)) 6 $200k–$400k $400k

Zero G Talent's figures put the median at $400k.

The rhythm is punctuated by in-person meetups every two months, currently rotating between hubs and expanding globally. Between sessions, coordination happens in the open: GitHub issues, design docs, benchmark regressions posted in real time. There is no separate internal roadmap; the public one is the only one. That transparency forces clarity: if a feature cannot be justified in a design doc that withstands community scrutiny, it doesn't ship.

Remote contributors in the core fifty operate on the same cadence. The San Francisco office anchors cluster and orchestration work; Singapore drives TPU and AMD GPU performance tracks. Both locations feed the same monorepo, the same CI pipeline, the same nightly benchmarks. That shared infrastructure is the culture — not a value statement on a wall, but a daily constraint that aligns incentives without meetings.

The hiring pipeline selects for engineers who have already lived in this loop. Open roles (cloud orchestration, cluster administration, SRE, TPU and AMD performance engineering) all demand production-scale GPU fleet experience. Candidates who have only trained models on static clusters don't clear the bar. The work is the runtime, not the model; the product is the layer that makes any model run on any chip, reliably, at the scale the telemetry already proves.

The Operating Philosophy

Inferact's philosophy reads directly from the vLLM project it stewards. Simon Mo, co‑founder and lead contributor, frames the mission bluntly: "I fundamentally believe that open source, especially how vLLM itself is structured, is critical to the AI infrastructure in the world. What we want to do with Inferact is support, maintain, steward, and push forward the open-source ecosystem." That statement, made in a January 2026 interview, matches the project's measurable trajectory. As of that interview, vLLM had crossed 2,000 contributors on GitHub, ranked among the fastest‑growing open‑source projects by GitHub's own metrics, and sustained 50‑plus regular contributors opening pull requests daily. The CI bill alone exceeds $100,000 per month — roughly $1 million annually, a line item the team treats as a quality gate rather than overhead.

The technical values follow from the problem space. Mo describes the shift from traditional ML serving ("clockwork" deterministic batches) to LLM inference where "each request looks differently… continuously filling and coming in" and "the language model itself will decide when does it stop." That dynamism forced two first‑class design pillars: scheduling and memory management. PagedAttention, the project's signature memory innovation, emerged from treating variable‑length sequences as a scheduling problem rather than a batching afterthought. The principle shows up in release discipline: "We want to make sure every single commit is well tested… people can deploy at not thousands but potentially millions of GPUs across the world in different environments."

Community use is explicit, not aspirational. The team collaborates directly with model vendors ("we often get help from these model vendors") and runs those same meetups, now expanding globally. That cadence mirrors contributor velocity: half a million GPUs running vLLM 24/7 means feedback loops close in hours, not quarters. The M×N abstraction — one inference layer connecting any model to any hardware, turns chip diversity and model‑architecture divergence from fragmentation into a product moat.

Curiosity sparked the project. Mo admits he "didn't really think it's the most important problem in the world back in the day. I just wanted to have a hands‑on experience on how this actually works." That origin story persists in hiring signals. Six Member of Technical Staff roles (Cloud Orchestration, Cluster Administration, Site Reliability, TPU Performance, AMD GPU Performance (both hubs)) sharing a common compensation band. The roles demand deep systems fluency across the stack: kernel‑level GPU optimization, distributed scheduling, and production‑grade reliability at million‑GPU scale. Technical depth is the filter; adaptability is the daily requirement.

Nothing in the public record (founder interviews, GitHub activity, job postings) ties Inferact to defense or space contracts. The company builds a universal inference layer for the open‑source AI ecosystem. If defense or space workloads run on vLLM, they do so as downstream users, not design partners. The culture visible in commit logs and contributor counts is high‑tempo, ownership‑driven, and ruthlessly practical about reliability at scale.

The Hiring Bar and What It Signals

Inferact's public job postings reveal a hiring profile clustered around a single title (Member of Technical Staff) across six specializations: those three tracks, TPU performance engineering, and AMD GPU performance engineering in both locations. That band alone signals the bar: these are not junior positions. The titles imply deep systems fluency (Kubernetes internals, accelerator runtime optimization, bare-metal cluster ops), and the compensation matches what top-tier AI infrastructure teams pay for engineers who can own a substrate layer end to end.

Broader industry trends reinforce what that profile suggests. McKinsey now asks graduate applicants to collaborate with its internal AI tool, Lilli, in final-round interviews, assessing how candidates "think, judge and collaborate with an AI tool rather than their technical AI knowledge." UK recruitment specialists told the Guardian in 2025 that "an affinity and competence with AI was becoming a crucial part of the selection process" for top roles. Indeed's Talent Scout agent, released in preview September 2025, uses job descriptions and hiring-manager notes to automate the pattern-matching that used to consume recruiter hours. The direction is consistent: elite technical hiring increasingly filters for judgment with AI tooling, not just raw coding ability.

For Inferact specifically, the roles map to the company's focus on AI-driven systems. TPU and AMD GPU performance engineering at the MTS level means extracting throughput from heterogeneous accelerator fleets, work demanding fluency in compiler stacks (XLA, ROCm), memory hierarchy tuning, and the failure modes of distributed training at scale. Cluster administration and cloud orchestration at the same tier imply ownership of the control plane: scheduling, quota, multi-tenancy, and the reliability contracts that demanding customers require. Site reliability engineering at this pay band means designing SLOs that survive constrained environments, not just paging rotations.

A CMU study on sponsorship bias in elite hiring (Supreme Court clerkships, 2022) found that candidates sponsored by senior women with tenure were more likely to advance than those sponsored by senior men — a signal that longevity in a discriminatory environment gets read as competence. While the authors cautioned the findings may not generalize beyond elite occupations, the mechanism is relevant: in tight-knit, high-trust engineering cultures, referral weight often correlates with the referrer's proven track record of shipping hard things. Inferact's "Member of Technical Staff" title (flat, non-hierarchical) signals "you own a technical area completely" rather than "you manage people."

Public review sites show almost no footprint for Inferact as of mid-2025. That silence is itself a signal: the company is small, hiring selectively, and its employees either haven't bothered to review or are bound by confidentiality clauses common in sensitive work. The board data shows 16 salaried roles tracked across the current posting cycle. That headcount signal — single-digit dozens, not hundreds, means any cultural assessment based on aggregate review scores would be statistically meaningless even if the reviews existed.

What the postings don't reveal, and no public review fills in, is the friction side of that model. Research on culture formation in technical organizations suggests what that compensation structure likely produces on the ground. The MIT Sloan "building culture from the middle out" study distinguishes "big-C" culture (the values leadership publishes) from "small-c" culture, the unwritten rules teams teach new hires. In a company where every engineer earns principal-level pay, the small-c culture tends to coalesce around autonomy and output velocity rather than presence signaling. SHRM India research notes that by 2027 most organizations there will shift to continuous feedback, AI-assisted evaluation, and two-way goal setting, trends mirroring what high-performing technical shops already practice informally: short loops, project-level retrospectives, and peer-driven calibration rather than annual HR rituals.

The same SHRM research warns that continuous feedback only reduces stress when managers are trained to give it; otherwise it becomes a surveillance loop. The MIT study found midlevel leaders often feel pressured to "endorse cultural norms rather than enrich them," a recipe for unwritten rules that erode fairness and belonging.

The only grounded conclusion is structural: Inferact has built a compensation and titling system that targets engineers who've already operated at staff/principal scope, and it has located those roles in three geographies (San Francisco, Singapore, Remote) that imply distributed ownership of hardware-software stacks. Whether the daily experience matches the "high-tempo environment that rewards technical depth and adaptability" described in the company's own framing remains an open question — one that only a current or former employee can answer, and none has done so on the record.

The commit log keeps scrolling. Half a million GPUs wait for no one.


Working in frontier tech? Zero G Talent tracks the openings: see every open Inferact role, browse frontier tech jobs, the companies hiring, and the people building the field.

Ready to Start Your Space Career?

Browse frontier jobs and find your next opportunity.

View frontier Jobs