Six‑Person AI Startup Soff Hires Four, Tests With Two‑Day Work Trial
Inside the Screening Gauntlet
Soff, a six-person Y Combinator startup (S24 batch, San Francisco) operating an industrial supply line under the Westgate Supply brand, is hiring for four roles. Its interview process — a one-minute video pitch, a quick phone screen, then a two-day work trial — tests whether candidates can execute in a high-velocity, in-person sales and operations environment.
The company's public job postings show no requirement for applied AI experience or industrial domain knowledge for these roles. The Business Development Representative posting explicitly states: "No prior industrial or distribution experience required; we'll teach you the industry." The screen evaluates hustle, communication with blue-collar buyers, and the ability to ramp fast on industrial supply fundamentals.
The process mirrors a broader shift in hiring at major AI labs since 2024. The AI Product Manager role barely existed before then; today, companies from OpenAI to Anthropic run dedicated loops that bear little resemblance to traditional PM interviews. Estimation questions and whiteboarding have mostly vanished. In their place: a dedicated AI product sense round, technical drill-downs in every loop, deeper behavioral probing, and — at the sharp end — a live prototype build while an interviewer watches.
The standard loop now runs four to five stages. A recruiter screen confirms baseline fluency. A hiring manager conversation tests whether you understand the company's actual product, not the generic AI narrative. The product sense round has split in two: thirty minutes on a conventional product prompt, then a vibe-coding session where you build a working prototype in a chatbot interface and narrate your trade-offs in real time. Token cost, retrieval strategy, latency budgets, and hallucination mitigation surface as follow-ups even in loops nominally scoped as "product." The technical depth round demands ML systems literacy at PM level: RAG versus fine-tuning, eval metrics that bridge precision and recall to business outcomes, build-versus-buy decisions for model infrastructure. You are not expected to write production code or train models. But you are expected to understand the machinery well enough to make informed product decisions about it.
The behavioral round has become the highest-failure stage at several labs. Anthropic's cultural screen is described by candidates as a therapy session focused on ethics and AI safety; you can pass every other round and still be rejected there. Vanilla STAR narratives read as rehearsed at Meta, Netflix, and OpenAI. Interviewers push for the exact metric, how you knew you were right, what you would do differently now. They are testing for the judgment that manages probabilistic systems, where latency spikes, hallucinations, and accuracy drift are not bugs but inherent properties of the product.
Four Open Roles: What Soff Is Looking For
Soff lists four open positions across its Y Combinator and Standout job pages, all tied to the company's pivot from selling software to distributors to becoming the distributor itself, operating under the Westgate Supply brand. The roles cluster around go-to-market and operations, reflecting a six-person team (S24 batch, San Francisco, in-person only) that needs to build revenue motion. Only the Business Development Representative role carries a full public specification; the other three appear by title alone on Standout.work.
Business Development Representative (Founding GTM)
The Y Combinator posting for this role is the most detailed artifact Soff has published. It is a founding BDR seat for Westgate Supply, the company's industrial supply line targeting fabricators, builders, and maintenance shops that maintain physical infrastructure. The metrics are explicit: 80–100+ outbound calls per day to purchasing managers, shop owners, and plant buyers; 100+ qualified RFQs per week; ownership of the account through first order. CRM hygiene in Close is non-negotiable, as is real-time script iteration: keep what converts, kill what doesn't.
The qualitative filters are equally specific. Soff wants "hard workers and fast learners who want a serious career in sales," comfortable with rejection, repetition, and pressure. Candidates must speak confidently with buyers who have ordered parts for 20+ years, not just executives. That requirement is waived; the company says it will teach the industry. The interview loop reflects this: the same three-step screen before an offer. The company states it moves fast.
Founding Ops
Listed on Standout.work with no public description. In a six-person team running a physical supply line, the title suggests a generalist who can build process from zero. Candidates should expect the same in-person, San Francisco requirement and the same two-day work trial.
Sales Development Representative (Founding SDR)
Also on Standout.work without a published spec. The distinction from the BDR title suggests a top-of-funnel focus: list building, cold outreach, qualification, and handoff. The BDR posting notes "Work closely with sales leadership to iterate on ICP, messaging, and outreach." The same industrial buyer profile applies: blue-collar owners, not C-suite. The same volume discipline applies.
Founding GTM
The fourth Standout listing, "Founding GTM," is ambiguous: it could be a catch-all for a growth lead who owns the full funnel, or a duplicate of the BDR/SDR motion under a broader title. Given the team size (six), it may be a single hire expected to design and run the go-to-market engine. The BDR posting promises "Fast Growth: There's no fixed ladder here. Perform on the phones and you can explore other parts of the business that interest you."
What the Cluster Reveals
Three of four roles are revenue-facing; one is operational. None are engineering or research titles, despite the "AI Agents for Distributors" tagline. The company is hiring for the distribution business it runs today. The same gauntlet tests whether you can do the work, with the buyers, in the cadence, starting now.
Hidden Criteria: Signals That Get You Past the Screen
Research on AI hiring reveals a consistent pattern: candidates who advance past technical screens at frontier companies demonstrate a cluster of less-visible capabilities that determine whether an AI system actually ships and stays useful in production.
Design thinking emerges as the single most distinctive soft skill among roles with the highest AI talent share, according to LinkedIn's Economic Graph analysis. That's a systematic approach to problem-solving that centers the user or customer. Operations management — what LinkedIn calls operational excellence — appears alongside it: the ability to instrument, monitor, and continuously improve a deployed system.
Communication ethics rounds out the top tier. As AI systems make more decisions that affect people directly, the ability to document assumptions, surface failure modes, and explain limits to non-technical stakeholders becomes a core engineering responsibility. Gartner found 61 percent of technology hiring managers rate critical thinking and teamwork as essential in the AI era, and NACE's 2026 survey showed employers rank communication, teamwork, and critical thinking above AI-specific technical skills in importance.
Domain translation is a hidden multiplier. "Domain translation decides whether a technical investment turns into a result the business can measure," notes the Analytics Insight analysis of global capability centers. McKinsey's Generative AI and the Future of Work report identifies "adaptability, coping with uncertainty, and synthesizing information" as the traits most correlated with employment and income gains in AI-exposed roles.
Written, asynchronous communication matters more than most applicants realize. Distributed teams, especially those building agent infrastructure across time zones, rely on short messages that convey intent, trade-offs, and open questions without a meeting. "Reading tone and intent correctly in a short message carries real weight," the GCC analysis observes.
Judgment over automation is a final filter. What matters is judging a model's output, spotting its limits, and flagging when a decision needs a human check. Prompt engineering persists as a proxy for this skill: translating between the machine and the human, knowing where the translation fails.
Ethical reasoning completes the set. With 60 percent of IT instructors now incorporating ethics discussions into AI coursework per a 2024 EDUCAUSE survey, and compliance and responsible AI moving into technical teams' daily work, a candidate who can articulate a concrete bias test they designed or a privacy review they triggered signals readiness for the regulatory and reputational reality of deployed agents.
These signals surface in follow-up questions, take-homes that ask for a post-mortem instead of a benchmark, and reference calls that probe for judgment under pressure.
Candidate Playbook: How to Tailor Your Application
Research on asynchronous video platforms like HireVue (used by over 700 employers) shows that AI scoring gates the human review: candidates below the threshold never reach a recruiter. Unstructured, rambling responses score poorly even when delivered confidently. The fix is not more polish; it is more structure.
Start with a story bank. Standard prep guidance recommends six to eight STAR narratives covering leadership, pressure deadlines, teamwork conflict, failure and learning, influencing others, innovation, commercial awareness, and stakeholder management. Write the four-word prompt for each ("Dissertation / 3 days / schedule + supervisor / 78% first") so you can glance at it during the 30-second think window and hold structure under pressure.
Use the full think time. Most candidates start recording the moment they feel ready. A 10-second structured pause before a perfect answer outscores an immediate, unstructured response every time. During that window, jot S-T-A-R plus your key points. This prevents mid-answer drift, one of the most common failure modes in scored video interviews.
Speak at 140–160 words per minute. Under recording pressure, people push to 180–200 WPM, which sounds rushed and reduces comprehension. Record yourself answering three STAR questions on your phone. Measure the pace. Deliberately slowing down by 15 percent is almost always an improvement. Eliminate filler words (um, uh, like, so, basically, you know), which HireVue's vocal analysis weights heavily negative. Replace each with a deliberate one-second pause. Silence reads as confidence; fillers read as nerves.
End each answer with a clean closing sentence. The final 10 seconds are disproportionately memorable in human review and weighted in AI tone analysis. "That result reinforced my belief that structured planning under pressure separates good outcomes from great ones." Then stop. No trailing off, no repetition.
For the written application, specificity beats volume. Automated scoring flags responses that could apply to any employer versus those containing employer-specific content. Cite a specific company initiative, recent deal, or design choice that matches a problem you have solved. Link it to the exact role and your stated career direction. Generic "passion for AI" language scores the same as silence.
If the process includes a work simulation or take-home, treat it like a production deliverable: readable code, clear assumptions, a one-page trade-off memo. Consistency across formats signals reliability.
Practice the full loop once before the real session. Record, watch, note posture, eye contact, pace, fillers. It is uncomfortable. It is also the single most effective preparation technique documented in the prep literature. Candidates who do it pass at higher rates; candidates who skip it rarely clear the AI threshold.
Why Soff's Bar Matters: Industry Context
The money flowing into AI startups has not slowed. Forty-nine U.S. companies raised $100 million or more in 2024, and the pace held in 2025 — eight companies raised multiple mega-rounds, while Anthropic alone closed two rounds exceeding $1 billion. Capital is abundant; deployed talent is not.
| Company | Role / Context | Metric | Value |
|---|---|---|---|
| Soff (Westgate Supply) | Business Development Representative | Base Salary | $80K–$100K |
| Soff (Westgate Supply) | Business Development Representative | OTE (uncapped) | $140K–$180K |
| Anthropic | Staff Research Engineer | Salary Band | $500K–$850K |
| Databricks | Strategic Sales Leaders | Salary Band | $350K–$600K |
| OpenAI | Funding Round | Raise | $40B |
| OpenAI | Valuation | Valuation | $300B |
| Cerebras Systems | Funding Round | Raise | $1.1B |
| Cerebras Systems | Valuation | Valuation | $8.1B |
| Cursor | Funding Round | Raise | $2.3B |
| Cursor | Valuation | Valuation | $29.3B |
| xAI | Series E | Raise | $20B |
| Merge Labs | Seed Round | Raise | $250M |
That scarcity shows up on hiring boards. Anthropic posted 43 roles in the past seven days, with staff research engineer bands running $500,000 to $850,000. Databricks added 55 roles in the same window, its strategic sales leaders priced at $350,000 to $600,000. These are expansion signals from companies racing to put models into production. The knowledge half-life in AI has shrunk to months. A researcher who published a seminal paper two years ago may already be working on obsolete architecture.
Enterprises feel the same pressure. Worker access to AI rose 50 percent in 2025, and the number of companies with 40 percent or more of their AI projects in production is set to double in six months. Yet only 34 percent of organizations say they are truly reimagining their business. Two-thirds report productivity gains, but just 20 percent have turned AI into revenue growth. The gap between pilot and production is where the hiring bar hardens.
Investors are explicit about the trade-off. Marell Evans of Exceptional Capital said companies increasing AI spend will pull from labor budgets. Rajeev Dham of Sapphire Ventures agreed that 2026 budgets shift resources from people to AI. Jason Mendel of Battery Ventures called 2026 the year agents move from augmenting humans to automating work. An MIT study estimates 11.7 percent of jobs are already automatable with AI. Employers are eliminating entry-level roles and citing AI as the reason for layoffs.
The skills gap is the most cited barrier to integration. Deloitte found education — not role redesign — was the number-one way companies adjusted talent strategies. Only one in five companies has a mature governance model for autonomous agents. Token costs have dropped 280-fold in two years, yet monthly bills run into tens of millions because usage exploded faster than costs declined. Infrastructure strategies built for cloud-first are collapsing under production-scale deployment, forcing a shift to hybrid: cloud for elasticity, on-premises for consistency, edge for immediacy.
That screen is a leading indicator for the kind of execution-focused hiring that scales. As more capital chases fewer deployment-ready operators, the screen will only tighten.
Working in AI? Zero G Talent tracks the openings: see every open Databricks role, browse AI jobs, openings at Anthropic, and the people building the field.