Skip to main content
← artificial intelligence

Working at Speak: Culture, Pace and Who Thrives

By Daniel Reyes•

How Work Gets Done

Speak employs roughly 40 people building AI language-learning products from two offices, San Francisco and Tokyo, separated by 16 time zones. The company keeps a low profile: no engineering blog, no culture deck, no conference talks breaking down sprint cadence or review process. What exists in public view is a 2026 Built In snapshot, one Glassdoor review, and the roles posted on its job board. Together they sketch an outline; the interior stays private.

Built In describes "flexibility, structured hours, and a supportive culture" alongside "startup-style breadth of responsibilities that can heighten workload and time pressure." The summary admits responsibilities "vary by team and peak periods due to resourcing and scheduling complexities," the closest thing to a public acknowledgment that a small team across the Pacific creates coordination friction. The environment is "generally flexible," balance "achievable," but the caveats do the real work: resourcing is thin, scheduling complex, every role wide by design.

The job board reinforces that breadth. Current San Francisco postings, Zero G Talent's data shows:

Role Salary Band
Senior Backend Engineer $170k–$280k
Staff Product Designer $140k–$270k
Senior Android Engineer $140k–$260k
AI Product Engineer $140k–$250k
Applied ML Engineer, Speech $140k–$240k
Summer 2026 Full-stack Intern $6k–$10k/month

Board-wide median: $245k across eight salaried roles. Zero Tokyo-based roles appear live; either that office hires through other channels, its headcount is smaller, or roles aren't advertised here. The absence is a signal: candidates should ask how product decisions flow between sites before assuming symmetry.

Glassdoor shows one review for "SPEAK"; Indeed surfaces reviews for "Speak Speech Therapy," a different entity entirely. The scarcity means usual pattern-matching (promotion velocity, on-call load, design-vs-engineering tension) isn't possible from open sources. The postings reveal a stack heavy on mobile (Android), speech-specific ML, and product-engineering hybrids: the "AI Product Engineer" title suggests engineers who own user-facing features end-to-end rather than tossing models over a wall.

Decision-making structure, meeting cadence, and whether Tokyo operates as peer or satellite are not documented. The Built In line about "scheduling complexities" hints at the cost of the two-office model without quantifying it. For a candidate, the interview process is the primary data-gathering tool: ask how a feature ships when the designer is in SF and the speech engineer in Tokyo, ask what "structured hours" means for a team spanning the Pacific, ask where the last three product debates were settled. The answers will tell you more than any culture page.

Principles in Practice

Speak sits at roughly 40 people, the threshold where Bain Capital Ventures research finds founders typically codify operating principles. "The turning point usually occurs when the company reaches a headcount of approximately 50 people," the firm says, because "teams cannot learn what's important from fleeting one-on-one interactions or occasional meetings" once a company spans multiple offices. Speak's two-office structure puts it squarely in this transition.

Bain draws a sharp distinction between values and operating principles. "The word 'operating' implies (and inspires) action," the piece argues. "'Values' calls to mind a cliche poster you might hang on a wall... That's precisely what operating principles are not." Effective principles "show employees how to prioritize their work, and can help guide difficult decisions"; they become the language people use to settle disagreements and the criteria against which candidates are assessed. Bain's own five principles have persisted for over 30 years because they're "the firm's day-to-day expression of its values," not aspirational slogans.

At Speak's scale, the principles governing behavior are the ones founders repeat in product reviews, the trade-offs they visibly make, the hires they approve or reject. Brian Chesky, describing Airbnb's shift to a functional organization, frames it as "give ground grudgingly," resisting divisional structures that let teams "row in different directions." He argues the CEO should be chief product officer because "the most important thing a company does is make a product. If the CEO is not the expert in the product, then why are they the CEO?" Leaders "shouldn't just be 'managers'... they should also be in the details. If we were a military, like a battalion, the cavalry general should know how to ride a horse."

Chesky's "founder mode" — presence in details, setting pace and standards rather than delegating vision — maps to how small, product-obsessed teams operate before bureaucracy sets in. Airbnb's quality control illustrates the principle: "We're very hands-on with quality control. Most new services we actually vet and certify ourselves. We think reviews are important, but we don't want to put the entire burden on the user base." The economic logic is explicit: "You can start to see a bad Airbnb, and people then don't rebook on Airbnb, they don't come back."

Speak's job board — Senior Backend Engineer, Staff Product Designer, AI Product Engineer, Applied ML Engineer Speech — signals a product-first, technical-bar-first orientation. The salary bands and summer internship suggest a team investing in long-term craft, not rapid headcount growth. The two-office structure demands principles that work across time zones and cultures without becoming bureaucratic overhead.

No manifesto appears in available sources; the careers page doesn't enumerate a values list. That absence is itself a signal: at 40 people, operating principles still live in founders' daily decisions, determining what ships, what gets cut, who gets hired, and how the two teams stay aligned without management layers. Candidates should listen for those patterns in interviews, not expect a documented framework. The principles are the product decisions the team makes when no one is writing them down.

What the Bar Selects For

Speak's six open engineering and design roles tell a clear story before a candidate speaks to a recruiter. Every role demands a defined specialty: speech, Android, backend systems, or applied ML. The product design role sits at staff level. No junior generalist listings. This is a team hiring for immediate contribution, not potential.

That contribution matters because the denominator is small. With roughly 40 people split between two offices, each hire represents 2.5% of the company. A mis-hire isn't a rounding error; it's a visible drag on the roadmap. The roles reveal the product's technical center: speech recognition and synthesis, mobile delivery, and the connective tissue between model and user experience. Candidates who cannot demonstrate production-grade work in one of these domains will not clear the first screen.

The two-office structure adds a collaboration filter absent from job descriptions. Tokyo is not a satellite; it's the other half of the product, close to the Japanese and Korean learner bases driving Speak's core markets. An SF engineer who cannot design for async handoff, write documentation surviving a 16-hour time difference, or accept product direction from a Tokyo designer creates friction the team cannot absorb. The bar selects for people who have already worked this way: distributed teams, cross-cultural product decisions, the discipline to ship without a stand-up to unblock them.

Product judgment is the third leg. Roles are titled "AI Product Engineer" and "Staff Product Designer," not "ML Researcher" or "Visual Designer." The distinction is deliberate. Speak builds a consumer language app where model latency, interruption handling, and conversation flow are product decisions, not just engineering metrics. A candidate who optimizes BLEU score at the expense of perceived responsiveness has missed the job. The interview process must surface whether a candidate treats the model as a component they own end-to-end or a black box they hand off. The board shows no pure research roles. That is a signal.

Public record doesn't show Speak's exact interview loop: stages, take-homes vs. live coding, whether a Tokyo engineer joins SF panels. The broader market has moved toward AI-mediated first-round screens: 63% of job seekers surveyed by Greenhouse in April 2026 reported an AI interview, and two-thirds of recruiters planned to increase AI pre-screening, but nothing in first-party data indicates Speak follows that trend. A 40-person company with this salary band and role specificity typically runs a human-led process: hiring-manager screen, technical deep-dive in the candidate's specialty, product collaboration exercise, cross-office conversation. Candidates should prepare for that shape, not for a bot.

The through-line is autonomy with alignment. The team is too small for managers to decompose work into tickets. The bar selects for engineers and designers who can be given a problem space — "make the conversation feel natural when the user interrupts" — and return a shipped feature working in English and Japanese, on Android and iOS, with a model fitting the latency budget. Signals that carry: shipped consumer features, depth in a listed specialty, evidence of async collaboration across time zones, product instinct treating the model as design material. The rest is noise.

The Review Vacuum

Public review sites list multiple entries for "Speak," nearly all belonging to a speech-therapy provider (Speak Speech Therapy), not the AI language-learning company building from San Francisco and Tokyo. Indeed reviews for Speak Speech Therapy span 2022–2026 and describe an owner-led clinical operation: complaints about unpredictable leadership, missing benefits, high turnover, and disorganization sit alongside praise for a warm mission serving neurodivergent children and supportive internships. None reference a 40-person AI team, a Tokyo office, or an LLM-centered roadmap. Treating them as signal for the AI company would be a category error.

MIT Sloan's CultureX platform analyzed 1.4 million reviews across large organizations and found respect — measured by how employees describe being treated — is 18 times more predictive of overall culture scores than the average topic. Integrity and benefits weigh heavily; compensation matters less than benefits. But those models assume reviews describe the same entity. When a name collision mixes two workplaces, the signal collapses.

For the AI company Speak, verified public accounts are sparse. Glassdoor and Blind show minimal coverage as of late 2024. The careers page highlights "autonomy," "product-minded builders," and a two-office structure, but without attributed employee narratives, candidates have no independent check on whether claims hold. The board data shows active hiring: eight salaried roles posted recently with bands from $72k to $273k, Zero G Talent's figures put, (median $245k), which confirms growth but not culture.

Candidates should treat the review vacuum as information itself. A 40-person team across two time zones generates fewer public reviews than a 500-person clinic chain. Absence of negative Glassdoor threads doesn't prove health; it may just reflect scale. The most reliable next step is direct conversation: ask to speak with a current engineer in San Francisco and one in Tokyo about a recent product decision requiring cross-office alignment. Their answers, whether they describe a clear process or ad-hoc Slack threads, will reveal more than any aggregated rating.

Who Stays, Who Burns

Speak's scale across two offices creates structural demands filtering for a narrow profile. The salary bands (senior backend at $170k–280k, Zero G Talent reported, staff product design at $140k–270k, Zero G Talent found, applied ML speech at $140k–240k, median $245k) signal an expectation of senior, self-directed contribution. No junior ladder exists; the internship at $6k–10k/month is the only sub-senior entry point.

People who thrive share three traits. First, they operate without a manager's daily vector. At 40 people, no one assigns tickets; engineers and designers pick up the next highest-impact problem, ship it, and measure whether it moved the product. Second, they treat the Tokyo–San Francisco split as a design constraint, not a tax. The 16-hour offset means decisions made in one office must be legible to the other without synchronous debate; writers who document context, engineers who leave clear migration paths, designers who spec interactions exhaustively survive the handoff. Third, they care about language learning as a product, not an ML benchmark. The applied ML speech role at $140k–240k sits beside product engineers at $140k–250k; the company pays for model quality that ships, not papers that don't.

Profiles that burn out follow a predictable pattern. Candidates who need a defined career ladder, regular 1:1s with a people manager, or a promotion committee to validate impact will stall, because those structures don't exist. People who optimize for local maxima (polishing unused components, refactoring without a product hypothesis) consume disproportionate oxygen on a team this small. Anyone who treats Tokyo as "remote" rather than "the other half of the product team" creates friction the organization cannot route around; MIT research shows respect is 18 times more predictive of culture scores than the average factor, and cross-office disrespect shows up fast in a 40-person graph.

The two-office structure amplifies a known failure mode: reorganization anxiety. MIT data finds employees discuss reorgs negatively 97% of the time, and the fewer people who mention them, the higher the culture score. At Speak, a single hiring or departure shifts the org chart visibly. Candidates who interpret that fluidity as instability — rather than the natural texture of a small team shipping fast, will churn.

Job security never appears on corporate value statements (MIT found zero companies listing it), yet layoff mentions correlate inversely with culture scores. Speak's size means no redundancy buffer; a funding pause or product pivot hits everyone. The people who stay are the ones who would rather own that risk than trade it for an org chart with 500 engineers and a clear leveling guide.

The 16-hour handoff is the daily test: an SF engineer's evening push becomes the Tokyo designer's morning context, with no synchronous stand-up to bridge the gap. The median $245k band is the floor, not the ceiling, for people who make that handoff feel like craft instead of chaos.


Working in AI? Zero G Talent tracks the openings: see every open Speak role, browse AI jobs, the companies hiring, and the people building the field.

Ready to Start Your Space Career?

Browse artificial intelligence jobs and find your next opportunity.

View artificial intelligence Jobs