The 99% Benchmark: What Aside Actually Proved
In June 2026, a five-person Y Combinator company called Aside posted a 99.0 percent success rate on Online-Mind2Web, the Princeton-led benchmark that replaced static snapshots with 300 tasks across 136 live websites where pages change, logins expire, and timing matters, the at-inc/aside-benchmarks repository reported. Two hundred ninety-seven tasks passed; two failed; one was marked impossible by evaluators. Excluding that task, the rate rises to 99.3 percent. Every easy task passed. 142 of 143 medium tasks passed. 75 of 77 hard tasks passed: a 97.4 percent rate on the hardest tier.
The same run produced 75.5 percent perfect-task completion on the Odysseys benchmark, an 88.8 percent rubric-item pass rate, according to the at-inc/aside-benchmarks repository, and a 93 percent pass rate on BU Bench V1, the at-inc/aside-benchmarks repository's figures put. Those secondary scores matter because they test longer horizons and partial-progress scoring, not just binary success.
Until June, the best published scores clustered around 30 percent for open-source frameworks and roughly 60 percent for the leading commercial agents, Claude Computer Use 3.7 and OpenAI Operator. The field had stalled.
What separates Aside from the agents that came before is not the model — it's the scaffold. Aside is a custom Chromium build with the agent baked into the rendering engine, tab manager, and local storage. The agent sees the DOM directly, drives the browser natively, and keeps its memory and credentials on the device. An integrated password manager logs the agent in without ever exposing secrets to the language model. The default behavior is to run until the task finishes, no intermediate confirmations required.
The numbers are self-reported. Aside's team published the results on their own GitHub repository with their own grading pipeline. No independent lab has reproduced the run on the same 300 tasks with the same judge. Steel.dev, which maintains the leaderboard, notes that judge methodology (human, WebJudge, or custom agentic judge) can shift scores materially, and that automated scores should only be compared when traces and task-level results are public. Aside's release also cites Browser Use at 97.7 percent, the aside-benchmarks repo's data shows, GPT-5.4 at 92.8 percent, the aside-benchmarks repository found, and Claude Opus 4.8 at 84.0 percent, the aside-benchmarks repo's numbers put, all derived from Aside's own runs rather than from the vendors themselves.
Independent observers have flagged this. The benchmark's designers built Online-Mind2Web precisely to prevent the overfitting that plagued static suites; tasks are assigned in natural language with no site-specific scaffolding. But a producer choosing the task subset, the timeout, and the success criterion holds a structural advantage. The open question is whether a third party (an academic group or an outfit like Artificial Analysis) will replicate the protocol across every contender. Until then, the 99 percent is a manufacturer's claim: impressive if confirmed, unverified until it is.
Two Weeks, Three Million Views
The viral launch that followed those benchmark claims produced 3 million-plus views and 20,000-plus users within the first two weeks, per the company's Y Combinator job posting. The traffic forced a five-person team in San Francisco to ship a substantial update roughly one week after going live, a pace that signals both the intensity of early adoption and the rawness of the product early users encountered.
Piunikaweb reported on July 1, 2026, "it has only been about a week since Aside launched publicly." That first update addressed signup failures, Intel Mac compatibility problems, and a GPU rendering memory footprint the team cut by two-thirds versus the initial release. Early adopters flagged thermal throttling on Apple Silicon laptops; one user reported their M2 MacBook Air heating up enough that keeping Aside open alongside other apps became impractical. Another user surfaced the same lag in the update thread, confirming it wasn't an isolated environment issue.
What drove the traffic? The product's core pitch — "Give it your passwords, browsing history, and browser context" — lands differently than the integration-dependent agents that dominate the category. Aside operates inside the browser itself, navigating Gmail, Notion, Slack, Figma, and banking sites without API keys or OAuth handshakes. That promise, combined with the benchmark claims the company published (first place on Online-Mind2Web, BU-Bench-V1, and Odysseys), gave the launch a technical credibility that pure hype cycles lack. Y Combinator backing and angel investment from Cloudflare's Dane Knecht added signal for the developer-heavy audience that typically populates early browser betas.
The early user experience was explicitly framed as incomplete. The welcome email from the team read: "We're still early, so you may run into bugs or rough edges. If anything feels confusing or broken, please reply directly to this email. I read every message." That direct line to founders (still feasible at 20,000 users) created a feedback loop that produced the July 1 update's feature list: pinned tabs (the top-requested feature), expanded AI provider support including xAI's SuperGrok, Kimi, Z.ai, and custom local models via Ollama and LM Studio, Firefox import, reliable Chromebook bookmark migration, and a hover-triggered vertical tab sidebar.
YouTube reviewers in late July 2026 described the browser as "bare-bones" but praised the AI flexibility, particularly the ability to bring your own subscriptions or API keys, the built-in password manager with Bitwarden detection, and the task-oriented tab organization that creates a folder per agent run. One reviewer completed an Airbnb search in Dallas, watching the agent open multiple tabs, show its reasoning, and return a curated result with price, reviews, and walkability details. The same reviewer declined to switch daily drivers, citing rough edges and performance, but called the development speed "refreshing" and said they'd keep watching.
The traction also attracted media placement: 9to5Mac listed Aside among "AI browsers people actually use" after ChatGPT Atlas shut down, and TechCrunch included it in a July 22 roundup of the hottest Chrome and Safari alternatives. That coverage cycle — launch, viral metrics, rapid patch, benchmark claims, press picks — compressed what usually takes quarters into weeks. The question the numbers raise isn't whether Aside captured attention; it's whether a five-person team can convert 20,000 early adopters into a sustainable user base before the rough edges become a retention ceiling. The next section examines the hire meant to answer that.
Can a Five-Person Team Turn Hype into Revenue?
A LinkedIn posting on July 4, 2026, made the next move explicit: "Jack & Jill hiring Founding GTM Lead ($120K - $150K + 0.20% - 0.40% Equity) at a confidential company - YC-backed AI Browser Workforce." The description matches Aside (Y Combinator-backed, AI-first browser, viral June launch) and the timing aligns with the post-launch inflection point. The founders, Jun Kim, Chanhee Lee, and Sanghun Lee, had taken the product from a struggling YC sales assistant pivot to 3 million views and 20,000 users in two weeks. That traction forced the transition from founder-led selling to a dedicated go-to-market function.
The compensation range signals where Aside sits in the 2026 market. Refery's 2026 data puts a founding GTM lead at a seed or Series A startup at $180K–$280K base plus 0.5–1.5% equity, with on-target earnings of $250K–$420K. Aside's posted $120K–$150K base with 0.20–0.40% equity sits below that band, suggesting either a pre-seed valuation framework or a deliberate lean structure; the team is five people total, per Y Combinator's directory, and the company lists three open roles across design, marketing, and engineering. Variable compensation for IC-level GTM roles should run 50–70% of base; the posting does not specify an OTE figure, leaving the commission structure an open question for candidates.
| Role | Base Salary | Equity | On-Target Earnings |
|---|---|---|---|
| Market benchmark (seed/Series A) | $180K–$280K | 0.5–1.5% | $250K–$420K |
| Aside's posting | $120K–$150K | 0.20–0.40% | Not specified |
Timing is the critical variable. The research consensus is clear: the right trigger is revenue stage and motion clarity, not headcount. The $500K–$1M ARR window is the sweet spot for most VC-backed startups. Aside's viral launch generated users, not necessarily revenue; the browser is free, and the monetization model (enterprise licenses, usage-based API access, or a freemium tier) has not been publicly detailed. That ambiguity is the hire's first problem to solve. Before hiring, the diagnostic Refery recommends is blunt: can a non-founder close three to five deals in a quarter using a defined playbook? If the answer is no, the founder is hiring too early. Aside's founders have been selling the vision to investors and early users; whether they have codified a repeatable motion for paying customers is the pivot point.
The risk profile is steep. Refery's 2026 data shows roughly 45 percent of first GTM hires at seed-stage startups exit within 12 months. The single biggest differentiator in retention is stage-relevant pattern matching: candidates with prior $0–$10M ARR experience show 78 percent retention at 18 months, versus 41 percent for hires from late-stage public companies with no startup tenure — a 37-point gap Refery calls the largest differentiator in any non-founder hire they track. The "downshift trap" is real: a star from a $200M ARR company fails as the first AE at $500K ARR because the systems and inbound that made them successful do not exist yet. Aside's hire needs to be a builder, not a scaler — someone who can cold-call, design the ICP, write the first playbook, and instrument the CRM before a single campaign launches.
That build-out maps to the 2026 GTM playbook emerging across the data. Signal-based outbound produces better pipeline than volume-based outbound; timeline-based hooks pull a 10 percent reply rate, more than double the 4 percent from problem-based hooks. Multi-channel outreach is not optional: cold email, LinkedIn, and calling as a coordinated sequence. Infrastructure decides whether a campaign scales or collapses: unwarmed domains, unsafe LinkedIn connection volumes, and sequences without deliverability monitoring all produce campaigns that work in week one and quietly stop landing by month two. The first GTM lead at Aside inherits all of this as a blank sheet.
The revenue model decision sits upstream of the hiring plan. Prospeo's 2026 framework argues the hybrid model (product-led growth for acquisition, sales for expansion) is where most mid-market SaaS companies are landing. If Aside's deal size stays below $10K, a traditional sales team is unnecessary; invest in product, growth engineering, and RevOps. If enterprise contracts push above $50K, the motion shifts to account-based, and the first GTM lead needs to operate as a de facto head of sales until two AEs are repeatably at quota. Hiring a sales leader before a repeatable motion exists wastes six to nine months and often ends in a mis-hire. Aside's founding GTM lead will likely wear three to five hats simultaneously: BDR, AE, solutions engineer, RevOps analyst, and product marketer, exactly the profile Refery describes for the seed-to-early-Series-A stage.
The hire also sets the organizational DNA. Chris Hillock, writing on LinkedIn about first GTM leaders, notes they must spend 75 percent of their time listening, using the rest to adapt and tailor the pitch based on real-time customer needs. Rigidity guarantees conflict; the narrative shifts weekly at this stage. Candidates who push a rigid process and dominate the conversation are the wrong fit. The ideal hire matches the founder's grit and operates with a "yes, and" disposition, amplifying founder magic into a repeatable engine rather than replacing it with a playbook from a previous company.
Aside's viral moment bought attention. The GTM hire converts it into a business. The difference between a 2 percent reply rate and a 10 percent reply rate often isn't copy — it's whether the emails land in inboxes and whether the contacts are real. The difference between a campaign that scales and one that collapses is infrastructure built before the first send. The founding GTM lead's first 90 days will determine whether Aside's 20,000 users become a revenue funnel or a vanity metric. The market is watching; the role has grown twentyfold in two years, and the companies that automate their grunt work and level up are the ones that don't get buried.
The Browser That Keeps Secrets on Your Laptop
Aside's architecture rests on three pillars the company has made explicit: autonomous task execution, local-first intelligence, and hardware-secured credential management. The second pillar is the one that changes the threat model for any organization that cannot ship data to an external model provider. "Local-first means the 'Source of Truth' is the browser's IndexedDB or SQLite," the team wrote in a technical post. "The Result: 0ms latency. The AI isn't querying a distant database; it's querying a local one." In practice that means every task plan, every scraped DOM snapshot, every inferred user preference lives in an encrypted store on the device, mathematically disjoint from any external server.
The credential layer is where the design gets concrete. Most agentic browsers solve authentication by asking users to paste API keys or OAuth tokens into a cloud dashboard. Aside instead built a password manager that autofills credentials into the page's own login forms, mediated by the device's Secure Enclave. The LLM never sees the secret; it only sees the resulting authenticated session. Every credential use is logged in an immutable audit trail, and the agent's filesystem and network access are scoped per task, a sandbox with guardrails that can be inspected after the fact. Sensitive actions — payments, posts, message sends — are gated by a human-in-the-loop confirmation step that cannot be bypassed by the model.
Encryption goes beyond the usual TLS-at-rest. Aside encrypts the local datastore with post-quantum algorithms and binds the keys to the Secure Enclave, so a compromised OS cannot trivially exfiltrate session tokens or browsing history. The browser also adopts a "Bring Your Own" model for the reasoning engine: users can plug in a ChatGPT or Claude subscription, or their own API key, meaning the model provider never receives the user's proprietary context unless the user explicitly routes it there.
For defense contractors, satellite operators, and classified-adjacent research labs, this stack answers a specific compliance headache: how to let an agent move through internal tools (Jira, Confluence, proprietary telemetry dashboards) without sending screen scrapes or DOM trees to OpenAI or Anthropic. The local-first memory graph means the agent "remembers" which internal wiki holds the launch-vehicle spec sheet without that fact ever leaving the air-gapped laptop. The audit log gives a security officer a tamper-evident record of every credential invocation, satisfying the "who touched what" requirement of NIST 800-53 and ITAR-adjacent policies.
The trade-off is latency on the reasoning side. Because the heavy model still runs in the cloud (unless the user self-hosts a local LLM), each step incurs a round-trip to the provider, but only the prompt and the model's next action travel, not the raw page content or credentials. That is a deliberate reduction of the attack surface, not an elimination of it. Independent researchers have warned that AI browsers as a class remain in their security infancy; prompt-injection vectors that trick the agent into exfiltrating data via crafted page elements are still being cataloged. Aside's sandbox and HITL gate mitigate the blast radius, but they do not erase the underlying risk of a model hallucinating a destructive action that a tired analyst approves.
The company's own benchmark claims — 99 percent on Online-Mind2Web, BU-Bench-V1, and Odysseys — are self-reported and not yet replicated by a neutral lab. That matters for enterprise buyers who need third-party validation before green-lighting a browser that holds root-level session cookies. Until independent red-team results land, the architecture is a strong design proposal, not a certified control. For now, the organizations most likely to pilot Aside are those that can run their own evaluation in a staging environment, exactly the teams that already treat the browser as a privileged endpoint.
Where Aside Fits in a Crowded Field
Digital Applied's June 2026 taxonomy places six serious contenders in three buckets: distribution plays (ChatGPT Atlas, Perplexity Comet), rethink-the-browser plays (Arc, Dia), and privacy or power-user niches (Brave Leo, Opera Neon). Aside, the YC-backed startup that launched its Chromium-based browser in June 2026, blends the rethink-the-browser ambition with a privacy architecture that runs the model locally, a combination none of the incumbents currently ship at scale.
Arc, released by The Browser Company in 2022, proved there was appetite for a radically different browser UX. By April 2026 the company had officially stopped feature development on Arc, though it continues to ship security patches. Dia, launched in June 2025 as the AI-first successor, is now the team's sole focus. Reviewers describe Dia as the most interesting product in the category from a design perspective and the most risky from a staying-power perspective, a small team competing against OpenAI, Perplexity, and eventually Google. Arc retains an active user base and remains a stable choice for existing users, but it is not the natural recommendation for a fresh deployment. Aside's founders built their own Chromium fork rather than layer on top of Arc's codebase, betting that full control of the rendering engine matters more for agent reliability than inheriting Arc's Spaces, Boosts, or Easels, features Dia itself has skipped.
Microsoft Edge's Copilot integration represents the distribution play with the widest install base. The May 2026 Edge update added Copilot-powered tab summaries, browsing-history recall, AI-generated podcasts, quizzes, writing help, and mobile Vision features across iOS and Android. Microsoft confirmed a redesign with a Copilot-style UI, bringing rounded design and unified UI across Bing and other Microsoft AI surfaces. Edge's advantage is reach; its limitation is that Copilot sits as a sidebar overlay rather than an agent in the navigation layer. The five characteristics that separate AI-native browsers from AI-augmented ones (agent in the navigation layer, persistent cross-session memory, first-party model integration, URL-bar intelligence, task execution over Q&A) are largely absent from Edge's implementation. Aside's agent operates inside authenticated sessions across websites, internal tools, and local files without requiring per-site integrations, a capability Edge does not claim.
SigmaOS, launched in 2021 as a privacy-focused browser, added its AI capabilities (the local AI Eclipse feature and the Airis AI Assistant) only in December 2025. Sigma Agents entered public beta in June 2026 with MCP tool support and warehouse agent integrations, signaling a push toward agentic workflows inside dashboards and apps. The company is clearly investing heavily in AI columns, Sigma Agents, and MCP tools. But Sigma's agent autonomy remains narrower than Aside's benchmark profile: on the Online-Mind2Web leaderboard updated June 29, 2026, Aside scored 99.0 percent across three agentic browsing evaluations (Online-Mind2Web, BU-Bench-V1, Odyssey), ahead of the previously cited scores for Browser Use, GPT-5.4, and Claude Opus 4.8, and ChatGPT Atlas at 70.0 percent. Sigma does not appear on that leaderboard. For security-conscious users in defense, space, and regulated finance, Sigma's local-first posture is credible; but Aside's architecture keeps the source of truth in that local store, delivering instant latency because the AI queries a local database rather than a remote one.
Enterprise adoption is the real battle, and the scoreboard is visible. Atlas Enterprise and Arc for Teams are furthest along on SSO, data-loss prevention, policy controls, and audit logs. Comet and Dia are closing the gap. Brave Leo and Opera Neon trail. Aside has not yet published an enterprise SKU, but its local-first design (no data leaves the device unless the user opts in) addresses the indirect prompt injection and data leakage risks that TestGrid identified as top limitations for AI browsers in March 2026. The company's first go-to-market hire, made after the viral June launch, will define whether Aside pursues a bottom-up developer motion or a top-down enterprise play. The competitive window is narrow: Digital Applied expects feature parity between leaders to shift quarterly, and the deeper pattern is convergence: every AI browser is adding tab chat, agent mode, summary, and writing help. The differentiator in 2026 is which AI ecosystem you already live in. Aside's bet is that enough users want an ecosystem of one: their own browser, their own data, their own model.
The first GTM lead will walk into a San Francisco office where the founders still read every support email. The benchmark claim that started this cycle — 99 percent on a live-site suite no one else has cracked — sits on a GitHub repo waiting for a third party to verify it. If the replication comes back clean, the conversation shifts from "impressive if confirmed" to "how fast can you ship enterprise SSO." If it doesn't, the 20,000 users who showed up for the promise become the hardest crowd to win back. Either way, the next 90 days write the next chapter.
Working in frontier tech? Zero G Talent tracks the openings: see every open ASML role, browse frontier tech jobs, openings at Stripe, and the people building the field.