Skip to main content
← artificial intelligence

Evidence AI Agent Drives 33,000 Weekly Installs, Redefining BI

By David Yu•

The Evidence Agent and the Code-Native BI Moment

Evidence shipped an AI agent that reads your dashboards, writes SQL, and checks its work against the definitions your team already agreed on. The agent runs inside Evidence Studio, Claude, and ChatGPT. It is not a wrapper around a legacy BI tool. It is the first analytics interface built on the assumption that the semantic layer — metrics, filters, access rules, analytical playbooks — lives in a repository, not a graphical interface.

The launch arrived August 4, 2026. Weekly active teams on Evidence grew by a third last quarter, most of the gain landing after the April release. Weekly installs top thirty-three thousand. GitHub stars near seven thousand. Community members: twenty-five hundred. But the numbers are not the story. The story is what the agent reveals: a structural gap in every BI platform that preceded it. As analytics becomes conversational, the underlying architecture must shift from GUI dashboards to code-defined metrics, forcing data teams to rebuild their analytics infrastructure.

Evidence's bet is explicit: "Every data team is being asked to deliver a capable analytics agent to its organization. We expect the analytics agent will become the most important data product in every organization. We have a pretty specific view of how the best data products are built and maintained. They are defined in code, managed in a repo, and developed by humans and coding agents together." That view — code-native, repo-first, agent-ready — is now the dividing line in business intelligence.

The problem Evidence's founders describe is familiar to data teams facing a growing wave of vibe-coded content from across their organizations: HTML artifacts and one-off React apps built in Claude Desktop and other tools. Someone in finance or operations has a question, opens Claude Desktop and forty minutes later has a working HTML artifact or React app. They do not need to submit a ticket or wait for another round of changes to a dashboard. The logic is buried in JavaScript. The access controls are missing or recreated from scratch. The analysis knows nothing about the reports that already exist, the definitions the company agreed on, or the work another team completed last month. Data teams call this a nightmare. Users call it self-service that finally works.

Evidence's agent approaches the same user need from the opposite direction. Every dashboard and report in an Evidence project is machine-readable context. The agent translates a natural-language question into the underlying SQL, applies the project's filters, respects row-level security and page-level access control inherited from the project's single version-controlled permissions file, and returns an insight written in Evidence Markdown, which is code a SQL analyst can read, review, and merge into the canonical report without rebuilding it. The insight carries a distilled version of the original prompt so it can be regenerated. Users save insights, share them, and branch off them. The analytical logic never leaves the repo.

Context and skills (the agent's memory and its playbooks) are markdown files in an agent/ directory at the project root. Context files (always loaded) hold company targets, business processes, org charts, product changelogs. Skills (loaded on demand) encode analytical processes: which follow-up questions to ask, which steps to complete, which UI to present. Both are version-controlled. Both can be generated, extended, and maintained by any process that writes markdown to a repo, not just Evidence but any coding agent. The agent runs in the same runtime as the project, so the same SSO, SCIM, and audit-log integrations apply whether the user chats in Studio, Claude Desktop, or ChatGPT via the MCP server's get_context tool. There is no second permissions system to build.

This architecture — semantic layer as code, version-controlled, agent-accessible — is what makes the agent possible. Traditional BI platforms store their semantic layer in a proprietary metadata store behind a GUI. An LLM cannot "read" a Tableau workbook or a Power BI dataset the way it reads a markdown file with embedded SQL. It cannot trace a metric definition through a drag-and-drop lineage graph. It cannot branch a dashboard, modify the metric, open a pull request, and deploy a preview environment. The GUI is a black box to the model. Evidence's repo is a white box.

Why Code-Native Beats GUI for AI Agents

The architectural mismatch is straightforward. Traditional BI tools — Tableau, Power BI, Metabase, Looker — store metric definitions inside their application databases. A "Revenue" metric lives as a saved question or calculated field in the BI platform's internal metadata store. When an AI agent needs to answer a question about revenue, it cannot read that definition. It sees only the rendered dashboard or the SQL the tool generates at query time. The logic is opaque, versioned only by the platform's own audit trail, and bound to a proprietary GUI state that no external system can author.

Code-native BI inverts this. Lightdash reads metric definitions from dbt YAML files checked into Git. Evidence defines pages, queries, and skills as Markdown and SQL files in a repository. dvt represents every dashboard as a JSON document validated against a published schema. In each case the artifact (the metric, the report, the dashboard) is a structured file with a schema. An AI agent reads the same files a human reads. It writes the same files a human writes. There is no "AI generates UI clicks" indirection. The agent is an author of the artifact, not a puppet of the interface.

This distinction determines whether an analytics agent can reason reliably. Evidence's agent treats every page in an Evidence project as context. Definitions and skills are files the team writes. The agent answers by constructing queries against those definitions and returning insights the user can review. dvt's architecture goes further: because the dashboard spec is a structured JSON document, an agent posts a valid spec through the API or over MCP. The platform does not need to simulate a human clicking a filter panel. The agent produces the artifact directly.

The context layer research from Atlan makes the requirement explicit. Agents need enterprise-specific knowledge they cannot learn from training data: which metric definitions are authoritative, how entities map across systems, what data policies govern access, and how to trace answers back to source systems. GUI-first BI cannot provide this layer because its definitions are not addressable, not version-controlled, and not schema-validated outside the platform. Code-native BI provides it by default; the repository is the context layer.

dbt Labs has framed this as a deliberate architectural choice. Building a metric definition framework that supports reliable AI agents requires definitions that live in one place, are enforced at query time, and do not drift. Lightdash's governance is structural for this reason: metric definitions live in the dbt project, and Lightdash surfaces them without allowing a second copy to diverge. Metabase, by contrast, sits on top of dbt mart tables but does not enforce the metric logic the team wrote in dbt. The definition exists in two places. Drift is inevitable.

The GUI-first model has a genuine strength. Metabase's visual query builder lets a marketing manager filter customers by region and group by signup month without writing SQL. That approachability lowers the floor for who can get an answer. But it is also a ceiling. Adoption rates outside the core data team consistently run three in ten to four in ten. The majority of dashboards built are rarely opened. The GUI that enables exploration also traps the logic inside the tool.

As of 2026 every BI vendor is adding AI features. "Has AI" is no longer a distinguishing characteristic. What distinguishes is the architecture underneath: whether the AI writes a structured, reviewable artifact or generates GUI state that only the platform can interpret. That architectural question will matter more as AI authorship becomes more common, not less. The teams that move their definitions into code now will be the ones whose agents can actually reason about them.

Tableau's AI Push and Its Limits

Salesforce bought Tableau in 2019 and the AI integration has accelerated since. The platform now ships three distinct AI layers: Tableau Agent (rebranded from Einstein Copilot in August 2024), Tableau Pulse, and Ask Data. Einstein Discovery adds predictive modeling on top. All of it runs through the Einstein Trust Layer, which masks sensitive fields, logs every AI action, and respects Tableau Cloud permissions before any data leaves the governed environment. Box uses Pulse for security monitoring and says incident response times improved. The features work; analysts report faster insight generation and higher self-service adoption across business units.

But the architecture constrains what these features can become. Tableau Agent writes calculated fields, suggests chart types, and generates entire worksheets from natural language. It does not build dashboards. Dashboard Narratives remains in beta and the company acknowledges the gap. Ask Data answers single-hop questions against a published source; multi-hop reasoning breaks down. Copilot calculations fail on complex nested level-of-detail expressions. Pulse pushes automated digests via email or Slack, but it requires Tableau Cloud Advanced — a higher tier, and scheduled extracts with email subscriptions remain the cheaper workaround. The best experience is English-only; multi-language support for Pulse Q&A has expanded but the core agent still prefers English field names and aliases.

These are not teething problems. They follow from a GUI-first foundation. Tableau stores metric definitions, chart layouts, and dashboard state as serialized workbook objects. An AI agent that reasons about data needs to inspect, version, and recompute those definitions programmatically. A workbook file does not expose a clean semantic layer; it exposes a visual one. When Pulse detects an anomaly, it summarizes the change in plain language. It cannot rewrite the underlying metric because the metric lives inside a workbook, not in a repository where an agent can propose a diff, run tests, and open a pull request. The Einstein Trust Layer governs data egress, but it does not solve the representation problem: the agent sees the output of a calculation, not the calculation itself as code.

Latency numbers confirm the cloud-bound design. Average AI query response sits at two to three seconds for medium datasets because every request routes to Salesforce Einstein servers with schema context and sample values. Caching helps repeated questions, but the round trip is architectural. On-premises Tableau Server does not get feature parity with Tableau Cloud; the two feature sets do not move in lockstep. Organizations that keep data on-prem for compliance lose the AI layer entirely or must duplicate governance across environments.

Power BI faces the same constraint. Its Copilot integration leans on the same cloud inference model, the same semantic model serialized in PBIX files, the same dashboard-as-artifact paradigm. The price advantage — Power BI Pro at fourteen dollars versus Tableau Creator at seventy-five, matters for seat count, but it does not change the substrate. ThoughtSpot offers search-led analytics. Looker offers code-governed metrics through LookML. Qlik offers associative discovery. Each adds an AI veneer. None of them rewrites the storage format.

The pattern is clear: incumbents bolt an LLM interface onto a visualization engine. The interface can author worksheets, explain charts, and push alerts. It cannot refactor a metric definition across fifty dashboards, verify the change against a test suite, and deploy it with a commit hash. That operation requires the metric to live in code, in a repository, with a type system and a build pipeline. The GUI-first vendors have no path to that architecture without abandoning the file format that their entire ecosystem depends on.

The PE Data Imperative: Why This Architecture Matters Now

Private equity's data crisis is not a reporting annoyance; it is a structural drag on returns. Most PE firms struggle with data visibility across portfolio companies. Many still collect portfolio data manually. Flawed due diligence is frequently cited when M&A deals fall short. EY's 2026 Exit Readiness Study identified access to data and KPIs as a major exit challenge from a finance perspective. When a portfolio view is assembled by hand from dozens of sources, it arrives structurally late; by the time numbers are reconciled, the month is over and the decisions that depended on them have already been made on instinct.

The root cause is structural, not technical. A PE portfolio is, by design, a collection of independently built companies acquired at different times, each carrying its own ERP, CRM, billing and HR platforms. Fragmentation is the natural state; visibility is the thing you have to engineer on top of it. Five forces compound the problem: fragmented systems, acquisition activity that injects new reporting cultures faster than integration can absorb them, inconsistent reporting formats and cadences, spreadsheet dependency that bridges the gap between source systems and portfolio views, and manual consolidations stitched together by hand each period, which are slow, error-prone, and impossible to drill into when a number looks wrong.

The most expensive problem in any portfolio is that "revenue," "gross margin," "churn," and "EBITDA" mean different things at different companies. One firm nets out refunds; another does not. One counts bookings; another counts recognized revenue. The moment you try to compare or aggregate, the numbers quietly stop meaning the same thing, and no one notices until a board deck is wrong. There is rarely a single owner of a definition, a documented calculation, or a certified source. So every number is debatable, and meetings dissolve into reconciling figures instead of acting on them. The natural endpoint is several competing answers to the same question, each defensible, none authoritative. When the operating partner, the CFO, and the deal team each bring their own number, confidence in all of them erodes, and the portfolio reverts to gut feel.

This is where the architectural shift becomes urgent. The firms that have solved this did not buy their way out with a single product. They built a thin, deliberate layer of capability above the portfolio: centralized reporting environments that every company feeds into, cloud platforms (commonly Microsoft Fabric or Snowflake) that ingest from each company's systems without requiring rip-and-replace, semantic models where each KPI is defined exactly once so "EBITDA" is the same calculation whether it appears on a company dashboard or a portfolio scorecard, governance with clear ownership of definitions and sources, and executive scorecards that make companies genuinely comparable. Crucially, this does not mean standardizing every company's internal systems; it means integrating their data upward into a governed layer and standardizing the definitions, not the software. The bottleneck is definitions, not technology.

Now AI agents enter the picture. BCG's "Are You Generating Value from AI?" report found that only one in twenty companies creates substantial value from AI at scale; three in five generate zero material value. In the PE context, this is catastrophic: a forty-percent margin improvement that never materializes, a customer acquisition cost reduction that stays theoretical, a margin expansion thesis that dies in pilots. The ninety-five percent fail because they treat AI as a technology problem when it is fundamentally an operating model problem. They skip the foundational data work: mapping data flows, establishing data governance, creating the truth tables that any ML system requires to avoid producing confident hallucinations about the business.

The five percent that succeed do something fundamentally different. They start by obsessing over EBITDA impact, working backward from cash before any line of code is written. They build governance first, establishing who owns the decision to put a model into production, what data quality thresholds trigger a halt, and what production performance thresholds trigger a rollback. They centralize AI playbooks across their portfolio: if one software company successfully uses agentic AI to automate customer onboarding, that playbook gets adapted and deployed across other software companies in the fund. They treat AI as part of their value creation thesis from day one of a new acquisition, modeling how AI improvements flow through the financial model and setting acquisition multiples with AI-driven margin expansion baked in.

This is the connection to code-native BI. An AI agent can only reason about data that is defined in code that is version-controlled, auditable, with explicit semantic definitions that travel with the metric. GUI dashboards cannot provide this. They embed logic in click-paths and proprietary XML that no agent can read, no git history can track, and no portfolio-wide governance layer can enforce. When a PE firm needs "EBITDA" to mean the same thing across twelve portfolio companies so an agent can answer "which companies are trending below plan on EBITDA margin?" the answer must live in a repo, not a workbook. The repo-first architecture — metrics as code, definitions as files, skills as reusable components — is not a developer preference. It is the only substrate that lets an agent operate across a fragmented portfolio with governed, consistent semantics.

Time to value from AI in private equity doubled in a year: two-thirds of firms now see benefits inside twelve months. The gap between firms that can deploy governed, agent-ready analytics across their portfolio and those still reconciling workbooks is widening, and it maps almost exactly to data maturity. The next three to five years will sort PE firms by their competence in operational AI deployment. The ninety-five percent of portfolio companies that fail will drag down fund returns. The five percent that succeed will be the margin-expansion stories that justify the entire fund vintage. The differentiation will not be AI capability; it will be operating model discipline, the boring, difficult work of linking every AI initiative to cash, establishing governance before deployment, building repeatable playbooks, and centralizing portfolio learning. The firms that understand this now are already positioning their Operating Partners and their acquisition theses accordingly. The rest will have a reckoning in their 2026 and 2027 quarterly reviews when they realize that the "AI transformation" narrative in their board decks didn't materialize into fund returns.

The Repo-First Audit: What to Rebuild Before the Next Agent Wave

The Evidence Agent launch made the requirement concrete: an analytics agent reads your project as context, treats your definitions and skills as files you write, and returns answers you can review. That loop only works when metrics, models, and business logic live in version-controlled code, not in a dashboard authoring UI. Teams that want to run an agent in production need a repeatable way to verify their repo is ready. The research points to an eleven-phase audit pattern used for AI-assisted codebases: a discovery pass records what the repo claims to be, then structured tests validate each claim with citations back to the source files.

Start with the semantic layer. The modern data stack has consolidated around the warehouse, dbt, and a unified semantic layer, but consolidation does not equal readiness. Inventory every metric definition. Ask: does each KPI exist as a code object with explicit grain, filters, and ownership metadata? If the answer lives in a LookML view, a Tableau calculated field, or a Power BI DAX measure, the agent cannot see it. The audit should flag every metric that lacks a dbt metric definition, a SQLMesh macro, or an Evidence skill file. Count them. That number is your migration backlog.

Next, test the repository structure. The Evidence Agent expects a project where pages, queries, and components are markdown and SQL files organized in a predictable directory tree. Run a structure check: does the repo enforce a single source of truth for each dataset? Are transformations idempotent and declarative? Does CI run dbt build or sqlmesh plan on every pull request? The AWS guide for evidence-based repository audits recommends staged workflows — lint, type-check, contract-test, then deploy, with CLI automation that turns complex codebases into architectural insights. Adopt that pattern. A repo that cannot produce a clean build in under ten minutes will bottleneck agent iteration cycles.

Validate the contract layer. Agents reason about data through explicit contracts: column types, nullability, freshness SLAs, and semantic tags. The GitHub checklist for repository architecture stresses continuous refinement based on feedback and performance data. Translate that into a contract test suite. Every fact table and dimension should have a corresponding YAML or Python contract file that the CI pipeline enforces. When a schema change breaks a contract, the build fails, and the agent never serves a stale definition.

Check orchestration and observability. The 2026 modern data stack guides show Airflow and Dagster as the dominant orchestrators, but the agent era demands tighter feedback loops. Your orchestrator must expose run metadata (start time, end time, row counts, data quality scores) as queryable artifacts. The agent uses that metadata to explain why a number changed. If your orchestrator only logs to a UI, rebuild the logging layer to write structured JSON to the warehouse.

Assess team workflows. Code-native BI shifts the analyst's daily work from drag-and-drop to pull-request-driven development. The audit must surface skill gaps: who writes SQLMesh macros? Who reviews metric definition PRs? Who maintains the Evidence skill files? The board data shows Databricks hiring senior directors and sales leaders at three-fifty to six hundred thousand, evidence that enterprises are staffing for platform-scale data products, not dashboard maintenance. If your team lacks a dedicated analytics engineer, the agent becomes a black box nobody can debug.

Finally, run an agent smoke test. Point the Evidence Agent (available in Evidence Studio, Claude, and ChatGPT) at a staging deploy of your repo. Ask it five questions your stakeholders ask weekly. Score each answer: correct, hallucinated, or "I don't know." The failure modes map directly to audit findings: missing metric definitions, broken contracts, stale orchestration metadata. Fix the root causes, then re-run. That cycle — audit, fix, test — is the new release process for analytics.

The repo-first audit is not a one-time project. It is a recurring gate. Every quarter, re-run the discovery pass. The architecture diagram generated from your repository (tools like datadef.io now sync GitHub, GitLab, or Azure DevOps repos into living architecture.md files) should reflect the current state without manual updates. When the diagram matches the code, the agent can reason. When it drifts, the agent hallucinates. The choice is that simple.


Working in AI? Zero G Talent tracks the openings: see every open Databricks role, browse AI jobs, openings at Anthropic, and the people building the field.

Ready to Start Your Space Career?

Browse artificial intelligence jobs and find your next opportunity.

View artificial intelligence Jobs