Skip to main content
artificial intelligence

Greenboard Pays $200,000 to Define Correct for AI

By John Hugo

The $20 Million Bet

Greenboard raised a $15.5 million Series A led by Base10 Partners in May 2026, bringing total funding to $20 million, BusinessWire reported. Y Combinator, General Catalyst, Wayfinder Ventures, Commerce Ventures, Transpose Platform, Liquid2 Ventures, Kulveer Taggar, and a handful of strategic industry investors also participated — a signal that the backers who seeded the company in 2024 are doubling down before the broader market catches up.

Round Amount Date
Seed $4.5M May 2024
Series A $15.5M May 2026
Total $20M

Dave Feldman and Ed Schembor founded Greenboard in June 2023 after meeting as freshmen at Johns Hopkins and later leading engineering teams at Amazon, Hive AI, and Guideline. They went through Y Combinator's Winter 2024 batch. The seed pitch was straightforward: financial compliance runs on fragmented tools that store records after the fact but don't help employees make decisions in the moment. The Series A pitch is that the product now does both — and the traction proves it.

As of May 2026, Greenboard serves more than 500 financial institutions (RIAs, broker-dealers, private funds, fintechs, and wealth managers) with a customer retention rate above 99 percent. The 2026 T3/Inside Information Software Survey shows the platform earned the highest compliance-category score (8.43) with reported metrics including a 7-second average marketing review turnaround, 99.14% reduction in eCommunications false positives, 88% of customers consolidating tools, and 99%+ retention across more than 500 financial institutions. That growth came from consolidating what used to require three or more legacy systems: SEC 17a-4-compliant communications archiving, AI-driven supervision, marketing review automation, employee attestations, vendor due diligence, and code-of-ethics monitoring all live in one audit trail.

The funding will accelerate product development across supervision and workflow automation. But the raise also forces a harder engineering problem into the open: how do you test an AI system whose outputs aren't deterministic? The platform's core value depends on AI that analyzes communications, detects risk, and routes decisions, all in regulated environments where a false negative isn't a bug, it's a compliance failure. That testing challenge is why Greenboard is hiring its first dedicated AI automation engineer now.

Why the Legacy Stack Crumbled

The regulatory clock started ticking in earnest in 2024. The EU AI Act entered force on 1 August 2024 and became applicable on 2 August 2026, with prohibited practices and AI literacy obligations already live since February 2025. DORA applied from January 2025. NIS2 required national transposition by October 2024. CRA vulnerability reporting obligations begin September 2026, with full compliance due December 2027. These four frameworks converge on a two-year window where financial firms must prove they can govern, secure, and audit every AI system they deploy. Deloitte's 2025 Global ITAM Survey found 81 percent of organisations view compliance with these digital regulations as an opportunity to strengthen IT asset management practices, yet only 29 percent have formally integrated ITAM into cybersecurity planning.

Financial institutions regulated by the SEC and FINRA have long relied on fragmented tools: separate vendors for communications archiving, marketing review, and those functions. Each tool generates its own data silo, its own audit trail, its own review workflow. Greenboard's founders started the company after observing that firms were "keeping up with the growing complexity of securities compliance" by adding yet another point solution. The result: compliance teams drowning in manual review, unable to surface risk across channels, and facing examiners who expect unified evidence.

Regulatory enforcement has sharpened. GDPR fines reached €1.2 billion in 2024 alone, with insufficient technical and organisational measures cited as the primary breach cause. One in four organisations reported increased AI-related incidents (data breaches, model failures, governance lapses) in the past financial year. Over 90 percent lack robust AI governance frameworks, Deloitte found. The AI Act classifies AI systems into four risk tiers: prohibited, high-risk, limited-risk, and minimal-risk. Financial services use cases (credit scoring, algorithmic trading, fraud detection, customer onboarding) fall squarely into high-risk, triggering mandatory risk management, transparency, cybersecurity measures, and human oversight.

Legacy platforms cannot meet these requirements. They were designed for static rules: keyword filters, lexicon-based alerts, predefined workflows. They cannot explain why an AI system flagged a communication, cannot demonstrate the model's confidence level, cannot produce the audit trail the AI Act demands for high-risk systems. The regulation relies on harmonised standards to provide a "presumption of conformity," standards that are still being written. Firms need a compliance operating system that can adapt as standards finalise, not a rigid toolset that requires rip-and-replace every time guidance shifts.

Greenboard's architecture reflects this reality. The platform unifies archiving, firm compliance, employee compliance, marketing compliance, and vendor due diligence into a single system where tasks and records live together. AI analyzes communications, detects risk, and maps compliance in real time. The AI Act's transparency obligations (Article 50) apply from 2 August 2026. Provider marking of AI-generated content lands 2 December 2026. Standalone high-risk systems face full obligations 2 December 2027. Embedded high-risk AI in regulated products gets until 2 August 2028. The Digital Omnibus on AI, published as Regulation (EU) 2026/1744 and in force since 27 July 2026, moved four application dates but changed no substantive requirements. The deferral tracks harmonised standards and tooling; the obligations themselves remain intact.

This timeline creates a forcing function. Firms cannot wait for perfect standards. They need governed execution in real workflows now, including inventory and classification of every AI system, risk assessments for high-risk use cases, technical documentation, human oversight procedures, and runtime evidence that oversight actually happened. The legacy approach of annual audits and sample testing fails when regulators can demand proof for any decision, any model version, any day. Compliance software had to become AI-native not because AI is trendy, but because the regulation now treats AI as a distinct risk category that demands continuous, auditable control. Greenboard's bet is that the only way to meet that demand is a platform built from the ground up to produce the evidence regulators will ask for.

Inside the Engine: How Greenboard Detects Risk

Greenboard's platform ingests e-communications from email, Slack, Microsoft Teams, SMS, iMessage, marketing platforms, and Instant Bloomberg, covering every channel where regulated employees talk to clients or each other. The system archives these streams in a unified repository, then applies AI-powered supervision to replace the industry's default: manual review and random sampling. Legacy tools flag keywords and hope compliance officers catch the rest. Greenboard's models score every message for risk, surface the ones that matter, and suppress the noise.

The numbers are specific. The company reports a 99.14 percent reduction in eCommunications false positives versus traditional lexicon-based filters. Marketing reviews that once took hours now average seven seconds. Eighty-eight percent of customers eliminated at least one legacy tool after consolidating onto Greenboard. These figures come from that survey, where it earned that score (8.43) and a 99 percent-plus retention rate across those institutions.

That architecture shows up in GreenboardGo, the conversational layer released alongside the Series A. An employee asks a compliance question, "Can I share this performance chart with a prospect?", and the system answers from the firm's own policies, not generic regulatory guidance. If the question requires judgment, GreenboardGo routes it to the designated compliance owner, captures the decision, and logs the approval chain. The compliance team reviews, edits, and finalizes. No output ships without a human sign-off.

Expert-in-the-loop isn't a feature. It's the architecture, pairing human precision with AI recall. The AI prepares reports, answers questions, and routes decisions, but nothing is finalized without a compliance professional reviewing and approving it.

Early adopters report measurable efficiency gains. Root Financial estimates it reclaimed roughly 24 hours per week previously spent on marketing reviews and manual tasks. JMG Financial Group cut compliance onboarding time by 60 percent, saved more than ten hours weekly on communications surveillance, and replaced three separate legacy systems. The platform's marketing compliance automation, configurable employee compliance workflows, and firm-level collaboration tools are built on a single data layer, which means every action is documented and searchable without stitching together exports from disconnected tools.

The platform also automates personal trading monitoring and attestations, screens trades for conflicts of interest, tracks marketing content through approval workflows, and monitors regulatory changes across SEC, FINRA, and global regimes. Continuous risk assessment replaces the annual checkbox exercise. Firms shift from periodic sampling to ongoing surveillance, a structural change that regulators increasingly expect.

Greenboard does not build autonomous agents that replace compliance officers. The company explicitly rejects that model. Its thesis: GPU-powered computational statistics will prove better than humans at examining records for potential issues, freeing compliance people to make determinations on risk and mitigation. The AI expands the efficiency and agency of personnel in nuanced regulatory environments where judgment remains essential. That distinction — augmentation over automation — shapes every layer of the testing problem the engineering team now faces.

The Non-Deterministic Nightmare

Standard software testing assumes a contract: same input, same output. A unit test asserts expect(add(2,2)).toBe(4). The assertion passes or it fails. The oracle — the mechanism that decides correctness — is trivial. Generative AI breaks that contract at the root. Ask the model the same question twice and you may get two differently-worded correct answers. There is no single string to assert against.

The evidence is measurable. Thinking Machines Lab sampled 1,000 completions from Qwen3-235B at temperature 0 and still produced 80 unique answers. Researchers at Penn State University and Comcast AI Technologies ran five models across eight benchmark tasks at temperature 0 with fixed seeds and recorded accuracy variations up to 15 percent across runs, with a gap between best and worst possible performance reaching 70 percent. As one practitioner noted: "I was testing an AI feature that turned a conversation transcript into a short summary, and the model never wrote the summary the same way twice. There was no exact string I could assert against."

Temperature 0 does not fix it. OpenAI states plainly in its API documentation that Chat Completions are "non-deterministic by default" and that a seeded request returns output that is only "(mostly) deterministic." Temperature 0 forces greedy decoding so the model always takes the highest-probability token, but it leaves the batching and kernel behavior underneath untouched. Inference servers group your request with other traffic, and that changing batch size alters floating-point reduction order, so identical prompts diverge even at temperature 0. Seed is best-effort; a system fingerprint change can shift output. Four things cause the variance, and only one lives in a prompt file you control: sampling temperature and top-p, context and retrieval drift, prompt phrasing changes, provider model updates. Even at temperature=0, batching, floating-point non-associativity, GPU kernel behavior, and mixture-of-experts routing all still move the output.

Exact-match assertions fail because one question has many correct answers. This is the test oracle problem. A conventional assertion needs an oracle, a reliable way to decide whether a specific output is correct. Generative systems break that assumption because the correct output cannot be enumerated in advance. A plain exact-match assertion is the wrong tool from the very first run.

The failure modes cascade. False positives at scale: tests that assert exact strings or exact JSON shapes fail constantly as the model varies its output. Engineering ignores them; the test suite loses signal. Brittleness to model updates: a test that passes against model version A may fail against model version B even though both produce correct results. Every model update becomes a test-rewrite event. Audit-incompatible documentation: when auditors ask "how do you know your AI system works correctly?", the answer "we have 10,000 exact-match tests" is unconvincing and often false, since the tests are flapping anyway.

Traditional flaky-test advice (retry it, quarantine it, mark it skip) actively misleads here. That advice assumes the variance is infrastructure noise sitting on top of a stable, correct system. With non-deterministic AI, the variance is the system. Retrying a non-deterministic AI test hides your only signal, because the variance you just retried away might be the model behaving differently on a prompt that's still ambiguous or an assertion that's still too strict.

The practitioner rule: a flaky AI test almost always means one of two things: the assertion is too strict, or the prompt is too ambiguous. It is almost never a third thing. If the outputs agree in meaning but differ in wording, your assertion is too strict. If the outputs disagree in meaning or structure, your prompt is too ambiguous.

Regulatory pressure sharpens the stakes. The Act entered force in August 2024 and applies in stages through 2027. U.S. state AI laws are multiplying. With AI system regulation increasing, defensible evaluation is moving from a quality concern to a compliance one. Organizations testing AI systems with property-based assertions and statistical thresholds detect regressions five to seven times faster than organizations relying on exact-match tests, OpenAI's 2024 Evals framework data and Anthropic's Constitutional AI evaluation papers show.

Greenboard's platform does so for 500-plus financial institutions. Its AI outputs (risk flags, review priorities, exception reports) are non-deterministic by nature. That job description states the problem directly: "You'll also help us solve something most automation engineers have never had to solve, which is how to test AI systems whose output is not deterministic."

Building the Test Framework from Zero

Greenboard is not inheriting a test framework; it is writing one from zero. That job description makes this explicit: "You'll own how we test: the framework, the automated suites, and the release process, built from scratch and then embedded with the engineers who ship." That hire will not slot into an existing QA organization. They will define it, choose the tooling, and then sit alongside product engineers to make the framework part of daily shipping rather than a separate gate.

The scope is defined by what cannot break. The job description lists four critical paths: message archiving and supervision, marketing review routing, employee attestations, and audit trails. These are the regulatory load-bearing walls. A missed bug in any of them does not produce a bad user experience; it produces an incomplete books-and-records archive, a marketing piece that never routed for approval, an attestation that quietly failed to send. Greenboard's own posting frames it plainly: "Our customers are regulated financial institutions, which means a missed bug is not just a bad user experience. It can mean such an archive... Quality is part of the product we sell."

Testing the AI layer introduces a different kind of difficulty. The platform uses large language models to detect risk in communications, route decisions, and prepare compliance tasks. The job description calls out "golden datasets and quality bars for output that is not deterministic." The engineer they seek has "tested a system where output is not deterministic and had to define correct as a threshold rather than an exact match." This is not regression testing as traditionally practiced. It is evaluation infrastructure: curated representative inputs, measurable quality thresholds, and a feedback loop that catches drift before it reaches a compliance officer's screen.

Data integrity at scale compounds the problem. Greenboard ingests tens of millions of messages a day. The testing framework must validate end-to-end from that ingestion layer through to what a compliance officer actually sees. The job description emphasizes the need to "test a system you cannot fully see from the UI. Comfortable with REST APIs and querying the database directly, because verifying a feature here means checking what landed and what the retention policy did with it, not just what rendered." UI automation alone is insufficient; the engineer must write tests that reach the persistence layer and verify retention, ordering, and completeness.

The tooling choices reflect that reality. The posting asks for Playwright or Cypress suites "that survive a redesign, with opinions about selectors, fixtures, and what belongs in an integration test versus an end-to-end one." The emphasis on durability over speed suggests a team that has watched brittle test suites collapse under product iteration and decided not to repeat it. The engineer must also "read a feature spec and name the states nobody thought about: the employee who left mid-attestation, the message that arrives out of order, the marketing piece edited after approval." Edge-case modeling is treated as a core competency, not an afterthought.

Release ownership rounds out the mandate. The hire will own "bug triage, release sign-off, and the definition of deployment-ready at Greenboard." The posting wants someone who has "held a release before, explained why to someone who wanted it out, and been right." That is a cultural signal: the test framework is not a quality checkbox; it is the mechanism by which the company decides what ships. In a regulated domain where the cost of a false negative is a regulatory finding, that authority sits with the engineer who built the tests, not the manager who wants the feature.

Domain experience carries weight. The role asks for three-plus years testing software where being wrong triggers an audit: fintech, healthtech, payments. A SOC 2 audit background is listed as a specific qualification, as the engineer should know what evidence a test suite can produce for one. The posting also values time spent watching a compliance officer, an advisor, or a back-office team work. That proximity shapes the edge-case intuition the job demands.

The stack runs Node.js and React on AWS with a queue-based ingestion pipeline. Compensation for the AI automation engineer role:

Source Range
Ashby listing $150,000–$200,000 + equity
OpenTalent $200,000

The role is based in New York City. Interest in financial services is a plus, not a requirement, because the systems thinking matters more than the sector pedigree.

There is no existing playbook for this combination: regulated-data-scale, non-deterministic AI outputs, and a mandate to build the testing infrastructure from nothing. The person who takes the role will write the playbook.

What Frontier-Tech Teams Can Learn

The testing problem Greenboard faces — verifying systems that produce different outputs for identical inputs — is not unique to financial compliance. It is the central engineering blocker for every high-stakes domain adopting AI. U.S. defense AI spending for FY2026:

Source Amount
DoD AI/autonomy budget request $13.4B
Congressional enactment (autonomous/unmanned) $9.8B
Navy AI spending increase (22.7% YoY) $308M

Congress enacted $9.8 billion toward autonomous and unmanned systems. The Navy alone added $308 million in AI spending, a 22.7 percent year-over-year jump. Yet a veriprajna.com analysis of defense acquisition notes that traditional test and evaluation assumes deterministic behavior: the same inputs produce the same outputs. ML-based systems are probabilistic. That mismatch stalls fielding.

Aerospace has spent decades building certification frameworks for deterministic software. Standards like DO-178C for software and DO-254 for complex hardware, ARP4754B for system development, and ARP4761A for safety assessment assign Design Assurance Levels based on failure impact. The certification process follows the Swiss cheese model, with multiple defense layers acknowledging system flaws. EASA has chosen an incremental approach for different autonomy levels, with the second version of its concept paper for Level 1 and 2 machine learning applications under review. The FAA employs a bottom-up approach, collaborating closely with applicants through AI Roadmap and Technical Exchange Meetings since September 2023. The Overarching Properties Working Group, with NASA and the FAA, developed a framework of three high-level properties: Intent, Correctness, and Innocuity. EUROCAE/SAE WG-114/G-34 is developing ARP6983 to guide AI-enabled system development, introducing the concept of an ML Constituent that includes models, traditional software, and hardware.

"AI systems are inherently imperfect and often nondeterministic, so they will make errors," said Elizabeth Davison of The Aerospace Corp. "While AI excels in predictable tasks and structured environments, it requires robust algorithms, extensive training data and adaptive methodologies to perform effectively in uncertain or rapidly changing conditions."

Medical devices face parallel pressure. In January 2025, the International Medical Device Regulators Forum released a final document identifying 10 guiding principles for Good Machine Learning Practice, building on 2021 principles from the FDA, Health Canada, and the UK's MHRA. These principles consider the total product life cycle. Explainable AI techniques are widely used in medical assessment because they are crucial for ensuring AI-driven medical assessments are transparent and interpretable by doctors.

Defense programs confront additional constraints. A transformer model running on an A100 in a data center is useless on a UAV payload constrained by size, weight, and power. Thermal management degrades performance by half in enclosed platforms. Security-driven restrictions on open-source tooling within classified networks mean the MLOps pipeline that works in commercial cloud does not work on SIPR or JWICS. CMMC 2.0 compliance, the FY2026 NDAA AI controls, ITAR technical data ambiguity, and DFARS 252.204-7012 must be handled as one integrated compliance architecture. Shadow AI — undocumented AI tool usage by cleared personnel on networks connected to controlled unclassified information — is the biggest audit finding.

Robotics teams at companies like Shield AI, whose Hivemind autonomy software has piloted 26 vehicle classes, and Skydio, which uses computer vision for infrastructure inspection, face the same verification gap. Energy sector operators deploying AI for grid management and nuclear plant monitoring inherit the same probabilistic verification challenge. Biotech firms building AI-driven drug discovery and diagnostic tools must satisfy FDA expectations for model transparency and life-cycle data governance.

The lesson from Greenboard's hiring push is structural: the first dedicated automation engineer owns that framework, those suites, and that release process, then embeds alongside that team. That pattern — test infrastructure as a first-class product, not an afterthought — repeats across frontier tech. MLOps practices for dataset versioning, experiment tracking, and model registry throughout the entire life cycle are crucial for ensuring consistency and traceability. Digital twins move from visualization tools to operational decision systems, but building one that reduces unscheduled maintenance requires continuous sensor data integration, validated physics-informed models, and a maintenance planning interface that dispatchers trust enough to act on.

The talent gap is real. BCG found 70 percent of aerospace and defense companies cite AI recruitment as a core challenge. The large consultancies face the same talent gap they are hired to fill. The combination of ML engineering depth and domain fluency — whether aerospace, defense, robotics, energy, or biotech — is the scarce resource. Greenboard's bet on a single automation engineer to solve non-deterministic testing mirrors the choice every frontier-tech team faces: build the verification layer first, or ship untestable code.


The engineer who takes this role will not just write tests. They will decide what "correct" means for an AI system that lives inside a regulator's evidence request. The framework they build will determine whether Greenboard's AI-native compliance layer can scale without compromising the audit trails that 500-plus financial institutions already trust with their regulatory exposure. The first test they write will be the one that proves the system works — not once, but every time the model changes, the prompt shifts, or the regulator asks for proof.


Working in AI? Zero G Talent tracks the openings: see every open Databricks role, browse AI jobs, openings at Anthropic, and the people building the field.

Ready to Start Your Space Career?

Browse artificial intelligence jobs and find your next opportunity.

View artificial intelligence Jobs