Skip to main content
← artificial intelligence

90% of AI Agent Projects Fail. The Fix Isn't a Better Model.

By David Yu•

The Back-Office Bottleneck in Lean Fintech

In April 2025, Visa and Mastercard each launched agentic payment protocols, followed by Stripe's Agentic Commerce Suite in December 2025 — but the real pressure isn't at checkout. It's in the back office, where a mid-sized payment processor handling 50 million monthly transactions faces 100,000 genuine exceptions every thirty days, and no hiring plan can solve it.

Fintechs are deploying autonomous AI agents to run leaner back-office operations, but regulatory demands for auditability and the complexity of financial systems are forcing a new engineering paradigm centered on guardrails and traceability.

At a 2 percent exception rate, those 50 million transactions yield one million items needing human eyes. Even if 90 percent are false positives, 100,000 real exceptions land on desks monthly — far beyond any reasonable team capacity. Manual reconciliation at this volume carries a persistent error rate of 0.5 to 2 percent. For a company moving $1 billion monthly, that leaves $5 million to $20 million in transactions with uncertain status at any moment.

A reconciliation analyst costs $80,000 to $120,000 a year fully loaded in major markets. A 20-person team exceeds $2 million annually before opportunity cost. Skilled accountants spend 80 percent of their time on mechanical matching instead of fraud analysis, cash optimization, or financial modeling. Fatigue compounds the problem: error rates in repetitive data tasks quadruple after four continuous hours. In reconciliation, errors create more errors — a mismatched transaction this month becomes a discrepancy that confuses next month's matching.

Traditional automation hit a wall. Rule-based bots require exact matches on amount, date, reference number. Real data is messier: bank descriptions truncate, dates shift by timezone, reference numbers get reformatted. Robotic Process Automation scripts were brittle and failed with minor input changes. Siloed systems — ERPs, banking portals, spreadsheets — offered limited integration. Doubling volume doesn't just double workload; it creates cascading backlogs. Monday's exceptions aren't resolved before Tuesday's arrive. By month-end, teams face thousands of unmatched items, most resolving to timing differences or formatting inconsistencies.

Regulators and auditors expect clean reconciliations. When they find systematic discrepancies, they don't assume innocent timing differences — they assume control failures. A fintech with chronic reconciliation gaps faces extended audit timelines, higher fees, regulatory findings requiring remediation, and potential restrictions on growth or new products. In extreme cases, enforcement actions follow.

The operational pressure is no longer theoretical. It's a daily constraint on how fast fintechs can scale, launch products, and stay compliant. Teams that solve it aren't hiring their way out. They're building agents that reason across messy data, learn from every correction, and operate with the auditability regulators demand.

The Regulatory Mandate for Auditability

Federal financial regulators have made their position clear: existing laws apply to AI-driven decisions the same way they apply to human ones. The GAO's May 2025 report, mandated under Dodd-Frank, surveyed the landscape and found that OCC, the Federal Reserve, FDIC, CFPB, SEC, CFTC, and NCUA all rely on current statutory authority rather than drafting new AI-specific rules. That stance forces fintechs to build compliance into the architecture of every agent from day one, not as a retrofit.

The numbers back the pressure. OCC has issued 17 matters requiring attention related to AI use since fiscal year 2020. CFPB has brought six enforcement actions in the same window, including a 2022 case against a large bank whose automated fraud detection system unlawfully froze accounts. SEC charged at least eight parties in 2023 and 2024 for false and misleading statements about purported AI use. NCUA issued a document of resolution and a regional director letter to a credit union that let an AI-driven program instantly approve loans without income verification. These are not hypothetical warnings — they are enforcement actions with named institutions and documented failures.

Explainability sits at the center of every supervisory letter. The GAO report cites the Consumer Financial Protection Bureau and the banking regulators: insufficient explainability inhibits a financial institution's understanding of a model's conceptual soundness, blocks independent review and audit, and makes compliance with laws and regulations more difficult. When SEC examined roughly 30 registered investment advisers in 2023, examiners found most lacked comprehensive policies and procedures governing AI use. Several had misrepresented their use of AI. OCC's 2019–2023 review of seven large banks found that risk assessments did not explicitly capture risk factors and complexities unique to AI models, and the banks provided limited information on efforts to evaluate bias and fair lending issues.

The bias risk is quantified. Testimony cited by GAO showed that some AI models can infer loan applicants' race or gender from application data or create complex variable interactions that produce disproportionately negative effects on protected groups. The Financial Stability Oversight Council warned that as models grow more complex, identifying and correcting biases becomes increasingly difficult. CFPB's two circulars (May 2022 and September 2023) clarified adverse action notification requirements when creditors use AI or other complex credit models. The message: if an agent denies credit, the institution must be able to explain why in terms a regulator and a consumer can audit.

State and international regimes add layers. Colorado's Artificial Intelligence Act, passed May 2024 with an effective date of June 30, 2026, requires developers of "high-risk" AI systems to use reasonable care to protect consumers from known or reasonably foreseeable risks of algorithmic discrimination. Covered businesses must conduct impact assessments, provide disclosures, and maintain risk management policies. The EU AI Act applies to development, deployment, and use of AI in the EU regardless of company location. A fintech whose agent touches a Colorado resident or a European merchant now faces two distinct audit regimes on top of federal expectations.

Regulators themselves are adopting AI for supervision — identifying risks, supporting research, detecting potential legal violations and reporting errors. Most told GAO that AI outputs inform staff decisions but are not used as sole decision-making sources. That distinction matters: the same standard applies to the institutions they supervise. An agent that moves money or approves credit cannot be a black box. It must log every decision, every data source, every tool call, and every human approval in a form an examiner can trace without reverse-engineering the model.

OCC's expectations crystallize the design requirements: risk management programs, adequate data management programs, adequate privacy and cyber controls, third-party risk management programs. A final interagency rule now requires mortgage originators and secondary market issuers to implement quality control standards for automated valuation models. CFPB has set expectations for automated customer service technologies, fraud screening technologies, and credit and lending decision technologies. CFTC reminded regulated entities in December 2024 that Commodity Exchange Act obligations still apply when AI facilitates or monitors derivatives transactions.

The technical implication is unavoidable. An agent that reconciles payments, underwrites risk, or flags fraud must operate within scoped permissions, require human approval for material actions, and produce an immutable audit trail that maps each output to the input, the policy, and the person who authorized the policy. That is not a feature request. It is the minimum viable product for any fintech that wants to deploy autonomous agents without inviting an enforcement action.

The Runtime Blueprint: Building Guardrails for Autonomous Payments

The first agent is the easy part. Running agents on payment data takes isolated computers, scoped credentials, role-based access, approvals, an audit trail, evals, model routing, and fallbacks when a provider goes down. An engineering team can spend two quarters on that before the first agent touches real data. That assessment, from Runtime's own documentation, frames the problem the company was founded to solve. Gus Trigos, co-founder and CEO, arrived at this insight after building Mentum, a Y Combinator-backed startup that deployed AI agents for procurement and supply-chain workflows across Latin America's largest asset and wealth managers, handling $20 billion in AUM data and 120,000 transactions a month. Mentum was acquired by Nuvocargo in 2025. Before that, Trigos worked as a quant in BlackRock's AI Quantitative Investment Group. He founded Runtime in 2026 with Carlos Volante; the team sits at seven people in San Francisco.

Runtime's agent harness sits between payment operations teams and the infrastructure they already use. Agents reach tools through APIs, databases, CLIs, and MCP servers, and they use a browser for processor, network, and bank portals that have no API. Common connections include the processor and ledger, Salesforce, Zendesk, Jira, Snowflake, Postgres, Splunk, and Datadog. Each agent runs on its own computer in the customer's cloud — AWS, GCP, Azure, or self-hosted via Helm chart — so PCI data never leaves the environment. Card numbers, SSNs, and secrets are masked before any prompt or log. For the most sensitive work, teams can run open-weight models locally inside their network. The platform is SOC 2 audited and does not train on customer data.

Trigos said: "Agents start read-only, and you choose the steps that need a person — like releasing a held payout, applying a reserve, booking a ledger entry, or replying to your sponsor bank."

That human-in-the-loop design is deliberate. Every run is recorded end to end: the trigger, every query and tool call, what each returned, the approvals, the cost, and the result. A stuck payment touches support, payment ops, finance, and compliance. Every finding lands in the same governed memory, so no team investigates it twice. Agents learn from resolved cases and propose updates to their own skills: what support learns about a stuck payout, payment ops already knows.

Runtime publishes concrete metrics from production deployments:

Metric Result
Run volume increase 10× more runs
Cost per run −93%
Total weekly spend −29%
Model cost on routine enrichment −86%
Steps per case −40%
Model spend on routine path $0

The harness also produces audit-ready outputs. In one example, an agent analyzing a dispute queue output: "Dispute ratio is 1.4%, above the 0.9% VAMP threshold. KYB still matches, but the website changed owners on Sep 22 and 61% of disputes are 'not received'. Recommend a 25% rolling reserve." Behavioral drift detection caught a model shift in under five minutes; an automatic circuit breaker paused the agent, and the ops team reviewed the model before reactivation — preventing further revenue loss. Continuous fairness monitoring runs on all lending agents, generating real-time metrics and automated evidence packages for regulators, replacing expensive periodic audits with always-on accountability. Automated controls blocked unauthorized external data transfers, logged every data access event, and maintained a complete consent chain.

RuntimeAI implements all 10 governance layers identified by SACR research (Stanford + MVP Ventures). Competitors implement three to five. Only RuntimeAI covers every layer — from discovery to cryptographic signing. The company is also building an agent marketplace with supply-chain verification via SBOMs and one-click deploy with policy templates, a post-quantum security platform compliant with NIST FIPS 203/204/205, and an agentic enablement platform covering commerce, payments, identity verification, PII protection, memory, and fraud prevention.

For fintech engineering teams, the implication is clear: the infrastructure to run autonomous agents on money movement safely does not have to be built from scratch. It can be adopted, configured, and extended, provided the team understands the guardrails well enough to set them.

The Divergent Strategies of Payments Giants

While back-office agents reconcile ledgers, a parallel arms race is unfolding at the front office. The three biggest names in payments all made major moves on AI agent infrastructure in 2025 and 2026, signaling a new front in the payments war. Visa, Mastercard, and Stripe are not chasing the same problem. Their architectures reveal fundamentally different bets on where value accrues when software starts spending money.

Visa's strategy centers on making the existing card rail AI-aware from the top down. The company launched Visa Intelligent Commerce in April 2025, a global initiative combining scoped tokenized credentials for AI agents, behavioral and issuer-side authentication for machine-initiated payments, and integrations with major LLM platforms. The centerpiece is the Trusted Agent Protocol, unveiled in October 2025 in collaboration with Cloudflare. Built on HTTP Message Signature standards and aligned with WebAuthn, TAP lets agents pass verified identity, intent, and payment credentials to merchants with minimal changes to existing checkout flows. Visa's answer to the core problem — how does a merchant know an agent is legitimate and not a malicious bot — is a cryptographic framework that slots into today's merchant stack. By December 2025, Visa announced that hundreds of secure, agent-initiated transactions had been completed in live production environments with partners including Skyfire, Nekuda, PayOS, and Ramp. Over 100 partners globally are now working across its agentic commerce ecosystem, with more than 30 actively building in the Visa Intelligent Commerce sandbox. As of April 2026, Visa had enrolled 85-plus issuer partners in the TAP testing programme across Asia Pacific and Latin America.

Visa's fraud model treats TAP-authorized agent transactions as equivalent in risk to a card-present chip transaction. The liability shift is explicit: the merchant is protected from disputes on in-scope agent transactions.

Where Visa leads with merchant trust verification, Mastercard has leaned harder into the identity and consent layer. Mastercard launched Agent Pay in April 2025, built around a new concept called Mastercard Agentic Tokens. The program includes an explicit consent flow — the consumer approves not just the initial authorization but the ongoing agent relationship — and ties into Mastercard's existing biometric authentication and Decision Intelligence fraud prevention systems. Mastercard's partnership with FIS, announced in January 2026, aims to bring Know Your Agent (KYA) capabilities to issuing banks, helping them authorize and monitor agent-initiated transactions at scale. The company has executed agentic payments in Australia, the United States, India, and (as of April 2026) Hong Kong. By Q1 2026, Mastercard disclosed that Agent Pay is enabled across nearly all Mastercard cards globally. On June 10, 2026, Mastercard launched Agent Pay for Machines (AP4M) with Coinbase, Stripe, Adyen, and 27 additional founding partners, a rail explicitly designed for sub-cent economics where card interchange would be 1,000× the transaction value. Mastercard created a separate chargeback reason code, MC 4849 (Agent-Initiated Transaction), for Agent Pay transactions. Disputes filed under this code are assessed against the agent profile ID, not the merchant's standard dispute record, and merchants with high Agent Pay volume should monitor MC 4849 volume separately; it is not included in the standard chargeback ratio calculations used by acquirers for account standing.

Stripe occupies a critical bridging position. Its Shared Payment Tokens work with both Visa Intelligent Commerce and Mastercard Agent Pay, as well as with BNPL providers Klarna and Affirm. For merchants already on Stripe, adding Agentic Commerce Protocol support reportedly requires a single line of code. Stripe shipped its suite that same month, and its Agentic Commerce Protocol (ACP), co-developed with OpenAI, powers ChatGPT checkout using Shared Payment Tokens. Stripe's Machine Payments Protocol, launched in March 2026, settles payments on Tempo (a Layer 1 blockchain purpose-built for payments) but critically supports fiat payment methods alongside stablecoins. Stripe shifts chargeback liability from the merchant to the issuer, extending the same protection. For most B2C applications launching in 2026, Stripe's Agentic Commerce Suite is the lowest-friction starting point.

Transaction Type Stripe SPT Visa TAP Mastercard Agent Pay x402 (USDC/Base)
Standard Consumer ($50) ~1.9–2.9% + $0.30 ~1.5–2.5% + interchange ~1.5–2.5% + interchange ~$0.0002–0.001
Micropayment ($0.001 × 1M/mo) ~$1,000 + $300k fees (No) ~$1,000 + $300k fees (No) ~$1,000 + $300k fees (No) ~$1,200 gas + $0 fees (Yes)

The economics make clear why x402 and AP4M exist: the micropayment use case is completely unserved by standard card infrastructure. An AI agent paying $0.0003 per API call cannot use a credit card. x402, open-sourced by Coinbase in May 2025 and moved to the Linux Foundation's x402 Foundation on July 14, 2026, processed 165 million cumulative transactions by late April 2026 with approximately $50 million in settled USDC volume. In the 30 days preceding the Foundation's launch, x402 averaged 32 cents per transaction across roughly 75 million transactions totaling $24 million. Transactions valued at $1 or more accounted for 95% of total dollar volume as of mid-2026, up from 49% in early 2025. Cloudflare reported its network was processing 1 billion HTTP 402 "payment required" responses per day at the time of the launch. But x402 has no dispute mechanism. On-chain stablecoin transactions are final. A consumer whose AI agent makes an incorrect or unauthorized purchase via x402 has no equivalent to a credit card reversal. The tax question compounds the burden: each on-chain stablecoin transfer is a taxable event under current IRS guidance.

The protocols are not necessarily competitors; many are designed to compose. Google's Universal Commerce Protocol, unveiled at NRF 2026 with backing from Shopify, Walmart, Target, and over 20 other partners, plugs into AP2 for payment authorization and Anthropic's Model Context Protocol for tool access. Visa's Trusted Agent Protocol aligns with both ACP and x402. Google's documentation describes x402 as the stablecoin settlement extension for AP2, handling the on-chain payment leg while AP2 handles the authorization and mandate layer above it. Stripe's ACP/SPT is the integration layer that provisions the networks' agentic tokens; you build against Stripe and Stripe routes to the rails, the same relationship you already have with Stripe Checkout. For most builders, you integrate one primitive: Stripe SPT/ACP. You generally do not integrate Visa and Mastercard separately.

Yet the fragmentation remains real. Enterprises deploying agentic commerce must prepare for multi-protocol support. The current split between ACP, UCP, AP2, Trusted Agent Protocol, Agent Pay, and MPP means most enterprises will need to support two or three protocols simultaneously. Businesses should contact their payment processor about ACP, Trusted Agent Protocol, and Mastercard Agent Pay support. Worldpay and commercetools are already live with agent-capable integrations. The open question isn't your code: it's availability, and whether your buyers' agents and their card issuers actually participate in your market yet.

This front-office arms race (agents buying things for consumers) sits apart from the back-office automation driving the rest of this story. The card networks and Stripe are building rails for agents that spend. Fintechs running lean operations are building agents that reconcile, underwrite, and comply. The engineering stacks diverge at the liability layer: front-office protocols fight over who eats the chargeback; back-office systems fight over who signs the audit log.

The New Engineering Stack for Trustworthy Agents

But whether front-office or back-office, the engineering challenge is the same: getting agents to production. Nearly nine in ten enterprise AI agent projects fail to reach production, and the failure is rarely the model. It is context rot, absent guardrails, no state persistence, or feedback loops that do not close. That figure, cited across multiple 2026 AI engineering reports, frames the problem fintechs face when they move agents from demo to ledger reconciliation, fraud triage, or compliance queues.

The discipline that addresses this gap is harness engineering: the design of everything around a large language model that turns it into a working agent: instructions and rule files, tool and MCP access, sandboxed execution, enforcement hooks, and orchestration logic. Mitchell Hashimoto, creator of Terraform, published the foundational definition in February 2026. The OpenAI Codex team documented the practical consequences the same month: a three-engineer team used harness engineering to produce a million-line codebase at 3.5 pull requests per engineer per day, with zero manually typed code. The framing "Agent = Model + Harness," popularized by LangChain's Vivek Trivedy in early 2026 and extended by engineers at Anthropic and OpenAI, has become the working definition for teams shipping in regulated environments.

Every production harness has three architectural layers. The names vary across frameworks, but the functional requirements are consistent. BCG's September 2026 synthesis identifies five key elements for an effective agentic operating system: Specs, Constitution, Control Panel, Context Hub, and Quality Gates. SPD Technology's breakdown maps to four layers: Instructions & memory, Tools & access, Execution & enforcement, Orchestration & observability. Each layer addresses specific vulnerabilities in agent execution. Several core components exist explicitly to combat context rot: the progressive degradation of model reasoning as the context window fills with historical logs, raw file dumps, and repetitive tool outputs.

The harness, not the model, is usually why two teams get different results from the same AI agent: one team demonstrated moving a coding agent from outside the top 30 into the top 5 on the Terminal Bench 2.0 leaderboard by changing only the harness, with the underlying model held entirely fixed.

The evidence for harness leverage is measurable. A separate LangChain engineering team gained 13.7 points on the same benchmark (moving from 52.8 to 66.5 on Terminal Bench 2.0) solely by refining system prompts, tool definitions, and middleware while keeping the foundation model constant. On SWE-bench Pro, swapping the agent harness changed pass@1 more than many model upgrades do. Same model, different harness: 23 percent to 52 percent pass@1 on GLM-5.2, and 15 percent to 36 percent on Gemma 4 26B. Harness rankings barely transfer across models (rank correlation -0.05), so a small model in the right harness can approach a much larger model in the wrong one. A bare frontier model was verified at about 30 percent on ARC-AGI-3; Prime Agent's harness took Opus 5 to 95.5 percent. The gap is widest on long-horizon work, the exact profile of financial reconciliation and compliance tasks.

Engineers must distinguish between two primary technical mechanisms: hooks and sandboxing. Hooks are enforcement points that intercept agent actions before they reach external systems: approval gates, policy checks, PII filters. Sandboxing isolates the execution environment so that a hallucinated tool call cannot mutate production state. Tool orchestration refers to the harness logic that decides which tools an agent can invoke, the sequence of execution, and how subagents or distinct models divide responsibilities. Ten well-described, non-overlapping tools outperform fifty overlapping ones because every tool's name and description eats into the same context-window space the model uses to decide what to call next; more options just means more chances to guess wrong.

The five failure patterns cluster predictably. Context rot is the most common: the model starts hallucinating, losing track of earlier instructions, or producing outputs that contradict constraints stated 50,000 tokens back. State loss is the second: the next session starts blind, users get inconsistent behavior because the harness has no persistence layer. Absent guardrails account for the third category: the agent produces outputs that violate compliance requirements, expose PII, or generate content that triggers legal review. Brittle tool integration is the fourth: error handling in the tool integration layer is absent or naive. The fifth category is no feedback loop: the agent makes a mistake, the mistake recurs, nobody captures the failure pattern, nobody encodes a rule to prevent recurrence.

Guardrail engineering has moved from optional to default infrastructure. With the EU AI Act's high-risk obligations applying from August 2, 2026, and the OWASP Top 10 for LLM Applications now treated as the canonical taxonomy of risk, fintechs need runtime controls that are auditable by design. Deloitte's 2026 AI report found that only 20 percent of organizations have mature governance models for AI agents, even as over 80 percent of technical teams have pushed past planning into active testing or production. Gravitee's 2026 report found that only 14.4 percent of agents went live with full security and IT approval, even as a similar share of teams are actively testing or deploying agents.

The platform landscape reflects this urgency. Galileo provides an AI evaluation, observability, and guardrails platform that turns pre-production evaluations into production governance controls, with a centralized open-source control plane (Agent Control, released March 2026 under Apache 2.0) for managing agent behavior at scale. NVIDIA NeMo Guardrails offers an open-source collection of software tools and NIM microservices using Colang, a domain-specific language for dialog policy control. Guardrails AI provides a Python framework for validating and correcting model outputs using composable validators, with a community-driven Hub of reusable safety checks. Lakera Guard, now part of Check Point after a $300 million acquisition in November 2025, delivers a real-time AI security API protecting against prompt injection, jailbreaking, data leakage, and toxic content with sub-50ms latency. No single platform covers every attack surface alone; teams stack complementary tools where gaps remain.

Vendor lock-in prevention is a harness architecture decision. Teams that build their harness logic around OpenAI-specific API features, Anthropic-specific prompt structures, or provider-specific tool schemas face expensive rebuilds when model prices change, capability gaps appear, or a better model releases on a different platform. The lock-in prevention pattern is straightforward: the harness treats the model as a pluggable component. Context assembly, tool dispatch, guardrail logic, and state persistence live in the harness, not in provider-specific SDK calls.

The skills required reflect this stack. Harness engineers design the execution environment: context management, tool dispatch, guardrail enforcement, state persistence, and feedback loops. Eval engineers build the verification discipline that begins where the harness hands off: deterministic test execution, LLM-as-judge rubrics, and trajectory analysis against curated benchmark datasets. Product teams define calibrated autonomy, scoping the decision per task so an agent earns the right to act without review only where the verification around that specific task is actually strong enough to catch it if it's wrong. Humans act more as system supervisors: they set goals and boundaries up front, leave the agent to handle routine work, and intervene at critical checkpoints. From the banking industry, BCG found these new structures disrupt traditional workforce pyramids, shifting toward more demanding design and strategy roles.

SPD Technology's deployment for a US ticketing platform shows the stack in production: 47 custom skills, 6 subagents, and 19 slash commands turn the team's conventions into executable workflows; the /feature and /bugfix pipelines stop hard before PR creation; a test-fidelity guardrail catches AI-written tests with phantom or weakened assertions. BCG built a fully agentic platform for a large Southeast Asian bank using the same principles: advisors tripled time spent actively engaging with clients, delivering a 30 percent-plus uplift in wealth-advisor revenue productivity with an associated four-to-five-times improvement in customer conversion, while engineering efficiency increased five times and speed from design to deployment rose 50 percent.

When the examiner asks why the agent applied a 25% rolling reserve on September 22, the answer sits in the audit log: dispute ratio 1.4%, VAMP threshold 0.9%, website ownership change, 61% "not received" disputes — every input, every policy, every human approval traced to a line the model didn't write but the harness recorded.


Working in AI? Zero G Talent tracks the openings: see every open Stripe role, browse AI jobs, the companies hiring, and the people building the field.

Ready to Start Your Space Career?

Browse artificial intelligence jobs and find your next opportunity.

View artificial intelligence Jobs