Skip to main content
frontier

Six‑Gigawatt Power Shortfall Looms for PJM Grid by 2027

By Elena Petrova

The Technical Foundation for a Pause

Scale AI, the data-labeling and evaluation contractor that serves major frontier labs, the U.S. government, and Fortune 500 companies, has highlighted technical constraints that make continued hyper-scale data center expansion unsustainable. The company's platform supplies annotated data, red-teaming, and evaluation suites labs use to measure whether new architectures work. That vantage shows how training runs scale, not just in GPU count but in megawatts and interconnect bandwidth each run demands.

The concerns Scale AI raises have fueled a cross-sector debate — more than 100 local moratoriums, over 300 state data-center bills filed in the first six weeks of 2026, and a federal pause proposal, grounding political arguments in the technical constraints of power and networking. On power, analyses cite grid stresses pushing PJM Interconnection (the wholesale market serving 65 million people across 13 states) to project a six-gigawatt shortfall by 2027, CNBC's figures put the shortfall at six gigawatts. In that region alone, power supply costs jumped from $2.2 billion to $14.7 billion, Brookings reported, in a year, with data centers driving nearly two-thirds, Brookings's data shows, of the increase. New York residential electricity prices have climbed 68 percent, CNBC found, since 2019. Arizona data centers already draw 1.5 gigawatts. A former Arizona utility executive disclosed 42 projects on deck totaling more than 10,255 MW. New AI campuses built around Nvidia's Blackwell chips consume five times the power of predecessors: 100 megawatt-hours a month or more.

On networking, a less-public but equally binding bottleneck exists. Frontier training clusters span hundreds of thousands of GPUs. Collective communication patterns — all-reduce, all-gather, pipeline parallelism — saturate conventional Ethernet fabrics. Nvidia's Spectrum-X Ethernet, designed for this scale, addresses these constraints but deployment lags behind cluster growth. Without grid upgrades and network architecture to catch up, the next model generation will hit physics the industry didn't plan for.

More than 100 local communities have enacted moratoriums. More than 300 state data-center bills were filed during that period. New York Governor Kathy Hochul signed an executive order barring new hyperscaler builds of 50 megawatts or more for up to a year. Senator Bernie Sanders and Representative Alexandria Ocasio-Cortez introduced a federal moratorium bill. Florida Governor Ron DeSantis champions a community veto.

"We have a limited grid. You do not have enough grid capacity in the United States to do what they're trying to do," DeSantis said at a Villages, Florida event.

The company's hiring signals the pivot's seriousness. Recent postings include a Director of Engineering for Physical AI (San Francisco) and a Principal Architect in Washington, D.C., roles focused on hardware-infrastructure interface, not model training. Scale AI is staffing for the constraint it warned about.

The Grid Is Already Cracking

For two decades U.S. electricity consumption barely budged. AI has broken that flat line. Brookings reported in April 2026 that AI drives demand growth several times faster than the historical baseline. North American data center power draw doubled from 2,688 megawatts at end-2022 to 5,341 megawatts at end-2023, a jump MIT partly attributes to generative AI. Globally, data centers consumed 460 terawatt-hours in 2022, ranking 11th largest on the planet between Saudi Arabia and France.

Forecasts for 2030 converge around 900 to 1,200 terawatt-hours. A summary of key projections:

Projection Source Metric Value Horizon
Brookings (IEA base) Global data center electricity 945 TWh 2030
Brookings (Goldman Sachs) Capacity demand increase 160% 2030 vs 2023
Deloitte Critical power for data centers 96 GW 2026
Deloitte AI data center annual consumption 90 TWh 2026
Anthropic U.S. AI sector new capacity needed 50 GW 2028
Eric Schmidt (Congress) Additional data center power 29 GW / 67 GW 2027 / 2030
LBNL (DOE) U.S. data center consumption 325–580 TWh 2028
CNBC (Tract) Buckeye, AZ campus secured power 1.8 GW Planned
CNBC (Lancium) Abilene campus ramp 250 MW → 1.2 GW 2025 → 2026

Training runs are the most visible spike. Independent analysis finds each run for models exceeding 175 billion parameters consumes 324 to 1,287 megawatt-hours, and models are retrained repeatedly. Anthropic estimates a single frontier model will require five gigawatts by 2027. Former Google CEO Eric Schmidt testified that data centers will need 29 gigawatts more by 2027 and 67 by 2030.

Inference, not training, may ultimately dominate the load. Brookings cites estimates that 80 to 90 percent of AI computing power goes to inference. The IEA models 30 percent annual growth for AI server electricity, predominantly inference, accounting for nearly half the net increase in global data center consumption through 2030. A single advanced generative AI query used 2.9 watt-hours in 2024, nearly ten times a Google search, though newer measurements show median text queries at 0.24 to 0.3 watt-hours, with long reasoning or multimodal prompts far higher.

The permitting clock is the hardest constraint. New power plants and high‑voltage transmission lines take over a decade in the U.S. and EU, while China added more than 400 gigawatts of capacity in a single year. AI companies are asking for gigawatts in years; the grid answers in decades.

Geographic concentration magnifies the stress. The United States accounts for 45 percent of global data center electricity use. In northern Virginia, the world's largest data center market, Dominion Energy expects demand to grow 85 percent over 15 years with data center load quadrupling. Local shares are extreme: 42 percent of Frankfurt's electricity, nearly 80 percent of Dublin's. Campuses in development routinely target one gigawatt or more. A one-gigawatt campus draws the equivalent of 700,000 homes or a city of 1.8 million people, more annual power than the retail electric sales of Alaska, Rhode Island, or Vermont. Equinix campuses are rising from 100-200 megawatts to several hundred each.

Chip-level power density tells the same story. CPUs historically ran at 150-200 watts. AI GPUs hit 400 watts until 2022; 2023 state-of-the-art parts run at 700 watts; 2024 next-gen chips rate at 1,200 watts. Average rack density is projected to climb from 36 kilowatts in 2023 to 50 by 2027. Cooling consumes 38-40 percent of data center electricity; air-based cooling alone can claim up to 40 percent of a facility's total draw.

The U.S. grid has about 1,200 gigawatts of generation capacity. Data centers sit at roughly 2 percent today — 25 gigawatts. A move to 10 percent implies 120 gigawatts, a fivefold increase. Meanwhile, nuclear and coal retirements could remove 150 gigawatts of firm capacity. Add LNG export terminals, cryptocurrency mining loads larger than Houston, and a manufacturing onshoring wave, and the arithmetic stops balancing. The pause isn't about compute. It's about whether wires and generators can show up in time.

The Network Choke Point

For decades, traditional off-the-shelf Ethernet ruled enterprise and cloud networking. It is cheap, standardized, and highly effective for general-purpose, high-entropy web traffic. Generative AI's massive growth has fundamentally altered data center design. As distributed model training spans hundreds of thousands of GPUs, the scale-out network connecting them has become a first-order bottleneck.

To scale an AI factory to hundreds of thousands of GPUs, traditional architectures require adding tiers, moving from two-tier to three-tier fat-tree topology. Each tier adds latency, jitter, and failure domains. Nvidia's Spectrum-X Ethernet takes a different approach: a two-tier topology scales to 128,000 endpoints, a three-tier to 16 million, without the latency and jitter penalties of traditional multi-tier fabrics.

The root cause is traffic pattern. AI training traffic is low entropy. GPUs continuously synchronize through collectives — All-Reduce, All-Gather, All-to-All — creating few, very large, synchronized flows. This exposes three structural limitations of traditional Ethernet.

First, hash collisions and stragglers. Equal-cost multi-path routing doesn't account for real-time congestion, so large flows collide on one link while others sit idle. Because synchronous collectives finish only when the slowest flow completes, one congested path delays the collective and idles many GPUs. Under a worst-case RDMA bisection benchmark, traditional Ethernet's static ECMP routing collapses from flow-hash collisions, dropping some GPU pairs to 25 Gbps throughput.

Second, lossy versus lossless operation. Congestion overflows switch buffers, triggering packet loss and retransmission, causing delays that hurt AI performance. RoCEv2 deployments use Priority Flow Control to reduce loss, but pause frames propagate congestion, create head-of-line blocking, and can stall the fabric.

Third, slow congestion control. Protocols like Data Center Quantized Congestion Notification are hard to tune for synchronized AI bursts. Delayed or excessive reactions cause buffer buildup, underutilization, and latency spikes. Traditional Ethernet shows wide latency jitter with 99th-percentile tail latency reaching 22 microseconds.

The "noisy neighbor" problem compounds all three. Standalone, standard Ethernet achieves a 735-millisecond training step. With background noise traffic, step times inflate to 1.18 seconds, a 1.6x slowdown. Noisy-neighbor traffic from one job bleeds into another, collapsing an All-to-All collective's bandwidth by more than 80 percent.

Resilience is equally fragile. When a host-to-leaf link flaps, traditional Ethernet's software-based load balancers take over 1.08 seconds to recover and reroute traffic. Dropping 10 percent of leaf uplinks collapses collective bandwidth by 50 percent or more from routing asymmetry.

Nvidia introduced Spectrum-X Ethernet as a hardware-accelerated architecture designed for giga-scale AI factories. It co-designs high-performance switches and host-side network cards to deliver predictable low latency, high fabric utilization, and robust resilience under extreme load. Full hardware acceleration is a core structural requirement, not an add-on. The fundamental design principle separates hardware-accelerated control loops by scope, signal, and responsibility.

In-switch Adaptive Routing replaces traditional Ethernet's static, hash-based ECMP. Spectrum-X switches implement per-packet Adaptive Routing, dynamically steering each packet based on real-time port congestion rather than a fixed hash. Targeted Congestion Control operates differently: the switch generates Explicit Congestion Notification marks only when adaptive routing capacity is exhausted and the queue keeps growing.

At the host edge, the Plane Load Balancer is a dedicated hardware engine inside the Spectrum-X SuperNIC. Spectrum-X Multiplane technology decomposes a host's massive bandwidth — 800 Gbps from an 8-lane ConnectX SuperNIC — into four physically independent 200 Gbps planes. Separating congestion-control state per plane and combining it with local egress queue depth isolates congestion to the affected plane.

If Plane 2 fails, the ConnectX SuperNIC detects the RTT timeout, masks Plane 2, and redirects traffic across the three healthy planes in under 3 milliseconds, preserving 75 percent of line-rate bisection bandwidth. This resilience lets operators run training workloads at near-optimal efficiency before all infrastructure issues resolve, drastically reducing "Time-to-AI."

A Different Way to Scale

While the industry presses for a pause on hyper-scale build-outs, parts of the enterprise market demonstrate a different scaling logic, sidestepping centralized megaclusters for distributed, domain-specific AI deployment. The clearest example sits in pharmaceuticals.

Danish neuroscience specialist Lundbeck announced in August 2026 it would extend AI across its U.S. commercial organization through an expanded partnership with Eversana, the Chicago-based commercialization services group. Trade press describes Lundbeck as the first pharma partner publicly named for Eversana's AI Agency platform, launched in 2025 and now anchoring the vendor's offering. The agreement broadens Lundbeck's use from pilot to day-to-day U.S. operations, a scaling move adding zero megawatts to any single data center campus.

Lundbeck's leadership frames the shift as strategic focus, not retreat. Chief executive Charl van Zyl characterized a parallel decision to withdraw from 27 markets and hand commercial operations to partners (Swixx Group across 22 countries, Zuellig Pharma in Southeast Asia, NewBridge Pharmaceuticals in Saudi Arabia and the UAE) as "sharpening" commercial strategy. The transition affects 600 workers, roughly 10 percent of the workforce, with a one-time cost of $61 million and no change to 2025 guidance. Yet the same leadership, posting from Copenhagen in August 2026, described a "bionic company" ambition: "bringing together the expertise, curiosity and judgement of our people with AI-enabled tools and workflows that help us learn faster, ask better questions and make stronger decisions." That language — human judgement augmented by AI, not replaced by brute-force compute — maps to an alternative scaling paradigm. The company also disclosed AI experiments in clinical trials, suggesting the pharma value chain may absorb significant AI workloads without a 100-megawatt training cluster.

Microsoft has published a trusted AI framework positioning responsible AI as a scaling enabler, not a compliance afterthought. The logic is straightforward: enterprises that cannot audit model behavior, data lineage, or deployment risk won't put mission-critical workloads on opaque frontier models, no matter the GPU count. By making trust signals a first-class product feature, Microsoft bets the next AI adoption wave will be gated by governance readiness, not raw parameter count. That bet aligns with Scale AI's positioning — selling "proven data, evaluations, and outcomes to AI labs, governments, and the Fortune 500" — suggesting a converging view: the bottleneck is not silicon but verifiability.

Lundbeck's Eversana partnership and Microsoft's trust framework share an architecture: they push AI workloads to the edge of business processes (commercial execution, clinical workflows, regulated deployments) where model size matters less than integration depth, auditability, and domain fidelity. Neither requires a new gigawatt-scale campus. Both represent capital-efficient scaling paths the hyper-scale pause debate rarely acknowledges. If power and networking truly gate frontier AI, companies already scaling inside those constraints (pharma commercialization teams, regulated cloud tenants, clinical-trial operators) may signal what comes next.

Local Opposition Is Growing

Opposition has reached regional opinion pages and municipal hearings. In Kentucky, the Louisville Democratic Socialists of America chapter campaigns for restrictions on hyper-scale AI data centers within city limits, tying facilities to three concerns recurring in municipal hearings: water for evaporative cooling, substation capacity crowding out residential ratepayers, and tax abatements shifting burden to homeowners.

Louisville's effort isn't isolated. Comparable moratorium proposals have surfaced from northern Virginia to central Ohio, typically triggered when a developer files for a special-use permit or requests a utility-service agreement revealing the project's true scale. In several cases, planning commissions ordered independent grid-impact studies, a procedural delay functioning as a soft pause. The goal isn't to block AI development categorically but to force disclosure of power-purchase agreements, water-rights transfers, and noise-mitigation plans before a vote.

The pattern reveals a structural mismatch. Hyperscalers negotiate with utilities and state economic-development offices under NDAs, while land-use decisions rest with city councils and zoning boards lacking technical staff to evaluate a 300-megawatt load profile. That asymmetry is what local campaigns try to correct through public pressure and ballot initiatives embedding capacity caps in zoning codes.

Whether these measures survive legal challenge remains untested. State pre-emption laws in several jurisdictions strip municipalities of authority over utility infrastructure. But the political signal is clear: communities that once competed for data-center tax base now demand the environmental-review rigor applied to heavy industry. The pause may arrive first not from a federal moratorium but from a patchwork of local "no" votes making the next 500-megawatt campus harder to site than the last.

What Fills the Pause?

The pause buys time. The question is what fills it. Deloitte's 2026 predictions make the timeline clear: almost all AI computing this year still runs in giant data centers or on-prem enterprise racks, not phones or laptops. The on-prem hybrid market will hit $50 billion, but the non-factory edge slice stays under $5 billion until robotics scales, likely after 2030. Inference workloads will claim two-thirds of all compute by 2026, up from a third in 2023, and the inference-chip market will top $50 billion. Yet Deloitte forecasts most cycles still land on $200 billion of power-hungry chips inside $400 billion facilities. The edge shift is real; it is not imminent.

That doesn't make edge irrelevant. For latency-critical, privacy-sensitive, or bandwidth-constrained workloads, pushing inference to local devices saves network and server capacity while keeping data on-site. Startups already ship on-device multimodal models skipping the cloud entirely. The lever is workload triage: decide per task whether the energy cost of round-tripping to a central cluster outweighs the edge device's lower throughput. Right now, the cluster wins for most training and heavy inference. That balance shifts as chips get more efficient per watt.

Inside the cluster, cooling is the next frontier. Air-based cooling eats a comparable portion of a typical data center's electricity. Liquid cooling cuts that dramatically (Deloitte cites 90 percent power reduction versus air) and supports rack densities of 50 to 100 kilowatts, well above the 36-kilowatt 2023 average and 50-kilowatt 2027 projection. One inference-optimized rack already draws 370 kilowatts, nearly triple its training counterpart. Liquid loops can eliminate chillers. The catch: they need lots of water in regions where it's already scarce. Rooftop rainwater capture could cover up to a third of cooling demand. Geothermal offers a twofer (direct cooling plus firm power), but deployment is site-specific. The technology works; supply chain and integration are still maturing.

Chip-level gains compound facility gains. Moving power delivery to the backside of the die slashes losses by 30 percent. Optical interconnects transmit data at one-tenth the energy of copper. Three-dimensional stacking with carbon-nanotube circuits promises 10x efficiency. Neuromorphic designs merging memory and compute target the brain's 20-watt budget, orders of magnitude below today's gigawatt clusters. A new generation of AI chips trains models in 90 days on 8.6 gigawatt-hours, one-tenth the prior generation's draw. These aren't lab curiosities; they're shipping or taping out now.

Grid-enhancing technologies buy capacity without new steel. Dynamic line ratings deploy in three months, adding 10-30 percent throughput. Flexible AC transmission and reconductoring take 8-18 months for 50-100 percent gains. FERC Order 2023's "first-ready, first-served" cluster studies and ERCOT's two-year controllable-load program shorten interconnection queues stretching to seven years. If data centers curtail 1 percent of load annually, grid operators could absorb 126 gigawatts of new demand with minimal build-out. Sixty-eight percent of executives Deloitte surveyed expect demand flexibility to become standard currency for speed to market.

Colocation rewrites the siting playbook. Pennsylvania's $10 billion coal-to-gas conversion will double capacity on existing transmission, targeting a 2027 start. Surplus interconnection at retired plants hooks new generation in under a year versus five-plus for greenfield. One hyperscaler, a renewable developer, and private equity are investing $20 billion in an energy park with colocated load, generation, and storage, operational by 2026. Peaker and load-following gas plants hold roughly 50 gigawatts of idle backup capacity to firm renewables for adjacent data centers.

Small modular reactors remain the long pole. They're in early development and won't deliver zero-carbon baseload this decade. Long-duration storage faces similar timing. The nearer-term lever is algorithmic: smaller models, targeted fine-tuning, and test-time compute budgets matching the task. Deloitte's survey puts technological innovation at 82 percent impact for closing the gap. The alternatives exist. The constraint is integration speed — not invention.


Working in frontier tech? Zero G Talent tracks the openings: see every open Scale AI role, browse frontier tech jobs, the companies hiring, and the people building the field.

Ready to Start Your Space Career?

Browse frontier jobs and find your next opportunity.

View frontier Jobs