The Price of Electricity Belongs in Your AI Forecast
A CFO is walking through next year's forecast. The AI line has its own row now, separate from software, separate from headcount. Last year it was a rounding error. This year someone asks a simple question: what happens to this row if the unit price of a token doubles?
Nobody in the room has a good answer. The model everyone built assumes AI keeps getting cheaper, the way it always has. That is the assumption that just broke.
For two years, the AI story lived inside the CTO's budget: chips, models, pilots. That started changing a few weeks ago. Two states told data centers to slow down, and Dwarkesh Patel published a detailed analysis of why compute itself might get much more expensive. Patel is an independent writer and the host of a long-form podcast known for conversations with AI researchers, lab leaders, economists, and scientists. His work is useful here because he turns technical scaling claims into economic assumptions that finance teams can examine. Finance did not have to model any of this two budget cycles ago, and it does now.
The Bottleneck Was Always Power
Every few months, AI commentary finds a new bottleneck to worry about. First it was memory: could the machine hold the model. Then it was GPUs: could you actually get the chips. Then it was tokens: could you buy intelligence by the word, and how fast was the price falling.
Each of those constraints was real enough to dominate a news cycle in its turn, even though the floor underneath all of them turned out to be something else entirely.
The AI bottleneck was always power. Memory, GPUs, and tokens were just where the constraint happened to be visible first.
Every layer above electricity, the chips, the training runs, the inference APIs, the agent calling five tools before breakfast, sits on top of a grid that has to physically deliver power to a building. That grid used to have slack to spare. Enough of that slack has disappeared in the regions that matter most for AI buildout that two governors just said so out loud.
How Electricity Actually Reaches Your Token Bill
Here is where the popular version of the story gets oversimplified, and the chain matters more than the headline.
A CFO who buys API tokens is not directly exposed to household or retail electricity rates, and treating the two as the same number is a mistake worth correcting before it gets into a model. The path runs through several layers, and each one can absorb or amplify the pressure before it ever reaches an invoice. Electricity, generation, and grid capacity set how much it costs a data-center operator to run and expand a facility. That operating and capacity cost feeds into how much compute is available to rent and at what price. That compute price shows up in what a model or API provider charges, or in the margin and rationing decisions a provider makes instead of moving sticker price. Only after all of that does it land in an enterprise's AI budget and the ROI math behind it.
The chain, in order: electricity and grid capacity, then data-center cost, then compute supply and rental price, then vendor margin or token price, then your AI budget.
Every link in that chain has slack that can weaken or delay the pass-through. Chip supply, the capital available for new data centers, how efficiently a provider utilizes the GPUs it already has, model efficiency gains, competition among vendors, existing contracts, and a vendor's own willingness to absorb cost into margin instead of price can all blunt what would otherwise be a direct hit. A utility rate increase in one region does not automatically show up as a token price increase next quarter, which is why it functions better as a leading indicator worth watching than as a formula that can be plugged straight into a spreadsheet.
The practical takeaway for finance is to treat power and grid signals as an early warning system rather than a pricing model: track the physical inputs first, then model a range of vendor outcomes rather than assuming any single one is inevitable.
The Payroll-Sized Question
Here is the part finance teams are not modeling yet, and it is the one that turns this from an infrastructure story into a budget line.
As companies move from using AI tools to running actual agentic workflows, agents that draft, check, escalate, and act with less human handoff, the token bill stops looking like a software subscription and starts behaving like a labor cost. What follows is best treated as a scenario-planning ceiling rather than a benchmark or a prediction for every company: in a mature agentic operating case, external model, API, and token spend, plus the infrastructure that runs your agents, could approach 20 percent of payroll inside the function where agents have taken over the most work.
Define the fraction before you use it. The numerator is external AI spend, model calls, tokens, and agent infrastructure, rather than internal engineering headcount, while the denominator is payroll inside whichever scope you are actually planning for, a single function or the whole enterprise, depending on how far agentic adoption has spread. Model it in bands instead of a single number: 5 percent as an early case, 10 percent as a plausible middle case, and 20 percent as the stress case for a function that has gone deep on agentic workflows, the ceiling worth stress-testing against rather than the number you put in a base-case forecast.
Now put two facts together. Your AI budget is heading toward a payroll-sized line in the functions where agents do the most work, at the same time the unit price of the compute behind those tokens is under real upward pressure, for the reasons the chain above describes. A small, cheap budget line absorbs a price shock quietly. A line approaching payroll size, growing more expensive per unit at the same time, is one the CFO has to model explicitly rather than notice after the fact.
Why Compute Itself Might Get More Expensive
Dwarkesh's argument deserves care, because it compresses easily into a headline that oversells it. His essay from late July and the video conversation he recorded of the same material in early August both work through what would have to be true for an extreme pace of revenue growth to continue, treating that as an open conditional he is exploring rather than something he is forecasting will happen.
The setup, from his essay: Anthropic's revenue has been growing roughly tenfold a year, while the compute available to labs has been growing roughly threefold a year. Dwarkesh asks what would have to be true for that ten-times revenue growth to continue for another year. Something has to close the gap between revenue and capacity: fatter margins, a bigger share of compute spent serving customers instead of training the next model, or compute itself getting more expensive. His read is that all three are already happening, and margins alone cannot close a gap that large. Spot prices for the relevant chips are already up more than 40 percent from earlier this year, by his account, and the secure, guaranteed capacity labs actually need to rent, rather than opportunistic spot instances, is trading well above even that elevated spot price.
The mechanism underneath is the part worth sitting with. As models get smarter, they get better at turning a fixed amount of compute into revenue, so labs and buyers become willing to pay more for the same chip. Dwarkesh uses a thought experiment to size that effect: if a model that performed like a human-level software engineer could run on a single H100-equivalent chip, that chip should rent for something like fifteen times today's spot price, based on what a human software engineer's labor is worth. The math is meant to show how far willingness to pay could rise as models get better, a way of sizing the pressure rather than a prediction of what any chip will actually cost. The same mechanism means weaker, cheaper uses of AI get priced out first, while the best models command a premium for using scarce compute efficiently.
Dwarkesh is careful to flag the honest counterargument himself: scarcity forecasts have been wrong before, and markets eventually find ways to make scarce things abundant again. He is careful not to claim compute stays expensive forever; he is simply describing a regime, and it happens to be the regime your next two budget cycles have to survive.
Once you are running agents against a paid API, a lab's internal cost problem becomes part of your unit economics, filtered through the chain described above.
Two States Just Made the Constraint Physical
The permitting story makes the abstraction concrete, even though the two states took different actions rather than following the same script.
On July 14, New York's Governor Hochul signed Executive Order 62, pausing discretionary state permits for data centers drawing 50 megawatts or more while the state builds out grid standards and completes an environmental review. On August 3, Texas Governor Abbott directed state regulators to audit every data center project in the grid interconnection queue, blocking new connections until that audit is complete, in a state where the queue is enormous and dominated by data center requests.
New York is building permitting standards before letting new large facilities proceed, while Texas is auditing existing requests before letting them connect: different mechanisms rather than one shared policy, held together by the same underlying signal that the grid cannot currently absorb everything AI wants to build, which is why two governments with very different politics each decided that was worth pausing to check.
Scarcity Doesn't Disappear, It Moves
Stay Naive has argued before that scarcity migrates rather than vanishing when a key input gets cheap. When drafting got cheap, the value moved to judgment about what to draft. The same rule applies here, just at the level of physics instead of workflow.
Cheap tokens simply moved scarcity upstream, into chips, then into power, then into the actual grid capacity of actual states, rather than removing it from the system. The commons question sitting underneath all of it is old and unglamorous: the grid is shared infrastructure, funded and permitted collectively, while the upside from using it flows mostly to private companies. Someone has to decide how that gets rationed, and right now the deciders are governors, not markets.
This is also a good moment for beginner's mind. The comfortable assumption of the last few years, that AI gets exponentially cheaper forever the way storage and bandwidth did, was never a law of physics. It was a subsidy story, and Stay Naive called that subsidy out directly when it was still easy to enjoy without asking who pays for it later. A naive thinker questions the slide that everyone else has stopped questioning. The slide said intelligence gets cheaper every quarter. The floor underneath it just pushed back.
The Honest Argument on Both Sides
Give both sides their strongest case, because this is not a story with an obvious villain.
The case for building anyway: pausing has a real competitiveness cost. Not every megawatt in an interconnection queue is a serious project, some are speculative reservations, but the companies that secure real power capacity now will have an advantage for years. Treating electricity as infinite was always a fantasy, though assuming it is now instantly and permanently capped would be its own overcorrection, since grids do expand, generation does get added, and some of today's scarcity is temporary friction rather than a permanent ceiling.
The case for pausing and pricing it in: households and CFOs are not being unreasonable when they resist absorbing costs that were never fully modeled. A data center that raises local electricity rates or strains a grid is a cost that someone else did not agree to pay. Pausing to build real standards, rather than approving first and measuring later, is a rational response to a genuinely new scale of demand.
Both arguments are really responses to the same underlying fact, that power has turned out to be more finite than the last two years of pricing suggested. The companies racing to secure capacity now and the governments pausing to check the math are each taking that finite reality seriously, just from opposite sides of the calendar, and neither side is arguing that the constraint is permanent.
What This Means For Finance, Over Time
Near term, this budget cycle: AI cost assumptions built on 2024 and 2025 pricing are already stale. Vendor and cloud invoices are becoming a real source of forecast variance rather than the background noise they used to be. Pilots that were "basically free" stop clearing the ROI bar first, because they were never actually free, just subsidized.
Medium term, the next two to four years: Compute starts behaving like any other volatile input cost, the kind procurement already knows how to hedge and forecast for. Spend concentrates on the highest-value use cases (see also why the wrong AI metric can hide this) while low-value automation gets cut first when unit prices rise. Which cloud region and which power market you build on becomes a real cost variable rather than a technical footnote engineers decide on their own, and model choice and routing, cheap model versus frontier model for a given task, becomes a finance decision as much as an engineering one.
Longer term: AI stops looking like unlimited, ever-cheaper software and starts looking like an industrial input tied to energy markets, with the cost curves to match. Companies that secured power-aware capacity early will show a different cost structure than peers who did not, the same way manufacturers with locked-in energy contracts outperform peers during a price spike.
The AI Cost Translation Workflow
Most finance teams do not have a repeatable process for turning a physical signal, like a power price or a grid queue, into a budget decision. Here is one simple enough to hand to a finance intern.
A. Track the leading physical signals. Power prices and contracts, grid interconnection constraints, data-center capacity announcements, and GPU rental rates. These move before your vendor invoice does.
B. Translate signals into vendor scenarios. For each major AI vendor, sketch what happens under no pass-through, partial pass-through, and full pass-through, and note when existing contracts lock in current pricing versus when they come up for renewal.
C. Map AI spend by workflow and owner. List which use cases and workflows generate the spend, and who inside the company is accountable for each one. You cannot cut, route, or renegotiate what nobody owns.
D. Stress test the budget. Run the numbers at 5, 10, and 20 percent of the relevant payroll, combined with 1x, 2x, and 5x the current unit price of compute. That gives you a small grid of scenarios instead of one fragile point estimate.
E. Rank by verified business value. For each workflow, decide in advance whether a price shock means cutting it, routing it to a cheaper model, or renegotiating the contract behind it, ranking by what the workflow is actually worth rather than by how loudly its sponsor defends it.
F. Reforecast quarterly. Compute pricing is moving fast enough that an annual AI budget is already stale by the time it gets approved, so revisit the scenarios every quarter rather than once a year.
Reflection Point
If the bottleneck was always power, which line in your plan still treats AI like weightless software?