The Physics Problem Behind AI's Exploding Power Demands
As artificial intelligence workloads surge, data centers are consuming unprecedented amounts of electricity and water, driven by fundamental computational constraints that resist easy optimization.

Across America's fractured political landscape, data centers have emerged as an unlikely point of consensus—nearly everyone opposes them. These nondescript, climate-controlled facilities have become lightning rods for public anger as construction accelerates to support the artificial intelligence boom.
The scale of investment is staggering. Fitch Group Inc. projects that the world's five largest hyperscalers will invest $750 billion in data center construction this year, with roughly three-quarters of that directed toward AI infrastructure. Yet this expansion has collided with mounting concerns about job displacement and environmental strain. A Heatmap survey found that three-quarters of Americans now oppose new data center construction in their communities, while activists have staged more than 130 protests across dozens of states. The issue has even surfaced in midterm election discussions.
The underlying numbers justify the alarm. In 2023, data centers consumed approximately 4.4% of all U.S. electricity, a share expected to triple by 2028. Globally, data center power consumption is projected to double by 2030, reaching levels equivalent to Japan's entire annual electricity use. Some computer scientists warn that data centers could consume as much as 20% of worldwide electricity by 2035.
Water consumption mirrors this appetite for electricity. A single large data center can draw up to 5 million gallons of water daily—equivalent to a city of 50,000 residents. The United Nations estimates that global AI demand will consume Denmark's total annual water usage in the coming year alone.

Mechanical differences
The root cause lies in how AI fundamentally differs from traditional software. Conventional applications are written once and deployed across multitenant servers where adding new users incurs negligible additional costs. Generative AI has demolished this economic model, forcing the technology sector to grapple with thermodynamics, material constraints and resource scarcity.
"Traditional software has very low marginal cost because computation happens primarily on the user's device or cheaper multitenant infrastructure," explained Andrew Marshall, vice president of developer relations at Yugabyte Inc., which makes a distributed PostgreSQL database. "AI inference incurs a real computational cost for every interaction."
This distinction is crucial. Whereas conventional software executes a single code path for millions of users, each AI application performs distinct computations for each user. "Conventional software runs one code path for up to millions of users," Marshall noted. "An AI application runs a different computation for each one. That's what makes it worth paying for, and it's what defeats the caching and code reuse that give software its margins."
A single query to an AI assistant like OpenAI PBC's ChatGPT requires up to 10 times more electricity than a traditional Google search, according to the International Energy Agency. Creating a five-second video using generative AI models consumes as much electricity as running a household microwave continuously for more than an hour.
The sector is also experiencing Jevons Paradox, a phenomenon identified by 19th-century economist William Stanley Jevons: even when processing becomes more efficient, demand surges to the point that total consumption rises despite per-unit improvements. Hyperscalers, neocloud providers and numerous startups are pursuing ways to reduce AI's energy footprint, but progress faces headwinds. "Demand is still skyrocketing for more and more capacity," said Ramesh Chettuvetty, senior vice president of AI product and business at Lightbits Labs Ltd., which develops intelligent cache orchestration technology. Although Lightbits Labs claims its technology can boost the capacity of existing graphics processing units up to 16-fold, "there is not going to be any impact on [total] GPU capacity," he said. "It's not possible to meet demand right now."
AI's newness compounds the challenge. Forecasting demand remains largely guesswork, and despite token cost reductions over the past two years, AI projects remain unpredictable operational expenses rather than fixed capital investments. Gartner Inc. predicts that at least half of generative AI projects will exceed their budgets through 2028.
GPU tax
The AI cost crisis originates in the specialized hardware required for intensive calculations. Unlike traditional software running on general-purpose central processing units, generative AI relies almost exclusively on graphics processing units or Google LLC's tensor processing units. This specialization has created severe supply chain bottlenecks, converting AI infrastructure from a software engineering problem into a capital-intensive hardware acquisition competition.

GPU expenses stem from two interconnected factors: semiconductor fabrication costs and memory limitations. High-performance chips demand hundreds of manufacturing steps and specialized resources. AI chips must be fabricated using "ultrapure" water to eliminate microscopic silicon residue. A single semiconductor fabrication facility consumes up to 10 million gallons of ultrapure water daily. Since producing one gallon of ultrapure water requires roughly 1.5 gallons of tap water, a typical chip factory draws 15 million gallons of municipal water every day—equivalent to the consumption of approximately 33,000 households.
Then there is the "memory wall." An LLM must load its entire weight matrix—containing hundreds of billions of parameters—into local memory. To handle the intense data transfer rates these operations demand, hardware manufacturers deploy specialized, expensive high-bandwidth memory that stacks memory chips vertically to accelerate data movement.
This has triggered a severe global resource shortage. AI companies are currently purchasing an estimated 70% of the world's supply of high-end computer memory, according to The Atlantic, creating acute shortages elsewhere. Consumer computer memory and hard-drive storage prices have surged, with some laptop costs rising as much as 50%.
Despite astronomical hardware costs—individual high-end GPUs can exceed tens of thousands of dollars—actual production efficiency remains surprisingly low. GPU clusters typically operate at average utilization rates of just 10% to 12%.
Memory starvation drives this inefficiency. Because GPUs process data far faster than memory architectures can supply it, processors spend significant portions of their active cycles waiting for data retrieval. To compensate, operators frequently overprovision capacity at every layer, reserving clusters to handle sudden, unpredictable traffic spikes. The result is a skewed cost structure where operators pay full price for continuous, maximum-power hardware while using only a fraction of its capacity.
"AI gets expensive because you reserve capacity at every layer and use a fraction of it," Marshall said. "Demand is spiky, and nobody wants to be the layer that runs out."
Arthur Rasmusson, director of AI architecture at Lightbits Labs, recalled working with an LLM provider operating a billion-dollar cluster at 10% utilization most of the time to manage occasional traffic surges. "You might be shocked," he said, at typical utilization rates.
Training vs. inference
Model training represents the most resource-intensive phase of the AI lifecycle, but it is not the largest consumer of power and water over time. Training involves feeding massive datasets into a neural network to adjust its billions of parameters—a process requiring thousands of high-end processors running at maximum capacity for months.
Training's power requirements are extraordinary. The Economist reported that Meta Platforms Inc.'s Llama 3.1 model required 27.5 gigawatt-hours of energy to train, sufficient to power 7,500 American homes for a year.
However, training is a fixed, one-time capital event that can be distributed across millions of subsequent transactions. Inference—processing live workloads—represents the much larger expense. Jefferies Financial Group Inc. analyst Brent Thill has estimated that inference accounts for 96% of the energy consumed in AI data centers, according to The Economist.
This cost stems from two architectural limitations in contemporary deep learning: quadratic complexity and the autoregressive execution loop.
Traditional software is resource-efficient because it typically scales logarithmically or linearly. As input size increases, computational time required grows slowly. LLMs scale quadratically, requiring calculation of the mathematical relationship between every single word or token in a prompt and every other word. Doubling the input document size therefore quadruples the memory and computation required.

Autoregression exacerbates this scaling penalty. In this technique, a model predicts the next data point in a sequence using its own previous outputs as inputs. Autoregression compensates for computers' inability to reason like humans by mimicking thought processes using probability.
Humans can formulate an entire sentence mentally, but an LLM must execute a complete pass through its neural network to predict a single next token—a chunk of data the model uses to read, write and process information. The output is appended to the previous text, and the entire combined string is fed back into the model to predict the next token. It assumes that the future will follow past patterns.
To generate a 1,000-token response, the GPU must run its billions of parameters through a mathematical loop 1,000 times. Every interaction is a resource-intensive computation that cannot be easily cached or bypassed. Long conversations, document summaries or multiturn software development tasks can quickly become extremely computationally demanding.
The simplest approach to reducing costs and power consumption is asking the AI model to do less. Loading a model with millions of data points essentially wastes GPU capacity on calculations that could be performed on a desktop computer.
"The fact that a model can ingest hundreds of thousands or millions of tokens does not mean it should," said Varqa Abyaneh, founder and CEO of Opetek Ltd., developer of an AI reasoning system for capital markets.
LLMs excel at tasks like understanding ambiguous questions, decomposing complex problems and selecting analytical approaches. Conventional computers excel at performing calculations across millions of data points and can do so at much lower cost. Opetek's AI reasoning system, called Arius, separates the data the model requires from data that can be processed more cheaply in a Python program or Excel. This approach has yielded over 90% cost reductions in some financial scenarios.
Abyaneh said the approach applies to any data-intensive scenario. "The goal should not be to minimize reasoning," he said. "It should be to spend reasoning where reasoning creates value."
Agentic frontier
The financial and physical demands of AI inference intensify with the shift from simple, human-driven chat interfaces to autonomous agentic workflows. Unlike chatbots that await prompts, AI agents operate independently, executing multistep business workflows, calling tools and interacting directly with other software.
This shift dramatically expands variable token volume. Human users face physical limits on how quickly they can read and type, but machine-to-machine agentic loops can execute thousands of transactions in seconds. Anthropic PBC has estimated that multi-agent systems consume about 15 times as many tokens as a single-turn human chat session. International Data Corp. has projected that the number of actively deployed AI agents worldwide will exceed 1 billion by 2029, roughly 40 times as many as were in use last year.
Agents are so resource-hungry because they do not actually think but loop repeatedly through the same data. "If an engineering assistant is tasked with fixing an application bug, it runs a build, encounters a failure, and invokes local tools to investigate," Avichay Har-Tuv, finops team lead at CloudZone Inc., wrote in an article reviewed by SiliconANGLE.
"To make a decision, it pulls thousands of lines of verbose container logs, deep JSON structural payloads and identical database schemas, moving the entire block back into the cloud LLM's context window," he wrote. "If the first fix fails, the agent repeats the loop."
Each iteration causes the agent to retransmit the same database schemas and metadata across the network to a remote endpoint. "The overwhelming majority of data transmitted during these multiturn sessions is not high-value logical code or intellectual property, but infrastructure noise," Har-Tuv wrote.
Token waste

In a video tutorial on the YouTube channel Computerphile, Michael Pound, an associate professor of computer science at the University of Nottingham, demonstrated how such "token waste" can consume 60,000 to 100,000 tokens in minutes for a simple bug fix. Although tokens cost only a fraction of a cent each, costs—and power demands—accumulate rapidly across hundreds of tasks.
"It's easy for an agent to spin up a lot of cycles of time without you even asking for it," said Dave McCarthy, group vice president of cloud and datacenter infrastructure at IDC. "There aren't many circuit breakers."
Human inefficiency compounds the problem. AI is so new that few organizations have restructured the data and processes needed to help models perform at peak efficiency. Gartner said the budget overruns it forecast will stem largely from fundamental deficiencies such as poor architectural designs and inadequate operational controls.
"If an AI system doesn't know what a field means, which metric is authoritative, where the data came from or whether it can be trusted, it has to figure those things out while it's working," said Animesh Kumar, co-founder and chief technology officer of The Modern Data Company Inc., creator of a platform that contextualizes data. "Retries, unnecessary context and using a powerful model for a relatively simple task all add more compute."
Fragmented data sources introduce overhead by requiring AI models to pull information from multiple databases, increasing duplication and consuming tokens, said Michael Gale, chief marketing officer at EnterpriseDB Corp., which sells a commercial version of the open-source PostgreSQL database management system.
Gale likens a DBMS to a refrigerator, which is opened and closed occasionally. AI essentially accesses the refrigerator constantly, drawing power each time.
"In an AI world, you have to pull data literally 86,000 seconds a day," he said. He estimates that organizations can cut inference costs by 10% by vectorizing their data, allowing multiple sequential operations to execute in parallel. "If you can solve the energy consumption issue at the data layer, it gives you a lot more agility to handle some of the bigger stuff," he said.
Mitigations

Data center operators and model developers recognize that the current brute-force approach to scaling AI is economically and environmentally unsustainable. Numerous initiatives are underway to reduce power demands without sacrificing accuracy or performance.
The most promising near-term approach is quantization. In standard machine learning, model parameters are stored as highly precise 32-bit floating-point numbers. Quantization compresses these weights to as little as four bits. Reducing each parameter's size can shrink a model's memory footprint by over 80%, enabling networks to run on cheaper hardware without meaningful output quality loss.
At the architectural level, mixture of experts designs are delivering substantial operational savings in some cases. Instead of activating a massive, monolithic neural network for every query, an MoE model divides its parameters into specialized sub-networks or "experts." When a user submits a query, a routing algorithm determines which expert suits the task best and activates only that specific pathway.
If a user asks a coding question, for example, only the programming experts activate, leaving most of the network dormant. Google LLC researchers have estimated that MoE can reduce computation and data transfer volumes by 10- to 100-fold.
Model distillation uses a massive, high-performing model as a "teacher" to train a highly efficient, compact "student" model that can execute specific tasks at a small fraction of the cost. The student model learns to match the teacher's detailed probability patterns rather than processing raw data labels, allowing it to capture the deeper reasoning and nuances of the larger model.
Microsoft Corp. researcher Alexia Jolicoeur-Martineau has pioneered tiny recursive models that achieved success on complex logic tasks in biology and electrical engineering. Her work has demonstrated that highly structured, small-scale architectures can solve well-defined problems without the overhead of massive foundation models.
"Not every enterprise task requires the largest or most powerful model," said The Modern Data Co.'s Kumar. "Matching the model to the task can allow smaller or specialized models to handle a significant amount of work at lower cost." GPU leader Nvidia Corp. estimates that up to 70% of current LLM queries could be handled by SLMs without meaningful performance loss.
IDC's McCarthy said AI's relative newness means organizations struggle to understand how to use the technology efficiently and waste resources in the process. "Every time there's a new technology wave, we see people throwing the kitchen sink at it," he said. "You don't always need the latest and greatest GPU or model. But organizations don't have a lot of history to work with."
Some techniques that can yield the greatest efficiency benefits are already well understood. Retrieval-augmented generation reduces unnecessary steps by providing contextually relevant information. Persistent memory allows agents to reuse prior work rather than reconstruct context from scratch. Both are established techniques that cut processing overhead. "The greatest efficiency gains come from eliminating duplicated retrieval and repeated inference without weakening the quality of the context provided," said Yugabyte's Marshall.
Context optimization layers are also showing promise. Working from the assumption that prompts often contain redundant information, content optimization techniques intercept and streamline data before it reaches the LLM. The open-source Project Headroom pre-processes heavy payloads locally, strips out syntax boilerplate, isolates log files and substitutes lightweight cryptographic hashes for long text streams to reduce token consumption up to 95% without affecting accuracy.
Selecting the right model can also significantly impact costs and power consumption. "The answer is not to use less AI; it's to be smarter about where you use it, match the right models to each use case, and architect efficient context management and data processing," said Gonçalo Borrêga, senior director of product management at OutSystems Inc. "Flexibility is imperative. If another model can do the same job at a lower cost six months from now, companies should be able to switch without rebuilding their entire system."
Users should also consider return on investment. "What matters is whether value grows faster than cost," Borrêga said. "If an agent can reduce a two-hour process to three minutes, paying for inference can still produce a strong return."
Next-generation infrastructure
Although software optimizations help, achieving greater efficiencies depends more on overhauling physical infrastructure and computing hardware. Promising advances include specialized silicon architectures, predictive memory management, carbon-aware grid scheduling and advanced thermodynamic engineering.
A major hardware initiative involves shifting processing from general-purpose GPUs to specialized application-specific integrated circuits and TPUs. Google, which has been co-designing its own TPUs for over a decade, says its Ironwood custom inference chip is 30 times more energy-efficient than its earliest models.
Tech giants are also transitioning their data center fleets to direct liquid cooling and liquid immersion systems, which completely submerge servers in nonconductive synthetic oil that conducts heat but not electricity.
Numerous startups are rethinking how software interacts with hardware to eliminate processing bottlenecks. Two-year-old startup Mindbeam AI Inc. recently released an open-source inference framework that it says can run LLMs on commodity consumer CPUs, bypassing the GPU bottleneck entirely for certain workloads.
Mindbeam's approach constrains neural network weights to just three values, eliminating the complex floating-point multiplication operations that consume processing cycles. The firm said its approach delivers a 17- to 96-fold improvement in CPU throughput, while slashing memory consumption.
Lightbits Labs, ScaleFlux Inc. and FarmGPU Inc. are collaborating on an architecture to ease AI inference bottlenecks caused by limited GPU memory. LightInferra software stores and reuses key-value cache data across nonvolatile memory express storage and managed GPU inference infrastructure to predict when data will be needed and moves it closer to processors. Lightbits says it can triple inference requests on existing GPUs while cutting power and infrastructure costs by 65%.
Inferra by Lightbits Labs is a "predictive prefetch" algorithm that analyzes upstream and downstream signals to predict what data the processor will need next and streams granular memory blocks only when needed. The company, which will release Inferra on Sept. 9, said it can raise GPU utilization from an average of 10% to 12% to more than 75% while enabling 16 times more concurrent inference sessions on existing GPU infrastructure.
Groq Inc. has raised $650 million for a chip design called a language processing unit that accelerates AI inference by bypassing GPU bottlenecks, delivering extreme token-generation speeds. SambaNova Inc. has raked in $1 billion to build an inference chip that it said can speed GPU processing up to fivefold.
Cerebras Systems Inc. is hoping to steal some of Nvidia Corp.'s market share with a wafer-scale architecture that bypasses the memory constraints that slow conventional GPUs by keeping model weights in on-chip memory. The company says its approach can boost throughput by 500%.
Despite efforts to displace GPUs, "they retain a major advantage because of their mature software ecosystem, flexibility and ability to support rapidly changing models," said Arun Chandrasekaran, distinguished VP analyst at Gartner Inc.
AI giants are also shifting from static grid consumption to carbon-aware computing. Nvidia claims to have significantly reduced GPU power requirements in its latest Vera Rubin platform by compressing the numerical values that an AI model calculates to determine how strongly each word or token should relate to other tokens when generating a response, thereby minimizing unnecessary calculations and data movement. Its DSX MaxLPS framework dynamically coordinates power limits across racks and workloads, enabling data center operators to run up to 40% more GPUs within the same power budget.
Google's system for Carbon-Intelligent Compute Management, for another example, automatically analyzes day-ahead carbon intensity forecasts and generates schedules that limit computing resources available to flexible background workloads during peak grid strain.
The road ahead
As promising as these initiatives are, they may ultimately founder on the shoals of Jevons Paradox. In its most recent earnings announcement, Nvidia noted that its growth is constrained by supply, indicating that current demand is nearly limitless.
"Better inference efficiency and hardware improvements can offset" some of the growth in power demand, said Gartner's Chandrasekaran. However, "the likely outcome is more AI compute needs overall, even if the cost and energy consumed per prompt or task continue to fall."
The complexity of AI processing also defies simple solutions. An example is the release of DeepSeek V3 in late 2024. It was initially hailed as an environmental and financial breakthrough, thanks to algorithmic optimizations that allowed the model's final training run to be completed ten times faster than comparable models, with a proportionate drop in power and inference costs.
Yet any potential energy savings were immediately swallowed by the introduction of "reasoning" models such as DeepSeek R1, The Economist noted. Because those models use a methodical approach called Type 2 thinking, breaking a problem down, testing multiple approaches and validating its work before settling on an answer, they require significantly more processing time per query. The efficiency gains of V3 were quickly eaten up by the extended thinking times of R1.
Agents have a similar effect. Because they can search the web, write code and execute multistep tasks, a single request consumes orders of magnitude more energy than a simple chat query. Existing measurement frameworks do not yet account for the impact of idle-machine overhead, data-center cooling and network transport.
That means the economic and environmental costs of AI likely cannot be solved at any single layer of the stack. Improving efficiency requires a coordinated, full-stack approach that links hardware design, software optimization, data management and energy supply. Only then can the industry transition from the brute-force scaling of the past to an efficient, sustainable utility model.
"The question in front of us is 'Can we make the marginal cost of AI low as we scale AI usage?'" Chandrasekaran said. "We haven't solved for this yet."


