Escape Velocity

Artificial intelligence / ai/energy-efficient-inference

Energy-efficient AI inference

mappedTRL 6 (6 of 9)horizon 2030scritical gap open

Running trained AI models to answer queries at a small fraction of today's energy per generated token, across the whole stack from the chip to the cooling plant, so that wider use of AI does not need proportionally more electricity.

Scope

In scope: the energy drawn to serve a trained language model, measured per token or per query, including accelerator, memory, interconnect, host and cooling overhead. Out of scope: the energy of training (not mapped yet); the logic devices, data movement and cooling that inference depends on, covered by computing/beyond-cmos-logic, computing/low-energy-data-movement and enablers/heat-removal; and what AI is used for, covered by ai/ai-for-science.

Readiness
TRL 6 (6 of 9)
Efficiency gains are demonstrated in production serving (a 33x reduction in energy per median Gemini Apps prompt over one year, company-reported), but the order-of-magnitude hardware gaps below are at laboratory or prototype level.
Serves
Affordable and clean energy, Industry, innovation and infrastructure, Climate action
Last reviewed
2026-10-04
Curators
none yet: volunteer

Impact

What reaching the target would change, and for whom. No claim is stronger than the evidence it cites.

Benefitextrapolation

Across models, serving systems and hardware, efficiency gains in sight could cut the energy of AI inference 8 to 20 times. At 1 billion queries a day with 10% long queries, demand would fall from 1.7 GWh a day to 0.8 GWh a day with efficiency interventions.

Who: Data centers that serve AI models, and the electricity grids that supply them

Assumptions: A bottom-up model of production serving from token throughput, node power and overhead, for frontier-scale models (more than 200B parameters) on H100 nodes; not a measurement.

Unlocked by the target for Energy per token

Serves: Affordable and clean energy, Climate action

Metrics

Energy per token headline1.0 orders of magnitude to go

Energy drawn per generated token in language model inference, as defined by the source. The number belongs to one of three measurement boundaries, and the conditions of the value say which: the chip or accelerator alone; the system, as power at the plug (MLPerf Power measures full system power with a SPEC-approved power analyzer read through the Power-Thermal Daemon, PTD, arXiv:2410.12032); or the facility, the system times the power usage effectiveness of the data center (PUE, ISO/IEC 30134-2:2026). Numbers from different boundaries are not comparable. Lower is better.
Energy per token: log scale, one tick per order of magnitude; better to the righttargetnow
Current (2026-08-01)0.72 J
Target0.07 J
Limit–
Conditions. GPU energy over the inference window divided by output tokens, as the source defines it; strongly dependent on model size, batch size, context and output length. The current value is for a 1B-parameter dense model, the favorable end of the range.
Why this target. One order of magnitude below the current value under the same conditions. Oviedo et al. (Joule 2026) estimate 8 to 20 times line-of-sight energy reductions across models, serving systems and hardware; 10 times sits inside that range and needs no new device physics.
Note. At 10 output tokens the same setup measures 7.46 J/token, so the number is only comparable under stated conditions. as_of is the approximate preprint month; the abstract gives no date.

Gaps

Moving weights and cache costs more than the arithmetic

criticalengineeringlayer: systemactive

Each generated token streams the model weights and the key-value cache through memory, so memory traffic, not arithmetic, tends to bind the energy per token. The best published electrical die-to-die link measures 0.65 pJ/bit in a prototype, an interface-only figure. Closing the gap means fewer bytes moved (quantization, cache compression, sparse attention) and cheaper bytes.

Held open by: Low-energy data movement

Decoding is bound by memory bandwidth, not compute

highengineeringlayer: deviceactive

Small batches leave the arithmetic units idle while the memory system is saturated. The measured token energy falls from 7.46 to 0.72 J/token as output length grows from 10 to 512 tokens at batch 16, because fixed costs are amortized; batching gains shrink as context grows (6.31x at 512 tokens of context against 1.17x at 4K for 10 output tokens).

Held open by: Low-energy data movement

CMOS logic has a ceiling of about two hundred times today's efficiency

mediumfundamental limitlayer: principleopen

Ho, Erdil and Besiroglu estimate a ceiling of 4.7e15 FP4 operations per joule for CMOS microprocessors, roughly two hundred times current microprocessors, from switching, interconnect capacitance and leakage. The atlas has no sourced current FLOP-per-joule value for deployed accelerators yet, so this gap has no metric of its own.

Held open by: Beyond-CMOS logic

Delivered power and cooling bound token output

mediumengineeringlayer: deploymentactive

At deployment scale the binding constraint can move from peak compute to delivered data-center power, cooling capacity and PUE. Embedded microfluidic cooling supports about 1e3 W/cm2; accelerator power density keeps rising.

Held open by: High-flux heat removal

Idle capacity and small batches waste energy

highengineeringlayer: deploymentpromising

Production energy per prompt includes idle machine capacity and data-center overhead, not only active accelerator power. Google reports a 33x reduction in energy per median Gemini Apps text prompt over one year from software efficiency and clean-energy procurement (the abstract states the combined effect on energy and does not separate the two). Further gains depend on batching, routing and model choice.

Dependencies

Requires

  • Beyond-CMOS logic Arithmetic in the accelerator is bounded by the switching energy of CMOS logic, which is estimated to allow only about two hundred times more efficiency than current microprocessors.
  • Low-energy data movement Token generation reads the model weights and the key-value cache from memory for every token, so memory and chip-to-chip traffic carries a large share of the energy. Need: Energy per bit moved between memory and processor at or below 0.1 pJ, so that streaming weights costs a few watts per terabyte per second.
  • High-flux heat removal Delivered power and cooling capacity bound how many accelerators fit in a rack and therefore the tokens produced per site.

Required by

  • AI for science Screening and agentic loops run many model calls per experiment, so the energy and cost per token set how far discovery loops can scale.
  • High-bandwidth brain-computer interface Decoding runs a neural network on the neural signal; running it in a wearable or implanted device needs low energy per inference.

Arrows point from a technology to what it requires. Select a node to open it.

Evidence

In The Alan Machine

Source TOML · Page on GitHub · Suggest a correction