Evidence / oviedo2026energy
Energy use of AI inference, efficiency pathways, and test-time scaling
Oviedo, Felipe, Kazhamiaka, Fiodar, Choukse, Esha, Kim, Allen, Luers, Amy Lynd, Nakagawa, Melanie, Bianchini, Ricardo, Lavista Ferres, Juan
Joule, 2026
establishedmachine checkedarticle
Read the source doi:10.1016/j.joule.2026.102430
A bottom-up model with deployment assumptions, not a measurement: 0.31 Wh per query for frontier-scale models, not per token, so its finding carries no metric. Cited for the 8 to 20 times line-of-sight energy reduction the authors estimate.
Findings
Per query, not per token: bottom-up model estimate for frontier-scale models (>200B parameters) on H100 nodes (0.31 Wh, about 1128 J). No metric of the atlas is per query, so this finding carries no metric.
we estimate a median energy of 0.31 Wh/query
Cited by
Added 2026-10-04 by agent:claude-sonnet-5, checked 2026-10-04 · Source TOML · Report a problem