Research Obsessions

9 Published AI Water Figures, Reconciled: 0.26 mL to 500 mL Explained

AI water use per prompt: 0.26 mL to 500 mL, reconciled. Where the "bottle of water" claim came from, and the researcher's 33x correction.

TL;DR

Published per-prompt water figures for AI range from 0.26 mL (Google's Gemini disclosure) to 500 mL (the original "Making AI Less Thirsty" paper on GPT-3), a roughly 2,000x spread that exists because each number measures a different scope, model, and year. The viral "bottle of water per prompt" claim traces to an uncited 140 Wh energy assumption in a 2024 Washington Post graphic, not the peer-reviewed paper it's usually attributed to, which used about 4 Wh. Shaolei Ren, the paper's own lead author, has since walked the water figure back to roughly 15 mL total (about 5 mL if you only count on-site cooling), a 33x correction from the number still circulating. This page reconciles all nine published figures with source, date, and scope, walks through six comparisons with the arithmetic shown, and explains Scope 1 (on-site) vs. Scope 2 (grid electricity) water, the framing that causes most of the confusion.

Last verified: August 1, 2026

I spent a week trying to answer a question that should have a simple answer and does not: how much water does one AI prompt use.

Google says 0.26 mL. Sam Altman says 0.32 mL. A widely shared claim says 500 mL, a full bottle. A Substack investigation, a UC Riverside professor, and a Washington Post graphic are all arguing about scope, methodology, and whose 2020-era energy estimate is still doing the rounding. None of that argument shows up in the headlines. So we went and read the primary sources ourselves.

What's Inside

Why does one source say AI uses 0.26 mL per prompt and another says 500 mL? Where did the "bottle of water" number come from, and why doesn't it match the peer-reviewed paper people cite for it? What separates Scope 1 from Scope 2 water, and why does that explain most of the confusion? How does an AI prompt compare to a Google search or a Mistral response, with the math shown? Is there an official water figure for AI video? Sources for every number, below.

The Number Everyone Cites Depends on Which Scope You Mean

I pulled nine of the most-cited per-prompt water figures into one table, sorted by scope and date. Here's what emerged.

#FigureModelSourceDateScopeWhy it differs
1500 mL per 10-50 responses (~10-50 mL/response)GPT-3 (175B)Li, Yang, Islam, Ren, arXiv 2304.03271Apr 2023 (ACM-published Jul 2025)Total (Scope 1 + 2)Built on a since-superseded 2020 OpenAI energy estimate for GPT-3, not GPT-4 or current models
2~2 liters per 10-50 queries (~40-200 mL/query)GPT-3 (175B)The Times, citing a 2024 Microsoft energy paperOct 2024TotalMicrosoft revised GPT-3's real energy draw up ~4x, and water scales with energy in this methodology
3~500 mL per 100-word emailGPT-4Washington Post graphic, reconstructed via The Energy Crisis SubstackSept 2024Total, built on an uncited assumptionRests on an unpublished 140 Wh estimate, ~35x higher than the peer-reviewed paper's 4 Wh figure for GPT-3
40.32 mL (0.000085 gal)ChatGPTSam AltmanJun 2025On-site only (inferred)No published methodology. OpenAI hasn't responded to requests for one
50.26 mL ("five drops")Gemini AppsGoogle CloudAug 2025Disputed. Google calls it comprehensive, critics read it as onsite-onlyGoogle's report doesn't state whether offsite generation water is included
62.2 mL/request (US average)GPT-3Li/Ren onsite breakout, via The Register2023 dataOn-site onlyThe correct apples-to-apples comparator for Google and Altman's figures. Still ~8x Google's number
745 mL per responseMistral Large 2Mistral AI lifecycle reportJul 2025Total lifecycle, third-party auditedOnly figure here with independent verification (Carbone 4, ADEME, peer review)
81.2 mL/queryGPT-4oJegham et al., arXiv 2505.09598May 2025On-site + offsite, excludes manufacturingTransparent formula, varies by prompt length and model
9~15 mL total, ~5 mL onsite"Average chatbot," GPT-4-classShaolei Ren, via Andy Masley's Substack and AI WeeklyMay-Jul 2026Total vs. onsite, both givenThe researcher's own current best estimate. A 33x reduction from the 500 mL figure

I traced each figure back to its original document, checked the publish date, and confirmed the scope the authors themselves claimed, which got mildly tedious around figure six. That last piece is what most competing pages skip.

The per-prompt water figures published for AI span roughly 0.26 mL to 500 mL, a spread of nearly 2,000x, because each number answers a different question about model, scope, and year. Once everything is sorted by scope and updated to current-generation efficiency, the real range for a single average query on 2025-2026 infrastructure narrows to about 0.26-15 mL total, or 0.12-5 mL if you count only the water evaporated inside the data center.

Where the "Bottle of Water per Prompt" Number Came From

This took the longest to untangle, and it's the reason this page exists.

The claim that one AI prompt burns a full bottle of water gets repeated constantly, usually attributed to a 2023 peer-reviewed paper, "Making AI Less Thirsty" by Li, Yang, Islam, and Ren.

Not quite.

That paper says 500 mL, but for 10 to 50 medium-length GPT-3 responses, not one, based on a 2020-era OpenAI energy estimate of roughly 4 Wh per request.

The "one bottle per single prompt" version people share came from somewhere else entirely. A September 2024 Washington Post graphic assumed a single GPT-4 query burns about 140 Wh and derived roughly 500 mL of water from that, according to an investigation by The Energy Crisis Substack. That 140 Wh figure "doesn't appear in any peer-reviewed paper," per the same investigation, and runs roughly 35 times higher than the number the actual academic paper used for a comparable request.

Two separate corrections got mashed into one viral fact. Untangled, they don't agree with each other, let alone with the bottle claim.

Shaolei Ren, the original paper's lead author, has since revised his own estimate downward for current-generation models. In direct correspondence reported by Andy Masley's Substack in May 2026, Ren stated plainly: "The 2024 estimate was time-specific, assumption-based, and should not be used to describe general AI/ChatGPT or today's optimized systems." His current estimate for an average GPT-4-class chatbot query sits at roughly 15 mL total, or about 5 mL counting only water used inside the data center itself, corroborated independently by AI Weekly and GPUSmith.

\( \dfrac{500 \text{ mL}}{15 \text{ mL}} \approx 33 \). Ren's own revised figure is about 33x lower than the bottle claim still attributed to him. The bottle is dead twice over: once for using an outdated model, once for absorbing an unsourced energy number that was never his.

Of the eight competitor pages audited while building this table, none combine the full reconciliation, the traced newspaper origin, and Ren's 2026 walk-back in one place. ToolixLab's statistics page comes closest with a scope-separated table. GPUSmith is the only one we found citing Ren's walk-back directly. Neither includes the Washington Post trace or a video-generation section.

Scope 1 vs. Scope 2 Water, Explained Without Jargon

Here's the single distinction that resolves most of the "who's lying" energy around this topic.

Direct water (Scope 1) is water used at the data center itself, mostly to cool chips. Cooling towers evaporate water to do this, while closed-loop systems with air chillers barely lose any once filled. This is almost certainly what Sam Altman meant when he called AI water concerns "totally fake" and having "no connection to reality," per The Atlantic.

Indirect water (Scope 2) is water consumed generating the electricity that data center draws, since thermoelectric power plants use cooling towers too. Here's the catch a closed-loop marketing pitch tends to skip: those systems use 10 to 65 percent more electricity than cooling-tower systems, according to Ren, cited in The Atlantic. Going water-free onsite pushes the water cost upstream to the power plant.

Meta's own 2024 disclosure makes the scale concrete: its indirect water consumption was 19 billion gallons, 23 times its direct water consumption, mostly attributable to data centers. Nationally, LBNL's 2024 report puts US data centers' direct water use at about 17.4 billion gallons in 2023, against an indirect figure near 800 billion liters derived from the same report's 176 TWh electricity number. \( \dfrac{800 \text{ billion L}}{66 \text{ billion L}} \approx 12 \). Indirect water runs roughly 12 times direct nationally, which is why "data centers barely use water" and "data centers use enormous water" can both cite real LBNL numbers while describing entirely different pictures.

Jonathan Koomey's framing is the one worth remembering: "If you use water to make your cooling more efficient on-site, you will use less electricity. It's not a simple matter of water use bad on-site." The honest answer to "how much water does a data center use" depends on local climate, water supply, grid fuel mix, and cooling design. No single number is honest without those four attached.

Six Comparisons, With the Arithmetic Shown

Numbers without something to measure against don't mean much, so here's the math, not just the conclusion. I ran every division myself against the sourced inputs rather than trust a secondhand summary.

A Gemini prompt vs. a 2009 Google search. A 2009 search used about 0.3 Wh. Gemini's comprehensive figure is 0.24 Wh. \( \dfrac{0.30}{0.24} \approx 1.25 \). A Gemini prompt today uses roughly 80 percent of the energy a single Google search used in 2009, per Google's own comparison.

A Gemini prompt vs. watching TV. Google states its 0.24 Wh figure equals under 9 seconds of TV. [UNCERTAIN] Google doesn't disclose the TV wattage behind that comparison, so treat it as Google's own framing, not something independently recomputed.

Gemini prompts vs. one Mistral response. Mistral's own comparison for its 45 mL Le Chat figure is the water needed to grow a small pink radish, which is a genuinely strange unit of measurement to standardize on. Against Google's 0.26 mL: \( \dfrac{45}{0.26} \approx 173 \). It takes about 173 Gemini-style prompts to match one Mistral response's water footprint, mostly because Mistral's number is lifecycle-inclusive (training amortization plus inference) while Google's tracks closer to the marginal query alone.

Today's queries vs. the old bottle claim. Using Ren's 2026 revised total of 15 mL against the 500 mL bottle figure: \( \dfrac{500}{15} \approx 33 \). It takes about 33 of today's average chatbot queries to use as much water as the original viral claim attributed to one.

US data centers' indirect vs. direct water (2023). Direct: 17.4 billion gallons (~66 billion liters). Indirect: roughly 800 billion liters, per LBNL's report as cited by ToolixLab. \( \dfrac{800}{66} \approx 12 \). Indirect water runs about 12x direct at the national level, the same scope mismatch driving most public disagreement.

An 8-second Sora video vs. a Gemini text prompt. Sora 2.0 Pro's mean energy for an 8-second 720p clip is 418.5 Wh, against Gemini's 0.24 Wh, per arXiv 2607.04553. \( \dfrac{418.5}{0.24} \approx 1{,}744 \). An 8-second video takes roughly 1,744 times the energy of a single text prompt, consistent with the paper's broader cited range of 1,625x to 1,990x depending on the exact model and resolution measured.

What About Video and Image Generation?

Short answer: no AI company has published an official per-video water figure, and every number circulating for AI video is a derivation, not a disclosure.

Image generation is comparatively well studied. Hugging Face and Carnegie Mellon researchers estimate about 2.9 Wh per image on average, up to 11.5 Wh for the least efficient models, roughly comparable to fully charging a smartphone.

Video is the least standardized measurement here. The most rigorous published source we found, arXiv 2607.04553, estimates an 8-second 720p Sora 2.0 Pro clip at a mean of 418.5 Wh, rising to 1,313 Wh for a 12-second 1080p clip. Sasha Luccioni's Hugging Face estimate puts a Sora video around 90 Wh, roughly 30x a typical image.

[UNCERTAIN] No AI company, including OpenAI, has published an official per-video water figure. Every water number attached to AI video in circulation, including a "4+ liters per 10-second clip" estimate from one independent analysis, is a third-party extrapolation from energy estimates multiplied by an assumed water-use-effectiveness factor. Treat any specific per-video water figure as someone's derivation, not a disclosure, until that changes.

Frequently Asked Questions

How much water does one ChatGPT query use?

I couldn't find an OpenAI-published figure with a stated methodology. Sam Altman's cited number is 0.32 mL, but OpenAI hasn't published how that was calculated, and The Verge reports it didn't respond to requests for one. The more defensible range, drawing on independent academic work and Ren's 2026 revision, sits closer to 0.12-15 mL depending on scope.

Is the "AI uses a bottle of water per prompt" claim true?

No, and tracing where it came from was the most satisfying part of building this page. It's not what the peer-reviewed paper found. It came from a 2024 Washington Post graphic built on an uncited 140 Wh assumption, and the paper's own lead author has since said the underlying estimate "should not be used to describe general AI/ChatGPT or today's optimized systems."

What's the difference between Scope 1 and Scope 2 water use for AI?

Scope 1 is water evaporated on-site to cool the data center. Scope 2 is water used generating the electricity that facility draws from the grid. I spent most of my research time on this single distinction, because it's why closed-loop data centers can report near-zero direct water while using more electricity, and therefore more indirect water, than the cooling-tower systems they replace.

Does AI video generation use more water than text?

Almost certainly, based on the energy difference alone. An 8-second Sora clip runs roughly 1,744x the energy of a single Gemini text prompt by the most rigorous published estimate found. But no company has disclosed an official per-video water number, so any liters-per-video figure you encounter is a derivation, not a measured fact.

Why do water-per-prompt estimates vary by almost 2,000x?

"How much water does AI use" isn't one question. It bundles which model, what scope, what prompt length, and what year's hardware efficiency, plus the occasional modeled projection dressed up as a measurement. Sort the nine published figures by those four variables, as we did above, and the spread for current-generation models collapses from 2,000x down to about 60x.

If you've found a company disclosure we missed, or a number here that's since been updated, I'd genuinely like to know. This entire page exists because the version of this question that goes viral is wrong twice over. Tracing a viral number back to its actual source is the most transferable habit in critical AI literacy for product thinking, and this page is what it looks like in practice. That's worth fixing together 🤗