AI cost lab · field instrument

What does it cost to run an AI query?

Every AI answer runs on rented GPUs. This estimates the marginal compute cost to serve a query and shows it next to what providers charge — a modeled estimate with wide, cited ranges, not a measured figure.

A chat query ≈ 500 input + 500 output tokens — the rate is fixed; the token mix sets the totals below.
Compute cost
$0.0011
to compute one chat query
Model cost rate
$1.94/1M out · $0.32/1M in
marginal compute only — excludes training, R&D, free tiers
You pay
$0.02
published price for one chat query
API price rate
$25.00/1M out · $5.00/1M in
published list price — what providers charge per query
Price is about 13× the marginal cost for this chat query.
We assumed:
GPU rental ~$3.00/GPU-hr · throughput ~2000 tok/s per GPU · utilization 30% · overhead 1.4× · input ~6× cheaper to serve than output
Back-and-forth conversation (~1:1 in:out). Switch to Advanced to change these.

That gap is not pure margin — the price covers fixed costs the per-query figure leaves out (training, R&D, free tiers). It also depends on the workload: input tokens are far cheaper to serve than output, so a provider's cost per query depends heavily on the input:output mix — even as the price-to-cost gap stays wide. Read the full explainer →

How the compute cost is modeled, with every cited figure, is in the methodology.

Estimates with wide ranges (~10×). Figures are modeled from public GPU rental rates, throughput benchmarks, and published API prices.

← Back to the energy calculator