AI cost lab · field instrument
What does it cost to run an AI query?
Every AI answer runs on rented GPUs. This estimates the marginal compute cost to serve a query and shows it next to what providers charge — a modeled estimate with wide, cited ranges, not a measured figure.
- Compute to serve 1M output tokens: ~$0.30–$10 (frontier central ~$1.90) — marginal compute only.
- Published API price (flagship output): ~$15–$30 / 1M tokens.
- The gap: for a typical chat query the price is ~10–20× the modeled compute cost; input-heavy (RAG) workloads widen it slightly (input tokens are priced above their share of compute), while the cheapest open-weight models run at or below cost for short-output chat.
- Inference cost trend: falling ~10× per year for fixed capability (Epoch AI).
That gap is not pure margin — the price covers fixed costs the per-query figure leaves out (training, R&D, free tiers). It also depends on the workload: input tokens are far cheaper to serve than output, so a provider's cost per query depends heavily on the input:output mix — even as the price-to-cost gap stays wide. Read the full explainer →
How the compute cost is modeled, with every cited figure, is in the methodology.
Estimates with wide ranges (~10×). Figures are modeled from public GPU rental rates, throughput benchmarks, and published API prices.