On-demand, reserved, spot, and committed-use pricing in 2026

"Reserved vs on-demand" describes more than one contract shape. The rate comparison only works when you distinguish hours you consume from capacity you promise to pay for, and when you keep interruptible spot capacity out of the reserved comparison.

  • On-demand. You rent a GPU by the hour and pay for consumed capacity without a one-year minimum. It is the flexible baseline used in this post's reference table; the current API exposes both the instance rate and a normalized per_gpu_hourly value.
  • Reserved, 1-year. You commit to a GPU or capacity block for a 12-month term at a provider-specific rate. Some contracts bill a fixed capacity block whether it runs or not; others apply a discounted rate to committed usage. The contract's billing rule determines the break-even formula.
  • Spot or preemptible. You accept interruption in exchange for a lower, variable rate. Spot is useful for checkpointed training, elastic batch, and experiments, but it is not a substitute for a verified 1-year reserved row.
  • Committed-use or spend plans. AWS Savings Plans, Google Cloud committed-use discounts, and Azure reservations typically commit spend, instance family, or capacity under provider-specific rules. Treat them as a separate contract until the scope, region, and portability match the GPU reservation being modeled.

The key procurement question is not "what is the advertised discount?" It is "what share of the contracted capacity will be consumed by an SLA-bound workload?" A 1-year commitment can lower the rate for every used hour and still cost more in total if it bills idle hours. That makes utilization a billable-hours measure over the term, not an instantaneous GPU busy percentage.

Reserve the stable floor, not the peak

Reserve the baseline that must be available for predictable production inference, a scheduled training window, or a mature batch queue. Keep burst capacity on on-demand or spot until the workload proves it can absorb a longer commitment. The Reserved Instance Advisor helps size that floor by GPU, provider, and term.

Break-even utilization and payback math

Two related calculations answer two different procurement questions. A fixed-capacity commitment asks whether the annual bill is lower than paying only for consumed on-demand hours. An upfront-premium payback calculation asks how quickly a one-time fee is recovered when reserved hours are otherwise consumed.

Fixed-capacity break-even

Let O be the on-demand $/hr, R be the reserved $/hr, P be any upfront premium, and u be the share of full-time capacity consumed. A one-year capacity block has 8,760 hours (12 × 730), while on-demand consumption is u × 8,760 hours.

  • reserved_cost = P + R × 8,760
  • on_demand_cost = O × u × 8,760
  • u_break_even = (P + R × 8,760) ÷ (O × 8,760)

Reserved wins only when the first cost is lower than the second. If P = 0, the utilization threshold simplifies to R ÷ O; an upfront premium raises the required utilization.

Upfront-premium payback

When a contract charges a one-time premium but the reserved rate applies to hours you actually consume, use 730 hours per month and the expected utilization u:

payback_months = P ÷ ((O − R) × u × 730)

This is not the same as a fixed-capacity annual comparison. It measures how long the hourly discount takes to recover P; if the contract bills unused capacity, include the full 8,760 × R commitment instead.

Worked example — using the current CoreWeave H100 on-demand row and a clearly labeled illustrative reservation input:

  • Capacity clock: full-time capacity is 730 hr/month and 8,760 hr/year. At 60% billable utilization, consumed hours are 5,256/year, not 730 hours for the year.
  • On-demand baseline: the August 28, 2026 API row is $2.23/hr per H100 on CoreWeave, so 60% utilization costs 5,256 × $2.23 = $11,720.88/year.
  • Fixed-capacity test: a reserved quote of R with upfront premium P wins at that utilization only if P + 8,760R < $11,720.88. With no upfront premium, the reserved rate must be at or below about $1.338/hr to break even with that 60%-utilized on-demand bill. The table leaves the actual reserved quote pending because no verified 12m row is available in the live grid snapshot.
  • Illustrative payback only: if a contract hypothetically used R = $1.80/hr and P = $500, monthly savings at 60% utilization would be ($2.23 − $1.80) × 0.60 × 730 = $188.34, so payback would be 2.65 months. Those inputs illustrate the formula; they are not provider quotes and are not used in the reference table.
Utilization nuance: a dashboard showing 60% GPU occupancy is not automatically 60% procurement utilization. Use billable hour consumption over the commitment term, account for weekends and planned idle periods, and model the minimum capacity your SLA requires. A reservation sized to peak can fail even when the GPU is busy during the hours it is running.

The verified-data convention

The reference table combines two current data paths, each with a different job:

  • 1-year reserved. The 12m provider-native rows from /gpu-forward-pricing's /api/forward-pricing/grid, shown only when verified_at is present.
  • On-demand. The matching GPU/provider rows from /api/gpu-pricing?pricing_type=on-demand, normalized to per_gpu_hourly for multi-GPU instances.
  • Data pending. A missing or unverified value remains data pending. We do not substitute spot pricing, interpolate from another GPU, or invent a reserved quote; savings are calculated only when both comparable numeric values exist.

As of 2026-08-28, the live grid returned no verified 12m provider-native rows for H100, A100, or B200. The pending cells below are therefore an intentional data-quality result, not a zero discount.

Workload categories that justify a commit

The right reservation target is a workload with a measurable floor, a known operating window, or a business penalty for losing capacity. Use the following rubric before signing a 12-month term:

  • Predictable production inference — YES, reserve the stable floor. Customer-facing chat, RAG, ranking, and API serving with a repeatable daily traffic curve are the clearest candidates. Reserve the P50-P75 capacity that must meet the SLA, then keep burst traffic on on-demand or spot. Do not reserve the P99 peak unless the SLA makes that peak non-interruptible.
  • Scheduled multi-week training — YES, reserve the known run window. A 4-12 week training or fine-tuning run with a committed start date, end date, and GPU shape can justify a term-sized reservation. Size it to the planned run and a conservative schedule, not to the assumption that a future job will fill the remaining months.
  • Mature batch and evaluation queues — MAYBE, measure first. Repeated evaluation, rendering, or batch-inference jobs are reserve candidates when queue depth and execution hours are stable. Run an on-demand pilot for two or three billing cycles, calculate billable utilization, and compare it with the fixed-capacity threshold before converting the baseline.
  • Mixed production plus elastic workloads — YES, reserve the floor and burst. A fleet with a stable SLA-bound baseline and interruptible overflow should use two lanes: reserved capacity for the baseline and on-demand or spot for peaks, retries, and temporary experiments. This avoids paying a reservation rate for capacity that only exists during a few high-load hours.
  • Experimentation, benchmarking, and open-ended research — NO, stay flexible. Short-horizon work with unknown model choice, utilization, or funding horizon should stay on on-demand and spot. The flexibility premium is cheaper than a year of idle committed capacity; reserve only after the work stabilizes into a measurable floor.
  • Three details can change the answer even when utilization looks healthy: GPU generation affects migration risk, region affects the comparable rate and availability, and capacity shape affects whether an 8-GPU reservation can be reused by a one-GPU workload. Compare like-for-like GPU, provider, region, and per-GPU basis before treating a percentage as savings.

    Reference table: on-demand vs 1-year reserved $/hr for H100, A100, B200 across 4 providers

    This August 28, 2026 snapshot compares one like-for-like on-demand row for each GPU/provider pair across AWS, CoreWeave, RunPod, and Google Cloud. The reserved column reads from the 12m provider-native rows on /gpu-forward-pricing; because the current grid has no verified 12m row for H100, A100-80GB, or B200, every reserved cell is intentionally marked data pending.

    The 12-month savings % formula is (on-demand − reserved) ÷ on-demand. It is calculated only when both rates are numeric and comparable. A pending value is not zero, a spot quote, or an estimate.

    GPU Provider On-demand $/hr 1-yr reserved $/hr 12-month savings % Source
    H100 (Hopper, 80GB HBM3) AWS $12.29 data pending data pending Forward grid
    on-demand API · us-east-1 · p5.48xlarge ÷ 8
    H100 (Hopper, 80GB HBM3) CoreWeave $2.23 data pending data pending Forward grid
    on-demand API · US · H100 SXM5
    H100 (Hopper, 80GB HBM3) RunPod $2.49 data pending data pending Forward grid
    on-demand API · US/EU · H100 SXM 80GB
    H100 (Hopper, 80GB HBM3) Google Cloud $3.90 data pending data pending Forward grid
    on-demand API · us-central1 · a3-highgpu-1g
    A100-80GB (Ampere) AWS $5.1207 data pending data pending Forward grid
    on-demand API · us-east-1 · p4de.24xlarge ÷ 8
    A100-80GB (Ampere) CoreWeave data pending data pending data pending Forward grid
    on-demand API · no A100-80GB row
    A100-80GB (Ampere) RunPod data pending data pending data pending Forward grid
    on-demand API · no A100-80GB row
    A100-80GB (Ampere) Google Cloud $3.6739 data pending data pending Forward grid
    on-demand API · us-central1 · a2-ultragpu-8g ÷ 8
    B200 (Blackwell, 192GB) AWS $6.90 data pending data pending Forward grid
    on-demand API · us-east-1 · p6.48xlarge ÷ 8
    B200 (Blackwell, 192GB) CoreWeave $4.49 data pending data pending Forward grid
    on-demand API · US · HGX_B200_x1
    B200 (Blackwell, 192GB) RunPod $5.98 data pending data pending Forward grid
    on-demand API · US/EU · B200 192GB
    B200 (Blackwell, 192GB) Google Cloud $6.60 data pending data pending Forward grid
    on-demand API · us-central1 · a4-highgpu-8g ÷ 8
    Methodology and as-of: rates were read on 2026-08-28 from the current API payloads. Each numeric on-demand value is the lowest active per-GPU row in the listed region for that provider/GPU comparison; multi-GPU AWS and Google Cloud instances are divided by their GPU count. The source uses GCP as the provider label for the H100 row, normalized here to Google Cloud. The 12m provider-native tenor means a one-year forward/reservation row, not 730 hours per year and not a spot price. Because no verified 12m rows are currently returned for these three GPUs, pending cells stay pending rather than being interpolated or filled with spot pricing. Recheck /gpu-forward-pricing before signing a term sheet.

    Three things to take from the table:

    • Provider choice moves the on-demand baseline. The selected H100 rows run from $2.23/hr on CoreWeave to $12.29/hr per GPU on AWS; matching A100-80GB rows run from $3.6739/hr on Google Cloud to $5.1207/hr on AWS; B200 runs from $4.49/hr to $6.90/hr. Those are on-demand spreads, not reserved savings.
    • Pending is a procurement signal. A missing verified 12m row means the reservation spread is unknown. It does not mean the provider has no discount, that the discount is zero, or that an adjacent GPU's rate can be reused.
    • Commit only after the pair is comparable. Once a provider-native 12m row appears, compare it with the same GPU, provider, region, and per-GPU on-demand basis, then apply the savings formula and your measured utilization threshold.

    2026 procurement takeaway

    The August 28, 2026 data snapshot supports a disciplined sequence: measure the workload floor, match the contract shape to that floor, then compare a verified reserved row with the same on-demand basis. A headline discount is not enough to justify a 12-month commitment.

  • Reserve predictable production inference. Put the P50-P75 SLA-bound capacity on a one-year term only after billable utilization clears the fixed-capacity threshold. Keep unpredictable peaks on on-demand or spot rather than paying a reservation rate for the P99.
  • Reserve scheduled multi-week training carefully. A known 4-12 week run can justify a term-sized commitment, but size it to the confirmed run window and model the cost if the schedule slips. Do not assume an unbooked future training job will consume the remaining term.
  • Keep experimentation and open-ended research flexible. Unknown model choice, utilization, or funding horizon makes on-demand and spot more valuable than a nominal reserved discount. Convert only the stable floor after two or three observed billing cycles.
  • Use reserved plus spot for mixed workloads. Protect the non-interruptible baseline with reserved capacity and send checkpointable overflow to spot. Recompute the blend when traffic, region, or GPU generation changes.
  • 2026 rule of thumb

    If the provider has a verified 12m row and the fixed-capacity equation clears your measured utilization, reserve the stable floor. If the row is data pending, keep the decision provisional: use /gpu-forward-pricing for the next verification, use /calculator for the on-demand baseline, and run the Reserved Instance Advisor before signing a term sheet.

    Frequently asked questions

    What utilization rate makes a 1-year reserved GPU commitment break even?

    For a fixed-capacity commitment, break-even utilization is (upfront premium + reserved rate × 8,760 hours) ÷ (on-demand rate × 8,760 hours). With no upfront premium, a reserved rate of $1.338/hr breaks even against a $2.23/hr on-demand rate at 60% utilization; the exact threshold changes with the quoted reserved rate, upfront premium, region, and contract terms. Use consumed billable hours, not instantaneous GPU busy time, for u.

    What happens if my workload drops during a 1-year reserved term?

    You still owe the fixed-capacity commitment or the contracted minimum, so unused hours turn into an overage against on-demand. The practical mitigation is to reserve only the stable floor, start with 1 year instead of 3 years, and validate demand for two or three billing cycles before adding capacity. Check the provider's transfer, exchange, or resale rules; they vary by contract and are not assumed in the break-even math.

    Should I combine reserved capacity with spot, or use spot only?

    Combine them when part of the workload is SLA-bound and part is interruptible: reserve the stable baseline and burst above it on spot or preemptible capacity. Use spot only for fault-tolerant experiments, elastic batch, and research that can checkpoint or retry. A mixed plan usually costs less than reserving peak while protecting the capacity that cannot be preempted.

    Run the break-even math on your workload

    GridStackHub's calculator, optimizer, and reserved advisor all read from the same provider-native dataset this post uses for its on-demand baseline and pending 12m checks.