Runpod Review 2026: GPU Cloud Pricing, Per-Second Math

Runpod rents GPUs by the second across three product shapes: dedicated Pods for training and long-running jobs, Serverless endpoints that scale to zero for inference, and multi-node Clusters up to 64 GPUs. It also runs a Public Endpoints marketplace of pre-deployed Whisper, FLUX, Qwen, and Sora models called over HTTP. The pricing page lists thousands of GPUs across 30+ regions, Community Cloud and Secure Cloud tiers side by side, and per-second granularity on every product. There is no permanent free tier. The entry point is a card on file and an RTX A5000 Pod at $0.27/hour (about 8 GB VRAM, 24 GB tier).

What you take on yourself is everything around the meter: batching, autoscaling, observability, and long-term capacity. The published per-hour numbers are honest. The bill is built from per-second time slices, idle storage, and the choice between Community Cloud and Secure Cloud. Numbers below come from runpod.io/pricing and the Runpod product documentation. They were captured from the live server-rendered pricing page (which itself carries an “Updated” stamp), in USD as displayed.

In short

Runpod is a GPU cloud that bills by the second across three product shapes — Pods, Serverless and Clusters — plus a Public Endpoints marketplace. The entry point is an RTX A5000 Pod at $0.27/hour. There is no permanent free tier; multi-month capacity and Secure Cloud premiums go to a sales conversation.

Runpod is the off-the-shelf option for developers and small AI teams that want per-second GPU access without signing an enterprise contract. Pods cover training and long-running workloads. Serverless covers inference with workers that scale to zero. Clusters cover multi-node training up to 64 GPUs. Public Endpoints cover pre-deployed models over HTTP. One account, one meter.

The verdict in short

Per-second GPU access without a sales call, and one account for training (Pods) and inference (Serverless). A multi-node Cluster option appears once training grows past a single node, and per-hour pricing stays visible the whole way. The catalog is broad, and the Public Endpoints marketplace spins up a Whisper, FLUX or Qwen endpoint in minutes without infrastructure work.

VendorRunpod (runpod.io)
CategoryGPU cloud — Pods, Serverless, Clusters, Public Endpoints
Free tierNone; the entry point is a credit card and a Pod
Entry price$0.27 per hour (RTX A5000, 24 GB, Community Cloud)
Billing unitPer second on Pods, Serverless and Clusters; per-hour is the headline rate
Top of the catalogB300 at $7.89/hour and H200 at $4.59/hour in the >80 GB tier
StorageContainer Disk $0.10/GB/month; Volume Disk $0.10 running, $0.20 idle
Best forPer-second GPU rental for training and bursty inference without a sales conversation
Prices checkedLive pricing page (which carries an “Updated” stamp)
Table: DeciderStack. Figures from the vendor pricing pages.

The bill in three numbers

Three prices define most of the Runpod conversation, and all three are per-second:

  • $0.27 per hour for an RTX A5000 Pod (24 GB VRAM, 25 GB RAM, 9 vCPUs). That is the published entry price and the on-ramp for a developer who wants to spin up a GPU without talking to sales.
  • $2.72 per hour for an A100 on Serverless (80 GB tier, “high throughput, cost-effective”). That is the published mid-market inference rate on the worker that fits most LLM workloads.
  • $7.89 per hour for a B300 Pod (288 GB HBM3e, 251 GB RAM, 32 vCPUs). That is the current top of the published Community Cloud catalog.

Per-second granularity is the structural feature. A 12-minute training run on an RTX A5000 costs about $0.054, not $0.27. An idle Serverless worker costs nothing. The same per-second meter applies on Clusters and on Serverless, so the time you run is the time you pay for.

alt="RunPod
Runpod’s pricing page opens on the product split – Pods, Serverless, Clusters and a Public Endpoints marketplace – rather than on GPU rates. Screenshot of runpod.io/pricing (source: runpod.io).
ProductHeadline rateWhat the headline buys
RTX A5000 Pod (Community Cloud)$0.27 / hour24 GB VRAM, per-second meter, consumer-grade supply
A100 Serverless worker$2.72 / hour80 GB VRAM, scale-to-zero inference, per-second while active
B300 Pod (Community Cloud)$7.89 / hour288 GB HBM3e, 32 vCPUs, top of the published catalog
H200 SXM Cluster$4.31 / hourMulti-node training up to 64 GPUs, sales contact for sustained capacity
Prices checked against the vendor pages.

Run the math on the load you actually have

The published per-hour numbers do not become a bill until you attach a workload. Here is what four representative workloads cost at the rates the pricing page shows. We have not run paid workloads on Runpod; these are the published rates multiplied by the time the workload needs.

WorkloadSKUTimeBill (Community Cloud)
Twelve-minute fine-tuneRTX A5000 Pod, 24 GB0.2 hr≈ $0.054
Single H100 SXM training dayH100 SXM, 80 GB24 hr$83.76
1,000 FLUX.1 [dev] images, 1 MP eachPublic Endpoints FLUX.1 [dev]per-request$20.00
1-hour inference on an H100 PRO workerH100 (PRO) Serverless1 hr active, scales to zero$4.79
Prices checked against the vendor pages.

The first row is where per-second billing matters most. The same workload on a per-hour vendor at the same headline rate costs $0.27. On Runpod it costs about a fifth of that because the meter is per-second and the job ran for 12 minutes. The second row is where Runpod’s published H100 SXM rate ($3.49/hour) sits below Lambda Labs’ published on-demand H100 SXM ($2.99–$3.99/hour depending on region). That is a real comparison, not a marketing one. The third row is the Public Endpoints math: per-request pricing replaces per-GPU-hour entirely. The fourth row is the inference production case, where Serverless PRO is the tier to pay for and Flex is the tier to ignore once cold-start latency matters.

alt="RunPod's
The Pods rate table underneath it, on the Community Cloud, from the H200 at $4.59 an hour down to the 48 GB cards at $1.09. Detail of runpod.io/pricing (Pods section; source: runpod.io).

Where Runpod stops being the cheapest GPU on the menu

Runpod is the right answer for a specific buyer. There are at least four buyers where it is not:

  • Vast.ai is cheaper on the headline number. Spot-market peer-to-peer GPU rental runs lower per-hour rates on commodity SKUs because supply comes from consumer desktops and small hosts. The trade-off is reliability: host churn is real, throughput variance is wider, and the published SLA is thinner. Runpod is the place to go when the headline rate is not the single variable.
  • AWS, GCP and Azure still own the compliance posture. FedRAMP, HIPAA BAA at scale, SOC 2 with dedicated tenancy, and per-region data residency are surfaced on the hyperscaler pricing and product pages. The Runpod pricing page does not publish the same compliance matrix. Enterprise buyers typically have to ask. Where the compliance posture is the headline — healthcare, public sector, financial services with audit obligations — the hyperscalers stay.
  • Modal, Replicate and Banana Dev ship the managed inference product. Auto-scaling, request queuing, batching across concurrent requests, and observability show up as product features there, not as code you write. Runpod’s Serverless is the worker. The surrounding orchestration is on the buyer. For a team that wants to deploy a model and forget about the worker, the managed-inference platforms stay.
  • Lambda Labs stays the published-reserved incumbent. Lambda publishes reserved pricing on H100 and A100 SKUs and runs a single-product cloud without a marketplace layer. Where multi-month capacity is the headline and a published rate matters for budgeting, Lambda Labs still wins.

Public Endpoints, priced per request instead of per GPU-hour

Public Endpoints is the part of Runpod that breaks the per-hour mental model entirely. The marketplace exposes audio, image, language and video models behind an HTTP API, billed per character, per request, per megapixel, or per million tokens as published. The pricing page lists 60+ models at capture time, and the list changes weekly. A representative slice from the published pricing:

ModelCategoryPublished price
pruna / Whisper V3 LargeAudio$0.05 per 1,000 characters
resembleai / Chatterbox TurboAudio$0.00 per 1,000 characters
bytedance / Seedream 4.0 EditImage$0.0270 per request
google / Nano Banana Pro EditImage$0.14 per request
black-forest-labs / FLUX.1 [dev]Image$0.02 per megapixel
qwen / Qwen3 32B AWQLanguage$10.00 per 1M tokens
ibm / IBM Granite 4.0 H SmallLanguage$1.00 per 1M tokens
bytedance / Seedance 1.0 ProVideo (5 s, 480p)$0.12 per request
Alibaba / Wan 2.6 T2VVideo (5 s)$0.50 per request
OpenAI / SORA 2 Pro I2VVideo (4 s)$1.20 per request
Prices checked against the vendor pages.

The trade-off versus Replicate and Modal is the metering model. Public Endpoints prices per request, per character, per megapixel, per million tokens or per second — there is no single metering unit. Translating that to a per-second inference comparison means converting every line, which is the hardest part of an apples-to-apples benchmark. The published per-request numbers are competitive; the comparison math is on the buyer. Cold-start latency and sustained throughput against an equivalent Replicate or Modal deployment depend on the inference workload and the GPU class — run a representative pilot before sizing a contract.

Reserved and PRO capacity: what “contact sales” actually hides

Three things on Runpod’s pricing page point at a sales conversation instead of a published number:

  • Reserved Clusters. Reserved Cluster deployments (10,000+ GPUs, custom configurations, SLA-backed uptime, discounted rates for 1-month to 12-month+ terms) are contact-sales only across every SKU. The pay-as-you-go H200 SXM rate is published at $4.31/hour and the A100 SXM rate at $1.79/hour; the discounted rate is not.
  • Reserved Serverless workers. PRO-tier workers (H100, L40S, L40, 5090, 4090, RTX 6000 Pro) carry an availability premium for stronger cold-start and queue behavior. The published PRO rates are visible per hour; the contracted rate is not.
  • Reserved Pods. The published per-hour Pod rates are pay-as-you-go. Sustained multi-month workloads are negotiable.

The asymmetry matters. Lambda Labs publishes at least directional reserved pricing on H100 and A100. CoreWeave publishes enterprise tier indicators. Runpod keeps list rates for sustained capacity off the public page, which means the buyer has to open a sales conversation to find out what reserved costs. Cluster interconnect and reserved-tier pricing are quoted sales-side — request a multi-month quote for an H100 SXM 8-GPU cluster for an apples-to-apples comparison against Lambda Labs reserved pricing.

The break-even point against Lambda Labs

The honest comparison against Lambda Labs is per-second on the same SKU, on the same region, on the same billing model. The published H100 SXM rates:

alt="Two
The full Pods ladder against the Serverless PRO rates for the same silicon. Chart by DeciderStack, built from our capture of runpod.io/pricing.
VendorH100 SXM published rateBilling unitReserved pricing
Runpod Pod (Community Cloud)$3.49 / hourper secondcontact sales
Runpod Cluster (pay-as-you-go)contact salesper secondcontact sales
Lambda Labs Cloud (1-click cluster)$2.99 / hourper hour, 1-click clusterspublished contracted tiers
Prices checked against the vendor pages.

The break-even is on workload shape, not headline rate. For a 12-minute job, per-second billing wins by a factor of about five on the same headline number. For a 24-hour sustained workload, per-hour billing matches per-second. For a multi-month contracted deployment, Lambda Labs’ published tiers win because the Runpod page leaves those tiers to a sales conversation.

What we did not price:

  • Secure Cloud premiums (not visible at the SKU level for every GPU)
  • Network egress on either vendor
  • The cost of Lambda Labs’ 1-click cluster tier on a workload that uses reserved capacity underneath

We flag those lines instead of estimating them.

The bills that surprise people

The headline rate is not the whole bill. Three lines on Runpod are easy to miss until they show up:

  • Volume Disk doubles in price when idle. $0.10/GB/month while a Pod is running, $0.20/GB/month while the Pod is stopped. The Pod cost goes to zero when stopped; the volume does not. Bake this into any model that spins Pods up and down.
  • Community Cloud vs Secure Cloud. Most GPUs are listed in both tiers. Community Cloud is the cheaper tier and uses peer-supplied capacity. Secure Cloud runs in Tier 3/Tier 4 data centers with isolated tenancy at a small premium. The premium is not always visible at the SKU level on the public page.
  • Serverless Flex vs PRO. Flex workers are the cheaper Serverless tier. PRO workers are dedicated-capacity with stronger cold-start and queue behavior. The published PRO rate on H100 is $4.79/hour. Flex on H100 is a separate number that does not show up on the headline comparison. Production traffic that cares about cold-start belongs on PRO; batch and dev belong on Flex.

FAQ

Is there a free Runpod tier?

No. Runpod has no permanent free tier; sign-up credit may apply to new accounts. The lowest published entry is an RTX A5000 Pod at $0.27/hour, billed per second while it runs.

How does Runpod billing work?

Per second on Pods, Serverless and Clusters. The per-hour number on the pricing page is the headline; the meter is per-second, so a 12-minute job costs 12 minutes.

Community Cloud or Secure Cloud?

Same SKUs, different physical isolation. Community Cloud is peer-supplied capacity and is the cheaper tier. Secure Cloud runs in Tier 3/Tier 4 data centers with isolated tenancy at a small premium that is not always visible per SKU.

Does an idle Volume Disk cost anything?

Yes — $0.10/GB/month while a Pod is running and $0.20/GB/month while the Pod is stopped. It is the easiest line in the storage model to miss.

Can I get reserved pricing?

Reserved Clusters and Reserved Serverless workers are contact-sales only. The Runpod page leaves list rates for sustained capacity to a sales conversation. Lambda Labs and CoreWeave publish directional pricing on the page instead.

Bottom line

Runpod in 2026 is the strongest off-the-shelf option for developers and small AI teams that want per-second GPU access without a sales conversation. One account covers both training (Pods) and inference (Serverless). A multi-node Cluster option kicks in when training scales beyond a single node. Pricing is transparent at the per-hour level. The GPU catalog is wide, and the Public Endpoints marketplace is a real productivity multiplier for teams that want a Whisper, FLUX, or Qwen endpoint in five minutes without managing infrastructure.

The calculus breaks down once the workload is bottlenecked on something other than GPU availability. Managed inference as a product (Modal, Replicate), the deepest hyperscaler-grade compliance posture (CoreWeave, AWS, GCP), the cheapest spot price on a commodity GPU (Vast.ai), or the simplest one-product cloud with published reserved pricing (Lambda Labs) — each serves a different buyer. The Runpod value is the per-second billing across three product shapes from one vendor, with no sales conversation required to start.

For a developer or small team buying their first GPU cloud, the right entry is an RTX A5000 Pod at $0.27 per hour. It bills per-second, only while the Pod is running, with Volume Disk as the persistent storage tier. Add Serverless PRO workers (H100 PRO at $4.79 / hour) once inference is in production and cold-start matters. Add a Cluster (H200 SXM at $4.31 / hour, A100 SXM at $1.79 / hour) once training spans multiple nodes. Reserved capacity is contact-sales only and worth a quote for any sustained multi-month workload. Public Endpoints are the right call when the model you need is already deployed — FLUX for image, Whisper for audio, Qwen for language — and you would rather pay per request than operate a Pod.

If the GPU cloud shape looks right, the Bright Data review covers the developer-platform angle on a different scale. The ScrapingBee review walks through the same per-request versus per-hour trade-off on a smaller bill. For teams comparing across GPU providers, the same published-versus-contact-sales pattern shows up in our Proton VPN review on a different product category.

How we tested this: every GPU rate, storage tier and endpoint price comes from runpod.io/pricing, captured from the live server-rendered page (the page itself is stamped “Updated September 2026”), and is quoted in USD as displayed. We have not run a paid Runpod workload, so figures that depend on cold-start latency, sustained throughput, queue behavior, Secure Cloud premiums or cluster interconnect performance are flagged where they appear in the text instead of being estimated.

Scoring criteria and our correction policy are documented on the methodology page.

DeciderStack Editorial Team — we sign up for the tools we cover, run the workload the vendor sells them for, and publish the bill. Who writes here · How we test · Editorial policy

This article contains affiliate links. If you buy through them we may earn a commission at no extra cost to you. Commission never changes our scoring or the order of a ranking.