- Home
- Pricing
Commercial
Pay for the resources you actually use
Start on demand with no commitment, reserve capacity when you have a deadline, or take the whole platform private. Every rate is shown before you launch, and the rate you see is the rate you are billed.
Build
Usage basedBilled per second of compute and per token of inference. No minimum term.
- On-demand GPU instances
- Serverless inference endpoints
- Model Foundry tuning jobs
- Dataset and object storage
- Console, CLI and REST API
- Community and email support
Scale
ReservedCapacity held for you, quoted against real availability.
- Reserved GPU capacity
- Dedicated inference replicas
- Higher concurrency and rate limits
- Priority technical support
- Named point of contact
- Agreed response commitments
Private
CustomA delivery project: design, build, commission, support.
- Single-tenant infrastructure
- On-premises or air-gapped installation
- Your identity provider
- Customer-controlled model catalogue
- Documented commissioning and acceptance
- Managed operations optional
On published GPU rates: this site does not list per-hour GPU prices, because availability and rates change with capacity. Current rates appear in the console before you launch an instance, and reserved capacity is quoted against confirmed availability. We would rather show you a price we can honour than a headline number that turns into a conversation.
What is metered
How each service is billed
Every line below appears separately in your usage breakdown, attributed to the project and API key that generated it.
| Service | Billing unit | Notes |
|---|---|---|
| GPU instances | Per second, from start to stop | Rate depends on GPU class and whether capacity is on-demand or reserved |
| Block storage | Per GB-month | Persistent volumes, billed while they exist, including while the instance is stopped |
| Object storage | Per GB-month, plus request volume | Datasets, artefacts and document corpora |
| Fine-tuning jobs | Per second of underlying GPU time | No separate platform surcharge on top of compute |
| Serverless inference | Per input and output token | Scales to zero; you pay nothing when idle |
| Dedicated replicas | Per replica-hour | Reserved capacity, no cold start, predictable latency |
| Knowledge Cloud | Per page ingested, plus storage and query volume | OCR is included in the ingestion rate |
| Synthetic Data Studio | Per record generated, plus underlying compute | Privacy and utility checks are included, not an add-on |
| Egress | Per GB out of the region | No charge for data moving between services inside your project |
Cost control
Budgets that stop things, not just warn you
The most common cloud bill surprise is an instance somebody forgot to stop. Set a ceiling per project, get a warning at eighty percent, and choose whether the platform blocks new workloads at the limit or simply tells you.
Usage is attributed per API key, so when consumption spikes you can see which client or which service caused it rather than reconstructing it from timestamps.
- Per-project spend ceilings with warning and hard-stop options
- Live accrued cost visible while a job is still running
- Per-key attribution for internal chargeback
- Idle instance alerts when a GPU has been running unused
- Daily and monthly exports in CSV for your finance system
- No egress charge between services inside the same project
FAQ
Pricing questions
Why are there no GPU prices on this page?
Because we will not publish a rate we cannot honour on the day you try to use it. GPU pricing moves with hardware availability and capacity commitments. Current rates are shown in the console before you launch anything, and reserved capacity is quoted against real availability. You will always see the price before you commit, and it will be the price you are billed.
How is compute metered?
Per second, from the moment an instance is running to the moment it stops. Storage is billed per gigabyte-month. Inference is billed per input and output token. Everything is broken down by project and by API key in the console.
Is there a minimum commitment?
Not on the Build tier. On-demand compute and serverless inference have no minimum term. Commitments only apply to reserved capacity, dedicated hardware or private deployment, where we are physically holding resources for you.
Can I set a spending limit?
Yes, per project. You choose the ceiling and what happens at it — a warning at eighty percent, and either a hard stop on new workloads or notification only. Budgets are the first thing we recommend configuring on a new project.
How does billing work for a private deployment?
Differently, because it is a delivery project rather than metered consumption. It is quoted as a fixed scope covering design, hardware, installation and commissioning, plus an ongoing support or managed-operations agreement. There is no usage meter on hardware you own.
Start on demand, commit only when it makes sense.
Create an account and see live rates in the console before you launch anything.