Contact sales to request provisioned throughput
Provisioned throughput is currently unavailable for self-service. Contact our sales team to request a quote or scope a commitment.
Supported models
Provisioned throughput is available for the following models:
To request provisioned throughput for a model that isn’t listed, contact sales.
When to use provisioned throughput
Consider provisioned throughput when:- You’re moving high-volume production traffic off a proprietary model API and want to cut cost by switching to open models without giving up reliability.
- You need a defined SLA and committed capacity that serverless models can’t guarantee.
- Your traffic is steady and predictable enough to reserve capacity for, rather than competing for the shared serverless pool.
How PTUs work
A provisioned throughput unit (PTU) is the unit of capacity you buy. Each PTU is a fixed slice of guaranteed throughput for a model or model family, priced at a flat $0.05 per PTU per minute. You buy PTUs for the volume you expect to send, and as traffic flows through, it draws down your committed capacity. Input tokens, output tokens, and cached reads consume PTUs at different rates. Output tokens are more expensive to serve than input tokens, and cached reads are cheaper than fresh inputs. The exact conversion ratios are model-specific and defined in your contract. You don’t need to forecast a precise traffic mix. Whatever shape your traffic takes, it converts into a single normalized rate that draws down your PTU capacity: output-heavy or cache-light traffic consumes PTUs faster, while cache-heavy traffic consumes them more slowly. Traffic shape changes how quickly you consume PTUs, not the SLA. To estimate how many PTUs your workload needs, use the pricing calculator.SLA
Eligible traffic that fits within your purchased PTUs and the published product limits is backed by the following service level agreement (SLA):Eligible requests include all requests made within your purchased PTU capacity and the published product limits. Customer errors, invalid requests, authentication failures, client cancellations, traffic above your purchased capacity, and requests that violate our published abuse or product protection limits are not covered by the SLA.