Get started
Quickstart
Make your first API request in a few minutes.
Available models
Browse the catalog of serverless models and their rates.
Recommended models
See our picks by use case if you’re not sure where to start.
Inference APIs
Call chat, image, audio, embedding, and more through one API.
Batch inference
Run asynchronous workloads at up to 50% lower cost.
Rate limits
Most users can expect to use serverless inference without encountering rate limits. Rate limiting may occasionally occur with high request volumes or large bursts of traffic. Response times can vary depending on the model and current demand. When demand is high, Together may limit requests to maintain performance goals. If you receive a429 Too Many Requests response, reduce your request rate and spread out bursts. If you receive a 503 Service Unavailable response, wait briefly before retrying. Use exponential backoff for retries. See serverless rate limits for details.
For committed throughput and reliability guarantees, contact sales about provisioned throughput. If you need direct control over reserved hardware, use a dedicated endpoint.
Region selection
Together AI routes serverless requests across its own infrastructure, and the serving region is not selectable. If you need requests to run in a specific region (for latency, residency, or compliance reasons), use a dedicated endpoint and pin its hardware to your target region.Pricing
Serverless models bill based on usage, with no minimums and no provisioning cost. You pay per unit of work, with units determined by model type:- Chat, language, embedding, and rerank: Per input and output token.
- Image generation: Per megapixel of output, or the provider’s own per-request charge. See How image models bill.
- Video generation: Per second of output.
- Speech-to-text and text-to-speech: Per second of audio.
How image models bill
Image models bill one of two ways, and the Unit column in the image catalog shows which applies to a given model.- By megapixel: Models that Together AI serves directly bill on the number of megapixels you generate, scaled up when you exceed the model’s default step count. The listed price is a true rate, so doubling the pixels doubles the cost.
- By image (estimated): Most image models are hosted by an external provider. Together passes your request through to the provider, and the cost is whatever that provider bills for the request. The catalog quotes an estimate for one image, but the actual cost varies with resolution, quality, how many reference images you send, and features like grounded search.
Cached input discounts
Select serverless chat models bill cached input tokens at a steep discount. Caching is:- Automatic: There is no header, parameter, or account toggle to enable it. Send the same prompt prefix again and any portion that’s still warm in the shared cache is billed at the cached rate.
- Prefix-based: Only the longest matching prefix of your input counts as cached. Tokens after the first difference are billed at the standard input rate.
- Best-effort and short-lived: The serverless cache is shared across the fleet and entries are evicted as traffic shifts, so cache hits aren’t guaranteed and there’s no configurable retention window. For predictable cache behavior, use a dedicated endpoint, where prompt caching is enabled by default and scoped to your own replicas.
- Limited to supported models: Only models with a value in the Cached input pricing column on Chat models support cached input billing. Models without a cached price bill all input tokens at the standard rate.