Skip to main content
Serverless models are the fastest way to run inference on Together AI. Call any supported model instantly, with no minimum cost or provisioning latency. You pay only for the tokens, megapixels, or seconds of audio/video/speech you process. Serverless uses the same inference APIs as dedicated model inference, so you can prototype on serverless and move to reserved hardware later without changing your application code.

Get started

Quickstart

Make your first API request in a few minutes.

Available models

Browse the catalog of serverless models and their rates.

Recommended models

See our picks by use case if you’re not sure where to start.

Inference APIs

Call chat, image, audio, embedding, and more through one API.

Batch inference

Run asynchronous workloads at up to 50% lower cost.

Rate limits

Most users can expect to use serverless inference without encountering rate limits. Rate limiting may occasionally occur with high request volumes or large bursts of traffic. Response times can vary depending on the model and current demand. When demand is high, Together may limit requests to maintain performance goals. If you receive a 429 Too Many Requests response, reduce your request rate and spread out bursts. If you receive a 503 Service Unavailable response, wait briefly before retrying. Use exponential backoff for retries. See serverless rate limits for details. For committed throughput and reliability guarantees, contact sales about provisioned throughput. If you need direct control over reserved hardware, use a dedicated endpoint.

Region selection

Together AI routes serverless requests across its own infrastructure, and the serving region is not selectable. If you need requests to run in a specific region (for latency, residency, or compliance reasons), use a dedicated endpoint and pin its hardware to your target region.

Pricing

Serverless models bill based on usage, with no minimums and no provisioning cost. You pay per unit of work, with units determined by model type:
  • Chat, language, embedding, and rerank: Per input and output token.
  • Image generation: Per megapixel of output, or the provider’s own per-request charge. See How image models bill.
  • Video generation: Per second of output.
  • Speech-to-text and text-to-speech: Per second of audio.
Per-model rates are in the catalog tables, and on together.ai/pricing. If you don’t need real-time responses, some models are discounted up to 50% when run with batch workloads.

How image models bill

Image models bill one of two ways, and the Unit column in the image catalog shows which applies to a given model.
  • By megapixel: Models that Together AI serves directly bill on the number of megapixels you generate, scaled up when you exceed the model’s default step count. The listed price is a true rate, so doubling the pixels doubles the cost.
  • By image (estimated): Most image models are hosted by an external provider. Together passes your request through to the provider, and the cost is whatever that provider bills for the request. The catalog quotes an estimate for one image, but the actual cost varies with resolution, quality, how many reference images you send, and features like grounded search.
Not every provider states the output size its estimate is based on. Where the catalog shows a size, it comes from the provider’s own pricing note. Where it doesn’t, the provider quotes a single figure without naming an output size, and Together AI doesn’t infer one, because a model’s default request size isn’t necessarily the size the estimate was priced at. You can also see the basis in the console. Open the model’s page, and a per-megapixel model shows one price with a “per 1M pixels” tooltip, while a provider-charge model shows its price under “Pricing (estimated per image).” Before you commit to a high-resolution workload, generate one image at your target resolution and quality and check what it cost. A 4096×4096 image has roughly 16 times the pixels of a 1024×1024 one, and 64 times the pixels of a 512×512 one.

Cached input discounts

Select serverless chat models bill cached input tokens at a steep discount. Caching is:
  • Automatic: There is no header, parameter, or account toggle to enable it. Send the same prompt prefix again and any portion that’s still warm in the shared cache is billed at the cached rate.
  • Prefix-based: Only the longest matching prefix of your input counts as cached. Tokens after the first difference are billed at the standard input rate.
  • Best-effort and short-lived: The serverless cache is shared across the fleet and entries are evicted as traffic shifts, so cache hits aren’t guaranteed and there’s no configurable retention window. For predictable cache behavior, use a dedicated endpoint, where prompt caching is enabled by default and scoped to your own replicas.
  • Limited to supported models: Only models with a value in the Cached input pricing column on Chat models support cached input billing. Models without a cached price bill all input tokens at the standard rate.