Skip to main content
Most users can expect to use Together AI serverless inference without encountering rate limits. Rate limiting may occasionally occur with high request volumes or large bursts of traffic.

Demand and performance

Response times can vary depending on the model and current demand. Together aims to maintain high performance for serverless requests, including:
  • Tokens per second (TPS): The speed at which a model generates output tokens.
  • Time to first token (TTFT): The time between sending a request and receiving the first output token.
When demand is high, Together may limit requests to maintain these performance goals. You may receive a 429 Too Many Requests or 503 Service Unavailable response. Serverless performance is best-effort. For committed throughput and reliability guarantees, see provisioned throughput.

Handle errors during high demand

The response code tells you how to adjust your requests:
  • 429 Too Many Requests: Reduce your request rate. Spread requests out over time and avoid sending large bursts. Use exponential backoff when retrying.
  • 503 Service Unavailable: Wait briefly, then retry. Use exponential backoff if the error continues.
Limit retries to fit your application’s latency needs. For workloads that need committed throughput and reliability, consider provisioned throughput.

Get throughput guarantees

If your workload needs committed throughput and reliability guarantees, provisioned throughput reserves capacity for a selected model or model family with a defined service level agreement (SLA). Contact sales to discuss your workload and capacity requirements.