Demand and performance
Response times can vary depending on the model and current demand. Together aims to maintain high performance for serverless requests, including:- Tokens per second (TPS): The speed at which a model generates output tokens.
- Time to first token (TTFT): The time between sending a request and receiving the first output token.
429 Too Many Requests or 503 Service Unavailable response.
Serverless performance is best-effort. For committed throughput and reliability guarantees, see provisioned throughput.
Handle errors during high demand
The response code tells you how to adjust your requests:429 Too Many Requests: Reduce your request rate. Spread requests out over time and avoid sending large bursts. Use exponential backoff when retrying.503 Service Unavailable: Wait briefly, then retry. Use exponential backoff if the error continues.