Skip to main content
Serverless models are the fastest way to run inference on Together. You call any supported model through a shared per-token API, with no provisioning, no replicas to size, and no minimum cost. Pay only for the tokens you process.

Models

If you’re not sure which model to use, see Recommended models for our picks by use case.

Chat

Image

Vision

Video

Audio

Embedding

Rerank

Moderation

Serverless and dedicated model inference support different sets of models. See the dedicated model inference catalog for details.
For rate limits and pricing, see the Serverless overview.

Chat models

Chat model examples

Image models

Use our Images endpoint for image models. Calling image models requires a positive credit balance.
Prices for models billed by image are estimates, not rates. These models pass through the provider’s own per-request charge, which varies with the resolution, quality, and other parameters you send. Models billed by megapixel use the formula below. For more details, see How image models bill.

Image model examples

  • Blinkshot.io: A realtime AI image playground built with Flux Schnell.
  • Logo creator: A logo generator that creates professional logos in seconds using Flux Pro 1.1.
  • PicMenu: A menu visualizer that takes a restaurant menu and generates nice images for each dish.

Per-megapixel cost formula

This applies only to models whose Unit column reads megapixel. It does not describe how the others are billed. For these models, cost depends on the size of the generated image in megapixels and on the number of steps used, if that number exceeds the default steps shown alongside the unit.
  • Default pricing: The listed price is the rate for one megapixel at the default number of steps.
  • Using more or fewer steps: Costs are adjusted based on the number of steps used only if you go above the default steps. If you use more steps, the cost increases proportionally using the formula below. If you use fewer steps, the cost does not decrease and is based on the default rate.
Here’s a formula to calculate cost: Cost = MP × Price per MP × (Steps ÷ Default Steps) Where:
  • MP = (Width × Height ÷ 1,000,000).
  • Price per MP = the value in the Price column.
  • Steps = The number of steps used for the image generation. This is only factored in if going above default steps.
Resolution dominates that formula. At $0.04 per megapixel, a 1024×1024 image (1.05 MP) costs about $0.042, while a 4096×4096 image (16.8 MP) costs about $0.67.

Gemini 3 Pro Image pricing

Gemini 3 Pro Image’s cost tracks the resolution you request. The catalog above quotes the 1K and 2K estimate of about $0.134 per image. At 4K it is about $0.24 per image. Supported dimensions: 1K: 1024×1024 (1:1), 1264×848 (3:2), 848×1264 (2:3), 1200×896 (4:3), 896×1200 (3:4), 928×1152 (4:5), 1152×928 (5:4), 768×1376 (9:16), 1376×768 (16:9), 1548×672 or 1584×672 (21:9). 2K: 2048×2048 (1:1), 2528×1696 (3:2), 1696×2528 (2:3), 2400×1792 (4:3), 1792×2400 (3:4), 1856×2304 (4:5), 2304×1856 (5:4), 1536×2752 (9:16), 2752×1536 (16:9), 3168×1344 (21:9). 4K: 4096×4096 (1:1), 5096×3392 or 5056×3392 (3:2), 3392×5096 or 3392×5056 (2:3), 4800×3584 (4:3), 3584×4800 (3:4), 3712×4608 (4:5), 4608×3712 (5:4), 3072×5504 (9:16), 5504×3072 (16:9), 6336×2688 (21:9).

Vision models

If you’re not sure which vision model to use, start with Qwen3.5 9B (Qwen/Qwen3.5-9B). For model-specific rate limits, see Rate limits.

Vision model examples

Video models

Audio models

Use our Audio endpoint for text-to-speech models. For speech-to-text models see Transcription and Translations. Audio model examples

Embedding models

There are currently no embedding models offered via serverless.

Rerank models

There are currently no rerank models offered via serverless. Rerank models like mixedbread-ai/mxbai-rerank-large-v2 are only available with dedicated model inference.

Moderation models

There are currently no moderation models offered via serverless.