Improvements
New models available for fine-tuning
You can now fine-tune the following models:Qwen/Qwen3.6-27B.
New releasesImprovements
Python SDK realtime transcription
The Together Python SDK now includesclient.beta.realtime.transcription() for streaming speech-to-text over WebSocket. Install with pip install "together[realtime]". The session reconnects with audio replay on transient drops, exposes normalized events such as TranscriptDelta and TranscriptCompleted, and supports application-level failover across endpoints with RealtimeConnectionError (code="no_healthy_workers") and session.pending_audio().See Streaming transcription.Fine-tuning comparison metrics filtering
The fine-tuning comparison view now includes the same Metrics filtering control as the single-job Metrics tab. Adjust Sampling rate and Step range, then click Apply to re-fetch metrics for every selected job with matching filters.See View metrics in the dashboard.New releasesImprovements
Dedicated containers OpenAI-compatible endpoints
Dedicated container inference now supports HTTP server mode for synchronous, OpenAI-compatible endpoints. Run your worker without the--queue flag, wire an OpenAI route with Sprocket, and call it with the OpenAI SDK or plain HTTP, with no Together-specific request shapes.See Serve an OpenAI-compatible endpoint.Fine-tuning output object names
Fine-tune retrieve and list-checkpoints responses now include qualified Together model registry names alongside object IDs. OnGET /fine-tunes/{id}, use model_object_name and adapter_object_name (LoRA jobs) for the final artifacts in <project_slug>/<model_name> form. On GET /fine-tunes/{id}/checkpoints, each entry adds object_name with the same naming pattern (including -<step> or -adapter suffixes). Names are resolved on retrieve only, not on list jobs. If the project slug cannot be resolved, the name field falls back to the object ID.On a completed job in the fine-tuning jobs dashboard, Output model shows the same qualified model_object_name and links to the registry model page.See Model registry object IDs.Fine-tuning checkpoint CLI output
tg fine-tuning list-checkpoints now displays registry artifact IDs in the table output. The Registry Artifact column shows object_id@object_revision_id, and a copyable Registry artifacts block prints below the table.See List checkpoints.Fine-tuning preview CLI command
tg fine-tuning preview samples rows from an uploaded JSONL training file and shows how a base model tokenizes them before you start a job. The table output highlights masked tokens, trained spans, and truncation; pass --json for the full response.See Preview.Worker engine cache metrics
Dedicated endpoint Prometheus metrics now exposeworker_engine_kv_cache_utilization and worker_engine_cache_hit_rate gauges. Use them to track KV-cache pressure and prefix-reuse efficiency on deployment workers.See Monitor endpoints and deployments.Improvements
Cluster remediation mode override
tg beta clusters remediations approve now accepts a --mode flag to override the recommended repair action when you approve a node remediation. Choose reboot, quick reprovision, migrate to new host, or remove without changing the recommendation in the console first.See Node repair and the clusters CLI reference.Improvements
GPU cluster add-on CLI flags
tg beta clusters create and tg beta clusters update now expose flags for Headlamp and Slurm Web cluster add-ons.- Create: Pass
--headlamp-addonor--slurm-web-addonto enable an add-on at cluster creation. - Update: Pass
--headlamp/--no-headlampor--slurm-web/--no-slurm-webto toggle add-ons on an existing cluster.
MoE expert-LoRA target modules
Fine-tuning jobs on supported mixture-of-experts models can now target expert feed-forward layers instead of attention projections. Setlora_trainable_modules to w_up,w_gate,w_down to train a compact shared-factor adapter across experts, useful when adapting domain knowledge rather than attention patterns.See Target MoE expert layers.Jig environment variable collision validation
jig deploy now rejects configs where the same name appears in both [tool.jig.deploy.environment_variables] and secrets. The error lists each colliding name so you can remove the duplicate or unset the secret before redeploying.See Secrets.New models available for fine-tuning
You can now fine-tune the following models:zai-org/GLM-5.1.zai-org/GLM-5.
New releasesNew modelsImprovements
Fine-tuning metrics dashboard
Fine-tuning jobs now have a Metrics tab in the web dashboard that charts training progress. Metrics stream in live while a job runs and stay available after it completes.What’s new:- Live training curves: Watch training loss, learning rate, and gradient norm update as the job trains, without leaving the dashboard.
- Preference-tuning metrics: DPO and other preference jobs add reward accuracy, reward margin, chosen and rejected rewards and log-probabilities, and approximate KL.
- Interactive charts: Zoom and pan across all charts on a shared step axis, switch the x-axis between step and time, toggle linear or logarithmic scales, and sync a hover crosshair across every metric.
- Compare runs: Overlay metrics from multiple jobs to compare fine-tuning runs side by side.
New dedicated endpoint models
The following models are now available for deployment on dedicated endpoints:deepseek-ai/DeepSeek-V4-Flash.
Model revision validation status
Model and adapter revision APIs now returnvalidationStatus, lastValidatedAt, and validationErrors on each revision. Poll validation before you pin a revision in a deployment — a revision must reach REVISION_VALIDATION_STATUS_SUCCESS before it can deploy when referenced explicitly.See Check revision validation.Improvements
Privacy and training opt-in settings
Training opt-in and other privacy toggles now live under Privacy in Organization Settings. Organization admins control data-sharing settings for all traffic sent under the organization’s API keys.See Privacy and security.Dedicated endpoint deployment listing
Endpoint get and list responses now include at most the 10 newest deployment summaries per endpoint. Use the deployments list API to retrieve every deployment on an endpoint.See Manage endpoints and deployments.New releasesNew modelsDeprecationsImprovements
Dedicated model inference
Dedicated model inference (DMI) serves a model on reserved GPUs, giving you higher throughput, lower latency, and predictable performance with no hard rate limits. It uses the same inference APIs as serverless, so you can prototype on serverless and deploy on dedicated hardware without changing your application code.What’s new:- Deploy in one command:
tg beta endpoints deploycreates an endpoint, attaches a deployment, and routes all traffic to it. - New resource model: Compose models, configs, endpoints, deployments, and traffic splits to control how your model is served. See Concepts.
- A/B tests and shadow experiments: Split live traffic across variants, or mirror traffic to a new deployment without serving its responses.
- Autoscaling and serving controls: Scale on request or GPU metrics, scale to zero on idle, and tune serving behavior per deployment. See Configure autoscaling.
If you’re already using dedicated endpoints, you can migrate to the v2 API by following the migration guide.
New serverless models
The following models are now available on serverless:thinkingmachines/Inkling: 552.8B parameters, 524,288 context length, NVFP4 quantization. Pricing: $1.00 input / $4.05 output / $0.17 cached input (per 1M tokens).
Image inputs for evaluations
You can now evaluate vision-capable models on image datasets. Add animage_data_urls column to your evaluation dataset — a base64 image data URL, or a list of them — and the images are attached to the model and judge requests alongside the text prompt.See Prepare your dataset for details.Fine-tuning artifact registry IDs
Fine-tune job and checkpoint responses now include Together model registry IDs for completed artifacts. Usemodel_object_id and model_object_revision_id on completed jobs, adapter_object_id and adapter_object_revision_id on LoRA jobs, and object_id and object_revision_id on list checkpoints entries to reference weights in the model registry. The Together Python SDK (2.24+) and Together TypeScript SDK expose these fields on retrieve and list-checkpoints responses.Together Python SDK 2.24
Together Python SDK 2.24 is available. OIDC cluster SSH (tg beta clusters ssh) no longer requires a Together API key, and the bastion-to-target SSH hop skips host key prompts for ephemeral cluster nodes. See SSH into a cluster.New releasesImprovements
TogetherLink
TogetherLink is an open-source launcher that connects Claude Code, Codex, ChatGPT Desktop, Pi Code, and OpenCode to models hosted by Together AI. Install with one command and run your existing tools without hand-editing provider settings.See Configure Claude Code, Codex, and ChatGPT with Together AI models.Slurm cluster OIDC SSH
Slurm GPU clusters with OIDC enabled now support browser-based SSH through the Together CLI. Choose OIDC on the cluster details page to sign in without uploading an SSH key. Key-based SSH remains available.See Cluster management.Fine-tuning model limits API
The model limits endpoint now returnssupports_full_training so you can check whether a model supports full fine-tuning before submitting a job. When the field is false, the model is LoRA-only.See LoRA vs. full fine-tuning.New modelsImprovements
New serverless models
The following models are now available on serverless:google/gemma-4-12B-it: 262,144 context length.
GPU cluster health monitoring
Passive health checks now document the full set of monitored failure signals and their recommended repair actions. Slurm node unavailability surfaces as a warning-only alert without an automated repair action.See Health checks and Node repair.Deprecations
Model deprecations
The following models have been deprecated and are no longer available on serverless:Qwen/Qwen3-235B-A22B-Instruct-2507-tput. Available as an on-demand dedicated endpoint.meta-llama/Meta-Llama-3-8B-Instruct-Lite. Available as an on-demand dedicated endpoint.zai-org/GLM-5.1. Available as an on-demand dedicated endpoint.
New releases
TorchTitan training health check
You can now run a TorchTitan training health check on your GPU clusters. It runs a short training benchmark on one or more nodes and measures steady-state model FLOPs utilization (MFU) to validate end-to-end training throughput. This feature is in preview and runs on demand on NVIDIA B200 (Blackwell) nodes.See Health checks for the available tests, thresholds, and results.New releasesNew modelsDeprecations
Storage performance health check
You can now run a storage performance health check on your GPU clusters. It usesfio to validate data integrity and measure sequential read and write bandwidth on the cluster’s storage volumes, and it also runs automatically during cluster acceptance testing.See Health checks for the available tests, thresholds, and results.New serverless models
The following models are now available on serverless:LiquidAI/LFM2.5-8B-A1B: 32,768 context length. Pricing: $0.03 input / $0.12 output (per 1M tokens).
Model deprecations
The following model has been deprecated and is no longer available on serverless:LiquidAI/LFM2-24B-A2B. Recommended replacement:LiquidAI/LFM2.5-8B-A1B.
New releasesNew modelsImprovements
Provisioned throughput
Provisioned throughput is now available, allowing you to reserve inference capacity for frontier open models. Commit to a one-month-or-longer term, and Together commits to throughput and reliability targets for traffic within your purchased capacity.At launch, provisioned throughput is available forMiniMaxAI/MiniMax-M3 and zai-org/GLM-5.2.New serverless models
The following models are now available on serverless:HappyHorse/HappyHorse-1.0-T2V(text-to-video).
Reasoning API field
Thereasoning field is now symmetric for input and output on reasoning models. Pass prior assistant reasoning back under reasoning for preserved thinking and multi-turn tool calling. The older reasoning_content key is still accepted on input for backward compatibility.See Reasoning for details.Usage object shape
Theusage object varies by model. Reasoning models nest cached and reasoning token counts under prompt_tokens_details and completion_tokens_details, while some non-reasoning models return cached_tokens at the top level. Read both shapes defensively so clients do not silently report zero.See OpenAI compatibility for examples.Improvements
Project slug editing
Project Admins can now change a Project’s slug from Project Settings. The new slug takes effect immediately. You can also copy any Project’s slug from the Projects list in Organization Settings.Changing a slug can break API requests, scripts, and integrations that reference resources by their slug-qualified path. Update any references that rely on the old slug.See Projects for details.GPU cluster creation region selection
The create cluster flow now defaults the Region field to Any region. Together picks the region with the most available capacity for your GPU type at create time. Changing the GPU type resets the region to Any region and clears any selected shared volume.See the GPU Clusters quickstart for the full create flow.Improvements
New models available for fine-tuning
You can now fine-tune the following vision-language model:google/gemma-4-31B-it-VLM.
Cost analytics units view
Organization and project cost analytics now include a Measure control to switch between Cost ($) and Units. Choose Units to chart daily billable quantities (for example, tokens) by product, line item, project, or API key instead of dollar spend.See Usage limits & analytics for details.Improvements
Bring your own model: Transformers v5
BYOM fine-tuning now supports Hugging Face models built with Transformers v5 or earlier.New models
New serverless models
The following models are now available on serverless:google/flash-image-3.1-lite(Gemini 3.1 Flash-Lite Image).Qwen/Qwen3.6-35B-A3B-Lora: 262,144 context length.alibaba/happyhorse-1.1-i2v(image-to-video).alibaba/happyhorse-1.1-r2v(reference-to-video).alibaba/happyhorse-1.1-t2v(text-to-video).
Deprecations
Model deprecations
The following model has been deprecated and is no longer available on serverless:Qwen/Qwen3.5-397B-A17B.
Deprecations
Model deprecations
The following models have been deprecated and are no longer available on serverless:zai-org/GLM-5.1. Available as an on-demand dedicated endpoint.meta-llama/Meta-Llama-3-8B-Instruct-Lite. Available as an on-demand dedicated endpoint.google/gemma-3n-E4B-it. Available as an on-demand dedicated endpoint.Qwen/Qwen3-235B-A22B-Instruct-2507-tput. Available as an on-demand dedicated endpoint.meta-llama/Llama-Guard-4-12B.
ImprovementsPricing
Seedance 2.0 adds 4K video
ByteDance/Seedance-2.0 now supports a 4k resolution tier (up to 3840x2160), alongside the existing 480p, 720p, and 1080p tiers. Pass resolution: "4k" to generate at the new tier.Pricing for the higher tiers (per second of output):- 1080p: $0.40 (text/image-to-video), from $0.48 (video-to-video).
- 4K: $0.836 (text/image-to-video), from $1.050 (video-to-video).
New releasesImprovements
Automatic node repair for GPU clusters
GPU clusters now support passive health checks and automatic node repair. Passive checks monitor your nodes continuously in the background, and when they (or active checks) detect a node-level issue, the system generates a repair recommendation for you to review and accept from the new Repairs tab. Together then handles the cordon, drain, remediation, and node rejoin.See Health checks and Node repair for details.Estimate fine-tuning job cost via API
A new endpoint,POST /fine-tunes/estimate-price, returns the estimated total price of a fine-tuning job before you launch it, along with estimated training and evaluation token counts and your remaining credit limit. Call it from the Python SDK or TypeScript SDK with the same parameters you plan to submit to the create-job endpoint.See Fine-tuning pricing for details.New models available for fine-tuning
You can now fine-tune the following models:moonshotai/Kimi-K2.7-Code.moonshotai/Kimi-K2.6.
New releasesImprovements
Whoami API endpoint
UseGET /whoami to confirm which API key, Project, and Organization are authenticating a request. The response includes the Project slug used in dedicated endpoint model names.See Whoami for details.Early stopping for fine-tuning
Fine-tuning jobs now support early stopping, which halts training when validation loss stops improving. This reduces cost and helps avoid overfitting on long runs.Enable it by settingearly_stopping_enabled=true on job creation along with a validation_file and n_evals >= early_stopping_patience + early_stopping_warmup_evals + 1. Tune behavior with early_stopping_patience, early_stopping_min_delta, and early_stopping_warmup_evals.When training halts early, the job still finishes with status completed. The response sets early_stopped=true and exposes the winning checkpoint via early_stopping_best_step and early_stopping_best_metric.See Early stopping for details.Audio transcription upload limit
Direct (binary) audio uploads for transcription and translation are now capped at 80 MB per request. For larger files, host the audio at a public HTTPS URL and pass that URL as thefile field, which supports up to 1 GB. When sending a binary upload, place the model form field before the file field in the multipart body.See Transcribe audio for details.New releases
Attach LoRA adapters to a dedicated endpoint
You can now attach multiple LoRA adapters to a single LoRA-enabled dedicated endpoint so they share the same hardware, instead of deploying one endpoint per adapter. Manage bindings from the Python SDK, the TypeScript SDK, the CLI, or the API:together endpoints adapters add <endpoint_id> <endpoint_name>:<adapter_model_name>together endpoints adapters list <endpoint_id>together endpoints adapters remove <endpoint_id> <endpoint_name>:<adapter_model_name>
Deprecations
Model deprecations
The following model has been deprecated and is no longer available on serverless:zai-org/GLM-5. Recommended replacement:zai-org/GLM-5.2.
New modelsImprovements
New serverless models
The following models are now available on serverless:zai-org/GLM-5.2: 262K context length, FP4 quantization. Pricing: $1.40 input / $4.40 output / $0.26 cached input (per 1M tokens). Supports function calling and structured outputs.
Organization and Project role labels
Organization members now use the Admin and Developer labels, and Project collaborators now use Admin and Editor labels. Permissions are unchanged, but the labels make Organization-wide access and Project-scoped editing clearer.See Roles & permissions for details.Deprecations
Model deprecations
The following models are scheduled for deprecation and will no longer be available on serverless after June 29, 2026:Qwen/Qwen3.5-397B-A17B. Recommended replacement:MiniMaxAI/MiniMax-M3, available as an on-demand dedicated endpoint.
New models
New serverless models
The following models are now available on serverless:moonshotai/Kimi-K2.7-Code: 262,144 context length, FP4 quantization. Pricing: $0.95 input / $4.00 output / $0.19 cached input (per 1M tokens). Supports function calling and structured outputs.
New models
New serverless models
The following models are now available on serverless:MiniMaxAI/MiniMax-M3: 524,288 context length, FP4 quantization. Pricing: $0.30 input / $1.20 output / $0.06 cached input (per 1M tokens).
Deprecations
Model deprecations
The following model has been deprecated and is no longer available on serverless:mistralai/Voxtral-Mini-3B-2507. Available as an on-demand dedicated endpoint.
Pricing
Pricing update
The following changes are effective June 9, 2026:New cached input pricing (per 1M tokens):zai-org/GLM-5.1: $0.26 cached input (81% discount from $1.40 standard input).Qwen/Qwen3.5-397B-A17B: $0.35 cached input (42% discount from $0.60 standard input).
deepseek-ai/DeepSeek-V4-Pro (per 1M tokens):- Input: $2.10 → $1.74.
- Output: $4.40 → $3.48.
- Cached input: $0.20 (unchanged).
Improvements
Server-side validation for fine-tuning datasets
Files uploaded for fine-tuning now go through full server-side schema validation during ingestion, with the result exposed on the file object. Poll the Files API and readprocessing_status (COMPLETED, INVALID_FORMAT, or FAILED) plus validation_report to detect dataset issues programmatically before launching a job, like missing role fields or malformed conversation turns.Errors include a user-facing reason, so you can fix the dataset and re-upload without trial-and-error training runs. For example:Deprecations
Model deprecations
The following model has been deprecated and is no longer available on serverless:Qwen/Qwen3-Coder-480B-A35B-Instruct-FP8. Recommended replacement:MiniMaxAI/MiniMax-M2.7, available as an on-demand dedicated endpoint.
New releasesImprovements
Fine-tuning job metrics API
A new API endpoint,GET /fine-tunes/{id}/metrics, returns training metrics for a fine-tuning job (e.g. loss curves and other per-step values) so you can monitor progress programmatically without opening the dashboard. See the API reference and Fine-tuning training metrics for details.Slurm startup scripts for GPU Clusters
GPU clusters now support Slurm startup scripts (lifecycle hook scripts that run at node startup, job allocation, and job completion). Use them to install packages at boot, configure SSH sessions, or run per-job prolog and epilog actions across worker, login, and controller nodes. See Slurm startup scripts for details.Evaluations: Single-pass compare mode
Thecompare evaluator now accepts a disable_position_bias_correction parameter. By default, the judge runs each comparison twice (A→B then B→A) and reconciles verdicts to cancel position bias. Setting disable_position_bias_correction to true runs a single pass, cutting judge cost and latency in half. See AI evaluations for details.Billing documentation updates
Updated billing docs for multiple payment methods, separate invoice addresses, ACH payment behavior, auto-recharge limits with bank transfers, and prepaid-only access (no negative balance limits). See Payment methods & invoices, Credits, and Billing troubleshooting.Pricing
Pricing update
The following models have updated pricing, effective May 29, 2026. All usage from that date forward will be billed at the new rates (per 1M tokens):Qwen/Qwen3.5-9B: $0.10 → $0.17 (input), $0.15 → $0.25 (output).meta-llama/Meta-Llama-3-8B-Instruct-Lite: $0.10 → $0.14 (input), $0.10 → $0.14 (output).meta-llama/Llama-3.3-70B-Instruct-Turbo: $0.88 → $1.04 (input), $0.88 → $1.04 (output).
Deprecations
Model deprecations
The following model has been deprecated and is no longer available on serverless:black-forest-labs/FLUX.1-krea-dev.
New models
New serverless models
The following image and video models are now available on serverless:ImageByteDance/Seedream-5.0-lite.
alibaba/happyhorse-1.0-i2v(image-to-video).alibaba/happyhorse-1.0-r2v(reference-to-video).google/veo-3.1.google/veo-3.1-lite.
New dedicated endpoint models
The following models are now available for deployment on dedicated endpoints:google/gemma-3-1b-it.google/gemma-3-27b-it.google/gemma-3-27b-it-lora.google/gemma-4-31B-it-lora.google/medgemma-27b-text-it.allenai/Molmo-7B-D-0924.meta-llama/Llama-3.2-3B-Instruct.meta-llama/Llama-4-Scout-17B-16E-Instruct-FP8-Lora.Qwen/Qwen2.5-14B.Qwen/Qwen2.5-32B.Qwen/Qwen3-235B-A22B-Instruct-2507-FP8.Qwen/Qwen2-72B.arcee-ai/trinity-mini.BAAI/bge-base-en-v1.5.minimax/speech-2.8-turbo.rime-labs/rime-mist-v3.rime-labs/rime-mist-v3-omni.
Seedance 2.0 quickstart
A quickstart is now available for Seedance 2.0, ByteDance’s unified multimodal audio-video generation model. The guide covers text-to-video, image-to-video, video extension, and instruction-based editing.New releases
GPU Clusters: External OIDC authentication and RBAC
GPU clusters now support external OpenID Connect (OIDC) authentication, allowing each team member to access the cluster’s Kubernetes API using their organization’s identity provider — Google, Okta, Auth0, Microsoft Entra ID, and others.With OIDC enabled, access is managed through standard Kubernetes RBAC: admins bind permissions to individual user identities, and each user authenticates via their browser using SSO. This replaces shared kubeconfig credentials with per-user tokens, per-user audit trails, and clean revocation. Currently this feature is only supported for Kubernetes clusters.OIDC must be configured at cluster creation time. See Set up OIDC authentication for the full setup guide.New models
New serverless models
The following models are now available on serverless:Qwen/Qwen3.7-Max. Pricing: $2.50 input / $7.50 output (per 1M tokens).
Deprecations
Model deprecations
The following model has been deprecated and is no longer available on serverless:moonshotai/Kimi-K2.5.
New models
New serverless models
The following models are now available on serverless:pearl-ai/gemma-4-31b-it: 32,000 context length, INT8 quantization. Pricing: $0.28 input / $0.86 output (per 1M tokens).
PricingDeprecations
Model deprecations
The following models have been deprecated and are no longer available on serverless:deepseek-ai/DeepSeek-R1.deepseek-ai/DeepSeek-V3.1.Qwen/Qwen3-Coder-Next-FP8.
Upcoming pricing update
The following model will have updated pricing, effective May 21, 2026:google/gemma-4-31b-it: $0.20 → $0.39 (input), $0.50 → $0.97 (output) per 1M tokens.
New releasesNew modelsPricing
External collaborators for projects
You can now invite users from outside your organization to collaborate on a project. Enable Allow external collaborators on the project’s settings page, then add them like any other collaborator. The feature is currently in beta. See roles & permissions for more details.New serverless models
The following models are now available on serverless:alibaba/happyhorse-1.0-t2v: $0.24/sec at 1080p.ByteDance/Seedance-2.0: $0.16/sec at 720p.
New releasesImprovements
Together CLI v2.10
The Together CLI has been updated withtg as the canonical command name and a refreshed command tree. Subcommands are now clearer and more consistent across fine-tuning, endpoints, evals, files, clusters, and jig.See CLI reference for details.Speech-to-text and translation: new audio formats
The/v1/audio/transcriptions and /v1/audio/translations endpoints now accept .ogg, .opus, and .aac files in addition to .wav, .mp3, .m4a, .webm, and .flac.Speech-to-text: task field is now optional in verbose JSON responses
Thetask field has been removed from the required fields of AudioTranscriptionVerboseJsonResponse and AudioTranslationVerboseJsonResponse. Clients that previously asserted on its presence should treat it as optional.New releases
Slurm-on-Kubernetes v1.0 for all new Slurm clusters
All newly provisioned Slurm GPU clusters now run on a new Slurm-on-Kubernetes stack with significant reliability improvements. Existing clusters can be migrated in place.What’s new:- Self-healing worker daemons: The Slurm worker daemon is now supervised and auto-restarts on crash, so transient failures recover without operator intervention or impact on healthy nodes.
- Durable job accounting: Job history (
sacct) is now persisted on durable, PVC-backed storage. Restarts and pod reschedules no longer wipe accounting data. - Correct process tracking and cleanup: Job processes (including daemonized children) are tracked at the kernel cgroup level and reliably cleaned up at job completion. No more orphaned processes holding GPU memory or
/dev/shm. - Zombie reaping: A dedicated init process reaps orphaned children, preventing PID-table exhaustion from blocking new jobs.
- GPU state correctness: The Slurm GPU view is rebuilt fresh on every node start, eliminating “GPU not found” failures after pod reschedules.
- Per-cluster GPU utilization metrics: DCGM metrics are now exposed in your cluster’s Grafana dashboards for fine-grained utilization visibility.
Deprecations
Model deprecations
The following model has been deprecated and is no longer available on serverless:MiniMaxAI/MiniMax-M2.5.
Improvements
Text-to-speech: pronunciation_dict parameter
A newpronunciation_dict parameter is available for TTS requests. Pass a list of "<source>/<replacement>" rules (e.g., ["omg/oh my god"]) to override how the model pronounces specific tokens.Together Deployments: custom metric autoscaling
Deployments can now autoscale on any Prometheus metric exposed by your worker’s/metrics endpoint. Set metric = "CustomMetric" and provide a custom_metric_name (e.g., vllm:num_requests_running) along with a target to scale on application-specific signals.Improvements
Fine-tuning: new supported models
The following models are now available for fine-tuning:Qwen/Qwen3.6-35B-A3B.google/gemma-4-31B-it.google/gemma-4-26B-A4B-it.
New modelsPricing
DeepSeek-V4-Pro on serverless
deepseek-ai/DeepSeek-V4-Pro has been added to serverless.- Context length: 512,000.
- Pricing: $2.10 input / $4.40 output / $0.20 cached input (per 1M tokens).
- Quantization: FP4.
- Function calling and structured outputs supported.
New serverless models
The following models are now available on serverless:deepcogito/cogito-v2-1-671b.google/veo-3.1-test-debug.vidu/vidu-q3.vidu/vidu-q3-turbo.Wan-AI/wan2.7-i2v.Wan-AI/wan2.7-r2v.
Pricing update: no-packing fine-tuning jobs
We rolled out a pricing update for no-packing fine-tuning jobs. When the no-packing option is chosen, the number of training dataset tokens is now calculated aslen(dataset) * max_seq_length to account for the compute used by packing-free jobs.max_seq_lengthis configurable in both the SDK and UI.- Price prediction reflects these changes, so if no-packing is chosen you can control the cost of the job by adjusting the sequence length.
New modelsImprovements
Dynamic rate limits and prepaid billing
- Build Tiers 1–5, Scale, and Enterprise tier labels have been retired. Dynamic rate limits are now live for all users.
- Billing has moved to a fully prepaid model.
- Model-specific tier gates have been removed. The platform-wide $5 credit purchase is the only gate.
New serverless models
The following models are now available on serverless:moonshotai/Kimi-K2.6.
Pricing
Pricing update
The following model has updated pricing, effective April 15, 2026:google/gemma-3n-E4B-it: $0.02 → $0.06 (input), $0.04 → $0.12 (output) per 1M tokens.
Deprecations
Model deprecations
The following models have been deprecated and are no longer available:Qwen/Qwen3-VL-8B-Instruct.Qwen/Qwen3-235B-A22B-Thinking-2507.mistralai/Mixtral-8x7B-Instruct-v0.1.
New models
New models
New serverless models
The following models are now available on serverless:google/gemma-4-31B-it.zai-org/GLM-5.1.
Deprecations
Model deprecations
The following models have been deprecated and are no longer available:zai-org/GLM-4.5-Air-FP8.zai-org/GLM-4.7.Qwen/Qwen3-Next-80B-A3B-Instruct.
Deprecations
Model deprecations
The following model has been deprecated and is no longer available:meta-llama/Llama-4-Maverick-17B-128E-Instruct-FP8.
Pricing
Cached input token pricing
Cached input token pricing is now available:MiniMaxAI/MiniMax-M2.5: $0.06 per 1M cached input tokens (80% off standard input price).
New models
Deprecations
Model deprecations
The following models have been deprecated and are no longer available:mixedbread-ai/Mxbai-Rerank-Large-V2.moonshotai/Kimi-K2-Thinking.meta-llama/Llama-3.2-3B-Instruct-Turbo.moonshotai/Kimi-K2-Instruct-0905.
Deprecations
Model deprecations
The following models have been deprecated and are no longer available:black-forest-labs/FLUX.1-dev.black-forest-labs/FLUX.1-dev-lora.black-forest-labs/FLUX.1-kontext-dev.Qwen/Qwen3-VL-32B-Instruct.mistralai/Ministral-3-14B-Instruct-2512.Qwen/Qwen3-Next-80B-A3B-Thinking.Alibaba-NLP/gte-modernbert-base.BAAI/bge-base-en-v1.5.meta-llama/Meta-Llama-3.1-70B-Instruct-Turbo.meta-llama/Llama-Guard-3-11B-Vision-Turbo.meta-llama/LlamaGuard-2-8b.marin-community/marin-8b-instruct.nvidia/NVIDIA-Nemotron-Nano-9B-v2.
New models
New models
New models
New releases
Dedicated Container Inference launch
Together AI has officially launched Dedicated Container Inference (DCI), formerly known as BYOC. DCI lets you containerize, deploy, and scale custom models on Together AI.Deprecations
Model deprecations
The following models have been deprecated and are no longer available:togethercomputer/m2-bert-80M-32k-retrieval.Salesforce/Llama-Rank-V1.togethercomputer/Refuel-Llm-V2.togethercomputer/Refuel-Llm-V2-Small.Qwen/Qwen3-235B-A22B-fp8-tput.qwen-qwen2-5-14b-instruct-lora.meta-llama/Llama-4-Scout-17B-16E-Instruct.Qwen/Qwen2.5-72B-Instruct-Turbo.meta-llama/Meta-Llama-3.1-405B-Instruct-Turbo.BAAI/bge-large-en-v1.5.
New releases
Python SDK v2.0 general availability
Together AI is releasing the Python SDK v2.0, a new, type-safe, OpenAPI-driven client designed to be faster, easier to maintain, and ready for everything we’re building next.- Install:
pip install togetheroruv add together. - Migration guide: A detailed Python SDK Migration Guide covers API-by-API changes, type updates, and troubleshooting tips.
- Code and docs: Access the Together Python v2 repo and reference docs with code examples.
- Main goal: Replace the legacy v1 Python SDK with a modern, strongly-typed, OpenAPI-generated client that matches the API surface more closely and stays in lock-step with new features.
- Net new: All new features will be built in version 2 moving forward. This first version already includes beta APIs for our Instant Clusters.
New modelsDeprecations
New serverless models
The following models are now available on serverless:Qwen/Qwen3-Coder-Next-FP8.
Model deprecations
The following models have been deprecated and are no longer available:deepseek-ai/DeepSeek-R1-0528-tput.
Deprecations
Model redirects
The following models are now being automatically redirected to their upgraded versions. See our Model Lifecycle Policy for details.These are same-lineage upgrades with compatible behavior. If you need the original version, deploy it as a dedicated endpoint.
New models
Deprecations
Model redirects
The following models are now being automatically redirected to their upgraded versions. See our Model Lifecycle Policy for details.These are same-lineage upgrades with compatible behavior. If you need the original version, deploy it as a dedicated endpoint.
ImprovementsDeprecations
Prompt caching now enabled by default for dedicated model inference
Prompt caching is now automatically enabled for all newly created dedicated endpoints. This change improves performance and reduces costs by default.What’s changing:- The
disable_prompt_cachefield (API),--no-prompt-cacheflag (CLI), and related SDK parameters are now deprecated. - Prompt caching will always be enabled. The field is accepted but ignored after deprecation.
- Now: Field is deprecated; setting it has no effect (prompt caching is always on).
- February 2026: Field will be removed.
--no-prompt-cachein CLI commands has no effect. You can remove it.disable_prompt_cachefrom API requests has no effect. You can remove it.- SDK calls that set this parameter have no effect. You can remove it.
New models
Deprecations
Model deprecations
The following models have been deprecated and are no longer available:Qwen/Qwen2.5-VL-72B-Instruct.
Deprecations
Model deprecations
The following models have been deprecated and are no longer available:deepseek-ai/DeepSeek-R1-Distill-Llama-70B.meta-llama/Meta-Llama-3-70B-Instruct-Turbo.black-forest-labs/FLUX.1-schnell-free.meta-llama/Meta-Llama-Guard-3-8B.
Deprecations
Model redirects
The following models are now being automatically redirected to their upgraded versions. See our Model Lifecycle Policy for details.These are same-lineage upgrades with compatible behavior. If you need the original version, deploy it as a dedicated endpoint.
New releases
Python SDK v2.0 release candidate
Together AI is releasing the Python SDK v2.0 Release Candidate, a new, OpenAPI-generated, strongly-typed client that replaces the legacy v1.0 package and brings the SDK into lock-step with the latest platform features.- Install:
pip install together==2.0.0a9. - RC period: The v2.0 RC window starts today and will run for approximately one month. During this time we’ll iterate quickly based on developer feedback and may make a few small, well-documented breaking changes before GA.
- Type-safe, modern client: Stronger typing across parameters and responses, keyword-only arguments, explicit
NOT_GIVENhandling for optional fields, and richtogether.types.*definitions for chat messages, eval parameters, and more. - Redesigned error model: Replaces
TogetherExceptionwith a newTogetherErrorhierarchy, includingAPIStatusErrorand specific HTTP status code errors such asBadRequestError (400),AuthenticationError (401),RateLimitError (429), andInternalServerError (5xx), plus transport (APIConnectionError,APITimeoutError) and validation (APIResponseValidationError) errors. - New Jobs API: Adds first-class support for the Jobs API (
client.jobs.*) so you can create, list, and inspect asynchronous jobs directly from the SDK without custom HTTP wrappers. - New Hardware API: Adds the Hardware API (
client.hardware.*) to discover available hardware, filter by model compatibility, and compute effective hourly pricing fromcents_per_minute. - Raw response and streaming helpers: New
.with_raw_responseand.with_streaming_responsehelpers make it easier to debug, inspect headers and status codes, and stream completions via context managers with automatic cleanup. - Code Interpreter sessions: Adds session management for the Code Interpreter (
client.code_interpreter.sessions.*), enabling multi-step, stateful code-execution workflows that were not possible in the legacy SDK. - High compatibility for core APIs: Most core usage patterns, including
chat.completions,completions,embeddings,images.generate, audio transcription/translation/speech,rerank,fine_tuning.create/list/retrieve/cancel, andmodels.list, are designed to be drop-in compatible between v1 and v2. - Targeted breaking changes: Some APIs (Files, Batches, Endpoints, Evals, Code Interpreter, select fine-tuning helpers) have updated method names, parameters, or response shapes; these are fully documented in the Python SDK Migration Guide and Breaking Changes notes.
- Migration resources: A dedicated Python SDK Migration Guide is available with API-by-API before/after examples, a feature parity matrix, and troubleshooting tips to help teams smoothly transition from v1 to v2 during the RC period.
New models
New serverless models
The following models are now available on serverless:mistralai/Ministral-3-14B-Instruct-2512.
New models
New serverless models
The following models are now available on serverless:zai-org/GLM-4.6.moonshotai/Kimi-K2-Thinking.
New releasesNew models
Real-time text-to-speech and speech-to-text
Together AI expands audio capabilities with real-time streaming for both TTS and STT, new models, and speaker diarization.- Real-time text-to-speech: WebSocket API for lowest-latency interactive applications.
- New TTS models: Orpheus 3B (
canopylabs/orpheus-3b-0.1-ft) and Kokoro 82M (hexgrad/Kokoro-82M), supporting REST, streaming, and WebSocket endpoints. - Real-time speech-to-text: WebSocket streaming transcription with Whisper for live audio applications.
- Voxtral model: New Mistral AI speech recognition model (
mistralai/Voxtral-Mini-3B-2507) for audio transcriptions. - Speaker diarization: Identify and label different speakers in audio transcriptions with a free
diarizeflag. - TTS WebSocket endpoint:
/v1/audio/speech/websocket. - STT WebSocket endpoint:
/v1/realtime.
Deprecations
Image model deprecations
The following image models have been deprecated and are no longer available:black-forest-labs/FLUX.1-pro(calls to FLUX.1-pro will now redirect to FLUX.1.1-pro).black-forest-labs/FLUX.1-Canny-pro.
New releasesNew models
Video generation API and 40+ new image and video models
Together AI expands into multimedia generation with comprehensive video and image capabilities. Read more.- New video generation API: Create high-quality videos with models like OpenAI Sora 2, Google Veo 3.0, and Minimax Hailuo.
- 40+ image and video models: Including Google Imagen 4.0 Ultra, Gemini Flash Image 2.5 (Nano Banana), ByteDance SeeDream, and specialized editing tools.
- Unified platform: Combine text, image, and video generation through the same APIs, authentication, and billing.
- Production-ready: Serverless endpoints with transparent per-model pricing and enterprise-grade infrastructure.
- Video endpoints:
/videos/createand/videos/retrieve. - Image endpoint:
/images/generations.
Improvements
Improved Batch Inference API
- Streamlined UI: Create and track batch jobs in an intuitive interface. No complex API calls required.
- Universal model access: The Batch Inference API now supports all serverless models and private deployments, so you can run batch workloads on exactly the models you need.
- Massive scale jump: Rate limits are up from 10M to 30B enqueued tokens per model per user, a 3,000x increase. Need more? We’ll work with you to customize.
- Lower cost: For most serverless models, the Batch Inference API runs at 50% the cost of our real-time API, making it the most economical way to process high-throughput workloads.
New models
Qwen3-Next-80B models
New Qwen3-Next-80B models are now available for both thinking and instruction tasks.- Model ID:
Qwen/Qwen3-Next-80B-A3B-Thinking. - Model ID:
Qwen/Qwen3-Next-80B-A3B-Instruct.
Improvements
Fine-tuning: new large models supported
Enhanced fine-tuning capabilities with expanded model support. Read more.openai/gpt-oss-120b.deepseek-ai/DeepSeek-V3.1.deepseek-ai/DeepSeek-V3.1-Base.deepseek-ai/DeepSeek-R1-0528.deepseek-ai/DeepSeek-R1.deepseek-ai/DeepSeek-V3-0324.deepseek-ai/DeepSeek-V3.deepseek-ai/DeepSeek-V3-Base.Qwen/Qwen3-Coder-480B-A35B-Instruct.Qwen/Qwen3-235B-A22B(context length 32,768 for SFT and 16,384 for DPO).Qwen/Qwen3-235B-A22B-Instruct-2507(context length 32,768 for SFT and 16,384 for DPO).meta-llama/Llama-4-Maverick-17B-128E.meta-llama/Llama-4-Maverick-17B-128E-Instruct.meta-llama/Llama-4-Scout-17B-16E.meta-llama/Llama-4-Scout-17B-16E-Instruct.
Fine-tuning: increased maximum context lengths
DeepSeek models
- DeepSeek-R1-Distill-Llama-70B: SFT 8,192 → 24,576; DPO 8,192 → 8,192.
- DeepSeek-R1-Distill-Qwen-14B: SFT 8,192 → 65,536; DPO 8,192 → 12,288.
- DeepSeek-R1-Distill-Qwen-1.5B: SFT 8,192 → 131,072; DPO 8,192 → 16,384.
Google Gemma models
- gemma-3-1b-it: SFT 16,384 → 32,768; DPO 16,384 → 12,288.
- gemma-3-1b-pt: SFT 16,384 → 32,768; DPO 16,384 → 12,288.
- gemma-3-4b-it: SFT 16,384 → 131,072; DPO 16,384 → 12,288.
- gemma-3-4b-pt: SFT 16,384 → 131,072; DPO 16,384 → 12,288.
- gemma-3-12b-pt: SFT 16,384 → 65,536; DPO 16,384 → 8,192.
- gemma-3-27b-it: SFT 12,288 → 49,152; DPO 12,288 → 8,192.
- gemma-3-27b-pt: SFT 12,288 → 49,152; DPO 12,288 → 8,192.
Qwen models
- Qwen3-0.6B / Qwen3-0.6B-Base: SFT 8,192 → 32,768; DPO 8,192 → 24,576.
- Qwen3-1.7B / Qwen3-1.7B-Base: SFT 8,192 → 32,768; DPO 8,192 → 16,384.
- Qwen3-4B / Qwen3-4B-Base: SFT 8,192 → 32,768; DPO 8,192 → 16,384.
- Qwen3-8B / Qwen3-8B-Base: SFT 8,192 → 32,768; DPO 8,192 → 16,384.
- Qwen3-14B / Qwen3-14B-Base: SFT 8,192 → 32,768; DPO 8,192 → 16,384.
- Qwen3-32B: SFT 8,192 → 24,576; DPO 8,192 → 4,096.
- Qwen2.5-72B-Instruct: SFT 8,192 → 24,576; DPO 8,192 → 8,192.
- Qwen2.5-32B-Instruct: SFT 8,192 → 32,768; DPO 8,192 → 12,288.
- Qwen2.5-32B: SFT 8,192 → 49,152; DPO 8,192 → 12,288.
- Qwen2.5-14B-Instruct: SFT 8,192 → 32,768; DPO 8,192 → 16,384.
- Qwen2.5-14B: SFT 8,192 → 65,536; DPO 8,192 → 16,384.
- Qwen2.5-7B-Instruct: SFT 8,192 → 32,768; DPO 8,192 → 16,384.
- Qwen2.5-7B: SFT 8,192 → 131,072; DPO 8,192 → 16,384.
- Qwen2.5-3B-Instruct: SFT 8,192 → 32,768; DPO 8,192 → 16,384.
- Qwen2.5-3B: SFT 8,192 → 32,768; DPO 8,192 → 16,384.
- Qwen2.5-1.5B-Instruct: SFT 8,192 → 32,768; DPO 8,192 → 16,384.
- Qwen2.5-1.5B: SFT 8,192 → 32,768; DPO 8,192 → 16,384.
- Qwen2-72B-Instruct / Qwen2-72B: SFT 8,192 → 32,768; DPO 8,192 → 8,192.
- Qwen2-7B-Instruct: SFT 8,192 → 32,768; DPO 8,192 → 16,384.
- Qwen2-7B: SFT 8,192 → 131,072; DPO 8,192 → 16,384.
- Qwen2-1.5B-Instruct: SFT 8,192 → 32,768; DPO 8,192 → 16,384.
- Qwen2-1.5B: SFT 8,192 → 131,072; DPO 8,192 → 16,384.
Meta Llama models
- Llama-3.3-70B-Instruct-Reference: SFT 8,192 → 24,576; DPO 8,192 → 8,192.
- Llama-3.2-3B-Instruct: SFT 8,192 → 131,072; DPO 8,192 → 24,576.
- Llama-3.2-1B-Instruct: SFT 8,192 → 131,072; DPO 8,192 → 24,576.
- Meta-Llama-3.1-8B-Instruct-Reference: SFT 8,192 → 131,072; DPO 8,192 → 16,384.
- Meta-Llama-3.1-8B-Reference: SFT 8,192 → 131,072; DPO 8,192 → 16,384.
- Meta-Llama-3.1-70B-Instruct-Reference: SFT 8,192 → 24,576; DPO 8,192 → 8,192.
- Meta-Llama-3.1-70B-Reference: SFT 8,192 → 24,576; DPO 8,192 → 8,192.
Mistral models
- mistralai/Mistral-7B-v0.1: SFT 8,192 → 32,768; DPO 8,192 → 32,768.
- teknium/OpenHermes-2p5-Mistral-7B: SFT 8,192 → 32,768; DPO 8,192 → 32,768.
Fine-tuning: Hugging Face integrations
- Fine-tune any < 100B parameter CausalLM from Hugging Face Hub.
- Support for DPO variants such as LN-DPO, DPO+NLL, and SimPO.
- Support fine-tuning with maximum batch size.
- Public
fine-tunes/models/limitsandfine-tunes/models/supportedendpoints. - Automatic filtering of sequences with no trainable tokens (e.g., if a sequence prompt is longer than the model’s context length, the completion is pushed outside the window).
New releases
Together Instant Clusters general availability
Self-service NVIDIA GPU clusters with API-first provisioning. Read more.- New API endpoints for cluster management:
/v1/gpu_cluster: Create and manage GPU clusters./v1/shared_volume: High-performance shared storage./v1/regions: Available data center locations.
- Support for NVIDIA Blackwell (HGX B200) and Hopper (H100, H200) GPUs.
- Scale from single-node (8 GPUs) to hundreds of interconnected GPUs.
- Pre-configured with Kubernetes, Slurm, and networking components.
Improvements
Serverless LoRA and dedicated model inference support for evaluations
You can now run evaluations:- Using Serverless LoRA models, including supported LoRA fine-tuned models.
- Using dedicated model inference, including fine-tuned models deployed via dedicated endpoints.
New models
New modelsDeprecations
DeepSeek-V3.1
Upgraded version of DeepSeek-R1-0528 and DeepSeek-V3-0324. Read more.- Dual modes: Fast mode for quick responses; thinking mode for complex reasoning.
- 671B total parameters, with 37B active parameters.
- Model ID:
deepseek-ai/DeepSeek-V3.1.
Model deprecations
The following models have been deprecated and are no longer available:meta-llama/Llama-3.2-90B-Vision-Instruct-Turbo.black-forest-labs/FLUX.1-canny.meta-llama/Llama-3-8b-chat-hf.black-forest-labs/FLUX.1-redux.black-forest-labs/FLUX.1-depth.deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B.NousResearch/Nous-Hermes-2-Mixtral-8x7B-DPO.meta-llama/Llama-3.2-11B-Vision-Instruct-Turbo.meta-llama-llama-3-3-70b-instruct-lora.Qwen/Qwen2.5-14B.meta-llama/Llama-Vision-Free.Qwen/Qwen2-72B-Instruct.google/gemma-2-27b-it.meta-llama/Meta-Llama-3-8B-Instruct.perplexity-ai/r1-1776.nvidia/Llama-3.1-Nemotron-70B-Instruct-HF.Qwen/Qwen2-VL-72B-Instruct.
Improvements
New models
New models
New releases
Together Evaluations framework
Benchmarking platform using LLM-as-a-judge methodology for model performance assessment. Read more.- Create custom LLM-as-a-judge evaluation suites for your domain.
- Supports
compare,classify, andscorefunctionality. - Compare models, prompts, and LLM configs; score and classify LLM outputs.
New models
Qwen3-Coder-480B
Agentic coding model with top SWE-Bench Verified performance. Read more.- 480B total parameters, with 35B active (MoE architecture).
- 256K context length for entire codebase handling.
- Leading SWE-Bench scores on software engineering benchmarks.
- Model ID:
Qwen/Qwen3-Coder-480B-A35B-Instruct-FP8.
New releases
New models
New releases
Whisper speech-to-text APIs
High-performance audio transcription that’s 15x faster than OpenAI, with support for files over 1 GB. Read more.- Multiple audio formats with timestamp generation.
- Speaker diarization and language detection.
- Use the
/audio/transcriptionsand/audio/translationsendpoints. - Model ID:
openai/whisper-large-v3.