Requirements
- A completed fine-tuning job. See the quickstart for the full lifecycle.
- One of the job’s two registry identifiers for the trained weights. Look for them in the
tg fine-tuning retrieveoutput on a completed job:model_object_id(anml_...value): The model-registry ID. The CLI takes it directly, and the SDK and API take it inside aprojects/{project_id}/models/{model_id}path.model_object_name(a<project_slug>/<model_name>value): The qualified registry name. The CLI accepts it in place of the ID. The SDK and API don’t.
Deploy on a dedicated endpoint
Together AI serves fine-tuned models through dedicated model inference (DMI). A completed fine-tuning job is already a private model in your project, so you deploy it by its Model Object ID (or, from the CLI, by its qualified registry name) with no separate upload step. Deployment and lifecycle operations use thetg beta CLI.
The dedicated model inference commands require Together CLI version
2.24.0 or later. Install or upgrade with uv tool install "together[cli]" (or uv tool upgrade "together[cli]"), and check your version with tg --version.- CLI
- SDK
- UI
The CLI’s Pass the name in full, including the project slug (Once the deployment is When you’re done, delete the endpoint and its deployment to stop billing. The CLI’s To pause billing without deleting anything (for example, if you plan to use the endpoint again later), scale the deployment to zero instead. See Delete resources for the full teardown order.
deploy command creates the endpoint, attaches a deployment, and routes all traffic to it in one step. Pass either the fine-tune’s model_object_name or its model_object_id. The CLI resolves the name to the same model when exactly one model in your project matches it:CLI
acme-corp/my-model-abc123). The bare model name doesn’t resolve, and neither does model_output_name.Poll until the deployment reaches DEPLOYMENT_STATE_READY:CLI
READY, the endpoint serves at its endpoint string (your-project-slug/my-finetuned-endpoint, printed by the deploy output). Pass it as the model parameter and point the base URL at https://api-inference.together.ai/v1:cURL
rm --force removes both in one step:CLI
If
deploy reports that the model has more than one deployment profile, re-run it with --config <cr_...>. List a model’s profiles with tg beta models configs "<MODEL_OBJECT_ID>". See Choose a deployment profile.Run locally
To run your model outside Together, download the checkpoint by job ID:Choose a checkpoint type
Thecheckpoint parameter selects what to download. If you omit it, the API picks a default automatically: merged for LoRA jobs and model_output_path for full fine-tunes. You can also pass checkpoint_step, which downloads a specific intermediate step and overrides checkpoint. Valid values depend on how the job was trained.
model_output_path returns the raw training output directory before any merging. It works for both job types but is mainly useful for advanced workflows that need the unmodified artifacts: for LoRA jobs, prefer merged or adapter; for full fine-tunes, it’s the only option.
The output is a .tar.zst archive that uses ZStandard compression. On macOS, install zstd with Homebrew and decompress:
Python
--checkpoint-step <STEP_NUMBER> to tg fine-tuning download (or checkpoint_step=<STEP_NUMBER> to client.fine_tuning.content()). List checkpoints with tg fine-tuning list-checkpoints <JOB_ID>.
Model registry object IDs
When a job uploads its output artifacts to the Together model registry alongside the primary storage write, the job and checkpoint responses include the registry object and revision IDs, plus qualified object names in<project_slug>/<model_name> form. Use the IDs to reference the uploaded weights in downstream workflows, such as binding a fine-tuned adapter to a dedicated endpoint. The CLI’s deploy command also accepts the qualified name.
On a completed job in the fine-tuning jobs dashboard, the Output model field shows model_object_name and links to that model in the registry by model_object_id.
The object ID fields are omitted when the registry upload did not run or did not succeed. A name can still be present for a job whose upload never ran, but the CLI can’t resolve it, so treat the matching *_object_id as the signal that the artifact reached the registry. Object names are resolved on the fly on GET /fine-tunes/{id} and GET /fine-tunes/{id}/checkpoints only (not on list jobs). If the project slug cannot be resolved, the name fields fall back to the corresponding object ID.
On the job
GET /fine-tunes/{id} includes the final artifact IDs once the job reaches completed:
On checkpoints
GET /fine-tunes/{id}/checkpoints returns the same ID pair on each checkpoint entry:
Intermediate checkpoints carry the IDs from the upload at that step. Final model and adapter checkpoints in the list reuse the job-level IDs from the table above.
Troubleshooting
- No Model Object ID yet: The job hasn’t reached
completed, somodel_object_idisn’t populated. Poll status withclient.fine_tuning.retrieve(id=...)until it’s done. See Monitor a fine-tuning job for the polling pattern. endpoints_v1_create_access_disabled(HTTP 403): You’re calling the retired v1 endpoints API (client.endpoints.create(...)ortg endpoints create). Deploy on dedicated model inference withtg beta endpoints deployinstead. See Migrate from v1.deployreports multiple deployment profiles: Re-run with--config <cr_...>. List a model’s profiles withtg beta models configs "<MODEL_OBJECT_ID>".Model <MODEL_OBJECT_NAME> not foundwhen deploying by name: Passmodel_object_namein full, including the project slug, and check that the job reports amodel_object_id. Without one, the weights never reached the registry, so neither identifier resolves. If the name still fails, deploy bymodel_object_id.Multiple models found for "<MODEL_OBJECT_NAME>": More than one model visible in your project has the same final name segment. Deploy bymodel_object_id, which is unambiguous.- Either identifier fails with a key from another project: The model belongs to the project the job ran in, and both identifiers resolve only within your key’s project. Use a key scoped to the project that owns the model.
- Deploy fails because the base isn’t supported: Not every base model can be hosted for dedicated inference. Confirm the base appears in the supported models list before training (
-Referencemodels often can’t be deployed). - 404 on inference: Point the base URL at
https://api-inference.together.ai/v1and pass the endpoint string (your-project-slug/<endpoint_name>, printed by the deploy output) as themodelparameter, not the Model Object ID.
Next steps
Upload a custom model
Upload your own model weights from outside the Together catalog.
Manage endpoints
Inspect, start, stop, update, and delete dedicated endpoints.
Configure autoscaling
Tune replica bounds and autoscale a deployment on the metric that fits your workload.