Skip to main content
Reasoning fine-tuning adapts a model that supports chain-of-thought reasoning. By providing reasoning or reasoning_content alongside the final assistant response, you shape how the model thinks through problems before producing an answer. This page covers the reasoning data shape, supported models, and launch parameters.
Reasoning models should always be fine-tuned with reasoning data. Training a reasoning model without it can degrade its reasoning ability. If your dataset doesn’t include reasoning, use an instruct model instead.

Supported models

The following models support reasoning fine-tuning. See supported models for context lengths and batch limits.

Prepare your data

Prepare data in a JSONL file. Each assistant message should carry the chain of thought in a reasoning (or reasoning_content) field and the final answer in content.

Conversational format

When fine-tuning reasoning models on conversational data, only the last assistant message is trained on by default. For multi-turn reasoning, split the conversation so each assistant message is the final message in its own example.

Preference format

For preference fine-tuning, both outputs carry reasoning. See preference tuning for the broader DPO workflow.

Validate and upload

Upload your data using the Together Python/TypeScript SDK or the Together CLI:

Launch the job

LoRA is the default. Pass lora=False for full fine-tuning.
For details on every available parameter, see the API reference.

Watch and deploy

Reasoning jobs use the same lifecycle as text jobs:
  • Poll the job with the SDK or CLI. Expect 10 to 30 minutes for a LoRA job on an 8B model with a few thousand examples.
  • Deploy the result on a dedicated endpoint.
  • Call the endpoint with the same chat-completions shape. The model emits reasoning_content alongside content for clients that surface it. See Inference → Reasoning for details.