Authentication and billing
When using Together AI through Hugging Face, you have two options for authentication:- Direct requests: Add your Together AI API key to your Hugging Face user account settings. Inference requests are sent directly to Together AI, and billing is handled by your Together AI account.
- Routed requests: If you don’t configure a Together AI API key, your requests are routed through Hugging Face and authenticated with a Hugging Face token. Billing for routed requests is applied to your Hugging Face account at standard provider API rates. You don’t need a Together AI account for this option.
- Go to your Hugging Face inference provider settings.
- Add your Together AI API key under the Together AI provider.
- Optionally, set your preferred provider order, which controls the display order in model widgets and code snippets.
You can browse all Together AI models on the Hub and try them directly in the model page widget.
Installation
Install the Hugging Face client for your language:Chat completions
Passprovider="together" to route the request to Together AI. The API key can be a Together AI key (direct requests) or a Hugging Face token (routed requests).
In
@huggingface/inference v3 and later, the HfInference class is a deprecated alias slated for removal. Use InferenceClient instead.reasoning_content field alongside content. The Hugging Face clients only forward their own parameter set, so you can’t pass Together-specific parameters like reasoning through them. If you need to control reasoning, call the Together API directly.
You can swap in any Together AI text model on the Hub.
OpenAI client library
You can also reach Together models through Hugging Face’s OpenAI-compatible router with the OpenAI Python client. The router requires a Hugging Face token. Append:together to the model ID to pin the provider:
Python