> ## Documentation Index
> Fetch the complete documentation index at: https://docs.together.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Spend controls

> Manage Together AI spend with balance cut-offs, auto-recharge, cost attribution, and programmatic monitoring.

Together AI provides balance cut-offs to stop API access, auto-recharge to keep prepaid workloads funded, and cost analytics to investigate spend. Your [organization](/docs/organizations) is the billing boundary: usage across its projects and members is billed to the organization.

## Methods for controlling spend

| Control | What it does | Availability |
| - | - | - |
| Hard cut-off at \$0 | Stops API access when the prepaid balance reaches zero. | Prepaid accounts. |
| In-arrears cut-off | Pauses access when usage passes an agreed negative balance. | Qualifying contracted enterprise accounts. |
| Auto-recharge | Purchases credits before the balance runs out. | Prepaid accounts with a credit or debit card as the default payment method. |
| Serverless access | Turns serverless inference off or on. | Contracted accounts, through your solutions architect. |

Organization and project budgets and spend alerts are [coming soon](#coming-soon-budgets-and-alerts). To monitor usage and costs, see [Usage limits & analytics](/docs/billing-usage-limits).

### Hard cut-off at \$0

If you don't have a sales contract with Together AI, your account runs on [prepaid credits](/docs/billing-credits). API access stops when your organization's balance reaches zero and resumes after you add credits. Customers with a sales contract continue to operate under their contract terms.

<Warning>
  Metering has a short delay, so some usage can arrive after the balance reaches zero. You remain responsible for that usage. The cut-off does not guarantee an exact spending cap.
</Warning>

### In-arrears cut-off

Qualifying enterprise customers billed monthly can set a cut-off at an agreed negative balance. Once usage passes that threshold, access pauses for the rest of the billing period. [Contact sales](https://www.together.ai/contact-sales) to find out whether your account qualifies.

### Turn serverless inference access off or on

If you have a contract, ask your solutions architect to turn serverless access off when you need to stop serverless spend, and to turn it back on later.

## Keep prepaid workloads funded

Prepaid services consume your organization's credit balance. You can [buy credits manually](/docs/billing-credits#buy-credits) for predictable, temporary, or experimental workloads, or enable [auto-recharge](/docs/billing-credits#auto-recharge-credits) for production workloads that need to keep running.

Auto-recharge purchases credits when your balance falls below a threshold you set. Choose the threshold based on your consumption rate: higher-throughput workloads need a larger buffer between a recharge and balance exhaustion. Auto-recharge keeps purchasing credits as you consume them, so it does not cap spend.

Auto-recharge requires a credit or debit card as your default payment method. Setting an Automated Clearing House (ACH) bank account as the default turns auto-recharge off, even if you have a card saved. See [Payment methods & invoices](/docs/billing-payment-methods).

## Attribute spend to projects and API keys

[Projects](/docs/projects) separate resources and API keys by team, environment, application, or workload. Choose project boundaries that match how you assign cost ownership. For example, separate production, staging, and experimentation workloads into their own projects.

Cost analytics can group organization usage by project, and each project has its own cost analytics view.

<Note>
  Projects share the organization's credit balance. A separate project does not get its own prepaid balance or spending cap, and turning off auto-recharge affects funding for the entire organization, including production projects.
</Note>

Within a project, create separate, descriptively named [API keys](/docs/api-keys-authentication) for distinct services or workloads, such as `prod-chat-api` and `prod-agent-worker`. Cost analytics attributes inference usage to individual API keys. API keys do not currently support independent spending limits.

## Investigate spend with cost analytics

Use [cost analytics](/docs/billing-usage-limits#cost-analytics) to view costs or billable units, such as tokens, over a selected date range. You can group usage by product, line item, project, or API key. Project and API key grouping are in beta.

To investigate an increase in spend, start with organization-level cost analytics and group by project. Open the responsible project's cost analytics, then group by API key or line item to identify the workload or product driving the change. Filter individual series to narrow the investigation. Filtering is unavailable when grouping by line item.

## Monitor spend programmatically

The [billing usage API](/reference/billing-usage) returns organization billing usage as cost-annotated line items through `GET /v1/billing/usage`. Use it to build internal dashboards, scheduled reports, anomaly detection, custom alerts, or chargeback systems.

<Note>
  The billing usage API is in beta and requires access to be enabled for your organization. [Contact support](https://portal.usepylon.com/together-ai/forms/support-request) to request access. The response shape can change during beta.
</Note>

Each line item includes the product, quantity, unit price, and cost. Inference line items can include `project_id` and `api_key_id`, so you can join usage to an internal service catalog to map it to the owning team.

The API returns usage for the key's entire organization, and any project's key can read it. Review the [access scope](/reference/billing-usage) before integrating it into internal systems. Send `organization_id` in your requests: it is optional during beta but will be required at general availability.

During beta, current-month usage can be up to 1 hour behind, and prior-month usage up to 24 hours behind. All timestamps are in Coordinated Universal Time (UTC).

## Coming soon: budgets and alerts

<Note>
  Organization and project budgets and spend alerts are planned features. They are not currently available to configure.
</Note>

**Spend alerts** will notify you when organization or project usage reaches thresholds you set, without interrupting your workloads. Until alerts are available, use the [billing usage API](/reference/billing-usage) for custom monitoring.

**Budgets** will stop spending at an organization or project limit. They suit workloads that can tolerate interruption, such as training jobs, development environments, and agent loops that should stop if spend runs away. Until budgets are available, the balance cut-offs above are the only hard stops.

## Recommended setups

### Production workloads

Use a dedicated production project, descriptive API keys, auto-recharge, and cost analytics. Set the recharge threshold high enough for your consumption rate, and integrate the billing usage API with your internal monitoring if you need custom reports or alerts.

### Development and experimentation

Use separate non-production projects and cost analytics to track consumption. If the entire organization can tolerate interruption, keep auto-recharge off and add credits manually. The zero-balance cut-off affects all projects in the organization.

### Larger organizations

Align projects with teams, applications, or environments, and use descriptive API keys per workload. Combine organization and project cost analytics with the billing usage API to feed existing financial operations (FinOps) reporting and ownership data.
