For the complete documentation index, see llms.txt. This page is also available as Markdown.

Rate Limits

Understanding API rate limits and how to manage them effectively.


Current Limits

Resource
Limit

Requests per minute (RPM)

60

Tokens per minute (TPM)

2,000,000

These limits apply per API key. Each key you create has its own independent rate limit.


Need Higher Limits?

Our standard limits support most development and production workloads. We're committed to supporting growing projects while ensuring platform stability for all users.

Shared rate limits scale with your total spend. The tiers below are best-effort and are not guaranteed throughput:

Total spend (USD)
Requests per minute
Tokens per minute

Default

60

2,000,000

$250

up to 600

20,000,000

$1,000

2,500

30,000,000

$2,500

5,000

180,000,000

Beyond the top tier, or when you need guaranteed throughput, we move you to dedicated inference.

For enterprise workloads requiring higher limits:

📧 Contact us: support@tensorx.ai

Include in your request:

  • Your use case and expected scale

  • Current bottlenecks you're experiencing

  • Your account email

Enterprise clients needing volume beyond the top tier, or guaranteed throughput rather than best-effort shared limits, move to dedicated inference. We'll work with you to find the right balance for your needs.


What Happens When You Hit a Limit

When you exceed rate limits, the API returns a 429 Too Many Requests error:

The error message includes:

  • Which limit you hit (requests or tokens)

  • Your current limit and remaining count

  • When the limit resets


Checking Your Rate Limit Status

Every API response includes headers showing your current usage:

Header
Example Value
Description

x-ratelimit-api_key-limit-requests

60

Max requests per minute

x-ratelimit-api_key-remaining-requests

45

Requests remaining this minute

x-ratelimit-api_key-limit-tokens

2000000

Max tokens per minute

x-ratelimit-api_key-remaining-tokens

1850000

Tokens remaining this minute

Check these headers to monitor your usage before hitting limits.


Handling Rate Limits

Retry with Exponential Backoff

The best practice is to retry with increasing delays:

Spread Out Requests

If you're making many requests, add small delays between them:


Tips to Stay Under Limits

Tip
How It Helps

Batch similar requests

Fewer API calls

Cache responses

Don't repeat identical queries

Use streaming

One request for long outputs

Set appropriate max_tokens

Avoid generating unnecessary tokens

Queue requests

Smooth out traffic spikes


Monitoring Your Usage

Check your request patterns in your Usage Dashboard:

  • See request counts over time

  • Identify peak usage periods

  • Spot patterns that might cause rate limiting

Last updated

Was this helpful?