Overview
How might we design rate limit warnings so people can trust and act on AI output?
When to use
- Essential for applications using AI APIs with rate limits, developer tools, and platforms where proactive limit management prevents service interruptions.
When to skip
- Truly unlimited internal deployments where limits never apply.
- Background batch jobs better served by queues than interactive warnings.
- When limits are so high that constant nags train users to ignore them.
Rules
Failing only after submit with a cryptic HTTP 429 and no retry guidance.
Warnings that push upgrade without showing remaining allowance.
Inconsistent units between the warning and the billing page.
Blocking the composer with no estimate of when capacity returns.
Evidence
| Product | Implementation |
|---|---|
| ChatGPT / Claude Pro | Usage and limit messaging when approaching plan caps. |
| Cursor | Premium model and agent usage warnings against plan quotas. |
| API consoles (OpenAI, Anthropic) | Rate limit headers and dashboard alerts for RPM/TPM. |
| Perplexity | Pro/feature gates when free limits are exhausted. |