Rate limits
Per-endpoint request caps on the LlamaParse API, the same on every plan and counted per organization; the 429 responses they return, the other situations that return 429, and how to back off.
Normal use never reaches these caps. They exist to stop accidental floods, such as a retry loop without a delay, and they cap requests per second, not how much work runs in parallel. For per-job caps see File and request limits; for how many jobs run at once on each plan see Plans.
| Endpoint | Limit | Scope |
|---|---|---|
Parse job creation, POST /api/v2/parse and POST /api/v1/parsing/upload | 500 requests per 10 seconds, each route counted separately | Organization |
File upload, POST /api/v1/beta/files | 50 per second | Project |
| Classify and Split job creation | 40 per second | Organization |
| Fetching one job (Parse, Extract, Classify, Split) | 80 per second | Organization |
| Listing jobs (Parse, Extract, Classify) | 5 per second | Organization |
| Every other route | 80 reads or 40 writes per second | Organization |
Limits are the same on every plan, and they count per organization: creating more API keys or projects does not raise them. Enterprise customers can have custom limits.
The 429 response
Section titled “The 429 response”A request over the limit is refused with 429 Too Many Requests and a body such as:
{ "detail": "Rate limit exceeded. The file upload limit is 50 requests per second." }The response carries no Retry-After header. Wait briefly and retry; see
Backing off.
Three other situations return 429 and are not rate limits:
- Plan object cap.
"You've already reached the maximum number of projects...", also for indexes, users, data sources, and extraction agents. Delete unused resources or upgrade. - Too many Extract jobs running.
"Resources exhausted. Please wait for existing jobs to complete."Wait for running jobs to finish, then retry. - Index sync called in rapid succession. A second sync while one is already running returns
409 Conflictinstead. Wait and retry.
Backing off
Section titled “Backing off”The official SDKs retry 429 and 5xx responses on their own with exponential backoff. The
Python client makes up to five attempts by default and raises RateLimitError after that:
from llama_cloud import LlamaCloud
client = LlamaCloud(max_retries=8) # reads LLAMA_CLOUD_API_KEY from the environmentFor code that calls the API directly, wrap submissions in a backoff loop with jitter:
import randomimport time
from llama_cloud import LlamaCloud, RateLimitError
client = LlamaCloud()
def with_backoff(call, attempts=6, base=1.0, cap=60.0): """Retry `call` on 429 with exponential backoff and full jitter.""" for attempt in range(attempts): try: return call() except RateLimitError: if attempt == attempts - 1: raise time.sleep(random.uniform(0, min(cap, base * 2**attempt)))
file = with_backoff(lambda: client.files.create(file="FILE_PATH", purpose="parse"))job = with_backoff(lambda: client.parsing.parse( file_id=file.id, tier="cost_effective", version="latest", expand=["markdown"]))Design for it
Section titled “Design for it”- Submit in bulk. For many files, one batch replaces thousands of individual job-creation requests.
- Don’t poll in a tight loop. Fetching a job is capped at 80 per second and listing jobs at
5 per second across your organization. Use
wait_for_completion()in the SDKs or webhooks to be told when a job finishes, and page job lists rather than re-listing. - Spread uploads over time. The Parse limit is a 10-second window, so a burst of 500 followed by a pause is fine; a steady 60 per second is not.
If you need more, email support@runllama.ai with your organization, the endpoint, the request rate you need, and for how long. Enterprise customers go through their account manager.