Skip to content
Guide
General

Rate limits

Per-endpoint request caps on the LlamaParse API, the same on every plan and counted per organization; the 429 responses they return, the other situations that return 429, and how to back off.

Normal use never reaches these caps. They exist to stop accidental floods, such as a retry loop without a delay, and they cap requests per second, not how much work runs in parallel. For per-job caps see File and request limits; for how many jobs run at once on each plan see Plans.

EndpointLimitScope
Parse job creation, POST /api/v2/parse and POST /api/v1/parsing/upload500 requests per 10 seconds, each route counted separatelyOrganization
File upload, POST /api/v1/beta/files50 per secondProject
Classify and Split job creation40 per secondOrganization
Fetching one job (Parse, Extract, Classify, Split)80 per secondOrganization
Listing jobs (Parse, Extract, Classify)5 per secondOrganization
Every other route80 reads or 40 writes per secondOrganization

Limits are the same on every plan, and they count per organization: creating more API keys or projects does not raise them. Enterprise customers can have custom limits.

A request over the limit is refused with 429 Too Many Requests and a body such as:

{ "detail": "Rate limit exceeded. The file upload limit is 50 requests per second." }

The response carries no Retry-After header. Wait briefly and retry; see Backing off.

Three other situations return 429 and are not rate limits:

  • Plan object cap. "You've already reached the maximum number of projects...", also for indexes, users, data sources, and extraction agents. Delete unused resources or upgrade.
  • Too many Extract jobs running. "Resources exhausted. Please wait for existing jobs to complete." Wait for running jobs to finish, then retry.
  • Index sync called in rapid succession. A second sync while one is already running returns 409 Conflict instead. Wait and retry.

The official SDKs retry 429 and 5xx responses on their own with exponential backoff. The Python client makes up to five attempts by default and raises RateLimitError after that:

from llama_cloud import LlamaCloud
client = LlamaCloud(max_retries=8) # reads LLAMA_CLOUD_API_KEY from the environment

For code that calls the API directly, wrap submissions in a backoff loop with jitter:

import random
import time
from llama_cloud import LlamaCloud, RateLimitError
client = LlamaCloud()
def with_backoff(call, attempts=6, base=1.0, cap=60.0):
"""Retry `call` on 429 with exponential backoff and full jitter."""
for attempt in range(attempts):
try:
return call()
except RateLimitError:
if attempt == attempts - 1:
raise
time.sleep(random.uniform(0, min(cap, base * 2**attempt)))
file = with_backoff(lambda: client.files.create(file="FILE_PATH", purpose="parse"))
job = with_backoff(lambda: client.parsing.parse(
file_id=file.id, tier="cost_effective", version="latest", expand=["markdown"]
))
  • Submit in bulk. For many files, one batch replaces thousands of individual job-creation requests.
  • Don’t poll in a tight loop. Fetching a job is capped at 80 per second and listing jobs at 5 per second across your organization. Use wait_for_completion() in the SDKs or webhooks to be told when a job finishes, and page job lists rather than re-listing.
  • Spread uploads over time. The Parse limit is a 10-second window, so a burst of 500 followed by a pause is fine; a steady 60 per second is not.

If you need more, email support@runllama.ai with your organization, the endpoint, the request rate you need, and for how long. Enterprise customers go through their account manager.

Note for AI agents: this documentation is built for programmatic access. - Overview of all docs: https://developers.llamaindex.ai/llms.txt - Any page is available as raw Markdown by appending index.md to its URL — e.g. https://developers.llamaindex.ai/llamaparse/parse/getting_started/index.md - Agent-friendly REST search APIs live under https://developers.llamaindex.ai/api/ — search (BM25 full-text), grep (regex), read (fetch a page), and list (browse the doc tree). See https://developers.llamaindex.ai/llms.txt for parameters. - A hosted documentation MCP server is available at https://developers.llamaindex.ai/mcp. If you support MCP, you can ask the user to install it for browsing these docs directly (an alternative to the REST API). Setup: https://developers.llamaindex.ai/for-agents/mcp/ - Other LlamaIndex tooling for agents — the LlamaParse Platform MCP server, agent skills and plugins, and the n8n node — is mapped at https://developers.llamaindex.ai/for-agents/