Skip to content
Guide
For Agents
Integrations

Use LlamaParse from the LlamaIndex Framework

Feed Parse output into a LlamaIndex Framework index or agent, and use Extract results as structured tool output, with the current llama-cloud SDK.

The LlamaIndex Framework is our open-source Python toolkit for building agents and retrieval-augmented generation (RAG) over your own data. Its built-in readers handle clean text well. For scanned PDFs, forms, spreadsheets and slide decks, run a Parse job first and hand the result to the framework: the quality of what goes into an index decides the quality of every answer it gives.

Both products use the same API key. Set LLAMA_CLOUD_API_KEY in your environment and install the two packages:

Terminal window
pip install llama-index "llama-cloud>=2.8"
export LLAMA_CLOUD_API_KEY=llx-...

parsing.parse() uploads the file, waits for the job to finish, and returns the result. With expand=["markdown"] the result carries one Markdown string per page. Wrap each page in a framework Document and build an index as you would from any other text:

from llama_cloud import LlamaCloud
from llama_index.core import Document, VectorStoreIndex
client = LlamaCloud() # reads LLAMA_CLOUD_API_KEY
file = client.files.create(file="data/report.pdf", purpose="parse")
result = client.parsing.parse(
file_id=file.id, tier="agentic", version="latest", expand=["markdown"]
)
pages = result.markdown.pages
documents = [
Document(text=p.markdown, metadata={"page": p.page_number})
for p in pages
if p.success
]
index = VectorStoreIndex.from_documents(documents)
query_engine = index.as_query_engine()
response = query_engine.query("What does this document say about pricing?")
print(response)

Pages that failed to parse come back with success set to False and no Markdown, which is why the comprehension filters on it. Keeping the page number in metadata lets retrieval results point back to the page they came from.

The same pattern works for a folder of files: run one Parse job per file, collect the documents, and build a single index. For larger collections, Index keeps an index in sync with your data sources and serves retrieval over the API, so nothing needs to be rebuilt locally.

Extract returns JSON shaped by your schema. A framework agent can call it as a tool, so the agent decides when to pull structured data out of a document:

from llama_cloud import LlamaCloud
from llama_index.core.agent.workflow import FunctionAgent
from llama_index.llms.openai import OpenAI
client = LlamaCloud()
def extract_invoice(path: str) -> dict:
"""Extract vendor, invoice number, date and total from an invoice file."""
file = client.files.create(file=path, purpose="extract")
job = client.extract.run( # creates the job and waits for it to finish
file_input=file.id,
configuration={
"data_schema": {
"type": "object",
"properties": {
"vendor": {"type": "string"},
"invoice_number": {"type": "string"},
"date": {"type": "string"},
"total": {"type": "number"},
},
}
},
)
return job.extract_result
agent = FunctionAgent(
tools=[extract_invoice],
llm=OpenAI(model="gpt-4.1-mini"),
system_prompt="You answer questions about invoices by extracting their fields.",
)

See Getting started with Extract for schema options and Building agents for the agent side.

NeedUse
Clean text files, Markdown, HTMLThe framework’s built-in readers
Local parsing with no API keyLiteParse, the open-source parser
Scans, tables, charts, forms, 130+ formatsParse, as above
Structured fields from documentsExtract, as above
A managed index that stays in sync with a data sourceIndex
Note for AI agents: this documentation is built for programmatic access. - Overview of all docs: https://developers.llamaindex.ai/llms.txt - Any page is available as raw Markdown by appending index.md to its URL — e.g. https://developers.llamaindex.ai/llamaparse/parse/getting_started/index.md - Agent-friendly REST search APIs live under https://developers.llamaindex.ai/api/ — search (BM25 full-text), grep (regex), read (fetch a page), and list (browse the doc tree). See https://developers.llamaindex.ai/llms.txt for parameters. - A hosted documentation MCP server is available at https://developers.llamaindex.ai/mcp. If you support MCP, you can ask the user to install it for browsing these docs directly (an alternative to the REST API). Setup: https://developers.llamaindex.ai/for-agents/mcp/ - Other LlamaIndex tooling for agents — the LlamaParse Platform MCP server, agent skills and plugins, and the n8n node — is mapped at https://developers.llamaindex.ai/for-agents/