Skip to content
Guide
Parse
Guides

Retrieving Results

Use the expand parameter to choose what a Parse job returns — markdown, text, items, metadata, full files, images or download URLs — with the common patterns in Python, TypeScript, Go, Java and the CLI.

expand controls what comes back from a Parse job. By default the API returns only the job itself (status, ID, error message), no parsed content. You add expand values to opt in to the data you want, and you can ask for a different set later against the same job without re-parsing.

This page shows how to use expand: every value, the patterns most callers use, and the limits. For what each returned field contains, see the response format.

These return parsed data in the response body, no download step.

ValueWhat you getNotes
markdownMarkdown per page
markdown_fullThe whole document as one markdown string
textPlain text per page
text_fullThe whole document as one plain-text string
itemsStructured items per page: tables, headings, figures, lists
metadataPer-page metadata: confidence, orientation, watermark, speaker notes
formsPer-page form trees (beta)Jobs with processing_options.forms: "enrich" only; never on fast
job_metadataJob execution metadata: state transitions and timing
usageCredits billed, under job.usageSee Credits billed for a job

These add a presigned download URL under result_content_metadata instead of returning the content. Use them when the result is large and you want to check its size first, stream it to storage, or fetch a file format such as XLSX or PDF. The entry shape and the file each one yields are in Download URLs.

ValueFileNotes
markdown_content_metadataPer-page markdown, JSON
markdown_full_content_metadataThe whole document as one .md
text_content_metadataPer-page text, JSON
text_full_content_metadataThe whole document as one .txt
items_content_metadataItems tree, JSON
metadata_content_metadataPage metadata, JSON
forms_content_metadataForms, JSONJobs with processing_options.forms: "enrich" only
images_content_metadataSaved images, each with its own URLWhich images exist depends on output_options.images_to_save
xlsx_content_metadataTables as a workbookJobs with tables_as_spreadsheet.enable: true only
output_pdf_content_metadataThe rendered PDFJobs with output_options.save_output_pdf: true only
raw_words_content_metadataWord-level boxes, JSONLJobs with output_options.additional_outputs: ["word_bbox"] only
result_content_metadata.grounded_itemsGrounded items sidecar, JSONLIncluded automatically when granular_bboxes is set; nothing to add to expand

Every value works on every tier, with two exceptions: forms needs the form pass, which the fast tier does not run, and a fast job pinned to a version older than 2026-06-15 has no markdown or items at all. See Limitations.

The SDK snippets below assume a client and an uploaded file, as in the quickstart. Per-page results can contain a failed page ("success": false) with no content fields, so the loops check success before reading; see Page arrays.

The simplest pattern. Most LLM pipelines only need this.

result = client.parsing.parse(
file_id=file.id,
tier="agentic",
version="latest",
expand=["markdown"],
)
for page in result.markdown.pages:
if page.success:
print(page.markdown)

For pipelines that need programmatic access to tables, headings, and figures. Items already include an md field per element, so you can build LLM input from items alone, but requesting markdown alongside gives you a ready-made per-page markdown string without reassembling it.

result = client.parsing.parse(
file_id=file.id,
tier="agentic",
version="latest",
expand=["markdown", "items"],
)
# Per-page markdown for the LLM (or use markdown_full for a single string)
llm_input = "\n\n".join(p.markdown for p in result.markdown.pages if p.success)
# Walk the items tree for tables
for page in result.items.pages:
if not page.success:
continue
for item in page.items:
if item.type == "table":
print(f"Table on page {page.page_number}: {len(item.rows)} rows")

For multimodal pipelines that want one markdown blob and the page screenshots as separate downloadable images.

result = client.parsing.parse(
file_id=file.id,
tier="agentic_plus",
version="latest",
output_options={"images_to_save": ["screenshot"]},
expand=["markdown_full", "images_content_metadata"],
)
# One big markdown blob for the LLM
print(result.markdown_full)
# Each page screenshot as a downloadable URL
for image in result.images_content_metadata.images:
print(f"{image.filename}: {image.presigned_url}")

Every per-page representation carries page_number, so you can line up markdown, items and metadata for the same page. Build a lookup keyed by page number rather than zipping arrays, since a failed page is still present in each array but has no content.

result = client.parsing.get(
job_id="JOB_ID", expand=["markdown", "items", "metadata"]
)
confidence = {m.page_number: m.confidence for m in result.metadata.pages}
items_by_page = {p.page_number: p.items for p in result.items.pages if p.success}
for page in result.markdown.pages:
if not page.success:
print(f"page {page.page_number} failed: {page.error}")
continue
n_items = len(items_by_page.get(page.page_number, []))
print(page.page_number, confidence.get(page.page_number), n_items, "items")

You don’t have to ask for everything you might ever need at parse time. Run a Parse job with one set of expand values, then call get later with a different set, without re-running the job.

This is the right pattern when:

  • You don’t know yet what you need. Run with expand=["markdown"], decide later you also want the items tree, fetch with expand=["items"] against the same job ID. Cheaper than re-parsing.
  • You’re working with large files. Use expand=["markdown_content_metadata"] first to check size_bytes and decide whether to download.
  • You’re resuming work on a previously parsed job. Pass the old job ID and an expand list to retrieve any representation the job produced.
from llama_cloud import LlamaCloud
client = LlamaCloud() # reads LLAMA_CLOUD_API_KEY from the environment
# Step 1: parse with a minimal expand
result = client.parsing.parse(
upload_file="doc.pdf",
tier="cost_effective",
version="latest",
expand=["markdown_full"],
)
print(result.job.status)
print(result.markdown_full)
# Step 2: later, retrieve the text representation of the same job
text_result = client.parsing.get(
job_id=result.job.id,
expand=["text_full"],
)
print(text_result.text_full)

images_content_metadata returns every saved image. To fetch a subset, pass their file names in image_filenames; the names come from an earlier full listing or from image items in the items tree. The REST parameter is a single comma-separated string.

result = client.parsing.get(
job_id="JOB_ID",
expand=["images_content_metadata"],
image_filenames="image_0.png,image_5.jpg",
)
for image in result.images_content_metadata.images:
print(f"{image.filename}: {image.presigned_url}")

expand=usage returns the credits billed for a Parse job, under job.usage rather than at the top level. Combine it with content values: expand=markdown,usage returns both.

result = client.parsing.get(job_id="JOB_ID", expand=["usage"])
print(result.job.usage.credits)
{
"job": {
"id": "JOB_ID",
"status": "COMPLETED",
"usage": {
"credits": 30.0
}
}
}

The current fast version returns markdown and items like every other tier. Only a fast job pinned to a version older than 2026-06-15 is text-only: expand=markdown on such a job is rejected with 400 and a message of the form "Markdown expansion is not available for FAST tier jobs on version ...", and items comes back null. Pin version: "latest" or the current fast version to avoid it.

The fast tier runs no form pass. Requesting processing_options.forms: "enrich" on a fast-tier job returns a validation error, so forms and forms_content_metadata exist only on jobs at cost_effective or higher. See Tiers for the full list of what each tier produces.

Presigned URLs expire after a limited time. Download files promptly after retrieving them, or call get again with the same expand to get fresh URLs.

  • Response format — every field each expand value returns, item types, bounding-box conventions
  • Tiers — what each tier produces
  • Configuring Parse — expand is a top-level key, not nested in output_options
  • Output Options — what you can ask Parse to generate, which then becomes available via expand: image assets, tables as spreadsheet, printed page numbers
  • Billing and Usage — plan credits and organization-wide usage, alongside the per-job credits expand=usage returns
Note for AI agents: this documentation is built for programmatic access. - Overview of all docs: https://developers.llamaindex.ai/llms.txt - Any page is available as raw Markdown by appending index.md to its URL — e.g. https://developers.llamaindex.ai/llamaparse/parse/getting_started/index.md - Agent-friendly REST search APIs live under https://developers.llamaindex.ai/api/ — search (BM25 full-text), grep (regex), read (fetch a page), and list (browse the doc tree). See https://developers.llamaindex.ai/llms.txt for parameters. - A hosted documentation MCP server is available at https://developers.llamaindex.ai/mcp. If you support MCP, you can ask the user to install it for browsing these docs directly (an alternative to the REST API). Setup: https://developers.llamaindex.ai/for-agents/mcp/ - Other LlamaIndex tooling for agents — the LlamaParse Platform MCP server, agent skills and plugins, and the n8n node — is mapped at https://developers.llamaindex.ai/for-agents/