Skip to content
Guide
Parse
Guides

Parse response format

Field-by-field reference for what a Parse job returns — the job object, per-page markdown, text, items, metadata and forms, item types, bounding-box conventions, the grounded-items sidecar, image metadata, and download URLs.

A Parse job result is one JSON object. job is always present; every other field appears only when you ask for it with expand. This page is the reference for what each field contains. For which expand values to combine, how to poll, and the common retrieval patterns with code, see Retrieving Results. For the request options that produce each output, see Configuring Parse.

{
"job": { "id": "JOB_ID", "status": "COMPLETED", "...": "..." },
"markdown": { "pages": [ ... ] },
"markdown_full": "# Complete Document\n\n...",
"text": { "pages": [ ... ] },
"text_full": "Complete Document\n\n...",
"items": { "pages": [ ... ] },
"metadata": { "pages": [ ... ], "document": { ... } },
"forms": { "pages": [ ... ] },
"job_metadata": { ... },
"images_content_metadata": { "total_count": 3, "images": [ ... ] },
"result_content_metadata": { "markdown": { ... }, "grounded_items": { ... } }
}
FieldTypeFilled by expandContents
jobobjectalwaysJob status and identifiers. See Job object.
markdownobjectmarkdownPer-page markdown. See Markdown pages.
markdown_fullstringmarkdown_fullThe whole document as one markdown string.
textobjecttextPer-page plain text. See Text pages.
text_fullstringtext_fullThe whole document as one plain-text string.
itemsobjectitemsPer-page structured items. See Items pages.
metadataobjectmetadataPer-page and document-level metadata. See Metadata.
formsobjectformsPer-page form trees. See Forms (Beta).
job_metadataobjectjob_metadataJob execution metadata: state-transition timestamps and timing. Not the job configuration.
images_content_metadataobjectimages_content_metadataSaved images with download URLs. See Images.
result_content_metadataobjectany *_content_metadata valueDownload URLs for result files. See Download URLs.

Fields you did not request are null or absent. Two tier rules apply: forms is never populated on the fast tier, because the form pass does not run there, and a fast job pinned to a version older than 2026-06-15 produces text only, so expand=markdown is rejected with 400 and items comes back null. See Tiers.

job is present on every response, including while the job is still running.

FieldTypeMeaning
idstringParse job identifier. Pass it to later get calls.
statusstringPENDING, RUNNING, COMPLETED, FAILED or CANCELLED. Content fields are populated only once the status is COMPLETED.
project_idstringProject the job belongs to.
namestring or nullOptional display name.
tierstring or nullTier the job ran with (fast, cost_effective, agentic, agentic_plus).
created_at, updated_atstring or nullTimestamps.
error_messagestring or nullError details when status is FAILED.
usageobject or null{ "credits": 30.0 }. Requires expand=usage. credits is null until billing has recorded the job, which can trail completion by a short time.
user_metadataobject or nullThe key/value tags you attached to the request, returned verbatim.

Every per-page representation has the same outer shape: a pages array with one entry per parsed page, in document order. page_number is 1-based and refers to the page’s position in the source document, so it is the key to join markdown, items, metadata and forms entries for the same page, and to match a page screenshot.

RepresentationPage carries successPage carries page_width / page_height
markdown.pages[]yesno
text.pages[]nono
items.pages[]yesyes
metadata.pages[]nono
forms.pages[]yesyes

Where a page carries success, a page that could not be processed is replaced by a failed-page entry with the same page_number and none of the content fields:

{ "page_number": 2, "success": false, "error": "..." }

Check success before reading content fields on those pages. In the typed SDKs the page is a union of the success and failure shapes, so the check also narrows the type. Failed pages count toward the job’s allowed_page_failure_ratio (see timeouts and failure conditions).

expand=markdown returns one entry per page. markdown_full returns the same content joined into a single string instead.

FieldTypeMeaning
page_numberinteger1-based page number.
markdownstringPage content as markdown. Tables render as HTML <table> blocks or markdown pipe tables depending on output_options.markdown.tables.output_tables_as_markdown.
headerstring or nullPage header text, in markdown, when one was detected.
footerstring or nullPage footer text, in markdown, when one was detected.
line_numbersarray or nullPrinted gutter line numbers mapped to offsets in markdown. Only with output_options.markdown.annotate_line_numbers.
successtrueAlways true on a successful page.
{
"page_number": 1,
"success": true,
"markdown": "# Heading\n\n## Subheading\n\nContent with **formatting**...",
"header": "Page header",
"footer": "LlamaIndex 2026",
"line_numbers": [
{ "line_number": "22", "start_index": 0, "end_index": 34 }
]
}

When output_options.markdown.annotate_line_numbers is on, Parse detects physical left-gutter line numbers, removes the printed labels from the content, and maps each one to the text it labelled:

FieldTypeMeaning
line_numberstringThe printed value, kept as a string because documents print 22 and A-3 alike.
start_indexintegerZero-based, inclusive offset into that page’s final markdown.
end_indexintegerZero-based, exclusive offset.

Offsets count UTF-16 code units, matching JavaScript String.slice. The array is omitted when no printed line can be mapped confidently.

expand=text returns the plain-text layer of each page with no markup. With spatial text options set, whitespace preserves the visual layout.

FieldTypeMeaning
page_numberinteger1-based page number.
textstringPlain text of the page.
{ "text": { "pages": [ { "page_number": 1, "text": "Extracted plain text content..." } ] } }

expand=items returns the layout tree: every element Parse detected on the page, typed and in reading order. Use it when you need tables as data, figures as files, or the position of an element on the page.

FieldTypeMeaning
page_numberinteger1-based page number.
page_width, page_heightnumberPage size in points. Bounding boxes on this page are in the same units.
itemsarrayThe item objects below, in reading order.
revisionsarray or nullWord tracked changes and comments. Only with output_options.markdown.annotate_revisions. See Revisions.
successtrueAlways true on a successful page.
{
"page_number": 1,
"page_width": 612.0,
"page_height": 792.0,
"success": true,
"items": [
{ "type": "heading", "level": 1, "value": "Document Title", "md": "# Document Title", "bbox": [ { "x": 72.0, "y": 60.0, "w": 300.0, "h": 24.0 } ] },
{ "type": "table", "rows": [["Header1", "Header2"], ["Row1", "Data1"]], "html": "<table>...</table>", "csv": "Header1,Header2\nRow1,Data1", "md": "| Header1 | Header2 |\n|---|---|..." }
]
}

Every item carries type, md (its markdown rendering) and bbox (an array of bounding boxes, or null). The other fields depend on the type.

typeRepresentsType-specific fields
textA paragraph or other run of body textvalue (plain text)
headingA headinglevel (1–6), value
tableA table, including chart data when chart parsing is onrows, html, csv, merged_from_pages, merged_into_page, parse_concerns — see Tables
imageA figure, photo or diagramurl (link to the extracted image), caption
listAn ordered or unordered listordered (boolean), items (nested text and list items)
codeA code blockvalue (the code), language (identifier or null)
linkA hyperlinktext (display text), url
headerThe page’s running headeritems (the items inside it)
footerThe page’s running footeritems (the items inside it)

header and footer are containers: their items array holds the headings, text, images and so on that sit in the header or footer region, each a normal item. The markdown page object exposes the same regions as plain strings in its header and footer fields.

FieldTypeMeaning
rowsarray of arraysCell values row by row. Each cell is a string, a number or null.
htmlstringThe table as an HTML <table>. Keeps merged cells (colspan / rowspan).
csvstringThe table as CSV.
mdstringThe table as a markdown pipe table. Cannot express merged cells.
merged_from_pagesarray of integers or nullWhen merge_continued_tables is on, the page numbers whose table fragments were merged into this one.
merged_into_pageinteger or nullOn a fragment that was merged elsewhere: the page where the full merged table starts.
parse_concernsarray or nullQuality flags from table extraction. Each entry has type (for example inconsistent_row_cell_count) and human-readable details.

With specialized chart parsing enabled, the data read from a chart is returned as a table item on that page, so the series values are available in rows like any other table. The chart parsing example walks through loading one into a DataFrame.

When output_options.markdown.annotate_revisions is on, an items page that contains Word tracked changes or reviewer comments also carries a revisions array:

{
"page_number": 1,
"items": [ ... ],
"revisions": [
{
"type": "deleted",
"target": "within thirty (30) days",
"content": "within thirty (30) days",
"author": "J. Reviewer",
"target_bbox": { "x": 120.4, "y": 302.1, "w": 96.2, "h": 11.0 },
"revision_bbox": { "x": 460.0, "y": 298.5, "w": 140.0, "h": 24.0 },
"start_index": 512,
"end_index": 536
}
],
"success": true
}
FieldTypeMeaning
typestringOne of inserted, deleted, formatted, moved_from, moved_to, comment.
targetstringThe page text the revision applies to.
contentstringThe revision or comment content.
authorstring or nullReviewer name, when available.
target_bboxobjectBounding box of the target text: x, y, w, h in page points.
revision_bboxobjectBounding box of the printed revision balloon.
target_spansarray or nullPresent when the target is discontinuous: each span has its own target, target_bbox, start_index and end_index.
start_index, end_indexinteger or nullOffsets of the target in that page’s final markdown, inclusive start and exclusive end.

Items, form fields, revisions and images carry bounding boxes that locate them on the page. On items and form fields bbox is an array, because one element can occupy several rectangles: a paragraph that wraps around a figure, or a form field whose fillable area is split.

FieldTypeMeaning
x, ynumberPosition of the box’s top-left corner.
w, hnumberWidth and height.
rnumber, optionalClockwise rotation of the visible text in degrees, around the box’s center. Omitted when unrotated.
labelstring or nullLabel for the box, when one applies.
confidencenumber or nullConfidence score for the box, when available.
start_index, end_indexinteger or nullOptional character offsets into the item’s text.

Conventions, the same for every box type:

  • Units are page points, not pixels and not normalized. The page’s page_width and page_height (on the items page, forms page and grounded-items sidecar) give the frame the box lives in. A US Letter page is 612 × 792.
  • Origin is the top-left corner of the page, x increasing to the right and y increasing downward, matching image coordinates.
  • Boxes are relative to the page the item sits on. Read page_width and page_height from that page entry; pages in one document can differ in size.
  • x/y/w/h describe the unrotated rectangle; r rotates it. Parse has already composed the page’s original_orientation_angle (from metadata) into the box. Apply only r; never add the page orientation to it.

Worked example. On a 612 × 792 page, a word box of { "x": 72.0, "y": 100.0, "w": 35.0, "h": 12.0 } sits 72 points from the left edge and 100 points from the top. To draw it on a page screenshot of 1224 × 1584 pixels, scale by 1224 / 612 = 2 horizontally and 1584 / 792 = 2 vertically: the rectangle runs from pixel (144, 200) to (214, 224). When the two scale factors differ, rotate the corners by r first and scale after, or the box distorts. The granular bounding boxes example has a complete rendering helper that handles rotation and page orientation.

Camera photos. For photos corrected with camera_photo_correction, Parse maps each rotated text rectangle back into the original photo frame and exports the best oriented-rectangle approximation. A perspective quadrilateral cannot be represented exactly by one rectangle plus one angle, so small edge differences are possible on strongly skewed photos. If that approximation would extend beyond the photo boundary, Parse clips it to a bounded axis-aligned box and sets r to 0.

Item-level boxes are the finest grain inlined on the result. Setting output_options.granular_bboxes to any of "word", "line", "cell" makes Parse write a separate grounded-items file with per-line, per-word or per-table-cell boxes for every item. Its download URL is included automatically as result_content_metadata.grounded_items; nothing extra goes in expand.

The file is JSONL, one JSON object per line and one line per page, not a JSON array. Each row is either a success row or a failure row:

// Success row
{
"page_number": 1,
"page_width": 612,
"page_height": 792,
"success": true,
"items": [
{
"type": "text",
"md": "Hello world",
"bbox": [{ "x": 72.0, "y": 100.0, "w": 200.0, "h": 12.0 }],
"grounding": {
"source": "md",
"lines": [
{
"span": [0, 11],
"bbox": { "x": 72.0, "y": 100.0, "w": 200.0, "h": 12.0 },
"words": [
{ "span": [0, 5], "bbox": { "x": 72.0, "y": 100.0, "w": 35.0, "h": 12.0 } },
{ "span": [6, 11], "bbox": { "x": 110.0, "y": 100.0, "w": 40.0, "h": 12.0 } }
]
}
]
}
}
]
}
// Failure row — grounding could not be produced for this page
{ "page_number": 2, "success": false, "error": "..." }
FieldMeaning
items[].type, md, bboxThe same item fields as the inline items representation. Nested items (list entries) appear under items[].items.
items[].grounding.sourceWhich text the spans index into: md for the item’s markdown, caption for an image caption.
items[].grounding.lines[]One entry per text line: a bbox and a [start, end) span into the source text.
items[].grounding.lines[].words[]One entry per word, when "word" was requested: a bbox and a span.
items[].grounding.rows[row][col]For table items: per-cell bbox, span and lines, with null for missing cells, plus row_bboxes and column_bboxes.

Spans are UTF-8 byte offsets, not character offsets. Encode the source text as UTF-8 before slicing, or the offsets drift on any non-ASCII character:

md_bytes = item["md"].encode("utf-8")
for line in item["grounding"]["lines"]:
start, end = line["span"]
print(md_bytes[start:end].decode("utf-8"), line["bbox"])

For a complete worked example, from request to walking the per-word grounding, see the granular bounding boxes example.

expand=metadata returns per-page quality and provenance fields plus a document block.

FieldTypeMeaning
page_numberinteger1-based page number.
confidencenumber or nullHow confident Parse is in this page’s output, 0–1. With confidence_score_effort: "high", this reflects the high-effort assessment.
cost_optimizedboolean or nulltrue when Cost Optimizer routed this page to cost_effective.
triggered_auto_modeboolean or nulltrue when an auto_mode_configuration rule fired on this page.
original_orientation_angleinteger or nullRotation Parse applied to read the page, in degrees; 0 when none was needed. Already composed into every bounding box on the page.
printed_page_numberstring or nullThe page number as printed on the page ("i", "A-3"). Only with output_options.extract_printed_page_number.
watermarkstring or nullWatermark text detected on the page; several are joined with |. Absent when none. Version 2026-09-28 or later on cost_effective, agentic and agentic_plus, whatever output_options.watermark_handling is set to; see Watermarks.
speaker_notesstring or nullPresentation inputs only: the slide’s speaker notes.
slide_section_namestring or nullPresentation inputs only: the section the slide belongs to.
{
"page_number": 1,
"confidence": 0.985,
"cost_optimized": false,
"triggered_auto_mode": false,
"original_orientation_angle": 0,
"printed_page_number": null,
"watermark": "CONFIDENTIAL",
"speaker_notes": null,
"slide_section_name": null
}

metadata.document is populated when confidence_score_effort: "high" is set.

FieldTypeMeaning
confidencenumber or nullMean confidence across the pages the high-effort judge scored, 0–1.
confidence_breakdown.min_page_scorenumberThe lowest page score, the worst page in the document.
confidence_breakdown.scored_pagesintegerPages the judge scored.
confidence_breakdown.total_pagesintegerPages in the document.

expand=forms is populated only on jobs created with processing_options.forms: "enrich"; otherwise it is null. Every page of the document has an entry; a page with no detected form has "forms": [].

FieldTypeMeaning
page_numberinteger1-based page number.
page_width, page_heightnumber or nullPage size in points, the frame for field bbox values.
detected_form_typesarray of strings or nullForm types detected on the page; null when the page was not treated as a form.
formsarrayOne object per form detected on the page.
successtrueAlways true on a successful page.

Each form holds the same content twice: json, an ordered tree of nodes, and list, a flattened bullet list whose md drops straight into a prompt.

Node typeRepresentsFields
sectionA grouping printed on the form (Part III, box 15)id, label, items (child nodes in reading order)
fieldOne entry: text input, checkbox, select group or signature linefield, id, label, value, isEmpty, valueItems, bbox
tableA fillable gridid, label, columns, rows, bbox

How a field node reads depends on its field kind:

fieldvalueNotes
textThe entered text, verbatimA printed-but-blank field has isEmpty: true and no value.
checkboxboolean, checked or not
signatureboolean, signed or not
single_select, multi_selectnonevalueItems lists the options, usually checkbox fields with their own boolean value.

bbox on a field or table is an array of bounding boxes around the fillable area, in page points. In a table node, each cell of rows is a string, null for a printed-but-blank cell, or { "items": [...] } holding the cell’s own nodes (a checkbox column, for example). When the job also sets output_options.granular_bboxes, a node may carry an optional grounding object that locates its printed id, label and string value (and, for tables, columns and scalar cells) as per-line boxes with UTF-8 byte spans into that text.

{
"page_number": 1, "page_width": 612, "page_height": 792, "success": true,
"forms": [
{
"json": [
{ "type": "field", "field": "text", "id": "1", "label": "Wages, tips, other compensation", "value": "29,513",
"bbox": [ { "x": 349.2, "y": 96.5, "w": 114.0, "h": 12.2 } ] },
{ "type": "field", "field": "multi_select", "id": "13",
"valueItems": [
{ "type": "field", "field": "checkbox", "label": "Statutory employee", "value": true },
{ "type": "field", "field": "checkbox", "label": "Retirement plan", "value": false }
] }
],
"list": { "type": "list", "ordered": false, "md": "- [1] Wages, tips, other compensation: 29,513\n- [13]\n - [x] Statutory employee\n - [ ] Retirement plan", "items": [ ... ] }
}
]
}

The Python SDK exposes json as json_, valueItems as value_items and isEmpty as is_empty; the raw response and the TypeScript SDK use the names shown here. For a walker that reaches every field through nested sections and table cells, see the enriched forms example.

expand=images_content_metadata lists the images saved for the job, each with its own download URL. Which images exist depends on output_options.images_to_save.

{
"images_content_metadata": {
"total_count": 2,
"images": [
{ "index": 0, "filename": "image_0.png", "category": "screenshot", "content_type": "image/png",
"bbox": { "x": 0, "y": 0, "w": 612, "h": 792 }, "presigned_url": "https://..." },
{ "index": 1, "filename": "image_1.jpg", "category": "layout", "content_type": "image/jpeg",
"bbox": { "x": 72, "y": 300, "w": 468, "h": 210 }, "presigned_url": "https://..." }
]
}
}
FieldTypeMeaning
total_countintegerNumber of images returned.
images[].indexintegerPosition in extraction order.
images[].filenamestringFile name, for example image_0.png. Pass it to the image_filenames parameter to fetch a subset.
images[].categorystring or nullscreenshot (a full-page render), embedded (an image file inside the document) or layout (a figure cropped from the page).
images[].content_typestring or nullMIME type.
images[].bboxobject or nullWhere the image sits on its page: x, y, w, h in page coordinates, rounded to integers.
images[].presigned_urlstring or nullTemporary download URL.
images[].size_bytes—Deprecated; always null.

image items in the items tree carry their own url to the image they represent.

Rotated pages. A full-page screenshot is rendered in the page’s original orientation. To draw bounding boxes on it, also request expand=metadata and read original_orientation_angle for the same page_number (match on page number, not on image index). Either rotate the screenshot upright first, or apply the inverse page rotation to the box corners; apply only each box’s own r on top. The rendering recipe shows both.

Every *_content_metadata value adds an entry under result_content_metadata instead of returning the content inline. Each entry has the same shape:

{
"result_content_metadata": {
"markdown": { "size_bytes": 45678, "exists": true, "presigned_url": "https://..." },
"items": { "size_bytes": 67890, "exists": true, "presigned_url": "https://..." },
"grounded_items": { "size_bytes": 123456, "exists": true, "presigned_url": "https://..." }
}
}
FieldMeaning
size_bytesSize of the result file. Check it before deciding whether to download.
existsWhether the file was produced for this job.
presigned_urlTemporary download URL. Request the result again with the same expand to get a fresh one once it expires.

The key under result_content_metadata and the file you download:

expand valueKeyFile
markdown_content_metadatamarkdownPer-page markdown, JSON with the same pages shape as the inline field
text_content_metadatatextPer-page text, JSON
items_content_metadataitemsItems tree, JSON
metadata_content_metadatametadataPage metadata, JSON
forms_content_metadataformsForms, JSON. Jobs with processing_options.forms: "enrich" only
markdown_full_content_metadatamarkdown_full_content_metadataThe whole document as one .md file
text_full_content_metadatatext_full_content_metadataThe whole document as one .txt file
xlsx_content_metadataxlsxTables as a workbook. Jobs with tables_as_spreadsheet.enable: true only
output_pdf_content_metadataoutputPDFThe rendered PDF. Jobs with output_options.save_output_pdf: true only
raw_words_content_metadataraw_wordsOne JSON object per word with page number and box, as JSONL. Jobs with output_options.additional_outputs: ["word_bbox"] only
none neededgrounded_itemsThe grounded items sidecar, JSONL. Jobs with output_options.granular_bboxes set only

images_content_metadata is the exception: it returns the images object at the top level rather than an entry here.

Results are not paginated. A per-page representation comes back as the complete pages array in one response, and markdown_full and text_full as one string. For long documents, request the *_content_metadata variant, read size_bytes, and download the file instead of holding the content in the response.

Note for AI agents: this documentation is built for programmatic access. - Overview of all docs: https://developers.llamaindex.ai/llms.txt - Any page is available as raw Markdown by appending index.md to its URL — e.g. https://developers.llamaindex.ai/llamaparse/parse/getting_started/index.md - Agent-friendly REST search APIs live under https://developers.llamaindex.ai/api/ — search (BM25 full-text), grep (regex), read (fetch a page), and list (browse the doc tree). See https://developers.llamaindex.ai/llms.txt for parameters. - A hosted documentation MCP server is available at https://developers.llamaindex.ai/mcp. If you support MCP, you can ask the user to install it for browsing these docs directly (an alternative to the REST API). Setup: https://developers.llamaindex.ai/for-agents/mcp/ - Other LlamaIndex tooling for agents — the LlamaParse Platform MCP server, agent skills and plugins, and the n8n node — is mapped at https://developers.llamaindex.ai/for-agents/