Skip to content
Guide
Parse
Features

Layout and bounding boxes

How Parse returns page layout as typed items with bounding boxes, adds per-word, per-line, and per-cell boxes with granular_bboxes, and the coordinate conventions for drawing those boxes on a page.

Every Parse job on the current version of its tier can return an items tree: an ordered list of typed elements per page (headings, text, tables, images, headers), each with its markdown and an item-level bounding box. For citation highlighting and grounding, output_options.granular_bboxes adds boxes per word, per line, and per table cell, delivered as a JSONL sidecar. The page markdown also separates out running headers and footers, and crop_box strips page chrome before parsing.

  • Highlighting the exact span a citation points to in a PDF viewer.
  • Grounding extracted answers to a word, line, or table cell rather than a whole paragraph.
  • A side-by-side preview that syncs markdown text to the region on the source page.
  • Dropping repeated headers and footers, or keeping them separate from body text.
  • Reconstructing reading order and element types for downstream chunking.

A fast job pinned to a version older than 2026-06-15 produces text only, so markdown, items and granular boxes are not available there.

OptionTypeDefaultWhat it does
output_options.granular_bboxesarray of "word", "line", "cell"[]Compute boxes at the chosen levels. Written to a grounded_items JSONL sidecar whose download URL is auto-included on the result; nothing to add to expand.
output_options.images_to_savearraysaves layout when the output links to cropped imagesAdd "screenshot" to get a full-page render to draw boxes on. Retrieve with expand=["images_content_metadata"].
output_options.extract_printed_page_numberbooleanunsetReturn the page number as printed on the page (v, A-3) in metadata.pages[].printed_page_number, for citations.
crop_box (top-level)object of top, bottom, left, right ratios 0 to 1unsetGeometric crop applied to every page before parsing; the usual way to drop fixed headers, footers, and margin chrome.
output_options.spatial_textobjectunsetWhitespace-preserving text. do_not_unroll_columns keeps multi-column layouts side by side; preserve_layout_alignment_across_pages and preserve_very_small_text are the other flags.
output_options.additional_outputsarray[]"word_bbox" saves raw word-level boxes as JSONL (one object per word with page number and x/y/w/h), fetched via expand=["raw_words_content_metadata"].

Request expand=["items"] for the layout tree, expand=["markdown"] for per-page markdown with header and footer fields, and expand=["metadata"] for original_orientation_angle and printed_page_number.

Parse with word, line, and cell grounding, then fetch the sidecar and print the box of every word on page 1:

import json
import httpx
from llama_cloud import LlamaCloud
client = LlamaCloud() # reads LLAMA_CLOUD_API_KEY from the environment
result = client.parsing.parse(
file_id="FILE_ID", # uploaded with client.files.create(file=..., purpose="parse")
tier="agentic",
version="latest",
output_options={"granular_bboxes": ["word", "line", "cell"]},
expand=["items"],
)
sidecar = (result.result_content_metadata or {}).get("grounded_items")
rows = [json.loads(line) for line in httpx.get(sidecar.presigned_url).text.splitlines() if line.strip()]
page = next(p for p in rows if p["success"] and p["page_number"] == 1)
for item in page["items"]:
grounding = item.get("grounding")
if not grounding or "lines" not in grounding:
continue
md_bytes = item["md"].encode("utf-8") # spans are UTF-8 byte offsets
for line in grounding["lines"]:
for word in line.get("words") or []:
start, end = word["span"]
print(md_bytes[start:end].decode("utf-8"), word["bbox"])

Items come back per page in document order, each with a type, its md, and a bbox list in page points. Typical types are header, heading, text, table, and image:

{
"page_number": 1, "page_width": 612.0, "page_height": 792.0,
"items": [
{ "type": "heading", "level": 1, "md": "# NATION BY THE NUMBERS", "bbox": [{ "x": 72.0, "y": 100.0, "w": 200.0, "h": 12.0 }] },
{ "type": "table", "rows": [["Gross Costs", "$ (7,772.2)"]], "csv": "...", "html": "...", "md": "..." }
],
"success": true
}

A sidecar row has the same items plus a grounding block. For text items it holds lines[], each with a bbox and a [start, end) span into the item’s md (UTF-8 byte offsets, not character offsets), and words[] inside each line. For tables it holds rows[row][col] cells with their own boxes and spans, plus row_bboxes and column_bboxes:

{ "span": [0, 11], "bbox": { "x": 72.0, "y": 100.0, "w": 200.0, "h": 12.0 },
"words": [ { "span": [0, 5], "bbox": { "x": 72.0, "y": 100.0, "w": 35.0, "h": 12.0 } } ] }
  • Boxes are x, y, w, h in page points, in the frame given by that page’s page_width and page_height. That frame is the original viewport before orientation correction; a PDF’s /Rotate is already included.
  • r, when present, is a clockwise rotation in degrees around the box’s center, for text that is visually rotated. Apply only r. Never add metadata.pages[].original_orientation_angle to it; that angle describes how the page was turned to read it, and is 0 when no turn was needed.
  • Match metadata to a page by page_number, not by array index, especially with a page_ranges selection.
  • To draw on a screenshot, rotate each box’s corners in page units first, then scale into image pixels. Scaling before rotating distorts boxes when the horizontal and vertical scales differ.

The page markdown separates running chrome from the body: markdown.pages[].header and footer carry the detected page header and footer text, and header items appear in the tree. For chrome that sits at a fixed position, crop_box removes it before parsing instead.

Note for AI agents: this documentation is built for programmatic access. - Overview of all docs: https://developers.llamaindex.ai/llms.txt - Any page is available as raw Markdown by appending index.md to its URL — e.g. https://developers.llamaindex.ai/llamaparse/parse/getting_started/index.md - Agent-friendly REST search APIs live under https://developers.llamaindex.ai/api/ — search (BM25 full-text), grep (regex), read (fetch a page), and list (browse the doc tree). See https://developers.llamaindex.ai/llms.txt for parameters. - A hosted documentation MCP server is available at https://developers.llamaindex.ai/mcp. If you support MCP, you can ask the user to install it for browsing these docs directly (an alternative to the REST API). Setup: https://developers.llamaindex.ai/for-agents/mcp/ - Other LlamaIndex tooling for agents — the LlamaParse Platform MCP server, agent skills and plugins, and the n8n node — is mapped at https://developers.llamaindex.ai/for-agents/