Layout and bounding boxes
How Parse returns page layout as typed items with bounding boxes, adds per-word, per-line, and per-cell boxes with granular_bboxes, and the coordinate conventions for drawing those boxes on a page.
Every Parse job on the current version of its tier can return an items tree: an ordered list of typed elements per page (headings, text, tables, images, headers), each with its markdown and an item-level bounding box. For citation highlighting and grounding, output_options.granular_bboxes adds boxes per word, per line, and per table cell, delivered as a JSONL sidecar. The page markdown also separates out running headers and footers, and crop_box strips page chrome before parsing.
When to use it
Section titled “When to use it”- Highlighting the exact span a citation points to in a PDF viewer.
- Grounding extracted answers to a word, line, or table cell rather than a whole paragraph.
- A side-by-side preview that syncs markdown text to the region on the source page.
- Dropping repeated headers and footers, or keeping them separate from body text.
- Reconstructing reading order and element types for downstream chunking.
A fast job pinned to a version older than 2026-06-15 produces text only, so markdown, items and granular boxes are not available there.
Options
Section titled “Options”| Option | Type | Default | What it does |
|---|---|---|---|
output_options.granular_bboxes | array of "word", "line", "cell" | [] | Compute boxes at the chosen levels. Written to a grounded_items JSONL sidecar whose download URL is auto-included on the result; nothing to add to expand. |
output_options.images_to_save | array | saves layout when the output links to cropped images | Add "screenshot" to get a full-page render to draw boxes on. Retrieve with expand=["images_content_metadata"]. |
output_options.extract_printed_page_number | boolean | unset | Return the page number as printed on the page (v, A-3) in metadata.pages[].printed_page_number, for citations. |
crop_box (top-level) | object of top, bottom, left, right ratios 0 to 1 | unset | Geometric crop applied to every page before parsing; the usual way to drop fixed headers, footers, and margin chrome. |
output_options.spatial_text | object | unset | Whitespace-preserving text. do_not_unroll_columns keeps multi-column layouts side by side; preserve_layout_alignment_across_pages and preserve_very_small_text are the other flags. |
output_options.additional_outputs | array | [] | "word_bbox" saves raw word-level boxes as JSONL (one object per word with page number and x/y/w/h), fetched via expand=["raw_words_content_metadata"]. |
Request expand=["items"] for the layout tree, expand=["markdown"] for per-page markdown with header and footer fields, and expand=["metadata"] for original_orientation_angle and printed_page_number.
Example
Section titled “Example”Parse with word, line, and cell grounding, then fetch the sidecar and print the box of every word on page 1:
import json
import httpxfrom llama_cloud import LlamaCloud
client = LlamaCloud() # reads LLAMA_CLOUD_API_KEY from the environment
result = client.parsing.parse( file_id="FILE_ID", # uploaded with client.files.create(file=..., purpose="parse") tier="agentic", version="latest", output_options={"granular_bboxes": ["word", "line", "cell"]}, expand=["items"],)
sidecar = (result.result_content_metadata or {}).get("grounded_items")rows = [json.loads(line) for line in httpx.get(sidecar.presigned_url).text.splitlines() if line.strip()]
page = next(p for p in rows if p["success"] and p["page_number"] == 1)for item in page["items"]: grounding = item.get("grounding") if not grounding or "lines" not in grounding: continue md_bytes = item["md"].encode("utf-8") # spans are UTF-8 byte offsets for line in grounding["lines"]: for word in line.get("words") or []: start, end = word["span"] print(md_bytes[start:end].decode("utf-8"), word["bbox"])What you get
Section titled “What you get”Items come back per page in document order, each with a type, its md, and a bbox list in page points. Typical types are header, heading, text, table, and image:
{ "page_number": 1, "page_width": 612.0, "page_height": 792.0, "items": [ { "type": "heading", "level": 1, "md": "# NATION BY THE NUMBERS", "bbox": [{ "x": 72.0, "y": 100.0, "w": 200.0, "h": 12.0 }] }, { "type": "table", "rows": [["Gross Costs", "$ (7,772.2)"]], "csv": "...", "html": "...", "md": "..." } ], "success": true}A sidecar row has the same items plus a grounding block. For text items it holds lines[], each with a bbox and a [start, end) span into the item’s md (UTF-8 byte offsets, not character offsets), and words[] inside each line. For tables it holds rows[row][col] cells with their own boxes and spans, plus row_bboxes and column_bboxes:
{ "span": [0, 11], "bbox": { "x": 72.0, "y": 100.0, "w": 200.0, "h": 12.0 }, "words": [ { "span": [0, 5], "bbox": { "x": 72.0, "y": 100.0, "w": 35.0, "h": 12.0 } } ] }Coordinate conventions
Section titled “Coordinate conventions”- Boxes are
x,y,w,hin page points, in the frame given by that page’spage_widthandpage_height. That frame is the original viewport before orientation correction; a PDF’s/Rotateis already included. r, when present, is a clockwise rotation in degrees around the box’s center, for text that is visually rotated. Apply onlyr. Never addmetadata.pages[].original_orientation_angleto it; that angle describes how the page was turned to read it, and is0when no turn was needed.- Match metadata to a page by
page_number, not by array index, especially with apage_rangesselection. - To draw on a screenshot, rotate each box’s corners in page units first, then scale into image pixels. Scaling before rotating distorts boxes when the horizontal and vertical scales differ.
The page markdown separates running chrome from the body: markdown.pages[].header and footer carry the detected page header and footer text, and header items appear in the tree. For chrome that sits at a fixed position, crop_box removes it before parsing instead.
See also
Section titled “See also”- Granular bounding boxes example: the sidecar schema, table-cell grounding, and a complete screenshot-rendering helper
- Response format: grounded items sidecar, items pages and bounding boxes
- Configuring Parse: granular bounding boxes, crop box, and spatial text
- Forms and checkboxes for per-field boxes on form pages