Skip to content
Guide
Parse
Features

Tables

How Parse extracts tables from PDFs, scans, and images into markdown, HTML, structured rows, CSV, or an XLSX file, with the options for merged cells, multi-page tables, and borderless tables.

Parse recovers tables with their cell structure intact and returns each one in several forms at once: as part of the page markdown (pipe table or HTML), and as a typed table item in the items tree with rows, csv, html, and md representations. You choose the format with a few output options, and you can also have Parse write every table into an XLSX workbook.

  • Financial reports, 10-Ks, earnings decks, and audit reports with many tables across many pages.
  • Tables with merged header cells, which markdown pipe tables cannot represent.
  • Tables that continue across page breaks and should come back as one table.
  • Borderless or lightly formatted tables that a text-based parser misses.
  • Loading table data into pandas, a warehouse, or a spreadsheet rather than reading it as prose.

Pick cost_effective for mostly-text documents with simple tables, agentic for real tables and scanned pages, and agentic_plus for dense financial reports and complex tables. The fast tier runs no AI model, so pick cost_effective or higher when tables have merged cells, continue across pages, or lack borders.

OptionTypeDefaultWhat it does
output_options.markdown.tables.output_tables_as_markdownbooleanunsettrue emits markdown pipe tables; false keeps HTML <table> tags. Pipe tables are simpler but cannot represent merged cells; HTML can, with colspan.
output_options.markdown.tables.merge_continued_tablesbooleanunsetMerge a table that spans several pages into one. The merged table appears on the first page with merged_from_pages metadata.
output_options.markdown.tables.compact_markdown_tablesbooleanunsetRemove whitespace padding inside markdown table cells.
output_options.markdown.tables.markdown_table_multiline_separatorstringunsetSeparator for multi-line cell content in markdown tables, for example "<br>" to keep line breaks or " " to join.
output_options.tables_as_spreadsheet.enablebooleanunsetAlso write every table into an XLSX file, one sheet per table. Retrieve with expand=["xlsx_content_metadata"].
output_options.tables_as_spreadsheet.guess_sheet_namebooleantrueName each sheet from the table’s headers and surrounding text instead of Table_1.
processing_options.aggressive_table_extractionbooleanunsetTry harder to find table boundaries, including tables without visible borders. May add false positives.
processing_options.disable_heuristicsbooleanunsetTurn off outlined-table extraction and adaptive long-table handling when they produce wrong results.
output_options.granular_bboxesarray of "cell", "line", "word"[]Add a bounding box per table cell for highlighting. See Layout and bounding boxes.

Parse a report, collect every table from the items tree, and load each one into pandas from its csv:

import io
import pandas as pd
from llama_cloud import LlamaCloud
client = LlamaCloud() # reads LLAMA_CLOUD_API_KEY from the environment
result = client.parsing.parse(
file_id="FILE_ID", # uploaded with client.files.create(file=..., purpose="parse")
tier="agentic",
version="latest",
output_options={"markdown": {"tables": {"merge_continued_tables": True}}},
expand=["markdown", "items"],
)
tables = [
(page.page_number, item)
for page in result.items.pages
if page.success # a failed page carries no items
for item in page.items
if item.type == "table"
]
print(f"Found {len(tables)} tables across {len(result.items.pages)} pages")
for page_number, table in tables:
df = pd.read_csv(io.StringIO(table.csv))
print(f"page {page_number}: {len(df)} rows x {len(df.columns)} cols")

In the page markdown, a default agentic parse of the Quick Start report returns the summary table as HTML, keeping its merged header row:

<table>
<thead>
<tr><th colspan="3">Financial Measures (Dollars in Billions):</th></tr>
<tr><th></th><th>2024</th><th>2023*</th></tr>
</thead>
<tbody>
<tr><td>Gross Costs</td><td>$ (7,772.2)</td><td>$ (7,661.7)</td></tr>
<tr><td>Less: Earned Revenue</td><td>$ 652.9</td><td>$ 539.5</td></tr>
</tbody>
</table>

In the items tree, the same table is a typed item with the data in four forms:

{
"type": "table",
"rows": [["Financial Measures (Dollars in Billions):", "2024", "2023*"], ["Gross Costs", "$ (7,772.2)", "$ (7,661.7)"]],
"csv": "Financial Measures (Dollars in Billions):,2024,2023*\nGross Costs,$ (7,772.2),$ (7,661.7)",
"html": "<table>...</table>",
"md": "| Financial Measures (Dollars in Billions): | 2024 | 2023* |\n|---|---|---|..."
}

rows is a list of lists, csv loads straight into pandas.read_csv, and html and md are ready to display. With tables_as_spreadsheet enabled, expand=["xlsx_content_metadata"] adds a presigned download URL for the workbook.

Note for AI agents: this documentation is built for programmatic access. - Overview of all docs: https://developers.llamaindex.ai/llms.txt - Any page is available as raw Markdown by appending index.md to its URL — e.g. https://developers.llamaindex.ai/llamaparse/parse/getting_started/index.md - Agent-friendly REST search APIs live under https://developers.llamaindex.ai/api/ — search (BM25 full-text), grep (regex), read (fetch a page), and list (browse the doc tree). See https://developers.llamaindex.ai/llms.txt for parameters. - A hosted documentation MCP server is available at https://developers.llamaindex.ai/mcp. If you support MCP, you can ask the user to install it for browsing these docs directly (an alternative to the REST API). Setup: https://developers.llamaindex.ai/for-agents/mcp/ - Other LlamaIndex tooling for agents — the LlamaParse Platform MCP server, agent skills and plugins, and the n8n node — is mapped at https://developers.llamaindex.ai/for-agents/