Skip to content
Guide
Parse
Features

OCR and languages

How Parse reads scanned pages, phone photos, and image-only documents, with the OCR language list, camera-photo correction, and which Parse tier to pick for scans.

Parse runs OCR automatically on any page that has no text layer: scanned PDFs, photographed receipts and forms, image files (jpg, png, tiff, webp, heic, and more), and images embedded inside born-digital documents. Text that already exists in a PDF is read directly. Most scans need no configuration at all. Three options matter when they do: the OCR language list, camera-photo correction for phone pictures, and the tier, because fast runs no AI model over the page.

  • A scanned contract, claim, lab report, or intake form where every page is an image.
  • Photos of documents taken with a phone, which arrive tilted, surrounded by background, and unevenly lit.
  • Documents in a language other than English, or with mixed-language content, where OCR should know which languages to expect.
  • Born-digital PDFs with embedded images whose text you want captured, or deliberately skipped.

The Parse overview lists scanned pages at odd angles and handwritten annotations among the layouts Parse is built to handle. Image quality still bounds what OCR can recover from a scan, and an image-only file whose pages yield no text fails with NO_DATA_FOUND_IN_FILE (see Troubleshooting).

OptionTypeDefaultWhat it does
processing_options.ocr_parameters.languagesarray of language codesunsetLanguages OCR should expect. Order matters: put the primary language first. Only affects text read from images; native PDF text is unaffected.
input_options.image.camera_photo_correctionbooleanunsetFor jpg, png, webp, and heic/heif inputs, detects a photographed document, crops it, corrects perspective, and flattens uneven lighting and shadows before parsing. Clean scans and screenshots are left untouched, so it is safe on mixed batches.
processing_options.ignore.ignore_text_in_imagebooleanunsetSkip OCR text from embedded images, for example logos or watermark graphics inside an otherwise digital document. It turns OCR off for the job, so use it only on born-digital files; a scanned page would come back without its text.
processing_control.job_failure_conditions.fail_on_image_ocr_errorbooleanunsetFail the whole job if OCR fails on any image. By default an OCR failure yields empty text for that image and the job continues.
output_options.spatial_textobjectunsetWhitespace-preserving text output for scanned forms and receipts where position carries meaning. Flags: preserve_layout_alignment_across_pages, preserve_very_small_text, do_not_unroll_columns. Retrieve with expand=["text"].

Language codes follow the ParserLanguages enum in the API reference: more than 80 values, including en, fr, de, es, it, pt, nl, ru, ar, hi, ja, ko, ch_sim, and ch_tra.

InputTier
Scanned pages, photos, multi-column scansagentic, the default recommendation; agentic_plus for dense or mission-critical scans
Plain scanned text where raw text or spatial text is enoughfast, the cheapest tier. It runs text extraction and OCR with no AI pass, so it does not accept agentic_options or processing_options.forms
Mixed documents where only some pages are scansagentic or agentic_plus with Cost Optimizer

Parse a photographed French invoice, telling OCR to expect French first and English second:

from llama_cloud import LlamaCloud
client = LlamaCloud() # reads LLAMA_CLOUD_API_KEY from the environment
result = client.parsing.parse(
file_id="FILE_ID", # uploaded with client.files.create(file=..., purpose="parse")
tier="agentic",
version="latest",
input_options={"image": {"camera_photo_correction": True}},
processing_options={"ocr_parameters": {"languages": ["fr", "en"]}},
expand=["markdown", "text", "metadata"],
)
print(result.markdown.pages[0].markdown)

input_options.image only applies to image files; it is ignored for a PDF. The languages list is used on every page that goes through OCR, whatever the file type.

OCR output looks the same as a born-digital parse: markdown per page with headings and tables, and text per page. The text view keeps the page’s spacing, which is why it suits receipts and forms. This excerpt is the first page of the Quick Start report in the text view:

1 EXECUTIVE SUMMARY TO THE FY 2024 FINANCIAL REPORT OF THE U.S. GOVERNMENT
NATION BY THE NUMBERS
Net Cost:
Gross Costs $ (7,772.2) $ (7,661.7)
Less: Earned Revenue 652.9 $ 539.5

Two per-page metadata fields are useful on scans. original_orientation_angle records, in degrees, the rotation Parse applied to read a crooked page (0 when none was needed), and confidence (0 to 1) is the parser’s own estimate for the page, so you can route low-confidence scans to a human:

{ "page_number": 1, "confidence": 0.985, "original_orientation_angle": 0, "watermark": "CONFIDENTIAL" }

Watermark text stamped across a page is reported in watermark on version 2026-09-28 and later of the cost_effective, agentic, and agentic_plus tiers; output_options.watermark_handling (move_to_end, move_to_start, or remove) controls where it goes in the markdown.

Note for AI agents: this documentation is built for programmatic access. - Overview of all docs: https://developers.llamaindex.ai/llms.txt - Any page is available as raw Markdown by appending index.md to its URL — e.g. https://developers.llamaindex.ai/llamaparse/parse/getting_started/index.md - Agent-friendly REST search APIs live under https://developers.llamaindex.ai/api/ — search (BM25 full-text), grep (regex), read (fetch a page), and list (browse the doc tree). See https://developers.llamaindex.ai/llms.txt for parameters. - A hosted documentation MCP server is available at https://developers.llamaindex.ai/mcp. If you support MCP, you can ask the user to install it for browsing these docs directly (an alternative to the REST API). Setup: https://developers.llamaindex.ai/for-agents/mcp/ - Other LlamaIndex tooling for agents — the LlamaParse Platform MCP server, agent skills and plugins, and the n8n node — is mapped at https://developers.llamaindex.ai/for-agents/