OCR and languages
How Parse reads scanned pages, phone photos, and image-only documents, with the OCR language list, camera-photo correction, and which Parse tier to pick for scans.
Parse runs OCR automatically on any page that has no text layer: scanned PDFs, photographed receipts and forms, image files (jpg, png, tiff, webp, heic, and more), and images embedded inside born-digital documents. Text that already exists in a PDF is read directly. Most scans need no configuration at all. Three options matter when they do: the OCR language list, camera-photo correction for phone pictures, and the tier, because fast runs no AI model over the page.
When to use it
Section titled “When to use it”- A scanned contract, claim, lab report, or intake form where every page is an image.
- Photos of documents taken with a phone, which arrive tilted, surrounded by background, and unevenly lit.
- Documents in a language other than English, or with mixed-language content, where OCR should know which languages to expect.
- Born-digital PDFs with embedded images whose text you want captured, or deliberately skipped.
The Parse overview lists scanned pages at odd angles and handwritten annotations among the layouts Parse is built to handle. Image quality still bounds what OCR can recover from a scan, and an image-only file whose pages yield no text fails with NO_DATA_FOUND_IN_FILE (see Troubleshooting).
Options
Section titled “Options”| Option | Type | Default | What it does |
|---|---|---|---|
processing_options.ocr_parameters.languages | array of language codes | unset | Languages OCR should expect. Order matters: put the primary language first. Only affects text read from images; native PDF text is unaffected. |
input_options.image.camera_photo_correction | boolean | unset | For jpg, png, webp, and heic/heif inputs, detects a photographed document, crops it, corrects perspective, and flattens uneven lighting and shadows before parsing. Clean scans and screenshots are left untouched, so it is safe on mixed batches. |
processing_options.ignore.ignore_text_in_image | boolean | unset | Skip OCR text from embedded images, for example logos or watermark graphics inside an otherwise digital document. It turns OCR off for the job, so use it only on born-digital files; a scanned page would come back without its text. |
processing_control.job_failure_conditions.fail_on_image_ocr_error | boolean | unset | Fail the whole job if OCR fails on any image. By default an OCR failure yields empty text for that image and the job continues. |
output_options.spatial_text | object | unset | Whitespace-preserving text output for scanned forms and receipts where position carries meaning. Flags: preserve_layout_alignment_across_pages, preserve_very_small_text, do_not_unroll_columns. Retrieve with expand=["text"]. |
Language codes follow the ParserLanguages enum in the API reference: more than 80 values, including en, fr, de, es, it, pt, nl, ru, ar, hi, ja, ko, ch_sim, and ch_tra.
Which tier to pick for scans
Section titled “Which tier to pick for scans”| Input | Tier |
|---|---|
| Scanned pages, photos, multi-column scans | agentic, the default recommendation; agentic_plus for dense or mission-critical scans |
| Plain scanned text where raw text or spatial text is enough | fast, the cheapest tier. It runs text extraction and OCR with no AI pass, so it does not accept agentic_options or processing_options.forms |
| Mixed documents where only some pages are scans | agentic or agentic_plus with Cost Optimizer |
Example
Section titled “Example”Parse a photographed French invoice, telling OCR to expect French first and English second:
from llama_cloud import LlamaCloud
client = LlamaCloud() # reads LLAMA_CLOUD_API_KEY from the environment
result = client.parsing.parse( file_id="FILE_ID", # uploaded with client.files.create(file=..., purpose="parse") tier="agentic", version="latest", input_options={"image": {"camera_photo_correction": True}}, processing_options={"ocr_parameters": {"languages": ["fr", "en"]}}, expand=["markdown", "text", "metadata"],)
print(result.markdown.pages[0].markdown)input_options.image only applies to image files; it is ignored for a PDF. The languages list is used on every page that goes through OCR, whatever the file type.
What you get
Section titled “What you get”OCR output looks the same as a born-digital parse: markdown per page with headings and tables, and text per page. The text view keeps the page’s spacing, which is why it suits receipts and forms. This excerpt is the first page of the Quick Start report in the text view:
1 EXECUTIVE SUMMARY TO THE FY 2024 FINANCIAL REPORT OF THE U.S. GOVERNMENT
NATION BY THE NUMBERS Net Cost: Gross Costs $ (7,772.2) $ (7,661.7) Less: Earned Revenue 652.9 $ 539.5Two per-page metadata fields are useful on scans. original_orientation_angle records, in degrees, the rotation Parse applied to read a crooked page (0 when none was needed), and confidence (0 to 1) is the parser’s own estimate for the page, so you can route low-confidence scans to a human:
{ "page_number": 1, "confidence": 0.985, "original_orientation_angle": 0, "watermark": "CONFIDENTIAL" }Watermark text stamped across a page is reported in watermark on version 2026-09-28 and later of the cost_effective, agentic, and agentic_plus tiers; output_options.watermark_handling (move_to_end, move_to_start, or remove) controls where it goes in the markdown.
See also
Section titled “See also”- Configuring Parse: OCR languages and camera photos
- Recipes: multi-language OCR in Python, TypeScript, Go, Java, and the CLI
- Tiers for the full comparison, including what
fastdoes and does not return - Supported document types for the image and scan formats Parse accepts
- Forms and checkboxes for scanned forms that need field values, not just text