Skip to content
Guide
Extract
Guides

Metadata Extensions

Advanced extraction features including citations and confidence scores for enhanced data extraction workflows.

LlamaExtract offers several advanced features that provide additional metadata and insights alongside your extracted data. These extensions are available under Advanced Settings in the UI and return schema-level metadata in the extract_metadata field of the response.

Citations provide the source information for every extracted field, allowing you to trace back exactly where each piece of data came from in the original document.

How it works: For every leaf-level field in your schema, citations return:

  • The page number where the information was found
  • The verbatim text that was used to extract the field value
  • Bounding box coordinates (x, y, w, h) indicating the exact location of the cited text on the page
  • Page dimensions (width, height) to help you render the bounding boxes accurately

Citations appear in the API response under extract_metadata.field_metadata.document_metadata and in the LlamaCloud UI. See Extract response format for the field reference.

Example API response structure (scalar fields):

"extract_metadata": {
"field_metadata": {
"document_metadata": {
"phone": {
"citation": [
{
"page": 1,
"matching_text": "(555) 123-4567",
"bounding_boxes": [
{
"x": 177,
"y": 82,
"w": 318,
"h": 43
}
],
"page_dimensions": {
"width": 612,
"height": 792
}
}
]
}
}
}
}

Array fields: Citations attach at the leaf sub-field level, not the array item level. The field_metadata tree mirrors the structure of your extracted data, with each leaf value replaced by its citation metadata.

For a schema like key_facts: list[KeyFact] where KeyFact has a fact: str field, the metadata structure is:

"extract_metadata": {
"field_metadata": {
"document_metadata": {
"key_facts": [
{
"fact": {
"citation": [
{
"page": 3,
"matching_text": "Revenue grew 114% year-over-year",
"bounding_boxes": [{ "x": 50, "y": 200, "w": 400, "h": 20 }],
"page_dimensions": { "width": 612, "height": 792 }
}
]
}
},
{
"fact": {
"citation": [
{
"page": 7,
"matching_text": "Operating expenses increased to $3.2B",
"bounding_boxes": [{ "x": 50, "y": 310, "w": 380, "h": 20 }],
"page_dimensions": { "width": 612, "height": 792 }
}
]
}
}
]
}
}
}

Note: the citation path is field_metadata.document_metadata.key_facts[i].fact.citation, not field_metadata.document_metadata.key_facts[i].citation. Each array element in the metadata corresponds positionally to the same element in the extracted data.

Usage: Set cite_sources: true in the configuration to enable this feature.

Use cases:

  • Compliance and audit requirements
  • Fact-checking and verification workflows
  • Understanding extraction quality and accuracy
  • Building custom highlighting/annotation features using bounding box coordinates

confidence is the overall score for a field value, from 0 to 1. Higher scores indicate greater confidence that the value is correct. Each field has its own score, so you can send uncertain values to human review while accepting the rest of a document. A high score is still an estimate, not a guarantee of correctness.

Set confidence_scores: true in the configuration to include scores in extract_metadata.field_metadata. When retrieving a job, request expand=extract_metadata to include that metadata in the response. See Extract response format.

The metadata follows the same nested structure as your schema. For example, a document with a total field and an items array could return:

{
"extract_metadata": {
"field_metadata": {
"document_metadata": {
"total": {"confidence": 0.97},
"items": [
{"amount": {"confidence": 0.62}}
]
}
}
}
}

The score for the first item’s amount is at extract_metadata.field_metadata.document_metadata.items[0].amount.confidence. With a review threshold of 0.8, your application would send that amount to review and accept the total. The threshold is your application’s decision.

Choose a review threshold using a sample of your own documents with known correct values. Start with a candidate threshold, then check both the fields sent to review and those accepted automatically. Count how many errors the threshold catches, how many incorrect values pass through, and how much review it creates. Raise the threshold to review more fields, or lower it to reduce review volume.

Use stricter thresholds for fields where an incorrect value is costly. Long free-text fields, such as summaries and descriptions, can score lower than short factual fields. Evaluate them separately rather than assuming one cutoff works for every field type.

Calibration describes how closely scores match observed correctness. Calibration varies by tier and version, so a score of 0.8 does not always mean an 80% probability of correctness. Validate your cutoff on the tier and version you use, and check it again when you change versions or start processing different kinds of documents.

Reasoning metadata is available for Extract versions through 2026-03-31. Newer versions do not return per-field reasoning strings.

If your application depends on reasoning strings in extract_metadata.field_metadata, pin configuration.version to 2026-03-31. For newer versions, use citations and confidence scores when you need provenance or review signals.

⚠️ Important: Citations and confidence scores will significantly slow down extraction processing time. Enable these features only when the additional metadata is essential for your use case.

For complete examples of how to configure and use these extensions with both the Python SDK and REST API, see the Configuring Extract page.

The configuration section includes:

  • Complete Python SDK examples with extension settings
  • REST API curl command examples
  • Configuration reference table with all available options

Quick reference for extensions:

import time
from llama_cloud import LlamaCloud
client = LlamaCloud(api_key="your_api_key")
file_obj = client.files.create(file="path/to/your/document.pdf", purpose="extract")
file_id = file_obj.id
job = client.extract.create(
file_input=file_id,
configuration={
"data_schema": {"type": "object", "properties": {}}, # your extraction schema
"tier": "agentic",
"cite_sources": True,
"confidence_scores": True,
},
)
# Poll for completion
while job.status not in ("COMPLETED", "FAILED", "CANCELLED"):
time.sleep(2)
job = client.extract.get(job.id)
Note for AI agents: this documentation is built for programmatic access. - Overview of all docs: https://developers.llamaindex.ai/llms.txt - Any page is available as raw Markdown by appending index.md to its URL — e.g. https://developers.llamaindex.ai/llamaparse/parse/getting_started/index.md - Agent-friendly REST search APIs live under https://developers.llamaindex.ai/api/ — search (BM25 full-text), grep (regex), read (fetch a page), and list (browse the doc tree). See https://developers.llamaindex.ai/llms.txt for parameters. - A hosted documentation MCP server is available at https://developers.llamaindex.ai/mcp. If you support MCP, you can ask the user to install it for browsing these docs directly (an alternative to the REST API). Setup: https://developers.llamaindex.ai/for-agents/mcp/ - Other LlamaIndex tooling for agents — the LlamaParse Platform MCP server, agent skills and plugins, and the n8n node — is mapped at https://developers.llamaindex.ai/for-agents/