---
title: Extract response format | Developer Documentation
description: Reference for Extract v2 job responses, extracted data, citations, confidence scores, and usage credits.
---

Extract jobs return `extract_result` with data matching your schema. Request `extract_metadata` for citations and confidence scores, and `usage` for billed credits. See the [Extract API reference](https://developers.llamaindex.ai/reference/resources/extract/) for the complete job object.

## Job envelope

| Field                      | Type                   | Description                                                                                                                           |
| -------------------------- | ---------------------- | ------------------------------------------------------------------------------------------------------------------------------------- |
| `id`                       | string                 | Job identifier, prefixed `ext-`. Pass it to `GET /api/v2/extract/{job_id}`.                                                           |
| `status`                   | string                 | Job state. See [Job status](#job-status).                                                                                             |
| `file_input`               | string                 | The file ID (`dfl-...`) or parse job ID (`pjb-...`) the job ran on.                                                                   |
| `project_id`               | string                 | Project the job belongs to.                                                                                                           |
| `created_at`, `updated_at` | datetime               | Creation and last-update timestamps.                                                                                                  |
| `error_message`            | string or null         | Error details when `status` is `FAILED`.                                                                                              |
| `extract_result`           | object, array, or null | The extracted data, shaped by your `data_schema`. See [extract\_result](#extract_result).                                             |
| `extract_metadata`         | object or null         | Citations, confidence scores, and parse details. Requires `expand=extract_metadata`. See [extract\_metadata](#extract_metadata).      |
| `configuration`            | object or null         | The configuration the job ran with. On retrieval, requires `expand=configuration`. See [Saved configurations](#saved-configurations). |
| `configuration_id`         | string or null         | Saved configuration ID used for the job, if any.                                                                                      |
| `usage`                    | object or null         | Credits billed against the job on retrieval. Requires `expand=usage`. See [Usage and credits](#usage-and-credits).                    |

### Expanding the response

`GET /api/v2/extract/{job_id}` returns `extract_result` by default and omits the heavier fields. Add one or more `expand` values to include them. The create and cancel responses take no `expand` and always include `configuration`.

| `expand` value     | Adds                                        |
| ------------------ | ------------------------------------------- |
| `extract_metadata` | Per-field citations and confidence scores   |
| `configuration`    | The resolved configuration the job ran with |
| `usage`            | Credits billed against the job              |

The list endpoint, `GET /api/v2/extract`, accepts `expand=configuration` and `expand=extract_metadata` and returns `items`, a `next_page_token` for the next page, and an optional `total_size`.

## Job status

| `status`    | Meaning                                         |
| ----------- | ----------------------------------------------- |
| `PENDING`   | Queued, not yet started.                        |
| `RUNNING`   | Actively processing.                            |
| `COMPLETED` | Finished; `extract_result` is populated.        |
| `FAILED`    | Terminated with an error; read `error_message`. |
| `CANCELLED` | Cancelled by the user.                          |

Poll until the status is one of the three terminal values, or let `client.extract.wait_for_completion(job.id)` do it for you.

## extract\_result

`extract_result` contains the fields and value types defined by your `data_schema`. Extract applies the schema to the document and returns one object. Arrays in your schema remain arrays inside that object.

```
{
  "company_name": "NVIDIA Corporation",
  "fiscal_year": 2025,
  "revenue": 130497
}
```

For missing values, see [Required and optional fields](/llamaparse/extract/guides/configuring-extract/#required-and-optional-fields/index.md).

## extract\_metadata

`extract_metadata` requires `expand=extract_metadata`. Enable `cite_sources` or `confidence_scores` in the configuration to include those fields.

| Field            | Description                                              |
| ---------------- | -------------------------------------------------------- |
| `field_metadata` | Per-field citations and confidence scores.               |
| `parse_job_id`   | ID of the Parse job used for extraction, when available. |
| `parse_tier`     | Tier of that Parse job, when available.                  |

### field\_metadata

Read field metadata from `field_metadata.document_metadata`. The metadata follows your schema, including nested objects and arrays. For example, metadata for `items[0].amount` is at `document_metadata.items[0].amount`.

```
{
  "field_metadata": {
    "document_metadata": {
      "company_name": {
        "citation": [{"page": 1, "matching_text": "NVIDIA CORPORATION"}],
        "confidence": 0.99,
        "parsing_confidence": 1.0,
        "extraction_confidence": 0.99
      }
    }
  }
}
```

### Citations

Set `cite_sources: true` to return a `citation` array for each field. Each citation identifies a source location:

| Key               | Description                              |
| ----------------- | ---------------------------------------- |
| `page`            | 1-based page number.                     |
| `matching_text`   | Source text for the extracted value.     |
| `bounding_boxes`  | Locations on the page as `{x, y, w, h}`. |
| `page_dimensions` | Page size as `{width, height}`.          |

Turbo citations include `page` and `matching_text` only. Spreadsheet mode does not return citations or confidence scores. See [Citations](/llamaparse/extract/guides/extensions/#citations/index.md).

### Confidence scores

Set `confidence_scores: true` to return `confidence`, `parsing_confidence`, and `extraction_confidence`. Use `confidence` as the combined score from 0 to 1. See [Confidence scores](/llamaparse/extract/guides/extensions/#confidence-scores/index.md) for choosing a threshold.

## Saved configurations

`configuration_id` identifies the saved configuration used for the job. It is `null` for inline configurations. Request `expand=configuration` to retrieve the resolved parameters, including the schema, tier, extensions, and concrete version. See [Using saved configurations](/llamaparse/extract/examples/using_saved_configurations/index.md).

## Usage and credits

`usage`, returned with `expand=usage`, describes what the job consumed:

| Key               | Description                                                                                                                                                                            |
| ----------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `credits`         | Total credits billed against the job: the sum of the two components below, counting only those already recorded.                                                                       |
| `extract_credits` | Credits billed for the extraction itself.                                                                                                                                              |
| `parse_credits`   | Credits billed against the Parse job that Extract created for this job. `null` when `file_input` was a Parse job you created yourself, because those credits belong to that Parse job. |

`usage` is `null` until completion. Credit values can remain `null` until billing records them. Poll again for pending values. When several Extract jobs use the same Parse job, count `parse_credits` once when totaling usage. See [Check the credits a job billed](/llamaparse/extract/api/#5-check-the-credits-a-job-billed/index.md).

## Fetch the full response

For a completed job, request all three expansions with its job ID:

```
import os
from llama_cloud import LlamaCloud


client = LlamaCloud(api_key=os.environ["LLAMA_CLOUD_API_KEY"])
job = client.extract.get("JOB_ID", expand=["extract_metadata", "configuration", "usage"])


print(job.extract_result)
print(job.extract_metadata.field_metadata.document_metadata)
print(job.usage)
```
