---
title: Split response format | Developer Documentation
description: Field-by-field reference for the Split job response — the job envelope and statuses, result segments with category, 1-indexed page numbers and confidence level, uncategorized-page handling, the beta endpoint differences, and how to feed segments into Parse, Extract, or Classify.
---

A Split job returns one object, the split job, from `POST /api/v1/split/jobs`, `GET /api/v1/split/jobs/{split_job_id}`, and the cancel endpoint. The SDKs return the same object from `client.split.create()` and `client.split.get()`. This page lists every field in that object and what each one means. For defining categories and running jobs, see [Getting started](/llamaparse/split/getting_started/index.md); for the request side, see the [Split API reference](https://developers.llamaindex.ai/reference/resources/split/).

Split is currently in beta and is subject to breaking changes.

## Job envelope

| Field                      | Type           | Description                                                                                                                                                                      |
| -------------------------- | -------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `id`                       | string         | Job identifier, prefixed `spl-`. Pass it to `GET /api/v1/split/jobs/{split_job_id}`.                                                                                             |
| `status`                   | string         | Job state. See [Job status](#job-status).                                                                                                                                        |
| `file_input`               | string         | The file ID or parse job ID the job ran on.                                                                                                                                      |
| `document_input_type`      | string         | What `file_input` refers to: `file_id`, `parse_job_id`, or `url`.                                                                                                                |
| `categories`               | array          | The categories the job split against, each with `name` and an optional `description`. Echoed from the request or the saved configuration.                                        |
| `splitting_strategy`       | object         | The strategy the job ran with: `allow_uncategorized`, `custom_instructions`, and `min_pages_per_split`. See [Uncategorized pages and merging](#uncategorized-pages-and-merging). |
| `parse_tier`               | string or null | Parse tier requested for the job. `null` means the default, `fast`. Not used when `file_input` is a completed Parse job.                                                         |
| `parse_config_id`          | string or null | Saved parse configuration ID requested for the job, if any. Takes precedence over `parse_tier`.                                                                                  |
| `target_pages`             | string or null | Page selection requested for the job (1-based, `"1-3,7"`), if any. Accepted only when `file_input` is a completed Parse job.                                                     |
| `configuration_id`         | string or null | Saved configuration ID used for the job, if any.                                                                                                                                 |
| `result`                   | object or null | The segments. Present only when `status` is `completed`. See [result.segments](#resultsegments).                                                                                 |
| `error_message`            | string or null | Why the job failed, when `status` is `failed`.                                                                                                                                   |
| `transaction_id`           | string or null | The idempotency key supplied at creation, if any. Reusing a key returns the original job.                                                                                        |
| `project_id`               | string         | Project the job belongs to.                                                                                                                                                      |
| `user_id`                  | string         | User who created the job.                                                                                                                                                        |
| `created_at`, `updated_at` | datetime       | Creation and last-update timestamps.                                                                                                                                             |

## Job status

Split statuses are lowercase, unlike the uppercase statuses of Extract and Classify jobs.

| `status`     | Meaning                                                           |
| ------------ | ----------------------------------------------------------------- |
| `pending`    | Queued, not yet started.                                          |
| `processing` | Actively processing.                                              |
| `completed`  | Finished; `result` is populated.                                  |
| `failed`     | Terminated with an error; `result` is null, read `error_message`. |
| `cancelled`  | Cancelled by the user; `result` is null.                          |

Poll `GET /api/v1/split/jobs/{split_job_id}` until the status is one of the three terminal values.

## result.segments

`result` has a single field, `segments`, an array with one element per detected document section. Each segment has:

| Field                 | Type              | Description                                                                                                                                 |
| --------------------- | ----------------- | ------------------------------------------------------------------------------------------------------------------------------------------- |
| `category`            | string            | The `name` of the category the segment was assigned to, or `uncategorized` when no category matched and `allow_uncategorized` is `include`. |
| `pages`               | array of integers | The 1-indexed page numbers that make up the segment.                                                                                        |
| `confidence_category` | string            | Confidence in the assignment: `high`, `medium`, or `low`.                                                                                   |

Page numbers are those of the input document, so with `target_pages` set they still refer to the original pages. Split detects where one document ends and the next begins, so adjacent segments can share a category: a bundle of three invoices in a row comes back as three `invoice` segments, not one. From [Getting started](/llamaparse/split/getting_started/#get-the-results/index.md):

```
{
  "id": "spl-abc123...",
  "status": "completed",
  "result": {
    "segments": [
      {
        "category": "invoice",
        "pages": [1, 2, 3],
        "confidence_category": "high"
      },
      {
        "category": "contract",
        "pages": [4, 5, 6, 7, 8],
        "confidence_category": "high"
      }
    ]
  }
}
```

## Uncategorized pages and merging

Three `splitting_strategy` options change which pages appear in `result.segments`:

| Option                | Values               | Effect on the result                                                                                                                               |
| --------------------- | -------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------- |
| `allow_uncategorized` | `include` (default)  | Pages that match no category are grouped into segments with `category` set to `uncategorized`. Every page of the document appears in some segment. |
|                       | `forbid`             | Every page must be assigned to one of your categories; no `uncategorized` segment is returned.                                                     |
|                       | `omit`               | Pages that match no category are left out of `result.segments`, so the union of `pages` across segments can be smaller than the document.          |
| `min_pages_per_split` | integer, default `1` | Segments shorter than this are merged into an adjacent segment. `1` disables merging.                                                              |
| `custom_instructions` | string or null       | Free-form guidance that steers where boundaries fall; it combines with the category descriptions rather than replacing them.                       |

The values the job actually ran with are echoed back in the response’s `splitting_strategy`, so a reader of the result can tell whether missing pages were omitted on purpose.

## Beta endpoint differences

The beta endpoints, `POST /api/v1/beta/split/jobs` and `GET /api/v1/beta/split/jobs/{split_job_id}` (`client.beta.split` in the SDKs), return the same `id`, `status`, `categories`, `configuration_id`, `result`, `error_message`, `project_id`, `user_id`, `created_at`, and `updated_at`. They differ in how the input is reported:

| Beta response                                                                                        | Replaces                                                 |
| ---------------------------------------------------------------------------------------------------- | -------------------------------------------------------- |
| `document_input` — an object with `type` (`file_id`) and `value` (the file ID)                       | `file_input` and `document_input_type`                   |
| no `splitting_strategy`, `transaction_id`, `parse_tier`, `parse_config_id`, or `target_pages` fields | the echoed strategy, idempotency key, and parse settings |

`result.segments` is identical on both endpoints.

On both endpoints a request carries either `configuration_id` or an inline `configuration`, not both. `GET /api/v1/split/jobs` lists jobs with the same `items`, `next_page_token` and optional `total_size` shape as the other products, filterable by `status` and `job_ids`; in the SDKs, `client.split.list()` returns a cursor you can iterate directly.

## Using segments downstream

A segment’s `pages` array is the handle for the next step. The file ID in `file_input` can be reused, so the usual pattern is one downstream job per segment, scoped to that segment’s pages with a 1-based `target_pages` string built from the segment, such as `"4-8"`:

- **Extract**: set `target_pages` in the Extract configuration and pick the `data_schema` by `category`. The [resume book example](/llamaparse/extract/examples/split_and_extract_resume_book/index.md) builds the string as `f"{min(segment.pages)}-{max(segment.pages)}"`.
- **Parse**: set `page_ranges.target_pages` in the Parse configuration; see [Page ranges](/llamaparse/parse/guides/configuring-parse/#page-ranges/index.md).
- **Classify**: set `parsing_configuration.target_pages` on the classify job when a segment needs a finer label than its Split category.

Because `pages` is a list, not a range, build the string from its values rather than assuming the pages are contiguous.

## Fetch the full response

```
import os
import time
from llama_cloud import LlamaCloud


client = LlamaCloud(api_key=os.environ["LLAMA_CLOUD_API_KEY"])


file_obj = client.files.create(file="path/to/bundle.pdf", purpose="split")


job = client.split.create(
    file_input=file_obj.id,
    configuration={
        "categories": [
            {"name": "invoice", "description": "A commercial document requesting payment for goods or services"},
            {"name": "contract", "description": "A legal agreement between parties outlining terms, conditions, obligations, and signatures"},
        ],
        "splitting_strategy": {"allow_uncategorized": "include"},
    },
)


# Poll until the job reaches a terminal state
job = client.split.get(job.id)
while job.status not in ("completed", "failed", "cancelled"):
    time.sleep(2)
    job = client.split.get(job.id)


print(job.status, job.splitting_strategy)
if job.result is None:
    print(f"No result: {job.error_message}")
else:
    for segment in job.result.segments:
        pages = f"{min(segment.pages)}-{max(segment.pages)}"
        print(segment.category, pages, segment.confidence_category)
```
