Split response format
Field-by-field reference for the Split job response — the job envelope and statuses, result segments with category, 1-indexed page numbers and confidence level, uncategorized-page handling, the beta endpoint differences, and how to feed segments into Parse, Extract, or Classify.
A Split job returns one object, the split job, from POST /api/v1/split/jobs,
GET /api/v1/split/jobs/{split_job_id}, and the cancel endpoint. The SDKs return the same object
from client.split.create() and client.split.get(). This page lists every field in that object
and what each one means. For defining categories and running jobs, see
Getting started; for the request side, see the
Split API reference.
Split is currently in beta and is subject to breaking changes.
Job envelope
Section titled “Job envelope”| Field | Type | Description |
|---|---|---|
id | string | Job identifier, prefixed spl-. Pass it to GET /api/v1/split/jobs/{split_job_id}. |
status | string | Job state. See Job status. |
file_input | string | The file ID or parse job ID the job ran on. |
document_input_type | string | What file_input refers to: file_id, parse_job_id, or url. |
categories | array | The categories the job split against, each with name and an optional description. Echoed from the request or the saved configuration. |
splitting_strategy | object | The strategy the job ran with: allow_uncategorized, custom_instructions, and min_pages_per_split. See Uncategorized pages and merging. |
parse_tier | string or null | Parse tier requested for the job. null means the default, fast. Not used when file_input is a completed Parse job. |
parse_config_id | string or null | Saved parse configuration ID requested for the job, if any. Takes precedence over parse_tier. |
target_pages | string or null | Page selection requested for the job (1-based, "1-3,7"), if any. Accepted only when file_input is a completed Parse job. |
configuration_id | string or null | Saved configuration ID used for the job, if any. |
result | object or null | The segments. Present only when status is completed. See result.segments. |
error_message | string or null | Why the job failed, when status is failed. |
transaction_id | string or null | The idempotency key supplied at creation, if any. Reusing a key returns the original job. |
project_id | string | Project the job belongs to. |
user_id | string | User who created the job. |
created_at, updated_at | datetime | Creation and last-update timestamps. |
Job status
Section titled “Job status”Split statuses are lowercase, unlike the uppercase statuses of Extract and Classify jobs.
status | Meaning |
|---|---|
pending | Queued, not yet started. |
processing | Actively processing. |
completed | Finished; result is populated. |
failed | Terminated with an error; result is null, read error_message. |
cancelled | Cancelled by the user; result is null. |
Poll GET /api/v1/split/jobs/{split_job_id} until the status is one of the three terminal
values.
result.segments
Section titled “result.segments”result has a single field, segments, an array with one element per detected document section.
Each segment has:
| Field | Type | Description |
|---|---|---|
category | string | The name of the category the segment was assigned to, or uncategorized when no category matched and allow_uncategorized is include. |
pages | array of integers | The 1-indexed page numbers that make up the segment. |
confidence_category | string | Confidence in the assignment: high, medium, or low. |
Page numbers are those of the input document, so with target_pages set they still refer to the
original pages. Split detects where one document ends and the next begins, so adjacent segments
can share a category: a bundle of three invoices in a row comes back as three invoice segments,
not one.
From Getting started:
{ "id": "spl-abc123...", "status": "completed", "result": { "segments": [ { "category": "invoice", "pages": [1, 2, 3], "confidence_category": "high" }, { "category": "contract", "pages": [4, 5, 6, 7, 8], "confidence_category": "high" } ] }}Uncategorized pages and merging
Section titled “Uncategorized pages and merging”Three splitting_strategy options change which pages appear in result.segments:
| Option | Values | Effect on the result |
|---|---|---|
allow_uncategorized | include (default) | Pages that match no category are grouped into segments with category set to uncategorized. Every page of the document appears in some segment. |
forbid | Every page must be assigned to one of your categories; no uncategorized segment is returned. | |
omit | Pages that match no category are left out of result.segments, so the union of pages across segments can be smaller than the document. | |
min_pages_per_split | integer, default 1 | Segments shorter than this are merged into an adjacent segment. 1 disables merging. |
custom_instructions | string or null | Free-form guidance that steers where boundaries fall; it combines with the category descriptions rather than replacing them. |
The values the job actually ran with are echoed back in the response’s splitting_strategy, so a
reader of the result can tell whether missing pages were omitted on purpose.
Beta endpoint differences
Section titled “Beta endpoint differences”The beta endpoints, POST /api/v1/beta/split/jobs and GET /api/v1/beta/split/jobs/{split_job_id}
(client.beta.split in the SDKs), return the same id, status, categories,
configuration_id, result, error_message, project_id, user_id, created_at, and
updated_at. They differ in how the input is reported:
| Beta response | Replaces |
|---|---|
document_input — an object with type (file_id) and value (the file ID) | file_input and document_input_type |
no splitting_strategy, transaction_id, parse_tier, parse_config_id, or target_pages fields | the echoed strategy, idempotency key, and parse settings |
result.segments is identical on both endpoints.
On both endpoints a request carries either configuration_id or an inline configuration, not
both. GET /api/v1/split/jobs lists jobs with the same items, next_page_token and optional
total_size shape as the other products, filterable by status and job_ids; in the SDKs,
client.split.list() returns a cursor you can iterate directly.
Using segments downstream
Section titled “Using segments downstream”A segment’s pages array is the handle for the next step. The file ID in file_input can be
reused, so the usual pattern is one downstream job per segment, scoped to that segment’s pages
with a 1-based target_pages string built from the segment, such as "4-8":
- Extract: set
target_pagesin the Extract configuration and pick thedata_schemabycategory. The resume book example builds the string asf"{min(segment.pages)}-{max(segment.pages)}". - Parse: set
page_ranges.target_pagesin the Parse configuration; see Page ranges. - Classify: set
parsing_configuration.target_pageson the classify job when a segment needs a finer label than its Split category.
Because pages is a list, not a range, build the string from its values rather than assuming
the pages are contiguous.
Fetch the full response
Section titled “Fetch the full response”import osimport timefrom llama_cloud import LlamaCloud
client = LlamaCloud(api_key=os.environ["LLAMA_CLOUD_API_KEY"])
file_obj = client.files.create(file="path/to/bundle.pdf", purpose="split")
job = client.split.create( file_input=file_obj.id, configuration={ "categories": [ {"name": "invoice", "description": "A commercial document requesting payment for goods or services"}, {"name": "contract", "description": "A legal agreement between parties outlining terms, conditions, obligations, and signatures"}, ], "splitting_strategy": {"allow_uncategorized": "include"}, },)
# Poll until the job reaches a terminal statejob = client.split.get(job.id)while job.status not in ("completed", "failed", "cancelled"): time.sleep(2) job = client.split.get(job.id)
print(job.status, job.splitting_strategy)if job.result is None: print(f"No result: {job.error_message}")else: for segment in job.result.segments: pages = f"{min(segment.pages)}-{max(segment.pages)}" print(segment.category, pages, segment.confidence_category)