Skip to content
Guide
Split

Split response format

Field-by-field reference for the Split job response — the job envelope and statuses, result segments with category, 1-indexed page numbers and confidence level, uncategorized-page handling, the beta endpoint differences, and how to feed segments into Parse, Extract, or Classify.

A Split job returns one object, the split job, from POST /api/v1/split/jobs, GET /api/v1/split/jobs/{split_job_id}, and the cancel endpoint. The SDKs return the same object from client.split.create() and client.split.get(). This page lists every field in that object and what each one means. For defining categories and running jobs, see Getting started; for the request side, see the Split API reference.

Split is currently in beta and is subject to breaking changes.

FieldTypeDescription
idstringJob identifier, prefixed spl-. Pass it to GET /api/v1/split/jobs/{split_job_id}.
statusstringJob state. See Job status.
file_inputstringThe file ID or parse job ID the job ran on.
document_input_typestringWhat file_input refers to: file_id, parse_job_id, or url.
categoriesarrayThe categories the job split against, each with name and an optional description. Echoed from the request or the saved configuration.
splitting_strategyobjectThe strategy the job ran with: allow_uncategorized, custom_instructions, and min_pages_per_split. See Uncategorized pages and merging.
parse_tierstring or nullParse tier requested for the job. null means the default, fast. Not used when file_input is a completed Parse job.
parse_config_idstring or nullSaved parse configuration ID requested for the job, if any. Takes precedence over parse_tier.
target_pagesstring or nullPage selection requested for the job (1-based, "1-3,7"), if any. Accepted only when file_input is a completed Parse job.
configuration_idstring or nullSaved configuration ID used for the job, if any.
resultobject or nullThe segments. Present only when status is completed. See result.segments.
error_messagestring or nullWhy the job failed, when status is failed.
transaction_idstring or nullThe idempotency key supplied at creation, if any. Reusing a key returns the original job.
project_idstringProject the job belongs to.
user_idstringUser who created the job.
created_at, updated_atdatetimeCreation and last-update timestamps.

Split statuses are lowercase, unlike the uppercase statuses of Extract and Classify jobs.

statusMeaning
pendingQueued, not yet started.
processingActively processing.
completedFinished; result is populated.
failedTerminated with an error; result is null, read error_message.
cancelledCancelled by the user; result is null.

Poll GET /api/v1/split/jobs/{split_job_id} until the status is one of the three terminal values.

result has a single field, segments, an array with one element per detected document section. Each segment has:

FieldTypeDescription
categorystringThe name of the category the segment was assigned to, or uncategorized when no category matched and allow_uncategorized is include.
pagesarray of integersThe 1-indexed page numbers that make up the segment.
confidence_categorystringConfidence in the assignment: high, medium, or low.

Page numbers are those of the input document, so with target_pages set they still refer to the original pages. Split detects where one document ends and the next begins, so adjacent segments can share a category: a bundle of three invoices in a row comes back as three invoice segments, not one. From Getting started:

{
"id": "spl-abc123...",
"status": "completed",
"result": {
"segments": [
{
"category": "invoice",
"pages": [1, 2, 3],
"confidence_category": "high"
},
{
"category": "contract",
"pages": [4, 5, 6, 7, 8],
"confidence_category": "high"
}
]
}
}

Three splitting_strategy options change which pages appear in result.segments:

OptionValuesEffect on the result
allow_uncategorizedinclude (default)Pages that match no category are grouped into segments with category set to uncategorized. Every page of the document appears in some segment.
forbidEvery page must be assigned to one of your categories; no uncategorized segment is returned.
omitPages that match no category are left out of result.segments, so the union of pages across segments can be smaller than the document.
min_pages_per_splitinteger, default 1Segments shorter than this are merged into an adjacent segment. 1 disables merging.
custom_instructionsstring or nullFree-form guidance that steers where boundaries fall; it combines with the category descriptions rather than replacing them.

The values the job actually ran with are echoed back in the response’s splitting_strategy, so a reader of the result can tell whether missing pages were omitted on purpose.

The beta endpoints, POST /api/v1/beta/split/jobs and GET /api/v1/beta/split/jobs/{split_job_id} (client.beta.split in the SDKs), return the same id, status, categories, configuration_id, result, error_message, project_id, user_id, created_at, and updated_at. They differ in how the input is reported:

Beta responseReplaces
document_input — an object with type (file_id) and value (the file ID)file_input and document_input_type
no splitting_strategy, transaction_id, parse_tier, parse_config_id, or target_pages fieldsthe echoed strategy, idempotency key, and parse settings

result.segments is identical on both endpoints.

On both endpoints a request carries either configuration_id or an inline configuration, not both. GET /api/v1/split/jobs lists jobs with the same items, next_page_token and optional total_size shape as the other products, filterable by status and job_ids; in the SDKs, client.split.list() returns a cursor you can iterate directly.

A segment’s pages array is the handle for the next step. The file ID in file_input can be reused, so the usual pattern is one downstream job per segment, scoped to that segment’s pages with a 1-based target_pages string built from the segment, such as "4-8":

  • Extract: set target_pages in the Extract configuration and pick the data_schema by category. The resume book example builds the string as f"{min(segment.pages)}-{max(segment.pages)}".
  • Parse: set page_ranges.target_pages in the Parse configuration; see Page ranges.
  • Classify: set parsing_configuration.target_pages on the classify job when a segment needs a finer label than its Split category.

Because pages is a list, not a range, build the string from its values rather than assuming the pages are contiguous.

import os
import time
from llama_cloud import LlamaCloud
client = LlamaCloud(api_key=os.environ["LLAMA_CLOUD_API_KEY"])
file_obj = client.files.create(file="path/to/bundle.pdf", purpose="split")
job = client.split.create(
file_input=file_obj.id,
configuration={
"categories": [
{"name": "invoice", "description": "A commercial document requesting payment for goods or services"},
{"name": "contract", "description": "A legal agreement between parties outlining terms, conditions, obligations, and signatures"},
],
"splitting_strategy": {"allow_uncategorized": "include"},
},
)
# Poll until the job reaches a terminal state
job = client.split.get(job.id)
while job.status not in ("completed", "failed", "cancelled"):
time.sleep(2)
job = client.split.get(job.id)
print(job.status, job.splitting_strategy)
if job.result is None:
print(f"No result: {job.error_message}")
else:
for segment in job.result.segments:
pages = f"{min(segment.pages)}-{max(segment.pages)}"
print(segment.category, pages, segment.confidence_category)
Note for AI agents: this documentation is built for programmatic access. - Overview of all docs: https://developers.llamaindex.ai/llms.txt - Any page is available as raw Markdown by appending index.md to its URL — e.g. https://developers.llamaindex.ai/llamaparse/parse/getting_started/index.md - Agent-friendly REST search APIs live under https://developers.llamaindex.ai/api/ — search (BM25 full-text), grep (regex), read (fetch a page), and list (browse the doc tree). See https://developers.llamaindex.ai/llms.txt for parameters. - A hosted documentation MCP server is available at https://developers.llamaindex.ai/mcp. If you support MCP, you can ask the user to install it for browsing these docs directly (an alternative to the REST API). Setup: https://developers.llamaindex.ai/for-agents/mcp/ - Other LlamaIndex tooling for agents — the LlamaParse Platform MCP server, agent skills and plugins, and the n8n node — is mapped at https://developers.llamaindex.ai/for-agents/