---
title: Use LlamaParse with LangChain | Developer Documentation
description: Turn Parse output into LangChain Document objects for splitting, embedding and retrieval, using the current llama-cloud SDK.
---

LlamaParse is framework-agnostic: it returns Markdown, text or JSON over the API, and any framework that works with text can consume it. For [LangChain](https://python.langchain.com/), the shortest path is to run a Parse job with the `llama-cloud` SDK and wrap each page in a LangChain `Document`.

Terminal window

```
pip install langchain-core langchain-text-splitters "llama-cloud>=2.8"
export LLAMA_CLOUD_API_KEY=llx-...
```

## Parse a file into LangChain documents

```
from langchain_core.documents import Document
from llama_cloud import LlamaCloud


client = LlamaCloud()  # reads LLAMA_CLOUD_API_KEY


file = client.files.create(file="data/report.pdf", purpose="parse")
result = client.parsing.parse(
    file_id=file.id, tier="agentic", version="latest", expand=["markdown"]
)


documents = [
    Document(
        page_content=p.markdown,
        metadata={"source": "data/report.pdf", "page": p.page_number},
    )
    for p in result.markdown.pages
    if p.success
]
```

Pages that failed to parse come back with `success` set to `False` and no Markdown, so the comprehension skips them. From here the documents go through the usual LangChain steps, for example a text splitter and a vector store:

```
from langchain_text_splitters import MarkdownHeaderTextSplitter


splitter = MarkdownHeaderTextSplitter(
    headers_to_split_on=[("#", "h1"), ("##", "h2")]
)
chunks = [
    chunk
    for doc in documents
    for chunk in splitter.split_text(doc.page_content)
]
```

Markdown is the output to prefer here: Parse preserves headings and tables as Markdown structure, which header-aware splitters use to keep sections together.

## Structured data with Extract

Extract returns JSON in your schema, which maps directly onto a LangChain tool or a structured-output step. Run the job with `client.extract.run(file_input=FILE_ID, configuration={"data_schema": DATA_SCHEMA})`, which waits for completion, and pass `job.extract_result` to the chain. See [Getting started with Extract](/llamaparse/extract/sdk/index.md) for the schema options.

## See also

- [Parse getting started](/llamaparse/parse/getting_started/index.md)
- [Retrieving results](/llamaparse/parse/guides/retrieving-results/index.md) for the other output representations
- [Use LlamaParse from the LlamaIndex Framework](../llamaindex-framework/)
