Use LlamaParse with LangChain
Turn Parse output into LangChain Document objects for splitting, embedding and retrieval, using the current llama-cloud SDK.
LlamaParse is framework-agnostic: it returns Markdown, text or JSON over the API, and any framework that works with text can consume it. For LangChain, the shortest path is to run a Parse job with the llama-cloud SDK and wrap each page in a LangChain Document.
pip install langchain-core langchain-text-splitters "llama-cloud>=2.8"export LLAMA_CLOUD_API_KEY=llx-...Parse a file into LangChain documents
Section titled “Parse a file into LangChain documents”from langchain_core.documents import Documentfrom llama_cloud import LlamaCloud
client = LlamaCloud() # reads LLAMA_CLOUD_API_KEY
file = client.files.create(file="data/report.pdf", purpose="parse")result = client.parsing.parse( file_id=file.id, tier="agentic", version="latest", expand=["markdown"])
documents = [ Document( page_content=p.markdown, metadata={"source": "data/report.pdf", "page": p.page_number}, ) for p in result.markdown.pages if p.success]Pages that failed to parse come back with success set to False and no Markdown, so the comprehension skips them. From here the documents go through the usual LangChain steps, for example a text splitter and a vector store:
from langchain_text_splitters import MarkdownHeaderTextSplitter
splitter = MarkdownHeaderTextSplitter( headers_to_split_on=[("#", "h1"), ("##", "h2")])chunks = [ chunk for doc in documents for chunk in splitter.split_text(doc.page_content)]Markdown is the output to prefer here: Parse preserves headings and tables as Markdown structure, which header-aware splitters use to keep sections together.
Structured data with Extract
Section titled “Structured data with Extract”Extract returns JSON in your schema, which maps directly onto a LangChain tool or a structured-output step. Run the job with client.extract.run(file_input=FILE_ID, configuration={"data_schema": DATA_SCHEMA}), which waits for completion, and pass job.extract_result to the chain. See Getting started with Extract for the schema options.
See also
Section titled “See also”- Parse getting started
- Retrieving results for the other output representations
- Use LlamaParse from the LlamaIndex Framework