Skip to content

Extract Data from Financial Reports - with Citations

LlamaExtract pulls structured fields from an SEC 10-K filing with cite_sources enabled, returning per-field citations to verify each value.

Given complex files like financial reports, contracts, invoices etc, Llama Extract allows you to make use of an LLM to extract the information relevant to you, in a structured format.

In this example, we’ll be using LlamaExtract to extract structured data from an SEC filing (specifically, the filing by Nvidia for fiscal year 2025).

On top of simple data extraction, we’ll enable citations for each extracted field. This lets us verify values against the source document and see where schema descriptions or the system prompt need to be tightened.

The example we go through below is also replicable within Llama Cloud as well, where you will also be able to pick between a number of pre-defined schemas, instead of building your own.

!pip install llama-cloud>=2.1

To get started, make sure you provide your Llama Cloud API key.

import os
from llama_cloud import LlamaCloud, AsyncLlamaCloud
# Could use async or sync clients, this page will use the sync client
client = LlamaCloud(api_key=os.environ["LLAMA_CLOUD_API_KEY"])

When using LlamaExtract via the API, you provide your own schema that describes what you want extracted from your documents. Here, we define a schema for extracting key information from SEC filings.

from pydantic import BaseModel, Field
from enum import Enum
class FilingType(str, Enum):
ten_k = '10 K'
ten_q = '10-Q'
ten_ka = '10-K/A'
ten_qa = '10-Q/A'
class FinancialReport(BaseModel):
company_name: str = Field(description="The name of the company")
description: str = Field(description="Short description of the filing and what it contains")
filing_type: FilingType = Field(description="Type of SEC filing")
filing_date: str = Field(description="Date when filing was submitted to SEC")
fiscal_year: int = Field(description="Fiscal year")
unit: str = Field(description="Unit of financial figures (thousands, millions, etc.)")
revenue: int = Field(description="Total revenue for period")

Download a PDF and Extract Data with Citations

Section titled “Download a PDF and Extract Data with Citations”

We’ll enable cite_sources in the extraction config so we can trace every value back to its origin in the document.

import time
import requests
url = 'https://raw.githubusercontent.com/run-llama/llama_cloud_services/refs/heads/main/examples/extract/data/sec_filings/nvda_10k.pdf'
response = requests.get(url)
if response.status_code == 200:
with open('/content/nvda_10k.pdf', 'wb') as f:
f.write(response.content)
print("PDF downloaded successfully.")
else:
print(f"Failed to download. Status code: {response.status_code}")
file_obj = client.files.create(
file="/content/nvda_10k.pdf",
purpose="extract",
)
file_id = file_obj.id
job = client.extract.create(
file_input=file_id,
configuration={
"data_schema": FinancialReport.model_json_schema(),
"tier": "agentic",
"cite_sources": True,
},
)
# Poll for completion
while job.status not in ("COMPLETED", "FAILED", "CANCELLED"):
time.sleep(2)
job = client.extract.get(job.id)
print(job.extract_result)

You should get an output similar to this:

{
'company_name': 'NVIDIA Corporation',
'description': "The filing provides a detailed overview of NVIDIA's business as a full-stack computing infrastructure company, discusses various technologies including digital avatars and autonomous vehicles, outlines numerous risk factors affecting operations such as supply chain issues and geopolitical tensions, and describes employee stock purchase plans and related compliance requirements.",
'filing_type': '10 K',
'filing_date': 'February 26, 2025',
'fiscal_year': 2025,
'unit': 'millions',
'revenue': 130497
}
# Fetch metadata (not returned by default)
detailed = client.extract.get(job.id, expand=["extract_metadata"])
print(detailed.extract_metadata)

Citations are under field_metadata.document_metadata. The example below shows citations for each extracted field, with repeated locations omitted.

{
"field_metadata": {
"document_metadata": {
"company_name": {
"citation": [
{"page": 1, "matching_text": "NVIDIA CORPORATION"},
{"page": 2, "matching_text": "NVIDIA Corporation"}
]
},
"filing_type": {
"citation": [
{"page": 1, "matching_text": "FORM 10-K"},
{"page": 2, "matching_text": "Item 16. | Form 10-K Summary"}
]
},
"fiscal_year": {
"citation": [
{"page": 1, "matching_text": "For the fiscal year ended January 26, 2025"},
{"page": 6, "matching_text": "In fiscal year 2025, we launched the NVIDIA Blackwell architecture"}
]
},
"description": {
"citation": [
{"page": 4, "matching_text": "NVIDIA is now a full-stack computing infrastructure company with data-center-scale offerings that are reshaping industry."}
]
},
"filing_date": {
"citation": [
{"page": 51, "matching_text": "February 26, 2025"},
{"page": 86, "matching_text": "on February 26, 2025."}
]
},
"unit": {
"citation": [
{"page": 38, "matching_text": "($ in millions, except per share data)"},
{"page": 42, "matching_text": "($ in millions)"}
]
},
"revenue": {
"citation": [
{"page": 38, "matching_text": "Revenue for fiscal year 2025 was $130.5 billion"},
{"page": 52, "matching_text": "Revenue | $ 130,497"}
]
}
}
}
}

See Extract response format for response fields and usage credits.

In this example, we used LlamaExtract to extract data with citations from a financial document. To further customize and improve on the results, you can also try adding a system_prompt to the extraction config.

Note for AI agents: this documentation is built for programmatic access. - Overview of all docs: https://developers.llamaindex.ai/llms.txt - Any page is available as raw Markdown by appending index.md to its URL — e.g. https://developers.llamaindex.ai/llamaparse/parse/getting_started/index.md - Agent-friendly REST search APIs live under https://developers.llamaindex.ai/api/ — search (BM25 full-text), grep (regex), read (fetch a page), and list (browse the doc tree). See https://developers.llamaindex.ai/llms.txt for parameters. - A hosted documentation MCP server is available at https://developers.llamaindex.ai/mcp. If you support MCP, you can ask the user to install it for browsing these docs directly (an alternative to the REST API). Setup: https://developers.llamaindex.ai/for-agents/mcp/ - Other LlamaIndex tooling for agents — the LlamaParse Platform MCP server, agent skills and plugins, and the n8n node — is mapped at https://developers.llamaindex.ai/for-agents/