Extract Data from Financial Reports - with Citations
LlamaExtract pulls structured fields from an SEC 10-K filing with cite_sources enabled, returning per-field citations to verify each value.
Given complex files like financial reports, contracts, invoices etc, Llama Extract allows you to make use of an LLM to extract the information relevant to you, in a structured format.
In this example, we’ll be using LlamaExtract to extract structured data from an SEC filing (specifically, the filing by Nvidia for fiscal year 2025).
On top of simple data extraction, we’ll enable citations for each extracted field. This lets us verify values against the source document and see where schema descriptions or the system prompt need to be tightened.
The example we go through below is also replicable within Llama Cloud as well, where you will also be able to pick between a number of pre-defined schemas, instead of building your own.
!pip install llama-cloud>=2.1npm install @llamaindex/llama-cloud zodgo get github.com/run-llama/llama-parse-goimplementation("ai.llamaindex:llama-cloud:1.3.0")go install github.com/run-llama/llama-parse-cli/cmd/llp@latestConnect to Llama Cloud
Section titled “Connect to Llama Cloud”To get started, make sure you provide your Llama Cloud API key.
import osfrom llama_cloud import LlamaCloud, AsyncLlamaCloud
# Could use async or sync clients, this page will use the sync clientclient = LlamaCloud(api_key=os.environ["LLAMA_CLOUD_API_KEY"])import LlamaCloud from "@llamaindex/llama-cloud";
const client = new LlamaCloud({ apiKey: process.env.LLAMA_CLOUD_API_KEY!,});import ( "context"
llamacloud "github.com/run-llama/llama-parse-go")
ctx := context.Background()
// Reads LLAMA_CLOUD_API_KEY from the environmentclient := llamacloud.NewClient()import ai.llamaindex.llamacloud.client.LlamaCloudClient;import ai.llamaindex.llamacloud.client.okhttp.LlamaCloudOkHttpClient;
// Reads LLAMA_CLOUD_API_KEY from the environmentLlamaCloudClient client = LlamaCloudOkHttpClient.fromEnv();export LLAMA_CLOUD_API_KEY=llx-xxxxxxExtract Data with LlamaExtract
Section titled “Extract Data with LlamaExtract”Provide Your Custom Schema
Section titled “Provide Your Custom Schema”When using LlamaExtract via the API, you provide your own schema that describes what you want extracted from your documents. Here, we define a schema for extracting key information from SEC filings.
from pydantic import BaseModel, Fieldfrom enum import Enum
class FilingType(str, Enum): ten_k = '10 K' ten_q = '10-Q' ten_ka = '10-K/A' ten_qa = '10-Q/A'
class FinancialReport(BaseModel): company_name: str = Field(description="The name of the company") description: str = Field(description="Short description of the filing and what it contains") filing_type: FilingType = Field(description="Type of SEC filing") filing_date: str = Field(description="Date when filing was submitted to SEC") fiscal_year: int = Field(description="Fiscal year") unit: str = Field(description="Unit of financial figures (thousands, millions, etc.)") revenue: int = Field(description="Total revenue for period")import { z } from "zod";
const FilingType = z.enum(['10 K', '10-Q', '10-K/A', '10-Q/A']);
const FinancialReportSchema = z.object({ company_name: z.string().describe("The name of the company"), description: z.string().describe("Short description of the filing and what it contains"), filing_type: FilingType.describe("Type of SEC filing"), filing_date: z.string().describe("Date when filing was submitted to SEC"), fiscal_year: z.number().describe("Fiscal year"), unit: z.string().describe("Unit of financial figures (thousands, millions, etc.)"), revenue: z.number().describe("Total revenue for period"),});import llamacloud "github.com/run-llama/llama-parse-go"
// Define the schema as a JSON Schema literaldataSchema := map[string]*llamacloud.ExtractConfigurationDataSchemaUnionParam{ "type": {OfString: llamacloud.String("object")}, "properties": {OfAnyMap: map[string]any{ "company_name": map[string]any{"type": "string", "description": "The name of the company"}, "description": map[string]any{"type": "string", "description": "Short description of the filing and what it contains"}, "filing_type": map[string]any{"type": "string", "enum": []any{"10 K", "10-Q", "10-K/A", "10-Q/A"}, "description": "Type of SEC filing"}, "filing_date": map[string]any{"type": "string", "description": "Date when filing was submitted to SEC"}, "fiscal_year": map[string]any{"type": "integer", "description": "Fiscal year"}, "unit": map[string]any{"type": "string", "description": "Unit of financial figures (thousands, millions, etc.)"}, "revenue": map[string]any{"type": "integer", "description": "Total revenue for period"}, }}, "required": {OfAnyArray: []any{"company_name", "description", "filing_type", "filing_date", "fiscal_year", "unit", "revenue"}},}import ai.llamaindex.llamacloud.core.JsonValue;import ai.llamaindex.llamacloud.models.extract.ExtractConfiguration;import java.util.List;import java.util.Map;
// Define the schema as a JSON Schema literalExtractConfiguration.DataSchema dataSchema = ExtractConfiguration.DataSchema.builder() .putAdditionalProperty("type", JsonValue.from("object")) .putAdditionalProperty("properties", JsonValue.from(Map.of( "company_name", Map.of("type", "string", "description", "The name of the company"), "description", Map.of("type", "string", "description", "Short description of the filing and what it contains"), "filing_type", Map.of("type", "string", "enum", List.of("10 K", "10-Q", "10-K/A", "10-Q/A"), "description", "Type of SEC filing"), "filing_date", Map.of("type", "string", "description", "Date when filing was submitted to SEC"), "fiscal_year", Map.of("type", "integer", "description", "Fiscal year"), "unit", Map.of("type", "string", "description", "Unit of financial figures (thousands, millions, etc.)"), "revenue", Map.of("type", "integer", "description", "Total revenue for period")))) .putAdditionalProperty("required", JsonValue.from(List.of( "company_name", "description", "filing_type", "filing_date", "fiscal_year", "unit", "revenue"))) .build();# Define the schema as a JSON Schema literalDATA_SCHEMA='{ "type": "object", "properties": { "company_name": {"type": "string", "description": "The name of the company"}, "description": {"type": "string", "description": "Short description of the filing and what it contains"}, "filing_type": {"type": "string", "enum": ["10 K", "10-Q", "10-K/A", "10-Q/A"], "description": "Type of SEC filing"}, "filing_date": {"type": "string", "description": "Date when filing was submitted to SEC"}, "fiscal_year": {"type": "integer", "description": "Fiscal year"}, "unit": {"type": "string", "description": "Unit of financial figures (thousands, millions, etc.)"}, "revenue": {"type": "integer", "description": "Total revenue for period"} }, "required": ["company_name", "description", "filing_type", "filing_date", "fiscal_year", "unit", "revenue"]}'Download a PDF and Extract Data with Citations
Section titled “Download a PDF and Extract Data with Citations”We’ll enable cite_sources in the extraction config so we can trace every value back to its origin in the document.
import timeimport requests
url = 'https://raw.githubusercontent.com/run-llama/llama_cloud_services/refs/heads/main/examples/extract/data/sec_filings/nvda_10k.pdf'
response = requests.get(url)
if response.status_code == 200: with open('/content/nvda_10k.pdf', 'wb') as f: f.write(response.content) print("PDF downloaded successfully.")else: print(f"Failed to download. Status code: {response.status_code}")
file_obj = client.files.create( file="/content/nvda_10k.pdf", purpose="extract",)file_id = file_obj.id
job = client.extract.create( file_input=file_id, configuration={ "data_schema": FinancialReport.model_json_schema(), "tier": "agentic", "cite_sources": True, },)
# Poll for completionwhile job.status not in ("COMPLETED", "FAILED", "CANCELLED"): time.sleep(2) job = client.extract.get(job.id)
print(job.extract_result)const url = "https://raw.githubusercontent.com/run-llama/llama_cloud_services/refs/heads/main/examples/extract/data/sec_filings/nvda_10k.pdf"
const response = await fetch(url);
// Upload the file firstconst fileObj = await client.files.create({ file: response, purpose: "extract",});const fileId = fileObj.id;
let job = await client.extract.create({ file_input: fileId, configuration: { data_schema: z.toJSONSchema(FinancialReportSchema), tier: 'agentic', cite_sources: true, },});
// Poll for completionwhile (!['COMPLETED', 'FAILED', 'CANCELLED'].includes(job.status)) { await new Promise((r) => setTimeout(r, 2000)); job = await client.extract.get(job.id);}
console.log(job.extract_result);import ( "io" "net/http" "os" "time"
llamacloud "github.com/run-llama/llama-parse-go")
url := "https://raw.githubusercontent.com/run-llama/llama_cloud_services/refs/heads/main/examples/extract/data/sec_filings/nvda_10k.pdf"
// Download the sample filingresp, err := http.Get(url)if err != nil { log.Fatal(err)}defer resp.Body.Close()
out, err := os.Create("nvda_10k.pdf")if err != nil { log.Fatal(err)}if _, err := io.Copy(out, resp.Body); err != nil { log.Fatal(err)}out.Close()
// Upload the file firstf, err := os.Open("nvda_10k.pdf")if err != nil { log.Fatal(err)}defer f.Close()
fileObj, err := client.Files.New(ctx, llamacloud.FileNewParams{ File: f, Purpose: "extract",})if err != nil { log.Fatal(err)}
// Extract with cite_sources enabledjob, err := client.Extract.New(ctx, llamacloud.ExtractNewParams{ ExtractV2JobCreate: llamacloud.ExtractV2JobCreateParam{ FileInput: fileObj.ID, Configuration: llamacloud.ExtractConfigurationParam{ DataSchema: dataSchema, Tier: llamacloud.ExtractConfigurationTierAgentic, CiteSources: llamacloud.Bool(true), }, },})if err != nil { log.Fatal(err)}
// Poll for completionfor job.Status != "COMPLETED" && job.Status != "FAILED" && job.Status != "CANCELLED" { time.Sleep(2 * time.Second) job, err = client.Extract.Get(ctx, job.ID, llamacloud.ExtractGetParams{}) if err != nil { log.Fatal(err) }}
fmt.Println(job.ExtractResult.RawJSON())import ai.llamaindex.llamacloud.models.extract.ExtractCreateParams;import ai.llamaindex.llamacloud.models.extract.ExtractGetParams;import ai.llamaindex.llamacloud.models.extract.ExtractV2Job;import ai.llamaindex.llamacloud.models.extract.ExtractV2JobCreate;import ai.llamaindex.llamacloud.models.files.FileCreateParams;import ai.llamaindex.llamacloud.models.files.FileCreateResponse;import java.io.InputStream;import java.net.URI;import java.nio.file.Files;import java.nio.file.Path;import java.nio.file.Paths;
String url = "https://raw.githubusercontent.com/run-llama/llama_cloud_services/refs/heads/main/examples/extract/data/sec_filings/nvda_10k.pdf";
// Download the sample filingPath pdfPath = Paths.get("nvda_10k.pdf");try (InputStream in = URI.create(url).toURL().openStream()) { Files.copy(in, pdfPath);}
// Upload the file firstFileCreateResponse fileObj = client.files().create( FileCreateParams.builder() .file(pdfPath) .purpose("extract") .build());
// Extract with cite_sources enabledExtractV2Job job = client.extract().create( ExtractCreateParams.builder() .extractV2JobCreate( ExtractV2JobCreate.builder() .fileInput(fileObj.id()) .configuration( ExtractConfiguration.builder() .dataSchema(dataSchema) .tier(ExtractConfiguration.Tier.AGENTIC) .citeSources(true) .build()) .build()) .build());
// Poll for completionwhile (!job.status().equals("COMPLETED") && !job.status().equals("FAILED") && !job.status().equals("CANCELLED")) { Thread.sleep(2000); job = client.extract().get(ExtractGetParams.builder().jobId(job.id()).build());}
System.out.println(job.extractResult());URL="https://raw.githubusercontent.com/run-llama/llama_cloud_services/refs/heads/main/examples/extract/data/sec_filings/nvda_10k.pdf"
# Download the sample filingcurl -sL "$URL" -o nvda_10k.pdf
# Upload the file firstFILE_ID=$(llp files create --file nvda_10k.pdf --purpose extract | jq -r '.id')
# Extract with cite_sources enabledJOB_ID=$(llp extract create \ --file-input "$FILE_ID" \ --configuration "{\"data_schema\": $DATA_SCHEMA, \"tier\": \"agentic\", \"cite_sources\": true}" \ | jq -r '.id')
# Poll for completionwhile true; do JOB=$(llp extract get --job-id "$JOB_ID") STATUS=$(echo "$JOB" | jq -r '.status') case "$STATUS" in COMPLETED|FAILED|CANCELLED) break ;; esac sleep 2done
echo "$JOB" | jq '.extract_result'You should get an output similar to this:
{ 'company_name': 'NVIDIA Corporation', 'description': "The filing provides a detailed overview of NVIDIA's business as a full-stack computing infrastructure company, discusses various technologies including digital avatars and autonomous vehicles, outlines numerous risk factors affecting operations such as supply chain issues and geopolitical tensions, and describes employee stock purchase plans and related compliance requirements.", 'filing_type': '10 K', 'filing_date': 'February 26, 2025', 'fiscal_year': 2025, 'unit': 'millions', 'revenue': 130497}Inspect Citations
Section titled “Inspect Citations”# Fetch metadata (not returned by default)detailed = client.extract.get(job.id, expand=["extract_metadata"])print(detailed.extract_metadata)// Fetch metadata (not returned by default)const detailed = await client.extract.get(job.id, { expand: ["extract_metadata"] });console.log(detailed.extract_metadata);// Fetch metadata (not returned by default)detailed, err := client.Extract.Get(ctx, job.ID, llamacloud.ExtractGetParams{ Expand: []string{"extract_metadata"},})if err != nil { log.Fatal(err)}fmt.Println(detailed.ExtractMetadata.RawJSON())// Fetch metadata (not returned by default)ExtractV2Job detailed = client.extract().get( ExtractGetParams.builder() .jobId(job.id()) .addExpand("extract_metadata") .build());System.out.println(detailed.extractMetadata());# Fetch metadata (not returned by default)llp extract get --job-id "$JOB_ID" --expand extract_metadata | jq '.extract_metadata'Citations are under field_metadata.document_metadata.
The example below shows citations for each extracted field, with repeated locations omitted.
{ "field_metadata": { "document_metadata": { "company_name": { "citation": [ {"page": 1, "matching_text": "NVIDIA CORPORATION"}, {"page": 2, "matching_text": "NVIDIA Corporation"} ] }, "filing_type": { "citation": [ {"page": 1, "matching_text": "FORM 10-K"}, {"page": 2, "matching_text": "Item 16. | Form 10-K Summary"} ] }, "fiscal_year": { "citation": [ {"page": 1, "matching_text": "For the fiscal year ended January 26, 2025"}, {"page": 6, "matching_text": "In fiscal year 2025, we launched the NVIDIA Blackwell architecture"} ] }, "description": { "citation": [ {"page": 4, "matching_text": "NVIDIA is now a full-stack computing infrastructure company with data-center-scale offerings that are reshaping industry."} ] }, "filing_date": { "citation": [ {"page": 51, "matching_text": "February 26, 2025"}, {"page": 86, "matching_text": "on February 26, 2025."} ] }, "unit": { "citation": [ {"page": 38, "matching_text": "($ in millions, except per share data)"}, {"page": 42, "matching_text": "($ in millions)"} ] }, "revenue": { "citation": [ {"page": 38, "matching_text": "Revenue for fiscal year 2025 was $130.5 billion"}, {"page": 52, "matching_text": "Revenue | $ 130,497"} ] } } }}See Extract response format for response fields and usage credits.
What’s Next?
Section titled “What’s Next?”In this example, we used LlamaExtract to extract data with citations from a financial document. To further customize and improve on the results, you can also try adding a system_prompt to the extraction config.