Custom Vector Stores
Configure an index to export raw parsed output instead of writing to a managed vector store, then load the data into your own database.
By default, an index parses, chunks, embeds, and writes to a managed vector store automatically. If you’d rather own the vector store (because you’re self-hosting the platform with a database we don’t manage, or you simply want to keep your data in your own infrastructure), you can disable the managed export and download the raw parsed output instead.
In this mode, LlamaCloud still handles directory management, parsing, and sync orchestration. You take over chunking, embedding, and writing to your database.
Create an parse-only index
Section titled “Create an parse-only index”Pass vector_target="DISABLED" when creating the index. The pipeline falls back to a download-only export, so parsed output is written to object storage instead of a managed vector store.
index = await client.indexes.create( source_directory_id=directory.id, vector_target="DISABLED",)const index = await client.indexes.create({ source_directory_id: directory.id, vector_target: "DISABLED",});index, err := client.Indexes.New(ctx, llamacloud.IndexNewParams{ SourceDirectoryID: directory.ID, VectorTarget: llamacloud.IndexNewParamsVectorTargetDisabled,})if err != nil { log.Fatal(err)}import ai.llamaindex.llamacloud.models.indexes.IndexCreateParams;import ai.llamaindex.llamacloud.models.indexes.IndexCreateResponse;
IndexCreateResponse index = client.indexes().create( IndexCreateParams.builder() .sourceDirectoryId(directory.id()) .vectorTarget(IndexCreateParams.VectorTarget.DISABLED) .build());INDEX_ID=$(llp beta:indexes create \ --source-directory-id "$DIRECTORY_ID" \ --vector-target DISABLED | jq -r '.id')
echo "Index ID: $INDEX_ID"Wait for the index to reach ready the same way as a normal index (see the getting started guide).
Download the parsed output
Section titled “Download the parsed output”Each index writes its parsed output to an output directory, separate from the source directory you created. The output directory holds one file per source file: a JSON payload with per-page markdown (header, text, footer) and references to each page’s attachments, such as screenshots. The source directory still holds your original files (PDFs, DOCX, …), so read the payloads from the output directory.
Get the output directory’s ID from the index as output_directory_id. The same response carries export_config_id, which you use when keying records in your database. Then list the output directory’s files and download each payload from its presigned URL.
Each output directory file carries the fields you need to track it in your database:
id: theparsed_directory_file_id. It stays stable for a source file across syncs.file_id: the stored JSON payload. Pass it toclient.files.contentto get a presigned download URL.metadata: the metadata you set on the source file.updated_at: changes when the file is parsed again or its metadata changes.
import httpx
# The index points at its output directory and export config.idx = await client.indexes.get(index.id)
async with httpx.AsyncClient() as http: # Iterating the list call walks every page of results. async for f in client.beta.directories.files.list(idx.output_directory_id): if f.file_id is None: continue presigned = await client.files.content(f.file_id) response = await http.get(presigned.url) response.raise_for_status() pages = response.json()["parse"]["pages"] # Key your records on (f.id, idx.export_config_id). # f.metadata holds the source file's metadata.// The index points at its output directory and export config.const idx = await client.indexes.get(index.id);
// Iterating the list call walks every page of results.for await (const f of client.beta.directories.files.list(idx.output_directory_id)) { if (!f.file_id) continue; const presigned = await client.files.content(f.file_id); const response = await fetch(presigned.url); const pages = (await response.json()).parse.pages; // Key your records on (f.id, idx.export_config_id). // f.metadata holds the source file's metadata.}// The index points at its output directory and export config.idx, err := client.Indexes.Get(ctx, index.ID, llamacloud.IndexGetParams{})if err != nil { log.Fatal(err)}
// ListAutoPaging walks every page of results.iter := client.Beta.Directories.Files.ListAutoPaging(ctx, idx.OutputDirectoryID, llamacloud.BetaDirectoryFileListParams{})for iter.Next() { f := iter.Current() if f.FileID == "" { continue } presigned, err := client.Files.Content(ctx, f.FileID, llamacloud.FileContentParams{}) if err != nil { log.Fatal(err) }
resp, err := http.Get(presigned.URL) if err != nil { log.Fatal(err) }
var parsed map[string]any err = json.NewDecoder(resp.Body).Decode(&parsed) resp.Body.Close() if err != nil { log.Fatal(err) } // parsed["parse"] holds the pages. Key your records on // (f.ID, idx.ExportConfigID); f.Metadata holds the source file's metadata.}if err := iter.Err(); err != nil { log.Fatal(err)}import java.net.URI;import java.net.http.HttpClient;import java.net.http.HttpRequest;import java.net.http.HttpResponse;
import ai.llamaindex.llamacloud.models.beta.directories.files.FileListResponse;import ai.llamaindex.llamacloud.models.indexes.IndexGetResponse;import ai.llamaindex.llamacloud.models.files.PresignedUrl;
// The index points at its output directory and export config.IndexGetResponse idx = client.indexes().get(index.id());HttpClient http = HttpClient.newHttpClient();
// autoPager() walks every page of results.for (FileListResponse f : client.beta().directories().files() .list(idx.outputDirectoryId()).autoPager()) { if (f.fileId().isEmpty()) { continue; } PresignedUrl presigned = client.files().content(f.fileId().get());
HttpResponse<String> response = http.send( HttpRequest.newBuilder(URI.create(presigned.url())).build(), HttpResponse.BodyHandlers.ofString()); String parsed = response.body(); // parsed is the payload JSON; its "parse" field holds the pages. // Key your records on (f.id(), idx.exportConfigId()).}# The index points at its output directory.OUTPUT_DIRECTORY_ID=$(llp beta:indexes get --index-id "$INDEX_ID" | jq -r '.output_directory_id')
FILE_IDS=$(llp beta:directories:files list \ --directory-id "$OUTPUT_DIRECTORY_ID" | jq -r '.file_id // empty')
for FILE_ID in $FILE_IDS; do URL=$(llp files content --file-id "$FILE_ID" | jq -r '.url') curl -s "$URL" | jq '.parse'doneExport only what changed
Section titled “Export only what changed”Download and re-export every file on each run and you’ll do a lot of repeated work. After each sync, compare the output directory with what you last exported:
- New or changed files: pass the time of your last export as
updated_at_on_or_afterwhen listing the output directory. Only files that were parsed again, or whose metadata changed, come back. - Removed files: a sync deletes a removed source file’s output file. Any
parsed_directory_file_idin your database that’s missing from a full listing of the output directory is gone, and you should delete its records.
from datetime import datetime
last_export: datetime | None = load_last_export_time() # from your own bookkeepingrun_started = datetime.now().astimezone()
idx = await client.indexes.get(index.id)params = {"updated_at_on_or_after": last_export} if last_export else {}
async for f in client.beta.directories.files.list(idx.output_directory_id, **params): ... # download, chunk, embed, and replace this file's records
current_ids = { f.id async for f in client.beta.directories.files.list(idx.output_directory_id)}stale_ids = exported_file_ids() - current_ids # IDs you hold for idx.export_config_id... # delete the records for stale_ids
save_last_export_time(run_started)load_last_export_time, exported_file_ids, and save_last_export_time stand in for your own bookkeeping. The list_snapshots() method in the reference exporters returns the file IDs a sink currently holds.
Chunk, embed, and push to your store
Section titled “Chunk, embed, and push to your store”From there, the steps are the same as any custom RAG pipeline:
- Concatenate per-page markdown while tracking page boundaries by character offset.
- Chunk the text with your preferred chunker.
- Embed the chunks with your preferred embedding model.
- Write the chunks to your database, keyed on
(parsed_directory_file_id, export_config_id): the output directory file’sidand the index’sexport_config_id. Delete a file’s existing records before inserting its new chunks, so a file that’s parsed again replaces its earlier chunks instead of duplicating them.
Reference implementations
Section titled “Reference implementations”The index-v2-data-sinks repo contains complete, runnable reference exporters for several databases:
- MongoDB (Atlas vector search)
- PostgreSQL + pgvector
- Qdrant
- Pinecone
- Turbopuffer
- Azure AI Search
Each exporter handles index/collection provisioning, deterministic IDs for idempotent re-exports, per-file delete-then-insert, and snapshot listing for incremental syncs. The shared Exporter protocol in export/base.py is the contract to implement when adding a new sink:
class Exporter(Protocol): async def export(self, outputs, *, project_id, embeddings=None, embedding_model=None): ... async def delete_file(self, *, parsed_directory_file_id, export_config_id): ... async def delete_files(self, *, parsed_directory_file_ids, export_config_id): ... async def list_snapshots(self, *, export_config_id): ...Point your coding agent at the repo as a worked example, copy a sink you like as a starting point, or implement the protocol against the store of your choice.
Caveats
Section titled “Caveats”- Retrieval, chat, and file operations (
client.retrieval,client.chat) query the managed vector store. Withvector_target="DISABLED", those endpoints have nothing to query — you serve retrieval from your own database. - You’re responsible for cleaning up records in your database when files are removed from the source directory. See Export only what changed.