Skip to main content

Parse files into structured information

posthttps://api-atlas.nomic.ai/v1/parse

Parse a file into structured information.

Supports PDF files with configurable chunking strategies and optional embedding generation.

Request​

Body (required) application/json

file_urlstring*required

File URL to process. Supports two URL types: 1. Public URLs - accessible from the internet 2. nomic:// prefixed URLs - obtained from the /upload endpoint

optionsobject

Options to customize document parsing.

chunkingobject

Options that control how the document is split into chunks.

chunk_modestring

The method by which the document is split into chunks.

Possible values: page

ocr_systemstring

The OCR method used to extract text from the document.

Possible values: standard

content_extraction_modestring

The overall strategy for extracting content from the document. metadata: Disable all OCR. Only use embedded document text. hybrid: Use a VLM for tables, and run an OCR model on all bitmaps found in the document. ocr: Use a VLM for tables. Run an OCR model on full pages.

Possible values: metadata, hybrid, ocr

table_summaryobject

Options for generating table summaries.

enabledboolean

Whether to generate a summary of table content.

Default: false

figure_summaryobject

Options for generating figure summaries.

enabledboolean

Whether to generate a summary of figure content.

Default: true

Responses​

201The task id of the parsing task.
task_idstring*required

The id of the task.

403The user is not authorized to perform this action.
422Validation Error

Examples​

curl
curl -X 'POST' \
'https://api-atlas.nomic.ai/v1/parse' \
-H 'accept: application/json' \
-H 'Authorization: Bearer <token>' \
-H 'Content-Type: application/json' \
-d '{
"file_url": "string",
"options": {
"chunking": {
"chunk_mode": "page"
},
"ocr_system": "standard",
"content_extraction_mode": "metadata",
"table_summary": {
"enabled": false
},
"figure_summary": {
"enabled": true
}
}
}'