Parse files into structured information
posthttps://api-atlas.nomic.ai/v1/parse
Parse a file into structured information.
Supports PDF files with configurable chunking strategies and optional embedding generation.
Request
Body (required) application/json
file_urlstring*requiredFile URL to process. Supports two URL types:
1. Public URLs - accessible from the internet
2. nomic:// prefixed URLs - obtained from the /upload endpoint
optionsobjectOptions to customize document parsing.
chunkingobjectOptions that control how the document is split into chunks.
chunk_modestringThe method by which the document is split into chunks.
Possible values: page
ocr_systemstringThe OCR method used to extract text from the document.
Possible values: standard
content_extraction_modestringThe overall strategy for extracting content from the document. metadata: Disable all OCR. Only use embedded document text. hybrid: Use a VLM for tables, and run an OCR model on all bitmaps found in the document. ocr: Use a VLM for tables. Run an OCR model on full pages.
Possible values: metadata, hybrid, ocr
table_summaryobjectOptions for generating table summaries.
enabledbooleanWhether to generate a summary of table content.
Default: false
figure_summaryobjectOptions for generating figure summaries.
enabledbooleanWhether to generate a summary of figure content.
Default: true
Responses
task_idstring*requiredThe id of the task.
Examples
curl -X 'POST' \
'https://api-atlas.nomic.ai/v1/parse' \
-H 'accept: application/json' \
-H 'Authorization: Bearer <token>' \
-H 'Content-Type: application/json' \
-d '{
"file_url": "string",
"options": {
"chunking": {
"chunk_mode": "page"
},
"ocr_system": "standard",
"content_extraction_mode": "metadata",
"table_summary": {
"enabled": false
},
"figure_summary": {
"enabled": true
}
}
}'