Skip to main content

Embed Text

posthttps://api-atlas.nomic.ai/v1/embedding/text

Generates text embeddings

nomic-embed-text was trained to support these tasks:

  • search_document (embedding document chunks for search & retrieval)
  • search_query (embedding queries for search & retrieval)
  • classification (embeddings for text classification)
  • clustering (embeddings for cluster visualization)

In the Nomic API or Python client, specify your task with the task_type parameter (default is search_document if no task_type is provided)

Using nomic-embed-text with other libraries requires you to use a prefix to specify your embedding task. See our HuggingFace model card for details.

Request​

Body (required) application/json

textsarray<string>*required

A batch of text you want embedded.

modelstring

The model to use when embedding.

Possible values: nomic-embed-text-v1, nomic-embed-text-v1.5

task_typestring

The downstream task to generate embeddings for. Options are search_document, search_query, classification, and clustering.

Default: "search_document"

long_text_modestring

How to handle text longer than the model can accept.

Possible values: truncate, mean

max_tokens_per_textinteger

Maximum amount of tokens per text. Defaults to 8192 if long_text_mode is "mean", or the maximum model input size if long_text_mode is "truncate".

Default: 8192

dimensionalityinteger

Optionally reduce embedding dimensionality. Defaults to full-size embeddings if unspecified. Only applies to nomic-embed-text-v1.5.

Default: 768

Responses​

200Successful Response
embeddingsarray<array<number>>*required

The embeddings

usageobject*required

The embedding usage

prompt_tokensinteger*required

The number of non-generated tokens ingested.

total_tokensinteger*required

The total tokens used.

modelstring*required

The model used to produce the embeddings.

Possible values: nomic-embed-text-v1, nomic-embed-text-v1.5

422Validation Error

Examples​

curl
curl -X 'POST' \
'https://api-atlas.nomic.ai/v1/embedding/text' \
-H 'accept: application/json' \
-H 'Authorization: Bearer <token>' \
-H 'Content-Type: application/json' \
-d '{
"texts": [
"string"
],
"model": "nomic-embed-text-v1",
"task_type": "search_document",
"long_text_mode": "truncate",
"max_tokens_per_text": 8192,
"dimensionality": 768
}'
Nomic
from nomic import embed

output = embed.text(
texts=['document 1', 'document 2'],
model='nomic-embed-text-v1.5',
task_type='search_document',
)

print(output)