Embed Text
posthttps://api-atlas.nomic.ai/v1/embedding/text
Generates text embeddings
nomic-embed-text was trained to support these tasks:
search_document(embedding document chunks for search & retrieval)search_query(embedding queries for search & retrieval)classification(embeddings for text classification)clustering(embeddings for cluster visualization)
In the Nomic API or Python client, specify your task with the task_type parameter (default is search_document if no task_type is provided)
Using nomic-embed-text with other libraries requires you to use a prefix to specify your embedding task. See our HuggingFace model card for details.
Request
Body (required) application/json
textsarray<string>*requiredA batch of text you want embedded.
modelstringThe model to use when embedding.
Possible values: nomic-embed-text-v1, nomic-embed-text-v1.5
task_typestringThe downstream task to generate embeddings for. Options are search_document, search_query, classification, and clustering.
Default: "search_document"
long_text_modestringHow to handle text longer than the model can accept.
Possible values: truncate, mean
max_tokens_per_textintegerMaximum amount of tokens per text. Defaults to 8192 if long_text_mode is "mean", or the maximum model input size if long_text_mode is "truncate".
Default: 8192
dimensionalityintegerOptionally reduce embedding dimensionality. Defaults to full-size embeddings if unspecified. Only applies to nomic-embed-text-v1.5.
Default: 768
Responses
embeddingsarray<array<number>>*requiredThe embeddings
usageobject*requiredThe embedding usage
prompt_tokensinteger*requiredThe number of non-generated tokens ingested.
total_tokensinteger*requiredThe total tokens used.
modelstring*requiredThe model used to produce the embeddings.
Possible values: nomic-embed-text-v1, nomic-embed-text-v1.5
Examples
curl -X 'POST' \
'https://api-atlas.nomic.ai/v1/embedding/text' \
-H 'accept: application/json' \
-H 'Authorization: Bearer <token>' \
-H 'Content-Type: application/json' \
-d '{
"texts": [
"string"
],
"model": "nomic-embed-text-v1",
"task_type": "search_document",
"long_text_mode": "truncate",
"max_tokens_per_text": 8192,
"dimensionality": 768
}'
from nomic import embed
output = embed.text(
texts=['document 1', 'document 2'],
model='nomic-embed-text-v1.5',
task_type='search_document',
)
print(output)