kosareva.cloud

Yandex Embeddings Doc API — price in US dollars

Yandex · Text embeddings

Yandex Embeddings Doc turns documents into vectors for semantic search: knowledge bases, RAG, deduplication. It indexes the documents; search queries go to its pair, Yandex Embeddings Query. Connect it through the kosareva.cloud OpenAI-compatible API; companies outside Russia pay by USD invoice.

Modalities
→Input: text. Output: vector.
Price per 1M tokens
$0.17 input only
Fragment / vector
up to 8,192 tokens · 256 dimensions
Provider
Yandex

What Yandex Embeddings Doc is

Yandex Embeddings Doc is a Yandex text embedding model: it turns a document fragment of up to 8,192 tokens into a vector of 256 numbers for semantic search. It works in a pair with Yandex Embeddings Query: Doc embeds the documents, Query embeds the search queries against them. Through kosareva.cloud it is available via one OpenAI-compatible API: the same /v1/embeddings call as for OpenAI — just change base_url and the key.

Yandex models are not available on OpenRouter as of October 2026. Here a company outside Russia gets Yandex Embeddings Doc with an English invoice in US dollars from Kosareva Cloud LLC (Armenia) and pays by bank transfer. Requests to Yandex Embeddings Doc are processed in Russia, on Yandex infrastructure.

One key opens more than Yandex Embeddings Doc: the catalog has over a hundred models, including the other Yandex models. That matters more than it seems: models age within months, and moving to the next version is one line of code, not a new contract and a new integration.

Yandex Embeddings Doc strengths and weaknesses

Where it is good

$0.17 per million tokens, the lowest-priced embedding model in our catalog. Fragments of up to 8,192 tokens: we checked — it answers on 6,008 tokens and returns an error naming the limit on 12,008. That is four times longer than Gemini Embedding 2, so a long article or a section of a policy can be embedded whole, without chunking. Requests are processed in Russia.

Where it falls short

The vector is short: 256 dimensions against 3,072 for Gemini Embedding 2. A short vector saves storage and speeds up search, but separates texts with close meanings less well: on a base with hundreds of similar documents the difference will show. This is half of a pair: it embeds documents, queries to them go to Yandex Embeddings Query, and the two roles must not be mixed. There is no backup route, and you cannot swap in another model: vectors from different models are incompatible, so a replacement means rebuilding the base.

Who it suits. Indexing a knowledge base: policies, contracts, articles, product cards. Anything you put into the base once and then search.

What to compare it with. Against Gemini Embedding 2 ($0.25), also in our catalog: one and a half times cheaper and takes fragments four times longer, but the vector is twelve times shorter. Choose Yandex when price, fragment length and processing in Russia matter; Gemini when the base is large and you need fine distinctions.

How the Doc and Query pair works

Yandex publishes its embeddings as two models with the same price and the same 256-dimension vectors. Documents are embedded with Yandex Embeddings Doc, the user’s search queries with Yandex Embeddings Query, and the two sets of vectors are compared with each other. Indexing a base and searching it with Doc alone lowers search quality.

Both cost $0.17 per million tokens. In practice most of the spend goes to Doc: documents are long, queries are short.

Yandex Embeddings Doc price in US dollars

Input$0.17 per 1M tokens
Outputnot billed
Details256 dimensions · up to 8,192 tokens per fragment · for indexing documents

Approximate: the catalog price (14.14 ₽ per 1M tokens) divided by 85.5 ₽ per $1, ID Bank’s non-cash rate on 3 October 2026. Your balance is kept in rubles; a USD invoice is converted at ID Bank’s non-cash rate on the invoice date. You pay only for the text you send; the vectors you get back cost nothing extra.

What indexing costs on Yandex Embeddings Doc

A page of Russian text is about 700 tokens. Here are three common jobs at $0.17 per million tokens.

One 30-page contract40,000 tokens$0.01
Monthly re-index of 1,000 changed pagesabout 700,000 tokens$0.12
A knowledge base of 10,000 pagesabout 7,000,000 tokens$1.16

Calculated from the current kosareva.cloud price list, in US dollars at 85.5 ₽ per $1.

Yandex Embeddings Doc cost calculator

Set how many fragments you embed a month and how long they are — the monthly cost in US dollars updates at once, with no sign-up.

Fragments per month
01,000,000
Tokens per fragment
08,192
Per fragment
—
Tokens per month
—
Output
not billed
Monthly cost
—
Get an API key

At $0.17 per million tokens; approximate, at 85.5 ₽ per $1. Embeddings have no output price. Billed for actual usage.

How to connect Yandex Embeddings Doc

Use the standard OpenAI SDK with the kosareva.cloud base_url and the model yandex-embeddings-doc. The call is embeddings.create, not chat. In existing code you change exactly two lines: the address and the key.

Python, OpenAI SDK

from openai import OpenAI

client = OpenAI(
    api_key="YOUR_KOSAREVA_KEY",
    base_url="https://api.kosareva.cloud/v1",
)

response = client.embeddings.create(
    model="yandex-embeddings-doc",
    input=[
        "First document fragment",
        "Second document fragment",
    ],
)
for item in response.data:
    print(len(item.embedding))  # vector length

Python without the SDK: requests

import requests

response = requests.post(
    "https://api.kosareva.cloud/v1/embeddings",
    headers={"Authorization": "Bearer YOUR_KOSAREVA_KEY"},
    json={"model": "yandex-embeddings-doc", "input": ["First document fragment", "Second document fragment"]},
    timeout=120,
)
vectors = [row["embedding"] for row in response.json()["data"]]
print(len(vectors), len(vectors[0]))

cURL

curl https://api.kosareva.cloud/v1/embeddings \
  -H "Authorization: Bearer YOUR_KOSAREVA_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"yandex-embeddings-doc","input":"First document fragment"}'

The same key works with any client that supports an OpenAI-compatible embeddings endpoint: LangChain, LlamaIndex, n8n. You don’t need a separate key per model — one key opens the whole catalog.

Yandex Embeddings Doc endpoints and parameters

The gateway follows the OpenAI schema, so parameter names and the response format are the ones you know. One base address for every model in the catalog.

base_urlhttps://api.kosareva.cloud/v1
Endpoint/embeddings
Model nameyandex-embeddings-doc
AuthorizationAuthorization: Bearer <key>
Vector size256 numbers per fragment
Main parametersinput — a string or a list of strings per request
Model listGET /v1/models

Send fragments in batches: input takes a list of strings, and one request with a hundred fragments is faster and has less overhead than a hundred requests with one each. Split documents into pieces of 200–500 words: a longer fragment loses precision, a shorter one loses context.

Specifications

ProviderYandex
TypeText embeddings
Fragment lengthUp to 8,192 tokens
Vector size256 dimensions
InputText
OutputVector
Role in the pairDocuments (search queries go to Query)
APIOpenAI-compatible, /v1/embeddings
ProcessingIn Russia

What Yandex Embeddings Doc is good for

  • Indexing knowledge bases: policies, contracts, articles, product cards
  • Semantic search and RAG over documents in Russian, together with Yandex Embeddings Query
  • Finding duplicates and near-duplicates among documents

What to keep in mind with Yandex Embeddings Doc

The Doc and Query pair is not a vendor whim: documents and search queries are embedded differently, and if you index the base with this model and search with it too, search quality drops. Documents go to this model, queries to Yandex Embeddings Query. And settle the vector size in your database schema at the start: changing it later means rebuilding the whole base.

First, 8,192 tokens per fragment is a hard limit: we checked, and a longer fragment returns an error that names the limit rather than being cut. Split long documents before you send them.

Second, vectors from different models are incompatible. Doc and Query vectors are compared with each other, but not with vectors from any other model: switching models later means re-embedding the whole base.

Third, the balance is shared by all models in the catalog. You don’t have to decide in advance how much money to “put on Yandex Embeddings Doc”: the same funds go to any other model if you switch tomorrow. Usage per model is shown in the dashboard on a separate line, with the number of requests and tokens.

Yandex Embeddings Doc for companies outside Russia

Yandex Embeddings Doc, the document half of the Yandex embedding pair: 256-dimension vectors, up to 8,192 tokens per fragment, requests processed in Russia. Access through kosareva.cloud with an OpenAI-compatible API and a USD invoice for companies.

  • Register as “Company or sole proprietor” → “In another country”: country of registration, company registration number, tax ID if you have one.
  • In “Top up balance”, “Issue invoice (companies)” gives an English invoice in US dollars from Kosareva Cloud LLC (Armenia): minimum $50, due in 10 days.
  • Pay by bank transfer to our USD account at ID Bank. The balance is credited after the bank confirms the payment, usually in 1–3 business days.
  • Cards issued outside Russia are not accepted. Invoices are in USD; for EUR, write to us.
  • Pay only for actual usage, no subscription. One OpenAI-compatible API for Yandex Embeddings Doc and over a hundred other models with one key.

Requests to Yandex Embeddings Doc are processed in Russia, on Yandex infrastructure. If you are in the EU or send personal data of EU residents, assess the GDPR side before you send it. You are responsible for compliance with the rules of your jurisdiction.

Other Yandex models

Price per 1M input tokens. The full list with output prices — in the US dollar price table.

Yandex Embeddings Doc questions

How much does Yandex Embeddings Doc cost?

$0.17 per 1M input tokens: 14.14 ₽ in the catalog, converted at 85.5 ₽ per $1 (ID Bank’s non-cash rate on 3 October 2026). There is no output price: you pay only for the text you send.

How do I connect Yandex Embeddings Doc by API?

Register at kosareva.cloud, create a key, set base_url https://api.kosareva.cloud/v1 and the model yandex-embeddings-doc in the standard OpenAI SDK, and call the embeddings endpoint.

How long can a fragment be?

Up to 8,192 tokens; a longer fragment returns an error. Each fragment becomes a vector of 256 numbers.

Which model embeds the search queries?

Yandex Embeddings Query. Index documents with Doc and embed queries with Query: the pair is built to be used together, and the price is the same, $0.17 per million tokens.

Who is Yandex Embeddings Doc for, and where is it weaker?

Indexing a knowledge base: policies, contracts, articles, product cards. Its vector is short, 256 dimensions, so on a base with many similar documents it separates them less well than models with longer vectors.

Which models should I compare Yandex Embeddings Doc with?

With Gemini Embedding 2 ($0.25), also in our catalog: Yandex Embeddings Doc is one and a half times cheaper and takes fragments four times longer, but its vector is twelve times shorter.

Where are requests to Yandex Embeddings Doc processed?

In Russia, on Yandex infrastructure. If GDPR applies to you, assess the transfer before sending personal data. You are responsible for compliance with the rules of your jurisdiction.

How does a company outside Russia pay?

Register as “Company or sole proprietor” → “In another country”, then click “Issue invoice (companies)” in “Top up balance”. You get an English invoice in US dollars from Kosareva Cloud LLC (Armenia), minimum $50, due in 10 days. The balance is credited after the bank confirms the payment, usually in 1–3 business days. Cards issued outside Russia are not accepted.

Can I try Yandex Embeddings Doc on a small amount?

The minimum invoice is $50. Usage is billed per token with no subscription, and the unused balance does not expire. Indexing a 10,000-page knowledge base costs about $1.16.

Connect Yandex Embeddings Doc in a couple of minutes

Get an API key and call Yandex Embeddings Doc through an OpenAI-compatible embeddings endpoint. Companies outside Russia pay by USD invoice. Questions: ceo@kosareva.cloud or @kosareva_cloud in Telegram.