kosareva.cloud

Yandex Embeddings Query API — price in US dollars

Yandex · Text embeddings

Yandex Embeddings Query, the second half of the Yandex embedding pair: it turns search queries into vectors to match against a base built with Yandex Embeddings Doc. Connect it through the kosareva.cloud OpenAI-compatible API; companies outside Russia pay by USD invoice.

Modalities
→Input: text. Output: vector.
Price per 1M tokens
$0.17 input only
Fragment / vector
up to 8,192 tokens · 256 dimensions
Provider
Yandex

What Yandex Embeddings Query is

Yandex Embeddings Query is a Yandex text embedding model for the search side: it turns a user’s query into a vector of 256 numbers to match against documents embedded with Yandex Embeddings Doc. Input can be up to 8,192 tokens, though queries are usually a few dozen. Through kosareva.cloud it is available via one OpenAI-compatible API: the same /v1/embeddings call as for OpenAI — just change base_url and the key.

Yandex models are not available on OpenRouter as of October 2026. Here a company outside Russia gets Yandex Embeddings Query with an English invoice in US dollars from Kosareva Cloud LLC (Armenia) and pays by bank transfer. Requests to Yandex Embeddings Query are processed in Russia, on Yandex infrastructure.

One key opens more than Yandex Embeddings Query: the catalog has over a hundred models, including the other Yandex models. That matters more than it seems: models age within months, and moving to the next version is one line of code, not a new contract and a new integration.

Yandex Embeddings Query strengths and weaknesses

Where it is good

The same price as Doc, $0.17 per million tokens, and the same 256 dimensions, but trained for a different job: short queries and requests, not long texts. Paired with Yandex Embeddings Doc it gives noticeably more accurate search than one model on both sides. Requests are processed in Russia.

Where it falls short

Useless on its own: without a base built with Doc, there is nothing to match the queries against. The vector has 256 dimensions, as in the pair, with the same limits. There is no backup route, and replacing the model means rebuilding the base.

Who it suits. User queries to a knowledge base, a search box, a question in a chatbot — anything that searches the base.

What to compare it with. Compare it not with other models but with the “one model for both index and search” setup. The doc–query pair is split precisely because a document and a question about it are built differently.

How the Doc and Query pair works

Yandex publishes its embeddings as two models with the same price and the same 256-dimension vectors. Documents are embedded with Yandex Embeddings Doc, the user’s search queries with Yandex Embeddings Query, and the two sets of vectors are compared with each other. Embedding queries with Doc instead lowers search quality.

Both cost $0.17 per million tokens. Queries are short, so this half of the pair usually costs far less than indexing.

Yandex Embeddings Query price in US dollars

Input$0.17 per 1M tokens
Outputnot billed
Details256 dimensions · up to 8,192 tokens per fragment · for search queries

Approximate: the catalog price (14.14 ₽ per 1M tokens) divided by 85.5 ₽ per $1, ID Bank’s non-cash rate on 3 October 2026. Your balance is kept in rubles; a USD invoice is converted at ID Bank’s non-cash rate on the invoice date. You pay only for the text you send; the vectors you get back cost nothing extra.

What search costs on Yandex Embeddings Query

A search query is usually a few dozen tokens. Here are three examples at $0.17 per million tokens, with 30-token queries.

One search query30 tokens<$0.01
100,000 queries a month3,000,000 tokens$0.50
1,000,000 queries a month30,000,000 tokens$4.96

Calculated from the current kosareva.cloud price list, in US dollars at 85.5 ₽ per $1. For comparison, indexing a 10,000-page knowledge base (about 7,000,000 tokens) with Yandex Embeddings Doc costs about $1.16.

Yandex Embeddings Query cost calculator

Set how many queries you run a month and how long they are — the monthly cost in US dollars updates at once, with no sign-up.

Search queries per month
01,000,000
Tokens per query
01,000
Per query
—
Tokens per month
—
Output
not billed
Monthly cost
—
Get an API key

At $0.17 per million tokens; approximate, at 85.5 ₽ per $1. Embeddings have no output price. Billed for actual usage.

How to connect Yandex Embeddings Query

Use the standard OpenAI SDK with the kosareva.cloud base_url and the model yandex-embeddings-query. The call is embeddings.create, not chat. In existing code you change exactly two lines: the address and the key.

Python, OpenAI SDK

from openai import OpenAI

client = OpenAI(
    api_key="YOUR_KOSAREVA_KEY",
    base_url="https://api.kosareva.cloud/v1",
)

response = client.embeddings.create(
    model="yandex-embeddings-query",
    input=[
        "refund policy for corporate clients",
        "how to change the invoice details",
    ],
)
for item in response.data:
    print(len(item.embedding))  # vector length

Python without the SDK: requests

import requests

response = requests.post(
    "https://api.kosareva.cloud/v1/embeddings",
    headers={"Authorization": "Bearer YOUR_KOSAREVA_KEY"},
    json={"model": "yandex-embeddings-query", "input": ["refund policy for corporate clients", "how to change the invoice details"]},
    timeout=120,
)
vectors = [row["embedding"] for row in response.json()["data"]]
print(len(vectors), len(vectors[0]))

cURL

curl https://api.kosareva.cloud/v1/embeddings \
  -H "Authorization: Bearer YOUR_KOSAREVA_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"yandex-embeddings-query","input":"refund policy for corporate clients"}'

The same key works with any client that supports an OpenAI-compatible embeddings endpoint: LangChain, LlamaIndex, n8n. You don’t need a separate key per model — one key opens the whole catalog.

Yandex Embeddings Query endpoints and parameters

The gateway follows the OpenAI schema, so parameter names and the response format are the ones you know. One base address for every model in the catalog.

base_urlhttps://api.kosareva.cloud/v1
Endpoint/embeddings
Model nameyandex-embeddings-query
AuthorizationAuthorization: Bearer <key>
Vector size256 numbers per fragment
Main parametersinput — a string or a list of strings per request
Model listGET /v1/models

A user query is usually one string per request. input also takes a list, which is handy when you embed a set of test queries to check search quality. Compare the query vector with the document vectors built by Yandex Embeddings Doc, usually by cosine similarity.

Specifications

ProviderYandex
TypeText embeddings
Fragment lengthUp to 8,192 tokens
Vector size256 dimensions
InputText
OutputVector
Role in the pairSearch queries (documents go to Doc)
APIOpenAI-compatible, /v1/embeddings
ProcessingIn Russia

What Yandex Embeddings Query is good for

  • Search box queries and chatbot questions against a knowledge base
  • The query side of semantic search and RAG over documents in Russian, together with Yandex Embeddings Doc

What to keep in mind with Yandex Embeddings Query

Queries are short, so this half of the pair usually costs dozens of times less than indexing: you pay mainly for building the base, and search costs almost nothing. Budget by the volume of documents, not by the number of queries.

First, 8,192 tokens per input is a hard limit: a longer input returns an error rather than being cut. Ordinary queries never come near it, but if you build the query from a chat history, trim the history first.

Second, vectors from different models are incompatible. Doc and Query vectors are compared with each other, but not with vectors from any other model: switching models later means re-embedding the whole base.

Third, the balance is shared by all models in the catalog. You don’t have to decide in advance how much money to “put on Yandex Embeddings Query”: the same funds go to any other model if you switch tomorrow. Usage per model is shown in the dashboard on a separate line, with the number of requests and tokens.

Yandex Embeddings Query for companies outside Russia

Yandex Embeddings Query, the search half of the Yandex embedding pair: 256-dimension vectors for queries against a base built with Embeddings Doc, requests processed in Russia. Access through kosareva.cloud with an OpenAI-compatible API and a USD invoice for companies.

  • Register as “Company or sole proprietor” → “In another country”: country of registration, company registration number, tax ID if you have one.
  • In “Top up balance”, “Issue invoice (companies)” gives an English invoice in US dollars from Kosareva Cloud LLC (Armenia): minimum $50, due in 10 days.
  • Pay by bank transfer to our USD account at ID Bank. The balance is credited after the bank confirms the payment, usually in 1–3 business days.
  • Cards issued outside Russia are not accepted. Invoices are in USD; for EUR, write to us.
  • Pay only for actual usage, no subscription. One OpenAI-compatible API for Yandex Embeddings Query and over a hundred other models with one key.

Requests to Yandex Embeddings Query are processed in Russia, on Yandex infrastructure. If you are in the EU or send personal data of EU residents, assess the GDPR side before you send it. You are responsible for compliance with the rules of your jurisdiction.

Other Yandex models

Price per 1M input tokens. The full list with output prices — in the US dollar price table.

Yandex Embeddings Query questions

How much does Yandex Embeddings Query cost?

$0.17 per 1M input tokens: 14.14 ₽ in the catalog, converted at 85.5 ₽ per $1 (ID Bank’s non-cash rate on 3 October 2026). There is no output price: you pay only for the text you send.

How do I connect Yandex Embeddings Query by API?

Register at kosareva.cloud, create a key, set base_url https://api.kosareva.cloud/v1 and the model yandex-embeddings-query in the standard OpenAI SDK, and call the embeddings endpoint.

Do I also need Yandex Embeddings Doc?

Yes. Query embeds search queries; the documents you search must be embedded with Yandex Embeddings Doc. The two are built to be used together.

How long can a query be?

Up to 8,192 tokens; a longer input returns an error. Each query becomes a vector of 256 numbers.

Who is Yandex Embeddings Query for, and where is it weaker?

User queries to a knowledge base, a search box, a question in a chatbot. On its own it is useless: without a base built with Doc there is nothing to match the queries against.

How much does search cost compared with indexing?

Far less. 100,000 queries of 30 tokens cost about $0.50 a month, while indexing a 10,000-page base with Doc costs about $1.16. Budget by the volume of documents, not by the number of queries.

Where are requests to Yandex Embeddings Query processed?

In Russia, on Yandex infrastructure. If GDPR applies to you, assess the transfer before sending personal data. You are responsible for compliance with the rules of your jurisdiction.

How does a company outside Russia pay?

Register as “Company or sole proprietor” → “In another country”, then click “Issue invoice (companies)” in “Top up balance”. You get an English invoice in US dollars from Kosareva Cloud LLC (Armenia), minimum $50, due in 10 days. The balance is credited after the bank confirms the payment, usually in 1–3 business days. Cards issued outside Russia are not accepted.

Can I try Yandex Embeddings Query on a small amount?

The minimum invoice is $50. Usage is billed per token with no subscription, and the unused balance does not expire.

Connect Yandex Embeddings Query in a couple of minutes

Get an API key and call Yandex Embeddings Query through an OpenAI-compatible embeddings endpoint. Companies outside Russia pay by USD invoice. Questions: ceo@kosareva.cloud or @kosareva_cloud in Telegram.