kosareva.cloud

Alice AI Flash API — price in US dollars

Yandex · Chat and text generation

Alice AI Flash, the fast version of Alice by Yandex: a fifth of the flagship’s input price, a cache discount and a 65K context. Requests are processed in Russia. Connect it through the kosareva.cloud OpenAI-compatible API; companies outside Russia pay by USD invoice.

Modalities
→Input: text. Output: text.
Input / output per 1M
$1.64 / $3.27
Context
65K tokens
Provider
Yandex

What Alice AI Flash is

Alice AI Flash is the fast version of Alice, the Yandex language model. It has a 65,536-token context window, takes text in and gives text out, and belongs to the chat and text generation category. Through kosareva.cloud it is available via one OpenAI-compatible API: the same code as for OpenAI — just change base_url and the key.

Yandex models are not available on OpenRouter as of October 2026. Here a company outside Russia gets Alice AI Flash with an English invoice in US dollars from Kosareva Cloud LLC (Armenia) and pays by bank transfer. Requests to Alice AI Flash are processed in Russia, on Yandex infrastructure.

One key opens more than Alice AI Flash: the catalog has over a hundred models, including the other Yandex models. That matters more than it seems: models age within months, and moving to the next version is one line of code, not a new contract and a new integration.

Alice AI Flash strengths and weaknesses

Where it is good

$1.64 / $3.27 per million: five times cheaper than Alice AI on input and six times on output. Of the two Alice models, only this one has a real cache discount: $0.41 per million cached input tokens instead of $1.64. The same natural Russian.

Where it falls short

Half the context of Alice AI — 65,536 tokens, so a long document will not fit whole. Simpler answers: complex reasoning and code are better left to another model. The vendor does not publish the maximum answer length. There is no backup route.

Who it suits. High-volume tasks in Russian: chat answers, short texts, request classification, email drafts. Anything where speed and volume matter more than depth.

What to compare it with. Against Alice AI ($8.19 / $19.65): five times cheaper on input and simpler; for a high-volume stream the difference in quality is rarely worth paying five times more. Against YandexGPT 5 Lite ($3.27 for both input and output): half the input price, the same output price, twice the context and a cache discount.

Alice AI Flash among the Yandex models

Alice AI Flash has the lowest input price of the Yandex chat models in our catalog, $1.64 per million tokens, and shares the lowest output price, $3.27, with YandexGPT 5 Lite. Its 65K context is half of Alice AI’s 131K and twice the 32K of YandexGPT 5.1 Pro, 5 Pro and 5 Lite. It is also the only Yandex model with a cache discount.

Alice AI Flash price in US dollars

Input$1.64 per 1M tokens
Output$3.27 per 1M tokens
Cached input (read)$0.41 per 1M tokens
Cache writeat the regular input price

Approximate: the catalog price (140 ₽, 280 ₽ and 35 ₽ per 1M tokens) divided by 85.5 ₽ per $1, ID Bank’s non-cash rate on 3 October 2026. Your balance is kept in rubles; a USD invoice is converted at ID Bank’s non-cash rate on the invoice date. You pay only for actual usage.

What a typical task costs on Alice AI Flash

A price per million tokens says little on its own, so here are four common tasks at Alice AI Flash rates: $1.64 per million input tokens and $3.27 per million output tokens.

Chat or bot answer1,000 tokens in, 500 out<$0.01
30-page document review40,000 tokens in, a 2,000-token summary out$0.07
Article or sales proposal2,000 tokens of brief in, 4,000 out$0.02
A month of support: 1,000 requests1,500 tokens in and 400 out each$3.77

Calculated from the current kosareva.cloud price list, in US dollars at 85.5 ₽ per $1, without the cache discount: when the start of the prompt repeats, cached input lowers these sums further.

Alice AI Flash cost calculator

Move the sliders to your volume — the monthly cost in US dollars updates at once, with no sign-up.

Requests per month
0100,000
Input tokens
060,000
Output tokens
032,000
Input
—
Output
—
Per request
—
Tokens per month
—
Monthly cost
—
Get an API key

At $1.64 per million input tokens and $3.27 per million output tokens; approximate, at 85.5 ₽ per $1. Billed for actual usage.

How to connect Alice AI Flash

Use the standard OpenAI SDK with the kosareva.cloud base_url and the model alice-ai-flash. In existing code you change exactly two lines: the address and the key.

Python, OpenAI SDK

from openai import OpenAI

client = OpenAI(
    api_key="YOUR_KOSAREVA_KEY",
    base_url="https://api.kosareva.cloud/v1",
)

response = client.chat.completions.create(
    model="alice-ai-flash",
    messages=[{"role": "user", "content": "Hello!"}],
)
print(response.choices[0].message.content)

Python without the SDK: requests

import requests

response = requests.post(
    "https://api.kosareva.cloud/v1/chat/completions",
    headers={"Authorization": "Bearer YOUR_KOSAREVA_KEY"},
    json={
        "model": "alice-ai-flash",
        "messages": [{"role": "user", "content": "Hello!"}],
    },
    timeout=600,
)
print(response.json()["choices"][0]["message"]["content"])

cURL

curl https://api.kosareva.cloud/v1/chat/completions \
  -H "Authorization: Bearer YOUR_KOSAREVA_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"alice-ai-flash","messages":[{"role":"user","content":"Hello!"}]}'

The same key works with any client that supports an OpenAI-compatible endpoint: LangChain, LlamaIndex, Cursor, Continue, n8n. You don’t need a separate key per model — one key opens the whole catalog.

Alice AI Flash endpoints and parameters

The gateway follows the OpenAI schema, so parameter names and the response format are the ones you know. One base address for every model in the catalog.

base_urlhttps://api.kosareva.cloud/v1
Endpoint/chat/completions
Model namealice-ai-flash
AuthorizationAuthorization: Bearer <key>
Streamingstream: true — the answer arrives as it is generated
Main parameterstemperature, max_tokens, top_p, stop, tools
Model listGET /v1/models

Request long answers with stream: true: with a large max_tokens, a regular request stays silent until generation ends and can hit your HTTP client’s timeout. In streaming mode the first tokens arrive within seconds, and the timeout problem goes away.

Specifications

ProviderYandex
TypeChat and text generation
Context65K tokens
InputText
OutputText
APIOpenAI-compatible
ProcessingIn Russia

What Alice AI Flash is good for

  • Chatbots, assistants, Q&A and customer support in Russian
  • Copy, newsletters, product descriptions and content marketing in Russian

What to keep in mind with Alice AI Flash

Here the cache works, and that is its main advantage over Alice AI: keep the system prompt and the fixed part of the context unchanged at the start of the request, and repeated input is billed at $0.41 per million instead of $1.64 — a quarter of the price. On a stream of similar requests this makes a visible difference to the bill.

First, context runs out faster than it seems. In a dialogue the whole conversation goes into the input, not just the last message, so a long chat gets more expensive with every turn. If the history isn’t needed for the answer, trim it: savings on input tokens usually matter more than picking a cheaper model.

Second, the 65K-token window is shared between the request and the answer: the more you send, the less room the model has to generate. Set max_tokens on purpose, not “with a margin”.

Third, the balance is shared by all models in the catalog. You don’t have to decide in advance how much money to “put on Alice AI Flash”: the same funds go to any other model if you switch tomorrow. Usage per model is shown in the dashboard on a separate line, with the number of requests and tokens.

Alice AI Flash for companies outside Russia

Alice AI Flash by Yandex, the fast version of Alice: 65K context, cache discount, requests processed in Russia. Access through kosareva.cloud with an OpenAI-compatible API and a USD invoice for companies.

  • Register as “Company or sole proprietor” → “In another country”: country of registration, company registration number, tax ID if you have one.
  • In “Top up balance”, “Issue invoice (companies)” gives an English invoice in US dollars from Kosareva Cloud LLC (Armenia): minimum $50, due in 10 days.
  • Pay by bank transfer to our USD account at ID Bank. The balance is credited after the bank confirms the payment, usually in 1–3 business days.
  • Cards issued outside Russia are not accepted. Invoices are in USD; for EUR, write to us.
  • Pay only for actual usage, no subscription. One OpenAI-compatible API for Alice AI Flash and over a hundred other models with one key.

Requests to Alice AI Flash are processed in Russia, on Yandex infrastructure. If you are in the EU or send personal data of EU residents, assess the GDPR side before you send it. You are responsible for compliance with the rules of your jurisdiction.

Other Yandex models

Price per 1M input tokens. The full list with output prices — in the US dollar price table.

Alice AI Flash questions

How much does Alice AI Flash cost?

$1.64 per 1M input tokens and $3.27 per 1M output tokens; cached input is $0.41 per 1M. In the catalog that is 140 ₽, 280 ₽ and 35 ₽, converted at 85.5 ₽ per $1 (ID Bank’s non-cash rate on 3 October 2026). You pay only for the tokens you use.

How do I connect Alice AI Flash by API?

Register at kosareva.cloud, create a key, set base_url https://api.kosareva.cloud/v1 and the model alice-ai-flash in the standard OpenAI SDK.

What is the context size of Alice AI Flash?

65,536 tokens, shared between the prompt and the answer.

Does Alice AI Flash have a cache discount?

Yes. When the start of the request repeats (system prompt, fixed instructions), that part is billed at $0.41 per million tokens instead of $1.64. Writing to the cache costs the regular input price.

Who is Alice AI Flash for, and where is it weaker?

High-volume tasks in Russian: chat answers, short texts, request classification, email drafts. It has half the context of Alice AI and gives simpler answers; complex reasoning and code are better left to another model.

Which models should I compare Alice AI Flash with?

With Alice AI ($8.19 / $19.65): Alice AI Flash is five times cheaper on input and simpler. With YandexGPT 5 Lite ($3.27 for input and output): Alice AI Flash costs half as much on input, the same on output, and has twice the context and a cache discount.

Where are requests to Alice AI Flash processed?

In Russia, on Yandex infrastructure. If GDPR applies to you, assess the transfer before sending personal data. You are responsible for compliance with the rules of your jurisdiction.

How does a company outside Russia pay?

Register as “Company or sole proprietor” → “In another country”, then click “Issue invoice (companies)” in “Top up balance”. You get an English invoice in US dollars from Kosareva Cloud LLC (Armenia), minimum $50, due in 10 days. The balance is credited after the bank confirms the payment, usually in 1–3 business days. Cards issued outside Russia are not accepted.

Can I try Alice AI Flash on a small amount?

The minimum invoice is $50. Usage is billed per token with no subscription, and the unused balance does not expire.

Connect Alice AI Flash in a couple of minutes

Get an API key and call Alice AI Flash through an OpenAI-compatible endpoint. Companies outside Russia pay by USD invoice. Questions: ceo@kosareva.cloud or @kosareva_cloud in Telegram.