Gemini Embedding 2

Neural Network

Gemini Embedding 2 is a multimodal embedding model from Google for semantic search and RAG.

Main

/

Models

/

Gemini Embedding 2
2 703

Max answer length

(in tokens)

65 536

Context size

(in tokens)

23,57 ₽

Prompt cost

(per 1M tokens)

< 0.01 ₽

Image prompt

(per 1K tokens)

*Prices are shown for API usage via ECO providers.
Code example and API for Gemini Embedding 2We offer full access to the OpenAI API through our service. All our endpoints fully comply with OpenAI endpoints and can be used both with plugins and when developing your own software through the SDK.Create API key
Text generation
Embeddings
Javascript
Python
Curl
illustaration

How it works Gemini Embedding 2?

Gemini Embedding 2 is Google's first multimodal embedding model: it maps text and images into a single vector space, allowing image and text-based search to work within the same index. Via BotHub, the model is available through a unified OpenAI-compatible API: one key, pay-as-you-go in rubles, no VPN or foreign cards required, AES-GCM encryption, and access to over 250 models in one window. Vectors can be stored in Qdrant, pgvector, or any other vector database. Suitable for semantic search across knowledge bases and product catalogs, RAG assistants for documentation, clustering and deduplication of requests, and finding similar images and materials.

Frequently asked questions about Gemini Embedding 2

Can I use Gemini Embedding 2 results for commercial purposes?

You can use generated results for commercial purposes. You own all rights to the content you create. The only restriction: make sure your prompt does not include copyrighted third-party material. You are responsible for respecting the rights to any input data.

What can gemini-embedding-2 do and what tasks is it suitable for?

gemini-embedding-2 from google-gemini converts text and documents into numerical vectors. This is the foundation for semantic search, RAG responses based on knowledge bases, clustering, deduplication, and request classification. A context window of up to 65,536 tokens allows for vectorizing long fragments entirely without cutting them into small pieces.

In what quality and format does gemini-embedding-2 generate?

The output is not text, but a vector: it is stored in a vector database like Qdrant or pgvector and compared using cosine similarity. According to our data, the output limit is 2703 tokens; batch mode is available for mass indexing, and connection is via the unified OpenAI-compatible BotHub API.

Can gemini-embedding-2 reason before answering?

Reasoning mode is not noted in our data, and it is usually not required for embeddings: the model does not formulate an answer but encodes the meaning of the text into a vector. Logic on top of the results is built by your search or RAG pipeline, and the final answer is written by a chat model.

Does gemini-embedding-2 support function calling and JSON output?

Function calling and structured output are not noted in our data. For an embedding model, this is not a typical scenario: the API response already comes as a machine-readable array of numbers, which is immediately written to a vector database. It is more logical to assign function calling to a chat model alongside it, within the same API.

Can I upload images, audio, or video to gemini-embedding-2?

According to our data, the input accepts text, documents (including native formats), and images. Audio and video are not listed. Through the unified BotHub API, it is easy to build a pipeline where another model transcribes media, and this one handles vectorization.

How well does gemini-embedding-2 understand the Russian language?

There is no information about languages in our model specifications, so you should check the quality on your own data. Collect a small set of typical queries and documents, run indexing, and compare the output using cosine similarity — this is the most honest test for embeddings. Access from the Russian Federation without VPN, payment in rubles.

Support ServiceOpen from 10:00 to 18:00 MSK