EmbeddingGemma 2 โ€” Text Only โ€” FP32 โ€” Transformers / SentenceTransformers

A modular deployment export of Google DeepMind's EmbeddingGemma 2. Only the text backbone is retained; the audio encoder and audio projection are absent, as are the vision encoder and projection. No training, distillation, or quantization was applied.

FP32 weights are a lossless expansion of the upstream BF16 values. This does not recover precision absent from the original trained checkpoint.

Property Value
Runtime Transformers / SentenceTransformers
Stored floating-point tensors FP32 (all tensors checked)
Effective model parameters 271,002,624
Weight files 1084.06 MB, decimal
Inputs Text and code
Output 768 dimensions; MRL at 512, 256, 128
Context budget 8,192 tokens
License Apache 2.0

Weight size is not total runtime memory. 270M/440M in repository names are rounded deployment sizes.

Install

pip install "transformers>=5.19.0" "sentence-transformers>=6.1.0"

Tested with Transformers 5.19.0, SentenceTransformers 6.1.0, and PyTorch 2.14.1.

Text and code

import torch
from sentence_transformers import SentenceTransformer

model = SentenceTransformer(
    "jayyun98/embeddinggemma-2-text-270m-fp32",
    device="cpu",
    model_kwargs={"dtype": torch.float32},
)
queries = model.encode(
    ["What causes the northern lights?", "๋กœ์ปฌ ์ฝ”๋“œ ๊ฒ€์ƒ‰ ๋ชจ๋ธ์„ ์ฐพ๊ณ  ์‹ถ์–ด์š”."],
    prompt_name="SearchQuery",
    normalize_embeddings=True,
)
documents = model.encode(
    ["Charged particles from the sun cause the northern lights."],
    prompt_name="Document",
    normalize_embeddings=True,
)
print(model.similarity(queries, documents))

code_query = model.encode(
    "Find a Python function that sorts a list.",
    prompt_name="CodeRetrieval",
    truncate_dim=256,
    normalize_embeddings=True,
)

Use SearchQuery for search queries, CodeRetrieval for code-search queries, and Document for corpus items. For the MLX API, prepend the corresponding literal prefix from config_sentence_transformers.json, as shown above. For titled documents, use title: {title} | text: {content} without another prefix. Use matching dimensions for queries and documents, and normalize after truncation. This FP32 package stores FP32 weights; loading it as BF16 changes runtime precision.

Verification

CPU FP32 inference was tested in the package's named precision. Outputs were exactly equal to the full source model running in the same precision. Checks cover English/Korean search and document text, code queries, and 128/256/512-dimensional normalized vectors.

Fixture Minimum cosine vs source FP32 Maximum absolute difference
search_english_korean 1.000000000 0
documents 1.000000000 0
code 1.000000000 0

Measurements compare with the pinned full Google checkpoint. They are small numerical and loading checks, not MTEB results, a retrieval-quality evaluation, or a speed benchmark. The original prompts, pooling behavior, tokenizer and processor assets are retained. Processor metadata does not restore the removed encoder weights. Image, audio and video encoders are unavailable.

Attribution and license

Original weights and tokenizer/processor assets: Google DeepMind.

This is an independent derivative deployment export, not an official Google release. The Apache 2.0 LICENSE and NOTICE are included. See the original model card for training, intended use and limitations.

Downloads last month
11
Safetensors
Model size
0.3B params
Tensor type
F32
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for jayyun98/embeddinggemma-2-text-270m-fp32

Finetuned
(22)
this model

Collection including jayyun98/embeddinggemma-2-text-270m-fp32