feat(python): Add gemini text embedding function (#806)

Named it Gemini-text for now. Not sure how complicated it will be to support both text and multimodal embeddings under the same class "gemini"..But its not something to worry about for now I guess.
2026-07-07 21:10:41 +00:00 · 2024-01-13 12:08:55 +05:30
parent f0a654036e
commit 2f72d5138e
4 changed files with 187 additions and 2 deletions
--- a/docs/src/embeddings/default_embedding_functions.md
+++ b/docs/src/embeddings/default_embedding_functions.md
@@ -118,6 +118,42 @@ texts = [{"text": "Capitalism has been dominant in the Western world since the e
 tbl.add(texts)
 ```

+## Gemini Embedding Function
+With Google's Gemini, you can represent text (words, sentences, and blocks of text) in a vectorized form, making it easier to compare and contrast embeddings. For example, two texts that share a similar subject matter or sentiment should have similar embeddings, which can be identified through mathematical comparison techniques such as cosine similarity. For more on how and why you should use embeddings, refer to the Embeddings guide.
+The Gemini Embedding Model API supports various task types:
+
+| Task Type               | Description                                                                                                                                                |
+|-------------------------|------------------------------------------------------------------------------------------------------------------------------------------------------------|
+| "`retrieval_query`"     | Specifies the given text is a query in a search/retrieval setting.                                                                                         |
+| "`retrieval_document`"  | Specifies the given text is a document in a search/retrieval setting. Using this task type requires a title but is automatically proided by Embeddings API |
+| "`semantic_similarity`" | Specifies the given text will be used for Semantic Textual Similarity (STS).                                                                               |
+| "`classification`"      | Specifies that the embeddings will be used for classification.                                                                                             |
+| "`clusering`"           | Specifies that the embeddings will be used for clustering.                                                                                                 |
+
+
+Usage Example:
+
+```python
+import lancedb
+import pandas as pd
+from lancedb.pydantic import LanceModel, Vector
+from lancedb.embeddings import get_registry
+
+
+model = get_registry().get("gemini-text").create()
+
+class TextModel(LanceModel):
+    text: str = model.SourceField()
+    vector: Vector(model.ndims()) = model.VectorField()
+
+df = pd.DataFrame({"text": ["hello world", "goodbye world"]})
+db = lancedb.connect("~/.lancedb")
+tbl = db.create_table("test", schema=TextModel, mode="overwrite")
+
+tbl.add(df)
+rs = tbl.search("hello").limit(1).to_pandas()
+```
+
 ## Multi-modal embedding functions
 Multi-modal embedding functions allow you to query your table using both images and text.