fix: prevent duplicate data in FTS index (#728)

This forces the user to replace the whole FTS directory when re-creating the index, prevent duplicate data from being created. Previously, the whole dataset was re-added to the existing index, duplicating existing rows in the index. This (in combination with lancedb/lance#1707) caused #726, since the duplicate data emitted duplicate indices for `take()` and an upstream issue caused those queries to fail. This solution isn't ideal, since it makes the FTS index temporarily unavailable while the index is built. In the future, we should have multiple FTS index directories, which would allow atomic commits of new indexes (as well as multiple indexes for different columns). Fixes #498. Fixes #726. --------- Co-authored-by: Chang She <759245+changhiskhan@users.noreply.github.com>
2026-07-04 03:20:40 +00:00 · 2023-12-20 13:07:07 -08:00
parent 48a12e780c
commit a975cc0a94
2 changed files with 36 additions and 1 deletions
--- a/python/tests/test_fts.py
+++ b/python/tests/test_fts.py
@@ -83,6 +83,24 @@ def test_create_index_from_table(tmp_path, table):
    assert len(df) == 10
    assert "text" in df.columns

+    # Check whether it can be updated
+    table.add(
+        [
+            {
+                "vector": np.random.randn(128),
+                "text": "gorilla",
+                "text2": "gorilla",
+                "nested": {"text": "gorilla"},
+            }
+        ]
+    )
+
+    with pytest.raises(ValueError, match="already exists"):
+        table.create_fts_index("text")
+
+    table.create_fts_index("text", replace=True)
+    assert len(table.search("gorilla").limit(1).to_pandas()) == 1
+

 def test_create_index_multiple_columns(tmp_path, table):
    table.create_fts_index(["text", "text2"])