mirror of
https://github.com/lancedb/lancedb.git
synced 2026-09-30 00:45:37 +00:00
Adds opt-in `batch_size=40`: 80 non-null candidates use **2 requests instead of 80**. Default `batch_size=1` preserves the existing payload. Each batched question sees only its document and the shared query; instructions/criteria stay unchanged. Prompts referencing `state.document` need adaptation. Concurrency still limits requests; retries remain SDK-managed. Batch-local IDs and ordered results preserve duplicate documents; invalid IDs/probabilities raise without fallback. Validation: **124 automated tests passed**, plus live search integrations in both modes. A **100-query live subset test of the initial implementation** (6,851 candidate pairs; `jev-1.13.0`, SDK 0.7.1, concurrency 32) passed: unbatched/batched calls **6,851/199**, median **877/351 ms**, p95 **9,446/598 ms**. Hybrid Hit@5/Hit@10 was **83%/89% vs 84%/90%**; vector/FTS results and exact configuration are in the [test report](https://github.com/lancedb/lancedb/blob/typesafe-request-batching/python/benchmarks/typesafe_batching.md). Ruff and MkDocs passed. The simplified batching flow also passed a fresh four-query live check: 292 candidate pairs, with 292 unbatched versus 8 batched calls. Builds on #4209; motivated by [research #7](https://github.com/lancedb/research/pull/7).