Files
lancedb/python/benchmarks
Taylor a6dcbba49e feat(python): add opt-in request batching to TypeSafeReranker (#4316)
Adds opt-in `batch_size=40`: 80 non-null candidates use **2 requests
instead of 80**. Default `batch_size=1` preserves the existing payload.
Each batched question sees only its document and the shared query;
instructions/criteria stay unchanged. Prompts referencing
`state.document` need adaptation. Concurrency still limits requests;
retries remain SDK-managed. Batch-local IDs and ordered results preserve
duplicate documents; invalid IDs/probabilities raise without fallback.

Validation: **124 automated tests passed**, plus live search
integrations in both modes. A **100-query live subset test of the
initial implementation** (6,851 candidate pairs; `jev-1.13.0`, SDK
0.7.1, concurrency 32) passed: unbatched/batched calls **6,851/199**,
median **877/351 ms**, p95 **9,446/598 ms**. Hybrid Hit@5/Hit@10 was
**83%/89% vs 84%/90%**; vector/FTS results and exact configuration are
in the [test
report](https://github.com/lancedb/lancedb/blob/typesafe-request-batching/python/benchmarks/typesafe_batching.md).
Ruff and MkDocs passed. The simplified batching flow also passed a fresh
four-query live check: 292 candidate pairs, with 292 unbatched versus 8
batched calls.

Builds on #4209; motivated by [research
#7](https://github.com/lancedb/research/pull/7).
2026-09-24 13:53:52 +05:30
..