- `ValueSource::term_dictionary()` lets a source resolve its own term
ords (`TermOrdDictionary`, implemented by the sstable dictionary).
Terms and cardinality resolve ords through it; a registered name never
uses the physical dictionary of the same field.
- `ValueSource::memory_consumption()`: the growth observed across
`load_block` calls is charged to the aggregation memory limits.
- `ValueSourceProvider::for_segment` receives the column types the
aggregation accepts.
- A registered source whose type is not allowed is treated as absent
instead of silently falling back to the physical field.
- composite and multi_terms reject registered sources explicitly.
* Make ValueSource::load_block take &mut self
Value sources were held as `Arc<dyn ValueSource>` and cloned into their
collector, so `load_block` could only take `&self`. A computed source
could not keep per-segment state (caches, scratch buffers) across blocks.
Collectors were built in two phases: `build_nodes` pushed each
`XxxAggReqData` into per-kind vectors of `AggregationsSegmentCtx`, and
each collector builder cloned its request data out of those vectors by
index. The ctx copy had to stay around for name lookups, memory
accounting, and for the collectors that read their data back by index.
Each request node is turned into exactly one collector, so nothing
actually shares a value source. This change makes that ownership
explicit:
- `ValueSource::load_block` takes `&mut self`,
`ValueSourceProvider::for_segment` returns `Box<dyn ValueSource>`.
- The request tree (`AggNode { data: AggNodeData, children }`) owns the
request data. Building collectors consumes the tree, moving the
request data into the collectors instead of cloning it.
- `AggregationsSegmentCtx` only keeps the shared collect-time state
(context and block accessor). `PerRequestAggSegCtx`, `AggKind`,
`idx_in_req_data` and the `push_*`/`get_*_req_data` helpers are gone.
- Cardinality, percentiles, top_hits and missing_term collectors own
their request data instead of reading it from the ctx by index.
- Histogram requests are normalized once, when building the tree. The
flattened terms x histogram path is split into a borrow-only plan step
and a consuming build step, and no longer clones the histogram request
at finalization.
- Request data memory is charged exactly once per node. Previously it
was charged at every nesting level, and range, histogram, filter and
composite builders charged their request data a second time.
- `FilterAggReqData::evaluator` no longer needs an `Rc`, and the
filter, composite and multi_terms request data no longer derive Clone.
* CR comment
* CR comment
---------
Co-authored-by: Paul Masurel <paul.masurel@datadoghq.com>
* Changing the way aggregation access their value.
They now get values via a ValueSource abstraction.
The aggregation collector also gets the possibility to
register ValueSourceProvider describing value columns that are computed
on the fly.
Finally, segment aggregation that require a full column
now manipulates a Arc<dyn ColumnValue> directly.
* CR comments
* Clippy
* Fixing regression
---------
Co-authored-by: Paul Masurel <paul.masurel@datadoghq.com>