Commit Graph
3 Commits
Author SHA1 Message Date
Paul Masurel 090f12157c Prepare ValueSource for computed text sources
- `ValueSource::term_dictionary()` lets a source resolve its own term
  ords (`TermOrdDictionary`, implemented by the sstable dictionary).
  Terms and cardinality resolve ords through it; a registered name never
  uses the physical dictionary of the same field.
- `ValueSource::memory_consumption()`: the growth observed across
  `load_block` calls is charged to the aggregation memory limits.
- `ValueSourceProvider::for_segment` receives the column types the
  aggregation accepts.
- A registered source whose type is not allowed is treated as absent
  instead of silently falling back to the physical field.
- composite and multi_terms reject registered sources explicitly.
2026-10-05 15:15:19 +02:00
Paul MasurelandPaul Masurel 45bbb16542 Refactoring to make ValueSource::load_block take &mut self (#3147)
* Make ValueSource::load_block take &mut self

Value sources were held as `Arc<dyn ValueSource>` and cloned into their
collector, so `load_block` could only take `&self`. A computed source
could not keep per-segment state (caches, scratch buffers) across blocks.

Collectors were built in two phases: `build_nodes` pushed each
`XxxAggReqData` into per-kind vectors of `AggregationsSegmentCtx`, and
each collector builder cloned its request data out of those vectors by
index. The ctx copy had to stay around for name lookups, memory
accounting, and for the collectors that read their data back by index.

Each request node is turned into exactly one collector, so nothing
actually shares a value source. This change makes that ownership
explicit:

- `ValueSource::load_block` takes `&mut self`,
  `ValueSourceProvider::for_segment` returns `Box<dyn ValueSource>`.
- The request tree (`AggNode { data: AggNodeData, children }`) owns the
  request data. Building collectors consumes the tree, moving the
  request data into the collectors instead of cloning it.
- `AggregationsSegmentCtx` only keeps the shared collect-time state
  (context and block accessor). `PerRequestAggSegCtx`, `AggKind`,
  `idx_in_req_data` and the `push_*`/`get_*_req_data` helpers are gone.
- Cardinality, percentiles, top_hits and missing_term collectors own
  their request data instead of reading it from the ctx by index.
- Histogram requests are normalized once, when building the tree. The
  flattened terms x histogram path is split into a borrow-only plan step
  and a consuming build step, and no longer clones the histogram request
  at finalization.
- Request data memory is charged exactly once per node. Previously it
  was charged at every nesting level, and range, histogram, filter and
  composite builders charged their request data a second time.
- `FilterAggReqData::evaluator` no longer needs an `Rc`, and the
  filter, composite and multi_terms request data no longer derive Clone.

* CR comment

* CR comment

---------

Co-authored-by: Paul Masurel <paul.masurel@datadoghq.com>
2026-10-05 15:13:44 +02:00
Paul MasurelandPaul Masurel 1f9e49da6b Abstracting columnar from aggregation. (#3112)
* Changing the way aggregation access their value.

They now get values via a ValueSource abstraction.
The aggregation collector also gets the possibility to
register ValueSourceProvider describing value columns that are computed
on the fly.

Finally, segment aggregation that require a full column
now manipulates a Arc<dyn ColumnValue> directly.

* CR comments

* Clippy

* Fixing regression

---------

Co-authored-by: Paul Masurel <paul.masurel@datadoghq.com>
2026-09-30 19:12:20 +02:00