Skip to main content

prefer_label_runs

Function prefer_label_runs 

Source
fn prefer_label_runs(columns: &[&ArrayRef], rows: usize) -> bool
Expand description

Decides whether to group the rows of a batch into runs that share their label values.

Sampling adjacent row pairs keeps the decision independent of the batch size. Probe positions come from a xorshift sequence instead of a fixed stride, which would alias with periodic series layouts. A wrong guess only costs time: partition still validates every boundary, and the row-by-row path builds the labels of every row.