Struct DfRecordBatch

pub struct DfRecordBatch {
    schema: Arc<Schema>,
    columns: Vec<Arc<dyn Array>>,
    row_count: usize,
}

Expand description

A two-dimensional batch of column-oriented data with a defined schema.

A RecordBatch is a two-dimensional dataset of a number of contiguous arrays, each the same length. A record batch has a schema which must match its arrays’ datatypes.

Record batches are a convenient unit of work for various serialization and computation functions, possibly incremental.

Fields§

§schema: Arc<Schema>§columns: Vec<Arc<dyn Array>>§row_count: usize

Implementations§

§

impl RecordBatch

pub fn try_new( schema: Arc<Schema>, columns: Vec<Arc<dyn Array>>, ) -> Result<RecordBatch, ArrowError>

Creates a RecordBatch from a schema and columns.

Expects the following:

the vec of columns to not be empty
the schema and column data types to have equal lengths and match
each array in columns to have the same length

If the conditions are not met, an error is returned.

§Example


let id_array = Int32Array::from(vec![1, 2, 3, 4, 5]);
let schema = Schema::new(vec![
    Field::new("id", DataType::Int32, false)
]);

let batch = RecordBatch::try_new(
    Arc::new(schema),
    vec![Arc::new(id_array)]
).unwrap();

pub fn try_new_with_options( schema: Arc<Schema>, columns: Vec<Arc<dyn Array>>, options: &RecordBatchOptions, ) -> Result<RecordBatch, ArrowError>

Creates a RecordBatch from a schema and columns, with additional options, such as whether to strictly validate field names.

See RecordBatch::try_new for the expected conditions.

pub fn new_empty(schema: Arc<Schema>) -> RecordBatch

Creates a new empty RecordBatch.

pub fn with_schema(self, schema: Arc<Schema>) -> Result<RecordBatch, ArrowError>

Override the schema of this RecordBatch

Returns an error if schema is not a superset of the current schema as determined by [Schema::contains]

pub fn schema(&self) -> Arc<Schema>

Returns the [Schema] of the record batch.

pub fn schema_ref(&self) -> &Arc<Schema>

Returns a reference to the [Schema] of the record batch.

pub fn project(&self, indices: &[usize]) -> Result<RecordBatch, ArrowError>

Projects the schema onto the specified columns

pub fn num_columns(&self) -> usize

Returns the number of columns in the record batch.

§Example


let id_array = Int32Array::from(vec![1, 2, 3, 4, 5]);
let schema = Schema::new(vec![
    Field::new("id", DataType::Int32, false)
]);

let batch = RecordBatch::try_new(Arc::new(schema), vec![Arc::new(id_array)]).unwrap();

assert_eq!(batch.num_columns(), 1);

pub fn num_rows(&self) -> usize

Returns the number of rows in each column.

§Example


let id_array = Int32Array::from(vec![1, 2, 3, 4, 5]);
let schema = Schema::new(vec![
    Field::new("id", DataType::Int32, false)
]);

let batch = RecordBatch::try_new(Arc::new(schema), vec![Arc::new(id_array)]).unwrap();

assert_eq!(batch.num_rows(), 5);

pub fn column(&self, index: usize) -> &Arc<dyn Array>

Get a reference to a column’s array by index.

§Panics

Panics if index is outside of 0..num_columns.

pub fn column_by_name(&self, name: &str) -> Option<&Arc<dyn Array>>

Get a reference to a column’s array by name.

pub fn columns(&self) -> &[Arc<dyn Array>]

Get a reference to all columns in the record batch.

pub fn remove_column(&mut self, index: usize) -> Arc<dyn Array>

Remove column by index and return it.

Return the ArrayRef if the column is removed.

§Panics

Panics if `index`` out of bounds.

§Example

use std::sync::Arc;
use arrow_array::{BooleanArray, Int32Array, RecordBatch};
use arrow_schema::{DataType, Field, Schema};
let id_array = Int32Array::from(vec![1, 2, 3, 4, 5]);
let bool_array = BooleanArray::from(vec![true, false, false, true, true]);
let schema = Schema::new(vec![
    Field::new("id", DataType::Int32, false),
    Field::new("bool", DataType::Boolean, false),
]);

let mut batch = RecordBatch::try_new(Arc::new(schema), vec![Arc::new(id_array), Arc::new(bool_array)]).unwrap();

let removed_column = batch.remove_column(0);
assert_eq!(removed_column.as_any().downcast_ref::<Int32Array>().unwrap(), &Int32Array::from(vec![1, 2, 3, 4, 5]));
assert_eq!(batch.num_columns(), 1);

pub fn slice(&self, offset: usize, length: usize) -> RecordBatch

Return a new RecordBatch where each column is sliced according to offset and length

§Panics

Panics if offset with length is greater than column length.

pub fn try_from_iter<I, F>(value: I) -> Result<RecordBatch, ArrowError>
where I: IntoIterator<Item = (F, Arc<dyn Array>)>, F: AsRef<str>,

Create a RecordBatch from an iterable list of pairs of the form (field_name, array), with the same requirements on fields and arrays as RecordBatch::try_new. This method is often used to create a single RecordBatch from arrays, e.g. for testing.

The resulting schema is marked as nullable for each column if the array for that column is has any nulls. To explicitly specify nullibility, use RecordBatch::try_from_iter_with_nullable

Example:


let a: ArrayRef = Arc::new(Int32Array::from(vec![1, 2]));
let b: ArrayRef = Arc::new(StringArray::from(vec!["a", "b"]));

let record_batch = RecordBatch::try_from_iter(vec![
  ("a", a),
  ("b", b),
]);

pub fn try_from_iter_with_nullable<I, F>( value: I, ) -> Result<RecordBatch, ArrowError>
where I: IntoIterator<Item = (F, Arc<dyn Array>, bool)>, F: AsRef<str>,

Create a RecordBatch from an iterable list of tuples of the form (field_name, array, nullable), with the same requirements on fields and arrays as RecordBatch::try_new. This method is often used to create a single RecordBatch from arrays, e.g. for testing.

Example:


let a: ArrayRef = Arc::new(Int32Array::from(vec![1, 2]));
let b: ArrayRef = Arc::new(StringArray::from(vec![Some("a"), Some("b")]));

// Note neither `a` nor `b` has any actual nulls, but we mark
// b an nullable
let record_batch = RecordBatch::try_from_iter_with_nullable(vec![
  ("a", a, false),
  ("b", b, true),
]);

pub fn get_array_memory_size(&self) -> usize

Returns the total number of bytes of memory occupied physically by this batch.

Trait Implementations§

§

impl Clone for RecordBatch

§

fn clone(&self) -> RecordBatch

Returns a copy of the value. Read more

1.0.0 · source§

fn clone_from(&mut self, source: &Self)

Performs copy-assignment from source. Read more

§

impl Debug for RecordBatch

§

fn fmt(&self, f: &mut Formatter<'_>) -> Result<(), Error>

Formats the value using the given formatter. Read more

§

impl From<&StructArray> for RecordBatch

§

fn from(struct_array: &StructArray) -> RecordBatch

Converts to this type from the input type.

§

impl From<StructArray> for RecordBatch

§

fn from(value: StructArray) -> RecordBatch

Converts to this type from the input type.

§

impl Index<&str> for RecordBatch

§

fn index(&self, name: &str) -> &<RecordBatch as Index<&str>>::Output

Get a reference to a column’s array by name.

§Panics

Panics if the name is not in the schema.

§

type Output = Arc<dyn Array>

The returned type after indexing.

§

impl PartialEq for RecordBatch

§

fn eq(&self, other: &RecordBatch) -> bool

Tests for self and other values to be equal, and is used by ==.

1.0.0 · source§

fn ne(&self, other: &Rhs) -> bool

Tests for !=. The default implementation is almost always sufficient, and should not be overridden without very good reason.

§

impl RecordOutput for &RecordBatch

§

fn record_output(self, bm: &BaselineMetrics) -> &RecordBatch

Record that some number of output rows have been produced Read more

§

impl RecordOutput for RecordBatch

§

fn record_output(self, bm: &BaselineMetrics) -> RecordBatch

Record that some number of output rows have been produced Read more

§

impl StructuralPartialEq for RecordBatch

Auto Trait Implementations§

§

Struct DfRecordBatchCopy item path

Fields§

Implementations§

impl RecordBatch

pub fn try_new( schema: Arc<Schema>, columns: Vec<Arc<dyn Array>>, ) -> Result<RecordBatch, ArrowError>

§Example

pub fn try_new_with_options( schema: Arc<Schema>, columns: Vec<Arc<dyn Array>>, options: &RecordBatchOptions, ) -> Result<RecordBatch, ArrowError>

pub fn new_empty(schema: Arc<Schema>) -> RecordBatch

pub fn with_schema(self, schema: Arc<Schema>) -> Result<RecordBatch, ArrowError>

pub fn schema(&self) -> Arc<Schema>

pub fn schema_ref(&self) -> &Arc<Schema>

pub fn project(&self, indices: &[usize]) -> Result<RecordBatch, ArrowError>

pub fn num_columns(&self) -> usize

§Example

pub fn num_rows(&self) -> usize

§Example

pub fn column(&self, index: usize) -> &Arc<dyn Array>

§Panics

pub fn column_by_name(&self, name: &str) -> Option<&Arc<dyn Array>>

pub fn columns(&self) -> &[Arc<dyn Array>]

pub fn remove_column(&mut self, index: usize) -> Arc<dyn Array>

§Panics

§Example

pub fn slice(&self, offset: usize, length: usize) -> RecordBatch

§Panics

pub fn try_from_iter<I, F>(value: I) -> Result<RecordBatch, ArrowError>where I: IntoIterator<Item = (F, Arc<dyn Array>)>, F: AsRef<str>,

pub fn try_from_iter_with_nullable<I, F>( value: I, ) -> Result<RecordBatch, ArrowError>where I: IntoIterator<Item = (F, Arc<dyn Array>, bool)>, F: AsRef<str>,

pub fn get_array_memory_size(&self) -> usize

Trait Implementations§

impl Clone for RecordBatch

fn clone(&self) -> RecordBatch

fn clone_from(&mut self, source: &Self)

impl Debug for RecordBatch

fn fmt(&self, f: &mut Formatter<'_>) -> Result<(), Error>

impl From<&StructArray> for RecordBatch

fn from(struct_array: &StructArray) -> RecordBatch

impl From<StructArray> for RecordBatch

fn from(value: StructArray) -> RecordBatch

impl Index<&str> for RecordBatch

fn index(&self, name: &str) -> &<RecordBatch as Index<&str>>::Output

§Panics

type Output = Arc<dyn Array>

impl PartialEq for RecordBatch

fn eq(&self, other: &RecordBatch) -> bool

fn ne(&self, other: &Rhs) -> bool

impl RecordOutput for &RecordBatch

fn record_output(self, bm: &BaselineMetrics) -> &RecordBatch

impl RecordOutput for RecordBatch

fn record_output(self, bm: &BaselineMetrics) -> RecordBatch

impl StructuralPartialEq for RecordBatch

Auto Trait Implementations§

impl Freeze for RecordBatch

impl !RefUnwindSafe for RecordBatch

impl Send for RecordBatch

impl Sync for RecordBatch

impl Unpin for RecordBatch

impl !UnwindSafe for RecordBatch

Blanket Implementations§

impl<T> Any for Twhere T: 'static + ?Sized,

fn type_id(&self) -> TypeId

impl<T> Borrow<T> for Twhere T: ?Sized,

fn borrow(&self) -> &T

impl<T> BorrowMut<T> for Twhere T: ?Sized,

fn borrow_mut(&mut self) -> &mut T

impl<T> CloneToUninit for Twhere T: Clone,

unsafe fn clone_to_uninit(&self, dst: *mut T)

impl<T> Conv for T

fn conv<T>(self) -> Twhere Self: Into<T>,

impl<T, V> Convert<T> for Vwhere V: Into<T>,

fn convert(value: Self) -> T

fn convert_box(value: Box<Self>) -> Box<T>

fn convert_vec(value: Vec<Self>) -> Vec<T>

fn convert_vec_box(value: Vec<Box<Self>>) -> Vec<Box<T>>

fn convert_matrix(value: Vec<Vec<Self>>) -> Vec<Vec<T>>

fn convert_option(value: Option<Self>) -> Option<T>

fn convert_option_box(value: Option<Box<Self>>) -> Option<Box<T>>

fn convert_option_vec(value: Option<Vec<Self>>) -> Option<Vec<T>>

impl<T> FmtForward for T

fn fmt_binary(self) -> FmtBinary<Self>where Self: Binary,

fn fmt_display(self) -> FmtDisplay<Self>where Self: Display,

Struct DfRecordBatch

pub fn try_from_iter<I, F>(value: I) -> Result<RecordBatch, ArrowError>
where I: IntoIterator<Item = (F, Arc<dyn Array>)>, F: AsRef<str>,

pub fn try_from_iter_with_nullable<I, F>( value: I, ) -> Result<RecordBatch, ArrowError>
where I: IntoIterator<Item = (F, Arc<dyn Array>, bool)>, F: AsRef<str>,

impl<T> Any for T
where T: 'static + ?Sized,

impl<T> Borrow<T> for T
where T: ?Sized,

impl<T> BorrowMut<T> for T
where T: ?Sized,

impl<T> CloneToUninit for T
where T: Clone,

fn conv<T>(self) -> T
where Self: Into<T>,

impl<T, V> Convert<T> for V
where V: Into<T>,

fn fmt_binary(self) -> FmtBinary<Self>
where Self: Binary,

fn fmt_display(self) -> FmtDisplay<Self>
where Self: Display,

fn fmt_lower_exp(self) -> FmtLowerExp<Self>
where Self: LowerExp,

fn fmt_lower_hex(self) -> FmtLowerHex<Self>
where Self: LowerHex,

fn fmt_octal(self) -> FmtOctal<Self>
where Self: Octal,

fn fmt_pointer(self) -> FmtPointer<Self>
where Self: Pointer,

fn fmt_upper_exp(self) -> FmtUpperExp<Self>
where Self: UpperExp,

fn fmt_upper_hex(self) -> FmtUpperHex<Self>
where Self: UpperHex,

fn fmt_list(self) -> FmtList<Self>
where &'a Self: for<'a> IntoIterator,

impl<T> FromRef<T> for T
where T: Clone,

impl<T, U> Into<U> for T
where U: From<T>,

fn into_either_with<F>(self, into_left: F) -> Either<Self, Self>
where F: FnOnce(&Self) -> bool,

impl<T> Pipe for T
where T: ?Sized,

fn pipe<R>(self, func: impl FnOnce(Self) -> R) -> R
where Self: Sized,

fn pipe_ref<'a, R>(&'a self, func: impl FnOnce(&'a Self) -> R) -> R
where R: 'a,

fn pipe_ref_mut<'a, R>(&'a mut self, func: impl FnOnce(&'a mut Self) -> R) -> R
where R: 'a,

fn pipe_borrow<'a, B, R>(&'a self, func: impl FnOnce(&'a B) -> R) -> R
where Self: Borrow<B>, B: 'a + ?Sized, R: 'a,

fn pipe_borrow_mut<'a, B, R>( &'a mut self, func: impl FnOnce(&'a mut B) -> R, ) -> R
where Self: BorrowMut<B>, B: 'a + ?Sized, R: 'a,

fn pipe_as_ref<'a, U, R>(&'a self, func: impl FnOnce(&'a U) -> R) -> R
where Self: AsRef<U>, U: 'a + ?Sized, R: 'a,

fn pipe_as_mut<'a, U, R>(&'a mut self, func: impl FnOnce(&'a mut U) -> R) -> R
where Self: AsMut<U>, U: 'a + ?Sized, R: 'a,

fn pipe_deref<'a, T, R>(&'a self, func: impl FnOnce(&'a T) -> R) -> R
where Self: Deref<Target = T>, T: 'a + ?Sized, R: 'a,

fn pipe_deref_mut<'a, T, R>( &'a mut self, func: impl FnOnce(&'a mut T) -> R, ) -> R
where Self: DerefMut<Target = T> + Deref, T: 'a + ?Sized, R: 'a,

fn tap_borrow<B>(self, func: impl FnOnce(&B)) -> Self
where Self: Borrow<B>, B: ?Sized,

fn tap_borrow_mut<B>(self, func: impl FnOnce(&mut B)) -> Self
where Self: BorrowMut<B>, B: ?Sized,

fn tap_ref<R>(self, func: impl FnOnce(&R)) -> Self
where Self: AsRef<R>, R: ?Sized,

fn tap_ref_mut<R>(self, func: impl FnOnce(&mut R)) -> Self
where Self: AsMut<R>, R: ?Sized,

fn tap_deref<T>(self, func: impl FnOnce(&T)) -> Self
where Self: Deref<Target = T>, T: ?Sized,

fn tap_deref_mut<T>(self, func: impl FnOnce(&mut T)) -> Self
where Self: DerefMut<Target = T> + Deref, T: ?Sized,

fn tap_borrow_dbg<B>(self, func: impl FnOnce(&B)) -> Self
where Self: Borrow<B>, B: ?Sized,

fn tap_borrow_mut_dbg<B>(self, func: impl FnOnce(&mut B)) -> Self
where Self: BorrowMut<B>, B: ?Sized,

fn tap_ref_dbg<R>(self, func: impl FnOnce(&R)) -> Self
where Self: AsRef<R>, R: ?Sized,

fn tap_ref_mut_dbg<R>(self, func: impl FnOnce(&mut R)) -> Self
where Self: AsMut<R>, R: ?Sized,

fn tap_deref_dbg<T>(self, func: impl FnOnce(&T)) -> Self
where Self: Deref<Target = T>, T: ?Sized,

fn tap_deref_mut_dbg<T>(self, func: impl FnOnce(&mut T)) -> Self
where Self: DerefMut<Target = T> + Deref, T: ?Sized,

impl<T> ToOwned for T
where T: Clone,

fn try_conv<T>(self) -> Result<T, Self::Error>
where Self: TryInto<T>,

impl<T, U> TryFrom<U> for T
where U: Into<T>,