The Nautilus.Stats module provides descriptive statistics over
tensor[n, f32] vectors. All functions are polymorphic over the tensor
length n. The reductions return f32 scalars; rank_vec and
zscore_vec return a tensor of the same length, and the covariance and
correlation matrices return a square tensor over the variables.
(x - mean) / std; NaN throughout for a constant vector
For finite input rank_vec matches scipy.stats.rankdata's default
method="average", so rank_vec([3, 1, 4, 1]) is [3, 1.5, 4, 1.5] and the
ranks sum to n(n+1)/2 whether or not there are ties.
rank_vec counts each element against every other, so it is O(n^2) and is
markedly slower than the sort-based reductions in this module; prefer
quantile_vec or median_vec when a rank vector is not what you need.
[m, n](x: &tensor[m, n, f32], ddof: i64) -> tensor[m, m, f32]
m variables by n observations
correlation_matrix
[m, n](x: &tensor[m, n, f32]) -> tensor[m, m, f32]
Pearson correlation over the same layout
Each row is a variable and each column an observation, matching
numpy.cov's default rowvar=True. covariance_matrix over two stacked
rows therefore agrees entrywise with covariance_matrix_2 on the same two
vectors. correlation_matrix is
scale-invariant, so it takes no ddof.
Every function borrows its input (&tensor), so the same tensor can be
passed to several statistics without copy.
skewness_vec and kurtosis_vec use population (biased) moments,
matching the default behavior of scipy.stats.skew/kurtosis with
bias=True.
quantile_vec clamps q to [0, 1] and uses linear interpolation
between adjacent sorted values.
trimmed_mean_vec truncates (floor) the trim count, so small
proportions on short vectors may trim nothing.
zscore_vec and correlation_matrix divide by a standard deviation.
A constant vector (or a constant row) has zero spread, so the result is
NaN rather than 0, as numpy.corrcoef also does.
rank_vec does not propagate NaN.lt and eq are both false
against NaN, so a NaN element ranks as 0.5 and is counted by no other
element: the result is finite, wrong, and gives the caller no signal.
scipy.stats.rankdata returns all-NaN instead. Screen NaN before
ranking. zscore_vec, covariance_matrix and correlation_matrix all
propagate NaN correctly. Infinities rank correctly.
The whole module is f32; its functions do not accept f64 inputs.