upxo.statops.stattests module
Created on Fri May 31 13:32:03 2024
@author: Dr. Sunil Anandatheertha
Imports
from upxo.statops.stattests import test_rand_distr_autocorr from upxo.statops.stattests import test_rand_distr_runs from upxo.statops.stattests import test_rand_distr_chisquare from upxo.statops.stattests import test_rand_distr_kolmogorovsmirnov from upxo.statops.stattests import test_rand_distr_kullbackleibler
- upxo.statops.stattests.check_coord_distr_for_randomness(coords, method='by_distance', cor=1)[source]
- Parameters:
coords
method (method to choose. We can either use distance or the count. The) – valid options are ‘by_distance’ and ‘by_count’.
cor (cut-off radius. Used only when method = 'by_count'.)
Example
from scipy.spatial.distance import pdist, squareform import seaborn as sns
coords = np.random.random((10000,2))
from scipy.spatial import cKDTree coordtree = cKDTree(coords, leafsize=16, compact_nodes=True, copy_data=False, balanced_tree=True, boxsize=None) cut_off_radius = 0.25 npoints = [coordtree.query_ball_point(coord, cut_off_radius, p=2., return_length=True) for coord in coords] sns.histplot(npoints, kde=True, color=’gray’, kde_kws={‘linecolor’: ‘black’})
- test_results = test_rand_distr_autocorr(npoints,
apply_random_shuffle=True, alpha=0.05, plot_acf=False, print_msg=False)
test_results[‘random’] As expected, the distribution of npoints is non-randpom. Now, lets check for normality. # 1. Visual Inspection
# Histogram plt.hist(npoints, bins=’auto’, density=True, alpha=0.7) plt.xlabel(‘Value’) plt.ylabel(‘Density’) plt.title(‘Histogram of Data’) plt.show()
# Q-Q Plot stats.probplot(data, dist=”norm”, plot=plt) # Compare to standard normal plt.title(‘Q-Q Plot (Normal)’) plt.show()
- tests = {‘test_rand_distr_autocorr’: False,
‘test_rand_distr_runs’: True, ‘test_rand_distr_chisquare’: True, ‘test_rand_distr_kolmogorovsmirnov’: True, ‘test_rand_distr_kullbackleibler’: True}
distances_matrix = squareform(pdist(centroids)) triu_indices = np.triu_indices_from(distances_matrix, k=1) distances = distances_matrix[triu_indices] test_results = check_coord_distr_for_randomness(coords) test_results[‘random’]
- upxo.statops.stattests.test_rand_distr_autocorr(ARRAY, alpha=0.05, apply_random_shuffle=True, _min_array_size_=10, plot_acf=False, print_msg=False)[source]
Usage
from upxo.statops.stattests import test_rand_distr_autocorr
Example
from scipy.spatial.distance import pdist, squareform centroids = np.random.random((100,2)) distances_matrix = squareform(pdist(centroids)) triu_indices = np.triu_indices_from(distances_matrix, k=1) distances = distances_matrix[triu_indices] test_results = test_rand_distr_autocorr(distances,
apply_random_shuffle=True, alpha=0.05, plot_acf=False, print_msg=False)
test_results[‘random’]
Explanations
# AUTO-CORRELATION TEST TO CHECK FOR RANDOMNESS Checks for correlation between the values at different lags.
- upxo.statops.stattests.test_rand_distr_runs(ARRAY, alpha=0.05, print_msg=False)[source]
Usage
from upxo.statops.stattests import test_rand_distr_runs
Example
from scipy.spatial.distance import pdist, squareform centroids = [(1, 2), (3, 4), (5, 6), (7, 8)] distances_matrix = squareform(pdist(centroids)) triu_indices = np.triu_indices_from(distances_matrix, k=1) distances = distances_matrix[triu_indices] random = test_rand_distr_runs(distances) random
Explanations
# RUNS TEST FOR RABNDOMNESS Checks for randomness in the sequence of values.
- upxo.statops.stattests.test_rand_distr_chisquare(ARRAY, alpha=0.05, print_msg=False)[source]
Usage
from upxo.statops.stattests import test_rand_distr_chisquare
Example
from scipy.spatial.distance import pdist, squareform centroids = [(1, 2), (3, 4), (5, 6), (7, 8)] distances_matrix = squareform(pdist(centroids)) triu_indices = np.triu_indices_from(distances_matrix, k=1) distances = distances_matrix[triu_indices] random = test_rand_distr_chisquare(distances) random
Explanations
# CHI-SQUARE TEST FOR RANDOMNESS Checks if the data follows a uniform distribution.
- upxo.statops.stattests.test_rand_distr_kolmogorovsmirnov(ARRAY, alpha=0.05, print_msg=False)[source]
Usage
from upxo.statops.stattests import test_rand_distr_kolmogorovsmirnov
Example
from scipy.spatial.distance import pdist, squareform centroids = [(1, 2), (3, 4), (5, 6), (7, 8)] distances_matrix = squareform(pdist(centroids)) triu_indices = np.triu_indices_from(distances_matrix, k=1) distances = distances_matrix[triu_indices] random = test_rand_distr_kolmogorovsmirnov(distances) random
Explanations
# Kolmogorov-Smirnov Test Compares the empirical distribution of your data to a reference distribution, often used for hypothesis testing.
The K-S test compares the empirical distribution function of your data with a reference probability distribution (e.g., uniform distribution). It is useful for testing if a sample comes from a specific distribution.
- upxo.statops.stattests.test_rand_distr_kullbackleibler(ARRAY, bin_method='auto', alpha=0.5, print_msg=False)[source]
Options for bin_method
‘auto’: maximum of the ‘sturges’ and ‘fd’ estimators
- ‘fd’: Freedman Diaconis Estimator. Can be too conservative for small
datasets, but is quite good for large datasets.
- ‘scott’: Can be too conservative for small datasets, but is quite good
for large datasets. The standard deviation is not very robust to outliers. Values are very similar to the Freedman-Diaconis estimator in the absence of outliers.
- ‘rice’: It tends to overestimate the number of bins and it does not take
into account data variability
- ‘sturges’: This estimator assumes normality of data and is too
conservative for larger, non-normal datasets.
- ‘doane’: An improved version of Sturges’ formula that produces better
estimates for non-normal datasets. This estimator attempts to account for the skew of the data.
- ‘sqrt’: The simplest and fastest estimator. Only takes into account the
data size.
- NOTE: The above explanations for bin_method options are taken frm the
below reference.
https://numpy.org/doc/stable/reference/generated/numpy.histogram_bin_edges
Usage
from upxo.statops.stattests import test_rand_distr_kullbackleibler
Example
from scipy.spatial.distance import pdist, squareform centroids = [(1, 2), (3, 4), (5, 6), (7, 8)] distances_matrix = squareform(pdist(centroids)) triu_indices = np.triu_indices_from(distances_matrix, k=1) distances = distances_matrix[triu_indices] random = test_rand_distr_kullbackleibler(distances, bin_method=’auto’) random
Explanations
# Kullback-Leibler Divergence Measures how one probability distribution diverges from a reference distribution, providing a sense of how “random” your data is compared to a uniform distribution.
The KL divergence measures how one probability distribution diverges from a second, expected probability distribution. It is useful to compare the observed distribution with an expected random distribution.