utils — utilities
General helper utilities.
- bobaT.utils.make_color_map(attractors, palette='hls', set_colors=None)[source]
Function to make a matplotlib color map from a list of strings and the matplotlib default color palette :param list_of_strings: List of strings :type list_of_strings: list :param palette: Matplotlib default color palette :type palette: string :param set_colors: Dictionary of colors for each attractor if user wants to specify particular color :type set_colors: dictionary
- Returns:
cmap – Dictionary of colors for each string
- Return type:
dictionary
- bobaT.utils.binarized_umap_transform(binarized_data)[source]
Function to perform dimensionality reduction on binarized data using UMAP :param binarized_data: :return:
- bobaT.utils.binarized_data_dict_to_binary_df(binarized_data, nodes)[source]
Function to convert a dictionary of binarized data to a binary dataframe :param binarized_data: Dictionary of binarized data :type binarized_data: dictionary :param nodes: List of nodes in the transcription factor network :type nodes: list
- Returns:
binary_df – Binary dataframe of binarized data
- Return type:
dataframe
- bobaT.utils.binarize_data_df(data, nodes, threshold=0.5)[source]
This function generates a binary dataframe directly from a dataframe of continuous data, with the same index and column names
- bobaT.utils.state_bool2idx(state)[source]
Convert list of boolean values (T/F) to index (int, base 10)
- bobaT.utils.hamming_idx(x, y, n)[source]
Hamming distance between 2 states, where binary states are given by decimal code
- bobaT.utils.split_train_test(data, data_t1, clusters, save_dir, suffix='', random_state=1234)[source]
Split a dataset into testing and training dataset
- Parameters:
data (Pandas dataframe) – Dataset or first timepoint of temporal dataset to be split into training/testing datasets
data_t1 ({Pandas dataframe, None}) – Second timepoint of temporal dataset, optional
clusters (Pandas DataFrame) – Cluster assignments for each sample; see ut.get_clusters() to generate
save_dir (str) – File path for saving training and testing sets
suffix (str, optional) – Suffix to add to file names for saving, defaults to None
- Returns:
List of dataframes split into training and testing: data (training set, t0), test (testing set, t1), data_t1 (training set, t1), test_t (testing set, t1), clusters_train (cluster IDs of training set), clusters_test (cluster IDs of testing set)
- Return type:
Pandas dataframes
- bobaT.utils.split_train_test_crossval(data, data_t1, clusters, save_dir, folds=5, fname=None, random_state=1234)[source]
Split a dataset into testing and training dataset with multiple folds (e.g. for 5-fold validation)
- Parameters:
data (Pandas dataframe) – Dataset or first timepoint of temporal dataset to be split into training/testing datasets
data_t1 ({Pandas dataframe, None}) – Second timepoint of temporal dataset, optional
clusters (Pandas DataFrame) – Cluster assignments for each sample; see ut.get_clusters() to generate
save_dir (str) – File path for saving training and testing sets
fname (str, optional) – Suffix to add to file names for saving, defaults to None
- Returns:
None
- bobaT.utils.update_node(rules, regulators_dict, node, node_i, nodes, node_indices, state_bool, return_state=False)[source]
_summary_
- Parameters:
rules (_type_) – _description_
regulators_dict (_type_) – _description_
node (_type_) – _description_
node_i (_type_) – _description_
nodes (_type_) – _description_
node_indices (_type_) – _description_
state_bool (_type_) – _description_
return_state (bool, optional) – _description_, defaults to False
- Returns:
_description_
- Return type:
_type_
- bobaT.utils.get_clusters(data, data_test=None, is_data_split=False, cellID_table=None, cluster_header_list=None)[source]
- bobaT.utils.get_partial_stg(start_states, rules, nodes, regulators_dict, radius, on_nodes=[], off_nodes=[], pthreshold=0.0)[source]
Simulates the probabilistic graph/rules and returns a state transition graph and edge_weights
- Parameters:
start_states (list of int) – list of start state indices (int, base 10)
rules (dict) – dictionary of rules in the form {gene_name: [list of leaf probabilities]}
regulators_dict (_type_) – _description_
radius (_type_) – _description_
on_nodes (list, optional) – _description_, defaults to []
off_nodes (list, optional) – _description_, defaults to []
pthreshold (_type_, optional) – _description_, defaults to 0.
- Returns:
state transition graph and edge_weights
- Return type:
_type_
- bobaT.utils.get_ci_sig(results, group_cols=['gene'], score_col='score', mean_threshold=-0.3)[source]