utils — utilities

General helper utilities.

bobaT.utils.make_color_map(attractors, palette='hls', set_colors=None)[source]

Function to make a matplotlib color map from a list of strings and the matplotlib default color palette :param list_of_strings: List of strings :type list_of_strings: list :param palette: Matplotlib default color palette :type palette: string :param set_colors: Dictionary of colors for each attractor if user wants to specify particular color :type set_colors: dictionary

Returns:

cmap – Dictionary of colors for each string

Return type:

dictionary

bobaT.utils.binarized_umap_transform(binarized_data)[source]

Function to perform dimensionality reduction on binarized data using UMAP :param binarized_data: :return:

bobaT.utils.binarized_data_dict_to_binary_df(binarized_data, nodes)[source]

Function to convert a dictionary of binarized data to a binary dataframe :param binarized_data: Dictionary of binarized data :type binarized_data: dictionary :param nodes: List of nodes in the transcription factor network :type nodes: list

Returns:

binary_df – Binary dataframe of binarized data

Return type:

dataframe

bobaT.utils.binarize_data_df(data, nodes, threshold=0.5)[source]

This function generates a binary dataframe directly from a dataframe of continuous data, with the same index and column names

Parameters:
  • data (Pandas DataFrame) – DataFrame of data, genes by samples

  • nodes (list) – list of node names

  • threshold (float, optional) – Value between 0 and 1 to threshold the binarization, defaults to 0.5

Returns:

DataFrame of binary data, genes by samples

Return type:

Pandas DataFrame

bobaT.utils.idx2binary(idx, n)[source]

Convert index (int, base 10) to a binary str

Parameters:
  • idx (int) – index value of vertex (base 10)

  • n (int) – base for binary conversion

Returns:

binary version of index

Return type:

str

bobaT.utils.state2idx(state)[source]

Convert binary str to index (int, base 10)

Parameters:

state (str) – binary version of index

Returns:

index value of vertex (base 10)

Return type:

int

bobaT.utils.state_bool2idx(state)[source]

Convert list of boolean values (T/F) to index (int, base 10)

Parameters:

state (list of bool values) – Boolean version of state

Returns:

index value of vertex (base 10)

Return type:

int

bobaT.utils.hamming(x, y)[source]

Hamming distance between 2 states

Parameters:
  • x (int) – State 1 in binary

  • y (int) – State 2 in binary

Returns:

Hamming distance

Return type:

int

bobaT.utils.hamming_idx(x, y, n)[source]

Hamming distance between 2 states, where binary states are given by decimal code

Parameters:
  • x (str) – State 1 index

  • y (str) – State 2 index

  • n (int) – Number of nodes in network

Returns:

Hamming distance

Return type:

int

bobaT.utils.r2(x, y)[source]
bobaT.utils.split_train_test(data, data_t1, clusters, save_dir, suffix='', random_state=1234)[source]

Split a dataset into testing and training dataset

Parameters:
  • data (Pandas dataframe) – Dataset or first timepoint of temporal dataset to be split into training/testing datasets

  • data_t1 ({Pandas dataframe, None}) – Second timepoint of temporal dataset, optional

  • clusters (Pandas DataFrame) – Cluster assignments for each sample; see ut.get_clusters() to generate

  • save_dir (str) – File path for saving training and testing sets

  • suffix (str, optional) – Suffix to add to file names for saving, defaults to None

Returns:

List of dataframes split into training and testing: data (training set, t0), test (testing set, t1), data_t1 (training set, t1), test_t (testing set, t1), clusters_train (cluster IDs of training set), clusters_test (cluster IDs of testing set)

Return type:

Pandas dataframes

bobaT.utils.split_train_test_crossval(data, data_t1, clusters, save_dir, folds=5, fname=None, random_state=1234)[source]

Split a dataset into testing and training dataset with multiple folds (e.g. for 5-fold validation)

Parameters:
  • data (Pandas dataframe) – Dataset or first timepoint of temporal dataset to be split into training/testing datasets

  • data_t1 ({Pandas dataframe, None}) – Second timepoint of temporal dataset, optional

  • clusters (Pandas DataFrame) – Cluster assignments for each sample; see ut.get_clusters() to generate

  • save_dir (str) – File path for saving training and testing sets

  • fname (str, optional) – Suffix to add to file names for saving, defaults to None

Returns:

None

bobaT.utils.condense(G, directed=True, attractors=True)[source]
bobaT.utils.average_state(idx_list, n)[source]
bobaT.utils.inspect_state(i, stg, vidx, rules, regulators_dict, nodes, n)[source]
bobaT.utils.update_node(rules, regulators_dict, node, node_i, nodes, node_indices, state_bool, return_state=False)[source]

_summary_

Parameters:
  • rules (_type_) – _description_

  • regulators_dict (_type_) – _description_

  • node (_type_) – _description_

  • node_i (_type_) – _description_

  • nodes (_type_) – _description_

  • node_indices (_type_) – _description_

  • state_bool (_type_) – _description_

  • return_state (bool, optional) – _description_, defaults to False

Returns:

_description_

Return type:

_type_

bobaT.utils.prune_stg_edges(stg, edge_weights, n, threshold=0.5)[source]
bobaT.utils.get_nodes(vertex_dict, graph)[source]
bobaT.utils.get_clusters(data, data_test=None, is_data_split=False, cellID_table=None, cluster_header_list=None)[source]
bobaT.utils.get_leaves_of_regulator(n, index)[source]
bobaT.utils.get_avg_state_index(nodes, average_states, outfile, save_dir=None)[source]
bobaT.utils.get_reprogramming_rules(rules, regulators_dict, on_nodes, off_nodes)[source]
bobaT.utils.get_flip_probs(idx, rules, regulators_dict, nodes, node_indices=None)[source]
bobaT.utils.get_avg_min_distance(binarized_data, n, min_dist=20)[source]
bobaT.utils.get_attractors(atts, c_vertex_dict, vert_idx=None)[source]
bobaT.utils.get_partial_stg(start_states, rules, nodes, regulators_dict, radius, on_nodes=[], off_nodes=[], pthreshold=0.0)[source]

Simulates the probabilistic graph/rules and returns a state transition graph and edge_weights

Parameters:
  • start_states (list of int) – list of start state indices (int, base 10)

  • rules (dict) – dictionary of rules in the form {gene_name: [list of leaf probabilities]}

  • nodes (list of str) – list of nodes in network

  • regulators_dict (_type_) – _description_

  • radius (_type_) – _description_

  • on_nodes (list, optional) – _description_, defaults to []

  • off_nodes (list, optional) – _description_, defaults to []

  • pthreshold (_type_, optional) – _description_, defaults to 0.

Returns:

state transition graph and edge_weights

Return type:

_type_

bobaT.utils.get_ci_sig(results, group_cols=['gene'], score_col='score', mean_threshold=-0.3)[source]
bobaT.utils.get_perturbation_dict(attractor_dict, perturbations_dir, significance='both', save_full=False, save_dir='clustered_perturb_plots', mean_threshold=-0.3)[source]
bobaT.utils.reverse_dictionary(dictionary)[source]
bobaT.utils.reverse_perturb_dictionary(dictionary)[source]
bobaT.utils.write_dict_of_dicts(dictionary, file)[source]
bobaT.utils.get_attractor_dict(ATTRACTOR_DIR, filtered=True)[source]
bobaT.utils.print_graph_info(graph, vertex_dict, nodes, suffix='', dir_prefix='', plot=True, fillcolor='lightcyan', gene2color=None, layout=None, add_edge_weights=False, ew_df=None)[source]