load — loading data

Load networks and expression data.

bobaT.load.load_network(filename, remove_sources=False, remove_sinks=True, remove_selfloops=True, add_selfloops_to_sources=True, header=None)[source]

Load the transcription factor network from file :param filename: file path of CSV file of network (first column = input nodes, second column = output nodes) :type filename: str :param remove_sources: Remove nodes without inputs, defaults to False :type remove_sources: bool, optional :param remove_sinks: Remove nodes without outputs, defaults to True :type remove_sinks: bool, optional :param remove_selfloops: Remove self loops (e.g. A –> A), defaults to True :type remove_selfloops: bool, optional :param add_selfloops_to_sources: Add a self-loop saying that a source node regulates itself, defaults to True :type add_selfloops_to_sources: bool, optional :param header: Header in filename, defaults to None :type header: str, optional :return: Return a graph-tool Graph object and a vertex dictionary :rtype: [networkx.DiGraph, dict]

bobaT.load.prune_network(G, remove_sources=True, remove_sinks=False)[source]
bobaT.load.load_data(filename, nodes, log=False, log1p=False, sample_order=None, delimiter=',', norm='gmm', index_col=0, transpose=False, fillna=None)[source]

Read data from CSV

Parameters:
  • filename (str) – file path of CSV file of data (rows = genes, cols = samples)

  • nodes (list) – List of nodes in TF network

  • log (bool, optional) – log-transform the data, defaults to False

  • log1p (bool, optional) – log1p-transform the data, defaults to False

  • sample_order (list, None, or False, optional) – If None, will generate a dendrogram to order to the samples. If False, no clustering of the samples is done. Otherwise, reorder the samples in the data to the given order, defaults to None

  • delimiter (str, optional) – Delimiter of data CSV, defaults to “,”

  • norm ({"gmm","minmax", float, None}, optional) – Normalization method to use. If float, use quantile normalization, defaults to “gmm”

  • index_col (int, optional) – Index column of data, defaults to 0

  • transpose (bool, optional) – If True, transpose the data, defaults to False

  • fillna (scalar, dict, Series, or DataFrame, optional) – If not None, fill NAs in data with given value, defaults to None

Returns:

data

Return type:

Pandas DataFrame

bobaT.load.load_data_multiple(filenames, nodes, log=False, delimiter=',', norm='gmm')[source]

Load multiple datasets from a list of filenames

Parameters:
  • filenames (list) – List of file paths to CSVs of data

  • nodes (list) – List of nodes in network

  • log (bool, optional) – If True, log-transform, defaults to False

  • delimiter (str, optional) – Delimiter in data CSVs, defaults to “,”

  • norm ({"gmm","minmax", float, None}, optional) – Normalization method to use. If float, use quantile normalization, defaults to “gmm”

Returns:

data

Return type:

Pandas DataFrame

bobaT.load.load_rules(fname='rules.txt', delimiter='|')[source]

Load rules that have previously been generated from a txt file.

Parameters:
  • fname (str, optional) – file path to text file with rules, defaults to “rules.txt”

  • delimiter (str, optional) – Delimiter of rules text file, defaults to “|”

Returns:

rules, regulators_dict

Return type:

dict(), dict()