load — loading data
Load networks and expression data.
- bobaT.load.load_network(filename, remove_sources=False, remove_sinks=True, remove_selfloops=True, add_selfloops_to_sources=True, header=None)[source]
Load the transcription factor network from file :param filename: file path of CSV file of network (first column = input nodes, second column = output nodes) :type filename: str :param remove_sources: Remove nodes without inputs, defaults to False :type remove_sources: bool, optional :param remove_sinks: Remove nodes without outputs, defaults to True :type remove_sinks: bool, optional :param remove_selfloops: Remove self loops (e.g. A –> A), defaults to True :type remove_selfloops: bool, optional :param add_selfloops_to_sources: Add a self-loop saying that a source node regulates itself, defaults to True :type add_selfloops_to_sources: bool, optional :param header: Header in filename, defaults to None :type header: str, optional :return: Return a graph-tool Graph object and a vertex dictionary :rtype: [networkx.DiGraph, dict]
- bobaT.load.load_data(filename, nodes, log=False, log1p=False, sample_order=None, delimiter=',', norm='gmm', index_col=0, transpose=False, fillna=None)[source]
Read data from CSV
- Parameters:
filename (str) – file path of CSV file of data (rows = genes, cols = samples)
nodes (list) – List of nodes in TF network
log (bool, optional) – log-transform the data, defaults to False
log1p (bool, optional) – log1p-transform the data, defaults to False
sample_order (list, None, or False, optional) – If None, will generate a dendrogram to order to the samples. If False, no clustering of the samples is done. Otherwise, reorder the samples in the data to the given order, defaults to None
delimiter (str, optional) – Delimiter of data CSV, defaults to “,”
norm ({"gmm","minmax", float, None}, optional) – Normalization method to use. If float, use quantile normalization, defaults to “gmm”
index_col (int, optional) – Index column of data, defaults to 0
transpose (bool, optional) – If True, transpose the data, defaults to False
fillna (scalar, dict, Series, or DataFrame, optional) – If not None, fill NAs in data with given value, defaults to None
- Returns:
data
- Return type:
Pandas DataFrame
- bobaT.load.load_data_multiple(filenames, nodes, log=False, delimiter=',', norm='gmm')[source]
Load multiple datasets from a list of filenames
- Parameters:
filenames (list) – List of file paths to CSVs of data
nodes (list) – List of nodes in network
log (bool, optional) – If True, log-transform, defaults to False
delimiter (str, optional) – Delimiter in data CSVs, defaults to “,”
norm ({"gmm","minmax", float, None}, optional) – Normalization method to use. If float, use quantile normalization, defaults to “gmm”
- Returns:
data
- Return type:
Pandas DataFrame