Rule inference from two timepoints
Fit BoBa-T rules when expression is measured at two timepoints, so each starting state is paired with its state at a later time.
Starter notebook. The cells below are a scaffold: real function calls with placeholder paths/arguments and
TODOmarkers. Fill in your own data and run top-to-bottom. Anything markedTODOis a choice you need to make for your dataset.
Setup and imports
[ ]:
import os
import os.path as op
import time
import numpy as np
import pandas as pd
import matplotlib.pyplot as plt
import seaborn as sns
import bobaT as bb
sns.set_style('whitegrid')
plt.rcParams['figure.dpi'] = 100
import warnings
warnings.filterwarnings('ignore')
Configure paths
[ ]:
# Input data
DATA_DIR = './test_data'
# Output directories
OUTPUT_DIR = './output'
VAL_DIR = './output/validation'
ATTRACTOR_DIR = './output/attractors'
PERTURB_DIR = './output/perturbations'
os.makedirs(OUTPUT_DIR, exist_ok=True)
data_path = 'data.csv' # TODO: timepoint 0 expression
data_t1_path = 'data_t1.csv' # TODO: timepoint 1 expression (same genes)
cellID_table = 'cellID-lookuptable_x1.csv'
Load network and both timepoints
[ ]:
# Load the base network as a graph-tool graph
network_path = 'tf-lit-network.csv' # TODO: your base network CSV
graph, vertex_dict = bb.load.load_network(
f'{DATA_DIR}/{network_path}',
remove_sinks=False, remove_selfloops=True, remove_sources=False,
)
v_names, nodes = bb.utils.get_nodes(vertex_dict, graph)
data = bb.load.load_data(f'{DATA_DIR}/{data_path}', nodes, norm=0.3,
delimiter=',', log1p=False, transpose=True,
sample_order=False, fillna=0)
data_t1 = bb.load.load_data(f'{DATA_DIR}/{data_t1_path}', nodes, norm=0.3,
delimiter=',', log1p=False, transpose=True,
sample_order=False, fillna=0)
clusters = bb.utils.get_clusters(data, cellID_table=f'{DATA_DIR}/{cellID_table}',
cluster_header_list=['CellStage', 'class', 'CellProtocol'])
Split into training/test, keeping timepoints paired
Pass the second timepoint via data_t1 so the split stays paired.
[ ]:
(data_train, data_test, data_t1_train, data_t1_test,
clusters_train, clusters_test) = bb.utils.split_train_test(
data, data_t1=data_t1, clusters=clusters,
save_dir=f'{OUTPUT_DIR}/data_split',
)
Binarize
[ ]:
binarized_data = bb.proc.binarize_data(data_train, phenotype_labels=clusters,
save=True, save_dir=OUTPUT_DIR,
fname='binarized_data_train')
Fit rules using the paired timepoints
TODO: two-timepoint fitting uses
get_rules_scvelo, which takes bothdataanddata_t1. Confirm this is the intended entry point for explicit timepoints (vs. velocity) in your workflow.
[ ]:
rules, regulators_dict, strengths, signed_strengths = bb.tl.get_rules_scvelo(
data=data_train, data_t1=data_t1_train, vertex_dict=vertex_dict,
plot=True, show_plot=False, save_plot=True,
save_dir=f'{OUTPUT_DIR}/rules/rule_plots', threshold=0,
)
bb.tl.save_rules(rules, regulators_dict, fname=f'{OUTPUT_DIR}/rules.txt')
Next steps
Validate the rules, then find attractors.