Rule inference from two timepoints

Fit BoBa-T rules when expression is measured at two timepoints, so each starting state is paired with its state at a later time.

Starter notebook. The cells below are a scaffold: real function calls with placeholder paths/arguments and TODO markers. Fill in your own data and run top-to-bottom. Anything marked TODO is a choice you need to make for your dataset.

Setup and imports

[ ]:
import os
import os.path as op
import time
import numpy as np
import pandas as pd
import matplotlib.pyplot as plt
import seaborn as sns
import bobaT as bb

sns.set_style('whitegrid')
plt.rcParams['figure.dpi'] = 100
import warnings
warnings.filterwarnings('ignore')

Configure paths

[ ]:
# Input data
DATA_DIR = './test_data'

# Output directories
OUTPUT_DIR = './output'
VAL_DIR = './output/validation'
ATTRACTOR_DIR = './output/attractors'
PERTURB_DIR = './output/perturbations'
os.makedirs(OUTPUT_DIR, exist_ok=True)

data_path = 'data.csv'          # TODO: timepoint 0 expression
data_t1_path = 'data_t1.csv'    # TODO: timepoint 1 expression (same genes)
cellID_table = 'cellID-lookuptable_x1.csv'

Load network and both timepoints

[ ]:
# Load the base network as a graph-tool graph
network_path = 'tf-lit-network.csv'  # TODO: your base network CSV
graph, vertex_dict = bb.load.load_network(
    f'{DATA_DIR}/{network_path}',
    remove_sinks=False, remove_selfloops=True, remove_sources=False,
)
v_names, nodes = bb.utils.get_nodes(vertex_dict, graph)

data = bb.load.load_data(f'{DATA_DIR}/{data_path}', nodes, norm=0.3,
                         delimiter=',', log1p=False, transpose=True,
                         sample_order=False, fillna=0)
data_t1 = bb.load.load_data(f'{DATA_DIR}/{data_t1_path}', nodes, norm=0.3,
                            delimiter=',', log1p=False, transpose=True,
                            sample_order=False, fillna=0)

clusters = bb.utils.get_clusters(data, cellID_table=f'{DATA_DIR}/{cellID_table}',
                                 cluster_header_list=['CellStage', 'class', 'CellProtocol'])

Split into training/test, keeping timepoints paired

Pass the second timepoint via data_t1 so the split stays paired.

[ ]:
(data_train, data_test, data_t1_train, data_t1_test,
 clusters_train, clusters_test) = bb.utils.split_train_test(
    data, data_t1=data_t1, clusters=clusters,
    save_dir=f'{OUTPUT_DIR}/data_split',
)

Binarize

[ ]:
binarized_data = bb.proc.binarize_data(data_train, phenotype_labels=clusters,
                                       save=True, save_dir=OUTPUT_DIR,
                                       fname='binarized_data_train')

Fit rules using the paired timepoints

TODO: two-timepoint fitting uses get_rules_scvelo, which takes both data and data_t1. Confirm this is the intended entry point for explicit timepoints (vs. velocity) in your workflow.

[ ]:
rules, regulators_dict, strengths, signed_strengths = bb.tl.get_rules_scvelo(
    data=data_train, data_t1=data_t1_train, vertex_dict=vertex_dict,
    plot=True, show_plot=False, save_plot=True,
    save_dir=f'{OUTPUT_DIR}/rules/rule_plots', threshold=0,
)
bb.tl.save_rules(rules, regulators_dict, fname=f'{OUTPUT_DIR}/rules.txt')

Next steps

Validate the rules, then find attractors.