7. Receptome construction¶
Language: Python · Input: intracellular_network_{KO,WT}.csv,
intercellular_network.csv ·
Output: receptome.csv, receptome_wo_celltype.csv, MSCs_receptome.csv
The receptome links receptors to downstream target genes, restricted to receptors that actually take part in a significant ligand–receptor interaction. It is the structure the ridge models are trained on in step 9, and it is what makes the perturbation receptor-driven rather than TF-driven.
generate_receptome builds it in one call. It keeps every intracellular
receptor → target gene connection whose receptor is engaged by a ligand in the same receiving
cell type, and handles the fiddly part on its own: LR receptor names may be complexes
(Il1r1_Il1rap) while the intracellular network stores single genes, so each complex is split
into its subunits and each subunit matched separately. Node type suffixes such as |R are
stripped before matching, and the TF / TF_Score columns of the intracellular network are
dropped — the ridge model connects receptors directly to target genes.
Handoff: R → Python
The receptome is built in Python, but both of its inputs sit inside the R objects of step 4 — the intracellular networks in the TF object, the differential ligand–receptor table in the CrossTalkeR object. Export them once:
TF_results <- readRDS(file.path(TF_OUT, "TF_results", "result_TF_object.RDS"))
for (cond in c("KO", "WT")) {
net <- unique(TF_results@intracellular_network_condition[[cond]][,
c("Receptor", "celltype", "Target_Gene")])
write.csv(net, file.path(PERT, paste0("intracellular_network_", cond, ".csv")),
row.names = FALSE)
}
CTR_results <- readRDS(file.path(CTR_OUT, "LR_data_final.Rds"))
intercell_network <- CTR_results@tables$KO_x_WT
write.csv(intercell_network, file.path(PERT, "intercellular_network.csv"),
row.names = FALSE)
Subsetting to the three columns before writing matters: the full networks are several million
rows wide with TF and TF_Score, and unique() on the three columns that
generate_receptome reads shrinks each file from a few hundred MB to under 20 MB without
changing the result.
Build the receptome¶
from pathlib import Path
import pandas as pd
from intratalkerpy.perturbation.ut import generate_receptome
intracellular = {
cond: pd.read_csv(PERT / f"intracellular_network_{cond}.csv")
for cond in ("KO", "WT")
}
CTR_input = pd.read_csv(PERT / "intercellular_network.csv")
receptome = generate_receptome(
intracellular_networks = intracellular,
CTR_input = CTR_input,
celltypes = ["MSCs"],
out_path = PERT,
)
len(set(receptome["MSCs_receptome"]["Receptor"])) # candidate receptors in MSCs
intracellular_networks takes a single data frame, a list, or a condition-keyed mapping — the
networks are pooled and deduplicated, so passing both conditions at once replaces the manual
rbind of the two condition networks. If you run the TF analysis in Python rather than in R,
hand TFObj.intracellular_network_condition (or …_cluster) to it directly instead of reading
the CSVs.
Always pass celltypes together with out_path
celltypes=None means every cell type in the receptome, and with out_path set that saves
one CSV per cell type — 20 extra files for this data set. Restricting it to ["MSCs"] saves
only the MSC receptome.
Which CrossTalkeR table to pass
CTR_input is CTR_results@tables$KO_x_WT, the result table of the differential KO vs. WT
analysis, so the receptome covers the interactions of that comparison.
The function also accepts the interaction tables used as input of CrossTalkeR
(Outputs of combine_LR_and_TF / combine_LR_and_TF_complexes): it reads only type_gene_A,
type_gene_B, gene_B and target, and uses the presence of an LR pair rather than its score.
What you get¶
generate_receptome returns the tables as a dictionary and, with out_path set, writes each
one as <key>.csv:
Key / file |
Columns |
Used by |
|---|---|---|
|
|
reference table, all cell types |
|
|
ridge input pooled over cell types |
|
|
MSCs_receptome.csv is the file we use here for the perturbation simulation; its unique
receptors define the candidates that can be perturbed in step 9 — 38 in
this run, of which the publication simulates 24. The exact count follows the TF filtering
thresholds of step 4, so expect it to differ if you change
logfc or pval.