scieee AI-readable full text Open interactive document viewer

conformeR: conformalized differential expression analysis of multi-condition single-cell data

Leclerc, Justine; Seiler, Christof

Abstract

Differential expression (DE) analysis in multi-condition single-cell transcriptomics poses several challenges: extensive multiple testing across tens of thousands of genes and cells, the need for interpretable inference at the cell-type and population level, and robustness to model assumptions. conformeR tackles these issues by combining conformal inference with counterfactual prediction to generate valid p-values for DE testing. Rather than performing gene-wise analyses, conformeR exploits conditional dependencies between genes to predict counterfactual expression levels, from which conformal prediction intervals for treatment effects are constructed. These intervals are then used to derive p-values that quantify evidence against the null hypothesis of no differential expression. P-values are subsequently transformed into local false discovery rates, then into their frequentist counterparts, and aggregated across cells and patients to achieve rigorous and interpretable FDR control at the cell-type level. At its current stage, conformeR shows encouraging results on both toy and real datasets. Ongoing work focuses on improving computational efficiency, strengthening integration with existing counterfactual prediction tools, and scaling to larger datasets. This work was first presented at EuroBioC 2025 in Barcelona.

Full text

conformeR: conformalized differential expression analysis of multi-condition single-cell data Justine Leclerc PhD candidate, Center of Experimental Rheumatology, University Hospital Zurich, University of Zurich Joint work with Christof Seiler EuroBioC 2025 – September 17th–19th, Barcelona 1 / 15 Motivation Goal: Detect differential expression in multi-condition single-cell datasets. Requirements: ✓Control the False Discovery Rate (FDR) ✓Gene-level p-values per cell type, aggregated over patients ✓Preferably use distribution-free methods conformeR combines conformal inference and counterfactual prediction for robust, valid p-values. 2 / 15 ©Meet Mrs. and Mr. Smith We observe gene expression in cells from Mrs. and Mr. Smith: •Gene log-counts: YA,YBfor genes Aand B •Same cell-type, under two conditions: control (T= 0) and treatment (T= 1) Model: YA∼ N(0,1),YB=βTYA+ε, ε ∼ N(0,1) τ(YA)=(βT=1 −βT=0)·YA Research question: is gene Bdifferentially expressed under treatment? 3 / 15 Predicting the counterfactual Idea: Predict ˆ YBas if the cell were in the other condition. {Implementations available in conformeR: •Calling existing packages lemur (Ahlmann-Eltze, C., Huber, W., 2025), cfcausal (Lei, L. and Cand` es, E., 2021) •Built-in prediction model using conformal inference Prediction of the counterfactual assumes the unobserved condition can be predicted by the expression of other genes in the same sample, cell-type and condition. YB(T=t)∼YA(T=t) 4 / 15 Predicting the counterfactual Learn P(YB|YA,T) to make predictions. βT=0 =βT=1 = 1 τ(YA)=0 βT=0 = 1, βT=1 =−3 τ(YA) = −4·YA 5 / 15 Conformal testing: p-values ¬Test hypothesis H0:FYB|YA,T=0 =FYB|YA,T=1 ¬Conformal prediction intervals ˆ C∆,α over grid A ⊂ [0,1] for ∆ = (Yobs B−ˆ YB(0), T= 1 ˆ YB(1) −Yobs B,T= 0 ¬Test statistic and p-value C=X α∈A 1{0∈ˆ C∆,α},p=1 + C 1 + |A| 6 / 15 Prediction intervals for ∆ τ(YA)=4.04,p=1+5 1+99 = 0.06 τ(YA)=0,p=1+70 1+99 = 0.71 7 / 15 Interpretation of p-values conformeR p-values reflect signal presence. τ(YA)= 0 ⇒left-skewed τ(YA)=0⇒uniform 8 / 15 Cell-type level inference: Step 1 Step 1: Patient-wise transformation of p-values to Fdr •Convert p-values to local false discovery rates (lfdr) using qvalue::qvalue (Storey, Bass, Dabney, Robinson 2025). •Estimate marginal (frequentist) Fdr: Pi:pi<α lfdr(pi) PN i=1 1{pi≤α}−−−−→ N→∞ E[lfdr(p)|p≤α] = Fdr(α) 9 / 15 Conformal inference of the counterfactual cfcausal (Lei, L. and Cand` es, E., 2021) requires two conditions to be fulfilled: Assumptions 1. Stable Unit Treatment Value Assumption (SUTVA) 2. Strong ignorability (YB(1),YB(0)) ⊥T|YA ,Conditioning on post-treatment variable YA How big is the impact of the strong ignorability violation on the prediction? Future implementation with lemur bypasses the problem Cell-type level inference: step 1 Transform cells p-values to local false discovery rates lfdr(p) = P(H0|p) π0=P(Hi= 0) Conformal prediction intervals ¬Define ∆ = (Yobs B(1) −ˆ YB(0), T= 1 ˆ YB(1) −Yobs B(0), T= 0 ¬Prediction sets over grid A ⊂ [0,1] ˆ C∆,α =     YB(1) −ˆ CYB(0),α(Yobs A),T= 1 ˆ CYB(1),α(Yobs A)−YB(0),T= 0