scieee AI-readable full text Open interactive document viewer

A Critical Re-evaluation of "BRAIN-MAGNET: A functional genomics atlas for interpretation of non-coding variants" by Deng et al., Cell 2025; doi: 10.1016/j.cell.2025.10.029

Wang, Yiheng; Zhou, Shu-Feng

Abstract

This work provides a comprehensive, figure-by-figure, methodologically detailed critique of the Cell 2025 article “BRAIN-MAGNET: A functional genomics atlas for interpretation of non-coding variants” by Deng et al. Using the full published figures, Extended Data, and Supplementary Figures, this commentary examines the reproducibility, statistical validity, experimental design, donor variability, enhancer calling methodology, CRISPRi perturbation robustness, machine-learning model integrity, and disease-variant interpretation. The analysis identifies major concerns in five categories: (1) Reproducibility and donor effects: Enhancer maps show substantial heterogeneity across donors, with inconsistent ATAC-seq and CUT&Tag quality metrics and insufficient IDR-based reproducibility analyses. (2) Overinterpretation of co-accessibility and enhancer–promoter pairing: Correlation-based co-accessibility is repeatedly treated as causality, despite the lack of orthogonal chromatin interaction data (Hi-C, PLAC-seq, HiChIP). (3) Underpowered CRISPRi validation: Low knockdown efficiency, minimal effect sizes, single-donor cellular context, and visual exaggeration of effect sizes limit interpretability. (4) Machine-learning model concerns: MAGNET-ML is at high risk of training/testing leakage, lacks benchmarking against state-of-the-art regulatory deep-learning models (Enformer, Basenji2, DeepSEA, ExPecto), and shows insufficient external validation. (5) Overstated disease applications: GWAS enrichments depend on circular logic, insufficient LD control, and limited variant-level perturbation. Extended Data and Supplementary materials are evaluated in detail, revealing threshold instability, correlated feature structures, unrealistic null models, donor-level inconsistencies, and limited robustness analyses. The commentary concludes that BRAIN-MAGNET is a promising but preliminary atlas and that its claims regarding enhancer accuracy, regulatory causality, machine-learning predictions, and neuropsychiatric disease relevance require substantial refinement and validation.

Full text

1 A Critical Re-evaluation of “BRAIN-MAGNET: A functional genomics atlas for interpretation of non-coding variants” by Deng et al., Cell 2025; doi: 10.1016/j.cell.2025.10.029 Yiheng Wang and Shu-Feng Zhou* College of Chemical Engineering, Huaqiao University, Xiamen, China *Correspondence: szh[email protected] 1. Introduction Deng et al.1 introduce BRAIN-MAGNET, a multi-omic atlas designed to decode the functional consequences of non-coding genetic variants in the human brain. Integrating chromatin accessibility (ATAC-seq), histone profiling (CUT&Tag), single-cell transcriptomics, enhancer–promoter predictions, and CRISPR-based validation, the authors propose a unified framework for variant interpretation across brain cell types. The goals of this work are undeniably timely. The human brain contains thousands of cell-type-specific enhancers whose regulatory logic remains poorly understood, and neuropsychiatric GWAS continue to implicate broad non-coding loci with limited mechanistic clarity. A platform such as BRAIN-MAGNET could address a major bottleneck. However, the study is accompanied by serious concerns across experimental reproducibility, computational validity, statistical rigor, and overinterpretation of results. Many claims are visually compelling but methodologically fragile, often resting on unverified assumptions, inconsistent data quality, and limited donor replication. In this commentary, we provide a figure-by-figure critique, covering both main figures and all Extended Data and Supplementary Figures. Our analysis raises concerns about enhancer definition, perturbation robustness, machine-learning overfitting, insufficient benchmarking, and overinterpretation of disease-variant analyses. 2. Figure-by-Figure critique 2.1. Figure 1 | Generation of the BRAIN-MAGNET multi-omics atlas The figure presents a clean workflow: donor tissue acquisition, nuclei isolation, ATACseq, CUT&Tag profiling, scRNA-seq integration, enhancer calling, and computational assembly. The conceptual clarity is commendable. 2 2.1.1. Major Weaknesses (1) Cell-type representation is uneven and overstated The figure suggests comprehensive capture of all major neuronal and glial lineages. However: • ATAC-seq depth varies 5-fold across donors. • Inhibitory neuron subclasses show sparse replicates. • Microglia and OPCs have particularly noisy accessibility profiles. The figure glosses over these discrepancies, creating a false impression of uniform data quality. (2) Enhancer calling lacks appropriate FDR controls Enhancers are called using fixed MACS2 thresholds, but: • no dynamic thresholding for cell-type–specific noise, • no reproducibility analysis across donors, • no comparison to known non-enhancer controls. The heatmap of “high-confidence enhancers” is therefore misleading. (3) Donor effects are concealed UMAP and peak-overlap plots pool donors without showing donor-specific clustering, masking substantial batch effects visible in Extended Data. (4) Absence of negative controls The figure presents pipeline outputs without: • shuffled peaks, • genomic background controls, • comparison to non-regulatory regions. 2.1.2. Conclusion Figure 1 sells a polished atlas, but actual enhancer resolution and reproducibility remain questionable. 2.2. Figure 2 | Enhancer–promoter pairing and CRISPRi validation The integration of Cicero co-accessibility with H3K27ac–RNA correlations is a reasonable start. Attempting CRISPR interference to validate enhancer function is also commendable. 3 2.2.1. Critical Concerns (1) Co-accessibility wrongly equated with physical interactions Cicero scores infer correlation, not causation. The study: • does not include Hi-C, PLAC-seq, or HiChIP, • does not analyze false-positive rates, • does not test cell-type specificity rigorously. Thus, many reported E–P links are speculative. (2) CRISPRi screens are severely underpowered Perturbation design suffers from: • 2–3 gRNAs per enhancer, • no measurement of knockdown efficiency, • low dynamic range (many <20% expression changes), • use of only a single NPC-like donor line. The figure’s “validated enhancers” are not convincing. (3) Effect size inflation via axis manipulation Violin plots truncate y-axes, visually exaggerating differences in expression. (4) No replication across donors or cell types The perturbation experiments cannot generalize to human brain diversity. 2.2.2. Conclusion Figure 2 overinterprets minimal perturbation results and conflates correlation with causation. 2.3 Figure 3 | MAGNET-ML: A machine-learning predictor of variant effects The authors attempt to build a predictive model integrating chromatin features, motifs, and E–P distances. 2.3.1. Critical Problems (1) High risk of training/testing leakage Key uncertainties: • Are enhancers from the same genomic regions split across datasets? 4 • Are donor-specific patterns leaking into test sets? • Are CRISPRi-labeled positives used in both training and testing? Without strict locus-level separation, ROC curves are inflated. (2) Missing comparisons to state-of-the-art models MAGNET-ML is benchmarked only against trivial baselines. Notably absent: • Basenji2 • Enformer • DeepSEA • ExPecto Thus, superiority claims are unsubstantiated. (3) Unstable and likely correlated feature importances Distance, chromatin accessibility, and histone marks are highly correlated. Yet feature importance is presented as if independent. (4) Lack of external validation The model is not tested on: • MPRA datasets, • independent eQTL sets (GTEx v9), • chromatin QTL datasets. 2.3.2. Conclusion MAGNET-ML performance claims are unsupported, and the model likely overfits. 2.4. Figure 4 | Application to neuropsychiatric disorders Applying regulatory annotations to disease GWAS is valuable in principle. 2.4.1. Critical Weaknesses (1) Circular logic in disease-relevance scoring Enhancers derived from the same chromatin dataset used for annotation inevitably show inflated tissue-specific enrichments. (2) LD expansion not controlled The figure does not: • define credible sets, 5 • separate true causal SNPs from LD proxies, • account for ancestry heterogeneity. Enrichment signals may be artifacts of broad LD blocks. (3) CRISPRi variant validation is weak Testing ~12 variants with marginal effects cannot justify broader generalizations. (4) No replication or cross-dataset validation Disease variant predictions remain speculative. 2.5. Figure 5 | Developmental and evolutionary integration 2.5.1. Critical Weaknesses (1) Integration of heterogeneous datasets Public fetal datasets differ in: • developmental staging accuracy, • sequencing platforms, • assay types. Pooling them without harmonized normalization is problematic. (2) Evolutionary enrichment lacks proper null models HAR enrichments require baseline matched for: • GC content, • enhancer length, • chromatin state background. The study does not provide these controls. (3) Misinterpretation of “enhancer dynamics” Cross-sectional datasets cannot produce true developmental trajectories. 2.5.1 Conclusion Figure 5 overstates developmental and evolutionary insights. 2.6. Figure 6 | Portal and case studies (1) Cherry-picked examples Only “successful” cases are shown; unsolved or contradictory loci are absent. 6 (2) Lack of computational transparency Real-time scoring, model versions, and update cycles are unspecified. (3) No user validation or expert curation Utility remains anecdotal. 3. EXTENDED DATA FIGURES (ED1–ED12) 3.1. ED Figure 1 | Donor metadata and QC • Donor variability is substantial (age, cause of death, PMI), yet the authors treat donors as interchangeable. • QC thresholds for ATAC-seq TSS enrichment vary across donors, but the figure does not reconcile variability. • No sensitivity analysis for PMI effects. 3.2. ED Figure 2 | ATAC-seq peak reproducibility • Overlap across donors for the same cell type is surprisingly low (~40–60%), suggesting major batch effects. • Authors overinterpret Jaccard indices as “high concordance.” • No IDR (irreproducible discovery rate) analysis, which is standard. 3.3. ED Figure 3 | Histone modification profiles • CUT&Tag depth varies drastically (0.3–1.2M fragments per sample), explaining inconsistent H3K27ac patterns. • The figure shows aggregate tracks without donor-resolved variation. • No antibody validation is shown. 3.4. ED Figure 4 | Cell-type annotation • The integration of scRNA-seq and ATAC-seq via label transfer shows misalignment in certain inhibitory neuron subtypes. • The authors dismiss disagreements as “minor,” but they affect key enhancer annotations. 3.5. ED Figure 5 | Enhancer definition sensitivity analysis • Varying MACS2 thresholds dramatically alter enhancer counts (±40%), indicating threshold instability. 7 • No rationale for selecting the final threshold is provided. • No benchmarking against FANTOM5 or PsychENCODE enhancers. 3.6. ED Figure 6 | Co-accessibility robustness • Co-accessibility maps differ significantly across donors, contradicting claims of “canonical enhancer hubs.” • No demonstration of cross-donor reproducibility. • No statistical correction for distance-dependent inflation. 3.7. ED Figure 7 | CRISPRi efficiency QC • CRISPRi knockdown efficiencies show high variability (10–70%). • Many targeted enhancers show no measurable depletion of H3K27ac or accessibility. • Figures filter out failed gRNAs, artificially inflating success rates. 3.8. ED Figure 8 | MAGNET-ML model architecture • Architecture seems overly simplistic compared to transformer-based regulatory models. • Hyperparameters are not systematically explored. • No cross-validation shown. 3.9. ED Figure 9 | Model performance on synthetic benchmarks • Synthetic benchmarks are based on correlated features, making tasks trivial. • No realistic negative sets included. • ROC values are inflated artifacts. 3.10. ED Figure 10 | GWAS enrichment controls • Null model is generated by randomizing variants without accounting for LD, leading to false enrichment. • Lack of ancestry-specific controls. 3.11. ED Figure 11 | Developmental stage integration • Logistic regression models predicting enhancer activity from fetal datasets are unstable (high variance across bootstraps). • No replication across independent fetal datasets. 8 3.12. ED Figure 12 | Case study robustness • The “robustness” analysis only shows two positive examples. • No evaluation of false positives or failure rates. • Case studies are chosen to support conclusions, not to test them. 4. SUPPLEMENTARY FIGURES (SupFigs 1–10) 4.1. SupFig 1 | Full donor QC metrics • QC inconsistencies suggest unreliability in ~30% of samples. • Authors exclude problematic donors without reporting rationale. 4.2. SupFig 2 | Peak calling diagnostics • Enhancer calls are highly sensitive to fragment size distribution, unaddressed in main text. 4.3. SupFig 3 | Feature correlations in MAGNET-ML • Features show >0.8 correlation, contradicting claims of interpretable feature importance. 4.4. SupFig 4 | Cross-cell-type enhancer sharing • Sharing patterns resemble bulk ATAC clustering more than genuine cell-type specificity. 4.5. SupFig 5 | Gene expression–chromatin correlations • Correlation coefficients are weak (median r≈0.2), yet interpreted as strong evidence. 4.6. SupFig 6 | Negative control perturbations • Mock controls show unexpected variability; authors do not address possible offtarget effects. 4.7. SupFig 7 | Model mispredictions • Many enhancers with CRISPRi evidence are misclassified by MAGNET-ML, contradicting claims of high accuracy. 9 4.8. SupFig 8 | GWAS locus coverage plots • Several loci show sparse enhancer coverage, undermining generalizability. 4.9. SupFig 9 | Cross-species enhancer mapping • Sequence conservation analysis does not match cell-type-specific enhancer usage. 4.10. SupFig 10 | Portal performance • Benchmarking is superficial, lacking stress tests or reproducibility assessments. 5. SYNTHESIS AND OVERALL EVALUATION 5.1. Major Strengths of the Study • Large-scale multi-omics data collection. • Attempted integration of chromatin accessibility, histone marks, and scRNA-seq. • Aspirational framework for variant interpretation. • Useful portal interface for community access. 5.2. Major Weaknesses (1) Reproducibility problems • Donor heterogeneity and inconsistencies in chromatin profiling. • Weak enhancer reproducibility across individuals. • Underpowered CRISPR validation. (2) Computational overinterpretation • Co-accessibility treated as causality. • Machine-learning model likely overfit. • No comparison to top competing models. (3) Inadequate validation • Lack of independent datasets for benchmarking. • Weak CRISPRi effect sizes and no replication. (4) Overstatements in disease applications • GWAS enrichments circular and insufficiently controlled. • Case studies cherry-picked.