Tl;dr Statistical generalization isn’t scientific understanding—don’t confuse prediction with insight. Foundation models like GET may predict gene expression patterns from clean data and learn patterns that appear biologically meaningful, but that’s not the same as understanding transcriptional regulation. Its outputs should be treated as hypotheses to interrogate—not definitive answers.
This week’s AI ∩ Bio: Reading the Revolution series features a recent Nature paper on “General Expression Transformer” (GET), a deep learning model trained to predict gene expression across 213 human cell types using DNA sequence and chromatin accessibility data. It’s ambitious: a transformer architecture that claims to learn the “grammar” of transcription, generalize across diverse cell types, and model long-range interactions between enhancers and promoters, as well as between transcription factors (TFs). In one case study, the authors link a leukemia-associated genetic variant (a SNP) to a disrupted protein–protein interaction between TFs, using AlphaFold structural modeling.
The aim is compelling—bringing together machine learning, epigenomics, and protein structure prediction. GET outperforms previous models like Enformer on several tasks, including prediction of reporter assay results (like MPRA) and identifying enhancer–promoter relationships. It also offers interpretability features, such as motif-level and region-level importance scores, which are an improvement over traditional “black box” models. And yes, it suggests potential utility: in theory, GET could aid in prioritizing noncoding variants or mapping regulatory networks.
But here’s the catch: GET requires both DNA sequence and chromatin accessibility data from the specific cell type of interest. You can’t simply input a variant file (like a VCF from whole-genome sequencing) and get back useful predictions. At best, you can explore hypotheses using accessibility data from public reference tissues—helpful for interpreting genome-wide association study (GWAS) results, but not yet practical for clinical use.
The generalization here is statistical—not biological, and certainly not clinical. GET was trained on harmonized, high-quality single-cell datasets from healthy human tissues. It performs well on similar data it hasn’t seen before—but that’s a narrow slice of human biology. Clinical samples, especially those from inflamed, cancerous, or drug-altered environments, often have chromatin landscapes and transcriptional programs that fall outside the model’s training distribution. High correlation on held-out healthy data doesn’t guarantee reliability under pathological conditions. The model is rigorous in its computation—but its clinical readiness remains speculative.
More in this series
- ChatNT: The future of biological assistants—or a mirage in a lab coat?
- GET: A Foundation Model for Transcription, Still Between Promise and Proof
- X-Atlas/Orion: Your Model is Only as Good as Your Training Data
- From Better Models to Better Questions: A Pathology AI Rethink
- The Model Isn’t the Magic: How Cytoland shows that domain expertise—not just deep learning—is what makes AI in biology work.
- Boltz-2: How much can 3D structure really tell us about molecular binding energetics?
- Investigating the volume and diversity of data needed for generalizable antibody–antigen ΔΔG prediction
- Beyond Binding: Rethinking Drug Design in the Age of AI and Structural Biology
- Testing the Physics Beneath the Predictions Beyond RMSD: What AlphaFold3 Really Understands
- Hype, Hurdles, and Hepatotoxicity: A Bold Step for AI-Designed Drugs, But Still Miles to Go
- When Complexity Misleads
- What Are Genomic Transformers Actually Learning?
- mRNABench and the Future of AI in Biology: Why Domain Knowledge Wins
- OncoGAN Creates Synthetic Cancer Genomes, Opening New Paths for Privacy-Preserving Precision Oncology
- Readable Rules, Testable Models: A New Grammar for Virtual Cells
- Multimodal CustOmics: Fusing Pathology Images and Tumor Genomics for Next-Gen Cancer Diagnostics
- Beyond Perturbation Simulations: PDGrapher Shows a Faster Way to Identify Actionable Targets
- Can AI design epigenetic anti-aging strategies?
- What happens when physicians use GPT-4 for diagnosis
- Can generative AI predict emergent phenomena?
- DeepSomatic and the question of how AI learns from itself
- Toward Mechanism-Centric Interpretability in Genomic Machine Learning
- From AlphaEvolve to DeepEvolve: What We’re Learning About Machine-Led Scientific Discovery
- Kosmos and the Culture of Discovery
- From Embeddings to Insight
- From Prediction to Explanation: How BIOREASON Reframes Genomic AI as a Reasoning Problem
- AI Models Need Better Truth—Platinum Pedigree Shows How
- From Evolutionary Intolerance to Clinical Insight: What popEVE Teaches Us About Missense Variants
- Interpretable Latent Spaces, Messy Biology: What AUTOENCODIX Teaches Us About Autoencoders in the Wild
- The Gene Ontology Knowledgebase in 2026
To early-career scientists: Always know the limits of your tools. Foundation models can learn statistical patterns that look biological, without understanding the underlying mechanisms. A model trained on chromatin accessibility may capture useful correlations, but not the dynamic, causal logic of transcription. And interpretability tools like motif importance scores are only meaningful if they lead to testable, falsifiable predictions that hold up under rigorous testing. Treat every output as a hypothesis to challenge.

Comments
Powered by WP LinkPress