This year’s Gene Ontology (GO) update is a reminder that infrastructure choices shape scientific conclusions, getting to the heart of this foundational tool for understanding biology at a time when omics, enrichment analyses, and AI models increasingly rely on GO as biological “ground truth.”
Are you new to Gene Ontology? See the PDF for a deeper dive.
What actually changed (2022–2025)
A few highlights that matter in practice:
- Major ontology cleanup: hundreds of new terms added, thousands of imprecise or redundant terms obsoleted.
- Human Functionome v2.0: a reviewed, integrated annotation set now covering ~84% of human genes, reducing enrichment clutter while preserving biological relevance.
- GO-CAMs scaled up: >1,500 expert-curated causal pathway models linking gene activities with evidence, moving beyond flat gene lists toward mechanistic flow.
Why this paper matters for diagnostics, AI, and innovation leaders
GO and AI models share something important: both are compressions of complex biology.
- Gene Ontology is a structured compression of current biological knowledge, but lack explicit biological context.
- Genomic language models are statistical compressions of high-dimensional data, but lack biological grounding.
- GO provides a curated prior (a biological sanity check) but it abstracts away context (cell state, disease, rewiring). In cancer, that context is often the signal. Used well, GO disciplines thinking and prevents nonsense. Used naively, it produces answers that look rigorous, but are nonsensical.
- The opposite risk exists with genomic language models. They learn dense embeddings that can capture patterns not explicitly labeled, but the derived “understanding” is not mechanistic by default; it’s statistical compression. And they can overindex on historical data distributions, which can amplify biases.
An intriguing option (and a common one in many recent AI Bio papers), is to combine Gene Ontology with language models.
For example:
Language model → propose;
Gene Ontology → check.
Use a language model to propose functional/interaction hypotheses from data & Gene Ontology to flag things like contradictions and flag known process vs believable novelty vs likely nonsense.
More in this series
- ChatNT: The future of biological assistants—or a mirage in a lab coat?
- GET: A Foundation Model for Transcription, Still Between Promise and Proof
- X-Atlas/Orion: Your Model is Only as Good as Your Training Data
- From Better Models to Better Questions: A Pathology AI Rethink
- The Model Isn’t the Magic: How Cytoland shows that domain expertise—not just deep learning—is what makes AI in biology work.
- Boltz-2: How much can 3D structure really tell us about molecular binding energetics?
- Investigating the volume and diversity of data needed for generalizable antibody–antigen ΔΔG prediction
- Beyond Binding: Rethinking Drug Design in the Age of AI and Structural Biology
- Testing the Physics Beneath the Predictions Beyond RMSD: What AlphaFold3 Really Understands
- Hype, Hurdles, and Hepatotoxicity: A Bold Step for AI-Designed Drugs, But Still Miles to Go
- When Complexity Misleads
- What Are Genomic Transformers Actually Learning?
- mRNABench and the Future of AI in Biology: Why Domain Knowledge Wins
- OncoGAN Creates Synthetic Cancer Genomes, Opening New Paths for Privacy-Preserving Precision Oncology
- Readable Rules, Testable Models: A New Grammar for Virtual Cells
- Multimodal CustOmics: Fusing Pathology Images and Tumor Genomics for Next-Gen Cancer Diagnostics
- Beyond Perturbation Simulations: PDGrapher Shows a Faster Way to Identify Actionable Targets
- Can AI design epigenetic anti-aging strategies?
- What happens when physicians use GPT-4 for diagnosis
- Can generative AI predict emergent phenomena?
- DeepSomatic and the question of how AI learns from itself
- Toward Mechanism-Centric Interpretability in Genomic Machine Learning
- From AlphaEvolve to DeepEvolve: What We’re Learning About Machine-Led Scientific Discovery
- Kosmos and the Culture of Discovery
- From Embeddings to Insight
- From Prediction to Explanation: How BIOREASON Reframes Genomic AI as a Reasoning Problem
- AI Models Need Better Truth—Platinum Pedigree Shows How
- From Evolutionary Intolerance to Clinical Insight: What popEVE Teaches Us About Missense Variants
- Interpretable Latent Spaces, Messy Biology: What AUTOENCODIX Teaches Us About Autoencoders in the Wild
- The Gene Ontology Knowledgebase in 2026
A note to early-career scientists
Impact doesn’t only come from novelty. It comes from:
- caring about definitions and evidence,
- understanding the assumptions baked into your tools,
- and knowing where structure helps, and where it hides uncertainty.
Bridge discovery with discipline, and insight with infrastructure, and you’ll do work that lasts.
Skip to PDF content
Comments
Powered by WP LinkPress