The PNAS Perspective by Tiwary et al. takes on one of the hardest open questions in modeling-driven science: can generative AI predict emergent phenomena?
The authors trace a careful path through the foundations of both computational chemistry and generative modeling, bridging statistical mechanics concepts like force fields and free energy landscapes with architectures including autoencoders (AEs), generative adversarial networks (GANs), flow-based diffusion models, and large language models (LLMs).
Their central argument deserves attention: models capable of predicting emergence must embed physical laws, not merely fit datasets. Statistical mechanics, thermodynamics, and quantum constraints aren’t optional. The bright spots in the field are already moving this way. Reinforcement learning grounded in the principle of maximum caliber, diffusion models inspired by nonequilibrium thermodynamics, and hybrid frameworks like AlphaFlow and AF2RAVE all point toward a new synthesis: physics as foundation, generative AI as engine.
Yet the conditional structure of biological and chemical systems sets hard limits. Most training sets collapse critical variables (temperature, solvent composition, ionic strength, and conformational heterogeneity) into latent noise. Without explicit conditioning, models risk conflating context-dependent behavior with sequence- or structure-intrinsic features. What the field needs next are frameworks that make those assumptions explicit: guidance on when each class of model is appropriate, how to diagnose failure, and how to measure progress beyond visual plausibility or interpolation accuracy.
Leadership angle
For those leading or investing in AI-driven science, the message is clear: the next leap won’t come from larger models alone, but from tighter coupling between representation and reality. The teams that will lead this next wave are those fluent in both the language of data AND the laws that govern it.
More in this series
- ChatNT: The future of biological assistants—or a mirage in a lab coat?
- GET: A Foundation Model for Transcription, Still Between Promise and Proof
- X-Atlas/Orion: Your Model is Only as Good as Your Training Data
- From Better Models to Better Questions: A Pathology AI Rethink
- The Model Isn’t the Magic: How Cytoland shows that domain expertise—not just deep learning—is what makes AI in biology work.
- Boltz-2: How much can 3D structure really tell us about molecular binding energetics?
- Investigating the volume and diversity of data needed for generalizable antibody–antigen ΔΔG prediction
- Beyond Binding: Rethinking Drug Design in the Age of AI and Structural Biology
- Testing the Physics Beneath the Predictions Beyond RMSD: What AlphaFold3 Really Understands
- Hype, Hurdles, and Hepatotoxicity: A Bold Step for AI-Designed Drugs, But Still Miles to Go
- When Complexity Misleads
- What Are Genomic Transformers Actually Learning?
- mRNABench and the Future of AI in Biology: Why Domain Knowledge Wins
- OncoGAN Creates Synthetic Cancer Genomes, Opening New Paths for Privacy-Preserving Precision Oncology
- Readable Rules, Testable Models: A New Grammar for Virtual Cells
- Multimodal CustOmics: Fusing Pathology Images and Tumor Genomics for Next-Gen Cancer Diagnostics
- Beyond Perturbation Simulations: PDGrapher Shows a Faster Way to Identify Actionable Targets
- Can AI design epigenetic anti-aging strategies?
- What happens when physicians use GPT-4 for diagnosis
- Can generative AI predict emergent phenomena?
- DeepSomatic and the question of how AI learns from itself
- Toward Mechanism-Centric Interpretability in Genomic Machine Learning
- From AlphaEvolve to DeepEvolve: What We’re Learning About Machine-Led Scientific Discovery
- Kosmos and the Culture of Discovery
- From Embeddings to Insight
- From Prediction to Explanation: How BIOREASON Reframes Genomic AI as a Reasoning Problem
- AI Models Need Better Truth—Platinum Pedigree Shows How
- From Evolutionary Intolerance to Clinical Insight: What popEVE Teaches Us About Missense Variants
- Interpretable Latent Spaces, Messy Biology: What AUTOENCODIX Teaches Us About Autoencoders in the Wild
- The Gene Ontology Knowledgebase in 2026
Mentorship angle
For early-career scientists, this is an invitation to think rigorously about foundations.
Learn the physics as well as the Python.
Understand how bias enters your data and what it does to inference. The next breakthroughs won’t come from models that memorize reality, but from those that explain it, and can then predict new emergent phenomena.

Comments
Powered by WP LinkPress