I hope you will forgive me this week, because this week’s paper isn’t an AI paper per se… it’s a methods and data generation paper, but one that sits squarely at the intersection of AI and biology. While much of the field is focused on building more powerful models, even the best architectures can’t reliably infer causality, dose effects, or subtle regulatory responses from observational data alone. Controlled perturbations are essential for training models that work towards explanation, not just correlation. 

In this preprint, the team at Xaira Therapeutics, a next-gen biotech building AI-native drug discovery tools, introduces FiCS Perturb-seq, a scalable, industrialized platform for dose-aware single-cell CRISPR. And the open source dataset they generated using this method, X-Atlas/Orion, is something truly remarkable.  

Scientific Insight

The core breakthrough here isn’t a novel discovery, it’s the industrialization of perturbation data generation. FiCS Perturb-seq standardizes the full pipeline: fixation, cryopreservation, FACS enrichment, automation, and high-density single-cell loading. This enables X-Atlas/Orion, a dataset of 8 million dual-sgRNA cells with deep, reproducible transcriptomic profiles—built not for exploration alone, but for training dose-aware, causal foundation models. A key analytic innovation: using sgRNA abundance as a proxy for knockdown strength, enabling dose-dependent transcriptional analyses that move beyond the binary logic of most Perturb-seq studies.

For R&D leaders, this paper is a great example for how you can scale experimental biology to meet the needs of machine learning. It’s not about a single step, it’s about asking the right questions about what data is really needed and designing the system to generate that data at scale: QC before sequencing, consistent cell handling, and deep coverage per perturbation. 

More in this series

  1. ChatNT: The future of biological assistants—or a mirage in a lab coat?
  2. GET: A Foundation Model for Transcription, Still Between Promise and Proof
  3. X-Atlas/Orion: Your Model is Only as Good as Your Training Data
  4. From Better Models to Better Questions: A Pathology AI Rethink
  5. The Model Isn’t the Magic: How Cytoland shows that domain expertise—not just deep learning—is what makes AI in biology work.
  6. Boltz-2: How much can 3D structure really tell us about molecular binding energetics?
  7. Investigating the volume and diversity of data needed for generalizable antibody–antigen ΔΔG prediction
  8. Beyond Binding: Rethinking Drug Design in the Age of AI and Structural Biology
  9. Testing the Physics Beneath the Predictions Beyond RMSD: What AlphaFold3 Really Understands
  10. Hype, Hurdles, and Hepatotoxicity: A Bold Step for AI-Designed Drugs, But Still Miles to Go
  11. When Complexity Misleads
  12. What Are Genomic Transformers Actually Learning?
  13. mRNABench and the Future of AI in Biology: Why Domain Knowledge Wins
  14. OncoGAN Creates Synthetic Cancer Genomes, Opening New Paths for Privacy-Preserving Precision Oncology
  15. Readable Rules, Testable Models: A New Grammar for Virtual Cells
  16. Multimodal CustOmics: Fusing Pathology Images and Tumor Genomics for Next-Gen Cancer Diagnostics
  17. Beyond Perturbation Simulations: PDGrapher Shows a Faster Way to Identify Actionable Targets
  18. Can AI design epigenetic anti-aging strategies?
  19. What happens when physicians use GPT-4 for diagnosis
  20. Can generative AI predict emergent phenomena?
  21. DeepSomatic and the question of how AI learns from itself
  22. Toward Mechanism-Centric Interpretability in Genomic Machine Learning
  23. From AlphaEvolve to DeepEvolve: What We’re Learning About Machine-Led Scientific Discovery
  24. Kosmos and the Culture of Discovery
  25. From Embeddings to Insight
  26. From Prediction to Explanation: How BIOREASON Reframes Genomic AI as a Reasoning Problem
  27. AI Models Need Better Truth—Platinum Pedigree Shows How
  28. From Evolutionary Intolerance to Clinical Insight: What popEVE Teaches Us About Missense Variants
  29. Interpretable Latent Spaces, Messy Biology: What AUTOENCODIX Teaches Us About Autoencoders in the Wild
  30. The Gene Ontology Knowledgebase in 2026

To early-career scientists: this paper highlights the importance of infrastructure for discovery. The authors didn’t invent fixation or dual guides—they made them scalable, automatable, and reliable enough to support next-generation modeling. If you’re working at the interface of wet lab and ML, remember: your model is only as good as your training data. The signal (i.e., what am I really measuring and what can it actually tell me), the structure, and the variability you build into your system will define what the model learns and what it misses.