I hope you will forgive me this week, because this week’s paper isn’t an AI paper per se… it’s a methods and data generation paper, but one that sits squarely at the intersection of AI and biology. While much of the field is focused on building more powerful models, even the best architectures can’t reliably infer causality, dose effects, or subtle regulatory responses from observational data alone. Controlled perturbations are essential for training models that work towards explanation, not just correlation.
In this preprint, the team at Xaira Therapeutics, a next-gen biotech building AI-native drug discovery tools, introduces FiCS Perturb-seq, a scalable, industrialized platform for dose-aware single-cell CRISPR. And the open source dataset they generated using this method, X-Atlas/Orion, is something truly remarkable.
Scientific Insight
The core breakthrough here isn’t a novel discovery, it’s the industrialization of perturbation data generation. FiCS Perturb-seq standardizes the full pipeline: fixation, cryopreservation, FACS enrichment, automation, and high-density single-cell loading. This enables X-Atlas/Orion, a dataset of 8 million dual-sgRNA cells with deep, reproducible transcriptomic profiles—built not for exploration alone, but for training dose-aware, causal foundation models. A key analytic innovation: using sgRNA abundance as a proxy for knockdown strength, enabling dose-dependent transcriptional analyses that move beyond the binary logic of most Perturb-seq studies.
For R&D leaders, this paper is a great example for how you can scale experimental biology to meet the needs of machine learning. It’s not about a single step, it’s about asking the right questions about what data is really needed and designing the system to generate that data at scale: QC before sequencing, consistent cell handling, and deep coverage per perturbation.
To early-career scientists: this paper highlights the importance of infrastructure for discovery. The authors didn’t invent fixation or dual guides—they made them scalable, automatable, and reliable enough to support next-generation modeling. If you’re working at the interface of wet lab and ML, remember: your model is only as good as your training data. The signal (i.e., what am I really measuring and what can it actually tell me), the structure, and the variability you build into your system will define what the model learns and what it misses.
