In molecular design, we often prioritize what’s measurable over what’s meaningful.

For decades, binding affinity has served as a cornerstone of early-stage drug discovery, not because it captures biological function in full, but because it’s one of the few properties we can quantify systematically and optimize across large libraries.

Now, as AI models generate binders faster than we can validate them, we must ask: What exactly are we optimizing for? And what datasets are we training on?

What we know:

  • Strong binding doesn’t guarantee efficacy
  • Residence time and conformational flexibility can matter more than affinity
  • Cellular context — target expression, pathway crosstalk, and off-target interactions — often dictates outcome in clinical applications

Yet much of the public data — and many AI training sets — still orbit around Kd, IC₅₀, and docking scores. These are abundant and easy to label, but they capture only a narrow slice of pharmacological reality (and we’re not even accounting for the fact that these measurements are highly dependent on the specific conditions- buffer, temperature, etc).

If we train models on what’s easy to measure, we shouldn’t be surprised when they generate molecules that impress in silico — and disappoint in vivo.

The problem isn’t that binding doesn’t matter. It does. The problem is that binding isn’t biology.

Toward More Meaningful Models

To do better, we’ll need to:

  • Incorporate multi-parametric data: kinetics, permeability, metabolism, toxicity, immune activation
  • Train models to include mechanism and uncertainty, not just affinity
  • Elevate datasets that link structure to systems, not just structure to scores

The best work ahead won’t just generate molecules; it will surface better models about how they work, and where they fail in the journey from discovery to clinic.