Leadership in Biotech

Author: Kristin Gleitsman Page 2 of 7

Scientist. Research and Development Leader. Collaborative Problem Solver.
Passionate about building belonging, so that together we can accomplish important, hard things.

Illustration of a desk with a figure from a recent paper about the pLM

From Embeddings to Insight

Protein language models (pLMs) learn from raw amino-acid sequences and turn each protein into a numerical “embedding.” Unlike AlphaFold (which is trained to predict 3D structure), pLMs are trained only on sequence patterns, then reused for tasks like fast homology search, function hinting, or variant triage.

The appeal of pLMs is that they might:

  • Identify remote homology
  • Assist in function annotation, especially for uncharacterized proteins
  • Predict mutation effects
  • Serve as general-purpose inputs for fold prediction, domain classification, etc.

This study benchmarks 14 pLMs to evaluate how well they capture biological similarity along three axes:

  • Sequence
  • Structure
  • Function

For each axis, the authors compare distances between embeddings to a “ground truth” similarity metric (see below). They test whether models reflect these similarities in two ways:

  • Inherent information: Does the raw embedding distance between two proteins correlate with their similarity?
  • Extractable information: Can a small model trained on the embeddings predict the similarity score?

Scientific Insight

This paper provides a systematic, well-designed benchmark for protein language models, useful for those developing or deploying these models computationally.

But it defines “success” in terms of alignment with thorny labels, which is a limitation of the field at the moment. See PDF below.

With that caveat, the results are compelling:

Out-of-the-box, small models often perform as well as large ones.

  • “Size-performance paradox”: model size doesn’t guarantee better embedding quality for biological similarity.

Larger models encode more information, but it’s hidden.

  • You need to train a model on the embeddings to “extract” biological signal.

Task-specific models don’t generalize.

  • Fine-tuning a model for one problem (e.g., enzyme function) distorts the embedding space and reduces general usefulness.

Leadership Angle

Bigger isn’t automatically better (we’re seeing this with genomic models as well). But also be mindful of what you are benchmarking to, and its relevance to what you care about.

Mentorship Angle

Let pLMs speed the front end of discovery, then make the science real.

  • Start with a testable claim in plain language—e.g., “Nearest neighbors in embedding space enrich for shared catalytic residues better than a sequence-identity cutoff at the same recall.”
  • Pre-register thresholds; include hard negatives (same fold, different function; low identity, same mechanism) & use controls (shuffle labels, etc)
  • Report where it fails—disordered regions, multi-domain proteins, complexes, and say why you think it failed
  • Document assumptions (databases, structure confidence, training leakage)

The win isn’t a pretty plot; the win is turning speed into better experiments that change what you do at the bench.

Skip to PDF content
Illustration of a desk with a figure from a recent paper about the Kosmos AI agent

Kosmos and the Culture of Discovery

AI Scientists are all the rage these days, and that excitement ramped up a notch or ten this week with the announcements of the Kosmos preprint from prominent AI researchers (notably, FutureHouse). The Kosmos system takes on an ambitious question:

Can an AI not only assist with science, but do science on its own?

Designed to read literature, analyze data, and generate new hypotheses in 12-hour autonomous runs, Kosmos reports nearly 80% statement accuracy and the equivalent of six months of human research per cycle.

It’s an extraordinary technical achievement…

…and…

…one that forces us to ask what, exactly, counts as scientific discovery?

Scientific Insight

At its core, Kosmos is a multi-agent system. One agent searches the literature, another analyzes data, and a coordinating model stitches their findings together into a cohesive research narrative.

The architecture is impressive and elegant, but it also reveals a key limitation.

Kosmos optimizes for coherence—for ideas that fit neatly together—rather than for falsifiability or experimental test.

The result is a system that can produce consistent and compelling stories, but not yet the self-correcting friction that turns a story into durable scientific insight.

Leadership Angle

For those of us leading R&D organizations, Kosmos is both inspiring and instructive. It shows how far autonomous reasoning has come. And it also demonstrates how easily coherence can masquerade as progress.

In the context of industrial scientific research, this lesson feels particularly relevant. Our job isn’t to chase automation for its own sake (although driving down cost is certainly a constant imperative), it’s to develop products that are safe, effective, and hold up in the real world.

To accomplish this task, we need to design scientific teams where human judgment and machine synthesis elevate the best of what each brings to the table.

Our new AI teammate is here, and in order to figure out how to integrate them safely and effectively with your human team, learning to manage them effectively is absolutely critical.

Mentorship Angle

For early-career scientists, Kosmos highlights part of what the future of science will look like, so pay attention to what these AI ‘scientists’ can and cannot deliver, and how they evolve.

Right now, Kosmos is fast, thorough, and tireless, but optimized to find coherence. The craft of science still lives in that space of productive stupidity and intellectual humility: the messy, uncertain, human part where you argue with data (and with your fellow scientists), question assumptions, and let yourself be wrong. AI can’t automate that part (at least not yet).

If Kosmos points to a future of machine collaborators, then the most valuable skill you can build now is learning how to think with them—and sometimes, against them.

Generative image of a group of four colleagues hugging and laughing together in an office

From Loneliness to the Connection Operating System

When Luis Velasquez and I first wrote about loneliness as a leadership challenge in Harvard Business Review, the response was gratifying. Leaders weren’t only intrigued, they were relieved. It put words to something many of us had felt but hadn’t named: that loneliness isn’t just showing up as a societal epidemic, it’s revealing itself in how teams work, lead, and connect.

For the last several weeks, I’ve been exploring what it actually looks like to build connection into the operating fabric of a team in a series of weekly posts. I’ve shared stories from my own work — not grand gestures, but small, durable practices that quietly shape belonging:

  1. Shared identity through rituals like our Discovery team onsite — infinity peanuts, Pictionary Telephone, and team tees that make us laugh and remember who we are.
  2. Collaboration by design, not department — forming peer-led GenAI learning groups that built trust across silos while upskilling the company.
  3. Leading with humanity, even when it means sharing a mistake — like the story I told new hires at Guardant about the time I failed early and was met with grace.
  4. Designing for belonging through structure — team-driven “About Me” rituals and annual culture calibrations that keep connection alive as a shared ethos, not a top-down edict.
  5. Sustaining connection at the top through Rising Women in Biotech — a circle that started with a single email and became a space for truth, trust, and perspective.

Each of these moments — personal, practical, sometimes a little silly — reinforces the same lesson: connection scales through intention, not intensity.

Connection is not a workplace perk or an HR initiative. It’s infrastructure: the quiet architecture beneath how people see each other, how learning happens, and how work gets done.

And now, we have new tools to help us do it better.

AI can help instead of hinder connections

AI can’t create trust, but it can help us remember, retell, and reinforce it — capturing stories, surfacing themes, and prompting reflection before the moments slip away.

If loneliness is the epidemic of our time, connection is our shared antidote.

And leaders, especially those of us building at the edge of science and technology, have both the opportunity and the responsibility to design for it.

Even though this is my last post in this series, I’d love to keep the conversation going. Please reach out and connect if this is a space you also care deeply about!

Illustration of a desk with figures from recent papers about DeepEvolve and AlphaEvolve

From AlphaEvolve to DeepEvolve: What We’re Learning About Machine-Led Scientific Discovery

Over the past year, AlphaEvolve (DeepMind) and DeepEvolve (Liu et al.) have taken on one of science’s most audacious challenges: can machines not only execute discovery but originate it? Both leverage large language models (LLMs) to iteratively evolve algorithms through feedback and automated evaluation.

AlphaEvolve reframed discovery as code evolution: LLMs proposing edits, testing them, and optimizing against performance scores. DeepEvolve extends this paradigm by integrating deep research: literature retrieval, structured reasoning, multi-file implementation, and automatic debugging. Together, they signal a shift from code-writing tools to more active scientific collaborators and raise foundational questions about what counts as “understanding.”

Scientific Insight

AlphaEvolve’s achievements are real and impressive: it discovered a novel algorithm for 4×4 matrix multiplication, improving on a benchmark that had stood since 1969. But its limitations are equally instructive. When optimization is tethered entirely to scalar reward signals, models risk learning how to perform rather than how to understand.

DeepEvolve addresses some of this by incorporating retrieval-augmented reasoning and grounded iteration, showing measurable gains across nine domains—from molecular property prediction to polymer engineering. Yet the deeper challenge persists: these systems can refine heuristics, but they cannot yet articulate principles; the boundary between discovery and reward chasing remains unresolved.

Leadership Angle

For R&D leaders in life sciences or diagnostics, these systems are both technical achievements and strategic case studies.

Peter Drucker once warned: “What gets measured gets managed—even when it’s pointless to measure or manage it.” In the context of these papers, the risk is clear: if we measure only performance, we may optimize into blind alleys. The opportunity lies in designing systems (and organizations!) that ask better questions, not just produce better scores.

Mentorship Angle

For early-career scientists, the message isn’t to fear or avoid these tools (please don’t, these tools are amazing!), but to understand their limitations. AlphaEvolve and DeepEvolve can automate exploration, but perhaps not judgment. They can generate thousands of hypotheses in hours, but still depend on human insight to distinguish signal from noise. And require careful thought in setting up reward systems that uncover real meaning in the context you care about.

In a world where even machines can “research,” the defining trait of good science will be the discipline to ask whether improvement is meaningful, not just measurable. That’s the work of real discovery and no algorithm can evolve that for us (at least not yet!).

Photo of three women holding up wine glasses silhouetted by a sunset at dusk

Case Studies in Connection: Even Leaders Need Connection

The higher you rise in your career, the fewer peers you have …and the pressure to appear composed and decisive only grows. It’s no wonder they say it’s lonely at the top.

One of the most meaningful antidotes for leadership loneliness for me has been Rising Women in Biotech.

We’re a small group of women leaders in life sciences and diagnostics who meet quarterly. On the surface, it looks like a leadership roundtable. In reality, it’s become a circle of trust, a place to bring career transitions, tough decisions, and the sticky points that are hard to navigate alone.

The funny thing is, it started with a simple email I sent out. Some of the women I knew well, others I had only met once or twice. We were joking about that email recently and how something so small became a group that now feels indispensable (at least to me).

That’s the power of peer circles: they don’t need to be complicated. They just need intention and trust.

Leaders often underestimate how much their own connection needs shape the cultures they create. If we’re isolated, we risk cascading isolation down the org.

But if we model the opposite — building our own scaffolds of trust and perspective — we give others permission to do the same.

AI could make this easier too:

  • Peer Circle Curator: imagine downloading your contacts from LinkedIn, uploading them to an AI tool, and having it suggest potential circle members who share values, challenges, or experiences, including people you might not think to invite.
  • Reflection Prompts: AI can also help capture what you take away from these conversations, turning raw insights into stories or reminders you can carry back to your teams.

Culture mirrors the top. If leaders are too busy to be human, teams will be too. But when we invest in our own connection, we create the conditions for everyone else to thrive.

Question for you: Who would you include if you sent that one email to start your own circle of trust?

Illustration of a desk with a figure from a recent paper about interpretable deep learning applications

Toward Mechanism-Centric Interpretability in Genomic Machine Learning

Rather than reviewing a paper, this week’s post takes a broader view on where we are at with respect to interpretability in genomics ML models.

As the field continues to rigorously interrogate the most recent ML genomics models, it has become clear that it’s incredibly easy to fool ourselves about what these models are learning, what they can predict, and how to do better.

One component of improving on the current state is to get more serious about interpretability. Most interpretability efforts remain retrospective—feature rankings, attention maps, or gradient plots that rationalize outputs but rarely reveal how, or whether, the model’s reasoning aligns with biology.

If we care about mechanism (and we should, because this is how models become more extensible and useful), we need a shift in stance. Interpretability should not be a gloss applied at the end of analysis, it should be part of how models are built, tested, and revised.

Here I posit that there are four questions worth asking of every architecture and dataset to help us move in that direction.

What are we interpreting: mechanisms, predictions, or confounds?

Each target demands a different standard of evidence. Mechanistic interpretability seeks causal structure; predictive interpretability seeks justification; artifact detection seeks bias. Without distinguishing them, we risk mistaking coherence for truth.

What biological hypotheses are encoded in the model architecture?

Every design choice carries an implicit worldview: MLPs flatten dependencies; GNNs canonize known graphs; transformers elevate context as signal. These are not neutral—they shape what the model is capable of discovering, and what it will systematically miss.

Can multimodal data be used to falsify interpretations?

Adding data layers isn’t just about increasing modeling power. Done correctly, an additional modality can act to challenge the others, serving as an independent test of whether the model’s inferences hold up under a different lens.

How can interpretability inform model iteration?

Used well, interpretability is diagnostic. It surfaces blind spots: missing biological priors, unrepresentable hierarchies, or architectural constraints that obscure mechanism. Those failures are invitations to refine both model and experiment.

Why it matters

Interpretability is not a transparency feature; it’s a scientific claim about correspondence between computation and biology.
And like any scientific claim, it must be testable, falsifiable, and revised in light of evidence.

Four coworkers sit at a table looking at a laptop together. They are happy and enjoying the interaction

Case Studies in Connection: Design for Belonging

Belonging doesn’t just happen, it’s designed through the routines that shape how teams work together.

One of our simplest, most consistent rituals is that we start every team meeting with a short “About Me” from a rotating member. It’s two minutes, one slide, sometimes a photo. Over time, it’s transformed us from a collection of experts into a community of known people.

We know whose kid just started a new school, who’s training for a triathlon, whose garden is exploding with chile peppers. We even discovered an improbable number of birders on our team this way. These details accumulate into trust.

Another anchor is our annual culture and operating-rhythm review. Once a year we pause to ask:

  • Are our meetings effective?
  • Are we working on the right things?
  • Do our shared values still guide us?

Then we commit to concrete changes based on what we discuss. Sometimes that means reworking meeting cadence, clarifying decision roles, or adding new rituals.

Building shared practices

Here’s what’s important: these practices weren’t decreed from the top. The idea for “About Me” came from the team itself. And while I create the space for our annual review, I haven’t led the session for the past two years, team members have. Culture in our team is not a leadership initiative, it’s a shared ethos that everyone owns and shapes.

AI could quietly strengthen practices like these by:

  • Sense-making: surfacing themes across culture check-ins and feedback.
  • Memory: tracking commitments and reminding teams months later what’s evolved.
  • Support: suggesting new “About Me” prompts or reflections to keep stories fresh and inclusive.

Connection scales when it’s built into the way we work, not as something extra, and not as something handed down, but as the operating system we all help maintain.

Question for you: What’s one small practice your team repeats that quietly turns strangers into collaborators?

Illustration of a desk with a figure from a recent paper about DeepSomatic

DeepSomatic and the question of how AI learns from itself

DeepSomatic, published this month in Nature Biotechnology, represents a milestone for cancer genomics: a deep-learning method that detects somatic small variants across both short- and long-read sequencing data. Built on Google’s DeepVariant framework, it bridges Illumina, PacBio HiFi, and Oxford Nanopore datasets and introduces CASTLE, a new multi-platform benchmark of six tumor–normal cell lines made openly available to the community. For anyone working in precision oncology, the technical ambition here is remarkable: one model spanning technologies, sample types, and variant classes.

Scientific Insight

DeepSomatic converts paired tumor–normal reads into tensor “images” that feed a convolutional neural network capable of distinguishing somatic, germline, and reference variants. The model outperformed leading tools such as Strelka2 and ClairS across variant types and variant allele frequencies, and it maintained accuracy across multiple sequencing chemistries. Beyond its raw performance, the CASTLE dataset fills a major gap in the field: creating a real benchmark for long-read somatic variant detection where none previously existed.

Scientific Rigor Note

Like many GenAI systems, DeepSomatic may fall pray to non-obvious data leakage, and would benefit from more explainability. Some of its evaluation data overlap with the model’s own training inputs, raising the risk of circular benchmarking bias, and the study offers little insight into why the network makes its calls.

Leadership & Mentorship Reflection

Building trustworthy AI in medicine requires independent data, transparent reasoning, and humility about limitations that are baked into how these models work.

For early-career scientists, this paper is a case study in responsible ambition: innovate boldly, share your data openly, and interrogate your own benchmarks. AI or not, progress comes from rigorous, open science that understands its own limitations.

Photo of a man sitting in the woods, contemplating what he's looking at on his laptop. Photo by Hamza Tighza.

Case Studies in Connection: Lead with Humanity

At my previous company, I used to share a story during onboarding.

It wasn’t about a big win. It was about a mistake.

In my first two months, I made a critical error. The details matter less than the response: my leaders at the time handled it with grace, empathy, and a forward-looking mentality. We didn’t sweep it under the rug, but we also didn’t spiral into blame. We focused on understanding what went wrong and where systemic improvements could make us stronger.

I told this story for two reasons:

  1. We’re human, and we’re all learning. Especially in R&D on the bleeding edge of innovation: no one has done this before, mistakes will happen.
  2. We tackle challenges together. Forward-looking, constructive, and with an eye toward improving the system, not punishing the individual.

By openly sharing my own mistake, I wanted new team members to know that psychological safety here is real. You don’t have to hide mistakes (and you really shouldn’t). You can admit them, learn, and grow, and in fact, that’s how we become resilient together.

That’s what “modeling humanity” means. It’s not oversharing. It’s showing that being human is allowed, and that’s how teams actually thrive.

Augmenting your humanity with AI

Now imagine if AI could help leaders do this more often and more effectively.

  • Onboarding Story Builder: Capture and distill real mistakes into stories of learning and resilience that new hires hear on day one.
  • Conversation Practice Lab: Roleplay tough moments (admitting uncertainty, apologizing, giving feedback) so leaders and teams build muscle memory for psychological safety.
  • Reflection Nudges: Tools like Microsoft 365 Copilot could surface subtle signals from our own communications, when humility or openness is showing up and when it’s missing, giving us all a chance to do better.

AI can help us practice, remember, and retell the moments that remind teams that humanity is not a weakness.

Question for you: If you could AI-ify one thing about modeling humanity as a leader, what would you want AI to support? Join the conversation on LinkedIn.

Illustration of a desk with a figure from a recent paper about generative AI and emergent phenomena

Can generative AI predict emergent phenomena?

The PNAS Perspective by Tiwary et al. takes on one of the hardest open questions in modeling-driven science: can generative AI predict emergent phenomena?

The authors trace a careful path through the foundations of both computational chemistry and generative modeling, bridging statistical mechanics concepts like force fields and free energy landscapes with architectures including autoencoders (AEs), generative adversarial networks (GANs), flow-based diffusion models, and large language models (LLMs).

Their central argument deserves attention: models capable of predicting emergence must embed physical laws, not merely fit datasets. Statistical mechanics, thermodynamics, and quantum constraints aren’t optional. The bright spots in the field are already moving this way. Reinforcement learning grounded in the principle of maximum caliber, diffusion models inspired by nonequilibrium thermodynamics, and hybrid frameworks like AlphaFlow and AF2RAVE all point toward a new synthesis: physics as foundation, generative AI as engine.

Yet the conditional structure of biological and chemical systems sets hard limits. Most training sets collapse critical variables (temperature, solvent composition, ionic strength, and conformational heterogeneity) into latent noise. Without explicit conditioning, models risk conflating context-dependent behavior with sequence- or structure-intrinsic features. What the field needs next are frameworks that make those assumptions explicit: guidance on when each class of model is appropriate, how to diagnose failure, and how to measure progress beyond visual plausibility or interpolation accuracy.

Leadership angle

For those leading or investing in AI-driven science, the message is clear: the next leap won’t come from larger models alone, but from tighter coupling between representation and reality. The teams that will lead this next wave are those fluent in both the language of data AND the laws that govern it.

Mentorship angle

For early-career scientists, this is an invitation to think rigorously about foundations.

Learn the physics as well as the Python.

Understand how bias enters your data and what it does to inference. The next breakthroughs won’t come from models that memorize reality, but from those that explain it, and can then predict new emergent phenomena.

Page 2 of 7

Powered by WordPress & Theme by Anders Norén