Leadership in Biotech

Author: Kristin Gleitsman Page 1 of 7

Scientist. Research and Development Leader. Collaborative Problem Solver.
Passionate about building belonging, so that together we can accomplish important, hard things.

Illustration of a desk with a figure from a recent paper about GeneOntology

The Gene Ontology Knowledgebase in 2026

This year’s Gene Ontology (GO) update is a reminder that infrastructure choices shape scientific conclusions, getting to the heart of this foundational tool for understanding biology at a time when omics, enrichment analyses, and AI models increasingly rely on GO as biological “ground truth.”

Are you new to Gene Ontology? See the PDF for a deeper dive.

What actually changed (2022–2025)

A few highlights that matter in practice:

  • Major ontology cleanup: hundreds of new terms added, thousands of imprecise or redundant terms obsoleted.
  • Human Functionome v2.0: a reviewed, integrated annotation set now covering ~84% of human genes, reducing enrichment clutter while preserving biological relevance.
  • GO-CAMs scaled up: >1,500 expert-curated causal pathway models linking gene activities with evidence, moving beyond flat gene lists toward mechanistic flow.

Why this paper matters for diagnostics, AI, and innovation leaders
GO and AI models share something important: both are compressions of complex biology.

  • Gene Ontology is a structured compression of current biological knowledge, but lack explicit biological context.
  • Genomic language models are statistical compressions of high-dimensional data, but lack biological grounding.
  • GO provides a curated prior (a biological sanity check) but it abstracts away context (cell state, disease, rewiring). In cancer, that context is often the signal. Used well, GO disciplines thinking and prevents nonsense. Used naively, it produces answers that look rigorous, but are nonsensical.
  • The opposite risk exists with genomic language models. They learn dense embeddings that can capture patterns not explicitly labeled, but the derived “understanding” is not mechanistic by default; it’s statistical compression. And they can overindex on historical data distributions, which can amplify biases.

An intriguing option (and a common one in many recent AI Bio papers), is to combine Gene Ontology with language models.

For example:
Language model → propose;
Gene Ontology → check.

Use a language model to propose functional/interaction hypotheses from data & Gene Ontology to flag things like contradictions and flag known process vs believable novelty vs likely nonsense.

A note to early-career scientists

Impact doesn’t only come from novelty. It comes from:

  • caring about definitions and evidence,
  • understanding the assumptions baked into your tools,
  • and knowing where structure helps, and where it hides uncertainty.

Bridge discovery with discipline, and insight with infrastructure, and you’ll do work that lasts.

Skip to PDF content

Books That Inspire Me and Keep Me Honest

It’s Christmas Eve today, one of those rare pauses in the year that naturally invites reflection.

The inbox quiets, urgency loosens its grip, and there’s just enough space to notice what’s been carrying you and what hasn’t been worth carrying at all.

This part of my bookshelf came together slowly, but it centers around a single idea I’ve had to learn more than once:

Leadership (and life) gets better when you stop pretending time, energy, and attention are infinite and start choosing deliberately.

These books are about accepting finitude without shrinking ambition. They reject the lie that better leaders simply optimize harder, move faster, or squeeze more into already-full lives.

Make no mistake: these books are not about doing less. They’re about choosing with presence and leading in a way that is both humane and high-leverage.

They also ask an uncomfortable and liberating question (to borrow from Mary Oliver):
What is it you plan to do with your one wild and precious life?

Across different angles and disciplines, they converge on the same truths:

  • Time is not something to be mastered.
  • Attention is not neutral.
  • Busy is a signal, not a status symbol.

Learning how to discern and choose deliberately is a core discipline of leadership and of a life well lived.

When you stop pretending there’s time for everything, you’re forced to confront what actually matters. From that place, you can become more intentional about where you show up fully, where you say no cleanly, and where you let go without guilt.

These books helped me loosen my grip on the fantasy that better leaders simply move faster or carry more. They reframed leadership as a practice of discernment: noticing what deserves care, imagination, and energy and what does not. After all, limits are creative pressure if you frame them correctly. Small teams and startups often have outsized impact not just because of agility, but because of the focus and discipline they’re forced to adopt.

They also remind me to zoom out from my sometimes maniacal focus on work and achievement. While work is absolutely a central source of meaning in my life, meaning also comes from my family, the relationships I choose to nurture, and the parts of my identity that exist independent of any job title.

So in this reflective time of year, I hope we can all step back and embrace:

  • the permission to slow without disengaging
  • the relief that ambition doesn’t require self-erasure
  • the clarity that judgment beats busyness
  • the validation that presence is not weakness, it’s leadership maturity

Kristin’s Bookshelf: Working with Humans

These books didn’t emerge from a single chapter of my career, but they all circle the same truth: working with humans starts with upgrading yourself. Not through performance, but through honesty, emotional range, and perspective.

The Stone & Heen books (Thanks for the Feedback and Difficult Conversations) came into my life during a period where I was learning — repeatedly and sometimes painfully — that I can’t control all that much except how I relate to a situation. That shift from “I don’t like this, it shouldn’t be happening” to “what’s actually going on here, and what are my options?” was one of the hardest pivots I’ve ever made.

These books helped me step back from the visceral reaction and see more of the chessboard so I could choose my next move with intention, not just instinct.

A little later came the reckoning with power dynamics…

For most of my career, I’ve had a deep allergy to organizational politics — especially the kind where polish is valued more than substance. My whole professional identity has been built around getting sh*t done. First alone, then with a team, then through leaders I developed. Power, by contrast, felt performative, slippery, and vaguely repulsive.

That’s why 7 Rules of Power is a book I love to hate and hate to love, depending on the day. There’s an uncomfortable honesty in it: the kind you want to look away from but can’t. The rules aren’t aspirational, but they ring true. And learning to navigate the reality of how power works, without losing sight of the kind of culture I want to build, is probably a practice I’ll be refining for the rest of my career.

Which brings me to Unlocking Leadership Mindtraps

This book reminds me how easily we flatten reality into false binaries — right/wrong, good/bad, smart/stupid — especially under stress. Humans crave simplicity, but leadership rarely grants it. We live in a messy, colorful world where multiple things can be true at once. The more I can hold that complexity, the more effective (and compassionate) I can be with others.

So while I put these books under the category of “Working with Humans,” they’re really about something deeper: how to work with yourself so you can work better with other messy, complicated, brilliant humans, most of whom you have very little control over.

These books helped me widen my lens, soften my reactivity, and engage the reality in front of me rather than the one I wish existed. And from that place, the work becomes more honest, more humane, and more impactful.

Illustration of a desk with a figure from a recent paper about AUTOENCODIX

Interpretable Latent Spaces, Messy Biology: What AUTOENCODIX Teaches Us About Autoencoders in the Wild

This week in AI ∩ Bio I dug into AUTOENCODIX, an open-source framework that stress-tests autoencoders (AEs) on real multi-omics data.

The punchline: no single architecture wins, reconstruction scores can mislead, and “interpretable” latent spaces inherit every bias baked into our ontologies. This paper provides exactly the kind of clarity we need to effectively apply these models in diagnostics and biomarker discovery.

AUTOENCODIX is a new open-source framework that tries to bring order to the AEs chaos in multi-omics, allowing the user to test multiple AEs through the same pipeline, then compare not just loss curves but how useful the learned embeddings actually are for biology and prognosis. (AEs explained in carousel)

Scientifically, a few themes stood out:

  1. they show how tuning β in VAEs affects performance; low β favors reconstruction; high β imposes compact, disentangled latent spaces.
  2. across TCGA and single-cell cortex data, no AE architecture consistently outperforms others. Good reconstruction doesn’t guarantee useful embeddings. Ontix, the biologically structured AE, wires decoder layers to known pathways or chromosomes, making latent dimensions interpretable. But robustness varies and depends on learning rate; and the results hint at artifactual learning (see comments).

Diagnostics-leadership perspective

This paper is a reminder to separate infrastructure from insight.

AUTOENCODIX is essentially AE infrastructure: it standardizes data handling, model training, and evaluation so you can ask disciplined questions instead of chasing whichever architecture is trending.

The results also challenge the reflex to equate fancier models with better clinical value: PCA remains a very strong baseline, and ontology-based models only shine when the chosen ontology matches the question and is treated carefully as a potential source of bias, not ground truth.

For leaders deciding where to invest, the take-home is: fund frameworks that make comparisons fair and reproducible, and judge models by task-relevant endpoints and robustness across cohorts—not by reconstruction loss or aesthetic latent plots.

For early-career scientists

There’s a quieter lesson here about how to work with powerful tools without giving up your scientific spine.

The authors don’t present a magical autoencoder that “solves” multi-omics; instead, they map trade-offs, show when tuning helps and when it doesn’t, and surface uncomfortable findings like decreased robustness after hyperparameter optimization for ontology-based VAEs.

If you’re building a career in computational or experimental biology, papers like this are an invitation to open the hood: run the benchmarks, break the assumptions, test models on tasks you actually care about, and treat interpretability as something you design and stress-test—not something you assume.

Skip to PDF content
Photo of "Managing for Happiness", "Where the Action Is" and "Rituals Roadmap" books sitting on a wooden table

“Ways of Working” Books

What “Ways of Working” Means in My Leadership

These books came into my leadership journey at moments when I was learning something essential: culture isn’t built in the big offsites or the annual strategy decks, it’s built in the patterns of daily work.

And the work of leadership gets substantially easier when you design practices that empower your team, reduce unnecessary friction, and help people feel enrolled in the work. This is especially true in scientific environments, where independence of thought, intellectual rigor, and high-stakes decision-making are core to advancing the mission.

Managing for Happiness

Managing for Happiness, by Jurgen Appelo, came to me at a point in my career when I needed to scale myself. My team was growing, the work was accelerating, and I could no longer be the gravitational center for every decision or conversation. This book gave me pragmatic ways to build systems where people could thrive without feeling like I needed to be “in” every decision, and where autonomy, clarity, and shared ownership weren’t ideals but defaults. It reframed “team motivation” not as cheerleading, but as conditions design.

Where the Action Is

Where the Action Is, by J. Elise Keith, became indispensable during one of the most demanding leadership moments of my career: leading a mission-critical core team with a newly acquired company. I knew I needed to up my game. Meetings were no longer routine touchpoints; they were the operating system for how two cultures, two workflows, and two sets of expectations would come together to deliver.

This book changed everything about how I prepare, facilitate, and follow up. It helped me redesign meetings as leverage points with clear decision rights, deliberate sequencing, and the connective tissue needed to prevent swirl, rework, or misinterpretation.

It also gave me a new diagnostic lens: the realization that what looks like interpersonal tension is often just poor meeting design, unclear decision flow, or broken communication architecture.

Rituals Roadmap

Rituals Roadmap, by Erica Keswin, gave me permission to lead with more humanity. It reminded me that rituals aren’t fluffy or indulgent, they’re structure for connection in high-intensity environments. Small, repeatable practices shape belonging, meaning, and team identity far more than grand gestures ever will.

These ideas became especially powerful as I scaled Discovery across locations and new hires. We were spread across time zones, functions, and scientific domains, yet rituals helped us feel like one team.

These books changed how I think about the everyday work of leadership. They taught me that ways of working are not trivial operational details. They are culture, they are strategy, and they are the invisible architecture that determines whether people feel connected or isolated, aligned or confused, energized or depleted.

And when designed with intention, they make the whole system feel more human, more coherent, and more capable of doing truly great work.

Illustration of a desk with a figure from a recent paper about popEVE

From Evolutionary Intolerance to Clinical Insight: What popEVE Teaches Us About Missense Variants

This week in AI ∩ Bio, we look at “Proteome-wide model for human disease genetics”, which introduces popEVE, an unsupervised model for prioritizing missense variants across the human proteome.

The key idea: give each amino-acid change a calibrated severity score based on how evolutionarily and statistically “intolerant” it looks, without pretending to directly predict pathogenicity. But even though the model is not trained on pathogenicity, the scores it produces are correlated with pathogenicity, and often clustered in biologically plausible regions.

Scientific Insight

popEVE combines two sequence-based models (EVE and ESM-1v) with real-world human variation from ~460,000 genomes (UK Biobank + gnomAD).

EVE and ESM-1v together give a composite view of how “acceptable” a variant is:

  • EVE is a variational autoencoder trained on multiple sequence alignments (MSAs) across species. It captures deep evolutionary conservation
  • ESM is a large-scale unsupervised transformer model trained on unlabeled protein sequences. It captures local structural and biophysical coherence

They are combined (with some weighting) into a raw deleteriousness score per variant.

Then, they look at Human Variant Depletion: Presence/Absence in gnomAD/UKBB, which provides empirical evidence of human constraint.

For each protein position, they ask:
“Across 460,000 people, how many unique missense variants were observed here?”

If a position has:

  • Many unique variants → assumed tolerant in humans
  • Few or no variants → suggests purifying selection, i.e., likely deleterious if mutated

They feed this site-level variant density into a Gaussian Process calibration layer, which:

  1. Adjusts the EVE+ESM score to reflect gene-specific constraint
  2. Produces a calibrated score that is comparable across genes

The resulting score correlates with disease severity and clusters in known functional domains and interfaces, and in singleton diagnostic cases, ~80% of truly causal variants land in the model’s top 10 candidates.

Importantly, the authors are clear about limits: no non-coding variants, no epistasis or tissue context, and potential confounding from sequencing and coverage artifacts.

Leadership Angle

For diagnostics leaders, popEVE is a tool for smarter triage, not a final verdict.

It offers a transparent, cross-gene notion of constraint that can sharpen variant review in rare disease and carrier screening workflows, if we remember exactly what it measures: statistical intolerance, not clinical truth.

Mentorship Angle

For early-career scientists, this paper demonstrates disciplined scope and thoughtful integration.

The authors explicitly state what popEVE does not do, then show where its scores line up with known biology and clinical patterns.

If you’re building a career in biotech, this is the sweet spot: models that are powerful because you understand the biology, the data generation, and the limits of what any score can tell you about a real patient.

Skip to PDF content
Photo of The Design Thinking Toolbox and The Design Thinking Playbook on a wooden table, next to two small houseplants

From Sticky Notes to Systems: What Design Thinking Taught Me

Why I Bought These Books

If I’m remembering correctly, like most of the good books in my life, I first checked The Design Thinking Toolbox out of the library. I didn’t necessarily expect much, mostly just a quick skim for ideas. But I found myself bookmarking pages, jotting down thoughts on sticky notes, and eventually admitting the obvious: I needed to own it. And if I was going to own it, I might as well get The Design Thinking Playbook too.

These books became part of my early days building the Discovery function at Veracyte. One of the first clear priorities I was given was to design an “Innovation Day” — which I immediately rebranded as Discovery Day. “Innovation” belongs to everyone, and I didn’t want a single team cornering that territory. Discovery, on the other hand, is an invitation.

The event eventually became an annual gathering of leaders from across our globally distributed organization: a chance to explore the trends shaping the future of diagnostics and translate those insights into Discovery priorities for the year ahead.

But in those first few weeks, all I knew was this: If people were going to fly across the world to sit in a room together, we couldn’t have them passively listen. The day had to be interactive. Engaging. Alive.

So I turned to design thinking for inspiration, and these books delivered.

They gave me concrete ideas for workshops, facilitation frameworks, and hands-on ways of bringing people into the conversation. They reminded me how to separate divergent and convergent thinking so we didn’t collapse creativity before it could breathe. They helped me draw out voices from across functions and levels, and design an experience that felt energizing rather than performative.

And the influence didn’t stop at Discovery Day.

What Design Thinking Means in My Leadership

Design thinking has become one of the engines of how I lead.

At its core, design thinking gives me a disciplined way to bring people together around complex problems. Not to perform collaboration, but to practice it. It pushes me to create rooms where people feel invited in, where ideas can breathe, and where we resist the urge to collapse creativity too quickly.

It’s also a structural reminder: explore before you evaluate; diverge before you converge; listen before you decide. That sequence has shaped how I architect conversations, how I build alignment across functions, and how I design cultures that don’t default to the loudest voice in the room.

It’s not about sticky notes.

It’s about designing practices and systems that surface insight, reduce fear, and make it easier for people to see what’s possible together.

Illustration of a desk with a figure from a recent paper about Platinum Pedigree

AI Models Need Better Truth—Platinum Pedigree Shows How

When I was in high school, I was obsessed with genetics. The Human Genome Project was in full swing, and it felt like the future was being written in real time. I told a family friend I wanted to become a geneticist. He smiled and said, “My cousin is at the NIH. They’ll finish the human genome before you finish college, so I wouldn’t bother.”

The project wrapped in 2003. But papers like this remind me how wrong that prediction was. Even after “finishing” the genome, we’re still uncovering what accuracy, completeness, and truth really mean.

The new Platinum Pedigree study pushes that frontier again.

Scientific Insight

This work builds one of the most comprehensive germline variant benchmarks to date, deep long-read sequencing across a 10-member family, combined with Mendelian logic.

By integrating PacBio HiFi, Oxford Nanopore Technologies, and Illumina and testing every variant against inheritance patterns, the authors defined 2.77 Gb of high-confidence genome (~200 Mb beyond prior benchmarks), including repeats, segmental duplications, and low-mappability regions.

The key innovation is biological grounding.

Each child inherits one haplotype from each parent; variants that obey those segregation patterns are kept, and those that don’t are removed. This yielded ~4.7M SNVs, 768k indels, 537k tandem repeats, and 24k structural variants as pedigree-consistent truth.

When DeepVariant was retrained on this truth set, error rates dropped by ~34% across challenging classes, especially indels and tandem repeats.

Better labels → better models.

Leadership Angle

For diagnostics leaders, this signals where the field is heading: stronger evidence standards, clearer definitions of “truth,” and biologically informed benchmarks rather than technology-constrained heuristics.

This strategy of combining multiple sequencing technologies and adjudicating discrepancies with inheritance is exactly how robust systems are built in uncertain environments.

It mirrors what clinical diagnostics now requires: pipelines that perform not just in easy regions, but in messy, clinically meaningful ones.

And it underscores a central lesson in AI-enabled diagnostics: your model is only as good as the ground truth you train it on.

The regions that are currently messy and difficult to map: that’s where new breakthroughs in understanding will occur.

Mentorship Angle

For early-career scientists, the lesson is craftsmanship. This paper doesn’t debut a flashy algorithm; it elevates the foundations. It asks simple but profound questions: Did this variant follow the rules of inheritance? If not, are we sure it’s real?

Your technical tools matter, but your willingness to interrogate assumptions matters more. If you want to build a meaningful career in genetics in this age of AI, stay curious about the scaffolding beneath the science.

Breakthroughs often start there.

Skip to PDF content
Image of a stack of books sitting on a wooden table in a living room.

Kristin’s Bookshelf

My husband has been nudging me for months to expand the book recommendations on my website. So this Saturday afternoon, I finally pulled apart the bookshelf in my bedroom and dusted off a few of my favorites.

It was oddly comforting to revisit these old friends, each one a marker of who I’ve become as a leader. Every book holds a lesson I had to learn the hard way: a moment when the situation demanded that I grow, shift, or rethink how I was showing up.

Looking at them together, I’m reminded that the core elements of leadership — culture, behavior, systems, and meaning — aren’t separate domains. They weave together. They shape how we decide, how we relate, and how we build environments where people can do their best work.

Still, because humans love categories, I sorted this shelf into five sections:

  1. Design Thinking
  2. Ways of Working
  3. Working Better With Humans
  4. Strategy & Systems
  5. Books That Inspire Me and Keep Me Honest

If you look closely, you’ll see the edges of dozens of sticky notes sticking out at odd angles. I’ve really lived with these books. But I don’t just read them once and walk away, I return to them. I re-read to remember, to recalibrate, and to trace the leadership muscles I’ve built and the ones I’m still strengthening. These books help me remember who I am, and who I want to be, especially in the moments that test both.

So I’d love for you to join me as I walk through my leadership bookshelf over the next few weeks — not as a list of recommendations, but as a guided tour of the ideas and practices that shaped me. My hope is that we come away with a fuller, more human view of what it means to lead, and how to tap into the best of our humanity.

Goodness knows we need it.

Illustration of a desk with a figure from a recent paper about BIOREASON

From Prediction to Explanation: How BIOREASON Reframes Genomic AI as a Reasoning Problem

While DNA foundation models like Evo2 and Nucleotide Transformer can encode genomic sequences into dense, information-rich embeddings, they still operate as black boxes—excellent at prediction, poor at explaining why. Large language models offer the opposite tradeoff: they excel at generating explanations but treat DNA as unstructured text, without any built-in understanding of motifs, regulatory grammar, or sequence constraints.

What BIOREASON Does

BIOREASON introduces a multimodal architecture that fuses:

  • A frozen DNA foundation model to encode biological sequence semantics
  • A fine-tuned LLM that ingests both the embeddings and natural-language context

This pairing enables:

  • Natural-language reasoning grounded (at least in theory) in genomic content
  • Generation of interpretable, mechanistic chains (variant → pathway → phenotype)
  • Improved predictive performance relative to either the DNA FM or LLM alone

But what do DNA embeddings “mean”?

Short answer: we don’t know—and that uncertainty is inherent to foundation models.

  • These embeddings are latent representations learned through massive unsupervised training.
  • They’re presumed to encode motifs, conservation, splicing signals, or regulatory cues because the model needed those features to solve its training task.
  • They are not human-interpretable.

BIOREASON treats these embeddings as a kind of “biological fingerprint,” trusting that an LLM can learn to reason over them with enough supervised examples. The gamble is that:

  • The DNA model has learned useful biological grammar
  • The LLM can exploit those learned signals to answer new questions

But there’s no explicit decoding or truth-checking of what the embeddings represent internally.

Why reasoning faithfulness still isn’t guaranteed

The <think> traces produced by BIOREASON are not probabilistic, validated, or causally guaranteed. The model “believes” its chain, but you must judge its soundness.

Anthropic’s “Reasoning Models Don’t Always Say What They Think” (Chen et al., 2025) shows why this matters: reasoning-tuned models often rely on subtle internal shortcuts, then fail to verbalize them, generating fluent but misleading explanations. BIOREASON inherits the same risk.

Key concerns:

  • Explainability ≠ faithfulness
    A coherent chain does not mean the model followed that chain internally.
  • Potential post-hoc rationalization
    The model may rely on correlations or dataset artifacts, then wrap them in a plausible narrative.
  • Compromised auditability
    If the chain isn’t faithful, transparency becomes performative rather than informative.
  • Hidden biases or shortcut features
    The model might use annotation frequency, ClinVar priors, or pathway prevalence without ever stating so.
  • Lack of mechanistic grounding
    True mechanistic understanding would require identifying which embedding dimensions or sequence contexts drove the decision. The <think> chain alone cannot provide this.

Let’s step back: What’s genuinely novel here?

Despite its limitations, BIOREASON introduces several meaningful advances for the field.

Fusion of Biological Foundation Models with Language Reasoning

Traditional models split into two camps:

  • Models that understand sequence biology (Enformer, Evo2, Nucleotide Transformer)
  • Models that generate explanations (GPT-style LLMs)

BIOREASON bridges these worlds:

  • Anchors reasoning in sequence-aware embeddings
  • Trains the LLM to produce structured, biologically grounded explanations

Why this matters

It reframes variant interpretation as causal narrative inference—a closer match to how human scientists reason.

Structured Explainability via <think> Tokens

Most genomics tools output scores or saliency maps. BIOREASON outputs reasoning.

  • <think> traces formalize a stepwise, human-auditable chain
  • Explanation becomes part of the training objective, not a reverse-engineered artifact

Why this matters

This is one of the first genomics models to explicitly train for mechanistic-style explanation.

A Real Multimodal Interface for Genomics

Multimodal architectures (image+text, audio+text) are flourishing, but genomics has lagged.

BIOREASON shows:

  • DNA sequences can be treated as semantic inputs
  • LLMs can generate biologically coherent outputs when grounded in embeddings

Why this matters

It opens the door to models that integrate DNA, RNA, protein, expression, and literature signals—moving us nearer to true AI lab partners.

Raises Critical Questions About Faithfulness

By making reasoning visible, BIOREASON forces the field to confront fundamental issues:

  • What does it mean for a model to “understand” a variant?
  • How do we measure explanation fidelity, not just fluency?
  • How can we prove the model’s logic is driven by sequence rather than language priors?

Why this matters

These questions will shape the evaluation standards for biological AI over the next decade.

Final Thought

BIOREASON’s contribution isn’t that it solves variant interpretation. It’s that it reframes the problem as reasoning, not classification. It pushes us closer to models that narrate mechanistic hypotheses—but it also reminds us why faithfulness, causal testing, and biological grounding matter just as much as model performance.

With stronger embeddings, uncertainty calibration, perturbation tests, and wet-lab validation, this line of work could become a cornerstone of how AI collaborates with scientists in the years ahead.

Page 1 of 7

Powered by WordPress & Theme by Anders Norén