How an embryo draws the line between body and germline

The molecules that form cell membranes help establish a boundary between cell types early in development which is essential for an animal's ability to produce offspring.

Mackenzie White | Whitehead Institute
September 7, 2026

Every embryo of an organism that reproduces sexually faces an early and consequential decision: which of its cell lineages will make up the mortal body, and which will become the germline, the lineage that eventually makes eggs or sperm and carries genetic information to the next generation. Get this wrong, and an organism can live but cannot reproduce and, evolutionarily, ceases to exist.

Like most embryos, fruit fly embryos begin as one mega-cell, the fertilized egg cell, full of the ingredients for future development. What’s different in fly embryos is that not all cells form at the same time. Before most of the embryo has been subdivided into individual cells, a small group of cells that will eventually produce eggs or sperm has already begun to form. These primordial germ cells emerge at the back end of the embryo, where small buds rise from its surface and pinch free from the shared interior.

A new study from the lab of Whitehead Institute Director Ruth Lehmann reveals that this process depends on preparations made before the buds appear. The researchers found that a protein, aptly named Germ Cell-less (GCL), produced at the posterior, organizes a distinct region of the embryo’s membrane so that the molecular machinery needed to separate the future germ cells can assemble.

The findings, published on July 22 in the Journal of Cell Biology, show how the local organization of lipids, the molecules that form cell membranes, helps to establish a boundary between the germline and the soma or body cells. This separation is essential for the animal’s ability to produce offspring.

How that separation begins depends on location. As the embryo’s nuclei travel outward to its surface, only those that reach the posterior go on to become germ cells, each sealing into a membrane and pinching away from the rest of the embryo as one of its first true cells. The germplasm, a specialized material the mother deposits at the posterior end, carries the factors that set these cells apart.

Previous work from the Lehmann lab showed that GCL is essential for this pinching-off. GCL directs the cell’s protein-disposal machinery to recognize and destroy Torso, a receptor that transmits signals from the embryo’s surface. Torso is best known for its role in the soma, where it triggers a signaling cascade switching on genes needed for head and tail development. Without GCL, Torso remains active at the posterior end, and germ cells usually fail to form. What Torso was doing to interfere with the process, however, remained unclear, since none of the hallmarks of its gene-activating pathways seemed involved.

To trace the signal downstream of Torso, the team, including lead author Mariyah Saiduddin, a graduate student in the Lehmann lab, used genetic experiments, live imaging, and optogenetics (a method that uses light to switch protein activity on or off with precise timing and spatial control). The results revealed a previously unknown and surprising role for Torso in early development: activating an enzyme called PI3K. That enzyme produces PIP3, a signaling lipid embedded in the embryo’s membrane.

Collaborating with Whitehead’s W.M. Keck Microscopy Innovation Center, the researchers developed image-analysis approaches to measure subtle changes in PIP3 signals in living embryos in real time.

In a normal embryo, PIP3 was abundant near the posterior end but excluded from the region where germ cells form. Without GCL, PIP3 spread into that region. Experimentally boosting PI3K activity sharply reduced germ cell formation, while dialing it down allowed extra cells to form. The lipid acted like a chemical switch: with PIP3 levels reduced by GCL’s action, Myosin II, a motor protein that generates contractile force, assembled into a ring-like structure to pinch each bud coming off of the embryo membrane at its base, separating each developing germ cell from the rest of the embryo. With too much PIP3 on the membrane, the Myosin structure was unable to assemble, the buds flattened, and germ cells failed to form.

“What was surprising was that the membrane was already being prepared before the nuclei reached it,” says Saiduddin. “We had thought of germ cell formation as beginning when the nuclei arrived at the surface, but GCL is acting earlier, preventing PIP3 from building up in advance.”

This study shows that the boundary between soma and germline isn’t drawn by turning genes on or off in the early embryo, but depends in part on locally shaping the chemistry of the cell membranes. In Drosophila, the antagonistic relationship between Torso and its destroyer, GCL, determines where in the membrane signaling molecules are positioned before the embryo even starts developing. This in turn provides constraints on where to exert mechanical force to round up and bud into a new cell, thereby helping to specify germ cell fate.

Because the signaling pathways involved here are conserved or present across the animal kingdom, the work offers broader clues about how local signaling cascades can organize membrane lipids and shape cell behavior and fate, beyond fruit flies.

The work also raises new questions about why the embryo uses more than one mechanism to protect germ cells from signals that would otherwise push them toward a somatic fate.

“This is a beautiful illustration of discovery science,” Lehmann says. “We solve one puzzle only to realize that there is so much more we do not understand.”

Paper: Mariyah Saiduddin, Juhee Pae, Asier M. Vidal, Martin L. Alani, Ruth Lehmann; GCL pruning of PIP3 establishes the soma-germline boundary. J Cell Biol 7 September 2026; 225 (9): e202604036. doi: https://doi.org/10.1083/jcb.202604036

Lung cancers can use two different mechanisms to evade KRAS-inhibiting drugs

Some tumor cells develop resistance by amplifying the KRAS gene, while others change their tumor type, MIT researchers have found.

Anne Trafton | MIT News
September 30, 2026

About 25 percent of lung adenocarcinomas have mutations of the gene KRAS, which drives uncontrolled cell growth. In recent years, the FDA has approved two KRAS inhibitors to treat patients with KRAS mutations. While these drugs can work well initially, tumors almost always develop resistance to them.

Usually, resistance emerges because cells reactivate KRAS activity, through mutations that prevent drug binding or by increasing KRAS expression that overpowers the effects of the inhibitor. However, in a new study, MIT researchers have modeled an alternative mechanism that cancer cells can use to become resistant to KRAS inhibition.

The researchers found that in some cases, lung tumors undergo transformation from adenocarcinoma to squamous cell carcinoma. Both of these tumor types are commonly found in the lungs, but they are thought to  arise from different cells and have different genetic profiles.

When this transition occurs, tumor cells no longer require KRAS, and appear to turn on alternative signaling pathways that help them continue to grow. Ongoing work to identify those pathways may reveal targets for new drugs that could help prevent resistance to KRAS inhibitors.

“The main takeaway is that there seem to be different routes of resistance to KRAS inhibitors, and so we need to be thinking about how we can address this,” says Carrie Rodriguez, an MIT graduate student and one of the lead authors of the paper.

Nicolas Mathey-Andrews PhD ’25 is also a lead author of the study, which appears today in the journal Nature Genetics. The paper’s senior author is Tyler Jacks, the David H. Koch Professor of Biology and a member of MIT’s Koch Institute for Integrative Cancer Research.

Tissue transformation

The two FDA-approved KRAS inhibitors both target a mutation called KRAS-G12C. These drugs are approved only for use in patients whose tumors have failed to respond to other drugs, and these patients usually have cancer that has spread beyond the lungs.

KRAS inhibitors are effective in about 35 percent of the patients who receive them. However, in those cases, the tumors almost always end up becoming resistant by generating additional copies of the KRAS gene or finding other ways to turn on the MAP kinase signaling pathway, which is usually triggered by KRAS and stimulates cell growth.

“Resistance to targeted therapies is a very serious problem,” Rodriguez says. “Sometimes these KRAS inhibitors can hold cancers at bay, but most cases do end up relapsing.”

A 2021 study from researchers at Dana-Farber Cancer Institute, which analyzed tumors from 17 non-small cell lung cancer patients treated with KRAS-G12C inhibition, identified secondary resistance mutations in a majority of patients. In two of these patients, however, the researchers found that tumors transformed from adenocarcinomas to squamous cell carcinomas, but they did not harbor obvious resistance mutations.

Both adenocarcinomas and squamous cell carcinomas are classified as non-small cell lung cancers (NSCLCs), which are the most common type of primary lung cancer. Adenocarcinomas, the most common type of NSCLCs, often originate from the surfactant-producing cells that line the lungs, while squamous cell carcinomas originate in the cells that line the central airways of the lungs.

Mutations of KRAS are found much more frequently in adenocarcinomas than in squamous cell carcinomas

In this study, the researchers set out to model the factors that might drive the transition from adenocarcinomas to squamous cell carcinomas. To do that, they engineered a mouse lung cancer model to express the mutation that is targeted by the FDA-approved KRAS inhibitors.

Following treatment with a KRAS-G12C inhibitor, tumors with genetic loss of Nkx2-1, which normally helps maintain alveolar epithelial identity, were able to undergo adeno-to-squamous transition. Turning on a transcription factor called DeltaNp63, which is overactive in many squamous cell carcinomas, also made this transition more likely. Another transcription factor known as SOX2 also helped stimulate the transition, but this gene could not initiate the transition on its own.

Paths to resistance

Tumors that underwent these tissue transformations did not acquire the mutations that typically boost KRAS expression in adenocarcinomas. Instead, KRAS signaling was shut off. The researchers hypothesize that these cells may turn on another signaling pathway that helps them to continue growing.

“There seem to be several different routes where you can get to squamous transformation, either through loss of lung-lineage-defining transcription factors, or overexpression of these squamous master regulators, SOX2 or DeltaNp63. Those resistant squamous tumors no longer respond to KRAS inhibition because they shut off the signaling or at least dampen it significantly,” Rodriguez says.

The researchers are now further exploring what happens to tumor cells as they transition to a squamous state, in hopes of identifying vulnerabilities that could be targeted with new drugs.

“Fundamentally this is a transition that’s poorly understood, and we were happy to see that we were able to model it,” Mathey-Andrews says. “Future directions that have an eye toward translation will utilize those models to understand the process and conditions by which histologic transformation occurs, and then also nominate potential targets downstream.”

The research was funded, in part, by the Koch Institute Support (core) Grant from the National Cancer Institute, a Ruth Kirschstein National Service Research Award, the National Institute of General Medical Sciences, and the Ludwig Center at MIT.

Biologists identify a cellular pathway that allows colorectal cancer to metastasize

They also found that obesity may put patients at higher risk for activation of the pathway. Drugs that block the pathway may help prevent metastasis.

Anne Trafton | MIT News
September 24, 2026

Most colon cancer deaths are caused by the spread of tumor cells beyond the colon, usually to the liver. In a new study, MIT biologists identified a cellular pathway necessary for colorectal cancer metastasis.

The pathway they identified, controlled by a protein known as YAP1, is normally involved in tissue repair. When activated in cancer cells, it promotes cell proliferation and migration. The researchers also found that a high-fat diet is more likely to turn on this pathway, through the production of fatty molecules called ceramides.

Drugs that block ceramide production could offer a new way to help prevent metastasis in patients diagnosed with colon cancer, the researchers say.

“We’ve found a pathway that we think is druggable. If we shut down the enzymes that make ceramides, tumor cells can’t switch on this regenerative program, and they largely fail to seed metastases in the liver,” says Omer Yilmaz, director of the MIT Stem Cell Initiative, a professor of biology at MIT and a member of MIT’s Koch Institute for Integrative Cancer Research. He is also a gastrointestinal pathologist and director of translational research in pathology at Beth Israel Deaconess Medical Center.

Yilmaz, Nilay Sethi, an associate professor of medicine at Harvard Medical School and Dana Farber Cancer Institute, and Alpaslan Tasdogan, head of the Institute for Tumor Metabolism and a professor in the Department of Dermatology at University Hospital Essen and the German Cancer Consortium (DKTK), are the senior authors of the study, which appears today in Science. MIT postdocs Swagata Goswami, Qiming Zhang, and Abdullah Burak Yildiz are the paper’s lead authors.

A hijacked pathway

In the United States, colon cancer is usually diagnosed at stage 2 or 3 — before the cancer has spread. However, even after successful surgery, up to a third of these patients will relapse with metastatic disease.

While scientists have identified many genetic mutations that drive the development of colon cancer, it’s unknown exactly what prompts them to spread beyond the colon.

“Many studies have looked for a genetic driver of metastasis and come up empty,” Yilmaz says. “There isn’t a defining mutational signature that separates metastatic cells from the primary tumor, which points to metastasis being driven largely by changes in which genes are switched on and off, rather than by new mutations.”

In this study, the researchers sought to identify epigenetic programs that enable colon cancer cells to metastasize. Using tumor organoids from mouse models of several types of colon cancer and from patients with colorectal cancer, they found that metastatic cells shared one key feature: activation of the YAP1 program.

YAP1 is a protein that works with partner factors to switch on genes related to development, stem cell maintenance, and regeneration. In normal tissue, it is active during fetal development, and after injury, to promote healing.

In the gut, that repair response runs through a rare, fetal-like cell type, which normally appears only briefly to rebuild the intestinal lining after damage. YAP1 has been linked to cancer for years, but the new work shows that diet-derived lipids push tumor cells into this specific regenerative state — and that the state itself is what licenses metastasis.

“The regenerative program that we described is generally observed in the gut when there is severe injury or infection and the gut needs to regenerate. We see the tumor cells hijack this program to drive metastatic progression,” Goswami says.

Activation of this set of genes helps cancer cells to break free from the original tumor site and spread to other locations in the body. For colon cancer, the most common site of metastasis is the liver, followed by the lungs.

In mouse studies, the researchers also found that cancer cells in animals fed a high-fat diet turned on YAP1 to a greater extent than mice fed a healthy diet. A high-fat diet, the researchers found, triggers activation of enzymes that produce ceramides, a type of lipid. Ceramides then release the molecular brake that normally keeps YAP1 inactive, allowing it to move into the nucleus and switch on its target genes.

Preventing metastasis

The researchers showed that genetically targeting YAP1, or the genes involved in ceramide production, markedly reduced the spread of colon cancer to the liver in mice.

To determine if YAP1 is also involved in metastasis in humans, the researchers analyzed RNA sequencing data from patients with colorectal cancer. They found that YAP1 was more active in metastatic cancer cells, and that patients with higher body mass index (BMI) showed higher expression of the genes activated by YAP1 than normal-weight patients. Patients with higher levels of those genes also had lower survival rates.

“We don’t think that the YAP1 program is specific to obesity. It’s just that it becomes accentuated in obesity, and that may account for why obesity is known to drive the progression of colorectal cancer,” Yilmaz says.

They now plan to develop drugs that inhibit two of the enzymes involved in ceramide production, DEGS1 and DEGS2, in hopes that such drugs could help prevent colon cancer metastasis.

The researchers caution that the findings do not yet translate into dietary advice for patients who have already been diagnosed, and that any drug targeting ceramide synthesis will have to clear a high bar for selectivity, since these lipids are also essential in healthy tissues.

The research was funded by the National Institutes of Health/National Cancer Institute, the MIT Stem Cell Initiative, a Koch Institute Frontier grant, and the NRW Junior Research Program.

Avoiding a sticky situation: how cells stop messenger RNAs from clumping together

Although cells contain many thousands of mRNA molecules crowded into a tiny space, these molecules typically don't form the clumping aggregates often seen when they're removed from their native context for study. New research from the Jain Lab suggests that evolution has shaped mRNA sequences to reduce unwanted interactions.

Alice McCarthy | Whitehead Institute
September 14, 2026

Messenger RNA, or mRNA, is best known for carrying genetic instructions from DNA to the cellular machinery that makes proteins. But all RNA molecules have another, less appreciated property: they are naturally sticky.

When RNA is removed from cells and studied in the laboratory, the molecules readily interact with one another and can clump together, or aggregate. This creates a biological puzzle. Cells contain many thousands of mRNA molecules crowded into a tiny space, yet their mRNAs do not routinely form the large aggregates that their physical properties would seem to favor. This puzzling behavior is important for cell function: unwanted RNA interactions can prevent mRNAs from being accessible to make proteins, and large RNA aggregates can be toxic to cells.

Now, researchers at Whitehead Institute have uncovered one way cells may have evolved to avoid these sticky situations. Their findings suggest that evolution has shaped mRNA sequences to reduce unwanted interactions with other mRNAs. The work could potentially inform the development of RNA therapeutics, allowing drug designers to learn from evolution when selecting RNA sequences.

The study, led by Whitehead Institute Member Ankur Jain, also an associate professor of biology at MIT, and Marco Todisco, a postdoctoral researcher in his lab, reveals a previously unrecognized constraint on the evolution of genetic sequences: DNA must not only encode functional proteins, but also produce mRNAs with physical properties that help keep them soluble inside the cell.

The researchers’ findings were published in the Proceedings of the National Academy of Sciences on September 14.

Scientists studying purified RNA have known for decades that the molecules can readily aggregate. But if the chemistry of RNA makes these interactions so favorable outside cells, Todisco wondered, why shouldn’t the same thing happen inside them?

To investigate, Todisco and colleagues focused on Escherichia coli, or E. coli, a bacterium whose biology has been extensively studied. Researchers have detailed information about which mRNAs are present in an E. coli cell, how many copies of each are present, and what those molecules look like — making it possible to model the behavior of an entire collection of cellular mRNAs, known as the transcriptome.

Todisco developed computer simulations that tracked individual mRNA molecules and predicted how they would behave at concentrations similar to those found inside a cell. Based on the physical chemistry of RNA alone, the simulations indicated that the molecules should aggregate. Longer mRNAs were particularly prone to joining these clusters.

The team then tested that prediction experimentally. With help from Christalyn Ausler, a research technician in the Jain lab, they purified mRNA from E. coli cells. The authors found that without proteins and other components normally surrounding it in the cell, the purified mRNA readily aggregated. Using RNA sequencing, the researchers identified which mRNAs were enriched in the aggregates and found that their properties closely matched the simulations.

Those findings strengthened the original puzzle: If mRNA has such a strong physical tendency to associate with other RNA molecules, how have cells managed that problem?

The researchers found part of the answer written into the genetic sequences themselves.

DNA carries the instructions for building proteins using three-letter sequences called codons, each of which specifies an amino acid, one of the building blocks of proteins. But the genetic code contains redundancy: most amino acids can be encoded by more than one codon. That means an enormous number of different DNA and mRNA sequences can produce exactly the same protein. Evolution therefore has some freedom in which sequence it uses.

The researchers took advantage of that flexibility to computationally create alternative versions of E. coli mRNAs. They changed the RNA sequences while preserving the proteins they encoded — in effect creating alternative evolutionary paths in which an organism could make the same proteins from different mRNAs.

When the researchers compared these randomized sequences with naturally occurring E. coli mRNAs, they found that the natural sequences were less prone to interacting with one another. Native mRNAs tended to fold onto themselves and to avoid exposing sticky stretches, reducing opportunities to form stable and unwanted interactions with other mRNA molecules.

In other words, stickiness is unavoidable, but evolution appears to have favored not just DNA sequences that make the right proteins, but sequences that produce less sticky mRNAs along the way.

The researchers found evidence that the phenomenon extends beyond bacteria. When they performed sequence analyses on a set of abundant human mRNAs, they found similar signatures: naturally occurring human RNA sequences also showed a reduced potential for aggregation compared with alternative sequences.

The discovery adds another dimension to the familiar “central dogma” of molecular biology, in which information flows from DNA to RNA to protein. Scientists have traditionally thought about genetic sequences largely in terms of their ultimate product: the protein. The new findings suggest that evolution also has to contend with the physical behavior of the mRNA intermediate.

That may not be the whole solution. Although naturally occurring mRNA sequences are less prone to aggregation, purified mRNA still readily clumps together. The fact that widespread aggregation is not normally seen inside healthy cells suggests that proteins or other cellular components may provide an additional layer of protection by buffering inappropriate RNA-RNA interactions. These findings help build a broader picture of how cells may use multiple mechanisms to manage RNA interactions and prevent inappropriate aggregation.

The findings could also have implications for the growing field of mRNA therapeutics. When researchers design synthetic mRNAs for vaccines and other treatments, they can choose among different codons that ultimately produce the same protein. The new work suggests that how the resulting RNA folds and interacts with other RNAs may be another important consideration when selecting those sequences.

“What this gives us is a new lens for thinking about how genomes evolve,” Jain says. “Sequences aren’t evolving just to encode functional proteins; there is an additional constraint to produce mRNAs that are less likely to stick to one another. These findings could also provide guiding principles for designing the sequences used in mRNA vaccines and other RNA therapeutics.”

Notes:

Todisco, M., Ausler, C., and Jain, A. “Maintaining transcriptome solubility constrains mRNA sequence composition.” Proceedings of the National Academy of Sciences (2026). https://doi.org/10.1073/pnas.2622980123.

Funding:

This work is supported by grants from the NIH (R35GM151111), Chan Zuckerberg Initiative (DAF2022-250422), Richard and Susan Smith Family Foundation, and the David and Lucile Packard Foundation. The content is solely the responsibility of the authors and does not necessarily represent the official views of the National Institutes of Health.

AI model decodes the language cells use to communicate

New research from the Li Lab at the Whitehead Institute characterizes the signaling pathways whose patterns of gene activity shape cell's structure and function, enabling cells in an organism to take on specific, specialized roles.

Whitehead Institute
September 4, 2026

As an embryo develops from a small cluster of stem cells, those once “blank slate” cells begin to take on more specialized roles like brain, liver, or muscle cells, and organize themselves into three-dimensional structures such as tissues and organs.

The fate of each cell — what type of specialized cell it will become — depends on which genes are turned on or off in the cell. These patterns of gene activity shape the cell’s structure and function, enabling it to take on a specific role in the body.

But this decision isn’t up to individual cells. They constantly send and receive chemical signals to and from neighboring cells, which help them understand where they are, what stage of development they are in, and what they should become. These messages spread through multi-step sequences called signaling pathways, which translate external signals into specific changes in gene activity in the cell.

For researchers, being able to retrace the sequence of instructions a cell has received would offer a powerful way to understand how tissues develop and how these processes go awry in disease.

But this has been difficult to achieve because scientists have long assumed that the effects of signaling pathways vary widely across cell types, meaning they would need to map each pathway separately in each cell type — an arduous and painstaking process.

Now, in a new study led by Whitehead Institute Member Pulin Li and graduate student Nicholas Hutchins, researchers have discovered that each signaling pathway leaves behind a unique “fingerprint” — a distinctive pattern of gene activity that reflects the particular signals the cell has encountered.

Importantly, these fingerprints are consistent across different cell types for the same signaling pathway, which means that instead of mapping each cell type separately, scientists can reconstruct signaling histories across many cell types using these pathway-specific fingerprints.

This discovery was made possible through a machine learning model called IRIS. This model can detect fingerprints of different signaling pathways and pinpoint which signals a cell received at different stages of development inside an embryo, even for cell types it hasn’t encountered before.

This AI-driven approach marks a major advance over traditional methods, which require researchers to experimentally test every pathway in every possible cell type, and opens up the possibility to comprehensively map the signaling histories of every cell inside a mouse or human embryo at an unprecedented scale.

“Think of voice recognition systems like Siri, which are trained mainly in English, but then use that training to help them recognize other languages,” says Li, who is also an assistant professor of biology at the Massachusetts Institute of Technology (MIT). “This is called transfer learning, and this is why IRIS can work across many different cell types.”

The researchers’ detailed findings, published in the journal Nature Methods on Sept. 8, could accelerate stem cell engineering for regenerative medicine and improve the creation of organoids — miniature, 3-D models that mimic real organs — for studying disease mechanisms and testing new drugs. This is because once researchers learn the pattern of signals that drives a stem cell to become a specific cell type, they can recreate those signals to control the fate of stem cells in a lab or medical setting.

IRIS is a neural network-based model, which, in essence, is an AI-driven system designed to recognize patterns in complex data, similar to how our brains spot patterns. It examines a cell’s overall gene activity and estimates which signaling pathways were likely “on” at specific times during development.

Li and Hutchins trained IRIS on a large experimental dataset that measured how thousands of human embryonic stem cells responded to dozens of combinations of six major signaling pathways at multiple stages of development. This created a comprehensive map, or atlas, of how signaling combinations influence cell behavior.

The team then tested IRIS using single cells from mouse embryos during gastrulation, a stage when cells are rapidly branching into different fates. IRIS could accurately predict when and where specific signaling pathways would activate in cells destined to become part of heart, gut, muscle, and spinal cord tissue. This variation in signaling molecules is what guides cells to form the right structures in the right places.

With this approach, scientists can not only begin to understand the fundamentals of cell-to-cell communication, but also gain a practical roadmap for guiding stem cells into specific, functional cell types in the lab.

When the researchers employed IRIS to identify signals needed to create a cell type critical for lung development, the model predicted that activating a specific signaling pathway would encourage lung-specific development. Experiments in mouse embryos confirmed the model’s prediction. By identifying the precise combinations of signals that drive lung cell development, they can more reliably generate accurate, lab-grown models of lung tissue.

“In these ways, IRIS is helping us decode the language cells use to talk to each other at a much faster rate than we could realistically achieve through experiments,” Hutchins says.

These improved models would allow researchers to study diseases like asthma, lung cancer, and pulmonary fibrosis, in which lung tissue becomes scarred, often without a known cause. They can then use these models to test potential therapies, and ultimately design regenerative treatments that can repair the damaged tissue.

Notes:

Hutchins, N. T., Meziane, M., Lu, C., Mitalipova, M., Fischer, D. S., & Li, P. (2026). Reconstructing signaling histories of single cells via perturbation screens and transfer learning. Nature Methods. https://doi.org/10.1038/s41592-026-03213-8

Research reported in this press release was supported by the National Institutes of Health grant number DP2HD108777 awarded to P.L., which funded 30% of the project’s costs. Additional support was provided by the National Science Foundation Graduate Research Fellowship, MathWorks Graduate Fellowship, Allen Family Philanthropies, and Centurion Foundation. This work was made possible by the Whitehead Institute’s Human Stem Cell Core, in collaboration with Maya Mitalipova.

The content is solely the responsibility of the authors and does not necessarily represent the official views of the National Institutes of Health.

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

Lillian Eden | Department of Biology
August 27, 2026

A protein’s function is determined by its structure, and structure — the way a protein folds — is determined by its sequence of amino acids, the building blocks of proteins.

Many methods for designing novel proteins, including examples that could bind to a disease-causing molecule in our cells, involve a two-step process: The structure comes first, and then a machine-learning framework generates a repertoire of sequences that could potentially adopt that structure.

In nature, many different amino acid sequences can fold into the same structure. At the same time, one amino acid sequence can potentially adopt different structures depending on the protein’s flexibility or a functional trigger. Therefore, when researchers use artificial intelligence to design new proteins, the challenge is to guide AI to “see” that there are many potentially useful answers — that many sequences can adopt the same fold

“For years, the field has measured success by asking whether a model can reproduce the protein sequence that evolution happened to select — our work shows that this isn’t the best metric for protein design,” says Amy E. Keating, Department of Biology head, Jay A. Stein (1968) Professor of Biology, professor of biological engineering, and senior author of a paper recently published in PNAS.

PottsMPNN, a new machine-learning framework developed in the Department of Biology, incorporates the physical principles that govern protein structure and stability, improving sequence generation and the ability to predict how mutations will affect a protein’s stability. In other words, the model has a better understanding of the sequence-energy landscape, meaning the relationship between the identity of each amino acid and the stability of the protein.

Adding this framework to a protein design pipeline will allow researchers to design structurally feasible proteins with sequences that don’t resemble those of any native protein.

“If we’re thinking about a completely novel, designed structure, there would be no native sequence to compare it to,” says graduate student and lead author Foster Birnbaum. “What we actually care about is how likely the generated sequences are to fold into the desired structures, how well the model understands the sequence-energy landscape, and how well it can predict the effect of mutations on the stability of the protein.”

Beyond the noise 

In the same way that AI has recently powered some dramatic social changes, so too has machine learning impacted the pace and breadth of fundamental biological research. Only recently has it become possible to reliably use a computational model to generate a protein structure or sequence. Perhaps the most widely used model today, however, was released in 2022.

“For a field that’s moving as fast as machine learning in biology, that model has not been surpassed — we’ve been trying to understand why that is, and what it is about that model that makes it so useful,” Birnbaum says.

Birnbaum was first interested in strategic applications of something researchers call “noise,” or adding variations to a protein structure during training. Noise decreases the tendency of the model to overly mimic native sequences, increasing the diversity of structures for which it’s able to generate sequences.

PottsMPNN also uses a pairwise distribution to capture interactions between amino acids. The ability to account for the physical interactions between all 20 possible sequence options at a pair of positions in the protein is a key reason that PottsMPNN more accurately models the sequence-energy landscape than other methods.

Finally, Birnbaum says, they introduced sets of evolutionarily related sequences into training the PottsMPNN framework to teach the model how different sequences can adopt the same folded structure.

Birnbaum acknowledges that in trying to shift away from adhering to native sequences, incorporating evolutionary information is, in some ways, still a reliance on them. But PottsMPNN succeeded in demonstrating that as the model depends less and less on native sequences, structural compatibility and energy prediction, including for novel proteins, improve.

Protein design in the age of AI

“Once we can design any protein we want, that enables us to do a potentially scary amount of biological engineering,” Birnbaum says. “It’s a difficult task, but I’m really optimistic about this century’s progress in biology.”

Birnbaum hopes that the model could be further improved and fine-tuned for a specific task, which has in the past led to better predictions, for example, on the outcome or consequence of a particular mutation.

Ultimately, according to Keating, “Our methods move the field toward designing useful new-to-nature proteins for diverse applications while providing a stronger foundation for future advances.”

Cells use a little-known molecule to protect themselves from iron overload

Research from the Jain and Henry Labs in the Whitehead Institute and Koch Institute respectively have discovered that cells are protected from destructive levels of iron buildup by small molecules called polyamines, which act like storage lockers, safely holding the metal in a non-reactive state until cells need it.

Shafaq Zia | Whitehead Institute
August 14, 2026

Iron is essential. Our cells need it to produce energy, carry oxygen throughout the body, and power countless chemical reactions that sustain life. But this metal has a dark side. When too much of it is left free inside cells, it can trigger destructive reactions that break down DNA, proteins, and even cell membranes.

Now, Whitehead Institute Member Ankur Jain and Whitney Henry, an investigator in MIT’s Department of Biology and the Koch Institute for Integrative Cancer Research, together with graduate student Pushkal Sharma have discovered that cells rely on an unexpected protector against this threat: small molecules called polyamines.

The researchers’ detailed findings, published Aug. 14 in the journal Cell, reveal that polyamines act like storage lockers for iron, safely holding the metal in a non-reactive state until cells need it.

These findings solve a decades-old mystery about why cells maintain such extraordinarily high levels of polyamines and uncover a previously unknown defense mechanism that protects cells from toxic iron overload.

This work could also help scientists develop better cancer treatments, by allowing iron overload to trigger cancer cell death. It could also offer new clues about diseases like early-onset Parkinson’s disease, in which mutations affect polyamine levels within neurons.

The Jain Lab studies RNA, the intermediary between DNA and the tiny molecular machines called proteins that perform most of the essential tasks inside cells. The lab is particularly interested in how RNA folds, misfolds, and sometimes clumps inside cells.

Jain and Sharma first began studying polyamines because these molecules bind to RNA and help shape its structure. However, they suspected that polyamines must be playing other roles inside cells: they’re among the most abundant small molecules within cells, present at levels comparable to ATP, the molecule cells use as their energy currency.

“We’ve known that without polyamines, cells stop growing and dividing,” says Jain, who is also an associate professor of biology at the Massachusetts Institute of Technology (MIT). “But their best-known function only requires a small fraction of the polyamine levels cells actually have.”

To uncover polyamines’ hidden function inside cells, the researchers used a large-scale genetic approach that allows them to screen the entire genome at once, rather than testing genes one-by-one, in order to find out which cellular processes are impacted when polyamine levels are changed within cells.

The screen revealed that when cells have reduced levels of polyamines, a protein called GPX4 becomes essential for survival. GPX4 is known to prevent harmful chemical reactions that damage the fatty molecules that make up cell membranes.

The team also found that cells with lower polyamine levels have higher amounts of another protein that acts as an iron sponge and keeps the metal in a mineralized form. Together, these findings led the researchers to hypothesize that polyamines might be helping keep iron in a safe, non-reactive state within cells.

To test this idea, they developed a new fluorescent sensor that would allow them to measure chemically reactive iron inside living cells. The new sensor causes living cells to glow based on the amount of chemically reactive iron they contain, allowing researchers to track any changes under a microscope in real time.

The team paired the new iron sensor with another sensor they had previously developed that measures polyamine levels within cells. By employing them simultaneously, they observed a striking pattern: as polyamine levels dropped within cells, the amount of chemically reactive iron went up, offering new evidence that polyamines play a key role in preventing toxic iron build up inside cells.

Beyond answering a fundamental biological question, these findings could have implications for cancer treatment. Cancer cells often rely on high polyamine levels to support their rapid growth and division. However, cancer drugs designed to lower polyamine levels to stop cell division have had limited success.

“We saw that when polyamine levels fall, cells rely on GPX4 to protect themselves from iron toxicity,” says Sharma, who is also the first author of the study. “This could mean that combining drugs that lower polyamine levels with those that block GPX4 might be more effective for killing cancer cells than targeting either pathway alone.”

The discovery may also have implications beyond cancer. Mutations in genes that help move polyamines around cells are linked to a rare form of early-onset Parkinson’s disease, and scientists have long observed unusually high levels of iron in the brains of Parkinson’s patients.

While it is still unclear whether excess iron directly contributes to neuron death in Parkinson’s, the discovery that polyamines help buffer reactive iron inside cells offers a possible explanation for this link and opens new directions for future investigation.

In addition, the researchers expect the new iron sensor to be a valuable tool for other scientists. By allowing them to track chemically reactive iron inside living cells, it could power new discoveries in aging, cancer, and neurodegeneration.

“There are a lot of promising future directions for this work,” Jain says. “It’s exciting to think about how these tools and findings could help answer further questions about disease pathways and potentially help design better therapies.”

Sharma, P., Keys, H. R., Mansell, R. P., Stark, J., Girard, L., Ausler, C., Anderson, R., Müller, S., Imada, S., Pires, I. S., Kunchok, T., Waite, M., Yuan, B., Deik, A., Ferro, L., Hammond, P. T., Rodriguez, R., Pandelia, M., Henry, W. S., & Jain, A. (2026). Polyamines buffer labile iron to suppress ferroptosis. Cell. https://doi.org/10.1016/j.cell.2026.07.040

Research reported in this press release was supported by the National Institutes of Health under grant number R35GM151111 awarded to A.J., which funded 25% of the project’s costs. Additional support was provided by the Bumpus Foundation and the Pew Charitable Trusts. The content is solely the responsibility of the authors and does not necessarily represent the official views of the National Institutes of Health.

How the Toxoplasma parasite adapts to crowded conditions in host cells

A new study by Whitehead Institute researchers identifies how the widespread parasite Toxoplasma gondii survives in crowded environments in infected cells, which helps it persist in long-lasting brain cysts.

Mackenzie White | Whitehead Institute
August 19, 2026

Toxoplasma gondii, or Toxoplasma, is a parasite that infects hundreds of millions of people around the world. Although cases are often mild, it can cause severe symptoms in in people with weakened immune systems, and in developing fetuses. It can also persist for years by forming long-lived cysts in tissues, allowing infection to become chronic.

During chronic infection, hundreds of Toxoplasma parasites can pack into a tissue cyst inside a brain or muscle cell. That crowded life carries a cost: Nutrients become harder to obtain, waste accumulates, and energy-producing reactions can become damaging.

How Toxoplasma reshapes its metabolism to keep growing under such strained conditions has been unclear. But a new study from the lab of MIT Associate Professor Sebastian Lourido, a member of the Whitehead Institute for Biomedical Research, identifies a parasite-specific protein that helps coordinate this response. The protein, named TgPRO, allows Toxoplasma to manage oxidative stress — the buildup of reactive oxygen molecules that can damage cells — by controlling genes involved in energy production and iron use.

The open-access findings, published on Aug. 11 in the journal Cell, reveal the first dedicated regulator of metabolic gene expression identified in apicomplexans, the group of parasites that includes Toxoplasma and the organisms that cause malaria. The study, led by co-first authors and Lourido lab affiliates Christopher Giuliano PhD ’26, a recent graduate student in biology, and Chinmay Kalluraya, a current graduate student in biology, reveals a previously unknown way that parasites regulate metabolism. The findings also point to a possible therapeutic strategy: Inhibiting pathways controlled by TgPRO could make Toxoplasma more vulnerable to antiparasitic drugs that induce oxidative stress, though this approach remains to be tested.

One gene at a time

To discover the genes that support Toxoplasma’s ability to live in crowded cells, the researchers used a genome-wide CRISPR screen to compare Toxoplasma growing at low and high densities. The screen tests the effects of turning off genes one by one at both population densities in order to determine which genes are essential specifically in crowded conditions. It highlighted pathways that make or recycle NAD and NADP, molecules important for energy production and defending against oxidative damage. It also pointed to TgPRO, a previously unstudied protein that was especially important when parasites became crowded.

“A genome-wide screen was a powerful way to ask how crowding affects parasite fitness,” Kalluraya says. “TgPRO emerged as very important at high density. Because almost nothing was known about it, we wanted to understand what it was doing.”

Parasites lacking functional TgPRO accumulated more reactive oxygen molecules and struggled to compete at high density. Experiments showed that the loss of TgPRO disrupted the mitochondrion — the structure that supplies much of a cell’s energy — and changed how parasites processed glucose and other nutrients. Providing additional iron or restoring an important chemical balance inside the mitochondrion improved parasite growth, connecting TgPRO’s effects to iron-dependent energy metabolism.

The team then traced the response to a molecular mechanism. TgPRO is an RNA-binding protein, meaning it attaches to the molecular messages (RNAs) that cells use to make proteins. The researchers found that it binds and stabilizes a select set of messages involved in nutrient use, mitochondrial activity, and the assembly of iron-sulfur clusters, small structures that many enzymes need to function. The experiments connected the original observation — that some parasites faltered only when crowded — to a precise interaction between a regulatory protein and its RNA targets.

“One of the really nice elements of the story is our ability to connect it all the way through — from the original observation and genome-wide screen to the metabolic consequences and the direct interaction between TgPRO and its target RNAs,” Lourido says.

The researchers found that lowering oxygen levels also reduced oxidative stress and partially restored the growth of parasites without TgPRO. Toxoplasma is commonly grown in laboratories at atmospheric oxygen levels, which are considerably higher than those found in most animal tissues. The result suggests that oxygen conditions can strongly shape parasite metabolism, and the researchers caution others studying Toxoplasma to take this into consideration.

Connecting TgPRO to chronic infection

After testing the role of TgPRO in artificially crowded settings, the team also tested whether TgPRO matters during chronic infection, when Toxoplasma forms cysts in the brain. Mice infected with parasites lacking functional TgPRO developed smaller brain cysts, suggesting TgPRO supports parasite growth in the naturally dense environment of a chronic-stage cyst.

“The chronic stage is still somewhat elusive,” Giuliano says. “Showing that TgPRO affects cyst growth suggests that these same metabolic changes are needed in the brain and gives us clues about how the parasites persist there for months or years.”

TgPRO bears little resemblance to the proteins that regulate similar metabolic programs in mammals, yeast, and bacteria, yet it controls many of the same kinds of genes that these organisms adjust when cells face oxidative stress or changing nutrient conditions. This is an example of convergent evolution: Distantly related organisms evolved different molecular machinery to solve a similar biological problem.

That convergence suggests that coordinating these metabolic pathways may be a fundamental requirement for cells adapting to stress.

Altogether, the study establishes a new paradigm for how apicomplexan parasites regulate their metabolism, and advances the foundation for investigating how Toxoplasma persists inside its hosts.

This work was supported by National Institutes of Health grants and by a Burroughs Wellcome Fund grant awarded to S.L. M.A.S is funded by an Early Career Award from the Wellcome Trust. C.R.H. is funded by a Sir Henry Dale Fellowship from the Wellcome Trust and the Royal Society. J.K. is supported through funding by a generous donor advised by CARIGEST SA and acquired by D.S.-F.

DNA shaper steers nervous system development

A team in the Horvitz Lab have discovered that cohesin, a protein complex that helps shape the structure of the genome, is also critical for establishing the identity of some neurons. This could help scientists find a way to treat rare developmental disorders caused by mutations to the cohesin complex.

Jennifer Michalowski | McGovern Institute
August 3, 2026

A functional nervous system depends on the cooperation of many kinds of cells. So as developing organisms build their nervous systems, their neurons must take on different forms and functions to fulfill their designated roles. That carefully orchestrated process gives rise to thousands of different cell types in the human brain.

In the tiny worm known as C. elegans, the nervous system is far simpler, comprising a mere 118 classes of neurons.

At MIT, scientists in H. Robert Horvitz’s lab are studying the worms to learn about how nervous systems develop. Horvitz is the David H. Koch Professor of Biology at MIT, an Investigator at the McGovern Institute for Brain Research at MIT, and an Investigator at the Howard Hughes Medical Institute. His team has just discovered that a protein complex called cohesin, which helps shape the three-dimensional structure of the genome in both worms and humans, is critical for establishing some neurons’ identities as development unfolds.

The findings, reported July 31, 2026, in the journal Science Advances, could help scientists find a way to treat a rare developmental disorder called Cornelia de Lange syndrome, which is caused by mutations that interrupt the cohesin complex.

Model organism

Postdoctoral researcher Dongyeop Lee explains that C. elegans is a powerful model for studying neurodevelopment not just because its nervous system has been comprehensively mapped, but also because of the ease and speed with which scientists can study the function of its genes.

Because many of the worm’s genes have been retained through evolution, findings from studies of C. elegans often reveal important aspects of human biology. The current study began with worms that, because of a genetic mutation, make too many neurons of a certain type.

Adrenergic neurons, named for the kind of neurotransmitter they use to communicate with other neurons, are vital for enabling worms to respond to both their environment and their own internal state. Normally, C. elegans has just two pairs of adrenergic neurons: two RIM neurons and two RIC neurons. But the worms Lee studied had extras of both.

Takashi Hirose, a former member of the Horvitz lab, first observed this change in 2007.

Lee later continued the study and discovered that worms carrying a mutation in a gene called coh-1 have extra adrenergic neurons. The coh-1 gene encodes one part of the cohesin complex.

When Lee tested other mutations that disrupt cohesin, he found the same effect: Worms without fully functional cohesin had too many RIM neurons and too many RIC neurons.

Molecular switch

With a series of experiments designed to tease apart how cohesin impacts neurons’ identities, Lee discovered that cohesin cooperates with a gene-regulating protein called EOR-1 (known in humans as PLZF) to direct some neurons to develop into neurons that communicate with the inhibitory neurotransmitter GABA.

By reorganizing the structure of the genome, cohesin can change the way gene regulators like EOR-1 interact with DNA. Lee’s experiments showed that when either cohesin or EOR-1 couldn’t do its job, cells that should have become GABA-producing neurons become adrenergic neurons instead.

“What we found is that there are two alternative possible fates of certain neurons, and cohesin acts as a molecular switch that decides one of the possible neuronal fates,” Lee explains. “This means the structure of genomic DNA in the nucleus is important for neuronal fate determination.”

Disease connection

Lee adds that extra adrenergic neurons were not the only abnormality he observed in worms with cohesin mutations. Cohesin is important for shaping cells and tissues throughout the body. “The mutants have severe developmental defects,” Lee says. “They grow slowly. They don’t move well, and they also have defects in reproduction.”

Notably, the problems Lee saw in the worms echo aspects of Cornelia de Lange syndrome, a rare genetic disorder that impacts physical, cognitive, and behavioral development. Cornelia de Lange syndrome can be caused by mutations in cohesin genes, and Lee says that the discovery of how cohesin mutations affect worm development and behavior opens new opportunities to study the disease and search for potential therapeutic targets in C. elegans.

The Horvitz lab already has some promising leads. Taking advantage of the quick genetic screens that are possible in worms, Lee has found additional mutations that can counteract impaired cohesin, improving the health of worms with cohesin mutations. The team is now working to identify the genes where these suppressor mutations occur, so they can investigate whether they might make good therapeutic targets in humans.

Meanwhile, the team is also exploring a potential role for cohesin in shaping the fates of other neuron types, as well as searching broadly for additional molecules that work with cohesin to guide development. “We expect we have opened up a new biology,” Lee says. “This paper is just the beginning.”

Paper.

Why are some bacterial genes high in purines?

In certain species of bacteria, the answer lies in shielding RNA transcripts from a quality-control factor called Rho. Understanding the requirements for expressible sequences is critical for expression engineering of therapeutic agents.

Lillian Eden | Department of Biology
July 2, 2026

In the study of bacteria, a longstanding dogma held that two molecular machines — RNA polymerase, which leads the way in transcribing DNA into RNA, and ribosomes, which bring up the rear translating RNA into proteins — worked so closely in tandem that they were effectively attached.

This close coupling of transcription and translation in bacteria was thought to be fundamental to gene expression in part because the trailing ribosome could shield nascent gene products from an effective and omnipresent quality-control protein called Rho.

In bacteria that exhibit something called runaway transcription, however, the polymerase instead speeds ahead, unhitched from its protective ribosome. Inexplicably, however, in bacteria that exhibit this runaway transcription, such as Bacillus subtilis, Rho targeted primarily noncoding, useless RNA products.

New research from the Department of Biology reveals that the secret to Rho’s quality-control specificity lies in the sequence composition of nucleotide bases that make up coding strands of DNA.

“We started with a hypothesis that Rho was regulated by sequence, but the fact that the sequence alone was enough to protect any gene in the entire B. subtilis genome from Rho was really surprising,” says Julia Dierksheide PhD ’26, a graduate student in the Li Lab and first author of a paper recently published in Nature Microbiology. “That’s a really diverse range of sequences — what sequence feature is shared by every single gene in the genome?”

Barricading with bias

Rho serves as a termination factor, meaning that it is a crucial mechanism for preventing bacteria from wasting precious resources by making RNA transcripts that serve no purpose.

All the information a bacterial cell needs is encoded in its DNA, which is made up of two strands of nucleic acids. These strands twist together to form a double helix, with genetic information codified in pairs of bases: purines guanine and adenine are matched with pyrimidines cytosine and thymine, respectively. Any sequence that gives rise to RNA transcripts is stored in complement to a parallel, noncoding strand, meaning that a large portion of genetic material is transcriptionally useless.

Coding DNA strands in certain bacteria were known to be significantly higher in purines guanine and adenine compared to the rest of the bacterial genome. The researchers found that this purine bias alone shields productive mRNA transcripts from Rho-mediated termination.

“I love having a big, complicated dataset and trying to reduce that to biological meaning,” Dierksheide says. “It seems like Rho itself has been broadly shaping the evolution of the B. subtilis genome to create these sequence composition biases.”

Bacterial species that, over generations, have lost Rho no longer exhibit this strong purine bias.

Rho also serves as a regulatory factor in bacteria becoming motile, forming biofilms, or sporulating, all of which are critical for biology and survival. The purine bias could also provide a layer of protection against the insertion of foreign DNA, for example, when a viral bacteriophage infects bacteria.

“Bacteria exist as single cells, so everything that they do, they have to do through gene expression,” Dierksheide says. “Understanding the fundamental details about how gene expression works, how a cell encodes all the information it needs to survive in the nucleotide sequence of the genome, is really exciting.”

Future directions

Although the exact mechanism underlying Rho’s specificity remains unclear, these results crack an underlying code in the composition of bacterial genomes.

Dierksheide said she hoped to perform a similar screen to characterize Rho’s specificity in Escherichia coli, which diverged from B. subtilis on the evolutionary tree an estimated 2 billion years ago and still exhibits coupled transcription-translation, where the transcribing RNA polymerase is closely followed by a translating ribosome.

The high sequence specificity of B. subtilis Rho is crucial for the protection of its runaway RNA polymerase, in which that molecular machine speeds ahead of the ribosome. A systematic comparison to E. coli Rho could help reveal how this heightened stringency arose.

This information will be critical for engineering diverse bacterial species for applications including the production of therapeutic agents. Other bacterial species, such as B. subtilis, may be better models for this process because they have abundant secretion pathways, according to Dierksheide, making it much easier to produce and isolate proteins in large quantities.

“Our findings reveal an important criterion for successful sequence design that must be considered in expression engineering,” says associate department head, associate professor of biology, and Howard Hughes Medical Institute investigator Gene-Wei Li, the lead author of the study. “There are so many cryptic messages in the genome, like the purine bias, and we are just beginning to be able to decipher what they mean.”