AI-Powered Project Seeks to Predict Protein Dysfunction in Neurodegenerative Diseases

Decoding the language of these proteins could intercept the disease before it starts
Halfmann describes the potential of understanding how disordered proteins interact to prevent neurodegeneration.
Mark

So the core problem here is that we can see protein damage after it's already happened, but we can't predict when it's going to start?

Mimi

Exactly. By the time symptoms show up or imaging catches the damage, you've already lost a lot of neurons. Halfmann's team is trying to build an early warning system.

Luke

But the data doesn't exist yet, right? That's why they need to generate 10 billion measurements. We don't actually know if AI can learn this pattern once it has the data.

Mimi

True. But Halfmann's lab has already shown that protein behavior in yeast can predict what happens in human cells. This is scaling that approach.

Mark

Why yeast? Why not just work with human cells from the start?

Mimi

Yeast is simpler, faster, cheaper. You can test thousands of variants quickly. Then they'll validate the predictions in actual human neurons through the Berkeley team.

Luke

And the timeline—they're funded for two years to generate data, then another two years to validate therapeutics. That's still a long way from a patient getting a preventive treatment.

Mimi

Right. This is foundational work. But if it works, it changes the game from treating symptoms to preventing disease.

Mark

What about the diseases they're not focusing on first? Will this approach work for Parkinson's and Alzheimer's too?

Mimi

That's the stated goal. They're starting with FTLD because it shares features with ALS, but the framework should apply across diseases involving protein misfolding.

Luke

Should apply. We won't know until they try. And we don't know yet whether the AI can actually decode this language, even with perfect data.

  • Neurodegenerative diseases share a silent prologue — proteins misfolding and clumping long before any scan or symptom can reveal the harm already done.
  • Current AI can model stable proteins with precision, but the shape-shifting intrinsically disordered proteins that drive neurodegeneration remain stubbornly beyond its reach.
  • The Stowers Institute's Randal Halfmann will use a cell-level measurement technology his own lab invented to produce over 10 billion data points — a dataset that has never existed at this scale.
  • A coalition spanning UC Berkeley, Johns Hopkins, Brown, Emory, and Texas A&M is pooling expertise in human neuron generation, protein interaction mapping, and deep learning to decode what no single institution could alone.
  • If the models succeed, patients may one day learn their risk of neurological disease early enough to pursue preventive treatments or join clinical trials before significant damage accumulates.

Among the most elusive frontiers in medicine is the moment before illness declares itself — the quiet molecular unraveling that precedes decades of neurological loss. A $28.6 million federal initiative called NATIVE-ID now brings together researchers from across the United States to train artificial intelligence on the behavior of proteins that have no fixed shape, the very proteins implicated in Alzheimer's, Parkinson's, ALS, and Huntington's disease. By generating more than 10 billion measurements of protein aggregation across 50,000 proteins, the project seeks to give science something it has never had: the ability to anticipate dysfunction before the damage begins.

Every major neurodegenerative disease — Alzheimer's, Parkinson's, ALS, Huntington's — carries the same molecular signature: proteins that lose their shape, clump together, and destroy cells. By the time that damage becomes visible, it is often already vast. A new $28.6 million federal initiative aims to intervene far earlier, by teaching artificial intelligence to predict which proteins are about to go wrong.

The core difficulty lies in a class of proteins that defy easy study. Roughly one-third of all human proteins have no fixed structure — they shift constantly among shapes, and that instability makes them capable of triggering neurodegeneration. Existing AI handles stable proteins well, but these intrinsically disordered proteins, or IDPs, remain largely opaque to current models. Randal Halfmann of the Stowers Institute for Medical Research in Kansas City will anchor the effort, generating the foundational dataset the AI needs to learn. Using a cell-level measurement technology his lab developed in 2018, his team will test 50,000 proteins across more than one million samples, producing over 10 billion individual measurements of how proteins aggregate inside living cells — a scale that has never before been attempted.

The project, called NATIVE-ID, unites scientists from UC Berkeley's Innovative Genomics Institute, Johns Hopkins, Brown, Emory, Texas A&M, and other institutions. Each partner contributes something distinct: human neurons derived from varied genetic backgrounds, tools to measure protein interactions within those neurons, and the deep-learning frameworks needed to make sense of it all. The team will begin with frontotemporal lobar degeneration, a condition that shares biological roots with ALS, building on Halfmann's earlier work on Huntington's disease and the ALS-linked protein TDP-43.

The initiative is structured in two phases. The first two years will focus on building the datasets and models. The second phase will turn toward validating potential therapeutics and developing earlier detection methods. The broader ambition is transformative: if researchers can predict when and in whom proteins are likely to misfold, patients could pursue preventive care or enroll in clinical trials before neurological damage becomes irreversible. NATIVE-ID is part of BIOGAMI, a wider ARPA-H program aimed at understanding and ultimately controlling harmful protein aggregation across the full spectrum of neurodegenerative disease.

Alzheimer's, Parkinson's, ALS, and Huntington's disease all share a common signature: proteins that lose their shape, clump together, and destroy cells. By the time a doctor can see these changes clearly on a scan or in a lab test, the damage is often already extensive. A new research initiative, backed by $28.6 million in federal funding, aims to catch that process much earlier—before the cascade begins—by training artificial intelligence to predict which proteins are about to malfunction.

The challenge is that roughly one-third of all proteins in the human body don't have a fixed structure. These intrinsically disordered proteins, or IDPs, shift constantly among different shapes, and that flexibility is precisely what makes them dangerous. They can tangle with other proteins in ways that trigger neurodegeneration. Current AI systems excel at predicting the behavior of stable proteins, but IDPs remain largely opaque. "Treatments for neurodegenerative diseases remain extremely limited," said Randal Halfmann, an investigator at the Stowers Institute for Medical Research in Kansas City. "Decoding the language of these IDP interactions could allow us to create therapeutic IDPs to 'intercept' problematic interactions that drive diseases like Alzheimer's."

Halfmann's laboratory will anchor the new project, called NATIVE-ID, by generating the massive dataset that AI systems need to learn. His team will receive approximately $4.1 million over two years from the Advanced Research Projects Agency for Health, or ARPA-H. The work will unfold across more than one million samples, producing over 10 billion individual measurements of how proteins aggregate. To accomplish this, Halfmann will test 50,000 different proteins, each expressed separately in yeast cells under conditions designed to replicate what happens inside human cells as we age. He will use a technology his own lab developed in 2018 called Distributed Amphifluoric FRET, or DAmFRET, which measures protein self-assembly inside living cells with cellular precision. "Direct measurements of protein aggregation at cellular resolution have not previously been done at this scale," Halfmann said.

The NATIVE-ID project brings together scientists from UC Berkeley, Johns Hopkins University, Brown University, Emory University, Texas A&M University, and other institutions. UC Berkeley's Innovative Genomics Institute leads the effort. Each partner contributes a distinct capability: some will generate human neurons from different genetic backgrounds, others will measure protein interactions in those neurons, still others will develop the deep-learning frameworks needed to interpret the data and run large-scale computational simulations. Alejandro Sánchez Alvarado, president and chief scientific officer at the Stowers Institute, emphasized that neither component of the collaboration could succeed alone. "A deep learning model that is to decode the language of disordered proteins needs rigorously obtained data at scale, something that has yet to exist," he said. "Randal's laboratory will provide what has been missing: more than 10 billion measurements of aggregation, taken one living cell at a time, across 50,000 proteins. Our colleagues at the Innovative Genomics Institute will bring the human neurons in which the model's predictions must ultimately hold."

The team will initially focus on frontotemporal lobar degeneration, or FTLD, a neurodegenerative disease that shares genetic and biological features with ALS. The work builds on Halfmann's prior research: in 2023, his lab became the first to experimentally determine the structure of the initiating step in amyloid formation associated with Huntington's disease. His team has also studied TDP-43, a protein strongly implicated in both ALS and FTLD. Those earlier studies examined hundreds of protein sequences. This new effort will scale that work to 50,000. "Those were sort of pilots for this new undertaking," Halfmann said.

The potential payoff extends beyond basic science. If researchers can predict when proteins are likely to misfold and at what age disease onset might occur, patients could seek preventive treatments or enroll in clinical trials before significant neurological damage accumulates. "If we can better predict the probabilities and onset ages of disease, it could allow many more people to seek preventive or early-stage treatments or enroll in clinical trials," Halfmann said. The current award funds an initial two-year phase focused on generating foundational datasets and models. A second phase, also planned for 24 months, will focus on validating potential therapeutics and developing methods to detect protein dysfunction earlier in the disease process. The work is part of BIOGAMI, a broader ARPA-H initiative aimed at understanding and ultimately controlling harmful protein aggregation across multiple neurodegenerative conditions.

If we can better predict the probabilities and onset ages of disease, it could allow many more people to seek preventive or early-stage treatments or enroll in clinical trials.
— Randal Halfmann, Stowers Institute
Decoding the language of these IDP interactions could allow us to create therapeutic IDPs to 'intercept' problematic interactions that drive diseases like Alzheimer's.
— Randal Halfmann
Quieres la nota completa? Lee el original en News-Medical ↗
Contáctanos FAQ