AI Framework Designs Novel Antibodies Against Gastric Cancer Target CLDN18.2

All of that filtering was computational. None of it has been tested experimentally.
The researchers generated 313 candidate antibodies using AI predictions, but actual laboratory validation remains ahead.
Mark

So the researchers built an AI system to design antibodies. What's the actual problem they're trying to solve?

Mimi

Gastric cancer is hard to treat, and even the best drug we have—zolbetuximab—only extends survival by a couple of months. The researchers want to design better antibodies against the same target, CLDN18.2. But making antibodies is slow and expensive using traditional methods.

Luke

Right, but let's be precise about what they did. They didn't design a whole new antibody from scratch. They took an existing antibody scaffold—zolbetuximab—and used AI to redesign just one small part of it, the CDRH3 loop. That's a much narrower problem.

Mimi

Exactly. The CDRH3 is the part that usually does the binding. So if you can optimize that loop while keeping everything else the same, you might get better binding without losing stability or manufacturability.

Mark

How many candidate sequences did they end up with?

Mimi

They generated 50,000 sequences, then filtered them down to 313 that met multiple criteria—good predicted binding, good charge properties, low predicted immunogenicity, and so on.

Luke

But here's the thing: all of that filtering was computational. They never actually made any of these antibodies. They used AlphaBind to predict binding, but they're explicit that AlphaBind is a relative ranking tool, not a calibrated measure of actual affinity. And they couldn't confirm from their models whether these antibodies would cross-react with CLDN18.1, which is a real clinical concern.

Mark

So the 313 candidates are just predictions?

Mimi

Yes. They used AlphaFold3 to model the 3D structure and molecular dynamics simulations to check stability, but none of it has been tested experimentally.

Luke

And they tested the framework on HER2 as a proof of concept, which is good. But when they compared cdrGPT to other antibody language models, it wasn't uniformly better. It was better at some metrics, worse at others. That's honest reporting, but it means this isn't a silver bullet.

Mark

What would it take to actually use one of these antibodies as a drug?

Mimi

You'd have to synthesize it, express it in cells, test it for binding to CLDN18.2 in a lab, test it for cross-reactivity with CLDN18.1, test it in cell-based assays, test it in animals, and eventually in humans. All of that is ahead.

Luke

The researchers acknowledge all of this. They're not claiming they've discovered new drugs. They're claiming they've built a tool that can narrow the search space from millions of possibilities to a few hundred worth testing. That's useful, but it's not the same as validation.

Mark

Is the AI actually learning something about antibodies, or is it just pattern-matching?

Mimi

It's trained on 637,887 real antibody sequences from a database, so it's learning statistical patterns about how natural antibodies are built. That's different from random generation. But whether those patterns translate to function in a new context—that's what experiments will tell.

Luke

And the model was fine-tuned using reinforcement learning to optimize for multiple objectives at once. That's more sophisticated than just pattern-matching. But again, the proof is in the lab.

  • Gastric cancer kills relentlessly, and even the best available antibody therapy — zolbetuximab — extends survival by only a few months, making the hunt for better drugs genuinely urgent.
  • Traditional antibody engineering forces scientists to trade one desirable property against another, making it nearly impossible to optimize binding strength, stability, and immune safety all at once.
  • cdrGPT attacks this bottleneck by training a language model on hundreds of thousands of real antibody sequences, then steering it with reinforcement learning toward candidates that satisfy multiple criteria simultaneously.
  • From 50,000 generated sequences, a multi-step computational filter — checking binding affinity, molecular behavior, and immunogenicity risk — reduced the field to 313 viable candidates in a matter of hours.
  • The critical gap remains: none of these antibodies have been synthesized or tested in a lab, and whether they can distinguish their target from a nearly identical protein is still an open and clinically vital question.

In the long struggle against gastric cancer — a disease that claims hundreds of thousands of lives each year and resists most treatments — researchers have turned to artificial intelligence not as a cure, but as a compass. A team has built a system called cdrGPT that learns the grammar of antibody design from nearly 640,000 natural sequences, then uses that knowledge to propose new molecular candidates targeting CLDN18.2, a protein that cancer has made vulnerable. The result is 313 computationally promising antibodies — not yet proven in a laboratory, but representing a faster, more principled way to begin the search.

Gastric cancer is one of the world's deadliest malignancies, and in China alone it accounts for nearly four in ten new cases globally each year. Treatments remain limited, but one target has earned genuine clinical credibility: a protein called CLDN18.2, which sits on the surface of stomach cancer cells and is largely absent from healthy tissue. The antibody drug zolbetuximab, which locks onto this protein, extended median survival in advanced gastric cancer patients from 15.5 to 18.2 months when added to chemotherapy — proof that the target is real, and that better drugs aimed at it could matter enormously.

The problem is that engineering improved antibodies the conventional way is slow and difficult. Optimizing for binding strength, stability, low toxicity, and reduced immune risk all at once — using traditional methods — is an exercise in competing trade-offs. A research team built cdrGPT to change that calculus. The system trained a deep-learning model on 637,887 real antibody sequences, learning the statistical patterns nature uses to build these molecules. It then applied reinforcement learning to steer the model toward sequences with strong predicted binding to CLDN18.2 and good practical properties — sequences that wouldn't clump, wouldn't be cleared too quickly, and wouldn't provoke unwanted immune responses.

Starting from 50,000 generated candidates, the team ran a multi-step computational filter: predicted binding affinity, molecular charge and hydrophobicity, and immunogenicity risk. After all that screening, 313 sequences survived. They clustered into three distinct groups, with one cluster closely resembling zolbetuximab itself. Seven of the most promising candidates were then modeled using AlphaFold3, and molecular dynamics simulations confirmed that the predicted binding complexes remained stable over time.

The researchers are candid about what this is and isn't. No antibody has yet been made in a lab, tested against cancer cells, or evaluated in an animal. The computational tools used for ranking are relative guides, not guarantees. A particularly important gap is selectivity: the models couldn't confirm whether these new antibodies would distinguish CLDN18.2 from its near-identical cousin CLDN18.1 — a distinction that matters clinically, since cross-reactivity could cause off-target harm. What cdrGPT offers is a tractable pipeline: a way to compress a vast search space into a manageable set of candidates worth testing, generated quickly and evaluated against multiple criteria at once. For a disease where better options are desperately needed, that is a meaningful step — even if the finish line still lies somewhere in the laboratory ahead.

Gastric cancer kills more people than most cancers, and the options for treating it remain thin. In China alone, nearly four in ten of the world's new gastric cancer cases appear each year. Even with chemotherapy, survival is poor. Researchers have been hunting for better targets—proteins on cancer cells that drugs can lock onto—and one has emerged with real clinical weight: a protein called CLDN18.2, which sits on the surface of stomach cancer cells and barely exists anywhere else in the body.

Zolbetuximab, an antibody drug that targets CLDN18.2, proved itself in a major trial. When added to standard chemotherapy, it stretched median overall survival from 15.5 months to 18.2 months in patients with advanced gastric cancer. That's real. It's also a proof of concept: CLDN18.2 works as a target. But making better antibodies against it has been a bottleneck. The conventional way of engineering antibodies is slow, labor-intensive, and hard to optimize for multiple properties at once—you want binding strength, but also stability, low toxicity, and low risk of triggering an immune response. Balancing all of that simultaneously, using traditional methods, is like threading a needle in the dark.

A team of researchers built an artificial intelligence system called cdrGPT to solve this problem. The core idea is elegant: antibodies have a highly variable region called CDRH3—a loop in the antibody's heavy chain that often does much of the work of recognizing and grabbing the target. Instead of engineering this loop by hand, they trained a deep-learning model to generate new sequences for it. The model learned from a database of 637,887 real antibody sequences, absorbing the statistical patterns of how nature builds these molecules. Then they fine-tuned it using reinforcement learning, steering it toward sequences that would bind well to CLDN18.2 while also having good developability—meaning they wouldn't clump up, wouldn't be cleared too fast from the bloodstream, and wouldn't trigger unwanted immune responses.

Starting with 50,000 generated sequences, they filtered them through a multi-step screen. They checked predicted binding affinity using a computational model called AlphaBind. They measured net charge and hydrophobicity to predict how the antibody would behave in solution. They predicted how likely the sequence was to be recognized as foreign by the human immune system. After all that filtering, 313 sequences survived. These weren't random—they clustered into three distinct groups based on their molecular features, with one cluster containing sequences very similar to zolbetuximab itself. The researchers then used AlphaFold3, a structure-prediction tool, to model how seven of these candidates would actually fold and bind to CLDN18.2. The predicted structures held together. The binding interfaces looked plausible. Molecular dynamics simulations—essentially watching the molecules move for 100 nanoseconds—showed the complexes remained stable.

But here's the hard truth: this is all computational. No one has yet made these antibodies in a lab, tested them against actual cancer cells, or put them in an animal. The researchers are careful about this. They note that AlphaBind was used only as a relative ranking tool, not as a measure of true binding strength. They acknowledge that their immunogenicity assessment, based on predicting MHC class II binding, doesn't capture everything that makes an antibody look foreign to the immune system. They also flag a critical gap: they couldn't confirm from their computational models that these new antibodies would specifically recognize CLDN18.2 and not its close cousin CLDN18.1, which differs in just one region. That distinction matters clinically. Cross-reactivity with CLDN18.1 could cause off-target toxicity.

What cdrGPT does offer is a tractable pipeline. It can generate diverse antibody sequences within a fixed scaffold—the zolbetuximab framework—and prioritize them based on multiple criteria at once. The model ran on four NVIDIA A100 GPUs for training and generated 20,000 candidate sequences in about four hours on a single GPU. That's fast enough to be useful. The researchers also tested the framework on other targets, like HER2, to show it could generalize. When they compared cdrGPT to two other antibody language models, AbLang and AntiBERTy, it showed trade-offs: better at predicting certain properties like hydrophobicity, weaker at predicting immunogenicity risk. No single model dominates.

The next step is the bench. The 313 candidates, or at least the most promising of them, need to be synthesized, expressed, and tested for real binding to CLDN18.2. They need to be tested against CLDN18.1 to confirm selectivity. They need to be tested in cells and in animals. Only then will anyone know whether the AI's predictions hold up in the messy reality of biology. That work hasn't happened yet. What exists now is a computational strategy—a way to narrow a vast search space down to a manageable set of candidates worth testing. For a disease where better drugs are desperately needed, that's a meaningful step forward, even if it's not the finish line.

Although experimental validation will be needed, this work provides a computational strategy for antibody optimization and illustrates the potential of AI-guided sequence design to support therapeutic antibody discovery.
— Study authors, in abstract
The present findings remain computational and require experimental validation. The prioritized candidates should be tested for binding affinity, CLDN18.2/CLDN18.1 specificity, expression, stability, and functional activity.
— Study authors, in discussion
Fale Conosco FAQ