At the intersection of consumer technology and clinical medicine, researchers at MIT and Empirical Health have quietly demonstrated that the imperfect, interrupted data streaming from millions of Apple Watches may already contain the seeds of early disease detection. By adapting an AI architecture originally conceived at Meta — one that learns meaning rather than reconstructing missing information — they trained a system called JETS on nearly three million days of human health observation. The result is not a finished medical device, but a philosophical shift: incompleteness, long treated as a
MIT researchers train disease-detection AI on 3M days of Apple Watch data
Incomplete data still contains patterns worth learning from
Why does it matter that 85 percent of the data was unlabeled? Couldn't they have just thrown it away?
Because unlabeled data still contains patterns. A person's heart rate and sleep and activity tell a story even if you don't know whether they have hypertension. The old way required you to know the answer first. This way, the AI learns the underlying structure of health from everyone, then learns to recognize disease from the smaller group where we actually know the diagnosis.
So JEPA is about predicting what you can't see from what you can see.
Exactly. In a photograph, you see the sky and trees; you infer what's behind the tree even though you can't see it. Here, you see a person's heart rate pattern and sleep quality; you infer whether they have atrial fibrillation, even from gaps in the data.
Why is this better than just asking the Apple Watch to measure more things more often?
Because people don't wear their watches perfectly. They charge them, they take them off, they forget them. You can't force perfect data collection. But you can build AI that works with imperfect data as it actually exists in the real world.
The accuracy numbers—70 to 87 percent—is that good enough to actually use in medicine?
That depends on the condition and the use case. For screening, for flagging someone to see a doctor, it's promising. For diagnosis, you'd want higher confidence. But remember, these aren't accuracy numbers in the traditional sense. They're measures of how well the model ranks which patients are most likely to have the disease. That's often more useful than raw accuracy.
What happens next?
The real test is whether this works on new people, new watches, new data the model has never seen. And whether doctors will trust it enough to act on it. The potential is enormous—billions of hours of wearable data already exist. But potential and clinical reality are different things.
El Pulso
- 85% of real-world wearable health data has historically been discarded as too irregular or incomplete for traditional machine learning — a vast reservoir of human signal left untouched.
- JETS reframes the problem entirely, teaching AI to infer the meaning of missing health data rather than reconstruct it, borrowing an architecture from Meta's AI research to handle the chaos of how people actually wear devices.
- Tested across conditions from hypertension to atrial flutter, the model achieved AUROC scores between 70 and 87 percent — outperforming baseline models across most evaluations and earning acceptance at NeurIPS.
- The findings land not as a clinical product but as a proof of concept: the diagnostic potential already sitting in consumers' wrists may be far greater than medicine has yet dared to extract.
At the intersection of consumer technology and clinical medicine, researchers at MIT and Empirical Health have quietly demonstrated that the imperfect, interrupted data streaming from millions of Apple Watches may already contain the seeds of early disease detection. By adapting an AI architecture originally conceived at Meta — one that learns meaning rather than reconstructing missing information — they trained a system called JETS on nearly three million days of human health observation. The result is not a finished medical device, but a philosophical shift: incompleteness, long treated as a flaw in data, may be a condition the right intelligence can work with rather than around.
Researchers at MIT and Empirical Health have trained an AI system to detect disease from nearly three million days of Apple Watch data — a dataset so large and irregular that it demanded a fundamentally different approach to machine learning.
The system is built on JEPA, a Joint-Embedding Predictive Architecture originally proposed by Meta's chief AI scientist Yann LeCun. Rather than reconstructing missing data — imagining the hidden pixels in a blacked-out photograph — JEPA teaches an AI to understand what the missing data represents. This shift from reconstruction to representation has become central to a growing field of so-called "world models," distinct from the token-prediction logic behind systems like ChatGPT.
The MIT team adapted this architecture to address a stubborn problem: wearable health data is messy. Heart rate, sleep, activity, and respiratory metrics arrive at irregular intervals, and some measurements appear in fewer than one percent of daily readings. Traditional methods would have discarded roughly 85 percent of the dataset as unusable. Instead, their model — JETS — learned patterns from the full unlabeled dataset first, then refined itself on the smaller portion where medical histories were documented.
The dataset spanned 16,522 individuals across 63 daily metrics in five health categories. JETS converted each observation into a token, masked portions of the data, and trained itself to infer the meaning of what was hidden. When tested, it achieved 86.8% accuracy for hypertension, 81% for chronic fatigue syndrome, and 70.5% for atrial flutter — consistently outperforming baseline models across most conditions.
What the numbers point toward is larger than any single diagnosis: the millions of Apple Watches already worn inconsistently by ordinary people may hold far more medical insight than has ever been extracted — waiting only for an architecture willing to meet the data as it actually is.
Researchers at MIT and Empirical Health have trained an artificial intelligence system to detect disease using nearly three million days of Apple Watch data—a dataset so large and varied that it required a fundamentally different approach to machine learning than what typically works in medicine.
The project centers on a technique called JEPA, or Joint-Embedding Predictive Architecture, originally proposed by Yann LeCun when he was Meta's chief AI scientist. The core insight is deceptively simple: instead of trying to guess what missing data actually contains, teach the AI to understand what the missing data represents. Imagine a photograph with parts of it blacked out. Rather than trying to reconstruct the exact pixels beneath the black boxes, JEPA learns to embed both the visible and hidden regions into a shared conceptual space, then infers the meaning of what's hidden from the context of what's visible. This shift—from reconstruction to representation—has become foundational to a new field exploring what researchers call "world models," a departure from the token-prediction focus that powers systems like ChatGPT.
The MIT team adapted this architecture for a specific problem: wearable health data is messy. Heart rate, sleep, activity, respiratory metrics—these don't arrive in neat, regular intervals. Some measurements appear almost never; others show up nearly every day. Traditional machine learning requires complete, labeled datasets, which meant that roughly 85 percent of the researchers' data would have been discarded as unusable. Instead, they developed JETS (a self-supervised joint embedding time series foundation model), which learned patterns from the entire unlabeled dataset first, then fine-tuned itself on the smaller subset of data where medical histories were actually documented.
The dataset itself is substantial: 16,522 individuals contributing roughly three million person-days of observation, with 63 distinct daily or sub-daily metrics across five categories—cardiovascular health, respiratory health, sleep, physical activity, and general statistics. Some metrics appeared in only 0.4 percent of daily readings; others in 99 percent. The researchers converted each observation into a token (day, value, metric type), masked portions of the data, encoded it, and ran it through a predictor designed to infer the embedding of the missing patches.
When they tested JETS against baseline models, the results were striking. The system achieved 86.8 percent AUROC (a measure of how well it ranks likely cases) for high blood pressure, 81 percent for chronic fatigue syndrome, 86.8 percent for sick sinus syndrome, and 70.5 percent for atrial flutter. It didn't always outperform every comparison, but the advantages were consistent across most conditions evaluated. The paper was recently accepted to a workshop at NeurIPS, the major machine learning conference.
What matters here is not just the numbers but what they represent: a proof that incomplete, irregular wearable data—the kind that real people generate when they wear a smartwatch inconsistently—can still yield meaningful medical insights if you train the AI differently. The study suggests that the millions of Apple Watches already in people's pockets contain far more diagnostic potential than anyone has yet extracted, waiting only for the right architecture to unlock it.
Citas Notables
The model learns internal patterns from the complete dataset first, then fine-tunes on the smaller subset where medical histories are documented— Study methodology