Your watch says you slept seven hours and twenty minutes and that your recovery is “moderate.” A lab portal adds an automatically generated risk estimate beside one of your results. A headline announces that software now reads scans as well as specialists. Each of these numbers comes from a system trained on data, and each arrives with the quiet confidence of a measurement.

The question most people ask is whether to trust the number. It is the right question. It has also run through my entire career as a scientist. Before an experiment can answer anything, the measurement must earn our trust. Does it capture what we think it captures? How accurately? Where was it checked, and how often is it wrong? Only when those questions are satisfied does a result become evidence.

This piece begins “AI, Measured”, a series that asks those questions of artificial intelligence in healthcare. It is neither a case for AI nor a case against it. Some of these tools are remarkable, and the evidence shows it. Others are better at sounding certain than at being right. The difference rarely shows on the screen. It lives upstream, in what the system was given to learn from.

So, we start where my series often starts: with the measurement. Measurement comes before understanding and before any decision. AI does not get to skip that step.


A Model Learns from Worked Examples

Strip away the vocabulary and most health AI works the same way. The model is a pattern-finder. Many examples are shown, each paired with an answer; e.g., a chest X-ray and a radiologist’s verdict, or blood results and whether the patient was alive three years later. The algorithm adjusts itself until its output matches those answers, and then it is used where the answer is unknown.

The model never meets the patient. It meets numbers, images, and codes produced by devices, laboratories, and clinicians. It learns whatever pattern links those inputs to those answers. If the inputs were measured badly or the answers were wrong, it learns that too.

Think of an apprentice who learns to grade exams by copying an experienced examiner. If the examiner consistently marks one kind of answer wrong, the apprentice reproduces the error and looks excellent when checked against that examiner’s marks. Or picture a bathroom scale that always reads two kilograms heavy. You could study its numbers with great care, chart them, average them, even predict tomorrow’s reading perfectly, and you would still be two kilograms wrong. The error was never in the reading. It was in the scale.

These analogies have a limit. With enough data, models can partly cope with random noise, the kind that scatters readings around the truth. What they cannot see is systematic error, the kind that pushes readings one way or changes quietly over time. Statisticians have tested this directly. When measurement protocols differ between model development and clinical deployment, calibration degrades, and discriminatory power declines, compromising clinical utility. ¹ Even plain noise has a hidden cost: a model fed error-prone measurements can stay well calibrated, meaning its average risk estimates still match reality, while becoming much worse at telling individual patients apart. ²

Readers of the Biomarker Validity series will know this pattern from the other side. There, the usual failure was not a broken assay, but a valid measurement attached to the wrong biological question. When an AI model is trained on such a measurement, it does not correct the mismatch. It learns it, automates it, and repeats it for every patient it scores.

Sometimes the Model Learns the Hospital, Not the Patient

In 2018, researchers examined one year of laboratory records for 669,452 patients at two large Boston hospitals. ³ They analyzed how well 272 common laboratory tests predicted whether a patient would be alive three years later. Then they asked a strange question: what happens if you ignore the result entirely and look only at when the test was ordered?

The findings were striking. For 233 of the 272 tests, simply having the test ordered was associated with survival, regardless of the result. ³ For about two-thirds of the tests where the comparison models could be evaluated, order timing (the hour of the day, day of the week, or interval between orders) provided more predictive signal for three-year survival than the physiological test result value itself, i.e., the clock carried more information than the chemistry. ³

Why? The likely reason is that a blood draw ordered at 3:00 AM reflects an urgent clinical concern, whereas a routine morning panel reflects scheduled care. The clock captured physician worry, and that worry was highly predictive of patient outcomes. But it was hospital workflow, not biology.

A similar phenomenon appears in medical imaging. ⁴ When models were trained to detect pneumonia on chest X-rays from three distinct hospital systems, researchers discovered the models could identify which hospital an image came from with over 99.9% accuracy. That mattered, because pneumonia was present in 34.2% of images at one hospital and about 1% at the others. The model learned to recognize where an image came from, and the hospital with far more pneumonia, as a shortcut. When tested on a new hospital system, its diagnostic performance dropped substantially. ⁴

In both cases, nothing glitched. The data held a real pattern. It simply lived in how healthcare was delivered, not inside the patient.

Three Hidden Measurements Inside Every AI Result

An AI result rests on three separate measurements, and each can mislead on its own.

1. The Input

The input consists of raw numbers, codes, or images the model reads. A model learns what values mean during training, and if that meaning changes later, it has no way of knowing.

Imagine a model that learned what a given troponin result (a blood marker of heart injury) says about a patient. If a hospital upgrades to a more sensitive troponin test, the same numerical result no longer represents the same clinical reality, yet an unadjusted model keeps reading it the old way. Researchers cataloging clinical AI failures list this assay change as one recognized source of dataset shift. ⁵

This happens in practice. In Toronto, a system monitoring an in-hospital mortality-prediction model across seven hospitals and 143,049 patients detected harmful shifts in input data, including changes for two blood markers, i.e., brain natriuretic peptide (BNP) and D-dimer. ⁶ In another instance, a sepsis-alert model was deactivated at a US hospital in April 2020 after the COVID-19 pandemic changed the baseline link between fevers and bacterial infection. ⁵

2. The Label

The label is the target answer the model was trained to predict. Labels are measurements too, and they are usually created by humans or administrative records.

When a team reexamined how specialists graded diabetic eye disease in retinal photographs, they found that the single most common source of disagreement, accounting for 36% of discrepancies, was a missed microaneurysm, a tiny bulge in a retinal blood vessel. ⁷ Retraining the model with a small set of carefully adjudicated grades from retina specialists, together with higher-resolution images, raised the model’s area under the curve (AUC). The AUC answers a simple question: take one patient with disease and one without, and how often does the model correctly rank the sick one as more likely to be sick? A score of 0.5 is a coin toss; 1.0 is right every time. The AUC for moderate or worse disease rose from 0.934 to 0.986.⁷ Neither change made the algorithm itself smarter. One gave it better answers to learn from, the other a sharper input. Furthermore, health record data carry their own errors. Researchers who study bias in AI built from these records count measurement error and misclassification among its standard sources.

3. The Benchmark Ruler

The third measurement is the evaluation standard used to grade the model. Even if an AI model receives clean inputs and was trained on solid labels, we can still misjudge its performance if the metric we used to grade it is warped. To evaluate clinical AI, researchers compare its outputs against a reference answer key. But if that answer key contains errors, a perfectly accurate model will appear to fail. ⁸

Picture 1,000 people, 100 truly sick and 900 healthy, and a perfect model that flags exactly the 100 sick people. Now grade it with a key that never misses a sick person but mislabels 1 in 10 healthy people as sick. That key lists 190 “cases”: the 100 real ones, plus 90 healthy people marked by mistake. The perfect model rightly ignores those 90, so against the key it seems to catch only 100 of 190, about 53%. ⁸ Nothing about the model changed. Only the metric did.

Laboratory scientists know this territory well. In high-throughput omics profiling, results shift with laboratory conditions, reagent lots, and personnel. ⁹ These “batch effects” become dangerous when technical quirks align with the disease under study. In my experience with large-scale population serum lipidomics across 26 analytical batches, the rule is simple. Every run must carry shared reference samples. Benchmark evaluations show that scaling raw data into ratios against concurrent reference materials is the most reliable way to eliminate batch noise, even when technical artifacts and biological signals are completely tangled together. ¹⁰ Without this calibration, an AI model won’t learn the biology—it will learn the day the machine was tuned.

From Hospital Algorithms to the Wrist

This three-part structure—raw input, training target, and benchmark evaluation—doesn’t just apply to hospital systems or genomic pipelines. It plays out every minute on millions of human wrists.

Consumer wearables are a useful personal case study. A smartwatch is essentially a miniature health model: light sensors on your skin collect raw optical signals (the input), onboard algorithms process those signals (the model), and your screen displays a simplified health metric (the derived estimate). When we look at how well these devices perform, the exact same measurement rules apply.

A comprehensive living review synthesizing 24 systematic reviews across 430,000 participants highlights where this pipeline works—and where it breaks down. ¹³ Heart rate tracking holds up well, with readings averaging within ~3% of clinical reference standards. However, estimated aerobic fitness (VO2max) was overestimated by a mean of ~15% in resting tests, and consumer sleep trackers tended to overestimate total sleep time, with mean errors typically exceeding 10%.¹³ Most surprisingly, only about 11% of commercially released wearable devices were validated for even a single measurement outcome. ¹³

Performance as measured vs. usefulness as claimed

Performance as measured is what a model was shown to do, on which data, and against which answer key.

Usefulness as claimed is what it is said to do for you, here and now.

The first is a finding. The second is a promise — and the gap between them is often a measurement problem.

What Evidence Shows

None of this means AI in healthcare is failing. The trial evidence is real and growing. A review of 86 randomized controlled trials evaluating AI in active clinical practice found that 70 of them (81%) reported positive results on their main outcome, most often in detection tasks such as finding polyps during colonoscopy. ¹¹ However, 63% of these trials were conducted in a single center, about half measured diagnostic yield or performance rather than whether patients ended up healthier, and the authors caution that publication bias likely flatters the overall picture.¹¹

AI can process clinical data at speed, but it cannot transcend the fidelity of its inputs. A more sophisticated model cannot recover information the measurement never captured. If we want AI tools that deliver genuine health intelligence, we need rigorous, transparent, and validated measurements at every step of the pipeline.

What a good model can and cannot do

A good model can do remarkable things with a good measurement. It can find patterns too subtle for a person to see and apply them consistently, without fatigue.

What it cannot do is create information the measurement never captured or correct a systematic error it has no way of seeing. Without dedicated monitoring, a deployed model may not recognize when its input data have shifted, allowing performance to deteriorate silently while routine-looking outputs continue to appear. ¹,² This is why monitoring matters. In the Toronto study, a system that watched incoming data flagged harmful shifts, and updating the model when they appeared helped maintain its performance. ⁶ Better answer keys matter just as much, because a model’s reliability is constrained by the examples it learns from and the yardstick used to judge it.⁷,⁸

Open questions remain. How often should a deployed model be rechecked? Who is responsible for noticing when a laboratory changes its method? How should a device tell you which of its numbers are measured and which are estimated? These are quality-control questions, the kind laboratories have long answered for their instruments.

So, when the next AI-generated number arrives, on your wrist, in your lab portal, or in a headline, start with one question before any other.

What was it actually measuring?

It will not be the only question worth asking. It is the one that comes first, because no model, however clever, can rescue a bad measurement. That is where measurement becomes meaning, or fails to.


This article is for informational and educational purposes only and does not constitute medical advice. Scientific claims are grounded in the peer-reviewed literature cited; individual findings should be interpreted within the scope and limitations of those sources.

José C Bozelli Jr., PhD, is a scientist working at the intersection of biomarkers, omics, data science, and AI. He writes about health intelligence: how biological measurements become evidence, insight, and better decisions.

bozelli.ca

References

1. Luijken K, et al. Impact of predictor measurement heterogeneity across settings on the performance of prediction models: a measurement error perspective. Stat Med. 2019;38(18):3444–59. DOI: 10.1002/sim.8183

2. Khudyakov P, et al. The impact of covariate measurement error on risk prediction. Stat Med. 2015;34(15):2353–67. DOI: 10.1002/sim.6498

3. Agniel D, et al. Biases in electronic health record data due to processes within the healthcare system: retrospective observational study. BMJ. 2018;361:k1479. DOI: 10.1136/bmj.k1479

4. Zech JR, et al. Variable generalization performance of a deep learning model to detect pneumonia in chest radiographs: a cross-sectional study. PLoS Med. 2018;15(11):e1002683. DOI: 10.1371/journal.pmed.1002683

5. Finlayson SG, et al. The clinician and dataset shift in artificial intelligence. N Engl J Med. 2021;385(3):283–6. DOI: 10.1056/NEJMc2104626

6. Subasri V, et al. Detecting and remediating harmful data shifts for the responsible deployment of clinical AI models. JAMA Netw Open. 2025;8(6):e2513685. DOI: 10.1001/jamanetworkopen.2025.13685

7. Krause J, et al. Grader variability and the importance of reference standards for evaluating machine learning models for diabetic retinopathy. Ophthalmology. 2018;125(8):1264–72. DOI: 10.1016/j.ophtha.2018.01.034

8. Chavoshi M, et al. Impact of label noise from large language model-generated annotations on evaluation of diagnostic model performance. Radiol Artif Intell. 2026;8(2):e250477. DOI: 10.1148/ryai.250477

9. Leek JT, et al. Tackling the widespread and critical impact of batch effects in high-throughput data. Nat Rev Genet. 2010;11(10):733–9. DOI: 10.1038/nrg2825

10. Yu Y, et al. Correcting batch effects in large-scale multiomics studies using a reference-material-based ratio method. Genome Biol. 2023;24(1):201. DOI: 10.1186/s13059-023-03047-z

11. Han R, et al. Randomised controlled trials evaluating artificial intelligence in clinical practice: a scoping review. Lancet Digit Health. 2024;6(5):e367–73. DOI: 10.1016/S2589-7500(24)00047-5

12. Gianfrancesco MA, et al. Potential biases in machine learning algorithms using electronic health record data. JAMA Intern Med. 2018;178(11):1544–7. DOI: 10.1001/jamainternmed.2018.3763

13. Doherty C, et al. Keeping pace with wearables: a living umbrella review of systematic reviews evaluating the accuracy of consumer wearable technologies in health measurement. Sports Med. 2024;54(11):2907–26. DOI: 10.1007/s40279-024-02077-2

Author & Editorial Note: This article was conceived and editorially directed by the author, who defined its argument, audience, voice, narrative structure, and scientific framing. AI tools, including large language model assistants, supported literature discovery, synthesis, and citation checking, and generated draft prose within that direction. The author reviewed the evidence, edited the manuscript, and retains responsibility for the interpretation and final published text.