The science is sound. The pipeline surrounding it is not.
The science of molecular biomarkers is genuinely powerful. Mass spectrometry can now resolve thousands of proteins, lipids, and metabolites from a single blood draw. Multi-omics platforms can map the molecular state of a living system with precision that would have been unimaginable a decade ago. The tools work. The measurements are real. The problem is not that omics fails to measure biology — it’s that we routinely ask it to answer questions it hasn’t been validated to answer.
What isn’t working is the system that decides what those measurements mean — and what gets done with them.
When biomarker candidates don’t replicate, when clinical translation stalls, when fifty years of psychiatric biomarker investment produce no validated molecular diagnostic, the failure is not in the measurement. It is in how results are translated, which questions they are asked to answer, and whose interests are served by treating a discovery as a conclusion. The science is sound. The pipeline surrounding it is not.
The Pipeline That Rewards the Wrong Thing
Biomarker research runs through broadly three sequential stages: discovery, validation, and clinical translation. At each stage, the evidence requirement escalates. At each stage, the incentive structure of academic science pushes strongly in the opposite direction.
Grants, publications, promotions, and press coverage flow to discovery — finding the candidate, announcing the signal. The validation work that follows is slow, expensive, frequently negative, and published reluctantly by journals that prefer novelty to replication. This is not a moral failing. It is a design feature of how academic biomedical research is structured.
The consequences are measurable. Across fifty years of sustained investment in psychiatric biomarkers — spanning depression, schizophrenia, autism spectrum disorder, bipolar disorder, and post-traumatic stress disorder (PTSD) — the number of validated molecular biomarkers in routine clinical use is, by the accounting of a 206-citation systematic review in World Psychiatry, essentially zero.¹ Not a handful. Zero. A 2023 systematic review in The American Journal of Psychiatry catalogued 940 candidate biomarkers for autism spectrum disorder alone, across 280 studies — approximately 80% appearing in exactly one publication, with a median study enrollment of 64 participants.² Candidates examined in multiple studies produced “mostly inconsistent results.”
This is not a run of bad luck. It is the predictable output of a system designed to produce discoveries rather than to validate them.
The Instrument Doesn’t Know What Day It Is — But the Data Does
Before the incentive problem, there is a measurement problem. And it operates silently, inside the data.
Mass spectrometers are sensitive enough to detect disease. They are also sensitive enough to detect which day it is, which technician ran the samples, and whether the internal standard lot changed last Tuesday. They do not distinguish between these things. Every source of variance enters the measurement together.
Think of a microphone set up to record a lecture — sensitive enough to catch every word, but also every HVAC cycle and hallway conversation. If the room gets noisier on Tuesdays for reasons nobody tracked, any analysis of lecture content carries a ghost of “Tuesday-ness” through it. Add enough Tuesdays to one sample group and Thursdays to the other, and you have a classifier that looks like it’s predicting your biological outcome. It is predicting your collection schedule.
In omics terms, the Tuesdays are freeze-thaw cycle count, sample preparation timing, reagent lot number, and instrument signal drift. A 2024 review in Genome Biology describes these batch effects as “notoriously common” and notes they “may result in misleading outcomes if uncorrected or hinder biomedical discovery if over-corrected.”³
The problem runs deeper. Pre-analytical handling — what happens to your sample between collection and analysis — can be equally consequential. Targeted lipidomics profiling shows that storage temperature and processing time are among the strongest determinants of data integrity.⁴ Lipidomics studies show that temperature-related storage conditions systematically alter key phospholipid species — with phosphatidylcholines degrading via endogenous lipase activity under benchtop conditions — changes within the magnitude of the biological signals biomarker discovery studies are designed to detect.⁵
There is a second, counterintuitive problem. The standard response to batch effects is computational correction. This logic holds when the technical structure and the biological structure are independent. They frequently are not. In a typical case-control study, cases and controls don’t arrive evenly across time. Run a normalization algorithm that equalizes processing batches, and it simultaneously equalizes the disease signal riding alongside them. The biomarker “works” statistically because correction encoded the batch-disease correlation as biological signal.³
The correction doesn’t just fail to solve the problem.
It manufactures the artifact.
A larger discovery cohort with the same uncontrolled batch structure does not fix this. At scale, this means the dataset is learning the experiment — not the biology. It produces a better-powered version of the same noise.
The False Discovery Backlog
The combination — discovery incentives, inadequate validation, and uncontrolled measurement variance — produces what might be called a false discovery backlog: an accumulating inventory of candidate biomarkers that have cleared peer review but have never been subjected to independent external validation.
A 2024 meta-analysis of 244 clinical metabolomics studies found that 72% of the 2,206 metabolites reported as statistically significant were identified by only a single study, and that statistical modeling suggested 85% were likely noise.⁶ When external groups attempt to replicate published omics candidates using independent platforms and cohorts, the results are consistently sobering. A review of replication and validation practice across the omics era finds that false-positive associations routinely clear discovery phase peer review, and that the gap between a statistically significant finding and a generalizable one is rarely acknowledged at the discovery stage.⁷ In aging biomarker research — one of the most intensively funded omics domains — a 2024 analysis in Nature Medicine found that no consensus exists on how omics-based biomarkers should be validated before clinical translation, despite decades of candidate generation.⁸
The pattern — discovery outpacing validation, compounded by measurement and interpretation challenges — is visible across entire subfields. Extracellular vesicle (EV) research in psychiatry illustrates it with unusual clarity. First, isolation. Current methods struggle to distinguish exosomes from microvesicles — a limitation confirmed across multiple separation approaches in human plasma, where no fully pure EV preparation has been demonstrated due to overlap in size and density with other particles⁹. Second, attribution. Claims of brain origin often rely on a single surface marker whose specificity has been heavily questioned. Third, validation. Many reported correlations with disease have not been subjected to functional validation. A 2025 pre-registered systematic review of 46 such studies found that a significant proportion lacked adequate characterization of what had actually been isolated.¹⁰
The exosome is not an anomaly. It is the field’s failure modes in miniature.
The problem has a name in the statistics literature. When the prior probability of a true discovery is low, when sample sizes are small, and when publication bias favors positive results, a majority of statistically significant findings can be false positives — not because researchers are making errors, but because the system is operating exactly as designed.¹¹ More discovery without validation infrastructure doesn’t accelerate the path to clinical tools. It deepens a backlog that no amount of additional mining will clear.
A signal is not a biomarker until it survives validation.
Where the Real Advantage Lies
The asymmetric opportunity in this space is hiding in plain sight — and it is not discovery. The organizations that will own the next decade in omics-based diagnostics are not the ones with the largest candidate libraries. They are the ones building validation infrastructure while competitors are still congratulating themselves on finding signals. That gap, between discovery capacity and validation capacity, is currently large, measurable, and almost entirely ignored by standard due diligence. Here is what closing that gap actually requires.
Longitudinal within-person tracking sidesteps the core structural problem of cross-sectional biomarker studies, which must separate signal from enormous between-person biological noise. Within-person designs use each person as their own reference — revealing signal that population snapshots cannot. Pre-analytical and analytical standardization is, at this stage, a genuine competitive advantage. Randomizing sample processing across analytical batches, rather than processing all cases before all controls, closes the most common pathway to batch-disease confounding. Most organizations do not do this. Multi-omics cross-validation — proteomics against metabolomics against genomics — functions not as a data volume argument but as a structural cross-check against false discovery. Mechanisms are more durable than individual candidate biomarkers.
The FDA-NIH BEST framework defines a biomarker as “a characteristic that is objectively measured and evaluated as an indicator of normal biological processes, pathogenic processes, or responses to an intervention.”¹² That definition contains a word — validated — doing enormous work in the gap between what gets published and what gets used.
What a validated predictive biomarker looks like is ApoB. The evidence linking it to cardiovascular risk, across large prospective cohorts, independent replication, and explicit comparison against competing markers, is what the process of validation actually produces.¹³ It is not impossible. It is rare because the system has never required it.
The organizations treating validation as a core competency rather than a downstream concern are, right now, genuinely rare. That won’t remain true indefinitely. But it is true now, and the window for building in that space is open.
The next generation of winners in this space will not be defined by how many biomarkers they discover, but by how many survive independent validation. Discovery scales easily. Validation does not. That asymmetry is where the current mispricing sits — and where durable advantage is being built.
The same gap has a consumer-facing expression that is less visible but equally consequential. Every time someone opens a wellness biomarker report, they encounter a number that doesn’t say whether it’s measuring where they are right now or predicting where they’re going. Those are two different questions with two different evidence requirements — and the distinction almost never appears in the report itself. That is the subject of Part Two: written for the people at the other end of the pipeline this article describes.
José Carlos Bozelli Jr., PhD, is a lipid biochemist, omics data scientist, and scientific writer. He advises biotech, CRO, and health-tech teams on lipidomics, large-scale omics data pipelines, and biomarker science — and translates complex molecular data into decisions for scientists, clinicians, and builders.
References
Abi-Dargham A, et al. Candidate biomarkers in psychiatric disorders: state of the field. World Psychiatry. 2023;22(2):236–262.
Parellada M, et al. Biomarkers in autism spectrum disorder. American Journal of Psychiatry. 2023;180(4):273–372.
Yu Y, et al. Assessing and mitigating batch effects in large-scale omics studies. Genome Biology. 2024;25:254 (PMC11447944).
Sens A, et al. Pre-analytical sample handling standardisation for reliable measurement of metabolites and lipids in LC-MS-based clinical research. Journal of Mass Spectrometry and Advances in the Clinical Lab. 2023;9:39–52.
Reis GB, et al. Stability of lipids in plasma and serum: Effects of temperature-related storage conditions on the human lipidome. Journal of Mass Spectrometry and Advances in the Clinical Lab. 2021;22:100218.
Cochran D, et al. A reproducibility crisis for clinical metabolomics studies. Trends in Analytical Chemistry. 2024 (PMID: 40236582; PMC11999569).
Perng W, et al. Find the needle in the haystack, then find it again: replication and validation in the ‘omics era. Metabolites. 2020;10(7):E274 (PMC7408356).
Moqri M, et al. Validation of biomarkers of aging. Nature Medicine. 2024;30(2):360–372.
Dong L, et al. Comprehensive evaluation of methods for small extracellular vesicles separation from human plasma, urine and cell culture medium. Journal of Extracellular Vesicles. 2020;10(2):e12044.
Tunset ME, et al. Blood-borne extracellular vesicles in psychiatric disorders: systematic review. Journal of Psychiatric Research. 2025. PROSPERO CRD42021277534.
Ioannidis JPA. Why most published research findings are false. PLoS Medicine. 2005;2(8):e124.
FDA-NIH Biomarker Working Group. BEST (Biomarkers, EndpointS, and other Tools) Resource. Silver Spring, MD: FDA; 2016.
Sniderman AD, et al. A meta-analysis of low-density lipoprotein cholesterol, non-high-density lipoprotein cholesterol, and apolipoprotein B as markers of cardiovascular risk. Circ Cardiovasc Qual Outcomes. 2011;4(3):337–345.
