The four questions above are this series' throughline: Part One explained why omics rewards discovery over validation, Part Two split biomarkers into "where you are" versus "where you're going," Part Three showed what actually clearing that bar looks like, and Part Four turned the same lens on your own data. This piece runs the same checklist against what's actually for sale: panels, CGMs, microbiome kits, age clocks, hormone tests, food-sensitivity panels, cancer screens. The data is neutral. Which path it lands on depends on whether anyone asked it these four questions first. Most haven't.
If you’re building a biotech or health-tech company in this space, treating patients through a functional or naturopathic lens, or just trying to decide whether your last $300 panel was worth it, you’re evaluating the same market through three different lenses, and almost nothing written about it speaks to all three at once. This piece audits seven categories of biomarker tests currently being sold across that market: what a founder should know before building here, what a clinician should know before recommending one, and what anyone paying for one should know before they buy.
For more than a decade, allergy and immunology societies in North America and Europe have said the same thing, in nearly the same words, about one of the most widely sold tests in wellness medicine. Food-specific IgG panels do not diagnose sensitivity, intolerance, or allergy to anything. The antibody they measure is not a sign of harm. It is a normal, expected response to food a person has safely eaten many times before, and its presence is, if anything, a marker of tolerance rather than its opposite. The American Academy of Allergy, Asthma & Immunology has formally endorsed this position.¹ The Canadian Society of Allergy and Clinical Immunology² and the European Academy of Allergy and Clinical Immunology³ have each published nearly identical conclusions. And the test remains easy to find, easy to order, and easy to build an elimination diet around.
This is not really a story about a bad test. It’s a story about a measurement that is mostly real — IgG antibodies to food genuinely exist, and the assay genuinely detects them — being asked to answer a question it was never built to answer. That gap, between what a test measures and what it’s sold as proving, is the single most useful lens for evaluating the consumer biomarker landscape in 2026, because it shows up almost everywhere, including in products built on real, often excellent, science. Most of what’s for sale right now isn’t fake. It’s just frequently answering a different question than the one on the label.
Four questions, asked plainly
Before spending money on any biomarker test, a panel, a kit, a subscription, four questions do almost all of the evaluative work, and almost no marketing page answers them directly. Does it say plainly whether it’s measuring what’s happening in you right now, or estimating what might happen to you later? Has anyone shown, independently, that the measurement gives the same answer at least thrice? Was it built to track how you change over time, or only to describe where you sit relative to other people? And does the report in your hands tell you, clearly, what it doesn’t know?
Apply those four questions to what’s actually being sold right now, and a real pattern shows up — not a verdict on an industry, but a map of where the science has landed, where it’s been stretched past what it can support, and where the test is running ahead of its own proof. Here’s that map, built from the categories people are actually buying and talking about.
The panel that bundles everything together
Among the most visible and heavily marketed categories in consumer testing is the comprehensive at-home panel: one blood draw, sometimes more than a hundred markers, delivered as a single report and sold as a single act of self-knowledge. That packaging is the category’s biggest evidentiary problem, because a hundred markers don’t sit at one evidence tier. They’re scattered across all four, inside the same PDF.
Apolipoprotein B is the clearest case of a marker in these panels that has genuinely arrived. A major 2026 update to U.S. cardiology guidelines now recommends apoB be used to guide treatment decisions in cardiovascular risk management, and recommends that every adult have Lp(a), a related, largely genetically fixed lipoprotein, measured at least once in their lifetime.⁴ Genetic and epidemiological evidence reviewed in JAMA Cardiology supports a likely causal relationship between elevated Lp(a) and atherosclerotic disease, which is part of why guidance has shifted toward recommending it.⁵ That’s not wellness-industry marketing. It’s the same evidentiary bar that moved ceramide testing into clinical labs, now clearing a second marker. But the same panel that includes apoB and Lp(a) often bundles them alongside proprietary wellness scores, inflammation composites, or hormone ratios with no comparable guideline behind them, and the report rarely tells you which is which. The product passes the first question for some of its rows and fails it for others, in the same document.
Continuous glucose, discontinuous evidence
Continuous glucose monitoring (CGM) is probably the single most talked-about tool among healthy people optimizing their metabolism, and as a functional readout, it earns the attention. Watching your own glucose rise and fall after a specific meal is a real, reproducible biological signal. The trouble starts when that real-time signal gets read as a predictor of future metabolic disease in someone who doesn’t have diabetes. A 2026 study in Nature Communications, tracking continuous glucose data across more than 3,600 people without diabetes, found that the metrics did capture real components of metabolic physiology, but its own authors concluded that a longer-term evaluation of actual health outcomes would be needed before the tool’s value for managing health in this population could be established.⁶ A separate 2026 systematic review of CGM use in non-diabetic populations found genuine improvements in average glucose and adherence to interventions, but concluded that evidence on measurement accuracy and meaningful variability remains limited.⁷ And researchers at Mass General Brigham, comparing CGM output to standard blood markers, found the device tracked actual glucose control closely in people with diabetes, but lost that correspondence in people without it, which is exactly where most of the predictive marketing is aimed.⁸ The tool works. The story being told about what the line on the app means for someone with normal glucose tolerance is the part still unproven.
A map of the gut, not a meal plan
Stool-based microbiome sequencing paired with personalized diet or supplement recommendations remains one of the most popular categories in consumer wellness testing, and the underlying biology isn’t in question. A landmark study published in Cell showed that an algorithm combining gut microbiome composition with diet, activity, and blood markers could predict an individual’s post-meal blood sugar response far better than standard nutritional advice, and a follow-up randomized trial confirmed that diets built on those predictions improved glucose control in practice.⁹ Gut microbial composition is real, measurable, and does interact with diet in ways that can be modeled. What’s typically missing in the commercial version of this idea is the rigor of that original study: continuous data collection, a validated predictive algorithm, and a randomized intervention testing whether the prediction actually changes outcomes. In 2025, an international, multidisciplinary panel of physicians and researchers published a consensus statement in The Lancet Gastroenterology & Hepatology concluding that current evidence is insufficient to widely recommend routine microbiome testing in clinical practice, and warning that the gap between commercial enthusiasm and proven clinical value risks wasting both money and trust.¹⁰ The test measures something real. The one-time personalized meal plan built on top of a single stool sample is, for now, closer to an informed guess than the validated intervention the underlying science has actually shown is possible.
A real number, an unclear instruction
Biological age testing, overwhelmingly dominated by DNA-methylation-based “epigenetic clocks,” is now close to a billion-dollar category on its own, and the population-level science behind it is genuinely strong, and still actively evolving at the highest levels of the field. A 2024 study in Nature Aging found that most existing epigenetic clocks aren’t actually built from the DNA sites most plausibly causal for aging, and proposed a new framework to better separate damaging from adaptive methylation changes, a sign that even the foundational measurement is still being refined by the researchers who built it.¹¹ What a 2025 commentary in the AMA Journal of Ethics flagged, working through an actual consumer case, is the gap one level down. A result can be personally meaningful to someone without being clinically actionable, and the commentary warns that this gap carries real risk of patient misunderstanding and psychological harm, urging companies to be more transparent about how reliable any single person’s number really is.¹² That concern has technical teeth. A 2025 preprint evaluating the reliability of eighteen epigenetic clocks found that some shifted meaningfully depending on lab protocol alone, with no change at all in the person being measured.¹³ The aging signal is real. What a healthy 45-year-old should do with the number, and which clock’s number to trust, is the part the field hasn’t agreed on yet.
Measuring something real, proving something else
Dried-urine and saliva-based hormone metabolite testing has become the signature diagnostic tool of functional and integrative medicine, and it deserves a more careful hearing than a quick dismissal. The underlying chemistry is sound: a peer-reviewed validation study found dried urine measurements of seventeen reproductive hormones and metabolites agreed closely with traditional liquid and 24-hour urine collection, the existing gold standard.¹⁴ Worth noting, though, is that this validation study was designed and authored by the scientists who created the test, not by an independent research group, a pattern that holds across most of the published evidence for this category. None of these tests currently carry FDA clearance as a diagnostic device. The measurement itself is probably sound. Whether a detailed hormone metabolite profile reliably leads to a better treatment decision than a simpler test would is a separate question, and it hasn’t yet been answered by the kind of independent, multi-site outcome research that moved ceramide testing into clinical use.
Ahead of the infrastructure
Blood-based multi-cancer early detection tests are the highest-stakes example in this entire landscape, because the claim being made, find cancer before it has symptoms, across dozens of types, from one blood draw, is the most consequential thing any consumer test promises. The clearest published numbers come from a large prospective study in The Lancet Oncology, which tested a methylation-based detection test in more than 5,400 patients referred for investigation of possible cancer symptoms.¹⁵ Specificity was excellent, above 98%. Sensitivity told a more complicated story: about 24% in stage I disease, climbing to roughly 95% by stage IV. That’s the central tension of this whole category in one statistic: the test is genuinely good at finding cancer that has already progressed, and considerably less good at the early-stage detection that’s the entire premise of the category, and the study itself was funded by the test’s own developer. What hasn’t moved as fast as the underlying science is the infrastructure around the claim. A symposium convened in September 2025 to discuss the path forward for these tests, and the consensus among the roughly two hundred researchers, clinicians, and public health officials in attendance, organized by Fred Hutch, was that the technology is not yet ready for wide deployment.¹⁶ This is the category where “coming up hot” is the most literal description available. The technology is real, the validation is incomplete, and the gap between the two carries the highest consequences of anything on this list.
What seven categories have in common
Run the same four questions across these categories and a consistent pattern emerges: most underlying measurements, apoB, glucose, DNA methylation, gut microbial composition, hormone metabolites, cell-free DNA, are scientifically defensible. The market’s story problem is usually a labeling failure, where a real number is used to answer a question, such as predicting future disease or guiding treatment, that it was never validated to answer. Only one category on this list, food-specific IgG, represents an interpretive failure rather than a technical one: the assay accurately measures exposure to food proteins, but it lacks a validated physiological link to the clinical intolerance it is sold to detect.¹,²,³ Every other failure here is a labeling failure, not a measurement failure.
None of this requires trusting any single company, or distrusting all of them. It requires the same four questions, asked consistently, whether you’re deciding what to build, what to recommend, or what to buy. A founder deciding where to invest in rigor, a clinician deciding what to keep ordering, and a person paying for the next panel are all asking a version of the same thing, and the answer doesn’t change depending on which chair you’re sitting in. That’s not skepticism. It’s literacy, and it scales across all three.
It also points toward where the real gap in this market sits. Almost everything reviewed here is built to answer where you are: a snapshot, a single draw, a single report. The harder, slower, and more valuable question is where you’re going, and that question doesn’t have a good answer yet in most of these categories. That’s probably exactly why it’s still unanswered.
Scientific note: This article cites peer-reviewed evidence, clinical guidelines, and reported professional consensus throughout. Sources are independently verifiable and listed below. Claims about individual tests are scoped to the published literature cited and do not constitute medical advice.
José Carlos Bozelli Jr., PhD, is a lipid biochemist, omics data scientist, and scientific writer. He advises biotech, CRO, and health-tech teams on lipidomics, large-scale omics data pipelines, and biomarker science, and translates complex molecular data into decisions for scientists, clinicians, and builders.
The content of this article is for informational and educational purposes only and does not constitute medical advice. Consult a qualified healthcare professional before making decisions based on biomarker results.
References
Bock SA. AAAAI support of the EAACI Position Paper on IgG4. J Allergy Clin Immunol. 2010;125(6):1410. DOI: 10.1016/j.jaci.2010.03.013.
Carr S, Chan E, Lavine E, Moote W. CSACI position statement on the testing of food-specific IgG. Allergy Asthma Clin Immunol. 2012;8:12. DOI: 10.1186/1710-1492-8-12.
Stapel SO, Asero R, Ballmer-Weber BK, et al; EAACI Task Force. Testing for IgG4 against foods is not recommended as a diagnostic tool: EAACI Task Force Report. Allergy. 2008;63(7):793–796. DOI: 10.1111/j.1398-9995.2008.01705.x.
Blumenthal RS, Morris PB, Gaudino M, et al. 2026 ACC/AHA/AACVPR/ABC/ACPM/ADA/AGS/APhA/ASPC/NLA/PCNA Guideline on the Management of Dyslipidemia: A Report of the American College of Cardiology/American Heart Association Joint Committee on Clinical Practice Guidelines. Circulation. 2026;153:e1154–e1276. DOI: 10.1161/CIR.0000000000001423.
Duarte Lau F, Giugliano RP. Lipoprotein(a) and its significance in cardiovascular disease: a review. JAMA Cardiol. 2022;7(7):760–769. DOI: 10.1001/jamacardio.2022.0987.
Bermingham KM, Smith HA, Duncan EL, et al. Associations of continuous glucose monitor derived time in range and glycaemic variability with diet, lifestyle, and demographics. Nat Commun. 2026;17:4496. DOI: 10.1038/s41467-026-70308-3.
Liao X, Li Y, Tang S, et al. Continuous glucose monitoring in non-diabetic populations: a systematic review of observational and interventional studies with meta-analysis. Eur J Med Res. 2026;31:397. DOI: 10.1186/s40001-026-03920-0.
Mass General Brigham. For people without diabetes, continuous glucose monitors may not accurately reflect blood sugar control. Press release, October 2025.
Zeevi D, Korem T, Zmora N, et al. Personalized nutrition by prediction of glycemic responses. Cell. 2015;163(5):1079–1094. DOI: 10.1016/j.cell.2015.11.001.
Porcari S, Mullish BH, Asnicar F, et al. International consensus statement on microbiome testing in clinical practice. Lancet Gastroenterol Hepatol. 2025;10(2):154–167. DOI: 10.1016/S2468-1253(24)00311-X.
Ying K, Liu H, Tarkhov AE, et al. Causality-enriched epigenetic age uncouples damage and adaptation. Nat Aging. 2024;4(2):231–246. DOI: 10.1038/s43587-023-00557-0.
Hauskeller M, Shore L. What are the most ethically salient implications of epigenetic age testing? AMA J Ethics. 2025;27(12):E828–833. DOI: 10.1001/amajethics.2025.828.
Sehgal R, Borrus D, Gonzalez J, et al. Biological versus technical reliability of epigenetic clocks and implications for disease prognosis and intervention response. bioRxiv. 2025. DOI: 10.1101/2025.10.13.682176. Preprint, not yet peer-reviewed.
Newman M, Curran DA. Reliability of a dried urine test for comprehensive assessment of urine hormones and metabolites. BMC Chem. 2021;15(1):18. DOI: 10.1186/s13065-021-00744-3.
Nicholson BD, Oke J, Virdee PS, et al. Multi-cancer early detection test in symptomatic patients referred for cancer investigation in England and Wales (SYMPLIFY): a large-scale, observational cohort study. Lancet Oncol. 2023;24(7):733–743. DOI: 10.1016/S1470-2045(23)00277-2.
Mapes D. Are we ready for multi-cancer detection tests? Fred Hutch News Service, September 22, 2025.
