The Fiction of the Representative Sample: Why Clinical Trial Populations Fail Real Patients

I’ve spent twenty years watching the gap widen between what we prove in a trial and what we see in the clinic. It isn’t a gap of nuance. It’s a chasm built on exclusion, convenience, and a stubborn refusal to confront the messiness of actual human biology. When a phase III trial reports a 30% reduction in some composite endpoint, I don’t ask about the p-value. I ask who was allowed through the door in the first place. The answer, almost always, is a group of people who bear little resemblance to the patient sitting across from me—someone juggling five chronic conditions, a tangled medication list, and a life that refuses to fit into a case report form.

A diverse group of people standing together, illustrating the variety of patients seen in real clinical practice

The Architecture of Exclusion

Clinical trials aren’t built to reflect the population. They’re built to detect a signal. That distinction matters more than most researchers are willing to admit. To maximise the odds of finding a treatment effect, we strip away variability. We exclude the elderly, the frail, the multimorbid, the pregnant, the obese, anyone with renal or hepatic impairment, and anyone taking a medication that might tangle with the investigational product. By the time the protocol has finished its work, the remaining cohort is a carefully curated subset: younger, healthier, pharmacologically pristine. The trial then declares the drug effective in a population that doesn’t exist outside the academic medical centre.

Take oncology. The median age of a cancer diagnosis in high-income countries hovers around 66. Yet the median age of participants in oncology registration trials routinely sits under 60, often under 55. Patients over 75 are nearly invisible. Those with an ECOG performance status above 1—meaning they spend more than a trivial amount of time resting—are excluded by design. When the drug reaches the market, the oncologist must extrapolate from data generated in fit 50-somethings to the 78-year-old with diabetes, mild heart failure, and a GFR that’s been drifting downward for a decade. That extrapolation isn’t science. It’s hope dressed in a white coat.

The Comorbidity Blind Spot

Real patients accumulate conditions over time. The typical 70-year-old in primary care manages three or four chronic diseases at once. Each condition drags along its own pathophysiology, its own medications, and its own capacity to alter drug metabolism, receptor sensitivity, and baseline risk. Clinical trials treat comorbidity as a contaminant. Protocols list exclusion criteria that read like a catalogue of what ails the actual population: chronic kidney disease, liver disease, heart failure, COPD, autoimmune disorders, psychiatric illness. The result is a trial population in which the prevalence of multimorbidity is a fraction of what it is in the target population. Efficacy and safety data generated in these artificial cohorts tell us almost nothing about what will happen when the drug is prescribed to someone whose body is already a battleground of competing pathologies.

The problem compounds when you consider polypharmacy. An older adult taking eight or nine medications isn’t unusual. Those medications interact with each other and with any new agent we add. Trials don’t study these interactions because they exclude patients who take interacting drugs. Post-marketing surveillance then becomes a slow, passive experiment conducted on an unsuspecting public. We learn about harms years after approval, often through case reports and retrospective analyses that carry far less weight than the original randomised evidence.

An older adult holding a weekly pill organiser with multiple medications, representing polypharmacy in real-world patients

The Convenience Sample in Disguise

Recruitment drives everything. Trials have to enrol enough participants to meet statistical power requirements, and they have to do it fast because time is money. The easiest patients to enrol are the ones well-connected to academic centres, with flexible schedules, who speak the dominant language and trust the research enterprise. This produces a demographic skew that’s well-documented but rarely corrected. Racial and ethnic minorities, rural populations, people with low health literacy, and those without reliable transportation or childcare are systematically underrepresented. The data that emerge apply most directly to the people least likely to suffer the greatest burden of disease—a bitter irony the research community has grown comfortable ignoring.

Consider cardiovascular trials. For decades, the evidence base for statins, antihypertensives, and antiplatelet agents was built on cohorts that were overwhelmingly white and male. Women were enrolled in numbers too small to permit meaningful subgroup analyses. When sex-specific data finally accumulated, we discovered differences in drug metabolism, side-effect profiles, and even treatment efficacy. The same pattern repeats across therapeutic areas. We run trials on narrow populations, generalise the results to everyone, and then act surprised when real-world outcomes diverge.

The Geography Problem

Globalisation of clinical trials has added a new layer of distortion. Sponsors increasingly conduct trials in regions where recruitment is faster and costs are lower. Eastern Europe, Latin America, and parts of Asia now supply a disproportionate share of trial participants. These populations may differ from the intended treatment population in genetic background, diet, environmental exposures, and baseline disease epidemiology. A drug tested mostly in one region is then prescribed in another, with no systematic effort to understand whether the results travel across contexts. Regulatory agencies accept this as routine. I find it indefensible.

The statistical tools meant to bridge this gap are weak. Subgroup analyses are underpowered and prone to false positives and false negatives. Meta-analyses aggregate trials that all share the same recruitment biases. Real-world evidence studies try to fill the void, but they lack randomisation and are vulnerable to confounding that no amount of propensity-score adjustment can fully eliminate. We’ve built an evidentiary system that is internally consistent but externally hollow.

A scientist reviewing data on a computer screen, reflecting the analytical but disconnected nature of trial evidence

The Consequences of Ignoring Heterogeneity

When a drug moves from trial to market, the real experiment begins. Adverse events that were too rare to surface in a few thousand carefully selected participants show up in tens of thousands of unselected patients. Efficacy that looked strong in a homogeneous cohort weakens or vanishes in a heterogeneous one. Drugs get pulled, labels get slapped with black-box warnings, and clinicians are left to manage the uncertainty they were promised the trial would resolve.

One underappreciated mechanism is the ecological fallacy in dosing. Trials determine a dose that works in the average participant, then apply that dose to everyone. But the average participant doesn’t exist. Real patients vary in body mass, organ function, genetic polymorphisms in drug-metabolising enzymes, and receptor sensitivity. A fixed dose that’s fine for a 70-kilogram trial volunteer with normal renal function may do nothing for a 120-kilogram patient or poison someone with unsuspected CYP2D6 poor-metaboliser status. We ignore these differences because the trial structure forces us to. The alternative—designing trials that explicitly model heterogeneity—is more complex and expensive, and therefore rare.

What Would Honest Trials Look Like?

The fix isn’t a small tweak to eligibility criteria. It’s a fundamental redesign of the whole trial enterprise. Pragmatic trials, which enrol broad populations and embed themselves in routine care, offer a partial answer. They sacrifice some internal validity for external relevance, and that trade-off is worth making more often. Registry-based randomised trials, which use existing health data infrastructure to identify, randomise, and follow participants, can include the very patients that traditional trials exclude. Adaptive trial designs can incorporate evolving knowledge about heterogeneity without requiring a new trial for every subgroup.

Regulatory insistence on more representative enrolment is necessary but not enough. The FDA’s guidance on diversity plans is a step, but it doesn’t change the economic incentives that push recruitment toward the most accessible populations. Funders must demand external validity as a condition of support. Journals must require transparent reporting of who was screened, who was excluded, and how the enrolled cohort compares to the target population on dimensions that matter. Reviewers must stop accepting the phrase “further research is needed” as a substitute for methodological rigour at the point of evidence generation.

Frequently Asked Questions

Why are clinical trials so restrictive if researchers know it limits generalisability?

Restrictive criteria serve the goal of reducing noise. Every source of variability—comorbidity, concomitant medications, age-related physiological changes—increases the variance of the outcome measure and makes it harder to spot a treatment effect. Researchers and sponsors optimise for a statistically significant result within a feasible budget and timeline. External validity is a secondary concern, often punted to post-marketing studies that may never happen. The incentive structure rewards clean answers over relevant ones.

Do all drugs show different effects in real-world populations compared to trial populations?

Not all, but the exceptions don’t excuse the rule. Some treatments have large effect sizes that survive the transition to unselected populations. Others have such a narrow therapeutic window that the trial-to-clinic gap becomes dangerous. The problem is we can’t predict which drugs will diverge until after they’re widely prescribed. A more systematic approach to studying heterogeneity during development would flag the drugs that need tighter prescribing guidance before they hit the market, instead of leaning on post-market surveillance as a safety net.

What can patients and clinicians do with the evidence we currently have?

Read the eligibility table of a trial before you read the results. Ask whether the participants resemble the person sitting in front of you. If the trial excluded everyone with kidney disease and your patient has an eGFR of 45, the evidence doesn’t apply directly. That doesn’t mean the drug can’t be used; it means the decision carries more uncertainty than the guideline admits. Demand that guideline developers and drug regulators make the limits of generalisability explicit. Clinical judgment means knowing when the evidence stops and the guesswork begins.

A Closing Refusal

I refuse to pretend that a trial of 3,000 handpicked participants tells me what I need to know about the millions who will swallow the pill. The research enterprise has spent decades perfecting a method that answers a narrow question with precision while ignoring the broader question that actually matters. Until we redesign trials to match the complexity of the patients we treat, we are not practising evidence-based medicine. We are practising convenience-based medicine, and the difference costs lives.