
I’m done pretending the person across from me matches the tidy profiles in a landmark trial. After twenty years in practice and research, I can count on one hand the patients who actually resembled those sanitized, optimized subjects. This isn’t a gripe about statistics. It’s a frustration with a system that erects cathedral-like evidence on a sand foundation, then expects us to apply it to flesh-and-blood people who are built from entirely different stuff.
Randomized controlled trials remain our least imperfect tool for nailing down causality. Nobody serious disputes their internal logic. But the question that shadows every prescription I write is simpler and more stubborn: what happens when you take a drug that shone in a pristine, artificial hothouse and release it into the uncontrolled, comorbid, polypharmacy jungle of actual human lives? Too often, we don’t have a clue. And acting like we do isn’t scientific caution—it’s a collective dodge.
The Architecture of Exclusion
Trial eligibility criteria aren’t neutral instruments for framing a research question. They’re sieves, and they methodically filter out the very patients who will end up getting the treatment. A 2015 BMJ Open analysis looked at 283 trials backing 189 new drug applications. The median number of exclusion criteria was 19. Typical reasons for exclusion—liver trouble, poor kidney function, psychiatric comorbidity, multiple concurrent medications—happen to define a hefty chunk of the population living with the disease under study.
Take heart failure. The foundational sacubitril/valsartan trials enrolled patients with a mean age of around 64, and they insisted on a run-in period to filter for tolerability. In my clinic, the average heart failure patient is 78, juggles 12 medications, has stage 3 chronic kidney disease, and would have been kicked out of that trial on at least three counts. I’m not attacking the drug. I’m attacking the lazy leap that says those trial results transfer cleanly to my patient. They don’t.

Age as a Systematic Blind Spot
Older adults swallow more medications than anyone else on the planet, yet they remain the least studied. A 2019 systematic review in JAMA Internal Medicine found that among RCTs in high-impact journals, over 40% explicitly barred patients above a certain age, and the average participant was often 10 to 20 years younger than the median age of the disease population. When I hand an 85-year-old with moderate dementia, sarcopenia, and orthostatic hypotension a drug tested in spry 60-year-olds with no cognitive slips, I’m not doing evidence-based medicine. I’m doing faith-based extrapolation.
The Renal and Hepatic Exclusion Habit
Organ dysfunction isn’t some exotic comorbidity; it’s the wallpaper of chronic disease. Yet trial protocols casually exclude anyone whose creatinine clearance dips below 30 mL/min or whose liver enzymes drift above twice the upper limit of normal. These routine exclusions leave a knowledge vacuum, and we clinicians fill it with guesswork. Dose adjustments for renal impairment often come from tiny pharmacokinetic studies—single-dose, small cohorts—not from real outcome data. The result? We treat some of our most fragile patients with dosing schemes that are essentially experimental, minus the ethical guardrails of a proper trial.
The Polypharmacy Paradox
Real patients pile up medications. Trial patients, by design, take as few as possible. Concomitant drug restrictions are standard, justified by the urge to isolate the investigational product’s signal. But that isolation is a scientific fantasy. Drug-drug interactions aren’t noise; they’re the main channel. When a shiny new anticoagulant gets tested in patients not touching amiodarone, verapamil, or strong CYP3A4 inhibitors, we learn zip about how it’ll behave in the actual atrial fibrillation crowd, where those drugs are everywhere. The trial answers a question no working clinician ever asked.
This paradox bites hardest in psychiatry and neurology, where polypharmacy is the everyday norm. Trials for atypical antipsychotics, for instance, usually exclude anyone on more than one psychotropic drug. But the real-world patients who get these prescriptions are often on three or four. Extrapolating efficacy and safety from squeaky-clean monotherapy trials to polypharmacy reality is a methodological leap that would earn a failing grade in any undergraduate science course.
Comorbidity: The Rule, Not the Exception
Multimorbidity is the signature health challenge of this century. Over two-thirds of people older than 65 live with two or more chronic conditions. Yet clinical trials keep treating diseases like they exist in separate glass boxes. A diabetes trial excludes heart failure patients. A COPD trial shuts out anyone with an anxiety disorder. These exclusions make a certain internal sense for the trial’s validity. But they produce evidence that is structurally useless for guiding care in the multimorbid patient.
When I treat a person who drags diabetes, coronary artery disease, depression, and osteoarthritis into the exam room, I’m not facing four separate diseases. I’m facing one snarled, interacting system. The evidence base, shattered into disease-specific silos, offers me no coherent map. Guidelines quarrel. Drug interactions pile up. The patient becomes a walking contradiction of the entire clinical trial enterprise.

The Run-In Period: A Trial Design That Erases Reality
Run-in phases are a methodological sleight of hand that warps generalizability. Here’s how it works: everyone starts on placebo or active treatment, and those who can’t stick with it, can’t tolerate it, or show an early response get tossed out before randomization. What’s left is a study population enriched for adherence and tolerability. The shiny effect size and tame adverse event profile you read about reflect that artificial enrichment.
When a new biologic for rheumatoid arthritis boasts a 70% response rate and whisper-quiet side effects, I need to know: 70% of whom? If 40% of the originally screened patients were washed out during the run-in, that number is deeply slippery. The drug will behave differently—often worse—when it hits an unselected clinic population. Run-in periods aren’t lies. But they get misread constantly, and the resulting overestimate of benefit plus underestimate of harm is a straight shot at patient safety.
Race, Ethnicity, and the Geography of Evidence
Clinical trial populations stay stubbornly unrepresentative of the global crowd that swallows the pills. A 2020 JAMA Network Open analysis showed that in trials supporting FDA drug approvals, Black participants were underrepresented relative to disease burden in almost every therapeutic area. Hispanic, Asian, and Indigenous representation is often worse. Genetic variation in drug metabolism—think CYP2C19 variants scrambling clopidogrel activation, or HLA-B*1502 and carbamazepine hypersensitivity—means efficacy and safety can swing wildly across populations. When trials run mostly in white European cohorts, we’re exporting uncertainty to the rest of the world.
This isn’t just a social justice talking point. It’s a basic pharmacology failure. Drugs get approved on data from a narrow genetic and environmental slice of humanity, then prescribed globally as if human biology were a flat, uniform surface. It isn’t.
Why This Persists: The Structural Incentives
No single villain is responsible for the phantom trial patient. The incentives are baked in. Sponsors hunger for clean data to smooth the regulatory path. Regulators want clear efficacy signals unclouded by comorbidity noise. Investigators need feasible recruitment. The result is a convenience collusion that churns out tidy evidence unfit for messy human beings.
Pragmatic trials—embedded in routine care, with wide eligibility gates—offer a partial antidote but remain a sliver of funded research. They’re harder to design, harder to bankroll, and harder to publish in journals that worship internal validity over external relevance. Until funding bodies and journals demand representativeness with the same ferocity they demand randomization, the phantom patient will keep rattling around our guidelines.
What the Clinician Must Do Now
I don’t have the luxury of waiting for the evidence base to be rebuilt brick by brick. Every day I face decisions for patients the trials never imagined. That demands a different kind of reasoning—one that treats trial results not as final answers but as wobbly starting points. I ask three questions with every major treatment decision:
First, who was actually studied? I skip the abstract. I go straight to the baseline characteristics table. If the mean age is 20 years younger than my patient, or if the exclusion list mirrors my patient’s chart, I downgrade my confidence in the applicability of the results.
Second, what was the absolute benefit? Relative risk reductions are seductive and often slippery. I calculate the number needed to treat from the control event rate and ask whether that absolute benefit justifies the treatment burden for this particular person, given their life expectancy and what matters to them.
Third, what’s the time horizon of harm? Trials are short. Most drugs for chronic conditions get tested over months to a couple of years. My patient may swallow the drug for decades. The harms that pile up over ten years—cognitive dulling, renal creep, cumulative toxicity—are invisible in a 52-week snapshot. I assume long-term safety data don’t exist, because usually they don’t.
Frequently Asked Questions
Why can’t researchers just include more diverse patients in trials?
They can, but the incentives tug the other way. Recruiting patients with multiple comorbidities is slower, costlier, and yields messier data—more adverse events, smaller apparent treatment effects. Sponsors, who bankroll most trials, aren’t rewarded for messy data. Regulators have started demanding diversity plans, but enforcement remains flimsy, and the definition of “diverse” often stops at race and gender, ignoring age, comorbidity, and polypharmacy.
Are pragmatic trials the answer?
They’re part of it. Pragmatic trials embedded in electronic health records can rope in broader populations and measure outcomes that actually matter to patients and health systems. But pragmatic designs sacrifice some internal validity, and they’re tough to blind, which lets bias creep in. They aren’t a replacement for explanatory RCTs—they’re a needed complement. The real fix is a portfolio of evidence types: RCTs, pragmatic trials, observational studies, n-of-1 trials—triangulated to build a fuller picture.
How should patients assess whether a trial applies to them?
Patients can ask their clinicians the same questions I laid out: Who was studied? What was the absolute benefit? What do we actually know about long-term safety? Beyond that, I’d tell patients to raise an eyebrow at any news report that trumpets a relative risk reduction without context. A 30% drop in heart attacks sounds dramatic; if the baseline risk was 3% over five years, the absolute reduction is less than 1%. That might still be worthwhile, but the decision belongs to the patient, armed with evidence placed in honest context.
Does this mean clinical trials are useless?
Not at all. The randomized trial is an indispensable tool. The problem isn’t the tool—it’s the arrogance with which we stretch its findings. A trial is a model, and as the statistician George Box put it, all models are wrong, but some are useful. Danger sets in when we mistake the model for the territory. Clinical trials are useful exactly to the extent that we grasp their limits and refuse to apply them blindly to patients they were never designed to mirror.
The phantom patient is a construct, but the harms that flow from treating real patients like phantoms are solid enough. Adverse drug reactions, therapeutic failures, wasted resources, and frayed trust—that’s the price of our collective refusal to demand evidence that reflects the world as it is, not as we wish it to be. Next time you read a trial, don’t fixate on whether the result was statistically significant. Ask whether the patients in that trial bear any resemblance to the person you’re trying to help. If they don’t, the p-value is the least of your worries.