Skip to main content
  1. Resources/
  2. Reports/

Animal Research Failures

18315 words·86 mins
Table of Contents

Animal Research Failures
#

road fork
Each fork in the road leads to a pre-defined destination
Credit: Gemini

Introduction
#

Animal research has long been a cornerstone of biomedical science, yet its failures have had devastating clinical, regulatory, and economic consequences. Despite rigorous testing in multiple animal species, many drugs and treatments that appeared safe and effective in animals have caused severe harm or death in humans. These failures underscore fundamental limitations in translating animal data to human outcomes, exposing critical gaps in preclinical research methodologies.

This report provides an in-depth synthesis of major animal research failures across clinical, regulatory, and economic dimensions, their root causes, regulatory evolution, ethical implications, and actionable recommendations to improve scientific rigor and human relevance in biomedical research.

[Editorial Note: Of possible historical significance is the circa 2000 ADAV report on this same topic: Do No Harm.]

Case Studies of Major Animal Research Failures
#

The translation of therapeutic candidates from preclinical safety evaluations to successful human clinical outcomes represents one of the most significant challenges in modern drug discovery and toxicology. For nearly a century, the standard paradigm of biomedical research has operated under the assumption that non-human mammalian models provide an essential, high-fidelity approximation of human physiology. This structural reliance has been codified into regulatory frameworks globally, requiring drug sponsors and chemical manufacturers to submit safety data from multiple animal species before initiating human trials.

However, systematic retrospective evaluations of this paradigm reveal a persistent translational deficit. Over 92% of therapeutic candidates that successfully clear preclinical animal-based safety and efficacy screenings subsequently fail when progressed into human clinical trials. This high attrition rate is primarily driven by unexpected human clinical toxicities that were undetected in animal testing, or by a complete lack of therapeutic efficacy in human patient populations.

Preclinical to Clinical Pipeline Preclinical to Clinical Pipeline
100 Preclinical Candidates 40 Eliminated (Pre-human)
60 Advance to Human Trials 54 Fail in Phases I-III
6 Approved for Clinical Use Only 6% Overall Success

An analysis of this developmental pipeline shows that while approximately 40% of potential drug candidates are eliminated during preclinical animal tests, the remaining 60% that enter clinical phases face a 90% failure rate. Within this clinical attrition envelope, approximately 40%–50% of candidates fail due to a lack of therapeutic efficacy at clinically tolerable doses, 25%–30% fail due to unmanageable clinical toxicities, and 10%–15% fail due to poor human absorption, distribution, metabolism, and excretion (ADME) profiles. These metrics show that traditional animal models frequently act as unreliable filters, introducing both false positives—which expose human volunteers to unanticipated clinical hazards—and false negatives, which can cause potentially therapeutic compounds to be discarded early in development. See Animal Research vs NAM Expenditures for further details.

The analyses indicates that the systemic reliance on non-human animal models has delayed biomedical progress, and shows that transitioning to human-biology-based New Approach Methodologies (NAMs) offers a more predictive and economically viable path forward.

The structure of each case study is:

  • Background: brief introduction to the topic
  • The Failure: the problems humans experienced from the drug
  • The Delay: adherence to the animal model prevented progress
  • The Progress: improvements via human-relevant science
  • Key Takeaways: statement regarding poor species translatability
  • Addendum: additional explanations providing greater detail

Thalidomide
#

Background: Thalidomide was introduced in the late 1950s as a sedative and anti-nausea medication, specifically marketed to pregnant women suffering from morning sickness1. Developed by the German pharmaceutical company Grunenthal, it was initially promoted as a safe alternative to barbiturates, which carried a higher risk of overdose2. Despite undergoing extensive animal testing across multiple species - including rodents, rabbits, dogs, hamsters, and primates - no significant teratogenic effects or maternal-fetal toxicities were observed1 3 4. This lack of adverse findings in animal studies led developers to declare thalidomide as exceptionally safe, even for use during pregnancy3.

The Failure: Thalidomide caused phocomelia (severe limb malformations) and other congenital defects, including absence of ears, shortened or absent arms and legs, and internal organ malformations, in over 10,000 infants worldwide before its withdrawal1 4. The most striking aspect of this disaster was that animal tests had completely failed to predict this severe teratogenicity, even when tested at doses far exceeding typical human exposure levels2. Subsequent attempts to replicate these teratogenic effects in pregnant rodents also failed under standard dosing conditions, revealing a catastrophic translational failure of preclinical testing4. The drug had been evaluated in approximately 10 strains of rats, 11 breeds of rabbits, 2 breeds of dogs, 3 strains of hamsters, and 8 species of primates, yet traditional preclinical protocols still failed to predict the human clinical risk2.

The Delay: The drug was withdrawn in November 1961 after birth defects became evident5, but the damage was irreversible, affecting thousands of children and families across 46 countries1 6. While alternative, human-relevant methodologies were not fully developed in the late 1950s, human in vitro tissue culture techniques and embryonic cell models were emerging7. These were largely ignored due to the regulatory focus on in vivo mammalian safety data1. This reliance on rodent models delayed the identification of thalidomide’s mechanism of action for decades, with the precise molecular pathways not fully characterized until the 21st century8.

The Progress: The disaster led to the implementation of stricter global regulations for reproductive and developmental drug testing, including mandatory teratogenicity testing in multiple species1. In the United States, Dr. Frances Kelsey’s refusal to approve thalidomide without further testing prevented widespread use and saved countless lives6. Modern human-centric cell biology and molecular assays eventually explained the mechanisms of thalidomide-induced teratogenesis, showing that human in vitro embryonic stem-cell assays can successfully predict these toxicities9. These human-focused methods have shown that thalidomide induces apoptosis in human embryonic fibroblasts10 while failing to do so in rodent embryonic cells (due to species-specific antioxidant defenses3 11), providing a predictive accuracy that standard animal models could not achieve12.

Key Takeaways: Animal models failed to predict human-specific developmental toxicity due to species differences in embryological development, antioxidant defense systems, and drug metabolism, underscoring the need for human-relevant testing methods1 2 13.

Addendum: The thalidomide disaster highlighted several species-specific physiological and molecular differences.

  • Antioxidant Defense System Divergence: Rodent embryonic cells, particularly mouse cells, often have greater glutathione-dependent antioxidant capacity than cells from thalidomide-sensitive species. Thalidomide-induced oxidative stress and glutathione depletion have been observed preferentially in sensitive species, and experimentally lowering glutathione can increase sensitivity in mouse cells11.

  • Pharmacokinetic and Embryonic Half-Life Disparities: Thalidomide is eliminated much faster in mice than in humans: approximately 0.5 hours in mice versus roughly 7.3 hours in humans in commonly cited pharmacokinetic comparisons. This extended exposure window in humans allows the compound to accumulate and exert prolonged teratogenic effect14.

TGN1412
#

Background: TGN1412, a superagonistic monoclonal antibody targeting CD28 for autoimmune diseases and rheumatoid arthritis, was developed by TeGenero Immunotherapeutics and extensively tested in cynomolgus and rhesus monkeys at doses up to 500 mg/kg, with no adverse effects observed1. CD28 is a co-stimulatory receptor on T cells, and TGN1412 was designed to selectively activate regulatory T cells while sparing effector T cells, a mechanism intended to suppress autoimmune responses without broad immunosuppression15.

The Failure: In a phase I human trial conducted in March 2006 at Northwick Park Hospital, London, six healthy male volunteers received TGN1412 at 1/500th the dose deemed safe in monkeys (0.1 mg/kg). Within minutes, all six suffered a catastrophic cytokine release syndrome (CRS), characterized by systemic organ failure, severe inflammation, and life-threatening cardiovascular collapse1 15. The volunteers experienced headache, myalgia, nausea, diarrhea, erythema, hypotension, and respiratory distress, with long-term complications including finger and toe necrosis16. The cytokine storm involved massive release of pro-inflammatory cytokines, including TNF-α, IFN-γ, IL-2, IL-6, and IL-8, which were not predicted by preclinical models17.

The Delay: The trial was immediately halted after the first volunteer received the drug, but the damage was irreversible, with all six volunteers requiring intensive care1. Standard human peripheral blood mononuclear cell (PBMC) assays used in preclinical testing had failed to detect the cytokine release potential of TGN1412, as these assays did not replicate the tissue-like conditions and cell interactions present in vivo18.

The Progress: This failure led to major revisions in guidelines for first-in-human trials of immunotherapeutics. Regulatory agencies now require staggered dosing (administering to one volunteer at a time with extended observation periods), enhanced monitoring for cytokine release, and improved in vitro assays that better mimic physiological conditions1. Subsequent research identified that CD4+ effector memory T cells—which express high levels of CD28 in humans but not in cynomolgus macaques or other preclinical species—were the primary source of the cytokine storm, explaining the species-specific failure19. Additionally, preculturing human PBMCs at high cell density was shown to reveal the cytokine release potential of TGN1412, leading to the development of more predictive human-relevant in vitro models20.

Key Takeaways: Species-specific differences in CD28 expression on CD4+ effector memory T cells and immune system activation thresholds explain the failure of animal models and standard PBMC assays to predict the human response. This case underscores the limitations of animal models in predicting human immune reactions, particularly for immunotherapeutics, and highlights the need for human-relevant testing methods and improved preclinical models that account for species differences in immune cell populations and activation pathways1 2 19 20.

Addendum: The six volunteers experienced catastrophic issues again due to poor translatability between species.

  • Wrong cellular target: The monkey was considered relevant because its CD28 receptor closely resembles the human receptor and TGN1412 bound it effectively. However, the critical difference was which cells expressed CD28: human CD4 effector-memory T cells retain substantial CD28 expression, whereas this population generally lacks CD28 in cynomolgus and rhesus monkeys. The monkeys therefore had the “right” receptor but not the same dangerous, cytokine-producing target-cell population21.

  • **Misleading dose calculation:" The starting human dose of 0.1 mg/kg was based largely on a 50 mg/kg monkey NOAEL, creating an apparently large safety margin. But the monkeys were biologically unable to reproduce the relevant human response, so their high tolerated dose was not a meaningful safety boundary. The calculation effectively relied on a fraction of the NOAEL rather than a true, conservative minimum anticipated biological effect level (MABEL) - a serious error for a potent immune agonist capable of activating human T cells22.

Fialuridine
#

Background: Fialuridine (FIAU), a synthetic fluorinated nucleoside analogue of thymidine, was developed as a potent inhibitor of hepatitis B virus (HBV) replication. It showed strong antiviral activity in preclinical studies, reducing serum HBV DNA levels by 70–95% in early clinical testing1 23. The drug was the deaminated product of fiacitabine (FIAC) and was extensively tested in mice, rats, dogs, and monkeys, with no significant toxicity observed1.

The Failure: In a Phase II clinical trial conducted by the National Institutes of Health (NIH) in 1993, Fialuridine caused severe, delayed toxicity in patients despite its promising antiviral effects. Within 13 weeks of treatment, seven patients developed hepatic failure characterized by progressive lactic acidosis, pancreatitis, myopathy, and peripheral neuropathy. Five of these patients died, and two survived only after emergency liver transplantation1 24. Histological examination revealed microvesicular steatosis and mitochondrial abnormalities in liver tissue, with minimal hepatocyte necrosis24.

The Delay: The trial was abruptly terminated after the first hospitalization, but six additional patients developed severe toxicity in the subsequent weeks, as the drug’s delayed mitochondrial damage manifested1 23. The toxicity was not immediate, allowing patients to initially tolerate the drug while mitochondrial dysfunction accumulated.

The Progress: The Fialuridine disaster led to mandatory mitochondrial toxicity testing for nucleoside analogues. Regulatory agencies now require evaluation of mitochondrial DNA depletion, polymerase-γ inhibition, lactate production, mitochondrial morphology, and delayed toxicity in human-relevant systems before advancing similar compounds1. Mechanistic studies revealed that Fialuridine triphosphate inhibits DNA polymerase-γ, leading to defective mitochondrial DNA replication and structural mitochondrial defects in human cells25. Critically, the human equilibrative nucleoside transporter 1 (hENT1) is expressed in mitochondrial membranes in humans but not in rodents, explaining the species-specific toxicity: hENT1 transports Fialuridine into human mitochondria, causing accumulation and mitochondrial damage, while rodent mitochondria lack this transporter26.

Key Takeaways: Species differences in mitochondrial transporter expression (hENT1) and mitochondrial metabolism masked toxicity risks in animal models, emphasizing the need for human-relevant testing methods and mitochondrial-specific safety assessments for nucleoside analogues1 2 25 26.

Addendum: Fialuridine failed for two interacting reasons: a delayed mitochondrial mechanism and a human-specific difference in how the drug reached mitochondria.

  • Delayed mitochondrial DNA toxicity: Fialuridine was converted into an active triphosphate metabolite that could interfere with mitochondrial DNA replication, probably through inhibition or misuse by mitochondrial DNA polymerase γγ. With prolonged exposure, mitochondrial DNA was depleted and mitochondria could no longer produce enough respiratory-chain proteins. The resulting energy failure produced lactic acidosis, microvesicular liver injury, hepatic failure, pancreatitis, myopathy, and neuropathy27.

  • Human-specific mitochondrial exposure: The animals did not reproduce the human level of mitochondrial exposure. One leading explanation is that humans express equilibrative nucleoside transporter 1, or ENT1, in mitochondrial membranes, whereas the tested rodents and nonhuman primates apparently do not express it there in the same way. Human mitochondrial uptake would increase the delivery of fialuridine or its phosphorylated metabolites to the site where mitochondrial DNA is replicated, making mtDNA depletion much more likely explaining how the drug could accumulate in animal tissues and still fail to produce the distinctive human syndrome28.

Troglitazone
#

Background: Troglitazone, a thiazolidinedione developed for type 2 diabetes, demonstrated strong efficacy in preclinical animal studies by effectively lowering glucose, insulin, and lipid levels1. As an insulin-sensitizing agent, it interacts with the peroxisome proliferator-activated receptor gamma (PPAR-γ) to help maintain glucose homeostasis, improve insulin action, and preserve β-cell function in both animal and human studies29. As the first approved drug in its class, it was hailed as a breakthrough for managing type 2 diabetes29.

The Failure: In clinical use, Troglitazone caused idiosyncratic hepatotoxicity, including severe liver failure and death. Despite its promising preclinical profile, reports of acute liver failure emerged shortly after its approval, leading to its withdrawal from the U.S. market in 200030. At least two dozen cases of acute liver failure, including deaths and cases requiring liver transplantation, were reported to the FDA30.

The Delay: The drug was withdrawn after reports of liver failure became undeniable, but many patients had already suffered irreversible harm, including chronic liver damage and the need for transplantation1 30.

The Progress: The Troglitazone disaster led to enhanced post-marketing drug safety surveillance protocols, including stricter liver enzyme monitoring requirements for new drugs1. It also spurred the development of systems pharmacology models to predict species differences in bile acid-mediated hepatotoxicity, with a focus on bile salt export pump (BSEP) inhibition as a key mechanism31. Research demonstrated that Troglitazone and its sulfated metabolite potently inhibit human BSEP - up to 10-fold more strongly than in preclinical species—leading to toxic bile acid accumulation in hepatocytes32 33. This accumulation disrupts mitochondrial function, causing mixed hepatocellular and cholestatic liver injury that was not observed in animal models33.

Key Takeaways: Human-specific bile acid transport differences and BSEP inhibition caused rare but severe liver toxicity, highlighting the limitations of animal models in predicting human metabolic and transporter-mediated responses1 2 32 33.

Addendum: Troglitazone failed for two main, interacting reasons: its toxicity was idiosyncratic and delayed, and human-specific differences in bile-acid handling and drug metabolism allowed liver injury that routine animal studies did not reveal.

  • Human-specific toxic metabolites and mitochondrial stress: Troglitazone was metabolized into chemically reactive products, including oxidation products associated with its distinctive chromane ring. These metabolites could bind cellular proteins and trigger oxidative stress, apoptosis, and mitochondrial dysfunction. Troglitazone itself also impaired mitochondrial function in experimental systems. However, reactive metabolites alone do not fully explain the injury: their formation did not consistently correlate with which cells were damaged, so the mechanism was probably a combination of metabolism, mitochondrial susceptibility, oxidative stress, and immune or genetic factors rather than one uniquely human metabolite34.

  • Bile-acid transport made human livers more vulnerable: Troglitazone and its sulfate metabolite inhibited hepatic bile-acid transporters, including the bile salt export pump. This could cause toxic bile acids to accumulate inside hepatocytes. Systems-pharmacology modeling suggests that toxic bile-acid accumulation may be approximately ten times higher in humans than in rats after comparable exposure, because bile-acid composition and transporter behavior differ between species. That provides a concrete explanation for why rats could show enzyme induction and enlarged livers without developing the severe human pattern of cholestatic and hepatocellular injury35.

Vioxx (rofecoxib)
#

Background: Vioxx (rofecoxib), a selective COX-2 inhibitor, was developed by Merck as a safer alternative to traditional nonsteroidal anti-inflammatory drugs (NSAIDs) for treating arthritis and chronic pain36. It was approved by the FDA in May 1999 after passing extensive animal safety tests across multiple species, including rodents and dogs, with no significant cardiovascular signals detected1.

The Failure: Vioxx was linked to a dramatically increased risk of myocardial infarction and stroke after long-term use. The APPROVe (Adenomatous Polyp Prevention on Vioxx) trial revealed that patients taking 25 mg/day of rofecoxib had a 2.80-fold increased relative risk of confirmed cardiac events (including myocardial infarction) and a 2.32-fold increased relative risk of cerebrovascular events (including ischemic stroke) compared to placebo (see Table 2 in footnote)37. By 2004, extrapolations estimated that Vioxx had caused 88,000 to 140,000 excess cases of serious coronary heart disease in the U.S. alone, from which an estimated 38,720 to 61,600 patients died38. Earlier, the VIGOR (Vioxx Gastrointestinal Outcomes Research) study had shown a 5-fold increased risk of acute myocardial infarction compared to naproxen (naproxen vs rofecoxib relative risk 0.2)36, but Merck initially attributed this to naproxen’s cardioprotective effects rather than Vioxx’s cardiotoxicity39.

The Delay: Vioxx was voluntarily withdrawn by Merck on September 30, 2004, in one of the largest drug recalls in history, but many patients continued to suffer serious cardiovascular complications due to the delayed recognition of its risks38. Internal documents later revealed that Merck had withheld critical cardiovascular risk data from regulators and the public for years, further delaying protective action40.

The Progress: The disaster led to greater transparency in clinical trial reporting and stricter requirements for chronic-use medications, including mandatory cardiovascular risk assessment in preclinical and clinical studies1. It also highlighted the mechanistic flaw of COX-2 inhibitors: they selectively reduce vascular prostacyclin synthesis—a vasodilator that prevents platelet aggregation—without disrupting COX-1-derived thromboxane synthesis in platelets, which promotes platelet aggregation and vasoconstriction. This imbalance tips the prostacyclin/thromboxane ratio in favor of thromboxane, creating a prothrombotic state that predisposes users to thrombosis, hypertension, and atherosclerosis41 42.

Key Takeaways: Preclinical animal studies were not designed to detect long-term cardiovascular risks, as standard toxicology protocols fail to account for species differences in COX-1/COX-2 expression and prostanoid balance. This case underscores the need for human-relevant testing methods and mechanism-based safety assessments that evaluate thrombotic potential in addition to traditional toxicity endpoints41 42.

Addendum: Vioxx (rofecoxib) was a selective COX-2 inhibitor that reduced pain and inflammation but was withdrawn after evidence showed an increased risk of myocardial infarction, stroke, and other cardiovascular events.

  • Prostacyclin–thromboxane imbalance: Rofecoxib’s cardiovascular hazard is best explained as a loss of vascular protection, not as a uniquely human phenomenon. COX-2 inhibition lowers endothelial prostacyclin PGI2, which normally promotes vasodilation and restrains platelet activation, while platelet COX-1—and therefore platelet thromboxane TXA2 - is largely spared. That imbalance can favor vasoconstriction, hypertension, and thrombosis. Animal models were capable of demonstrating parts of this mechanism, but conventional toxicology studies usually did not reproduce the clinical setting in which an acute plaque rupture, chronic vascular disease, inflammation, and prolonged exposure convert a modest prothrombotic shift into myocardial infarction or stroke43.

  • The 20-HETE mechanism: Three months of rofecoxib treatment increased circulating 20-HETE in mice by more than 120-fold and shortened tail-bleeding time. 20-HETE is a vasoconstrictor and can promote platelet activation; rofecoxib appeared to increase it partly by inhibiting its degradation. Standard preclinical regulatory packages of the era focused on gross toxicological endpoints rather than lipidomic biomarkers of cardiovascular risk, meaning that this hypercoagulant signaling cascade went undetected in animal assays. 20-HETE was not part of the standard safety endpoints, and the biomarker’s connection to human cardiovascular events was recognized only later44.

  • Ischemia/reperfusion and arrhythmia: Rofecoxib’s risk became especially apparent when a rat heart was placed under ischemia/reperfusion stress, a condition more analogous to an acute coronary event than routine toxicity testing. Chronic treatment increased ventricular fibrillation and produced an approximately 7.7-fold higher odds of acute mortality compared with pooled control groups. This is important evidence that the drug could worsen cardiac electrical instability and survival highlighting cardiotoxicity under ischemic injury. Despite this finding, standard preclinical regulatory packages did not mandate these specialized disease-state models, allowing the drug to proceed to clinical phases without identifying these risks. A larger failure was that such disease-state, thrombosis, and electrophysiological studies were not routine requirements for evaluating a drug whose principal toxicity was expected to be gastrointestinal and renal rather than cardiovascular45.

BIA 10-2474
#

Background: BIA 10-2474, a fatty acid amide hydrolase (FAAH) inhibitor, was developed by Bial for anxiety, Parkinsonism, and chronic pain. It underwent preclinical testing in mice, rats, dogs, and monkeys, with no serious adverse events reported at doses far exceeding those later used in humans2 46.

The Failure: In a Phase I trial conducted by Biotrial in Rennes, France, in January 2016, BIA 10-2474 caused deep brain hemorrhage and necrosis in five out of six volunteers receiving the highest dose (50 mg/day for 10 days). One volunteer died, and the remaining four suffered permanent neurological damage, including cognitive impairment, memory loss, and motor deficits[^467]. The syndrome was characterized by headache, cerebellar dysfunction, altered consciousness, and MRI evidence of acute, progressive neurologic injury47.

The Delay: The trial was halted immediately after the first volunteer was hospitalized, but the other four in the high-dose group had already received their sixth dose. All five developed severe, long-term neurological symptoms, with MRI scans revealing deep, necrotic, and hemorrhagic lesions in their brains47.

The Progress: The disaster led to major revisions in drug administration protocols, including stricter dose-escalation rules and real-time pharmacokinetic monitoring in first-in-human trials48. Investigations revealed critical flaws in the trial design: the dose-escalation scheme (20 mg → 50 mg → 100 mg) did not account for nonlinear pharmacokinetics observed at higher doses, and emerging data from lower-dose groups were not used to adjust the protocol46 48. Mechanistic studies identified off-target inhibition of lipid-metabolizing enzymes (e.g., ABHD6, PNPLA6, PLA2G6) as the likely cause of toxicity. Unlike other FAAH inhibitors, BIA 10-2474 disrupted lipid networks in human cortical neurons, leading to metabolic dysregulation and neurotoxicity49.

Key Takeaways: Toxicity studies failed to evaluate metabolites or off-target effects, and animal models did not predict the severe human neurotoxicity. The case underscores the need for comprehensive off-target profiling, mechanism-based safety testing, and adaptive trial designs that incorporate emerging pharmacokinetic and toxicity data2 49.

Addendum: BIA 10-2474 was a supposedly selective FAAH inhibitor whose repeated high-dose administration in a first-in-human trial caused an unexpected, delayed, and partly irreversible neurotoxic syndrome, including one death.

  • Failure of sentinel dosing due to delayed onset: The trial employed a staggered dosing design where two volunteers received the drug first, followed by the rest after a brief safety observation period. However, because the neurotoxicity was cumulative and delayed—manifesting after multiple days of repeated dosing rather than acutely—this standard safety measure failed to protect the subsequent volunteers. This highlights a critical blind spot in first-in-human trial designs when evaluating compounds with delayed toxicity profiles47.

  • The likely problem was off-target, cumulative neurotoxicity: BIA 10-2474 inhibited FAAH but was relatively nonselective at higher exposures. Chemical-proteomic studies found inhibition of additional serine hydrolases, including ABHD11, PNPLA6, PLA2G6, and related lipid-metabolism proteins that are expressed in the brain. Its major metabolite also showed off-target activity, raising the possibility that repeated dosing caused accumulation and disrupted neuronal lipid metabolism. The exact mechanism remains unresolved50.

HIV/AIDS Research
#

Background: HIV/AIDS research has relied heavily on non-human primates (NHPs), particularly chimpanzees, due to their susceptibility to HIV-1 infection and perceived similarities in disease progression timeline to humans. However, chimpanzees do not develop AIDS despite chronic infection, and their immune responses differ fundamentally from those of humans51.

The Failure: Despite decades of research and over 90 HIV vaccines successful in animal models (including NHPs and macaques), none have translated to human efficacy due to profound species-specific immune response disparities52. The STEP trial (2007) exemplified this failure: a Merck adenovirus serotype 5 (Ad5)-based vaccine that showed promise in macaques failed to prevent HIV-1 infection in humans and may have enhanced susceptibility in certain populations53.

The Delay: The lack of suitable animal models—chimpanzees fail to develop clinical AIDS, while macaques infected with SIV/SHIV exhibit divergent immunopathology—has delayed vaccine and cure development. High HIV variability, structural weaknesses in trial design, and funding instability have further compounded progress54.

The Progress: Ethical concerns, cost considerations, and the recognition of species-specific limitations have led to significant reductions in chimpanzee use. Research has increasingly shifted toward human-relevant models, including 3D human tissue models, organ-on-a-chip systems, and computational approaches, which better recapitulate HIV pathogenesis, viral reservoirs, and immune responses55.

Key Takeaways: Animal models fail to fully recapitulate human HIV/AIDS pathology—chimpanzees resist disease progression, and macaques exhibit species-specific immune mechanisms—highlighting the need for human-relevant, mechanism-based approaches and ethical alternatives55.

Addendum: Chimpanzees were valuable for showing that HIV-1 could establish persistent infection, but they were poor models for predicting the human transition from infection to AIDS because their virus-host interaction usually produced much less chronic immune damage.

  • The disease phenotype did not transfer: HIV-1-infected chimpanzees generally maintained relatively stable CD4\(^+\) T-cell counts, preserved T-helper-cell function, and did not develop the chronic immune activation, gut-barrier damage, opportunistic infections, or malignancies typical of human AIDS. This was not because the virus necessarily failed to replicate: chimpanzees could sustain persistent infection, but their immune systems often avoided the self-reinforcing cycle of T-cell activation, apoptosis, microbial translocation, and systemic inflammation that drives human disease. A small number of chimpanzees did progress, so the difference was one of probability and typical disease course, not absolute resistance56.

  • Immune regulation, rather than CD4 count alone, was decisive: The chimpanzee model underestimated the importance of chronic immune activation. In humans, ongoing activation causes excessive T-cell turnover, bystander apoptosis, impaired T-cell renewal, lymphoid-tissue damage, and progressive immune exhaustion. HIV-1-infected chimpanzees often showed limited chronic activation and greater resistance of T cells to activation-induced apoptosis, allowing them to preserve immune function even when some direct viral killing occurred. Therefore, a study that measured viral load and peripheral CD4 counts without reproducing human lymphoid-tissue and mucosal pathology could incorrectly suggest that HIV-1 was relatively benign57.

  • Species-specific virus-host molecular interactions confounded interpretation: HIV-1 and chimpanzee SIVcpz arose from related cross-species viruses, but viral proteins interact differently with human and chimpanzee restriction factors, antigen-presentation molecules, innate sensors, and cellular cofactors. Factors such as TRIM5\(\alpha\), APOBEC3 proteins, tetherin, and species-specific MHC alleles can alter viral replication, immune recognition, and inflammatory signaling. However, it is too strong to say that HIV-1 is simply “rapidly neutralized” in chimpanzee cells: chimpanzees are susceptible to persistent HIV-1 infection. The key problem is that the same viral replication can produce a different balance of immune activation, tissue injury, and immune control in the two species58.

  • Evolutionary incompatibility further limits interspecies translation: These restriction systems are products of millions of years of host-retrovirus co-evolution. Viral countermeasures that are effective in one primate species may interact differently with the corresponding proteins in another, changing not only whether the virus replicates but also how innate sensing, inflammation, immune control, and tissue damage develop. Consequently, the chimpanzee model could reproduce persistent infection without reliably reproducing the human pathogenic interaction that leads to progressive AIDS. The problem was therefore not simply that HIV-1 was blocked in nonhuman primates, but that evolutionary differences shifted the entire balance between viral replication and disease59.

  • Chimpanzee vaccine studies also exposed a separate limitation: Protection against one HIV-1 strain or subtype did not necessarily generalize to another. In this study, two chimpanzees previously protected against subtype-B challenges were both infected after challenge with subtype-E virus, demonstrating that intraclade protection could not be assumed to predict interclade efficacy60. [Editorial Note: The study’s ability to support broad conclusions is limited by its sample size of only two animals. More generally, chimpanzee studies - and animal studies in biomedical research overall - may suffer from limited sample sizes, artificial challenge conditions, and design features that reduce their ability to predict human outcomes.]

Cancer Drug Development
#

Background: Animal models, especially mice, are widely used in cancer drug development due to their genetic manipulability, rapid reproduction, and ability to grow human tumor xenografts61. However, mouse models often fail to recapitulate the complexity of human tumor microenvironments, including stromal interactions, immune system differences, and metabolic heterogeneity62.

The Failure: Many cancer drugs effective in animal models fail in human trials due to differences in tumor microenvironment, immune response, and drug metabolism61. For decades, drug candidates that may have had efficacy in human-specific microenvironments were discarded because they failed to shrink localized tumors in rodents, while compounds that successfully shrank rodent tumors advanced to clinical trials, only to fail in patients62. Notably, only a historically low single-digit percentage (e.g., ~3–8%) of oncology drugs that enter clinical trials ultimately gain approval, with poor translational predictivity of mouse models being a major contributor63.

The Delay: The reliance on animal models has slowed the development of effective cancer therapies and increased costs61. Traditional 2D cell cultures and murine xenografts lack the structural, biochemical, and mechanical complexity of human tumors, leading to false positives in preclinical screening and late-stage clinical failures63.

The Progress: Emerging in vitro and in vivo human systems, such as 3D patient-derived tumor organoids, spheroids, and vascularized tumor-on-a-chip platforms, offer improved predictive accuracy and reduced animal use64 65. These platforms allow researchers to test therapies on a patient’s own cells within a bioengineered human microenvironment, capturing tumor heterogeneity, vascular flow, and metastatic potential far more accurately than mouse models65 66. Vascularized tumor-on-a-chip systems, for example, replicate human tumor microenvironments by integrating perfusable vessels and stromal elements, enabling dynamic assessment of drug delivery and therapeutic response under physiologically relevant conditions66.

Key Takeaways: Animal models often fail to replicate human cancer biology; human-relevant models—such as patient-derived organoids and tumor-on-a-chip platforms—are critical for improving drug development success rates and reducing translational failures61 66.

Addendum: Translational inadequacies to humans make the animal model ineffective. As some scientists put it, ‘we have cured cancer in mice’, but not in humans. (See This frog bacterium wiped out cancer tumors in mice with a single dose and Spanish scientists cure pancreatic cancer in mice.)

  • The microenvironment is structurally different: Subcutaneous mouse xenografts often place human tumor cells in an anatomically artificial site and may omit the stromal, vascular, immune, and extracellular-matrix conditions that determine tumor growth, drug penetration, hypoxia, invasion, and resistance. Immunodeficient xenograft hosts are especially limited for testing immunotherapies because the human tumor is interacting with a severely altered or absent immune system. Orthotopic, genetically engineered, and patient-derived xenograft models can improve biological relevance, but none fully reproduces the human tumor microenvironment or metastatic process67.

  • Mouse models underrepresent metastasis and clinical heterogeneity: Many laboratory tumors are selected for rapid, uniform growth and are evaluated as localized masses, whereas human cancers evolve through branching clones, therapy-resistant subpopulations, dormancy, dissemination, and organ-specific metastasis. Consequently, a treatment can shrink a primary mouse tumor while failing against metastatic or resistant disease in patients. Patient-derived organoids can preserve more of the donor tumor’s genetic and phenotypic heterogeneity and allow testing across multiple drugs, but they still commonly lack intact blood vessels, immune cells, fibroblasts, nerves, and other components of the tumor microenvironment68.

  • Reproducibility and experimental design weaken translation: In 2012, Amgen researchers reported that only 6 of 53 prominent preclinical oncology findings could be reproduced, meaning that approximately 90% failed replication. This is a direct reproducibility failure, and it undermines confidence in those studies as a foundation for predicting clinical efficacy. The problem cannot be dismissed as merely a limitation in generalizing from mice to humans: if the result does not reliably recur even under preclinical investigation, its translational value is already compromised69 70. [Editorial Note: Former head of global cancer research at Amgen, Glenn Begley’s exclamation “It was shocking” to Reuters69 highlights the paucity of scientific rigor animal models inherently bear that no amount of procedural refinement71 is going to overcome. A poor model cannot reliably predict a biological outcome it does not reproduce. Don’t try to overcome a bad habit - just start a good new one with NAM!]

Stroke and Traumatic Brain Injury (TBI)
#

Background: Animal models, particularly rodent Middle Cerebral Artery Occlusion (MCAO), are used to study stroke and TBI due to their controlled experimental conditions52. A filament is inserted into the rodent brain to induce ischemia, to evaluate neuroprotective candidates.

The Failure: Animal models often fail to predict human neurological outcomes due to differences in brain anatomy, physiology, and recovery mechanisms52. Numerous experimental treatments have shown neuroprotective efficacy in animal models of stroke, but none have proven effective in Phase III human clinical trials72. Over 1,000 neuroprotective strategies have been tested in preclinical animal models, yet all have failed to translate to clinical efficacy, despite promising results in rodents73.

The Delay: Reliance on animal models has delayed the development of effective stroke and TBI treatments52. The systemic failure to control experimental bias in preclinical animal models led to an overestimation of efficacy, driving premature clinical trials and wasting hundreds of millions of dollars74. Poor study quality—including lack of randomization, blinding, and appropriate power calculations—has been identified as a major contributor to inflated efficacy estimates in animal studies74.

The Progress: Advanced imaging and human-based models, including organ-on-chip systems, are emerging as more predictive alternatives75. Human-centric methodologies, such as human cortical slice cultures, microfluidic blood-brain barrier (BBB) models, and human in vitro neurovascular units, have since improved the predictive accuracy of neuroprotective screenings75 76. These systems accurately model human cellular responses to ischemia and reperfusion under controlled conditions, avoiding the confounding variables inherent in animal surgeries.

Key Takeaways: Species-specific neurological differences limit animal model utility; human-relevant models are essential for progress52.

Addendum: Human-centered systems such as cortical slice cultures, microfluidic blood-brain barrier models, and human neurovascular-unit platforms provide controlled ways to study ischemia, reperfusion, barrier failure, neuronal injury, and vascular inflammation without the surgical and physiological confounders of whole-animal models.

  • The experimental animals do not resemble the patients: Most preclinical stroke studies have historically used young, healthy male rodents, whereas clinical stroke disproportionately affects older people of both sexes who often have hypertension, diabetes, obesity, hyperlipidemia, atherosclerosis, or cardiovascular disease. These conditions alter vascular structure, inflammation, blood-brain barrier integrity, infarct development, drug distribution, and endogenous repair. A systematic review found that only 11.4% of preclinical ischemic-stroke studies included an aged or comorbid animal, making it unsurprising that neuroprotective effects observed in healthy young rodents frequently disappear in clinically realistic populations77.

  • Anesthesia and uncontrolled physiology can manufacture apparent protection: Anesthetic agents used during rodent stroke surgery can independently reduce neurological injury, making an experimental drug appear more effective than it is. A systematic review found that anesthetic treatment reduced neurological injury by approximately 28% in experimental rodent stroke studies, and that this apparent protection failed in female animals and animals with comorbidities. Temperature, blood pressure, blood gases, glucose, pH, and cerebral perfusion must also be monitored carefully because hypothermia and other physiological changes can substantially alter infarct size and recovery. Thus, an intervention may appear neuroprotective because it was tested under an anesthetic or physiological state that already suppresses ischemic injury78.

  • Human neurovascular models can expose failures earlier: The neurovascular unit, rather than the neuron alone, is a central therapeutic target in stroke because endothelial cells, pericytes, astrocytes, neurons, immune cells, and extracellular matrix jointly determine barrier permeability, edema, inflammation, and drug delivery. Microfluidic human BBB and neurovascular-unit models can impose controlled hypoxia, hypoglycemia, and halted perfusion while measuring barrier leakage, mitochondrial dysfunction, ATP depletion, and cellular injury in real time. These platforms should not be presented as fully validated replacements for clinical trials, but they can serve as human-relevant filters that reveal poor BBB penetration, vascular toxicity, or loss of efficacy under ischemia-reperfusion conditions before a candidate is advanced into poorly predictive animal studies79. For example, a neurovascular-unit chip has reproduced stroke-associated reductions in barrier integrity, mitochondrial potential, and ATP while incorporating human endothelial cells, astrocytes, and neurons. These models are not complete replacements for clinical biology, but they can provide more human-relevant screening and mechanistic information than isolated neuronal cultures or highly artificial rodent paradigms80.

Toxicology: Dioxin, Asbestos, and Smoking
#

Background: Animal models have been used extensively in toxicology to assess chemical safety, including dioxin, asbestos, and smoking-related toxins52. However, species-specific differences in metabolism, receptor sensitivity, and anatomical structures fundamentally limit their predictive value for human toxicity.

The Failure: Animal models often fail to predict human toxicity accurately. For dioxin (TCDD), rodent strains exhibit significant enzyme induction at body burdens below 50 ng/kg, while human cells require ~10-fold higher concentrations due to impaired AhR binding function, leading to underestimation of human sensitivity and overestimation of rodent risk81. For asbestos, rodent models often develop less severe fibrosis and different lesion patterns compared to humans, with chrysotile asbestos causing minimal pathology in rodents despite its known human carcinogenicity82. In smoking-related COPD, rodent models produce emphysema-like lesions that are not close mimics of human disease, with species-to-species variation in pathology that limits translational relevance83.

The Delay: Overreliance on animal models has delayed the recognition of human health risks from environmental toxins52. For example, dioxin risk assessments based on rodent data failed to account for human-specific AhR polymorphisms and metabolic differences, leading to misclassification of human sensitivity81. Similarly, asbestos potency varies dramatically between species, with amphibole fibers (e.g., crocidolite) being far more potent in humans than in rodents, yet this was not fully appreciated until epidemiological studies revealed the discrepancy84.

The Progress: In vitro and computational toxicology models are increasingly used to improve predictive accuracy and reduce animal testing85. Human cell-based assays for AhR activation, 3D lung tissue models for asbestos fiber toxicity, and human airway organoids for smoking-induced damage now provide more human-relevant assessments of carcinogenic and fibrotic potential. These models better capture species-specific metabolic pathways, receptor binding affinities, and tissue responses that animal models miss.

Key Takeaways: Animal models often under- or overestimate human toxicological risks; alternative methods, such as human cell-based assays, 3D tissue models, and computational approaches, provide more reliable and ethical assessments52.

Addendum: Environmental toxicology illustrates a central limitation of animal testing: an animal may reproduce a general injury pattern while still misrepresenting human susceptibility, exposure, dose-response relationships, or latency. Human-relevant in-vitro systems, computational toxicology, epidemiology, and exposure-based modelling can therefore provide evidence that animal studies alone cannot reliably supply.

  • Dioxin toxicity is highly species-and strain-dependent: Dioxin toxicity is mediated largely through the aryl hydrocarbon receptor AhR, but human and rodent cells can differ substantially in receptor responsiveness, downstream gene regulation, metabolism, and tissue-specific effects. Comparative experiments found that TCDD produced different gene-expression profiles in human and mouse cells, demonstrating that similar receptor activation does not guarantee equivalent biological outcomes. Other studies have also reported large differences in TCDD susceptibility across species and mouse strains, with differences in both AhR affinity and the ability to convert an external dose into internal tissue concentrations. These disparities make direct extrapolation from rodent doses to human risk highly uncertain86 87.

  • Asbestos demonstrated that animal confirmation does not define human risk: Animal studies eventually showed that asbestos fibres can cause fibrosis, lung cancer, and mesothelioma, but they were poor quantitative predictors of human risk. Fibre deposition, clearance, lung anatomy, exposure duration, and fibre dimensions differ substantially between rats and humans. One comparative analysis found that rat inhalation experiments required fibre concentrations more than 100 times higher than those associated with lung-cancer risk in asbestos workers, and approximately 1,000 times higher to produce a comparable mesothelioma risk. Animal studies therefore supported the qualitative conclusion that asbestos is carcinogenic, while human epidemiology was essential for estimating its occupational danger88.

  • Smoking toxicity cannot be reduced to a single-animal exposure model: Cigarette smoke-induced rodent models can reproduce some inflammation, oxidative stress, emphysema, and vascular changes, but they do not reproduce the full human disease trajectory. Rodents generally fail to develop the severe chronic bronchitis, mucus-gland pathology, acute exacerbations, and progressive GOLD 3–4 COPD that characterize many human smokers. Their smaller and differently structured airways, distinct respiratory-cell populations, shorter exposure periods, and non-equivalent smoke-delivery systems also change particle deposition and metabolism. Human airway organoids, air-liquid-interface cultures, and lung-on-chip systems can therefore provide complementary human-relevant evidence for epithelial injury, inflammatory signalling, and smoke toxicity89 90.

COVID-19 Vaccines
#

Background: During the pandemic, various animal models, including non-human primates (NHPs) and rodents, were utilized in preclinical vaccine evaluation91. However, standard rodent models (such as wild-type mice) are inherently resistant to SARS-CoV-2 due to the incompatibility between the viral spike protein and murine ACE2 receptors92. Consequently, researchers were forced to rely on genetically modified rodents (e.g., transgenic mice expressing human ACE2) or artificially adapted viral strains to enable infection, highlighting fundamental species-specific susceptibility differences91 92. NHPs, such as rhesus and cynomolgus macaques, while having similarities to human ACE2 receptors also had significant differences which led to response incompatibilities93.

The Failure: Some vaccine candidates effective in animals failed in humans due to differences in viral binding, immune response, and vaccine-induced immunity52. For example, while mouse-adapted SARS-CoV-2 models showed protective efficacy, these results did not always translate to human trials due to divergent immune mechanisms and ACE2 receptor binding affinities92. Additionally, HLA class I and II differences between NHPs and humans limited the predictive value of T cell-based vaccine responses in primates93.

The Delay: The need for rapid vaccine development exposed the limitations of animal models in predicting human immune responses52. Mouse models required engineered human ACE2 expression to support infection, while hamster models—though susceptible to wild-type virus—exhibited disease severity patterns that did not fully mirror human COVID-1992.

The Progress: The pandemic accelerated the adoption of human-based models and computational methods to improve vaccine design and testing94. AI-driven approaches leveraged genomic data, protein structure analysis, and immune system modeling to rapidly identify vaccine candidates, predict antigenic targets, and optimize clinical trial design94 95. These computational pipelines enabled faster, more accurate vaccine development by integrating structural bioinformatics, machine learning, and immunoinformatics95.

Key Takeaways: Animal models have limited predictive value for human vaccine efficacy; human-relevant models (e.g., humanized mice, organ chips) and computational approaches (e.g., AI/ML for antigen prediction) are critical for future pandemic preparedness52.

Addendum: Animal models used during the COVID-19 emergency supplied rapid infection, immunogenicity, and challenge data, but they did not provide a complete forecast of human vaccine protection. Their limitations involved receptor biology, artificial infection conditions, restricted disease phenotypes, and immune responses that were less diverse than those of human populations.

  • Infection biology differed between species: Ordinary laboratory mice are relatively resistant to ancestral SARS-CoV-2 because mouse ACE2 interacts poorly with the viral spike protein. Researchers therefore used human-ACE2 transgenic mice or mouse-adapted viruses, but these modifications introduced new problems: ACE2 could be expressed in unnatural tissues, and adaptation could alter viral tropism or antibody sensitivity. Hamsters and nonhuman primates were more permissive, but they generally developed mild or moderate disease and did not reproduce the severe, multisystem COVID-19 seen in vulnerable humans. These differences complicated interpretation of vaccine protection, because preventing detectable virus in an artificial challenge model was not equivalent to preventing severe disease or transmission in people96.

  • Vaccine-induced immunity was not the same as human protection: Animal studies measured antibody titres, T-cell responses, viral RNA, and lung pathology, but these endpoints did not always predict the durability, breadth, or clinical effectiveness of human immunity. Laboratory animals were often genetically similar, pathogen-controlled, young, and immunologically inexperienced, whereas human vaccine recipients differed in age, prior infection, comorbidities, immune history, and exposure to evolving viral variants. NHP challenge studies also used small cohorts, high inoculum doses, and direct inoculation routes that did not reproduce ordinary human exposure. A systematic review concluded that animal models reproduced only a limited subset of human COVID-19 signs and symptoms, reinforcing the need to interpret vaccine-challenge results cautiously97.

  • Human-derived systems can test human viral tropism and immune-relevant responses: Human airway and lung organoids reproduce the differentiated epithelial cell types that SARS-CoV-2 infects in people, including ciliated, club, and alveolar type II cells. They have also reproduced differences in infectivity and host response between viral variants and can test antiviral or antibody activity in tissue architectures that more closely resemble the human respiratory tract than standard cell lines or animal tissues. Human organoids and organs-on-chips therefore provide direct platforms for examining human receptor usage, viral entry, epithelial injury, variant escape, and mucosal antiviral responses during vaccine and therapeutic development98.

  • Computational immunology can screen human-specific vaccine targets: Computational approaches can analyze circulating viral sequences against human HLA diversity, predict conserved B-cell and T-cell epitopes, model antibody-antigen interactions, and identify mutations likely to cause immune escape. During COVID-19, immunoinformatics approaches were used to prioritize multi-epitope vaccine designs from SARS-CoV-2 sequence data and estimate their potential coverage across human populations. These predictions require biological confirmation, but that confirmation can be performed through human-relevant assays such as peptide–HLA binding, human T-cell and B-cell activation, serum neutralization, and human airway or organoid systems rather than defaulting to animal models99.

Cross-Cutting Themes in Animal Research Failures
#

Despite spanning diverse therapeutic areas and disease models, the failure of animal research to predict human outcomes is not a series of isolated incidents, but rather the result of deep, systemic flaws. Across toxicology, immunology, and neuroscience, animal models consistently struggle with fundamental biological mismatches, methodological weaknesses, and institutional inertia. Recognizing these recurring patterns is essential for shifting the scientific paradigm away from entrenched, animal-dependent frameworks and toward more predictive, human-relevant methodologies.

Common Causes of Failure
#

  • Pharmacokinetics and Immunogenicity: Animal models often fail due to interspecies differences in drug metabolism and immune responses. For example, animals frequently mount anti-drug antibody (ADA) responses to both human and humanized monoclonal antibodies, altering drug exposure and confounding toxicity interpretation. These responses, though, are poor predictors of human immunogenicity100.

  • Experimental Bias and Stress Confounders: Small sample sizes, inadequate statistical power, failure to control for confounding variables, and stress-induced physiological changes in captive animals routinely undermine study validity as well as reproducibility101 52.

  • Publication Bias and Study Design: Preclinical animal studies often lack standardized reporting practices and suffer from severe publication bias favoring positive results, thereby undermining the quality and trustworthiness of the evidence base102.

Systemic Barriers to Human-Relevant Research
#

  • Regulatory Inertia: Resistance from regulators and sponsors to adopt non-animal New Approach Methodologies (NAM) slows progress. Driven by an entrenched bureaucratic culture, the reluctance to update legacy guidelines to human-relevant thresholds, upholds an outdated, inaccurate, ‘medieval’ environment103.

  • Funding Bias: Industry-funded research tends to favor sponsor outcomes, skewing results and undermining public trust in scientific findings104.

  • Academic Dogma: Career incentives, disciplinary silos, and communication barriers limit collaboration and innovation in research methods across various fields such as biology, mathematics, physics, and engineering105 106.

Successful Alternatives
#

  • In Silico Methods: Computational models simulate human biology and drug effects, increasingly outperforming animal models in predictive accuracy and early risk detection107.

  • In Vitro Methods: 2D coculture, 3D spheroids, organoids, and organ-on-chip models replicate human organ function and disease mechanisms more accurately than animal models, improving data relevance as well as reducing ethical concerns108 109.

  • Ex Vivo Assays: Living tissue cultures exposed to chemicals or drugs in controlled environments provide physiologically relevant data without the need for whole-animal use110.

Regulatory and Industry Shifts
#

The persistent translational failures of animal models have catalyzed a profound paradigm shift in biomedical and environmental research. This transition is no longer confined to academic debate; it is actively being codified into law, regulatory policy, and market economics. Driven by the urgent need for more predictive, human-relevant science, legislative bodies, regulatory agencies, and private investors are increasingly converging to mandate, fund, and adopt New Approach Methodologies (NAM). This momentum marks a critical turning point, systematically dismantling legacy animal-dependent frameworks in favor of more accurate, ethical, and efficient scientific standards.

FDA Modernization Acts 2.0 & 3.0
#

  • FDA Modernization Act 2.0 (2022): Authorized the use of non-animal New Approach Methodologies (NAM),such as in silico and in vitro models, for nonclinical drug testing, removing the strict federal mandate that animal studies must be conducted for FDA approval111.

  • FDA Modernization Act 3.0 (2025): Requires FDA to update regulations to reflect nonclinical testing reforms, replacing “animal” with “nonclinical” tests, and to publish an interim final rule within one year, effectively forcing the integration of NAM into legacy guidelines112.

EPA 2035 Mandate
#

  • The EPA has expanded acceptable NAM for chemical assessments and opened pathways for new method nominations.

  • The aim is to reduce vertebrate animal use in regulatory testing, with a long-term goal of eliminating mammalian testing by 2035113.

Adoption of New Approach Methodologies (NAM)
#

  • The FDA’s 2025 roadmap outlines a strategic shift to validated NAM, including AI-based computational models, organ-on-chip, and organoid-based testing, collaborating with NIH and ICCVAM for method validation114.

  • Industry adoption is growing, with companies developing organ-on-chip and computational modeling platforms. The organ-on-chip market is projected to grow from $157 million (2024) to $952 million (2030), and the organoid market from $1.4 billion (2025) to $4 billion (2035)112.

Funding Trends #

  • NIH: Substantial investments in developing, standardizing, and validating human-based NAMs, including the Complement-ARIE program to accelerate NAM for biomedical research and safety testing115.

  • Private Sector: Over 40 partners, from small developers to large pharma, are advancing NAMs. Significant investments in organ-on-chip and organoid technologies are driving market growth and regulatory acceptance112.

Ethical and Socioeconomic Dimensions
#

Beyond the scientific and translational limitations of animal models, their continued use carries profound ethical and socioeconomic consequences. The reliance on outdated animal paradigms not only perpetuates systemic animal suffering but also distorts global research priorities, exacerbates health inequities, and misallocates finite scientific resources. Addressing these dimensions is not merely a moral imperative; it is a fundamental requirement for building a more equitable, efficient, and trustworthy scientific enterprise.

Animal and Human Toll
#

  • Animal experimentation inherently causes undeniable and significant suffering. Forced confinement, disease induction, subjection to painful and traumatic processes, and usually eventual death, raise fundamental ethical concerns116.

  • Researchers and the biomedical society face moral dilemmas and psychological impacts from these practices, which also erode public trust in both the ethics as well as the science itself116.

  • Medical treatments based on flawed animal research adversely affects the patients being served resulting in disastrous clinical outcomes (as documented throughout this report).

  • Environment damage caused by animal experimentation can be devastating - see Environmental Impact of Animal Experimentation.

False Positives/Negatives
#

  • Animal models frequently yield false positives (effective in animals but failing in humans) and false negatives (missing human-specific risks), wasting resources and misleading research efforts52 117.

  • Such ‘hit-and-miss’ approaches lead to enormous monetary expense, diverting funds from more accurate, human-relevant methods - see Animal Research vs NAM Costs.

Research Equity and Global Representation
#

  • Research priorities disproportionately focus on diseases with immediate economic impact or those prevalent in high-income countries creating severe demographic gaps, underrepresenting the health needs of the Global South, and exacerbating health inequities118.

  • Discrepancies in how research value is defined, coupled with entrenched institutional biases, perpetuate a cycle where outdated animal models consume funding that could otherwise be directed toward human-relevant, equitable, and innovative health solutions116 118.

Actionable Recommendations
#

Addressing the systemic failures of animal research requires coordinated, multi-stakeholder action. Transitioning toward a more predictive, ethical, and efficient scientific paradigm cannot be achieved by any single group alone. By aligning the efforts of researchers, regulators, funders, and the public, we can accelerate the adoption of human-relevant New Approach Methodologies (NAMs), enforce rigorous ethical standards, and ultimately build a biomedical research ecosystem that truly serves human health.

For Researchers
#

  • Adopt microphysiological systems, computational modeling, and non-animal trial frameworks to improve predictive accuracy and reduce animal use119.

  • Follow ARRIVE guidelines for rigorous animal study design and reporting to ensure transparency and reproducibility120.

  • Prioritize animal welfare by adhering to AVMA euthanasia guidelines and providing environmental enrichment121 122.

For Regulators
#

  • Update and enforce ARAC and IACUC guidelines to ensure ethical and humane animal research practices123 124.

  • Promote transparency by publishing experiment evaluations online to build public trust and understanding125.

  • Support and validate NAMs in regulatory testing to advance the Three Rs (Replacement, Reduction, Refinement)126.

For Funders
#

  • Allocate resources to develop and validate alternative methods, such as organ-on-chip and computational models119.

  • Encourage adherence to ethical guidelines and support research that prioritizes animal welfare120 127.

  • Foster collaboration among researchers, regulators, and industry to accelerate NAMs development and adoption126.

For the Public
#

  • Educate and advocate for animal welfare and the widespread adoption of alternative research methods124.

  • Support ethical research practices and organizations promoting humane animal research standards125.

  • Engage in public policy initiatives to improve animal research regulation and oversight126.

References
#


  1. Progress Science Mauritius, “When Animal Testing Fails: 6 Famous Drug Disasters That Harmed Humans”, n.d.
    This article reviews six historical pharmaceutical disasters where drugs that passed preclinical animal safety trials subsequently caused severe harm or death in human clinical trials or post-market use. It highlights the fundamental physiological and metabolic limitations of animal models in predicting human toxicity, underscoring the urgent need for more human-relevant testing methodologies to prevent future tragedies. ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎

  2. Van Norman G., “Limitations of Animal Studies for Predicting Toxicity in Clinical Trials: Is it Time to Rethink Our Current Approach?”, JACC: Basic to Translational Science, 2019.
    This paper critically examines the high failure rate of drugs in clinical trials due to unforeseen toxicity, despite prior safety in animal models. The author argues that profound genetic, physiological, and metabolic differences between species render animal studies unreliable predictors of human adverse events, advocating for a paradigm shift toward human-based in vitro and in silico New Approach Methodologies (NAMs). ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎

  3. Yang, Joanna, Tantibanchachai, Chanapa, “Studies of Thalidomide’s Effects on Rodent Embryos from 1962-2008”, Embryo Project Encyclopedia, 2014.
    This historical review details the decades-long scientific struggle to replicate thalidomide-induced birth defects in rodent models following the human tragedy of the 1960s. It illustrates how the initial failure of animal models to predict human teratogenicity delayed the understanding of the drug’s mechanism, ultimately highlighting the highly species-specific nature of thalidomide’s toxic effects. ↩︎ ↩︎ ↩︎

  4. Vargesson, N., “Thalidomide-induced teratogenesis: history and mechanisms”, Birth Defects Research Part C: Embryo Today: Reviews, 2015.
    This review traces the history of thalidomide from its initial release and subsequent withdrawal due to severe birth defects, to its modern renaissance as a treatment for multiple myeloma and leprosy. It details the molecular mechanisms of its teratogenicity, particularly its binding to the protein cereblon, and discusses how understanding these pathways has informed safer drug development while exposing the limitations of historical animal testing. ↩︎ ↩︎ ↩︎

  5. Vargesson, Neil, and Trent Stephens, Thalidomide: History, Withdrawal, Renaissance, and Safety Concerns, Expert Opinion on Drug Safety, 2021.
    This comprehensive article explores the complex lifecycle of thalidomide, from its catastrophic impact on fetal development to its carefully regulated reintroduction for specific conditions. It emphasizes the critical safety concerns and strict risk management programs required today, serving as a cautionary tale about the dangers of relying solely on inadequate preclinical animal data for assessing teratogenicity. ↩︎

  6. Kelsey, F. O., “Problems Raised for the FDA By the Occurrence of Thalidomide Embryopathy in Germany, 1960-1961”, American Journal of Public Health, 1999.
    This historical account details the pivotal role of FDA medical officer Frances Kelsey in blocking the approval of thalidomide in the United States, thereby preventing a similar tragedy to the one in Europe. It discusses the regulatory gaps exposed by the event and how it fundamentally transformed drug approval processes, emphasizing the need for rigorous proof of safety and efficacy before human exposure. ↩︎ ↩︎

  7. Niethammer, M., Burgdorf, et al., In vitro models of human development and their potential application in developmental toxicity testing, Development, 2022.
    This paper discusses the development and application of advanced in vitro models, such as stem cell-derived embryoids and organoids, for assessing developmental toxicity. It argues that these human-relevant models can overcome the species-specific limitations of traditional animal tests, offering more accurate predictions of how chemicals and drugs affect human embryonic development. ↩︎

  8. Rehman, W., Arfons, L. M., and Lazarus, H. M., The rise, fall and subsequent triumph of thalidomide: lessons learned in drug development, Therapeutic Advances in Hematology, 2011.
    This article chronicles the dramatic history of thalidomide, analyzing the pharmacological and regulatory lessons learned from its initial failure and eventual successful repurposing. It highlights how modern pharmacovigilance, strict risk evaluation, and a deeper understanding of its mechanism of action have allowed its safe use, contrasting sharply with the inadequate preclinical testing of the past. ↩︎

  9. Meganathan, K., Jagtap, S., Wagh, V., et al., Identification of thalidomide-specific transcriptomics and proteomics signatures during differentiation of human embryonic stem cells, PLOS ONE, 2012.
    This study utilizes human embryonic stem cells to identify specific transcriptomic and proteomic signatures associated with thalidomide exposure during cellular differentiation. By mapping these human-specific molecular responses, the research demonstrates the utility of in vitro human cell models in detecting teratogenic risks that are often missed by conventional animal testing. ↩︎

  10. Knobloch, J., Shaughnessy, J. D., and Rüther, U., Thalidomide induces limb deformities by perturbing the Bmp/Dkk1/Wnt signaling pathway, The FASEB Journal, 2007.
    This research identifies the specific molecular pathway through which thalidomide causes limb deformities, demonstrating that it perturbs the Bmp/Dkk1/Wnt signaling cascade. The findings provide a mechanistic understanding of its teratogenicity, highlighting how human-relevant molecular studies can uncover toxicity pathways that are not conserved or easily observable in standard animal models. ↩︎

  11. Knobloch, J., Reimann, K., Klotz, L., Rüther, U., “Thalidomide resistance is based on the capacity of the glutathione-dependent antioxidant defense”, Mol Pharm, 2008.
    This study investigates the mechanisms behind species-specific resistance to thalidomide-induced teratogenicity, identifying the glutathione-dependent antioxidant defense system as a key protective factor in resistant species. This finding helps explain why certain animal models failed to show the severe birth defects seen in humans, underscoring the metabolic differences that complicate cross-species toxicity extrapolation. ↩︎ ↩︎

  12. Aikawa, N., Kunisato, A., Nagao, K., et al., Detection of Thalidomide Embryotoxicity by In Vitro Embryotoxicity Testing Based on Human iPS Cells, Journal of Pharmacological Sciences, 2014.
    This paper demonstrates the successful use of human induced pluripotent stem (iPS) cells in an in vitro embryotoxicity test to detect the harmful effects of thalidomide. The study validates human iPS cell models as a highly sensitive and reliable alternative to animal testing for predicting developmental toxicity and safeguarding human health during drug development. ↩︎

  13. Shanks, N., Greek, R., and Greek, J., Are animal models predictive for humans?, Philosophy, Ethics, and Humanities in Medicine, 2009.
    This critical review evaluates the scientific validity of using animal models to predict human responses to drugs and diseases. The authors conclude that due to significant genetic, physiological, and metabolic differences between species, animal models are fundamentally poor predictors of human outcomes, advocating for a transition to human-biology-based research methods. ↩︎

  14. Ito, T., Ando, H., Handa, H., “Teratogenic effects of thalidomide: molecular mechanisms”, Cell Mol Life Sci, 2011.
    This review elucidates the molecular mechanisms underlying the teratogenic effects of thalidomide, focusing on its interaction with the protein cereblon and the subsequent degradation of specific transcription factors essential for limb development. It highlights how modern molecular biology has finally explained the human-specific toxicity that baffled researchers using traditional animal models for decades. ↩︎

  15. Lin, C. H., and Hünig, T., Efficient expansion of regulatory T cells in vitro and in vivo with a CD28 superagonist, European Journal of Immunology, 2003.
    This study explores the use of a CD28 superagonist to efficiently expand regulatory T cells both in vitro and in vivo, initially showing promise for treating autoimmune diseases. However, it also sets the stage for understanding the severe, unanticipated immune responses that can occur when such potent immunomodulators are introduced, highlighting the complexities of translating in vitro findings to human in vivo contexts. ↩︎ ↩︎

  16. Suntharalingam, G., et al., “Cytokine Storm in a Phase 1 Trial of the Anti-CD28 Monoclonal Antibody TGN1412”, New England Journal of Medicine, 2006.
    This landmark paper describes the catastrophic Phase 1 clinical trial of TGN1412, where six healthy volunteers experienced a life-threatening “cytokine storm” despite the drug showing no toxicity in preclinical animal models, including non-human primates. It serves as a stark reminder of the limitations of animal testing in predicting severe human immunological reactions to novel biologics. ↩︎

  17. Stebbings, R., et al., “After TGN1412: Recent developments in cytokine release assays”, Journal of Immunotoxicity, 2012.
    Following the TGN1412 tragedy, this paper reviews the subsequent advancements in in vitro cytokine release assays designed to better predict the risk of cytokine storms in humans. It emphasizes the shift toward using human primary cells and more sophisticated in vitro models to assess the immunogenicity and safety of novel monoclonal antibodies before they enter clinical trials. ↩︎

  18. Römer, P. S., Berr, S., Avota, E., et al., “Preculture of PBMCs at high cell density increases sensitivity of T-cell responses, revealing cytokine release by CD28 superagonist TGN1412”, Blood, 2011.
    This research identifies a critical methodological flaw in the preclinical testing of TGN1412, demonstrating that preculturing human peripheral blood mononuclear cells (PBMCs) at high density is necessary to reveal the drug’s potent cytokine-releasing properties. The findings highlight how specific in vitro conditions can uncover human-specific toxicities that standard animal models and low-density cell assays fail to detect. ↩︎

  19. Eastwood, D., et al., “Monoclonal antibody TGN1412 trial failure explained by species differences in CD28 expression on CD4+ effector memory T-cells”, British Journal of Pharmacology, 2010.
    This study provides a mechanistic explanation for the TGN1412 trial failure, revealing that the severe cytokine storm was caused by species-specific differences in CD28 expression on human CD4+ effector memory T-cells, which are absent or functionally different in the animal models used. This underscores the critical importance of human-specific in vitro testing for immunomodulatory drugs. ↩︎ ↩︎

  20. Wegner, J., Hackenberg, S., Scholz, C. J., et al., “High-density preculture of PBMCs restores defective sensitivity of circulating CD8 T cells to virus- and tumor-derived antigens”, Blood, 2015.
    This paper builds on the lessons learned from the TGN1412 incident by demonstrating that high-density preculture of PBMCs can restore the sensitivity of human T cells to specific antigens in vitro. It advocates for optimized in vitro human cell culture techniques as essential tools for accurately evaluating the safety and efficacy of immunotherapies prior to clinical application. ↩︎ ↩︎

  21. Kenter, M.J.H., Cohen, A.F., “The return of the prodigal son and the extraordinary development route of antibody TGN1412 - lessons for drug development and clinical pharmacology”, Br J Clin Pharmacol, 2015.
    This article reflects on the extraordinary development and subsequent failure of the TGN1412 antibody, extracting critical lessons for clinical pharmacology and drug development. It emphasizes the need for rigorous, human-relevant preclinical safety assessments and cautious, step-wise dose escalation in early-phase human trials to prevent catastrophic immune reactions. ↩︎

  22. Attarwala, H., “TGN1412: From Discovery to Disaster”, J Young Pharm, 2010.
    This review chronicles the development of TGN1412 from its promising preclinical discovery to the disastrous Phase 1 clinical trial that left multiple volunteers critically ill. It analyzes the regulatory and scientific failures that allowed the trial to proceed, advocating for stricter reliance on human-based in vitro models to predict cytokine release syndrome before human exposure. ↩︎

  23. Institute of Medicine (US) Committee to Review the Fialuridine (FIAU/FIAC) Clinical Trials; Manning, F. J., Swartz, M., editors, “Review of the Fialuridine (FIAU) Clinical Trials”, National Academies Press (US), 1995.
    This comprehensive report investigates the fatal clinical trials of fialuridine (FIAU), a drug that caused severe, delayed liver failure and lactic acidosis in humans despite showing no such toxicity in extensive animal testing. The review highlights the critical failure of animal models to predict human-specific mitochondrial toxicity and calls for improved preclinical screening methods. ↩︎ ↩︎

  24. ScienceDirect Topics, Fialuridine, Immunology and Microbiology, n.d.
    This reference entry provides a scientific overview of fialuridine, an investigational nucleoside analogue for hepatitis B that was withdrawn from clinical trials due to unexpected and fatal hepatotoxicity in humans. It summarizes the drug’s mechanism of action and the subsequent scientific focus on its unique, species-specific mitochondrial toxicity that animal models failed to predict. ↩︎ ↩︎

  25. Lewis, W., et al., Fialuridine and its metabolites inhibit DNA polymerase gamma at sites of multiple adjacent analog incorporation, decrease mtDNA abundance, and cause mitochondrial structural defects in cultured hepatoblasts, Proceedings of the National Academy of Sciences, 1996.
    This study elucidates the molecular mechanism behind fialuridine’s fatal hepatotoxicity, demonstrating that the drug and its metabolites specifically inhibit human DNA polymerase gamma, leading to mitochondrial DNA depletion and structural defects in human liver cells. This human-specific mechanism explains why standard animal models, which lack this specific metabolic vulnerability, failed to predict the drug’s toxicity. ↩︎ ↩︎

  26. Lee, E. W., Lai, Y., Zhang, H., and Unadkat, J. D., Identification of the Mitochondrial Targeting Signal of the Human Equilibrative Nucleoside Transporter 1 (hENT1): Implications for Interspecies Differences in Mitochondrial Toxicity of Fialuridine, Journal of Biological Chemistry, 2006.
    This research identifies the specific mitochondrial targeting signal in the human equilibrative nucleoside transporter 1 (hENT1) that allows fialuridine to accumulate in human mitochondria, causing toxicity. The study highlights a crucial interspecies difference, explaining why animal models were protected from the drug’s effects and underscoring the necessity of human-specific in vitro testing for nucleoside analogues. ↩︎ ↩︎

  27. McKenzie, R., Fried, M.W., Sallie, R., et al., “Hepatic Failure and Lactic Acidosis Due to Fialuridine (FIAU), an Investigational Nucleoside Analogue for Chronic Hepatitis B”, N Engl J Med, 1995.
    This clinical report details the tragic outcomes of the fialuridine trials, where patients developed severe, often fatal hepatic failure and lactic acidosis after prolonged exposure. It serves as a primary clinical documentation of the drug’s delayed and species-specific toxicity, fundamentally challenging the reliability of preclinical animal safety data for this class of compounds. ↩︎

  28. Xu, D., Nishimura, T., Nishimura, S., et al., Fialuridine Induces Acute Liver Failure in Chimeric TK-NOG Mice: A Model for Detecting Hepatic Drug Toxicity Prior to Human Testing, PLOS Medicine, 2014.
    This study presents a novel chimeric mouse model (TK-NOG) engrafted with human hepatocytes that successfully replicated the fatal liver failure caused by fialuridine. It demonstrates the potential of humanized animal models and in vitro human liver systems to detect human-specific hepatotoxicity that traditional rodent models completely miss. ↩︎

  29. Vieira, R., Souto, S. B., Sánchez-López, E., et al., Sugar-Lowering Drugs for Type 2 Diabetes Mellitus and Metabolic Syndrome—Review of Classical and New Compounds: Part-I, Pharmaceuticals, 2019.
    This review examines various classical and novel sugar-lowering drugs for Type 2 diabetes, including their mechanisms, efficacy, and safety profiles. It contextualizes the historical challenges in diabetes drug development, where some compounds have been withdrawn due to unforeseen adverse effects, highlighting the ongoing need for rigorous, human-relevant safety assessments. ↩︎ ↩︎

  30. Graham, D. J., Green, L., Senior, J. R., and Nourjah, P., Troglitazone-induced liver failure: a case study, The American Journal of Medicine, 2003.
    This case study analyzes the withdrawal of troglitazone, a diabetes drug that caused idiosyncratic, fatal liver failure in a subset of patients despite passing standard preclinical animal safety tests. It highlights the difficulty of predicting rare, human-specific idiosyncratic drug reactions using traditional animal models and underscores the need for better in vitro human hepatotoxicity screening. ↩︎ ↩︎ ↩︎

  31. Yang, H., et al., “Systems Pharmacology Modeling Predicts Delayed Presentation and Species Differences in Bile Acid-Mediated Troglitazone Hepatotoxicity”, PLoS Computational Biology, 2014.
    This paper utilizes systems pharmacology modeling to explain the delayed and species-specific hepatotoxicity of troglitazone, focusing on its inhibition of the bile salt export pump (BSEP). The computational model successfully predicts the human-specific cholestatic potential of the drug, demonstrating how in silico approaches can identify toxicity risks that animal studies fail to capture. ↩︎

  32. Funk, C., et al., “Cholestatic potential of troglitazone as a possible factor contributing to troglitazone-induced hepatotoxicity: in vivo and in vitro interaction at the canalicular bile salt export pump (Bsep) in the rat”, Molecular Pharmacology, 2001.
    This study investigates the cholestatic potential of troglitazone by examining its interaction with the bile salt export pump (BSEP) in both in vivo and in vitro rat models. While it provides insights into the drug’s mechanism, it also highlights the limitations of relying solely on rat models, as the human BSEP exhibits different sensitivity, contributing to the drug’s idiosyncratic human hepatotoxicity. ↩︎ ↩︎

  33. Zhang, J., He, K., et al., “Inhibition of bile salt transport by drugs associated with liver injury in primary hepatocytes from human, monkey, dog, rat, and mouse”, Chem Biol Interact, 2016.
    This comparative study evaluates the inhibition of bile salt transport by various hepatotoxic drugs across primary hepatocytes from humans and multiple animal species. The findings reveal significant interspecies differences in drug sensitivity, emphasizing that human hepatocytes are the most reliable model for predicting human-specific drug-induced liver injury (DILI) and bile acid accumulation. ↩︎ ↩︎ ↩︎

  34. Vansant, G., Pezzoli, P., Monforte, J., “Gene Expression Analysis of Troglitazone Reveals Its Impact on Multiple Pathways in Cell Culture: A Case for In Vitro Platforms Combined with Gene Expression Analysis for Early (Idiosyncratic) Toxicity Screening”, Toxicology in Vitro, 2006.
    This research demonstrates how gene expression analysis in human in vitro cell cultures can reveal the complex, multi-pathway impacts of troglitazone, including early signs of idiosyncratic toxicity. It advocates for the integration of transcriptomics with in vitro human models as a powerful, early screening tool to identify hepatotoxic risks that traditional animal studies overlook. ↩︎

  35. Masubuchi, Y., “Metabolic and non-metabolic factors determining troglitazone hepatotoxicity: a review”, Drug Metab Pharmacokinet, 2006.
    This review synthesizes current knowledge on the metabolic and non-metabolic factors that contribute to troglitazone-induced hepatotoxicity. It discusses the role of reactive metabolites, mitochondrial dysfunction, and immune responses, concluding that the complex, idiosyncratic nature of this toxicity in humans cannot be adequately modeled in standard animal testing, necessitating advanced human-relevant in vitro systems. ↩︎

  36. Bombardier, C., et al., “Comparison of upper gastrointestinal toxicity of rofecoxib and naproxen in patients with rheumatoid arthritis”, New England Journal of Medicine, 2000.
    This landmark clinical trial compared the gastrointestinal safety of the COX-2 inhibitor rofecoxib (Vioxx) with naproxen, initially showing a favorable GI profile for rofecoxib. However, this study set the stage for later revelations about the drug’s cardiovascular risks, illustrating how a narrow focus on specific endpoints in clinical trials can mask other severe, systemic toxicities not predicted by preclinical animal models. ↩︎ ↩︎

  37. Bresalier, R. S., et al., “Cardiovascular Events Associated with Rofecoxib in a Colorectal Adenoma Chemoprevention Trial”, New England Journal of Medicine, 2005.
    This pivotal study revealed a significantly increased risk of serious cardiovascular events, such as heart attacks and strokes, in patients taking rofecoxib compared to placebo. The findings led to the drug’s worldwide withdrawal and highlighted a major failure of preclinical animal testing to predict the human-specific cardiovascular risks associated with selective COX-2 inhibition. ↩︎

  38. DeCanio, Samuel, Cost benefit analysis and the FDA: measuring the costs and benefits of drug approval under the PDUFA I-II, 1998–2005, Journal of Regulatory Economics, 2024.
    This economic analysis evaluates the cost-benefit dynamics of the FDA’s drug approval process during the PDUFA I-II era, a period that included the approval and subsequent withdrawal of rofecoxib. It discusses the regulatory pressures and systemic challenges in balancing rapid drug access with rigorous safety evaluation, emphasizing the high societal costs of approving drugs with undetected, severe adverse effects. ↩︎ ↩︎

  39. Jüni, P., et al., “Risk of cardiovascular events and rofecoxib: cumulative meta-analysis”, The Lancet, 2004.
    This cumulative meta-analysis aggregates data from multiple trials to demonstrate a clear, dose-dependent increase in cardiovascular risk associated with rofecoxib use. The study critically examines how early signals of cardiovascular toxicity were overlooked or underweighted, underscoring the need for more robust, human-relevant preclinical models to detect class-specific cardiovascular liabilities before large-scale human exposure. ↩︎

  40. Curfman, G. D., Morrissey, S., Drazen, J. M., “Expression of Concern: Cardiovascular Events Associated with Rofecoxib in a Colorectal Adenoma Chemoprevention Trial”, New England Journal of Medicine, 2004.
    This editorial expresses profound concern over the delayed recognition and communication of the cardiovascular risks associated with rofecoxib. It critiques the regulatory and scientific processes that allowed the drug to remain on the market despite emerging evidence of harm, highlighting the critical need for transparent, proactive safety monitoring and better predictive preclinical models. ↩︎

  41. McAdam, B. F., Catella-Lawson, F., Mardini, I. A., et al., Systemic biosynthesis of prostacyclin by cyclooxygenase-2 in humans: the human pharmacology of a selective inhibitor of COX-2, Proceedings of the National Academy of Sciences, 1999.
    This study investigates the human pharmacology of selective COX-2 inhibitors, demonstrating their profound suppression of systemic prostacyclin biosynthesis without affecting thromboxane production. This mechanistic insight explains the prothrombotic state and increased cardiovascular risk observed with drugs like rofecoxib, a human-specific physiological imbalance that was not adequately predicted by standard animal models. ↩︎ ↩︎

  42. Fitzgerald, G. A., “Coxibs and cardiovascular disease”, New England Journal of Medicine, 2004.
    This perspective article discusses the mechanistic link between COX-2 inhibitors (coxibs) and increased cardiovascular disease risk, focusing on the disruption of the balance between pro-thrombotic thromboxane and anti-thrombotic prostacyclin. It serves as a critical analysis of how the rofecoxib tragedy reshaped our understanding of NSAID safety and highlighted the limitations of animal models in predicting complex human cardiovascular physiology. ↩︎ ↩︎

  43. Funk, C., FitzGerald, G., “COX-2 Inhibitors and Cardiovascular Risk”, Journal of Cardiovascular Pharmacology, 2007.
    This review elaborates on the cardiovascular risks associated with COX-2 inhibitors, detailing the molecular and physiological mechanisms that lead to an increased risk of thrombosis and hypertension. It emphasizes that the human-specific nature of these vascular responses makes traditional animal testing an unreliable predictor of cardiovascular safety for this class of drugs. ↩︎

  44. Liu, J., Li, N., Yang, J., Hammock, B.D., “Metabolic profiling of murine plasma reveals an unexpected biomarker in rofecoxib-mediated cardiovascular events”, PNAS, 2010.
    This study uses metabolic profiling in murine models to identify potential biomarkers associated with rofecoxib-mediated cardiovascular events. While it attempts to find predictive markers in animals, the findings also underscore the ongoing challenge of translating murine metabolic responses to human clinical outcomes, reinforcing the need for human-specific in vitro and in silico models. ↩︎

  45. Brenner, G., et al., “Hidden Cardiotoxicity of Rofecoxib Can be Revealed in Experimental Models of Ischemia/Reperfusion”, Cells, 2020.
    This research demonstrates that the cardiotoxic effects of rofecoxib, which were missed in standard preclinical safety studies, can be unmasked using specialized experimental models of ischemia/reperfusion. It suggests that more rigorous, stress-tested in vitro and ex vivo human tissue models are necessary to reveal the “hidden” toxicities of drugs before they are approved for human use. ↩︎

  46. Temporary Specialist Scientific Committee (TSSC), “Report on the causes of the accident during a phase 1 clinical trial in Rennes in January 2016”, Agence Nationale de Sécurité du Médicament (ANSM), 2016.
    This official French government report investigates the tragic Phase 1 clinical trial of the FAAH inhibitor BIA 10-2474, which resulted in one death and severe neurological damage in several healthy volunteers. The report critically analyzes the preclinical data, concluding that the animal models used were inadequate for predicting the drug’s severe, human-specific neurotoxicity and off-target effects. ↩︎ ↩︎

  47. Kerbrat, A., et al., “Acute Neurologic Disorder from an Inhibitor of Fatty Acid Amide Hydrolase”, New England Journal of Medicine, 2016.
    This clinical case series describes the acute, severe neurological disorders observed in participants of the BIA 10-2474 Phase 1 trial, detailing the clinical presentation and progression of the toxicity. It serves as a stark clinical documentation of the catastrophic failure of preclinical animal testing to predict the profound neurotoxic potential of this specific FAAH inhibitor in humans. ↩︎ ↩︎ ↩︎

  48. Van der Stelt, M., et al., “Revisiting what went wrong in the clinical trial tragedy in France”, Chemical & Engineering News, 2017.
    This scientific commentary revisits the BIA 10-2474 clinical trial tragedy, analyzing the pharmacological and regulatory missteps that led to the disaster. It highlights the dangers of over-relying on animal data that showed no toxicity, arguing for the mandatory inclusion of human-relevant in vitro assays and more rigorous off-target screening before advancing novel neurological compounds to human trials. ↩︎ ↩︎

  49. van Esbroeck, A. C. M., et al., “Activity-based protein profiling reveals off-target proteins of the FAAH inhibitor BIA 10-2474”, Science, 2017.
    This study uses activity-based protein profiling to identify that BIA 10-2474, unlike other FAAH inhibitors, potently inhibits several off-target proteins, including MAGL and ABHD6. This discovery provides a plausible molecular explanation for the severe neurotoxicity observed in humans, demonstrating how advanced human-relevant biochemical assays can uncover dangerous off-target effects missed by traditional animal testing. ↩︎ ↩︎

  50. Bonifácio, M. J., et al., “Preclinical pharmacological evaluation of the fatty acid amide hydrolase inhibitor BIA 10‐2474”, Br J Pharmacol, 2020.
    This paper presents the original preclinical pharmacological data for BIA 10-2474, which showed a favorable safety profile in various animal models. In retrospect, this study is frequently cited as a prime example of how standard animal toxicology can fail to detect severe, human-specific off-target toxicities, highlighting the critical need for more predictive, human-based New Approach Methodologies. ↩︎

  51. Haigwood, N. L., and Walker, C. M., Commissioned Paper: Comparison of Immunity to Pathogens in Humans, Chimpanzees, and Macaques, Institute of Medicine (US) Committee on the Use of Chimpanzees in Biomedical and Behavioral Research, 2011.
    This commissioned paper compares the immune responses of humans, chimpanzees, and macaques to various pathogens, highlighting significant interspecies differences in immune system architecture and function. It concludes that while non-human primates share some similarities with humans, their immune responses are distinct enough to make them unreliable predictors of human vaccine efficacy or immunopathology. ↩︎

  52. Akhtar, A., “The flaws and human harms of animal experimentation”, Camb Q Healthc Ethics, 2015.
    This article provides a comprehensive critique of animal experimentation, arguing that it is not only ethically problematic but also scientifically flawed due to profound interspecies differences. The author presents evidence that reliance on animal models has frequently led to human harm in clinical trials and advocates for a decisive shift toward human-relevant, non-animal research methodologies. ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎

  53. Watkins, D. I., et al., “Nonhuman primate models and the failure of the Merck HIV-1 vaccine in humans”, Nature Medicine, 2008.
    This paper analyzes the failure of the Merck HIV-1 vaccine in human clinical trials, despite it showing promising results in non-human primate models. The authors demonstrate that the immune responses elicited in macaques did not translate to protective immunity in humans, underscoring the fundamental limitations of primate models in predicting human vaccine efficacy against complex viruses like HIV. ↩︎

  54. Sekaly, R. P., “The failed HIV Merck vaccine study: a step back or a launching point for future vaccine development?”, Journal of Experimental Medicine, 2008.
    This perspective discusses the implications of the failed Merck HIV vaccine trial, exploring whether the reliance on non-human primate models misled researchers about the vaccine’s potential. It argues that the failure should serve as a catalyst for developing more accurate, human-specific in vitro and computational models to evaluate vaccine candidates before proceeding to costly and risky human trials. ↩︎

  55. Tang, H., et al., “Human Organs-on-Chips for Virology”, Trends in Microbiology, 2020.
    This review highlights the emerging role of human organs-on-chips in virology research, demonstrating their ability to model complex human-virus interactions, such as SARS-CoV-2 infection, more accurately than traditional animal models. It emphasizes how these microphysiological systems can accelerate antiviral drug discovery and vaccine development by providing human-relevant data on viral pathogenesis and host responses. ↩︎ ↩︎

  56. Greenwood, E. J. D., et al., “Simian Immunodeficiency Virus Infection of Chimpanzees (Pan troglodytes) Shares Features of Both Pathogenic and Non-pathogenic Lentiviral Infections”, PLoS Pathogens, 2015.
    This study examines SIV infection in chimpanzees, revealing that their immune response shares features of both pathogenic and non-pathogenic lentiviral infections, differing significantly from human HIV infection. This finding further complicates the use of chimpanzees as predictive models for HIV/AIDS, highlighting the unique evolutionary adaptations that make human-specific research models essential. ↩︎

  57. Heeney, J., et al., “Immune strategies utilized by lentivirus infected chimpanzees to resist progression to AIDS”, Immunology Letters, 1996.
    This research investigates the unique immune strategies that allow chimpanzees to resist progression to AIDS despite lentivirus infection. By highlighting the distinct immunological mechanisms in chimpanzees that are absent in humans, the study underscores why chimpanzee models are poor predictors of human HIV disease progression and therapeutic responses. ↩︎

  58. de Groot, N. G., Bontrop, R. E., “The HIV-1 pandemic: does the selective sweep in chimpanzees mirror humankind’s future?”, Retrovirology, 2013.
    This paper explores the evolutionary history of HIV-1 and its simian precursors, questioning whether the genetic adaptations seen in chimpanzees can predict human responses to the virus. The authors conclude that the distinct evolutionary paths of humans and chimpanzees make such extrapolations unreliable, reinforcing the need for human-centric research models in HIV studies. ↩︎

  59. Sharp, P. M., and Hahn, B. H., “The evolution of HIV-1 and the origin of AIDS”, Philos Trans R Soc Lond B Biol Sci, 2010.
    This comprehensive review traces the evolutionary origins of HIV-1 from simian immunodeficiency viruses (SIVs) in chimpanzees and gorillas. While it provides crucial insights into the zoonotic transmission of the virus, it also highlights the vast genetic and immunological divergence between SIV in natural hosts and HIV-1 in humans, limiting the predictive value of animal models for human disease. ↩︎

  60. Girard, M., et al., “Failure of a human immunodeficiency virus type 1 (HIV-1) subtype B-derived vaccine to prevent infection of chimpanzees by an HIV-1 subtype E strain”, J Virol, 1996.
    This study documents the failure of an HIV-1 subtype B vaccine to protect chimpanzees from infection by a subtype E strain, highlighting the challenges of viral diversity and cross-protection. It illustrates the complexities of HIV vaccine development and the limitations of chimpanzee models in predicting the nuanced, strain-specific immune responses required for human vaccine efficacy. ↩︎

  61. Mak, I. W., Evaniew, N., and Ghert, M., Lost in translation: animal models and clinical trials in cancer treatment, American Journal of Translational Research, 2014.
    This review examines the high rate of failure in translating promising cancer treatments from animal models to successful human clinical trials. The authors attribute this “lost in translation” phenomenon to the oversimplified nature of animal tumor models, which fail to replicate the complex genetic heterogeneity and microenvironment of human cancers, advocating for patient-derived models. ↩︎ ↩︎ ↩︎ ↩︎

  62. Shao, C., et al., Modeling cancer patient populations in mice: complex genetic and environmental factors, Drug Discovery Today: Disease Models, 2008.
    This paper discusses the challenges of accurately modeling human cancer patient populations in mice, emphasizing the profound influence of complex genetic backgrounds and environmental factors on tumor biology. It argues that standard murine models are insufficient for predicting human drug responses, highlighting the need for more sophisticated, humanized, or patient-derived xenograft models. ↩︎ ↩︎

  63. Wong, C. H., Siah, K. W., and Lo, A. W., Estimation of clinical trial success rates and related parameters, Biostatistics, 2019.
    This large-scale statistical analysis estimates the success rates of drugs transitioning from preclinical development through various phases of clinical trials. The data reveal a stark attrition rate, particularly in oncology, underscoring the poor predictive validity of traditional preclinical animal models and the urgent need for more reliable, human-relevant screening technologies to improve clinical trial success. ↩︎ ↩︎

  64. Verduin, M., Hoeben, A., De Ruysscher, D., and Vooijs, M., Patient-Derived Cancer Organoids as Predictors of Treatment Response, Frontiers in Oncology, 2021.
    This review highlights the growing utility of patient-derived cancer organoids as highly predictive models for individualized treatment response. By retaining the genetic and phenotypic characteristics of the original tumor, these 3D in vitro models offer a superior alternative to traditional animal models for testing drug efficacy and guiding personalized cancer therapy. ↩︎

  65. Lipreri, M. V., Totaro, M. T., Baldini, N., and Avnet, S., From Spheroids to Tumor-on-a-Chip for Cancer Modeling and Therapeutic Testing, Micromachines, 2024.
    This article traces the evolution of in vitro cancer models from simple 2D cultures and spheroids to advanced “tumor-on-a-chip” microphysiological systems. It emphasizes how these complex, human-relevant models better replicate the tumor microenvironment and drug penetration dynamics, offering a more accurate and ethical alternative to animal testing for oncology drug discovery. ↩︎ ↩︎

  66. Chakrabarty, S., et al., A Microfluidic Cancer-on-Chip Platform Predicts Drug Response Using Organotypic Tumor Slice Culture, Cancer Research, 2022.
    This study presents a novel microfluidic cancer-on-chip platform that integrates organotypic tumor slice cultures to predict patient-specific drug responses. The platform successfully maintains the native tumor architecture and microenvironment, demonstrating superior predictive accuracy for chemotherapy efficacy compared to traditional animal models, paving the way for personalized oncology. ↩︎ ↩︎ ↩︎

  67. Rae, C., Amato, F., Braconi, C., “Patient-Derived Organoids as a Model for Cancer Drug Discovery”, International Journal of Molecular Sciences, 2021.
    This review discusses the application of patient-derived organoids (PDOs) in cancer drug discovery, highlighting their ability to recapitulate the genetic and histological features of human tumors. The authors argue that PDOs represent a paradigm shift in preclinical testing, offering a scalable, human-relevant alternative to animal models for identifying novel therapeutic targets and predicting clinical outcomes. ↩︎

  68. Foo, M. A., et al., “Clinical translation of patient-derived tumour organoids- bottlenecks and strategies”, Biomarker Research, 2022.
    This paper examines the challenges and strategies associated with translating patient-derived tumor organoids from the laboratory to clinical practice. While acknowledging bottlenecks such as standardization and scalability, the authors emphasize that overcoming these hurdles is essential to fully replace animal models and realize the potential of organoids in personalized cancer medicine. ↩︎

  69. Hawkes, N., “Most laboratory cancer studies cannot be replicated, study shows”, BMJ, 2012.
    This news report highlights a major study revealing that a significant majority of preclinical cancer research findings, largely based on animal models, cannot be replicated. This “reproducibility crisis” underscores the inherent variability and lack of robustness in animal studies, raising serious questions about their reliability as a foundation for human clinical trials. ↩︎ ↩︎

  70. Haelle, T., Dozens of major cancer studies can’t be replicated, Science News, 2021.
    This article discusses the widespread inability to replicate dozens of high-profile cancer studies, many of which relied heavily on murine models. It emphasizes the scientific and financial waste generated by irreproducible animal research and advocates for more rigorous, transparent, and human-relevant methodologies to ensure that preclinical findings are robust and translatable. ↩︎

  71. Hentze, H., Ibsen, E., “The “reproducibility crisis” of animal studies in oncology - How did we get here, and how can we resolve it”, Cancer Res, 2019.
    This commentary addresses the reproducibility crisis specifically within oncology animal studies, attributing it to factors like poor experimental design, publication bias, and biological variability. The authors propose solutions such as stricter reporting guidelines, increased use of human in vitro models, and a cultural shift toward validating findings in human-relevant systems before advancing to clinical trials. ↩︎

  72. O’Collins, V. E., et al., “The failure of animal models of neuroprotection in acute ischemic stroke to translate to clinical efficacy”, Translational Stroke Research, 2006.
    This systematic review analyzes the historical failure of neuroprotective drugs that showed remarkable efficacy in animal models of acute ischemic stroke but consistently failed in human clinical trials. The authors identify fundamental flaws in animal study design, such as the use of young, healthy animals without comorbidities, which do not reflect the human stroke population. ↩︎

  73. Schmidt-Pogoda, A., et al., “Why Most Acute Stroke Studies Are Positive in Animals but Not in Patients”, Annals of Neurology, 2020.
    This paper investigates the discrepancy between positive preclinical results in animal stroke models and negative outcomes in human trials. It highlights critical methodological differences, including the timing of drug administration and the lack of relevant comorbidities in animal models, arguing that these factors render traditional animal studies poor predictors of human stroke treatment efficacy. ↩︎

  74. Fisher, M., et al., “Update of the Stroke Therapy Academic Industry Roundtable Preclinical Recommendations”, Stroke, 2009.
    This update to the STAIR recommendations provides revised guidelines for preclinical stroke research, emphasizing the need for more rigorous, clinically relevant animal models. It calls for the inclusion of aged animals, comorbidities, and long-term functional outcomes, while also encouraging the integration of human in vitro models to improve the translational value of preclinical stroke research. ↩︎ ↩︎

  75. Ingber, D. E., “Human organs-on-chips for disease modelling, drug development and personalized medicine”, Nature Reviews Genetics, 2022.
    This comprehensive review by the pioneer of organ-on-a-chip technology details how these microfluidic devices are revolutionizing disease modeling and drug development. Ingber explains how organs-on-chips replicate the complex mechanical and biochemical microenvironments of human tissues, offering a highly predictive, human-relevant alternative to animal testing for personalized medicine. ↩︎ ↩︎

  76. Nzou, G., et al., “Multicellular 3D neurovascular unit model for assessing hypoxia and neuroinflammation induced blood-brain barrier dysfunction”, Scientific Reports, 2020.
    This study presents a sophisticated 3D in vitro model of the human neurovascular unit, designed to assess blood-brain barrier dysfunction under conditions of hypoxia and neuroinflammation. The model successfully replicates complex human pathophysiological responses, demonstrating the potential of advanced in vitro systems to replace animal models in neurological drug safety and efficacy testing. ↩︎

  77. McCann, S. K., and Lawrence, C. B., Comorbidity and age in the modelling of stroke: are we still failing to consider the characteristics of stroke patients?, BMJ Open Science, 2020.
    This critical review questions whether current preclinical stroke models adequately account for the age and comorbidities (e.g., hypertension, diabetes) typical of human stroke patients. The authors argue that the continued use of young, healthy animal models is a primary reason for the translational failure of neuroprotective therapies, urging a shift toward more clinically representative models. ↩︎

  78. Archer, D. P., Walker, A. M., McCann, S. K., Moser, J. J., and Appireddy, R. M., Anesthetic Neuroprotection in Experimental Stroke in Rodents: A Systematic Review and Meta-analysis, Anesthesiology, 2017.
    This meta-analysis reveals that the use of anesthetics in rodent stroke models, which are themselves neuroprotective, has confounded decades of preclinical research, leading to false-positive results that failed to translate to humans. This finding highlights a critical, often-overlooked methodological flaw in animal testing that underscores the need for more controlled, human-relevant in vitro systems. ↩︎

  79. Wevers, N. R., and De Vries, H. E., Microfluidic models of the neurovascular unit: a translational view, Fluids and Barriers of the CNS, 2023.
    This review explores the development and translational potential of microfluidic models of the neurovascular unit (NVU). It highlights how these “brain-on-a-chip” systems can accurately model human blood-brain barrier function and neuroinflammation, providing a superior, human-specific platform for drug screening and toxicity assessment compared to traditional animal models. ↩︎

  80. Wevers, N. R., Nair, A. L., Fowke, T. M., et al., Modeling ischemic stroke in a triculture neurovascular unit on-a-chip, Fluids and Barriers of the CNS, 2021.
    This study demonstrates the successful modeling of ischemic stroke using a triculture neurovascular unit-on-a-chip, incorporating human endothelial cells, astrocytes, and neurons. The model accurately recapitulates the cellular and molecular responses to ischemia, proving that complex human pathophysiology can be studied in vitro, reducing the reliance on inadequate animal stroke models. ↩︎

  81. Connor, K. T., and Aylward, L. L., Human response to dioxin: aryl hydrocarbon receptor (AhR) molecular structure, function, and dose-response data for enzyme induction indicate an impaired human AhR, Journal of Toxicology and Environmental Health, Part B: Critical Reviews, 2006.
    This review analyzes the human response to dioxin, demonstrating that the human aryl hydrocarbon receptor (AhR) is significantly less sensitive than those of commonly used animal models like mice and rats. This profound interspecies difference explains why animal toxicity data for dioxin often overestimate human risk, highlighting the necessity of human-specific in vitro assays for accurate risk assessment. ↩︎ ↩︎

  82. Mossman, B. T., et al., Pulmonary Endpoints (Lung Carcinomas and Asbestosis) Following Inhalation Exposure to Asbestos, Journal of Toxicology and Environmental Health, Part B: Critical Reviews, 2011.
    This paper reviews the pulmonary endpoints, including lung carcinomas and asbestosis, following inhalation exposure to asbestos in both animal models and human epidemiological studies. It discusses the challenges in extrapolating animal data to human risk, emphasizing the species-specific differences in lung clearance and fiber pathogenicity that complicate the use of animals for asbestos safety testing. ↩︎

  83. Wright, J. L., and Churg, A., Animal models of cigarette smoke-induced COPD, Chest, 2002.
    This review evaluates various animal models used to study cigarette smoke-induced chronic obstructive pulmonary disease (COPD). The authors conclude that while these models can replicate some features of the disease, they fail to fully capture the complex, chronic inflammatory and remodeling processes seen in human COPD, limiting their predictive value for therapeutic development. ↩︎

  84. Agency for Toxic Substances and Disease Registry (ATSDR), Toxicological Profile for Asbestos, U.S. Department of Health and Human Services, 2001.
    This comprehensive toxicological profile details the health effects, exposure pathways, and regulatory status of asbestos. It explicitly acknowledges the limitations and interspecies differences in animal studies used to assess asbestos toxicity, emphasizing the reliance on human epidemiological data and the growing need for human-relevant in vitro models to understand fiber pathogenicity. ↩︎

  85. Lee, S. Y., Lee, D. Y., Kang, J. H., et al., Alternative experimental approaches to reduce animal use in biomedical studies, Journal of Drug Delivery Science and Technology, 2022.
    This review surveys a range of alternative experimental approaches, including in vitro cell cultures, organoids, and in silico modeling, designed to reduce or replace animal use in biomedical research. It highlights the scientific advantages of these human-relevant methods, such as improved predictive accuracy and ethical benefits, advocating for their broader adoption in regulatory and preclinical testing. ↩︎

  86. Dere, E., Lee, A. W., Burgoon, L. D., Zacharewski, T. R., “Differences in TCDD-elicited gene expression profiles in human HepG2, mouse Hepa1c1c7 and rat H4IIE hepatoma cells”, BMC Genomics, 2011.
    This study compares the gene expression profiles elicited by TCDD (dioxin) in human, mouse, and rat hepatoma cell lines, revealing significant interspecies differences in transcriptional responses. The findings demonstrate that rodent cells are poor surrogates for predicting human-specific toxicological responses to environmental chemicals, underscoring the need for human-based in vitro testing. ↩︎

  87. Lawrence, G. S., Gobas, F. A., “A pharmacokinetic analysis of interspecies extrapolation in dioxin risk assessment”, Chemosphere, 1997.
    This pharmacokinetic analysis examines the challenges of extrapolating dioxin toxicity data from animals to humans for risk assessment purposes. The authors highlight profound differences in half-life, metabolism, and tissue distribution between species, concluding that direct extrapolation from animal models is highly uncertain and that human-specific in vitro data are essential for accurate risk evaluation. ↩︎

  88. Muhle, H., Pott, F., “Asbestos as reference material for fibre-induced cancer”, Int Arch Occup Environ Health, 2000.
    This paper discusses the use of asbestos as a reference material for studying fiber-induced cancer in animal models. It critically evaluates the limitations of these models, noting that the mechanisms of fiber clearance and tumor development in rodents differ significantly from humans, which complicates the regulatory classification and risk assessment of novel synthetic fibers. ↩︎

  89. Wright, J. L., Cosio, M., and Churg, A., “Animal models of chronic obstructive pulmonary disease”, American Journal of Physiology-Lung Cellular and Molecular Physiology, 2008.
    This review critically assesses the utility of animal models in COPD research, concluding that they are inadequate for replicating the full spectrum of human disease, particularly the irreversible airway remodeling and complex inflammatory profiles. The authors advocate for the development of more sophisticated, human-relevant in vitro models to improve the discovery of effective COPD therapies. ↩︎

  90. Agraval, H., and Chu, H. W., “Lung Organoids in Smoking Research: Current Advances and Future Promises”, Biomolecules, 2022.
    This article highlights the emerging role of lung organoids in smoking research, demonstrating their ability to model the complex cellular responses to cigarette smoke exposure in a human-relevant context. It argues that lung organoids offer a superior alternative to animal models for studying smoking-induced diseases like COPD and lung cancer, enabling more accurate mechanistic studies and drug testing. ↩︎

  91. Wang, S., Li, L., Yan, F., Gao, Y., Yang, S., and Xia, X., COVID-19 Animal Models and Vaccines: Current Landscape and Future Prospects, Vaccines, 2021.
    This review evaluates the various animal models used in COVID-19 research and vaccine development, highlighting their limitations in fully recapitulating human disease pathology and immune responses. It emphasizes the need to complement animal studies with human-relevant models, such as organoids and humanized mice, to ensure the safety and efficacy of vaccines before widespread human use. ↩︎ ↩︎

  92. Bao, L., et al., The pathogenicity of SARS-CoV-2 in hACE2 transgenic mice, Nature, 2020.
    This study investigates the pathogenicity of SARS-CoV-2 in human ACE2 transgenic mice, demonstrating that while these mice can be infected, their disease presentation and immune response differ significantly from human COVID-19. This highlights the inherent limitations of even genetically modified animal models in perfectly mimicking complex human viral infections. ↩︎ ↩︎ ↩︎ ↩︎

  93. Han, R., Su, L., and Cheng, L., Advancing Human Vaccine Development Using Humanized Mouse Models, Vaccines, 2024.
    This paper discusses the use of humanized mouse models, which are engrafted with human immune systems, to advance vaccine development. While acknowledging their improved relevance over standard mice, the authors note that these models still have limitations in fully replicating human immune complexity, underscoring the complementary need for in vitro human immune assays. ↩︎ ↩︎

  94. Ghosh, A., Larrondo-Petrie, M. M., and Pavlovic, M., Revolutionizing Vaccine Development for COVID-19: A Review of AI-Based Approaches, Information, 2023.
    This review explores how artificial intelligence and machine learning are revolutionizing COVID-19 vaccine development by predicting antigenicity, optimizing sequences, and modeling immune responses in silico. It highlights how AI-driven approaches can accelerate the discovery process and reduce reliance on traditional, time-consuming animal testing by identifying the most promising candidates computationally. ↩︎ ↩︎

  95. Olawade, D. B., Teke, J., Fapohunda, O., et al., Leveraging artificial intelligence in vaccine development: A narrative review, Journal of Microbiological Methods, 2024.
    This narrative review examines the broad applications of artificial intelligence in vaccine development, from target identification to clinical trial optimization. The authors emphasize that AI and in silico modeling are becoming indispensable tools for predicting vaccine efficacy and safety, offering a faster, more ethical, and human-relevant alternative to conventional animal-based preclinical testing. ↩︎ ↩︎

  96. Lee, C. Y., and Lowen, A. C., Animal models for SARS-CoV-2, Current Opinion in Virology, 2021.
    This review catalogs the various animal models used to study SARS-CoV-2, including mice, hamsters, and non-human primates. It critically evaluates their utility, noting that while they provide some insights into viral transmission and pathogenesis, none fully replicate the complexity of human COVID-19, highlighting the need for human organoid and in vitro models for comprehensive study. ↩︎

  97. Moothedath, M., Muhamood, M., Bhosale, Y. S., et al., “COVID and Animal Trials: A Systematic Review”, Journal of Pharmacy and Bioallied Sciences, 2021.
    This systematic review analyzes the preclinical animal trials conducted for COVID-19 therapeutics and vaccines, assessing their predictive value for human clinical outcomes. The findings reveal significant discrepancies between animal efficacy and human trial results, reinforcing the argument that animal models are unreliable predictors of human responses to novel viral pathogens and treatments. ↩︎

  98. Han, Y., Yang, L., Lacko, L. A., and Chen, S., Human organoid models to study SARS-CoV-2 infection, Nature Methods, 2022.
    This paper demonstrates the utility of human organoid models, particularly lung and intestinal organoids, in studying SARS-CoV-2 infection and testing antiviral drugs. It highlights how these 3D human tissue models accurately recapitulate viral entry, replication, and host responses, providing a highly predictive and ethical alternative to animal models for COVID-19 research. ↩︎

  99. Singh, J., Malik, D., and Raina, A., Immuno-informatics approach for B-cell and T-cell epitope based peptide vaccine design against novel COVID-19 virus, Vaccine, 2021.
    This study utilizes an immuno-informatics approach to design peptide-based vaccines against SARS-CoV-2 by computationally predicting B-cell and T-cell epitopes. It showcases how in silico methods can rapidly identify promising vaccine candidates with high human immunogenicity potential, significantly reducing the need for extensive, preliminary animal screening. ↩︎

  100. U.S. Food and Drug Administration (FDA), Roadmap to Reducing Animal Testing in Preclinical Safety Studies, 2024.
    This official FDA roadmap outlines the agency’s strategic plan to reduce, refine, and ultimately replace animal testing in preclinical safety studies with scientifically valid New Approach Methodologies (NAMs). It details the regulatory pathways, validation processes, and collaborative efforts required to integrate human-relevant in vitro and in silico models into standard drug development practices. ↩︎

  101. Bracken, M. B., Why animal studies are often poor predictors of human reactions to exposure, Journal of the Royal Society of Medicine, 2009.
    This article systematically reviews the reasons why animal studies frequently fail to predict human reactions to chemical and drug exposures. The author cites profound genetic, metabolic, and physiological differences between species, arguing that reliance on animal data is not only scientifically flawed but also poses significant risks to human health in clinical trials. ↩︎

  102. Pound, P., and Ritskes-Hoitinga, M., Is it possible to overcome issues of external validity in preclinical animal research? Why most animal models are bound to fail, PLOS Biology, 2018.
    This critical paper argues that the external validity of preclinical animal research is fundamentally compromised by the artificial conditions of the laboratory environment and the genetic uniformity of test subjects. The authors conclude that most animal models are inherently bound to fail in predicting human clinical outcomes, advocating for a paradigm shift toward human biology-based research. ↩︎

  103. Humane World for Animals, Why is outdated animal research still being used when better science exists?, n.d.
    This advocacy article questions the continued reliance on outdated animal research methods despite the availability of more accurate, human-relevant New Approach Methodologies (NAMs). It highlights the scientific, ethical, and economic benefits of transitioning to in vitro and in silico models, urging regulatory agencies and researchers to modernize their testing paradigms. ↩︎

  104. Robertson, C. T., The Money Blind: How to Stop Industry Bias in Biomedical Science, without Violating the First Amendment, American Journal of Law & Medicine, 2011.
    This legal and ethical analysis examines the pervasive influence of industry funding on biomedical research, including the perpetuation of animal models that may serve commercial interests rather than scientific truth. The author proposes regulatory and structural reforms to mitigate this bias, promoting transparency and the adoption of more objective, human-relevant research methodologies. ↩︎

  105. Pellmar, T. C., and Eisenberg, L. (Eds.), Barriers to Interdisciplinary Research and Training, in Bridging Disciplines in the Brain, Behavioral, and Clinical Sciences, National Academies Press, 2000.
    This report identifies the institutional, cultural, and funding barriers that hinder interdisciplinary research and training in the biomedical sciences. It argues that overcoming these silos is essential for fostering the collaboration between biologists, engineers, and computational scientists needed to develop and validate innovative, non-animal research technologies. ↩︎

  106. Dzeng, E., Entrenched biases and structural incentives limit the influence of interdisciplinary research, Impact of Social Sciences, 2014.
    This commentary discusses how entrenched academic biases and structural funding incentives often marginalize interdisciplinary research, particularly in the development of alternative toxicology methods. The author argues that reforming these systemic incentives is crucial to accelerating the adoption of human-relevant, non-animal models in mainstream biomedical research. ↩︎

  107. Viceconti, M., Henney, A., and Morley-Fletcher, S., In silico clinical trials: how computer simulation will transform the biomedical industry, International Journal of Clinical Trials, 2016.
    This forward-looking article explores the concept of “in silico clinical trials,” where computer simulations and digital twins are used to predict drug safety and efficacy in virtual human populations. The authors argue that this technology has the potential to revolutionize the biomedical industry by significantly reducing the need for both animal testing and early-phase human trials. ↩︎

  108. People for the Ethical Treatment of Animals (PETA), Alternatives to Animal Testing, n.d.
    This resource provides a comprehensive overview of the scientifically validated alternatives to animal testing, including in vitro cell cultures, organ-on-a-chip technology, and advanced computer modeling. It emphasizes that these methods are not only more ethical but also more reliable and predictive of human responses than traditional animal experiments. ↩︎

  109. Gao, Q., Chow, S. K., Shinohara, I., et al., Can alternatives to animal testing yield useful information regarding biological mechanisms and drug discovery?, Journal of Orthopaedic Translation, 2025.
    This review evaluates the scientific validity of alternatives to animal testing, demonstrating that methods like 3D cell cultures, organoids, and microfluidics can indeed yield profound insights into biological mechanisms and drug discovery. The authors conclude that these human-relevant models are not just ethical alternatives, but scientifically superior tools for translational research. ↩︎

  110. Kakkad, R., Exploring alternatives to animal testing in drug discovery, Drug Target Review, 2023.
    This article explores the rapidly evolving landscape of alternatives to animal testing in drug discovery, highlighting advancements in organ-on-a-chip, AI-driven predictive toxicology, and patient-derived organoids. It argues that the pharmaceutical industry’s increasing adoption of these technologies is driven by both ethical imperatives and the need for more accurate, human-predictive data. ↩︎

  111. U.S. Congress, FDA Modernization Act 2.0 (S.5002), 117th Congress, 2022.
    This landmark legislation amends the Federal Food, Drug, and Cosmetic Act to remove the mandatory requirement for animal testing before human clinical trials for new drugs. It explicitly allows the use of New Approach Methodologies (NAMs), such as cell-based assays, organoids, and computer models, marking a historic regulatory shift toward human-relevant safety testing. ↩︎

  112. Buntz, B., Senate clears FDA Modernization Act 3.0, aiming to align FDA regulations with nonclinical-testing reforms, Drug Discovery & Development, 2025.
    This news report details the progression of FDA Modernization Act 3.0, which seeks to further align FDA regulations with ongoing nonclinical testing reforms. The legislation aims to accelerate the regulatory acceptance and implementation of New Approach Methodologies (NAMs), reinforcing the U.S. commitment to phasing out outdated animal testing in favor of more predictive, human-based science. ↩︎ ↩︎ ↩︎

  113. Huber, E., and McCullough, S., New Approach Methodologies (NAMs): Why Scientific Rigor Matters More Than Ever, RTI International, 2024.
    This article emphasizes the critical importance of scientific rigor in the development and validation of New Approach Methodologies (NAMs). It argues that to successfully replace animal models, NAMs must be held to the highest standards of reproducibility, transparency, and predictive validity to gain the trust of regulators, industry, and the scientific community. ↩︎

  114. Gerke, S., Balamut, J., and Wagner, J. K., The FDA’s plan to phase out animal testing, Trends in Biotechnology, 2026.
    This analysis examines the FDA’s strategic initiatives and regulatory guidance aimed at phasing out animal testing in preclinical drug development. The authors discuss the challenges and opportunities associated with this transition, emphasizing the need for robust validation frameworks to ensure that New Approach Methodologies (NAMs) reliably protect human health. ↩︎

  115. National Institutes of Health (NIH), Complement Animal Research In Experimentation (Complement-ARIE), NIH Common Fund, 2024.
    This program overview describes the NIH’s Complement-ARIE initiative, a major funding effort designed to accelerate the development, validation, and regulatory adoption of New Approach Methodologies (NAMs). The program focuses on creating advanced human-relevant models, data hubs, and validation networks to ultimately replace animal testing in biomedical research. ↩︎

  116. American Anti-Vivisection Society, “Problems with Animal Research”, n.d.
    This resource outlines the scientific, ethical, and practical problems associated with animal research, citing high failure rates in clinical translation and the inherent physiological differences between species. It advocates for a complete transition to modern, human-relevant research methods, such as in vitro models and computational biology, to improve human health outcomes. ↩︎ ↩︎ ↩︎

  117. U.S. Government Accountability Office, “National Institutes of Health: Assessing Efforts to Improve Animal Research Could Lead to Greater Human Health Benefits”, 2026.
    This GAO report evaluates the NIH’s efforts to improve the relevance and reproducibility of animal research. It concludes that while some progress has been made, the NIH must do more to actively promote and fund the development of New Approach Methodologies (NAMs), as transitioning away from flawed animal models could yield significantly greater human health benefits. ↩︎

  118. Vors, L., Debil, F., Saint-Cyr, et al., “A socio-ecological approach to the determinants of animal health management: A scoping review”, PLOS, 2026.
    This scoping review applies a socio-ecological framework to understand the complex determinants of animal health management in research settings. While focused on animal welfare, it implicitly highlights the growing ethical and scientific pressures on the research community to transition toward non-animal methodologies that eliminate the need for such management altogether. ↩︎ ↩︎

  119. The Comply Guide, “FDA Non-Animal Testing Guidance Draft: Regulatory Shift 2026”, 2026.
    This guide analyzes the FDA’s draft guidance on non-animal testing, highlighting a significant regulatory shift expected to solidify in 2026. It details how the new guidelines will streamline the submission and acceptance of New Approach Methodologies (NAMs) data, providing a clearer pathway for pharmaceutical companies to replace traditional animal studies with human-relevant alternatives. ↩︎ ↩︎

  120. NC3Rs, “ARRIVE: Animal Research Reporting In Vivo Experiments”, n.d.
    The ARRIVE guidelines provide a comprehensive checklist for reporting in vivo animal research, aimed at improving the transparency, reproducibility, and quality of animal studies. While designed to maximize the value of animal research, the guidelines also highlight the inherent complexities and variability of animal models, reinforcing the argument for transitioning to more standardized, human-relevant in vitro methods. ↩︎ ↩︎

  121. AVMA, “AVMA Guidelines for the Euthanasia of Animals”, n.d.
    These guidelines, established by the American Veterinary Medical Association, outline the accepted methods for the humane euthanasia of animals used in research and other settings. The existence and continuous revision of such guidelines underscore the ethical burden and inherent harm associated with animal experimentation, driving the scientific push toward alternative, non-animal testing methods. ↩︎

  122. NIH, “Regulatory References for Animal Welfare”, n.d.
    This NIH resource compiles the key federal laws, regulations, and policies governing the humane care and use of laboratory animals. It serves as a foundational reference for institutional animal care and use committees (IACUCs), while also highlighting the extensive regulatory framework required to manage animal research, contrasting with the streamlined potential of New Approach Methodologies. ↩︎

  123. NIH, “Animal Research Advisory Committee (ARAC) Guidelines”, n.d.
    These guidelines from the NIH’s Animal Research Advisory Committee provide recommendations for the ethical and scientific conduct of animal research, including pain management and experimental design. The document reflects the ongoing effort to refine animal use (the “R” in 3Rs), while the broader scientific community increasingly advocates for complete replacement (the first “R”) via advanced in vitro technologies. ↩︎

  124. Princeton University, “Regulations and Resources for Animal Care and Use”, n.d.
    This institutional document outlines Princeton University’s policies and resources for the ethical care and use of animals in research, in compliance with federal regulations. It emphasizes the institution’s commitment to the 3Rs (Replacement, Reduction, Refinement), reflecting a broader academic trend toward minimizing animal use and exploring alternative, human-relevant research models. ↩︎ ↩︎

  125. Azilagbetor, D., Shaw, D., Elger, B., “Animal Research Regulation: Improving Decision-Making and Adopting a Transparent System to Address Concerns around Approval Rate of Experiments”, Animals (Basel), 2024.
    This paper examines the regulatory frameworks governing animal research approval, arguing for a more transparent and ethically robust decision-making system. The authors suggest that improving transparency and rigor in animal study approval could naturally accelerate the adoption of New Approach Methodologies by highlighting the limitations and ethical costs of traditional animal experiments. ↩︎ ↩︎

  126. European Commission, “Animals in Science: EU Regulations”, n.d.
    This overview details the European Union’s stringent regulations on the use of animals in scientific research, which are founded on the principle of the 3Rs (Replacement, Reduction, Refinement). The EU is a global leader in actively funding and promoting the development of New Approach Methodologies (NAMs) to ultimately replace animal testing, setting a regulatory precedent for other regions. ↩︎ ↩︎ ↩︎

  127. APA, “Guidelines for Ethical Conduct in the Care and Use of Animals”, n.d.
    These guidelines from the American Psychological Association outline the ethical principles for the care and use of animals in psychological and behavioral research. While emphasizing humane treatment and scientific justification, the guidelines also reflect the growing ethical scrutiny of animal research, encouraging researchers to consider and utilize non-animal alternatives whenever scientifically feasible. ↩︎