The Replication Crisis in Psychology

The Replication Crisis in Psychology: Foundations, Failures, and Reform

Abstract

Over the last decade, psychology has undergone an extensive methodological reckoning commonly referred to as the replication crisis. Large-scale replication efforts, preregistered studies, and meta-scientific audits have demonstrated that many highly cited and routinely taught findings fail to reproduce under rigorous conditions, or reproduce only with substantially smaller effect sizes than originally reported. This article provides an overview of the replication crisis, explains the structural causes of the crisis, and catalogues a set of foundational psychological findings that were historically central to education and theory but whose strong versions are now empirically unsupported. Emphasis is placed on methodological precision, conceptual clarity, and the distinction between genuine falsification and evidential attenuation.


Replication and Its Role in Psychological Science

Replication is the process by which an independent research group attempts to reproduce a previously reported finding. In psychology, replications are typically classified as:

  • Direct replications, which aim to reproduce the original methods as closely as possible.
  • Conceptual replications, which test the same theoretical claim using different operationalisations.

Historically, psychology relied heavily on conceptual replication and narrative coherence rather than systematic direct replication. This practice allowed findings to become theoretically entrenched and pedagogically canonical before their empirical robustness had been adequately tested.

What Is Meant by the “Replication Crisis”

The replication crisis does not imply that psychology is fundamentally unscientific, nor that most findings are false. Rather, it refers to a documented pattern in which:

  • Reported effects frequently fail to reappear in direct replications.
  • Replicated effects are substantially smaller than originally published estimates.
  • Statistical significance is achieved through analytic flexibility rather than a stable signal.

A pivotal moment occurred when the Open Science Collaboration attempted to replicate 100 findings from leading psychology journals. Fewer than half of the effects replicated using conventional criteria, and replicated effect sizes were, on average, approximately half the magnitude of the original reports. This result was not anomalous but instead indicative of systemic bias.

Structural Causes of the Crisis

Low Statistical Power

Historically, psychological studies relied on small samples, leading to unstable estimates and exaggerated effect sizes. Low power simultaneously increases false negatives and inflates the magnitude of false positives that survive publication. IEMT practitioners should note that this vulnerability is endemic amongst the alphabet therapies, with practitioners regularly overestimating their own personal therapeutic effectiveness.

Publication Bias

Journals preferentially publish statistically significant results, while null findings remain unpublished. This “file drawer problem” distorts the literature, making reported effects appear stronger and more consistent than they truly are. Basically, no one likes to publish something that turned out to be a waste of time, creating a strong positivity bias. This is something all IEMT practitioners ought to pay close attention to, as a Pollyanna mindset is common when sharing positive reports on social media with overly generalised conclusions. It is this that leads to the ridiculous paradigm whereby many alphabet therapists expound how all illnesses are largely optional, as one only needs to have a more positive mindset, and health shall surely follow.

Researcher Degrees of Freedom

Flexible decisions regarding data exclusion, outcome selection, covariate inclusion, and stopping rules allow researchers (often unintentionally) to capitalise on chance. When these practices are undisclosed, nominal p-values no longer correspond to their stated error rates.

HARKing (Hypothesising After Results Are Known)

Post hoc hypotheses presented as a priori predictions transform exploratory findings into seemingly confirmatory evidence, misleading readers about the strength of theoretical support.

“Disproven” Versus “Unsupported”: A Necessary Distinction

Few psychological claims are logically falsified in a strict sense (i.e. deliberate fraud). More commonly, replication reveals that:

  • The original effect was dramatically overstated.
  • The effect is highly context-dependent rather than general.
  • The original paradigm does not reliably produce the claimed outcome.

Foundational Psychology Findings Now Strongly Challenged

Ego Depletion (Self-Control as a Finite Resource)

For decades, ego depletion was taught as a central mechanism of self-regulation: exerting willpower in one task was said to reliably impair subsequent self-control. Large preregistered, multi-laboratory replications have failed to support this general effect, reporting effect sizes indistinguishable from zero under standardised protocols.

Current consensus holds that the strong resource-depletion model is unsupported. If self-control fatigue exists, it is smaller, less general, and more contextually constrained than originally proposed.

Facial Feedback via the Pen-in-Mouth Paradigm

The claim that facial muscle activation directly alters emotional experience became a staple classroom demonstration.

However, a large Registered Replication Report failed to reproduce the original effect. While facial feedback as a broad hypothesis remains debated, the canonical pen-in-mouth demonstration is no longer considered reliable evidence.

A young woman with shoulder length hair and a T shirt holds a pen horizontally in her mouth, smiling playfully. This...

Power Posing

A confident woman stands with her hands on her hips, looking slightly upward. She is wearing a blazer and trousers, and...

Expansive body postures were claimed to produce hormonal changes (increased testosterone, reduced cortisol) and behavioural risk-taking. Such nonsense became a staple for NLP trainers and TED talkers.

High-powered replications failed to detect these effects, and one of the original co-authors publicly withdrew support for the original claims. The strong biological and behavioural interpretation is therefore unsupported.

Social and Behavioural Priming (Elderly Prime Creates Slower Walking)

Perhaps the most influential example of social priming involved priming participants with elderly-related words and observing slower walking speed. Direct replication attempts failed to reproduce the effect, and subsequent scrutiny revealed broader fragility in behavioural priming claims. Contemporary views treat such effects as, at best, highly sensitive to demand characteristics and contextual cues.

Precognition Effects in Experimental Psychology

Claims that individuals could respond to future stimuli were published using standard statistical methods, exposing the vulnerability of those methods to false positives. Multiple preregistered replications failed to support the effects, and the case is now widely used as a methodological caution rather than substantive evidence, but this doesn't stop podcasters from using Daryl Bem's flawed study as "proof" that free will doesn't exist.

The Marshmallow Test as a Predictor of Life Success

Originally interpreted as evidence that early self-control robustly predicts later achievement, later analyses demonstrated that the predictive power of delay-of-gratification tasks is substantially reduced once socioeconomic and family factors are controlled. The popular causal narrative of stable “willpower” has therefore been strongly qualified.

The Stanford Prison Experiment

Long presented as decisive evidence for the power of situational forces, extensive reanalysis and archival work have revealed substantial experimenter involvement, demand characteristics, and methodological weaknesses. The study’s didactic value is now primarily historical rather than evidential.

The Mozart Effect

The claim that listening to Mozart increases intelligence has not survived replication. Meta-analytic work indicates only small, transient, task-specific performance changes consistent with arousal or mood rather than cognitive enhancement.


Methodological Reform and the Future of Psychological Science

In response to the replication crisis, psychology has implemented substantial reforms:

  • Preregistration of hypotheses and analysis plans.
  • Registered Reports, in which studies are accepted prior to knowing the results.
  • Open data, materials, and analytic code.
  • Large-scale, multi-laboratory collaborations.

These reforms aim to realign incentives toward accuracy, reduce false positives, and ensure that theoretical claims are proportionate to evidential strength.


Classic study / effectWhat was classically taughtWhat failed under replicationCurrent evidence-based position
Ego depletion (Baumeister et al.)Self-control relies on a finite internal resource that becomes depleted, impairing subsequent self-control across tasks.Large preregistered multi-lab replications found effects near zero using standard paradigms.The strong, general “resource depletion” model is unsupported; any effects (if present) appear small and task-dependent rather than reliably generalisable.
Power posing (Carney, Cuddy & Yap)Expansive posture increases testosterone, reduces cortisol, and increases risk-taking and confidence.An expansive posture increases testosterone, reduces cortisol, and increases risk-taking and confidence.Posture may influence subjective feelings in some contexts, but the biological and behavioural causal claims are unsupported.
Pen-in-mouth facial feedback (Strack et al.)Facial muscle activation directly alters emotional experience (e.g., “smiling” makes cartoons funnier).A Registered Replication Report failed to reproduce the canonical effect under preregistered conditions.The classic classroom paradigm is unreliable; broader facial feedback hypotheses remain debated but are not established by this demonstration.
Early delay ability reflects stable self-control that strongly predicts later achievement and well-being.Incidental semantic primes unconsciously alter overt behaviour in robust, predictable ways.Direct replications failed; effects appear highly sensitive to contextual cues, expectancy, and demand characteristics.Strong claims of unconscious behavioural control via priming are unsupported; where priming effects occur, they are typically small and context-dependent.
Unconscious goal priming (strong form)Subliminal or incidental cues can reliably activate complex goals and motivations that guide behaviour outside awareness.Preregistered studies often find weak, inconsistent, or null effects, with strong dependence on protocol details.Any effects appear fragile and constrained by boundary conditions; broad generalisations are not supported.
Stereotype threat (strong form)Making a negative stereotype salient reliably produces large performance decrements across testing contexts.Meta-analytic reassessments report smaller and more variable effects than early narratives implied, with concerns about publication bias.Effects may occur in specific conditions, but the claim of large, ubiquitous impairment is unsupported.
Marshmallow test (delay of gratification as “willpower”)Delay performance reflects multiple factors (environmental reliability, background variables, and early cognition) more than a simple trait-like “willpower” cause.Predictive associations substantially reduce or disappear after controlling for socioeconomic and family variables.Delay performance reflects multiple factors (environmental reliability, background variables, early cognition) more than a simple trait-like “willpower” cause.
Stanford Prison Experiment (SPE)Situational roles alone rapidly produce cruelty and abuse; individuals are secondary to the situation.Archival and methodological critiques indicate substantial experimenter influence, demand characteristics, and design weaknesses.Historically influential but not strong causal experimental evidence; best used as a case study in methodology and ethics rather than a definitive finding.
Mozart effect (intelligence enhancement)Listening to Mozart increases intelligence (often framed as a boost to IQ).Replications show only small, short-lived, task-specific changes consistent with arousal or mood.No evidence for intelligence enhancement; any performance effects are transient and non-specific.
Micro-expression lie detection (as a reliable skill)Brief facial cues (“micro-expressions”) can reliably reveal deception to trained observers.Controlled research typically finds accuracy near chance for deception detection using behavioural cues alone.No dependable behavioural lie-detection method exists; deception detection remains difficult and error-prone without corroborating evidence.

The replication crisis represents a maturation rather than a collapse of psychological science. Many once-celebrated findings have not withstood rigorous scrutiny, but this does not negate psychology’s scientific value. Instead, it demonstrates the necessity of methodological humility, transparency, and continual empirical reassessment. Modern psychology is increasingly characterised by slower, more cautious knowledge accumulation - an essential correction to earlier excesses.

5 1 vote
Article Rating
Subscribe
Notify of
guest
0 Comments
Oldest
Newest Most Voted
0
Would love your thoughts, please comment.x
()
x