HACKS VITAE

MINDSET & DISCIPLINE · 14 POINTS LOOKED INTO

Cognitive Biases, Checked
14 Famous Findings and Whether They Held Up

Anchoring, confirmation bias, the Dunning-Kruger effect, ego depletion, power posing, the marshmallow test and more — what happened when scientists ran the famous studies again, with a verdict on each.

HOVER A GLOWING POINT · DRAG TO TURN
  • Partly true 7
  • True 2
  • Disputed 4
  • Unverified 1
Published
September 17, 2026
Updated
September 29, 2026
Read
4 min
points
14

BACKGROUND · GEORGES DE LA TOUR, THE PENITENT MAGDALEN, C. 1640 · THE MET, OPEN ACCESS

THE SHORT VERSION

  1. Fourteen famous psychology findings, checked against the large projects that tried to repeat them. Two held up cleanly, seven are real but smaller or narrower than the famous version, four are genuinely disputed, and one has not been confirmed.
  2. Anchoring and hindsight bias held up when they were tested again; framing held up at about half its original size.
  3. The strong version of ego depletion did not replicate, and the famous "elderly priming" walking study has not been confirmed.
  4. One failed replication is not the end of a finding, and one success is not the end of the argument. The verdicts below say which way the weight of evidence currently leans, and why.

What we found

Our reading of the evidence on each of the 14 points, with where it comes from. Open any row, look at the source, and make up your own mind.

TrueAnchoring: a number mentioned in a question sways the answer

Held up. Described by Tversky and Kahneman in 1974. When the Many Labs project tested anchoring questions — where a high or low number is stated in the question itself — across 36 samples, the authors wrote that the original studies had produced underestimates of some effects, anchoring among them. That result is about numbers given in the question; it does not show that any random number you happen to see affects you.

SOURCE Klein et al., Many Labs 1, Social Psychology (2014).

TrueHindsight bias: after the fact, outcomes seem more predictable than they were

Held up. A 2021 close replication of Fischhoff's 1975 study 'found support for hindsight bias in retrospective judgments'.

SOURCE Chen et al., Journal of Experimental Social Psychology 96 (2021).

Partly trueConfirmation bias: we judge evidence more kindly when it suits us

Real, but it is a family of effects, not one finding. One well-measured version: a meta-analysis of 51 experiments found people rated identical information more favourably when it supported their political side, a bias its authors called robust and found on both sides. There is no single canonical 'confirmation bias study' to replicate.

SOURCE Ditto et al., Perspectives on Psychological Science (2019).

Partly trueFraming: the same choice described as gains or losses changes decisions

Held up, at about half the original size. Tversky and Kahneman showed in 1981 that framing the same problem differently produces 'predictable shifts of preference'. In Many Labs 1 the gain-versus-loss framing effect replicated, but the effect came back at roughly half the size originally reported.

SOURCE Tversky & Kahneman, Science (1981); Klein et al. (2014).

Partly trueSunk cost: people keep investing because of what they have already spent

Replicated as a small effect in a hypothetical scenario. Many Labs 1 tested a written sunk-cost scenario (based on a 2009 study by Oppenheimer, Meyvis and Davidenko). Across all samples combined it replicated, but only about half of the 36 individual samples found a significant effect. It was a hypothetical choice, not a decision with real money at stake.

SOURCE Klein et al. (2014).

Partly trueThe marshmallow test: children who wait for a second treat do better in life

Smaller than claimed. A 2018 study, focusing on children whose mothers had not completed college, found the link with later achievement was 'only half the size of those reported in the original studies and was reduced by two thirds' once family background, early ability and the home environment were taken into account.

SOURCE Watts, Duncan & Quan, Psychological Science (2018).

Partly trueThe bystander effect: nobody helps when others are around

Real, but usually misread. A meta-analysis confirmed that each individual is less likely to help when others are present, but found the effect weaker in dangerous emergencies. And when researchers studied CCTV footage of real public conflicts, 'in 9 of 10 public conflicts, at least 1 bystander' did something to help.

SOURCE Fischer et al., Psychological Bulletin (2011); Philpot et al., American Psychologist (2020).

Partly truePower posing: standing like a superhero changes your hormones and behaviour

The hormone and behaviour claims failed; the feeling replicated in that study. A 2015 replication with 200 people 'found no significant effect of power posing on hormonal levels or in any of the three behavioral tasks', but did find people felt more powerful. The original study's first author, Dana Carney, has since written: 'I do not believe that "power pose" effects are real.'

SOURCE Ranehill et al., Psychological Science (2015); Carney, 'My position on Power Poses'.

Partly trueFacial feedback: holding a pen in your teeth to force a smile makes cartoons funnier

The pen study did not replicate; a small effect of facial feedback seems real. A 17-lab replication of the pen study found a rating difference of just 0.03 units. A meta-analysis found facial feedback's overall effect 'significant but small', and a later multi-country study found the evidence 'less conclusive' for the pen-in-mouth method specifically.

SOURCE Wagenmakers et al. (2016); Coles, Larsen & Lench (2019); Coles et al. (2022).

DisputedLoss aversion: losses hurt more than equal gains

Genuinely contested. A 2018 review argued that current evidence 'does not support that losses, on balance, tend to be any more impactful than gains'. Studies with 17,720 people in total pushed back in 2020, finding people 'of all knowledge and experience levels were loss averse'. A 2024 meta-analysis of 607 estimates put the average loss aversion coefficient at 1.955 — losses weighing roughly twice as much as gains. A second 2024 meta-analysis, limited to studies that fitted prospect theory to choices between risky gambles, found that 'much of the data are of poor quality' and put the figure at 1.31 (95% CI 1.10 to 1.53). We lean towards loss aversion being real on average, because both meta-analyses put it above 1; its size is still argued over, with the two estimates at about 1.3 and about 2.

SOURCE Gal & Rucker (2018); Mrkva et al. (2020); Brown, Imai, Vieider & Camerer, Journal of Economic Literature (2024); Walasek, Mullett & Stewart, Journal of Economic Psychology (2024).

DisputedThe Dunning-Kruger effect: the least skilled are the most overconfident

The overestimation is real; its size and cause are disputed. In the 1999 study, the lowest scorers on average put themselves around the 62nd percentile when they were in the 12th. Critics argue much of the famous pattern is a statistical artefact, and one 2020 paper concluded the effect 'may be much smaller than reported previously'. A large 2021 study still found low performers were less able to judge whether their answers were right, in grammar and logical reasoning. Our reading: low performers overestimating themselves is well supported; how much of the classic pattern is real rather than statistical is an open question.

SOURCE Kruger & Dunning (1999); Gignac & Zajenkowski, Intelligence (2020); Jansen, Rafferty & Griffiths, Nature Human Behaviour (2021).

DisputedStereotype threat: reminding girls of a stereotype lowers their maths scores

Probably smaller than claimed. A 2015 meta-analysis of studies with schoolgirls found an average effect, but warned that 'publication bias might seriously distort the literature'. A pre-registered study of 2,064 Dutch students aged 13 to 14 then 'found neither an overall effect of stereotype threat on math performance' among girls, nor any moderated effect, and a 2019 meta-analysis of settings closer to real exams, covering women and ethnic minorities, concluded the effect 'may range from negligible to small'. We lean towards a small or negligible effect, because the pre-registered study and the publication-bias warning both point that way. The pre-registered study was about girls and maths; the 2015 meta-analysis covered girls on maths, science and spatial tests; the 2019 meta-analysis covered women and ethnic minorities on cognitive ability tests.

SOURCE Flore & Wicherts, Journal of School Psychology (2015); Flore, Mulder & Wicherts (2018); Shewach, Sackett & Quint, Journal of Applied Psychology (2019).

DisputedEgo depletion: willpower is a resource that runs out when you use it

The strong version failed; a very small effect is disputed. A 23-lab registered replication found an effect whose confidence interval 'encompassed zero'. A 36-site study led by Vohs and colleagues reported that its 'preregistered analyses did not find evidence for a depletion effect', though an exploratory analysis of its full sample found a small significant effect. A preregistered 12-lab study, meanwhile, found 'a small and significant ego depletion effect'. Our reading: the popular idea that willpower drains like a fuel tank is not supported; whether a tiny effect exists is still open.

SOURCE Hagger et al. (2016); Vohs et al., Psychological Science (2021); Dang et al., Social Psychological and Personality Science (2020).

UnverifiedElderly priming: reading words about old age makes people walk more slowly

The original finding has not been confirmed. A 2012 replication using automated timing 'failed to show priming'. In a second experiment the slower walking appeared 'only when experimenters believed participants would indeed walk slower' — pointing to the experimenters' expectations rather than the words. We found no large, pre-registered replication either way.

SOURCE Doyen, Klein, Pichon & Cleeremans, PLoS ONE (2012).

THE ARTICLE · 4 MIN

Cognitive biases are often presented as settled science. They are not all equally settled.

Psychologists have re-run many of their field’s most famous experiments — often with larger samples, in many labs at once, and sometimes with the analysis plan registered in advance. At least one came back stronger than first reported. Several came back smaller. A few have not come back convincingly at all.

This page checks fourteen of them. The cheat sheet on this page (“What we found”) has a verdict and the key study for each; the sections below explain how to read those verdicts.

Why this matters: the replication projects

In 2015, the Open Science Collaboration published an attempt to repeat 100 studies from three psychology journals. By one of its measures, “Thirty-six percent of replications had statistically significant results”, and the replication effects “were half the magnitude of original effects, representing a substantial decline.”

It is tempting to read that as “only 36% of psychology replicates”. That goes too far. The project used several measures of success, and 36% is only one of them. A fairer summary is narrower: many published effects were weaker than first reported, and some did not reappear.

The Many Labs projects took a different approach — testing a set of well-known effects across dozens of labs at once. In the second of them, covering 28 findings, the authors reported that 15, or 54%, “provided evidence of a statistically significant effect in the same direction as the original finding.”

How to read the verdicts

  • True — the finding replicated in large, independent tests.
  • Partly true — something real was found, but it is smaller, narrower, or different from the popular version.
  • Disputed — credible studies point in different directions. We say which way the evidence currently leans and why.
  • Unverified — the original finding has not been confirmed by a convincing replication.

Two cautions apply to every one of them. A single failed replication is not a disproof: it may have tested the idea differently, or been unlucky. And a single success is not the end of the argument either. What matters is the weight of evidence across several good tests — which is also why these verdicts could change.

What the fourteen show — our reading

The biases about judgement mostly held up, sometimes smaller. Anchoring and hindsight bias replicated clearly; framing replicated at about half size; the sunk cost effect replicated as a small effect in a hypothetical scenario.

The findings about subtle manipulations mostly shrank or failed. Ego depletion, elderly priming, power posing’s hormone effects and the pen-in-mouth smile all claimed that a small manipulation changes how people feel, perform or act. Those are the claims that struggled most when tested again.

Several famous findings are real but routinely misdescribed. The bystander effect is real for each individual, yet in real conflicts caught on camera someone usually helps. The marshmallow test found something, but much less than the popular story. The Dunning–Kruger effect has a real core, but how much of its classic pattern is real is still argued over.

Using this without over-correcting

It is tempting to swing from “psychology proves everything” to “psychology proves nothing”. Neither is right. The replication projects are themselves psychology working as it should — checking its own claims and changing its mind.

The practical rule for any study you hear about: ask whether it has been independently repeated, and how big the effect was the second time.

For the method of checking a study yourself, see how to check something online and our guide to numbers that mislead. The Dunning–Kruger chart is covered in Older Than You Think, Part 2.

Sources

Checked September 2026.

Related: Logical fallacies cheat sheet · Numbers that mislead · Overcoming survivorship bias

  • critical thinking
  • psychology
  • cognitive biases
  • replication
  • cheat sheet

SHARE & CITE

Hacks Vitae. "Cognitive Biases, Checked: 14 Famous Findings and Whether They Held Up." September 17, 2026. https://www.hacksvitae.com/life-hack/cognitive-biases-checked-14-famous-findings-and-whether-they-held-up

That's what we found. The rest is your call.

118 articles, each with its sources listed. Spotted something off? hacksvitae@gmail.com

Open the library