THE ARTICLE · 9 MIN
You do not need to be a scientist to read a study sensibly. You need to know which parts to look at, and which common assumptions to drop.
These nine questions work for most research you will meet in the news or on social media. Where we could, we used sources from outside medicine — this is a guide to reading research, not advice about any health decision.
1. Did you read past the abstract?
A scientific paper usually has the same parts. The University of Michigan Library’s guide describes them plainly: abstracts “are always found at the beginning of an article and provide a basic summary or roadmap to the article.”
The part that matters most for a careful reader is the methods section. As the same guide puts it: “It is necessary for other researchers to understand the methods used so that they can replicate the study.” Our advice follows from that: the methods are where you find out who was studied, how many, and what was actually done.
Then keep the results and the discussion apart. The guide notes that “The discussion section is where you will find the researcher’s interpretation of the results.” The results are what they found; the discussion is what they think it means.
2. What kind of study is it?
The biggest single question is whether the researchers assigned people to conditions, or simply observed them.
In an experiment with random assignment, chance decides who gets which condition. The open statistics textbook Introductory Statistics from OpenStax explains why that matters: “When subjects are assigned treatments randomly, all of the potential lurking variables are spread equally among the groups.” (That is a textbook simplification: in practice random assignment makes the groups similar on average rather than identical.) An observational study cannot do that, so a link it finds may be driven by something else — the trap covered in our guide to numbers that mislead.
Two other labels you will meet:
- A systematic review, in a definition quoted by KTDRR from John Last’s Dictionary of Epidemiology, is “the application of procedures that limit bias in the assembly, critical appraisal, and synthesis of all relevant studies on a particular topic.”
- A meta-analysis is “the statistical synthesis of the data from separate but comparable studies, leading to a quantitative summary of the pooled results.”
A meta-analysis combines results. It cannot make weak studies strong — which is our reading, and the reason the next question still applies.
3. Was it peer reviewed — and what does that mean?
The organisation Sense about Science describes peer review as a process that “subjects scientific research papers to independent scrutiny by other qualified scientific experts (peers) before they are made public.”
Two limits are worth knowing:
- It is not a fraud check. The same organisation’s guide says: “Peer review is not a fraud detection system.” It adds that referees “are likely to detect some wrongdoing, such as copying someone else’s research or misrepresenting data”, but deliberate fraud can pass.
- It is a start, not a finish. In its words: “Publication of a peer-reviewed paper is just the first step”.
Some studies appear first as preprints, before any peer review. arXiv, a preprint server covering physics, mathematics, computer science and several other fields, moderates submissions, but is explicit: “Please note that the arXiv moderation process is not a peer-review process.” A preprint can be excellent; it simply has not been through that check yet.
4. How big was the sample?
Small studies do more than miss real effects. A 2013 review in Nature Reviews Neuroscience by Button and colleagues warned that “low power also reduces the likelihood that a statistically significant result reflects a true effect”. In other words, a dramatic finding from a small study deserves more caution, not less.
5. What does the p-value tell you — and what doesn’t it?
In 2016 the American Statistical Association issued a statement on p-values, setting out six principles. They are short enough to quote in full:
- “P-values can indicate how incompatible the data are with a specified statistical model.”
- “P-values do not measure the probability that the studied hypothesis is true, or the probability that the data were produced by random chance alone.”
- “Scientific conclusions and business or policy decisions should not be based only on whether a p-value passes a specific threshold.”
- “Proper inference requires full reporting and transparency.”
- “A p-value, or statistical significance, does not measure the size of an effect or the importance of a result.”
- “By itself, a p-value does not provide a good measure of evidence regarding a model or hypothesis.”
The association’s executive director, Ron Wasserstein, summed up the spirit of it: “The p-value was never intended to be a substitute for scientific reasoning”.
The practical upshot: “statistically significant” does not mean “true”, and it does not mean “big”.
6. How big is the effect?
Because significance says nothing about size (principle 5 above), look for the effect size — how large the difference or relationship actually was — and how precisely it was estimated.
Psychology’s reporting standards, published by the American Psychological Association, ask authors to report “Effect-size estimates and confidence intervals on those estimates that correspond to each inferential test conducted, when possible”. If a study reports a significant result but no effect size, you cannot tell whether the finding matters in practice.
7. Has anyone repeated it?
In 2015 the Open Science Collaboration reported repeating 100 studies from three psychology journals. “Ninety-seven percent of original studies had statistically significant results. Thirty-six percent of replications had statistically significant results”. The replication effects “were half the magnitude of original effects, representing a substantial decline.” That 36% is one of several measures the project used, and what the results mean is itself DISPUTED. In 2016 Gilbert and colleagues argued that the data “are consistent with the opposite conclusion, namely, that the reproducibility of psychological science is quite high.” The original team replied that “both optimistic and pessimistic conclusions about reproducibility are possible, and neither are yet warranted.”
The other side matters too. The US National Academies of Sciences, Engineering, and Medicine caution that “nor does a single failed replication conclusively refute the original claims” — and, in the same sentence, that a successful replication does not guarantee the original was right. Our reading: what builds confidence is several good studies pointing the same way. Our cognitive biases check shows what that looks like for fourteen famous findings.
Two newer practices are designed to reduce these problems, and are worth looking for:
- Pre-registration. The Center for Open Science describes it as “specifying your research plan in advance of your study and submitting it to a registry”. The reason: “the same data cannot be used to generate and test a hypothesis”.
- Registered Reports. A journal reviews the study plan before any data exist and, if the plan is approved, “the journal virtually guarantees publication if the authors conduct the experiment in accordance with their approved protocol”. According to the Center for Open Science, “over 300 journals use the Registered Reports publishing format either as a regular submission option or as part of a single special issue” (as read in September 2026).
8. Who paid for it, and who wrote the headline?
Funding and conflicts of interest. In the wording of the International Committee of Medical Journal Editors, “The potential for conflict of interest and bias exists when professional judgment concerning a primary interest (such as patients’ welfare or the validity of research) may be influenced by a secondary interest (such as financial gain).” Its recommendations ask authors to state “Sources of support for the work, including sponsor names along with explanations of the role of those sources”. Look for a funding or disclosure statement near the end of the paper. Our advice: funding does not make a study wrong; it is a reason to read the methods more closely.
Press releases. A 2014 study in The BMJ by Sumner and colleagues examined press releases about health and biomedical research from UK universities and found that “Exaggeration in news is strongly associated with exaggeration in press releases.” It also found “there was little evidence that exaggeration in press releases increased the uptake of news”. That study looked only at health research. Our advice, in any field: go back to the paper, not the press release.
9. Has it been retracted?
Papers are sometimes retracted after publication.
One public record of retractions is the Retraction Watch database. In September 2023 Crossref announced that “the Retraction Watch database has been acquired by Crossref and made a public resource”. The database’s user guide notes that “The more fields you fill in, the more specific the search” — searching by the paper’s DOI or title is the quickest check.
A note on the “evidence pyramid”
You may see research drawn as a pyramid. Murad and colleagues describe the traditional version, with weaker designs such as case series near the bottom and systematic reviews at the top. It is a useful rough guide, but it has critics. A 2016 paper by Murad and colleagues argued that “Study design alone appears to be insufficient on its own as a surrogate for risk of bias”, and proposed to “remove systematic reviews from the top of the pyramid and use them as a lens through which other types of studies should be seen”. Our reading: a well-run observational study can be more informative than a poorly run experiment.
The nine questions, on one screen
- Did I read the methods, not just the abstract?
- Experiment or observation?
- Peer reviewed, preprint, or press release only?
- How many people or cases?
- Am I reading more into the p-value than it says?
- How big was the effect?
- Has it been repeated — and was it pre-registered?
- Who paid, and does the headline match the paper?
- Has it been retracted?
Sources
- University of Michigan Library, “Reading Scholarly Articles”.
- OpenStax, Introductory Statistics 2e, section 1.4, “Experimental Design and Ethics”.
- KTDRR / NCDDR, FOCUS Technical Brief 16, “The Campbell Collaboration” (definitions quoted from Last, 2001, via Chalmers, Hedges and Cooper, 2002).
- Sense about Science, I don’t know what to believe (first edition 2005, updated 2017).
- arXiv, “arXiv moderation” help page.
- Button, Ioannidis, Mokrysz, Nosek, Flint, Robinson and Munafò, “Power failure: why small sample size undermines the reliability of neuroscience”, Nature Reviews Neuroscience 14 (2013).
- American Statistical Association, statement on statistical significance and p-values (press release, 7 March 2016).
- Appelbaum and colleagues, “Journal Article Reporting Standards for Quantitative Research in Psychology”, American Psychologist 73 (2018).
- Open Science Collaboration, “Estimating the reproducibility of psychological science”, Science 349 (2015); Gilbert, King, Pettigrew and Wilson, comment, Science (2016); Anderson and colleagues, response to comment, Science (2016).
- National Academies of Sciences, Engineering, and Medicine, Reproducibility and Replicability in Science (2019).
- Center for Open Science, “Preregistration” and “Registered Reports”.
- International Committee of Medical Journal Editors, Recommendations: author responsibilities and conflicts of interest.
- Sumner and colleagues, “The association between exaggeration in health related science news and academic press releases: retrospective observational study”, The BMJ 349 (2014).
- Crossref blog post on acquiring the Retraction Watch database (12 September 2023); Retraction Watch Database User Guide.
- Murad, Asi, Alsawas and Alahdab, “New evidence pyramid”, Evidence-Based Medicine 21 (2016).
Checked September 2026.
Related: How to check something online: the SIFT method · Cognitive biases, checked · Numbers that mislead
- critical thinking
- science literacy
- research
- statistics
- replication
