← Field manual directory

Department of Reasonable Doubt

Document
FM–004
Revision
01
Updated
Status
Public preview

Interpreting scientific articles

A study is not a verdict.

“A study found” is the beginning of an inspection, not the end of a conversation. Please escort the headline to Methods. Its credentials will be checked there.

Purpose Assess whether a quantitative study's conclusions fit its design, results, and uncertainty. This is introductory guidance, not clinical or statistical consulting.

First: what has arrived on your desk?

Find the actual article, not just its press release. Check the title, authors, date, journal or repository, and version. Open linked corrections, expressions of concern, or retraction notices and read what they say. An identifier such as a DOI helps locate a work; it does not certify its quality.

This guide focuses on quantitative empirical articles, especially studies of interventions. The same questions about claims, methods, and limits travel well, but not every field uses trials or p-values. Qualitative research, mathematical proofs, and theoretical work need methods appropriate to their questions—not a compulsory laboratory coat.

Original research

Reports a study or analysis. Identify whether it is observational, experimental, qualitative, a model, or something else. “Original” describes its role in the evidence chain, not whether its conclusion is correct.

Review or meta-analysis

A review summarizes existing work. A systematic review uses explicit methods to find and assess studies; a meta-analysis statistically combines results. Check the search, exclusions, study quality, and whether combining those studies makes sense. Pooling does not disinfect bad inputs.

Preprint

A manuscript shared before journal peer review. Check for a later version and changes. It may be valuable evidence, but do not present it as having passed a review process it has not undergone.

Commentary or press release

Explains, argues, or promotes. It may help you find the research, but it is not the research itself. Follow the link, and check which findings belong to which document.

Document navigation / Recommended route

The abstract is the reception desk.

Read it for orientation. Do not let it conduct the entire investigation.

  1. Write down the question.

    Who or what was studied? What exposure or intervention? Compared with what? Which outcome, and over what period? Preserve those details when translating the result into ordinary language.

  2. Visit Methods.

    Find recruitment, group assignment, measurements, exclusions, and the analysis plan. These determine what the study can reasonably establish. Important details may be in supplements, a protocol, or a registration record.

  3. Read Results before accepting the explanation.

    Look for the primary outcome, group sizes, estimates, uncertainty, and missing observations. Read table footnotes and figure axes. Separate what was measured from the authors' explanation of why it happened.

  4. Return to Discussion with your notes.

    Do the conclusions keep the same population, outcome, and level of certainty? Are limitations acknowledged? Does an improvement in a laboratory marker become a claim about longer life? That promotion needs additional evidence.

The introduction supplies background, but its citations may be repeating someone else's assertions. Use the source-tracing checks rather than treating every sentence in a paper as a finding of that paper.

Methods / Where the claim earns its salary

Inspect how the answer was manufactured.

Design and comparison

Observational studies examine what happens without assigning the exposure; other differences between groups can help explain an association. Random assignment can reduce confounding, but poor allocation procedures, missing data, or biased measurement can still undermine a trial. Neither label ends the review.

Population and setting

Cells, mice, adult volunteers, and patients with a specific illness answer different questions. Random assignment is not random sampling from the public. Check who was excluded and whether the setting and follow-up resemble the claim you want to make.

Measurement

Was the outcome measured directly, self-reported, or inferred from a proxy? Did participants or assessors know the assigned group? Blinding may reduce some biases but is not always feasible. Ask how the study addressed the resulting risks.

Size and missing data

How many independent participants or units were analyzed—not merely how many measurements were taken? Repeated readings from one person are not extra people. Check withdrawals, exclusions, their reasons, and whether they differ between groups. A large sample does not fix systematic bias.

Plan versus discovery

Were the primary outcome and analysis specified before results were known? Compare the dated registration or protocol with the paper. Preregistration aids scrutiny; it is not a quality seal. Exploratory analyses can generate useful ideas when labeled as exploratory.

Many chances to find something

Testing many outcomes, subgroups, or analysis choices creates more opportunities for a striking result. Look for a justified plan and treatment of multiple comparisons. A promising subgroup is not automatically a confirmed exception; it may need new testing.

For intervention trials, these checks draw on Cochrane's guidance on risk of bias. This is an introductory inspection, not a substitute for a formal appraisal or specialist methods review.

Statistical translation / No ceremonial thresholds

Significant. In what sense?

Start with size, units, and the comparison.

“Better” could mean a tiny average difference or a change that matters in practice. Ask how much, compared with what, for how long, and at what cost or risk. Statistical significance does not measure practical importance.

Look for both relative and absolute effects, the baseline frequency, and the time period. A 50% reduction does not mean 50 fewer people out of every 100. Also check the measure: an odds ratio is not generally interchangeable with a risk ratio.

A confidence interval is not a warranty.

An estimate comes with uncertainty. A confidence interval describes statistical precision under the analysis's assumptions; a wide interval can leave substantially different effects compatible with the data. It does not automatically account for biased measurement, missing studies, or a poorly chosen model.

Technically, a 95% confidence procedure would produce intervals containing the true parameter in 95% of repeated applications under its assumptions. It does not mean 95% of participants improved, or assign a 95% probability to this particular fixed parameter falling inside the reported interval.

A p-value is not the probability the claim is false.

A p-value measures how unusual results at least as extreme as those observed would be under a specified statistical model, including a null hypothesis such as no difference. It is not the probability the study is wrong, that the hypothesis is true, or that the result was “caused by chance.”

A small p-value does not tell you an effect is large, important, unbiased, or likely to replicate. A large one does not prove no effect. Inspect the estimate, interval, design, and full reporting rather than allowing 0.05 to chair the meeting.

Sources: Cochrane, interpreting results, especially sections 15.3–15.4; American Statistical Association's six principles on p-values (2016, PDF).

Training simulation / No actual study

The headline has exceeded its authority.

Every study detail and numerical result below is invented for teaching. No real app, participants, or experiment are described. The interval is a stipulated example, not a calculation from an available dataset.

The promotional headline reads: “Bedtime app proven to give everyone 42 extra minutes of sleep.” Let us inspect the fictional paper it links to.

Question
Does the app improve self-reported sleep duration compared with a sleep-information leaflet after four weeks?
Design
80 adult volunteers, randomly assigned: 40 to the app, 40 to the leaflet. All complete follow-up. Participants know their assignment and report their sleep in diaries.
Primary outcome
Change in self-reported nightly sleep duration, specified before data collection.
App group
Mean increase: 0.7 hours (42 minutes).
Leaflet group
Mean increase: 0.2 hours (12 minutes).
Between-group estimate
0.5 hours (30 minutes) greater improvement with the app. Reported 95% confidence interval: −0.1 to +1.1 hours (−6 to +66 minutes).
Disclosure
The app maker funded the trial. The paper reports funding but does not establish whether the sponsor controlled analysis or publication.
Open the inspection report: four overclaims
  1. 42 minutes is the wrong comparison. It is the change within the app group. The leaflet group also improved. The estimated difference in changes is 42 − 12 = 30 minutes, not 42.
  2. “Proven” outruns the precision. The stated interval spans a small disadvantage, no difference, and a substantial benefit. It does not settle the size or direction of the effect. Equally, crossing zero does not prove the app has no effect.
  3. “Everyone” is not the study population. This is an average among adult volunteers over four weeks. It does not promise a benefit to every participant, children, or people using it for years.
  4. The measurement has limits. Participants knew their group and supplied diary estimates. Expectations or reporting differences could matter. Randomization helps the comparison; it does not make self-report an objective sleep measurement.

Funding deserves scrutiny of the sponsor's role, protocol, analysis, and publication rights. It is not a mathematical disproof of the result. Conversely, a disclosure is not a substitute for showing that the methods were protected from interference.

A defensible summary: “In this fictional four-week randomized study of adult volunteers, the app group reported an average 30-minute greater improvement than the leaflet group. The estimate was imprecise, with an interval from 6 minutes less to 66 minutes more. Unblinded self-report and the unresolved sponsor role warrant caution. The study does not establish a universal 42-minute benefit.”

Context review / Please retain previous findings

One paper does not get the whole building.

Ask how the result fits other relevant studies, careful systematic reviews, and current expert assessments. Have independent teams found similar effects with comparable methods? Are there plausible reasons for different results? Check the review's date and coverage rather than assuming the newest isolated paper automatically supersedes everything before it.

Reanalyzing the same data can check an analysis, but it is not replication with new observations. Several papers can also report on the same participants. Count independent evidence, not PDFs. Selective publication of striking findings can skew the accessible literature; a search returning only positive headlines is not a complete evidence review.

Look for protocols, accessible data or code where appropriate, and explanations of restrictions. Privacy or consent obligations can legitimately limit sharing. Missing open data is a limit on what you can verify, not automatic proof of fraud.

Know when to request a specialist.

If you cannot evaluate a key method, say so. Seek qualified independent interpretation rather than borrowing certainty from the abstract. A casual reader can spot mismatched claims without being able to audit every statistical model.

For personal healthcare decisions, do not start or stop treatment on the basis of this guide or a single paper; discuss applicable evidence, benefits, and harms with a qualified clinician. Other high-stakes decisions likewise deserve expertise appropriate to the field.

Sound research can contain uncertainty. Weak research can sound confident. Our corporate preference is for the kind that leaves its reasoning available for inspection.

Inspect this guide's sources.

These are methodological references, not evidence that the invented app works. Cochrane's guidance focuses on health interventions; other research questions require appropriate field-specific methods. The explanations, examples, and satirical commentary here are Suscorp's, not quotations from or endorsements by these organizations.

Published September 13, 2026. Educational public preview, not clinical or statistical consulting. Independent editorial review remains outstanding.

Revision is an expected outcome.

· Revision 01: introduced the standard manual format and monochrome interface. Existing methods guidance and fictional examples retained.

Revision numbering begins with this format edition; earlier unnumbered content was published September 13, 2026. Independent editorial review remains outstanding. Department names are fictional editorial identities, not claims of staffing or review.

Return to document header ↑ · Method and institutional disclosures