Skip to content
Briefs · researchUpdated

Hypothesis Stress-Test

Attack a hypothesis to find where it would break before you test it.

You walk away with

A list of the assumptions, confounds and the experiment that would falsify it.

Decidi convenes

Hypotheses die expensively when they are tested before they are attacked, so the attack comes first: Stephen Hawking’s seat asks whether the idea is precise enough to be wrong at all, the Statistician designs the confounds and power problems out before they design themselves in, and the Deep Researcher checks whether the answer already exists in prior work. The Data Skeptic hunts the alternative explanations for the result you are hoping for, the Devil’s Advocate tries to kill the idea outright, and the Steelman rebuilds the strongest surviving version — because the goal is a better hypothesis, not a dead one.

Recommended level: DeepThe newest, most capable models — for when being wrong is expensive.
What the council debates
Stress-test this hypothesis before I invest in testing it. Try to break it.

THE HYPOTHESIS:
[state it precisely — what you predict and why]
THE REASONING: [the logic and any prior evidence behind it]
HOW YOU PLAN TO TEST IT: [the experiment or study design, if any]

Debate:
1. Is the hypothesis precise and falsifiable, or too vague to test.
2. The assumptions it rests on, and which are shakiest.
3. The confounds and alternative explanations for any expected result.
4. What evidence would actually falsify it (not just confirm it).
5. The weakest link in the proposed test design.
6. The cheapest experiment that would most quickly support or kill it.

FINAL SYNTHESIS:
- The assumptions the hypothesis depends on, ranked by fragility.
- The most likely alternative explanation to rule out.
- The single experiment that would most efficiently test it, and the result that would falsify it.
What the debate sounds like

An illustrative excerpt for a hypothetical brief of this type: A behavioural-science graduate student hypothesises that daily gratitude journalling improves sleep quality, and plans a four-week study with volunteer undergraduates self-reporting both. Every real run is generated fresh from your specifics, debated across rounds, and audited before the verdict.

🌌
Stephen Hawking

First make it falsifiable in its current words: “improves sleep quality” — by what measure, how much, for whom, against what comparison? As stated, any positive drift confirms it and nothing can refute it, which makes it a hope, not a hypothesis. Commit now: a defined sleep measure, a minimum effect worth caring about, a comparison arm. A hypothesis that cannot specify the result that would kill it is not ready for data.

🔬
The Data Skeptic

The design as described manufactures its own result: volunteers who sign up for a gratitude study expect gratitude to work, and both variables are self-reported by the same person who knows which arm they are in. Expectancy plus demand characteristics will produce a positive finding with no real effect. The fix is structural — an active control doing neutral journalling, so both arms carry equal expectation and equal pen-and-paper time.

📈
The Statistician

Power arithmetic before a single participant: effects in this literature are small-to-moderate, and small effects with noisy self-report measures at a plausible volunteer sample size mean the study is more likely to produce a false negative or an inflated fluke than a trustworthy answer. Options: a bigger sample, a within-subject crossover design, or an objective sleep measure with less noise. Choose one deliberately — or the four weeks answer nothing either way.

🛡️
The Steelman

The surviving version, rebuilt: “Compared with neutral journalling, four weeks of gratitude journalling improves [defined sleep-quality score] by at least [named margin] in [defined population].” Testable, honest about the comparison, and pre-committed on the effect size. And the cheapest next step is not the full study — it is a two-week pilot to check whether participants actually do the journalling, because adherence is the assumption everything upstream silently rests on.

Run this brief as a council

Prefer drop-and-go? Use the Hypothesis Stress-Test tool — team pre-seated, included with Plus.

Questions people ask

Is this only for academic research, or for product and business hypotheses too?

The machinery is identical for “feature X will reduce churn” and for lab science: precision, falsifiability, confounds, the cheapest decisive test. Product hypotheses actually benefit more, because business experiments are routinely run with the design flaws this council is built to catch — self-selected samples, no control, metrics chosen after the fact.

What does “what evidence would falsify it” give me in practice?

Pre-committing to the disconfirming result is the single strongest protection against fooling yourself: after data arrives, every ambiguous outcome gets reinterpreted as support. The deliverable states, before you run anything, exactly which result kills the hypothesis — so the future argument with yourself is already settled.

My hypothesis survived — what does the output actually contain?

The assumptions ranked by fragility, the most likely alternative explanation you must rule out, the strongest rebuilt version of the hypothesis, and the single most efficient experiment with its falsifying result named. Surviving the council does not mean the idea is right — it means the test you run next will be worth running.