Skip to content
For professionals · Data ScientistsUpdated
📐

Decidi for data scientists

Find the leak in the analysis before it ships a bad decision.

Stress-test your work Chat free · no sign-up, no card

Have the methodology, the data and the conclusion challenged by independent models — so the result you present is one a skeptic couldn’t unwind.

Why data scientists use it
  • Hunt for data leakage, the confound and the spurious correlation a confident model would happily report.
  • Pressure-test the conclusion: does the analysis actually support it, or is it one inference too far?
  • Surface the bias in the data, the sampling problem and the population you didn’t actually measure.
  • Stress-test the model for overfitting and for whether the metric reflects the real objective.
  • Have a skeptic challenge the chart — what does it hide, what scale flatters it?
  • Cross-check the statistics so one model’s confident p-value doesn’t go unquestioned.
Stress-test before you ship
  • Data leakage, confounds and spurious correlation
  • Whether the conclusion follows from the analysis
  • Sampling, bias and the population you measured
  • Overfitting and metric-vs-objective alignment
  • The visualisation — what it hides or flatters
  • Statistical claims and significance
Adversarial passes we run
Data-leakage sweepConclusion-support testBias & sampling auditOverfitting checkStatistics cross-check
A worked example — the model result that looks too good

Say your churn model just hit an accuracy number that beats the previous baseline by a suspicious margin, the business team wants it in production this month, and a small voice is asking where the leakage is. The voice is usually right.

The Data Skeptic hunts the leakage first — the feature computed after the outcome, the train-test split that leaks time, the proxy that encodes the label — because a result that good is a bug until proven otherwise. The Methodologist audits the evaluation itself: is the metric the business one, is the baseline honest, does performance hold on the customers who matter rather than on average? The Statistician checks whether the margin over baseline survives its confidence interval, and the End-User Advocate asks what the retention team will actually do with a churn score — because a model nobody acts on has an accuracy of zero in practice. The Devil’s Advocate argues for shipping the simpler model. The verdict: the specific checks before production, ranked.

Stress-test your work now

Chat free · no sign-up, no card