MARKET WATCH
Loading live market prices…

Evidence Lab / Simulated evidence

Is it overfitted? We’ll tell you.

A strong backtest is a reason to ask harder questions. Evidence Lab challenges the result from different angles, then shows what held up, what weakened and what the data could not answer.

Historical, simulated tests. A grade is not a forecast or a decision to trade.

08 checksOne reasoned reportMethod 2026.10.02
EVIDENCE LAB · Edge Theory

The report starts with a verdict

A grade with reasons. Never a green light.

The letter summarises the evidence available for that saved version. Read the sentence, the leading reasons and the points lost before drawing a conclusion.

Not enough dataFewer than 30 completed trades? The report withholds the grade instead of dressing uncertainty up as a score.

ROBUSTNESS GRADESCORE RANGE · 0–100
A85+
B70–84
C55–69
D40–54
F0–39
Each lost point has a reason and a suggested next step.Grade bands · methodology 2026.10.02

The examination

Eight checks. Different ways a backtest can fail.

Each check has a specific question. A missing sample, unsupported comparison or unavailable market stays visible as unavailable.

01 — 04 Selection and repeatability

01

Luck test

When the strategy version and sample support it, compare the saved result with randomized entry timing on held-out trades.

02

Unseen data

Keep the final 30% of the history aside until you choose Reveal. The reveal is recorded for that strategy version and cannot be undone.

03

Walk-forward

Where the data supports it, repeat the tune-and-test process across rolling windows and compare how the rules behaved on the next window.

04

Overfitting meter

Record the strategy trials and adjust for repeated testing. When enough aligned trials exist, the report also checks how often in-sample winners weaken out of sample.

05 — 08 Stress and context

05

Bad-luck stress

Resample recorded trade outcomes to examine possible drawdowns and losing streaks. The report uses the stress depth available to your plan.

06

Fragility

Vary nearby settings and increase execution costs to see whether the result depends on one narrow configuration.

07

Other markets

Where plan and data allow, test unchanged rules on supported exchanges and related markets. Missing coverage is shown as unavailable, not filled with an estimate.

08

When it works

Break down recorded outcomes by session, weekday, available market regime and volatility to show where the historical result came from.

Evidence depth varies by membership. This report records the run counts, data range, costs and any checks that could not be completed.

The luck test

Could random timing have looked just as good?

Evidence Lab compares the observed strategy with randomized entry timing. Run counts follow the member’s plan: 200–1,000 comparisons across current tiers. The report records the actual count used for that run.

How the comparison is scored
HOW TO READ THE COMPARISONILLUSTRATION · NO RESULT SHOWN
Method diagram only · no strategy outcome or member data shown

A one-way decision

The unseen period stays unseen until you choose.

The final 30% of the tested history is held out. Reveal makes those results visible for that strategy version, records the action and cannot be reversed.

Development periodHeld out
LOCKED
Earlier historyLatest 30%
What happens when I press Reveal?

The hidden-period result is added to the saved report and the reveal time is recorded. Review the rules and report context first; that version cannot be re-locked.

Context travels with the result

The conditions are part of the evidence.

Reports include the assumptions needed to read a simulated result: the strategy version, test window, market, costs and execution limits.

REALISM

Fees, slippage, funding source and liquidity checks used in that run.

DATA

Market, venue, timeframe, candle count and historical range.

ENOUGH DATA?

Out-of-sample trades against a 100-trade reference, plus the minimum-track-record estimate; thin samples are called out.

METHOD

Versioned methodology and report integrity status.

RUN DEPTH · LUCK TEST200–1,000 randomized comparisonsBAD-LUCK STRESS200–2,000 simulated sequences

Ranges reflect Free, Pro and Elite evidence entitlements; unavailable checks remain labelled in each report.

Share the evidence, not the hype

A shareable card tied to the report.

When a member chooses to share, the Evidence Card carries an engine signature and the context needed to identify what was tested. A signature checks report integrity; it does not turn simulated results into a forecast.

See Community

The research behind the tests

Established methods. Plain-language results.

Evidence Lab uses published statistical checks alongside simulation and market comparisons. The full versioned method is available to read.

Deflated Sharpe & minimum track record Bailey & López de PradoProbability of Backtest Overfitting Bailey, Borwein, López de Prado & Zhu · 2015Sharpe ratio haircuts Harvey & Liu · 2015Robustness checks Monte Carlo stress & random-entry comparisons

Evidence Lab reports describe historical simulations under stated assumptions. They do not prove an edge, guarantee a result or predict future performance. The grade is one part of a decision, not permission to trade.