Glossary

Multiple-testing bias

Multiple-testing bias is the effect where testing many strategies guarantees that some will look profitable through chance alone. The more ideas you test, the higher the bar any one of them must clear to count as evidence.

In plain terms

Take twenty strategies with no edge whatsoever — pure coin flips, 25 trades each. Run them. Roughly two will come back looking like winners. Not because they are, but because that is what randomness does across twenty attempts.

Now imagine you tested those twenty ideas one at a time over three months, keeping the two that worked. You would have no way of knowing you had just found noise, because from the inside it feels like research.

Why it matters

This is the mechanism that turns diligence into self-deception. A trader who tests one idea and finds it works has weak evidence. A trader who tests fifty and finds three that work has weaker evidence, not stronger — but it feels like the opposite.

The problem is well established in quantitative finance. Bailey and López de Prado's deflated Sharpe ratio (Journal of Portfolio Management, 2014) exists specifically to adjust a reported Sharpe ratio for the number of trials behind it. Harvey and Liu's work on evaluating trading strategies argues that given how many factors have already been tested, new candidates should clear a substantially higher statistical hurdle than the conventional one.

How Mithos handles it

The platform tracks how many related strategy variations you have tested and raises the required bar accordingly. The penalty accumulates per family of related ideas, and it cannot be reset by restructuring a strategy to look like a fresh start — because from a statistical point of view, it is not one.

This is deliberately the least popular feature in the product. It makes strategies harder to pass, which is the entire point.

Related
Common questions

Answers to the usual ones

Is testing lots of ideas bad?
No. Testing lots of ideas and then judging the winners as if they were the only idea you tested is bad. Explore as widely as you like, as long as the accounting is honest.
Why can't I just start fresh?
Because the ideas you already rejected still happened. Discarding that history does not discard the selection effect it created.