Skip to content

Backtesting

How Many Trades for a Backtest? Sample Size and Uncertainty

How many trades a backtest needs depends on the metric, its variability, the desired precision and whether trades supply independent evidence. There is no universal minimum sample size. At a 50% observed win rate, 100 independent trades give an approximate 95% margin of ±9.8 percentage points; 400 give ±4.9 points. A larger count does not repair biased data, overfitting or missing market conditions.

Illustrative cover for backtest sample size showing growing groups of observations beside narrowing uncertainty bands, with the official PineConnector logo.
More independent observations improve precision. They do not repair biased data or repeated tuning.

The trading metrics library explains what each measure means. This page asks how uncertain that measure is. All numerical examples are illustrative calculations, not strategy results or a recommended testing threshold.

Backtest sample size at a glance

Scroll horizontally to read every column.

Question Useful calculation What else matters
How precise is the observed win rate? Standard error √(W(1 − W) ÷ N) Independence, breakevens and a stable outcome probability
How precise is the mean trade result? Standard error s ÷ √N Result variability, outliers and a consistent trade definition
Is a mean different from zero under a model? t = mean ÷ standard error Model assumptions and how the strategy was selected
Does the sample cover the intended use? Dates, conditions and independent evaluation records A trade count alone cannot answer this

How does uncertainty in win rate shrink with N?

Let N be all closed trades and W be winning trades divided by N. Breakeven trades remain in N but are not wins. Treat each observation as win versus non-win for this calculation.

Standard error of win rate = √(W(1 − W) ÷ N)

Approximate 95% interval = W ± 1.96 × standard error

The normal-approximation interval is a familiar estimate for a binomial proportion.[1] It assumes independent observations of a common win probability and becomes unreliable near the boundaries or with sparse outcomes. It can even return limits below zero or above one; Wilson or exact binomial intervals address that problem.

Those alternatives do not fix dependence between trades or a changing strategy. A confidence interval describes uncertainty under its sampling model. A 95% procedure is designed to cover the fixed underlying value in about 95% of repeated samples under that model.[2]

Worked example: 30, 100, 400 and 1,600 trades

Illustrative, not a recommendation. Each row assumes an observed 50% win rate among all closed trades and independent, identically distributed win/non-win outcomes. The table uses the normal approximation, not an exact interval.

Closed trades N Approximate 95% margin Interval around 50%
30 ±17.89 percentage points 32.11% to 67.89%
100 ±9.80 percentage points 40.20% to 59.80%
400 ±4.90 percentage points 45.10% to 54.90%
1,600 ±2.45 percentage points 47.55% to 52.45%

For N = 100, the standard error is √(0.5 × 0.5 ÷ 100) = 0.05. Multiplying by 1.96 gives 0.098, or 9.8 percentage points. The interval is 50% ± 9.8 points, not 50% ± 9.8% of 50%.

Illustrative comparison of approximate 95% win-rate margins at an observed 50% win rate: 100 independent trades give plus or minus 9.8 percentage points, 400 give 4.9 points and 1600 give 2.45 points.
Illustrative normal approximation for a win/non-win proportion. Four times as many independent observations halve the margin; no row establishes strategy validity.

The square root creates diminishing returns: halving the margin requires four times the observations if the outcome probability stays the same. Rearranging the formula gives a planning calculation:

N ≈ 1.96² × W(1 − W) ÷ m²

Here m is the desired margin written as a proportion. At W = 0.5 and m = 0.05, N ≈ 384.16, rounded up to 385 observations. That is a model-based calculation for a ±5-point margin, not a minimum number of trades that validates a strategy.

Win-rate precision also leaves trade size unresolved. A well-estimated win rate can coexist with uncertain average wins and losses. Expectancy requires both frequency and magnitude.

How do you calculate the standard error of mean R?

For every closed trade, divide its net result by its own initial risk to obtain an R-multiple. Keep the initial risk even if a stop later moved. The resulting values are R1 through RN, including zero-result trades.

Mean R = Σ Ri ÷ N; sample standard deviation s = √(Σ(Ri − mean R)² ÷ (N − 1))

Standard error of mean R = s ÷ √N; t = mean R ÷ (s ÷ √N)

The t-statistic here tests a null mean of zero. NIST gives the general one-sample statistic as the observed mean minus the hypothesised mean, divided by its standard error.[2] At least two observations and a non-zero sample standard deviation are needed for this form.

Standard deviation describes the spread of individual trade results. Standard error describes uncertainty in their estimated mean under the model. Replacing one with the other changes the question.

A second worked example: mean R and the t-statistic

Illustrative summary statistics, not backtest results. Suppose N = 100, mean R = +0.20R and sample standard deviation s = 1.50R.

  1. Standard error = 1.50 ÷ √100 = 1.50 ÷ 10 = 0.15R.
  2. t = 0.20 ÷ 0.15 ≈ 1.33.
  3. A rough two-standard-error interval is 0.20 ± 0.30R, or −0.10R to +0.50R.

The formal two-sided 95% interval is mean R ± t0.975,N−1 × standard error.[2] Using two as the multiplier is only a convenient approximation. The Student t distribution supplies the multiplier for N − 1 degrees of freedom.

If 400 independent trades had the same mean and standard deviation, the standard error would be 0.075R and t approximately 2.67. The rough interval would narrow to +0.05R to +0.35R. The additional trades do not have to preserve those summary statistics; this comparison isolates the effect of N.

The uncapped System Quality Number has the same algebra, √N × mean(R) ÷ s. A different name does not remove the t-statistic's assumptions.

Why are “30 trades” and “statistically significant trades” misleading?

Thirty is not a point where a backtest becomes representative. The figure echoes a textbook rule of thumb that a sample mean is approximately normal when n is at least 30,[3] a statement about one approximation, not about strategy evidence. In the illustrative win-rate table, 30 observations still leave an approximate margin of almost 18 percentage points. For mean R, uncertainty also depends on how widely results vary and whether a few extreme values dominate.

A textbook t interval is exact for independent, normally distributed observations.[3] With other distributions, its approximation needs care. Skewed outcomes, heavy tails and serial dependence can make a nominal interval misleading. Increasing N helps only insofar as the added observations provide relevant evidence.

Statistical significance applies to a specified test under a model. It does not attach to a trade count, establish an economically useful effect or certify an execution system. A small estimated mean can be statistically distinguishable from zero yet be sensitive to omitted trading costs.

Selection matters too. Choosing the best-looking result after trying many rule variations changes the interpretation of a single reported test. TradingView warns about selection bias and overfitting, and describes testing a frozen configuration outside the sample used for optimisation.[4]

Why do regime coverage and independence matter?

  • Many positions, one event: several correlated instruments entered together can supply repeated exposure to the same market movement. Counting tickets does not establish independent observations.
  • A narrow period: a large number of trades in one volatility or liquidity environment says little about omitted conditions.
  • Rule changes: pooling results from materially different entries, exits or sizing rules estimates a mixture, not one frozen strategy.
  • Partial exits: splitting a position into several closing deals increases record count without necessarily increasing independent decisions.
  • Changing behaviour: old and recent trades may no longer share a common probability or result distribution.

Report the date range and conditions alongside N, then inspect meaningful subperiods and rule versions. A segment with few trades remains uncertain even when the pooled total is large. The losing streak guide shows another consequence of assuming independent outcomes.

Coverage and a stable statistical model can conflict: a longer history may add conditions while also combining different distributions. Preserve that distinction instead of claiming that more calendar time automatically solves sampling uncertainty.

How do you build a usable trade sample from the platforms?

As of 25 September 2026, TradingView's strategy report exports trade data from the Trades tab. The default range retains individual records for the latest 9000 trades while aggregate metrics are unaffected.[4] Record the actual export window before calculating N or comparing the file with report totals.

  1. Freeze the specification: save the rule version, inputs, symbol, feed, timeframe, date range and cost assumptions.
  2. Define one observation: state how entries, partial exits and breakevens become closed trades.
  3. Reconcile the list: confirm the exported count and dates; identify exclusions instead of silently dropping inconvenient trades.
  4. Add initial risk: preserve each trade's original stop and currency R for an R-based analysis.
  5. Keep evaluation data separate: after tuning, evaluate the frozen rule on data that did not drive the choice.

The native MT5 Strategy Tester evaluates Expert Advisors and offers a historical forward split.[5] A later historical interval is still simulated evidence. It is different from observing new demo trades after the rule was frozen. See out-of-sample testing.

For PineConnector, strategy validation and execution verification require separate records. A condition being true, an alert triggering, webhook delivery, processing, an EA request, broker acceptance, a deal and a position are distinct events. The setup test requires checking the alert, processing record and broker trade.[6]

A successful demo order verifies that particular path and configuration. It does not add a missing historical sample or establish a future mean. Conversely, thousands of simulated trades do not prove that a broker-side stop matches a strategy-only exit.

Frequently asked questions

How many trades do you need for a backtest?

No fixed trade count is sufficient for every backtest. Required sample size depends on the metric, variability, desired precision and independence of observations. With a 50% observed win rate, the normal approximation gives a 95% margin of ±9.8 percentage points for 100 independent trades. Data quality, market coverage and independent evaluation still need separate checks.

Are 30 trades enough to test a strategy?

Thirty trades do not establish a general standard of adequacy. At a 50% observed win rate, 30 independent observations give an approximate 95% margin of ±17.89 percentage points. Uncertainty in the mean result also depends on trade-result variability. Thirty observations may expose a coding problem, but that task is different from estimating a stable strategy characteristic.

What is a statistically significant number of trades?

Statistical significance belongs to a specified hypothesis test, not to a universal number of trades. For a mean R tested against zero, t equals mean R divided by its standard error. The interpretation depends on the distribution, independence and strategy-selection process. A significant statistic does not establish future results or validate the TradingView-to-broker execution path.

Does increasing the backtest sample size reduce overfitting?

More independent observations can reduce sampling uncertainty, but adding trades does not by itself remove overfitting. A strategy repeatedly adjusted after inspecting the same history can exploit that history's peculiarities. Evaluate a frozen configuration on separate data, preserve the development record and check for future-data leakage. A larger biased sample remains a biased sample.

Reviewed 25 September 2026. Facts were checked against the linked sources on that date. Nothing in this article was tested on a trading account.

Related reading

Sources

  1. NIST/SEMATECH – e-Handbook: Confidence intervals for a population proportion, accessed 25 September 2026.
  2. NIST/SEMATECH – e-Handbook: Confidence Limits for the Mean, accessed 25 September 2026.
  3. Penn State Eberly College of Science – STAT 500 Lesson 5: Confidence Intervals, accessed 25 September 2026.
  4. TradingView – Pine Script v6 User Manual: Strategies, accessed 25 September 2026.
  5. MetaQuotes – MetaTrader 5 Help: Strategy Testing, accessed 25 September 2026.
  6. PineConnector – Test your setup, accessed 25 September 2026.

PineConnector executes the instructions you send it. It does not select trades, manage money, or hold funds. Trading carries risk, and past performance of any strategy does not indicate future results.


Leave a comment

Back To PiCo Blog

Ready when your strategy is

You bring the strategy.We bring the infrastructure.

Connect TradingView to MetaTrader, choose where MT5 runs and put the full PineConnector workflow through its paces from your first month.

Strategy and trading decisions remain yours. The MT5 environment can be ours.

PineConnector Edge

Run the full PineConnector workflow.

$59/mo at launch

Core plan · 1 connection · 1 hosted MT5 environment

Try Core for $7

7 days of Core for $7, then $59/month at the launch price unless you cancel.