Bayesian Optimization for Trading Strategies
Bayesian optimization is the most sample-efficient parameter search method. It builds a probabilistic model — a surrogate — of how a strategy's parameters map to performance, then uses that model to pick the most promising combination to test next. By learning from every backtest it runs, it concentrates trials where improvement is likely and finds strong parameter sets in far fewer runs than grid or random search.
Definition
Bayesian optimization — a search that models the relationship between parameters and the objective with a Gaussian Process, using the model's predictions and uncertainty to choose each next trial. backtester.run seeds the model with a few random points, then runs an ask/tell loop, snapping the model's continuous proposals onto your step grid.
How does Bayesian optimization work?
The algorithm follows a disciplined cycle that gets smarter with every result it collects. Here is each step in the sequence:
- Seed with random points. Run a handful of random parameter sets — backtester.run uses
min(5, budget)warm-up points — to give the surrogate model enough data to make its first meaningful predictions. Without these initial observations the Gaussian Process has no signal to learn from. - Fit a surrogate model. A Gaussian Process is fitted to all results seen so far. Unlike a simple regression, it outputs two quantities for every untested point: an expected objective score and an uncertainty estimate. Points near observed results have low uncertainty; unexplored regions have high uncertainty.
- Pick the next point via the acquisition function. An acquisition function combines the expected score and the uncertainty into a single value for each candidate point. High-scoring and uncertain-but-promising regions both score well. The point with the highest acquisition value is selected as the next trial. This is how the algorithm decides whether to explore or exploit — automatically, without any manual tuning.
- Update and repeat. Run the backtest at the chosen parameter set, feed the result back into the model (the "tell" step of the ask/tell loop), refit the Gaussian Process, and ask for the next point. Each iteration narrows the model's uncertainty in the regions that matter most. Repeat until the trial budget is spent.
- Rank by objective. Return the full leaderboard of parameter sets tested, sorted by the chosen objective metric. The best result from the search is typically reached in a fraction of the trials that grid or random search would need.
backtester.run uses scikit-optimize's Gaussian Process surrogate under the hood. The model's continuous proposals are snapped to your step grid before each backtest, so indicator periods stay at whole-number values and thresholds stay within valid ranges.
Explore vs exploit
Every sequential search faces the same core tension: should the next trial go somewhere new and uncertain — explore — or should it refine a region that already looks strong — exploit? A pure exploration strategy wastes trials in regions that turn out to be poor; a pure exploitation strategy converges prematurely on the first good-looking local peak and misses better solutions elsewhere. Bayesian optimization's acquisition function manages this tradeoff automatically. Early in the search, when the surrogate is uncertain almost everywhere, it favours exploration. As the model accumulates evidence and the uncertainty in promising regions falls, it shifts toward exploitation, refining the most competitive parameter sets with each remaining trial. This self-tuning behaviour is what makes Bayesian optimization so sample-efficient compared to grid and random search, which have no mechanism for learning from earlier results.
When should you use Bayesian optimization?
Bayesian optimization pays the greatest dividend when the cost per trial is high or the budget is tight. The specific conditions where it outperforms the alternatives are:
- Each backtest is slow. Long histories, intraday minute-bar data, or complex multi-indicator strategies all increase per-trial cost. When a single run takes seconds, saving 60% of the trials is meaningful.
- Your trial budget is tight. If you are on a plan with a 20- or 100-iteration cap, Bayesian search extracts more signal from those trials than random sampling would.
- The response surface is reasonably smooth. Bayesian optimization assumes that nearby parameter values produce similar outcomes — that there are gradients to follow. Strategies where performance changes sharply between adjacent grid points violate this assumption, and the surrogate model is less helpful.
- You have several parameters to tune. Grid search blows up combinatorially with more than two parameters. Random search handles many parameters well but ignores what it has already learned. Bayesian optimization navigates high-dimensional spaces efficiently by building a model.
One important caveat: for fast backtests — strategies that run in milliseconds on short daily-bar histories — random search may reach an equally good result in less wall-clock time. Bayesian optimization fits and queries the surrogate model between every trial, which adds overhead that is negligible relative to a slow backtest but significant relative to a near-instant one. When backtests are cheap, the model-free random approach often wins on raw throughput.
Limitations
Bayesian optimization is the most sample-efficient of the three methods, but it is not without weaknesses. First, the per-step modelling overhead — refitting the Gaussian Process after each trial — is non-trivial; for strategies where backtests complete in milliseconds, this overhead inverts the efficiency argument and random search may cover more of the parameter space in the same wall-clock time. Second, the Gaussian Process assumes some smoothness in the objective surface: strategies whose performance fluctuates sharply between adjacent parameter values provide a poor signal for the surrogate, reducing its ability to guide the next trial intelligently. Third, and most importantly for trading applications, Bayesian optimization only controls how the search is conducted — it does not address the fundamental problem that the winning parameters are selected by looking at historical results. Any search method, no matter how clever, produces a result that is biased upward and must be validated out-of-sample. Always run the best-found parameters through walk-forward analysis on data the search never touched, and treat a result that does not survive that test as evidence of overfitting, not a flaw in the search algorithm itself.
Run Bayesian optimization on your strategy
backtester.run runs a Gaussian-process Bayesian search across real market data, learning from each backtest to find strong parameter sets in the fewest trials — describe your strategy in plain English to begin.
Start free →Frequently Asked Questions
- What is Bayesian optimization?
- Bayesian optimization is a parameter search that builds a probabilistic model — a surrogate — of how parameters map to performance, then uses it to choose the most promising combination to test next. By learning from every result, it finds strong parameter sets in far fewer backtests than grid or random search.
- How does Bayesian optimization choose the next point?
- It fits a Gaussian Process to the results seen so far, which predicts both an expected score and an uncertainty for every untested point. An acquisition function balances exploring uncertain regions against exploiting regions that look good, and the point that scores highest on it is tried next.
- When is Bayesian optimization the right choice?
- Use it when each backtest is slow or your trial budget is tight, and you want the best result from the fewest runs. It works best when the response surface is reasonably smooth, so nearby parameter values tend to produce similar results.
- What are the limitations of Bayesian optimization?
- It adds modelling overhead between trials, so for fast backtests random search can find as good a result in less wall-clock time. It also assumes some smoothness in the response surface; for jagged, noisy surfaces the surrogate model is less helpful.
- How does backtester.run implement Bayesian optimization?
- It uses a Gaussian Process surrogate from scikit-optimize. The search starts with a few random points to seed the model, then alternates asking the model for the next point and telling it the result. Continuous proposals are snapped to your step grid so periods and thresholds stay valid.