Comparing Backtesting Tools

Backtesting tools are best compared on five things: how you express a strategy, the data model behind it, how orders are filled, what costs are applied by default, and whether the tool helps you validate the result. Feature lists rarely decide anything — two tools with near identical features can report very different returns for the same strategy, because their defaults disagree about what a realistic fill looks like. The comparisons below are written from firsthand use, and this page explains the criteria they are judged on.

What actually separates one backtesting tool from another

How you express a strategy

Code, a visual builder, or plain language. This decides how fast you can test an idea and how easily a rule can be misstated — a silent source of wrong results that no amount of engine quality fixes.

The data model

Which markets, how far back, what bar resolution, and whether the tool handles trading calendars, holidays, and delistings. A daily-bar-only tool cannot answer an intraday question honestly.

Execution realism

When an order fills, and at what price. Filling at the signal bar's close instead of the next bar's open manufactures returns that cannot be captured live — a form of look-ahead bias baked into the engine itself.

Cost assumptions

What fees and slippage are applied by default, and whether you can change them. A zero-cost default makes every high-frequency strategy look profitable, and is the most common reason a backtest fails to survive contact with a broker.

Validation support

Whether the tool helps you test out-of-sample, walk forward, or vary parameters. A tool that only produces one number per run makes overfitting easy and detecting it hard.

How the main approaches differ

ApproachStrategy is written asYou control costs & fillsSetup required
Python library (zipline, backtrader)Python codeFullyData pipeline, environment, maintenance
Charting platform (TradingView)Pine ScriptPartlyNone
Hosted engine (backtester.run)Plain English, compiled to a specPer strategyNone

None of these rows is the right answer on its own. A library is the correct choice when you need an assumption the vendor did not anticipate; a hosted engine is correct when the cost of building and maintaining that infrastructure exceeds what the extra control buys you.

Firsthand comparisons

Comparing results across tools

Running the same strategy on two tools and getting two answers is normal, and usually says more about the tools than the strategy. Before concluding anything from the gap, align four things:

  1. The same bars — identical symbol, timeframe, date range, and source.
  2. The same fill rule — most commonly, the next bar's open after a signal.
  3. The same costs — identical commission and slippage, explicitly set rather than left at whatever each tool defaults to.
  4. The same capital and sizing — position sizing differences alone can swing the result more than the entry logic does.

If a meaningful difference survives all four, it is worth investigating. Whichever tool you land on, the result still needs validation before it means anything — no engine protects you from overfitting.

Backtest without picking a library

backtester.run runs a patched zipline engine with explicit fee and slippage settings per strategy, so you get a library-grade execution model without building the data pipeline. Describe the strategy in plain English and read the tearsheet — see how to backtest a trading strategy for the full process.

Start free →

Frequently Asked Questions

What should I compare when choosing a backtesting tool?
Five things decide whether a backtest is trustworthy: how you express a strategy, the data model and its resolution, how orders are filled, what costs are applied by default, and whether the tool supports out-of-sample validation. Feature checklists matter far less than these.
Does the backtesting tool change the result?
Yes, often materially. The same strategy run on two tools can differ by a wide margin because of fill assumptions, bar timestamping, and default fees and slippage. A tool that fills at the signal bar's close rather than the next bar's open will report results that are not reachable in live trading.
Is a Python backtesting library better than a hosted platform?
It is a trade-off, not a ranking. A library gives full control over data, costs, and execution, at the price of building and maintaining that infrastructure. A hosted platform removes the setup work but fixes the assumptions the vendor chose. The right answer depends on whether those assumptions match your market and horizon.
Which backtesting library has the most realistic execution model?
Among the open-source Python libraries, zipline's event-driven loop with a trading calendar and explicit slippage and commission models is the most conservative by default. Backtrader is more permissive out of the box, which is flexible but makes it easier to build an unrealistically optimistic backtest without noticing.
Can I compare results across different backtesting tools?
Only if you align the assumptions first: the same bar data, the same fill rule, the same fees and slippage, and the same starting capital. Without that, a difference in reported return tells you about the tools' defaults rather than about the strategy.

© 2026 backtester.run · All rights reserved · support@backtester.run

Backtest results are hypothetical and do not guarantee future performance.