BACKTESTMARKET
Quants: Run This 7 Step Data Audit Before Backtesting on MT4/MT5
backtesting·

Quants: Run This 7 Step Data Audit Before Backtesting on MT4/MT5

A reproducible, audit-first checklist for quants validating MT4/MT5 backtest data. Preserve snapshots, run checksums, classify per-year gaps, smoke-test,...

By BacktestMarket Team
backtesting data qualitybacktesting data integritydata accuracy for testingimproving backtesting resultsbacktesting data strategiesdata reliability in trading

Researcher auditing historical trading data

High-quality backtesting data means completeness at every decision timestamp, the correct price type and resolution for your strategy, and a preserved snapshot you can reproduce months later. Before trusting any result, run a small reproducible audit: build an expected timestamp grid, classify gaps year by year, checksum your archive, and run a deterministic smoke test. Audit-ready sources, including BacktestMarket's minute-bar datasets, make this step faster but never optional.


TL;DR:

  • High-quality backtest data must include complete timestamps, accurate price types, and a preserved snapshot for reproducibility, requiring careful audit procedures.
  • Common data flaws like gaps, synthetic bars, timezone misalignments, and corporate-action mismatches can significantly distort results if not identified through targeted checks.
  • Filtering strategy requirements by resolution and data fields ensures you avoid unnecessary storage costs, with tick-level data needed only for microstructure-sensitive strategies.
  • Conducting a structured audit, including checksum verification, gap classification, and smoke tests, is essential before model calibration or optimization.
  • Choosing a data vendor should involve explicit documentation of coverage, adjustments, and support capabilities, with ongoing audits necessary to maintain data integrity over time.

Table of Contents

Why data quality matters for backtest accuracy

Bad data does not just add noise, it changes conclusions. A missing block of bars during a volatile session can erase the drawdown that would have triggered a stop, making a strategy look safer than it is. A feed that silently substitutes last price for mid price inflates fills on spread-sensitive systems. Omitted transaction costs do something similar: they turn a marginal edge into a false winner, because every basis point of spread or slippage you forget to model becomes a basis point of phantom profit.

The MQL5 reproducible audit demonstrated this directly: running the same deterministic strategy against three different feeds produced materially different results, even though all three claimed to cover the same instrument and period. The divergence was not random. It traced back to specific causes that the audit's authors recommend decomposing into three buckets:

  • Cost effect: differences in spread, commission, or slippage assumptions between feeds or models.
  • Price effect: differences in the underlying price series itself, such as bid versus mid versus last.
  • Trade effect: differences in which trades actually trigger, caused by gaps, timestamp shifts, or synthetic bars.

This decomposition matters because a single aggregate difference in profit factor tells you nothing about where to fix your pipeline. A strategy that is sensitive to the cost effect needs a better execution model. One sensitive to the price effect needs a different data source entirely. One sensitive to the trade effect usually has a gap or timestamp problem hiding in the feed.

Historical validity decays, too. Academic re-examinations of long-running predictive signals found that many published effects weaken when extended out of sample, a pattern consistent with data snooping rather than genuine structure. The practical defense is the same one serious researchers use for any dataset: keep an untouched out-of-sample window, run walk-forward tests instead of a single in-sample optimization, and write down every configuration you tried, not just the one that worked.

The metric to watch is not your headline return. It is robustness across regimes: does the strategy hold its edge in a low-volatility year and a crisis year, on a feed with gaps and on a clean one, with realistic costs and without them? A result that only survives under one narrow set of assumptions is not an edge, it is an artifact of the data you happened to use.

Backtests with large, well-constructed samples and documented lookback windows produce more reliable conclusions than small, undocumented ones. This is consistent with regulatory commentary on back-testing frameworks, which stresses clear lookback periods and remediation steps when results fall outside expectations.

Common data defects and how to spot them fast

Most backtest failures trace back to a small set of recurring defects. Knowing the signature of each one lets you catch it in minutes instead of discovering it after you have already optimized around bad data.

Gaps are the most common and the most misunderstood. Not every missing bar is a defect. Weekends, exchange holidays, and low-liquidity overnight windows produce legitimate gaps that a correct session calendar should explain. Per-year, per-session gap reporting is the only way to catch this. The MQL5 audit workflow recommends classifying every missing timestamp slot against an expected grid, then reporting defects by year rather than as one aggregate number.

Synthetic bars are the next most common issue. Some feeds fill missing intraday periods with flat OHLC bars (open, high, low, and close all identical) rather than leaving them absent. Others space intraday bars at daily intervals when the underlying request failed silently. Neither failure mode announces itself. You have to check for it directly.

  • Scan for runs of bars where open equals high equals low equals close, which almost never happens in real intraday trading.
  • Check timestamp spacing within each day: minute data with gaps of 1,440 minutes (a full day) instead of 1 minute signals a broken export, not a quiet market.
  • Compare volume or tick_volume fields against zero or null runs that align suspiciously with the flat-OHLC bars above.

Pro Tip: Validate by timestamp spacing and volume behavior, not by eyeballing price charts; a chart can look smooth even when the underlying bars are synthetic.

Timezone and session misalignment shows up as trades that trigger an hour early or late relative to the actual market open, usually because a feed's timestamps are in server time, UTC, or exchange time without saying which. Cross-check a handful of known session opens (London, New York, Tokyo) against your data's timestamps before you trust anything built on intraday session logic.

Corporate-action mismatches are a separate failure class entirely. A raw price series that is not adjusted for splits or dividends will show a return that never happened, and a series adjusted with an undocumented method may double-count or miss an adjustment. QuantLib's calendar reference treats corporate actions and exchange calendars as explicit, dated inputs rather than something to infer from price behavior, which is the right mental model: know exactly what adjustment method your vendor used and whether raw and adjusted series are both available.

Raw and adjusted price series diverging

Vintage issues are the subtlest. The same instrument, downloaded today versus downloaded two years ago, can carry different values because a vendor revised its construction methodology. A Review of Finance study on Fama-French factor data found that factor returns differ materially depending on when the data were downloaded, driven by retroactive methodological changes rather than new information. If you cannot say which vintage of a dataset your backtest used, you cannot reproduce it.

What resolution and fields your strategy actually needs

Not every strategy needs tick data, and buying more resolution than you need just adds storage and processing cost without improving your results. The right starting point is matching your strategy class to the minimum viable feed.

  1. Daily or statistical strategies (factor rotation, macro trend following, monthly rebalancing) generally do fine with daily OHLCV bars, provided adjusted and raw price series are both available and corporate actions are documented.
  2. Intraday bar-based strategies (breakout systems, session-based mean reversion, most retail algorithmic setups on MT4 or MT5) typically need minute bars at minimum, with volume or tick_volume included to filter low-liquidity periods.
  3. Microstructure-sensitive strategies (scalping, market making, anything that models fill probability or queue position) require tick-level or bid/ask data, because minute bars average away the spread dynamics these strategies depend on.

The field list matters as much as the resolution. Bid and ask, not just a single "price" column, are essential for anything that models spread cost directly. Last trade price matters for strategies built around actual executed trades rather than quoted prices. Tick-level timestamps, down to the millisecond where available, are necessary for queue-position or latency-sensitive models. Volume or tick_volume is useful even for slower strategies, since a liquidity filter can keep a backtest from trading through conditions that would never fill in reality.

A reasonable rule of thumb: if your strategy's edge is measured in multiple pips or percentage points, minute bars with a documented spread model usually suffice. If the edge is measured in fractions of a pip, you need tick or bid/ask data, because that is exactly the scale at which minute-bar averaging destroys the signal you are trying to measure. BacktestMarket's minute-bar guidance covers this boundary in more detail for traders deciding between the two.

Step-by-step audit workflow you can run today

The audit below takes a few hours the first time and minutes on every dataset after that, once scripted. It follows the sequence validated in the MQL5 reproducible audit, which showed that skipping any one of these steps let a flawed feed pass undetected until live trading exposed it.

  1. Preserve the raw archive before touching it. Copy the original export to a read-only location the moment you receive it. Every later step should operate on a copy, never the original.
  2. Compute and store a checksum. A simple hash of the raw file lets you prove, months later, that the data you are analyzing is the exact file you started with, not a silently re-exported or edited version.
  3. Build the expected timestamp grid. For the instrument's trading calendar and session hours, generate every timestamp that should exist. This grid is the reference against which every gap gets measured.
  4. Classify every missing slot, by year and by session. Mark each gap as benign (holiday, weekend, known illiquid window) or unexplained. Report defect rates per year, not as a single global percentage, since that is the only view that reveals a bad quarter hiding inside an otherwise clean decade.
  5. Sample-compare against an independent feed. Pull a few weeks of overlapping data from a second source and compare bar by bar. Material divergence here is a signal to decompose into cost, price, and trade effects rather than assume one feed is simply "wrong."
  6. Run a deterministic smoke-test strategy. A fixed, simple rule set (something like a moving-average crossover with no optimization) run against the new dataset and a known-good one will expose behavioral differences fast. If the smoke test trades differently on timestamp or gap handling alone, the underlying feed has a problem worth fixing before you build anything real on it.
  7. Freeze the exact snapshot and record its version. Once the dataset passes, lock it: store the file, its checksum, the download date, and the vendor's stated vintage together as one reproducibility record.

Pro Tip: Treat how you requested the data as part of the audit, not an afterthought. The MQL5 team found that range-based versus paged export requests against the same broker can return different datasets, so the request method itself needs to be documented alongside the data.

This sequence is deliberately boring. Each step is cheap and mechanical, which is the point: a reproducible audit should not depend on judgment calls that vary between researchers or between runs. BacktestMarket's MT5 data-quality workflow walks through the archive-preservation and checksum steps in platform-specific detail, and the 99% modeling quality guidance for MT4 covers the smoke-test step for traders working in that platform specifically.

Modeling costs, slippage, and liquidity without fooling yourself

A backtest that ignores transaction costs is not a weaker version of a good backtest, it is a different, usually wrong, question. Spread, commission, slippage, and market impact each behave differently and deserve separate models rather than one blended "cost assumption."

  • Spread should come from the data itself where bid/ask is available, snapshotted at the relevant timestamp rather than assumed constant across the day.
  • Commission is typically fixed or volume-tiered and is the easiest component to model correctly, since it usually comes straight from a broker's published schedule.
  • Slippage needs a synthetic curve built from observed fill behavior or tick data where available, since it grows nonlinearly with order size and volatility rather than staying flat.
  • Market impact matters most for larger orders relative to typical volume and is often the piece researchers skip entirely, which quietly favors strategies that would struggle at real size.

Returns measured against mid price alone overstate what a strategy could actually have captured, particularly for high-turnover systems that cross the spread frequently. The fix is not a single conservative cost assumption, it is a sensitivity test: run the same strategy across a range of spread and slippage assumptions, and across both calm and volatile periods, to see whether the edge survives realistic friction or only survives under the most generous one. Strategies whose profitability collapses between a 1-pip and a 2-pip slippage assumption are telling you something important about how fragile the result is.

Whatever cost model you land on, record it alongside the data vintage you used. A backtest result without its cost assumptions attached is not reproducible, even if the price data behind it is perfectly clean.

Reproducibility, vintages, and handling vendor revisions

A dataset that is more accurate today may still be the wrong dataset for a backtest meant to represent a historical decision point. The distinction is between corrected data and historically knowable data: a value a vendor revised last year, however accurate, was not available to a strategy making decisions in real time two years ago. Using the corrected version anyway quietly leaks information from the future into a simulation that is supposed to represent the past.

The Review of Finance study on Fama-French factor revisions makes the scale of this problem concrete: factor returns for the same historical period differ depending on when the data were downloaded, driven by methodological changes vendors made after the fact. Treat vendor revisions themselves as a tracked event, the same way you would track a code change in your strategy logic.

Practical snapshot management does not require elaborate infrastructure, just discipline:

  • Store the checksum of every raw export alongside the file itself, so you can prove later which exact bytes you used.
  • Freeze a clean export after each audit passes, separate from any working copy you continue to edit.
  • Keep the import script that loaded the data into your platform version-controlled, since an import process change can alter results as much as a data change can.
  • Stamp every backtest report with the dataset vintage, download date, and checksum, not just the instrument and date range.

When you publish or share a backtest, the dataset vintage belongs in the methodology section next to the strategy parameters, not buried in a footnote. A result that cannot be traced back to a specific, checksummed snapshot is a result nobody else can verify, including you, six months from now.

How to evaluate a data vendor before you buy

Choosing a historical data provider is a short checklist, not a leap of faith. Match the vendor's documentation against what your audit workflow actually needs, before you commit to a dataset you will later have to defend.

  • Coverage window and resolution: does the vendor state exactly which years and bar sizes are included, or do you have to find out after purchase?
  • Corporate-action documentation: is the adjustment method named explicitly, with raw and adjusted series both available?
  • Checksums and import-ready snapshots: can you verify the file you received matches what the vendor published, and does it import directly into MT4 or MT5 without manual reformatting?
  • Per-year gap reporting: does the vendor publish defect rates by year and session, or only a single global completeness percentage?
  • Engineer-level support: when a gap or anomaly shows up, can you reach someone who understands the data collection process itself?

Any of these should push you toward a different vendor or at minimum a longer independent audit before you trust the data.

Pro Tip: Match the vendor's documented coverage window against your strategy's stress-test regimes before buying. A dataset that covers the last five calm years tells you nothing about how your strategy behaves in a crisis, no matter how clean it is.

BacktestMarket's historical data catalog and its outlier-handling guidance for M1 data are built around exactly this checklist: documented coverage, checksummed exports, and import-ready bundles rather than a single completeness number.

Why audit discipline matters more than any single data source

Backtesting data quality gets treated as a solved problem by most retail traders, something you buy once and stop thinking about. It is not. The dataset you trust today will get revised, your vendor's adjustment method will change, and the gap you dismissed as benign may be sitting on the exact week your strategy's drawdown would have triggered. The habit that actually protects you is not finding a perfect vendor, it is running the same small audit on every dataset, every time, regardless of how reputable the source.

The uncomfortable part of this discipline is that it rarely changes your conclusion about whether a strategy works. Most of the time it confirms what you already suspected. Its value shows up in the rare case where it does not confirm your assumption, because that is exactly the case where skipping the audit would have cost you real money. BacktestMarket's blog and import guides exist to make that habit cheaper to maintain, not to replace it.

Run the audit checklist before you optimize anything. A strategy tuned against unaudited data is a strategy tuned against noise, however good the equity curve looks.

— Start

BacktestMarket: ready-to-import data built for the audit workflow

Every step in the audit above takes less time when the dataset arrives already clean. BacktestMarket sells complete minute-bar historical datasets across forex, metals, stock indices, bonds, and commodities, built for direct import into MT4 and MT5 without reformatting.

Healthcare Stocks

The company states that its data undergoes bar-by-bar verification with documented adjustments, which is the same standard this article argues every quant should demand regardless of vendor. Instead of spending a weekend writing grid and checksum scripts from scratch, you can start from a dataset built to pass that audit on the first try.

  • Browse the full Historical Data catalog for coverage by asset class and resolution.
  • Check healthcare sector stock data if your strategy needs sector-specific equity history.
  • Compare the Annual Plan, listed at 119 EUR per year, against a one-time dataset purchase depending on how often you need fresh vintages.
  • Review Expert Advisors and Indicators if your workflow also needs automated strategy tools alongside the raw data.

Support comes directly from the engineers who collect the data, which matters most the moment your own audit flags something you cannot explain. Check the catalog, run your checksum, and move on to testing instead of cleaning.

FAQ

What makes historical data "high quality" for backtesting?

High-quality data means completeness at every expected decision timestamp, the correct price type and resolution for your strategy, and documented corporate-action handling. It also means you can reproduce the exact dataset later, which requires a checksum and a recorded vintage, as described in the MQL5 reproducible audit.

Why does a small missing-data percentage still matter?

A low global missing percentage can hide a feed that is almost entirely absent during one critical month or volatility spike. Per-year, per-session gap classification, rather than one aggregate number, is the only way to catch this kind of concentrated defect.

Do I need tick data or will minute bars work?

Minute bars are usually sufficient for intraday bar-based strategies where the edge is measured in multiple pips or more. Strategies sensitive to spread dynamics, like scalping or market making, need tick-level or bid/ask data because minute-bar averaging erases the exact signal they depend on.

How should I handle a vendor's revised or "corrected" historical data?

Preserve the original vintage you used for any backtest meant to represent a historical decision point, since a later correction was not knowable at that time. The Review of Finance study on factor data revisions shows that methodology changes alone can materially shift historical returns.

Does BacktestMarket provide checksums or import-ready files?

BacktestMarket states that its datasets are built for direct import into MT4 and MT5 and undergo bar-by-bar verification with documented adjustments. Full coverage details by asset class are listed on the Historical Data page.

Sources

The claims and figures above draw on a reproducible audit methodology for MetaTrader data, peer-reviewed research on out-of-sample decay and factor-data vintages, and quant library documentation on corporate-action and calendar handling. Readers who want to replicate the audit workflow or review the underlying research can start with the sources listed below.

Recommended

Related resources

Explore BacktestMarket's Expert Advisor robots to put the ideas in this article into practice.

Newsletter

Stay updated

New datasets, expert advisors, discounts, and trading insights — straight to your inbox.

Cart

Your cart is empty

Add some products to get started.