
Clean forex data is a price history that is consistently timestamped in UTC, deduplicated, checked for anomalies, documented for source and format, and ready to import into a trading platform without manual fixes. If you are starting from a raw file right now, the first move is a chronology check: sort every row by timestamp and calculate your duplicate rate before you touch anything else. The rest of this guide walks through the full cleaning workflow and the audit checklist that follows it.
TL;DR:
- Duplicates should be identified with a tolerant rule, especially for near-identical timestamps within a few milliseconds, to avoid biasing backtests.
- Verify timestamps are strictly increasing and sort data chronologically if out-of-order, instead of discarding rows, to preserve data integrity.
- Maintain timestamps in UTC and record original feed and ingestion timestamps separately to support accurate reconciliation and debugging.
- Flag anomalies and outliers with reversible labels rather than deleting them, allowing for flexible strategy testing and data auditing.
- Run a comprehensive five-step audit, including timeline, deduplication, gap, and anomaly checks, before importing datasets into trading platforms like MT4 or MT5.
Table of Contents
- Why source, provenance, and formats matter for FX research
- Common raw-data defects and how to detect them
- How do you build a step-by-step forex data cleaning workflow?
- Preparing clean data for credible backtests
- An audit checklist for ready-to-import forex data
- Practical tradeoffs worth knowing before you clean your own data
- Ready-to-import datasets if you want to skip the cleanup
- FAQ
- Sources
Why source, provenance, and formats matter for FX research
Forex trades over the counter across a fragmented set of banks, brokers, and electronic venues, so there is no single authoritative tape the way there is for a listed stock. The BIS Triennial Central Bank Survey highlighted the extremely high volume of OTC turnover in April 2025 across 52 jurisdictions, a scale that underscores how many different venues are quoting slightly different prices at any given moment. That makes the question "where did this price come from" central to any backtest, not a footnote.
Format choice compounds the problem. Tick data preserves every quote change but is heavy to store and slow to query; minute bars (M1) compress that into open, high, low, close values that are far easier to work with but hide intrabar structure; platform exports (MT4/MT5), CSV, and Parquet each trade off portability against performance.
Before you clean anything, record (for practical guidance, see Backtesting - Cryptocurrency Trading Blog | AI Trading Strategies & Guides | Darkbot):
- The data source and venue, plus whether the quote is bid, ask, or mid.
- Feed timestamp versus the timestamp your system received the record.
- The session definition used (which market hours, which holiday calendar).
Common raw-data defects and how to detect them
Most forex datasets fail quietly. The defects below are the ones most likely to bias a backtest, and each has a quick diagnostic.
- Chronology errors: Check whether timestamps are strictly increasing. If
df['timestamp'].is_monotonic_increasingreturns false, re-sort by timestamp rather than discarding rows, since out-of-order arrival is usually a transport artifact, not bad data. - Duplicates: Compute a dedupe rate by grouping on symbol, timestamp, and bid/ask, then dividing duplicate rows by total rows. Exact matches are easy; near-duplicates within a few milliseconds need a tolerant rule rather than a strict equality check.
- Gaps and session errors: Build a gap table comparing expected bar counts per session against actual counts, then flag anything that lines up with a known holiday versus anything that does not, since the second case is more likely a feed outage. Daylight saving transitions are a frequent source of apparent one-hour gaps that are really timezone artifacts rather than missing data.
- Anomalies: Run a rolling z-score on price changes to flag statistical outliers, then label (never silently delete) anything beyond your threshold, and separately check field integrity (bid never exceeding ask, no negative volumes).
Pro Tip: Keep every flag as a reversible label in a separate column rather than deleting rows, so you can test strategies with and without the filtered points.
How do you build a step-by-step forex data cleaning workflow?
Turning raw ticks into backtest-ready bars is a sequence, and skipping a step tends to surface as a silent bias later rather than an obvious error now.
- Ingest: Pull from a persistent stream (WebSocket) where possible, or batch-ingest with logic that corrects for out-of-order arrival, as recommended in applied tick-cleaning workflows.
- Timestamp strategy: Normalize every record to UTC, but keep the original timestamp field alongside it, and also record your own ingestion timestamp for later debugging.
- Batch sorting: For high-throughput tick feeds, sort in cache-and-sort windows rather than resorting the entire file on each new record; this keeps ingestion fast without sacrificing chronology.
- Deduplicate: Apply exact-match keys first (symbol, timestamp, bid, ask), then a tolerant rule for microsecond-level collisions that exact matching misses.
- Label outliers: Flag statistical and contextual anomalies with a reversible marker rather than removing them outright.
- Resample: Build OHLC bars from midpoint or bid/ask depending on your execution model, and decide explicitly how to treat intervals with zero ticks rather than letting them default to a forward-filled flat bar.
Two things matter across every step: unify timestamps to UTC early, and never let a cleaning operation destroy the raw record it acted on.
Preparing clean data for credible backtests
A clean file is not the same as a backtest-ready one. Validation work still has to happen before you trust the results.
- Hold out a validation period that the strategy never saw during development, and avoid overlapping lookback windows that let information leak across samples.
- Favor histories that span multiple volatility regimes rather than one quiet year, since a strategy that only survives calm markets tells you little.
- Model execution realistically: include a spread profile, a slippage distribution, and fill heuristics rather than assuming every order fills at the quoted price.
Regulatory and institutional backtesting guidance supports this discipline directly. An SEC order on the ICC's back-testing framework describes separating validation from training data and favoring longer, regime-diverse histories to avoid biased inferences, a principle that transfers cleanly from margin-model testing to strategy backtesting.
Track your audit metrics as part of the dataset's metadata: dedupe rate, gap frequency, anomaly flag rate, and a check on volatility stability across sub-periods, alongside the sample's date bounds and source.
An audit checklist for ready-to-import forex data
Before you trust any dataset, run five checks: a chronology test, a dedupe rate calculation, a gap table against your session calendar, an anomaly flag rate, and an import smoke test into MT4 or MT5. Each has a one-line SQL or Pandas equivalent, from ORDER BY timestamp to a groupby().size() count on duplicate keys.
- Chronology and gap checks catch feed outages before they corrupt a backtest.
- A dedupe rate above a small fraction of total rows usually signals an ingestion bug worth fixing at the source.
- An import smoke test into MT4/MT5 confirms the file format matches what the platform expects, not just what the spec says it should.
Datasets documented since 2014 with published timestamp rules and direct engineer support, our audit guide for MT5 backtesting data walks through this exact checklist against a real file.
Pro Tip: Run the full checklist on a small sample before buying or importing a full history, since a bad timestamp rule shows up in the first thousand rows just as clearly as the millionth.
Practical tradeoffs worth knowing before you clean your own data
Labeling beats deleting almost every time. The only exception is a record you can prove is corrupt, like a negative volume or a crossed bid/ask that cannot represent a real market state. Everything else should stay in the dataset as a flagged, reversible row so you can rerun a strategy with different filter thresholds later.

Storage matters more than people expect once a dataset grows past a few years of tick data. Compressed Parquet, partitioned by symbol and date, keeps query times reasonable without the overhead of raw CSV.
On resolution, a simple rule of thumb holds up: use minute bars for strategies with holding periods of minutes or longer, and reserve tick data for genuinely high-frequency work where intrabar structure changes the outcome.
— Start
Ready-to-import datasets if you want to skip the cleanup
If building this pipeline yourself is not the point, clean datasets are available for purchase. Our GBPUSD 1mo dataset is a practical example to test the checklist above against: run the chronology and dedupe checks on it, confirm the gap table lines up with expected sessions, and try the import smoke test straight into MT4 or MT5.
Beyond single pairs, our Historical Data catalog covers forex, metals, bonds, and stock indices as complete, documented downloads, and our Annual Plan at 119 EUR per year gives ongoing access across that catalog. We also build Expert Advisors and Indicators for traders who want to automate the strategy side once the data side is settled.
- Download the sample and run the five-step audit checklist before anything else.
- Reach out to our engineering support team directly if a check flags something you cannot explain.
- Move to the Annual Plan when you need more than one instrument or a longer history.
FAQ
What timestamp format should clean forex data use?
Store timestamps in UTC to avoid ambiguity across sessions and daylight saving transitions, but keep the original feed timestamp in a separate field. This lets you reconstruct the source record if a downstream calculation ever looks wrong.
Should outliers be deleted or just flagged?
Flag outliers with a reversible label rather than deleting them, except for records that are provably corrupt, like a negative volume or a crossed bid and ask. This approach, recommended in applied tick-cleaning workflows, lets you rerun a strategy with different filters without losing the raw record.
How much history do I need for a credible backtest?
There is no single universal minimum, but regulatory backtesting guidance favors longer, regime-diverse samples with validation periods kept separate from training data. A history that only covers one calm year will understate risk in a volatile one.
What counts as a duplicate in tick-level forex data?
An exact duplicate shares the same symbol, timestamp, and bid/ask values, but near-duplicates within a few milliseconds of each other also need a tolerant matching rule. Computing a dedupe rate across both cases before cleaning tells you how serious the problem is in your specific feed.
Does BacktestMarket's data come ready for MT4 or MT5 without extra formatting?
Yes, datasets like GBPUSD 1mo are built for direct import into MT4 and MT5 alongside CSV and Parquet formats. Our Historical Data catalog documents the format and timestamp rules for each instrument so you can verify compatibility before importing.
Sources
- BIS Triennial Central Bank Survey (FX turnover, 2025)
- How to sort and clean forex API tick data for accurate market backtesting
- Order approving proposed changes to ICC Back-Testing Framework
Recommended
- Audit First MT5 Backtesting Data: Gap, Timestamp, Ready to Import
- 5 Audits Quants Must Run on Outlier Handled M1 Data Before MT4/MT5
- Fixing MT4 Missing Data: A Complete Recovery Guide
- How to Achieve 99% Modeling Quality in MT4 for Backtests
Related resources
Explore BacktestMarket's Expert Advisor robots to put the ideas in this article into practice.

