BACKTESTMARKET
Market Data Formats for Developers & Quants: SBE, Parquet, Dollar Bars
Algorithmic Lessons·

Market Data Formats for Developers & Quants: SBE, Parquet, Dollar Bars

A developer and quant primer on market data formats: SBE/FAST feeds, Parquet and CSV storage, tick/volume/dollar sampling, ingestion checks, and...

By BacktestMarket Team
financial data structureshow to choose data formatsmarket information formatsdata format standardsfinancial data modelsstructured market data

Developer comparing structured market data views

Market data formats fall into three practical categories: real-time serialization protocols that move data off exchange feeds, storage formats that archive it, and sampling structures that reshape it for analysis. For feeds, learn Simple Binary Encoding (SBE) and FAST first. For storage, learn Parquet and CSV. Everything else builds on those five.


TL;DR:

  • Market data formats are categorized into real-time protocols like SBE and FAST, storage formats like Parquet and CSV, and sampling structures such as dollar and volume bars.
  • Low-latency feeds typically use SBE for compact serialization, with FIX defining message content and FAST optimizing bandwidth over UDP, especially for L2 and L3 data.
  • Data storage favors Parquet for analysis due to its columnar compression, while CSV remains the universal format for compatibility, though it is less query-efficient at scale.
  • Event-based sampling methods like dollar bars provide more stable and noise-reduced inputs for machine learning models compared to fixed time bars.
  • Proper schema design with versioning, explicit units, and machine-readable definitions is essential for data reusability and consistent cross-system validation.

Table of Contents

Real-time feeds: SBE, FAST, FIX, and the L1/L2/L3 split

FIX is the messaging standard that defines what a market data message means. It does not define how those bytes travel on the wire, which is where encoding layers come in. Simple Binary Encoding (SBE) is the FIX Trading Community's preferred encoding for low-latency binary serialization, favored by venues that need compact, predictable-size messages parsed without heavy string handling. FAST takes a different tradeoff: it strips redundancy from repetitive fields to shrink stream size, which suits bandwidth-constrained multicast feeds rather than raw decode speed.

Feed granularity determines what you actually receive:

  • L1 gives best bid, best ask, and last trade, enough for a ticker but not for order flow analysis.
  • L2 adds depth of book, multiple price levels on each side, which lets you model liquidity and slippage.
  • L3 exposes individual orders rather than aggregated levels, giving full order book reconstruction.

Exchange feed specs typically define snapshots, incremental updates, and recovery mechanisms, and the JSE Market Data Gateway spec is a concrete example: it uses FAST over UDP with explicit field lists for best bid/offer, trade messages, and sequence-based recovery. Any consumer needs logic for out-of-sequence packets, gap detection on multicast, and a clear source of truth for timestamps, since exchange time and receipt time diverge under load.

Choosing a storage format: CSV, Parquet, and Avro compared

CSV remains the universal exchange format precisely because every tool reads it, and large historical products still ship that way: the NYSE TAQ dataset is distributed as gzip-compressed CSV mapped closely to the original live-feed events. But CSV is row-oriented and untyped, which makes it expensive to query at scale. Apache Parquet solves that with columnar storage and built-in compression, letting an analytical query scan only the columns it needs instead of parsing every row.

Avro sits between the two: row-based like CSV but with an embedded schema, which makes it a good fit for raw ingestion pipelines that will later be converted to Parquet for analysis. A practical pattern:

  • Ingest raw feed messages as Avro or newline-delimited JSON, since the schema travels with the data.
  • Validate records against expected types, ranges, and monotonic timestamps before anything downstream touches them.
  • Convert validated batches to Parquet for querying, keeping CSV exports on hand for tools that only accept flat files.

Pro Tip: Keep a PyArrow-compatible reader in your toolchain as a fallback; almost every Parquet quirk you hit is a version mismatch between writer and reader libraries.

Time bars versus event-based bars: tick, volume, and dollar sampling

Time bars sample the market at fixed intervals, one bar per minute regardless of activity. That is simple but statistically awkward: a quiet minute and a frantic minute get equal weight. Event-based bars sample on activity instead of the clock, and research on financial data structures argues they often produce more stable inputs for machine learning models than fixed time bars.

The main variants:

  1. Tick bars close after a fixed number of trades, smoothing out periods of low activity.
  2. Volume bars close after a fixed number of shares or contracts trade, useful when trade sizes vary widely.
  3. Dollar bars close after a fixed dollar amount changes hands, which stays more stable across price regimes than tick or volume bars because it adjusts automatically as price moves.
  4. Imbalance bars close when buy and sell volume diverge past a threshold, capturing informed-flow moments rather than raw activity.

Building dollar bars conceptually: sum notional value (price times size) trade by trade, and once the running total crosses your threshold, close the bar and reset the counter. Backtests on liquid instruments tend to favor dollar or volume bars for feature stability, while time bars still work fine for coarse, low-frequency signal work. For deeper implementation notes and code, minute-bar data structured this way reduces the noise that fixed-interval sampling introduces into a backtest.

Schema design and metadata: what makes a format reusable

A schema that only you can read is not a schema, it is a liability. Start with stable field names, explicit units (basis points versus percent, milliseconds versus microseconds), and enumerations instead of free-text status fields. Ship the schema itself in a machine-readable form, JSON Schema, Avro schema, or Protobuf definitions, so downstream teams validate against it rather than guessing from sample rows.

  • Version the schema explicitly and document every breaking change, not just additive ones.
  • Include sample records and a validation rule set alongside the schema file, not in a separate wiki.
  • Fix your timestamp and timezone convention once (UTC is the safe default) and never mix conventions across fields.

Pro Tip: A short README with a controlled vocabulary for field names does more for adoption than a longer spec nobody reads. When no formal industry standard covers your case, community-endorsed data standards still recommend a clear README plus machine-readable schema to keep the data reusable and FAIR-compliant.

How BacktestMarket packages historical data for import

The platform provides clean minute-bar historical intraday data across multiple asset classes, structured for direct import into trading platforms like MT4 and MT5 without a separate conversion step. Datasets are packaged as complete, ready-to-load sets per instrument, which shortcuts the ingest-validate-convert cycle most CSV or Parquet pipelines still need.

Before running any dataset through a backtest, check three things: bar-by-bar continuity (no unexplained gaps), timestamp alignment against your platform's own clock convention, and consistent adjustment handling on any back-adjusted futures series. Guidance on importing CSV data into MT4 walks through the mechanics of that final import step once the dataset itself checks out.

How BacktestMarket packages historical data for import — overview diagram

Where production data pipelines actually fail

Where production data pipelines actually fail — overview diagram

The real tradeoff is not latency versus features, it is ergonomics versus correctness, and most teams optimize the wrong one first. Out-of-order timestamps, duplicate trades near feed reconnects, and daylight-saving-time boundaries cause more production incidents than encoding choice ever does.

Before trusting a feed, replay a recorded session end to end, run constraint tests on every field (price greater than zero, monotonic sequence IDs, valid enum values), and feed the pipeline synthetic edge cases like a gap-fill or a duplicate snapshot. Fix ingestion correctness first, worry about shaving microseconds second.

— Start

Getting integration-ready historical data without building the pipeline yourself

Building a compliant ingest-validate-store pipeline from raw exchange feeds takes real engineering time. BacktestMarket shortcuts that by shipping clean minute-bar historical data across forex, metals, bonds, and stock indices since 2014, packaged for immediate use rather than raw exchange dumps you have to clean yourself.

Backtestmarket

  • Datasets arrive ready for direct import into MT4 or MT5, no manual reformatting step required.
  • The BTM Data Converter handles format conversion when you need a different structure than the default package.
  • Support comes directly from the engineers who build and maintain the datasets, not a general help desk.

For readers who want continuous access rather than one-off downloads, the Annual Plan covers ongoing data access at 119 EUR per year. Browse the full historical data catalog to see instrument coverage before you commit to a backtest architecture.

Sources

FAQ

What are the different types of market data?

Market data generally splits into real-time feed data (quotes, trades, order book depth), historical data used for backtesting and research, and reference data like instrument identifiers and corporate actions. Format-wise, feeds typically use binary encodings such as SBE, while historical archives commonly use CSV or Parquet.

What is L1, L2, and L3 market data?

L1 covers the best bid, best ask, and last trade price, the minimum needed for a quote display. L2 adds multiple price levels of depth on each side of the book, and L3 exposes individual orders rather than aggregated levels, enabling full order book reconstruction.

What are the 7 types of financial markets?

Definitions vary by source, but a commonly cited breakdown includes stock markets, bond markets, money markets, derivatives markets, forex markets, commodities markets, and cryptocurrency markets. Each has its own data formats and feed conventions, though the underlying serialization and storage approaches covered here apply broadly across them.

What is the 3-5-7 rule in trading?

The 3-5 rule is a risk management guideline, not a data format standard, and it is outside the scope of market data schemas covered here. It generally refers to capping risk per trade and per portfolio at set percentages, but readers should consult a dedicated risk management resource rather than a data formatting guide.

Why do quants prefer dollar or volume bars over time bars?

Event-based bars sample on market activity rather than the clock, which research on financial data structures suggests produces more stable inputs for machine learning than fixed time bars. Dollar bars in particular adjust for price changes over time, keeping bar frequency more consistent across different volatility regimes.

Recommended

Related resources

Explore BacktestMarket's historical data packs to put the ideas in this article into practice.

Newsletter

Stay updated

New datasets, expert advisors, discounts, and trading insights — straight to your inbox.

Cart

Your cart is empty

Add some products to get started.