September 28, 2026

Why Crypto Data Breaks: Gaps, Duplicates, Symbol Drift, and Exchange-Specific Edge Cases

featured image

Crypto data looks simple until you try to build something serious with it.

A trade has a price, size, symbol, and timestamp. A candle has an open, high, low, close, and volume. An order book has bids and asks.

But once you start collecting crypto exchange data across many venues and long historical periods, the difficult problems are rarely the obvious ones.

A missing candle may be perfectly valid.

Two identical-looking trades may both be real.

A familiar ticker may refer to a different market than it did two years ago.

An apparent gap may come from an exchange outage, maintenance window, delisting, trading suspension, low liquidity, API throttling, or an ingestion problem.

For data engineers, quants, AI trading teams, and backtesting teams, reliable crypto data therefore depends on much more than connecting to an exchange API.

It requires normalization, sequencing, reconciliation, metadata, and enough context to understand what actually happened at the venue.

CoinAPI handles this problem across more than 380 integrated exchanges, 599,000+ symbols, and over 632 TB of historical market data.

But the scale is only part of the problem.

The real challenge is preserving meaning while hundreds of exchanges behave differently.

One of the easiest mistakes in a historical data pipeline is assuming that every time interval should contain an OHLCV candle.

That is not necessarily true.

A candle is created from trades. If no trades occur during a given interval, there may be no candle to generate.

For a heavily traded BTC pair, a missing one-minute candle deserves investigation.

For a thinly traded altcoin, that same gap may be completely legitimate.

Consider several possible cases:

What you seeWhat it may actually mean
One isolated missing candleNo trades, short maintenance, or temporary feed interruption
Several consecutive missing candlesExchange downtime, maintenance, or API outage
No candles before a particular dateThe market had not been listed yet
Frequent gaps on an illiquid pairNo trading occurred during those intervals
Simultaneous gaps across many symbolsExchange-wide outage or connectivity problem
Gap affecting only one pairSuspension, delisting, inactivity, or pair-specific issue

The visual symptom is identical: a blank interval.

The explanation is not.

Suppose a backtesting engine sees a missing one-minute bar and simply copies the previous close forward.

That may look harmless.

But the resulting candle represents market activity that never happened.

Repeated across a dataset, synthetic bars can produce false price continuity and artificial liquidity. They can also distort:

  • volatility calculations
  • stop-loss behavior
  • execution simulations
  • indicator calculations
  • time-series models

Missing data should therefore be investigated before it is repaired.

Fields such as trades_count, volume_traded, data_start, data_end, data_trade_start, and data_trade_end can provide the context needed to determine whether a gap is abnormal.

The rule is simple:

Do not confuse an empty market interval with a broken data pipeline.

Duplicate detection sounds straightforward.

It often is not.

Imagine receiving two trades with:

  • the same symbol
  • the same price
  • the same size
  • the same timestamp

Are they duplicates?

Possibly.

But two separate trades can legitimately occur at the same price and size within the same timestamp resolution.

This becomes particularly difficult when an exchange provides low-precision timestamps or weak native identifiers.

Deduplicating on:

timestamp + price + size

can therefore remove valid transactions.

A more reliable pipeline should consider identifiers and sequencing information such as:

  • uuid
  • native id_trade, where available
  • symbol_id
  • time_exchange
  • time_coinapi
  • sequence, where supported

Even then, exchange behavior matters.

Some venues provide strong native trade IDs.

Others provide weaker identifiers.

Events can also arrive late or out of order.

Historical processing therefore needs reconciliation rather than a single simplistic duplicate rule.

This distinction is easy to miss.

Real-time feeds are designed primarily to deliver new market events quickly.

Historical datasets can be processed later with additional context.

That makes it possible to reconcile:

  • duplicate events
  • late-arriving events
  • out-of-order records
  • exchange-side corrections

As a result, a real-time stream captured locally may not always bit-for-bit match the final historical dataset.

That is not necessarily a contradiction.

It reflects two different objectives:

Real-time crypto data is an operational feed. Historical crypto data is a reconciled record.

For trading infrastructure, latency matters.

For research, reproducibility matters.

A serious data architecture needs to understand the difference.

Another subtle failure mode is treating timestamps as if every event has one universally correct time.

It does not.

CoinAPI distinguishes between:

time_exchange

The timestamp reported by the venue.

and:

time_coinapi

The timestamp associated with when CoinAPI received the event.

That distinction matters because exchanges differ in clock precision and delivery behavior.

Messages can be delayed.

Exchange clocks can behave differently.

Several events may share the same low-resolution timestamp.

Messages may arrive in a different order from their exchange timestamps.

For a basic chart, this may not matter much.

For event replay, latency analysis, or high-resolution backtesting, it matters enormously.

Suppose three trades have exchange timestamps that appear identical but arrive sequentially.

Sorting purely by time_exchange could destroy their observable arrival order.

Conversely, using only receipt time ignores when the venue says the event occurred.

A robust system may therefore need both.

In crypto, “the timestamp” is not one truth.

There is the venue's clock and the collector's clock, and confusing them can corrupt sequencing, latency measurements, and historical replay.

CoinAPI standardizes timestamps in UTC, giving downstream systems a consistent time representation while preserving the distinction between exchange and receipt timestamps.

BTCUSDT looks like an identifier.

It is really just a label.

Across the crypto market, the same asset may appear under different native naming conventions. Markets may be renamed. Tokens may migrate or rebrand. Contracts may expire or disappear. Native exchange symbols can also differ significantly across venues.

That makes ticker-based historical joins dangerous.

For reliable crypto historical data, an instrument is better understood through a combination of:

  • venue
  • instrument type
  • normalized base asset
  • normalized quote asset
  • exchange-native symbol
  • availability period

CoinAPI exposes normalized metadata such as:

FieldPurpose
symbol_idCoinAPI normalized instrument identifier
exchange_idVenue
symbol_typeSPOT, FUTURES, PERPETUAL, OPTION, and others
asset_id_baseNormalized base asset
asset_id_quoteNormalized quote asset
symbol_id_exchangeNative exchange symbol
asset_id_base_exchangeNative base naming
asset_id_quote_exchangeNative quote naming
data_startEarliest available data
data_endLatest available data

CoinAPI also provides symbol mapping between its normalized identifiers and exchange-native identifiers.

That distinction becomes critical when data from multiple venues or multiple years is combined.

A BTC/USDT spot market on Exchange A is not identical to BTC/USDT on Exchange B.

The prices may be similar.

The market is not.

Liquidity, matching-engine events, outages, spreads, order book depth, and trades all belong to the individual venue.

This is why collapsing venue-specific data into a ticker alone can destroy information that later models may need.

For historical research, a better conceptual key is:

venue + instrument type + normalized base asset + normalized quote asset + validity window

rather than simply:

BTCUSDT

There is another historical-data problem hidden in symbol metadata: survivorship bias.

Suppose you want to backtest a strategy across crypto markets between 2022 and 2025.

If you begin by downloading today's list of active symbols, you have already changed history.

Markets that were delisted during that period are gone from your universe.

Tokens that failed, lost liquidity, were removed by exchanges, or disappeared entirely may never enter your test.

The result can make historical performance appear stronger than it really was.

CoinAPI keeps inactive and delisted markets queryable through their symbol identifiers and exposes availability fields including:

  • data_start
  • data_end
  • data_trade_start
  • data_trade_end

These fields are not merely metadata for an API catalog.

They are controls against survivorship bias.

A reliable crypto historical data pipeline should reconstruct the instruments that existed at the time being studied, not only the ones that survived until today.

Imagine that an exchange stops sending updates for seven minutes.

Your chart now contains a seven-minute gap.

  • Was your ingestion service down?
  • Was CoinAPI disconnected?
  • Did the exchange enter maintenance?
  • Was one market suspended?
  • Did trading stop naturally?

Without cross-market context, the gap itself does not answer the question.

A useful diagnostic is to compare multiple symbols from the same venue.

If dozens of unrelated markets stop updating at approximately the same time, an exchange-level outage, maintenance event, or connectivity problem becomes more likely.

If only one market stops, investigate:

  • listing status
  • delisting status
  • pair suspension
  • market inactivity
  • pair-specific exchange problems

This leads to an important engineering principle:

A data pipeline needs to distinguish collector failure from venue reality.

The same visual gap can represent several completely different events.

And if an exchange genuinely stops publishing market updates, a data provider cannot manufacture real messages that never existed.

CoinAPI normalizes incoming market data, but it does not fabricate trades, quotes, order books, or candles during upstream silence.

OHLCV data is interval-based.

Order books are event-driven.

That distinction creates another common source of confusion.

A system might expect: one order book update every second.

But that is not how order books work.

If nothing changes at the exchange, there may be no update.

A liquid market may generate thousands of order book events per second.

A quiet market may produce none for a while.

Neither case is inherently wrong.

Historical reconstruction also requires order book updates to be processed according to the venue's update model and sequence.

Relevant fields can include:

  • symbol_id
  • time_exchange
  • time_coinapi
  • sequence
  • bids
  • asks
  • price
  • size
  • number_of_orders, where provided

On some feeds, an update with size = 0 represents removal of a price level.

Other fields can vary from exchange to exchange.

That is exactly why treating every venue as if it emits an identical order book format is dangerous.

For large-scale full-depth historical order book replay, CoinAPI Flat Files provide tick-level order book history, while REST historical order book access is better suited to snapshots and limited-depth queries.

A quiet order book is not necessarily broken.

It may simply mean the venue generated no book event.

Not every crypto data workload should use the same interface.

Trying to retrieve years of tick-level history through the same mechanism used for live application queries creates unnecessary complexity.

A useful division is:

WorkloadTypical access method
Latest market stateREST
Live market updatesWebSocket or FIX
Recent historical queriesREST
Large historical backtestsFlat Files
Full-depth historical order book replayFlat Files
AI and research workflowsAPIs, MCP, or Flat Files depending on scale

CoinAPI REST supports historical OHLCV periods ranging from 1SEC through long-duration aggregations, with historical queries supporting large record sets.

WebSocket is designed for live updates.

Flat Files are designed for bulk historical access and large-scale processing.

The interface should match the workload.

Otherwise, teams often end up building unnecessary local archives, issuing huge numbers of historical requests, or maintaining recovery logic that could have been avoided by using bulk data in the first place.

Reliable pipelines do not simply ingest data.

They test their assumptions about the data.

Here are several checks worth implementing.

For every symbol_id and candle period:

  1. Generate the expected time buckets.
  2. Compare them with the returned candles.
  3. Investigate missing intervals.
  4. Check data_start and data_end.
  5. Verify whether trades occurred.
  6. Compare the gap with other symbols from the same exchange.
  7. Check whether the market was listed, suspended, or delisted.

The important step is investigation before repair.

Track updates across multiple unrelated markets on each venue.

If a large portion of them stop updating simultaneously, classify the event differently from a single-symbol gap.

This can help distinguish exchange-side events from instrument-specific behavior.

Do not deduplicate purely using:

timestamp + price + size

Use stronger identifiers where possible, including:

uuid

id_trade

symbol_id

time_exchange

time_coinapi

sequence

The exact logic may vary depending on what a particular venue provides.

Before requesting a long historical range, verify:

  • the normalized symbol
  • the exchange
  • instrument type
  • native exchange symbol
  • data_start
  • data_end
  • listing lifecycle

This prevents queries from silently assuming that a market existed throughout the entire test period.

Keep exchange time and receipt time separate.

Normalizing everything into one timestamp field may make storage simpler, but it can remove information needed for later latency or sequencing analysis.

Connecting directly to one crypto exchange is relatively straightforward.

Connecting to hundreds of exchanges while maintaining consistent identifiers, timestamps, historical continuity, duplicate handling, and market lifecycle metadata is a different engineering problem.

CoinAPI provides normalization and reconciliation across heterogeneous exchange feeds, including:

  • normalized exchange, asset, and symbol identifiers
  • standardized UTC timestamps
  • duplicate handling
  • timestamp validation
  • out-of-order reconciliation
  • exchange-side corrections
  • metadata for active and inactive markets
  • historical availability windows
  • consistent market schemas

That layer becomes more important as the number of exchanges, instruments, and historical years increases.

Because the value of a crypto data provider is not simply API access.

It is making inconsistent exchange-specific data usable as one coherent data layer.

The most dangerous crypto data problems are often not obvious errors.

They are plausible-looking records with the wrong interpretation.

A missing candle gets filled even though no trade occurred.

Two real trades are collapsed into one.

A renamed symbol is joined to the wrong historical instrument.

A delisted token disappears from a backtest.

An exchange outage is interpreted as an ingestion failure.

An order book with no updates is marked stale even though nothing changed at the venue.

These problems are difficult because the data still looks reasonable.

Reliable market-data infrastructure must therefore preserve not only values, but context:

which exchange, which instrument, which timestamp, which lifecycle state, which sequence, and which historical availability window.

That is what turns raw crypto exchange data into crypto data that can actually support trading systems, AI models, research, and backtesting.

CoinAPI provides real-time and historical crypto data across hundreds of exchanges, with normalized symbols, trades, quotes, OHLCV, order books, metadata, exchange rates, indexes, and bulk Flat Files.

Use APIs for live and targeted queries.

Use Flat Files for large historical datasets, backtesting, machine learning, and full-depth historical analysis.

Explore CoinAPI Market Data and historical crypto datasets.

Start Building with CoinAPI

Recent Articles