AI Order Flow Analysis Model

Trading futures and cryptocurrencies requires understanding hidden signals that are not visible on regular charts. We develop an AI model for Order Flow analysis that helps identify true intentions of market participants. Our team delivers the project turnkey—from data collection to deployment and ongoing support, ensuring a reliable solution for your trading strategy.

AI Development Areas

Frequently Asked Questions

Latest works

  • Development of a web application for FEEDME
    Development of a web application for FEEDME
    1344
  • Development of an online store for the company FURNORO
    Development of an online store for the company FURNORO
    1306
  • B2B Advance company logo design
    B2B Advance company logo design
    753
  • Development of a web application for Enviok
    Development of a web application for Enviok
    1049
  • AIDER company logo development
    AIDER company logo development
    992
  • CRM development for Chasseurs
    CRM development for Chasseurs
    1097

You trade E-mini S&P 500 futures. The price stays in a narrow range for 15 minutes, but the order book volume grows. Regular indicators say nothing, while the Order Flow model sees: buyer volume aggressively exceeds seller volume, CVD is rising, the Footprint shows clusters at the 4500 level. Two minutes later, the price breaks the level from below — you know in advance it is a false breakout. Our AI Order Flow analysis model provides this advantage. Solutions are already deployed in 10+ projects, and our team has over 5 years of experience in AI/ML for financial markets.

How Does Aggressor Classification Work?

Every trade has an initiator — the aggressive side. According to the Lee-Ready algorithm, proposed over 30 years ago: a trade at a price higher than the previous is buyer-initiated, lower is seller-initiated. If the price is unchanged, we look at the previous movement. We fine-tune the classifier on labeled data, achieving up to 95% accuracy for cryptocurrencies (Binance) and 92% for futures (CME).

What Are Delta and CVD?

CVD (Cumulative Volume Delta) is a key Order Flow indicator:

Delta = Buyer_Volume - Seller_Volume CVD = Σ Delta over period 

Positive CVD with rising price = trend confirmation. Negative CVD with rising price = divergence, often preceding a reversal. We build CVD for multiple time windows (1s, 30s, 5min) and feed it as a feature into the model. Additionally, Absorption: when a large player absorbs aggressive orders without price movement. These are support/resistance levels that the model learns to detect.

Feature Engineering: From Ticks to Features

Tick data is transformed into features via rolling window aggregation:

def compute_order_flow_features(trades_df, window_seconds=60):
    features = {}
    trades_df['initiator'] = np.where(trades_df['side'] == 'buy', 1, -1)
    features['buy_volume'] = trades_df[trades_df.initiator==1]['volume'].rolling(f'{window_seconds}s').sum()
    features['sell_volume'] = trades_df[trades_df.initiator==-1]['volume'].rolling(f'{window_seconds}s').sum()
    features['cvd'] = features['buy_volume'] - features['sell_volume']
    features['trade_imbalance'] = features['cvd'] / (features['buy_volume'] + features['sell_volume'])
    features['avg_buy_size'] = (features['buy_volume'] / buy_count)
    features['avg_sell_size'] = (features['sell_volume'] / sell_count)
    features['large_buy_ratio'] = (large_buy_volume / total_volume)
    return features

Volume Profile — a histogram of volume at price levels. VPOC (Volume Point of Control) — level with maximum volume, used as support/resistance. Time and Sales analysis: clusters of large trades in a short time = large player entering a position.

Footprint CNN: When a Neural Network Reads Clusters

A Footprint Chart (Cluster Chart) combines Order Book and Order Flow: each candle is split into price levels, each level contains [buyer_volume × seller_volume]. We feed this data as a 3D tensor [time_bins × price_levels × 2] into a convolutional network:

class FootprintCNN(nn.Module):
    def __init__(self):
        super().__init__()
        self.conv1 = nn.Conv3d(1, 32, kernel_size=(3, 3, 2))
        self.conv2 = nn.Conv3d(32, 64, kernel_size=(3, 3, 1))
        self.flatten = nn.Flatten()
        self.fc = nn.Linear(64 * ..., 1)

The CNN learns to detect divergences and absorption that are inaccessible to linear models.

Comparison of Approaches: ML vs Classic Indicators

Method Accuracy (1-min forecast) Training Time Interpretability
Linear regression on CVD 55-60% 1 hour High
Gradient Boosting 60-65% 2-4 hours Medium
Footprint CNN 65-70% 1-2 days (GPU) Low

The CNN on footprint data outperforms regular technical indicators by 2-3 times in price direction prediction accuracy on short intervals.

Accuracy Comparison by Instrument (1-minute forecast)

Instrument Linear Regression Gradient Boosting Footprint CNN
E-mini S&P 500 57% 63% 68%
BTC/USD 55% 61% 66%
EUR/USD 59% 64% 70%

The CNN consistently outperforms classic ML models on all instruments, especially on volatile markets.

AI Model Deployment Process

  1. Data audit — assess tick data quality, select source, calculate required volume (minimum 3 months of history).
  2. Feature Engineering — develop Order Flow, Volume Profile, Stacked Imbalance features.
  3. Baseline model — Linear Regression or LightGBM for a quick baseline (3-4 weeks).
  4. Advanced model — Footprint CNN with backtesting and optimization (3-4 months).
  5. MLOps Pipeline — deploy model to production (Kubernetes, vLLM, Kafka for streaming).
  6. Monitoring and drift — automatic quality reassessment, alerting on metric drops.

What Is Included in the Result

  • Source code of the model and pipelines (Python, PyTorch/TensorFlow).
  • Documentation: architecture description, run instructions, API specification.
  • Team training (2-3 sessions of 4 hours each).
  • Technical support for 3 months after deployment.
  • Guarantee: if the model does not meet target metrics (ROC-AUC ≥ 0.7 on validation), we refine it free of charge.

Estimated Timelines

  • Order Flow Feature Engineering + baseline regression: 3-4 weeks.
  • Footprint CNN with backtesting and production pipeline: 3-4 months.
  • The cost is calculated individually after a data audit and defined KPIs. Contact us for a consultation — we will select the optimal solution for your infrastructure and budget. Order a pilot project: we will conduct a data audit and show a baseline model in 2 weeks.