Why Most AI Price Forecasting Models Fail?
Traders and funds often invest resources in developing complex models, but on real markets they show losses. The reason lies in three traps: look-ahead bias, ignoring transaction costs, and overfitting to historical data. For example, a simple moving average model may show a Sharpe ratio of 1.5 in backtest, but in live trading without accounting for slippage and commissions, this drops to 0.3. We have encountered projects where a model with IC 0.08 on validation showed Sharpe 0.2 in production — due to ignoring slippage. Our approach eliminates such surprises. We solve these issues through purged walk-forward validation, a realistic transaction cost model (Almgren-Chriss), and strict factor selection control.
How to Choose the Forecasting Horizon?
The practical goal is not the exact price in N days, but a signal with positive expected value after transaction costs. Even a model with MAPE 3% on S&P500 stocks is useless if the strategy's Sharpe ratio is < 0. The horizon determines the signal type:
- Intraday (minutes-hours): microstructure signals, order flow imbalance — typical return 0.5–1.5% per trade.
- Short-term (1-5 days): momentum, mean reversion — average IC 0.05–0.08.
- Medium-term (1-4 weeks): earnings, macro catalysts — IC can reach 0.12.
- Long-term (months): fundamental valuation, factor exposure — more stable but requires higher accuracy.
The optimal horizon depends on the instrument's liquidity and rebalancing frequency. For less liquid assets, shorter horizons are less reliable.
What is Purged Walk-Forward Validation?
Correct validation is key to a realistic backtest. We use purged walk-forward cross-validation:
- Training: t=0 to t=T
- Purge gap: T to T+embargo (eliminates look-ahead from overlapping labels)
- Test: T+embargo to T+embargo+H
- Embargo period: usually equal to the forecast horizon
Embargo period ensures that future information does not leak into the training set. This is critical for time series. Metrics: IC (Information Coefficient) — correlation between predicted and actual return ranks. IC > 0.05 is weak, IC > 0.10 is good. ICIR (IC Information Ratio) — signal stability. Strategy Sharpe ratio from the signal is the main practical metric. Efficient Market Hypothesis states markets are efficient, but in practice micro-anomalies exist and can be identified with correct validation.
For model selection, consider the volume and structure of data: if many instruments — LightGBM ranking, if a single time series — LSTM, if multi-instruments with known events — Temporal Fusion Transformer. With limited data, start with LightGBM.
Model Features and Architecture
Price-based (technical analysis):
- Returns: log returns for 1, 5, 10, 21 trading days.
- Momentum: 12-1 month momentum (Jegadeesh-Titman factor).
- RSI, MACD, Bollinger Band width — oscillators as functions of price.
- Volatility: realized volatility for 5/21/63 days.
Volume-based:
- Volume relative to 20-day average.
- Price × Volume (dollar volume).
- On-Balance Volume (OBV).
- VWAP deviation.
Fundamental (for stocks):
- P/E, P/B, EV/EBITDA.
- EPS growth YoY.
- Revenue growth.
- Debt/Equity.
Alternative data:
- Sentiment from Twitter/Reddit (NLP score).
- Google Trends for consumer stocks.
- Satellite imagery (retail parking lots, commodity stores).
- Job postings growth (Glassdoor, LinkedIn).
Comparison of main modeling approaches:
| Model | Strengths | Weaknesses | Application |
|---|---|---|---|
| LightGBM (ranking) | Fast, interpretable, resistant to overfitting | Does not handle sequences | Cross-sectional ranking, large universe |
| LSTM | Captures temporal dependencies | Slow training, requires clean data | Single instrument, time series |
| Temporal Fusion Transformer | Handles future covariates, multi-horizon | Complex tuning | Many instruments with known events |
LightGBM trains 10x faster than LSTM on tabular data — an advantage for rapid prototyping. For ranking tasks, we use LGBMRanker with objective='lambdarank'. Example configuration:
import lightgbm as lgb model = lgb.LGBMRanker( objective='lambdarank', n_estimators=500, learning_rate=0.05, max_depth=6 ) For single instrument time series, we use LSTM with 60 days of history:
model = Sequential([ LSTM(64, return_sequences=True, input_shape=(60, n_features)), Dropout(0.2), LSTM(32), Dropout(0.2), Dense(1) ]) Temporal Fusion Transformer — the best choice when known future covariates (earnings dates, macro events) and 100+ instruments are available.
Model quality is assessed not only by IC. We use a comprehensive set of metrics:
| Metric | Good Value | Interpretation |
|---|---|---|
| Information Coefficient | > 0.05 | Correlation of predictions with reality |
| ICIR | > 0.5 | Signal stability |
| Sharpe ratio (after TC) | > 1.0 | Strategy efficiency |
| Win rate | > 55% | Share of profitable trades |
From Model to Trading Strategy
Model → signal → position → PnL — a chain with multiple stages of loss:
- Signal generation: ranking score across stock universe (typically 500-1000 instruments).
- Portfolio construction: mean-variance optimization (Markowitz) or equal-weight deciles. Typical number of positions 20-50.
- Risk management: limits on sector/factor exposure, max position size 5%.
- Transaction cost model: bid-ask spread + market impact (Almgren-Chriss) — accounts for slippage, often 10-30 bps.
- Backtesting: with real TC and slippage — key! We use Zipline / Backtrader or custom backtester.
Common mistakes: survivorship bias (training only on existing stocks), look-ahead bias in fundamental data (use point-in-time), ignoring transaction costs. We document every assumption.
What's Included
- Documentation: dashboard with metrics (IC, Sharpe), model description, reproducible code.
- Model access: REST API or Python package with documentation.
- Team training: workshop on operation and retraining.
- Support: 3 months after deployment, including drift monitoring.
Get an engineer consultation for your project — we'll assess data and timelines for free.
Our Results
We have built models for several hedge funds and prop trading teams. Average savings on transaction costs are 20-30% compared to naive benchmarks. We guarantee IC > 0.05 on out-of-sample, Sharpe ratio > 1.0 after TC. Certified in AWS and GCP ML. Experience with LightGBM and PyTorch. Contact us for a project evaluation — we'll calculate timelines and cost for free.







