backtrader
Event-driven backtesting with bar-by-bar execution, complex order types, multiple analyzers,…
Walk-forward validation framework for trading strategies and ML models with time-series-aware splits, overfit detection, and regime-aware validation
$ npx -y skills add agiprolabs/claude-trading-skills --skill walk-forward-validation --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/walk-forward-validationContext preview
The summary Claude sees to decide when to auto-load this skill.
Walk-forward validation framework for trading strategies and ML models with time-series-aware splits, overfit detection, and regime-aware validation
name: walk-forward-validation description: Walk-forward validation framework for trading strategies and ML models with time-series-aware splits, overfit detection, and regime-aware validation
Walk-forward validation framework for trading strategies and ML models. Standard cross-validation (k-fold, random splits) fails catastrophically for financial time series because it introduces lookahead bias and ignores autocorrelation. This skill covers proper time-series validation techniques including rolling and expanding windows, purged cross-validation, combinatorial purged cross-validation (CPCV), and overfit detection metrics.
Standard k-fold CV assumes data points are independent and identically distributed (IID). Financial time series violate both assumptions:
1. **Lookahead bias** — Random splits let the model train on future data and predict past data, artificially inflating performance. 2. **Autocorrelation** — Adjacent observations are correlated. A random split that puts Monday in test and Tuesday in train leaks information. 3. **Regime dependence** — Markets shift between regimes. A model trained on a bull market and tested on a bull market tells you nothing about bear market performance. 4. **Label overlap** — If labels are computed over windows (e.g., 24h forward return), adjacent train/test samples share label computation periods, leaking information.
The train window has a fixed size and slides forward in time. This is preferred when you believe older data is less relevant (common in crypto).
Window 1: [===TRAIN===][=TEST=] Window 2: [===TRAIN===][=TEST=] Window 3: [===TRAIN===][=TEST=]
**Parameters:**
The train window starts at the beginning and expands forward. This uses all available historical data, which helps when data is scarce.
Window 1: [==TRAIN==][=TEST=] Window 2: [====TRAIN====][=TEST=] Window 3: [======TRAIN======][=TEST=]
**Parameters:**
| Factor | Rolling | Expanding | |---|---|---| | Data recency | Prioritizes recent data | Uses all history | | Regime changes | Better adapts to new regimes | May dilute recent regime | | Sample size | Fixed, may be small | Grows over time | | Crypto preference | Preferred for < 6mo horizons | Better for regime-stable models |
Remove training samples whose labels overlap with the test set's time range. If a label is computed as the 24h forward return starting at time `t`, any training sample where `t + 24h` extends into the test period must be purged.
def purge_train_indices(
train_idx: list[int],
test_start: int,
label_horizon: int,
timestamps: list[int],
) -> list[int]:
"""Remove train samples whose label windows overlap test period."""
test_start_time = timestamps[test_start]
return [
i for i in train_idx
if timestamps[i] + label_horizon < test_start_time
]Add a buffer gap between the end of training and start of testing to account for serial correlation that purging alone does not eliminate.
[===TRAIN===][--EMBARGO--][=TEST=]
Typical embargo sizes:
CPCV (Lopez de Prado, 2018) generates all possible train/test combinations from `N` groups while maintaining temporal ordering. This produces far more test paths than standard walk-forward, enabling statistical tests for overfitting.
**Key properties:**
See `references/methodology.md` for the full CPCV algorithm and formulas.
The observed Sharpe ratio must be adjusted for:
import numpy as np
from scipy.stats import norm
def deflated_sharpe_ratio(
observed_sr: float,
num_trials: int,
backtest_length: int,
skewness: float = 0.0,
kurtosis: float = 3.0,
) -> float:
"""Compute the probability that observed SR > 0 after deflation.
Args:
observed_sr: Annualized Sharpe ratio of the selected strategy.
num_trials: Number of strategies tested (including discarded ones).
backtest_length: Number of return observations.
skewness: Skewness of returns.
kurtosis: Excess kurtosis of returns.
Returns:
p-value (probability SR is genuinely > 0).
"""
sr_std = np.sqrt(
(1 - skewness * observed_sr + (kurtosis - 1) / 4 * observed_sr**2)
/ (backtest_length - 1)
)
# Expected max SR under null (Euler-Mascheroni approximation)
euler_mascheroni = 0.5772156649
expected_max_sr = norm.ppf(1 - 1 / num_trials) * (
1 - euler_mascheroni
) + euler_mascheroni * norm.ppf(1 - 1 / (num_trials * np.e))
dsr = norm.cdf((observed_sr - expected_max_sr) / sr_std)
return dsrA DSR below 0.95 suggests the observed performance is likely due to overfitting across the trials
A comprehensive collection of 68 ready-to-use trading, DeFi, and quantitative finance Agent Skills. Works with Claude Code, Cursor, Codex, Gemini CLI, and 30+ other tools.
Repo: agiprolabs/claude-trading-skills
Event-driven backtesting with bar-by-bar execution, complex order types, multiple analyzers,…
Solana token market data via Birdeye — prices, OHLCV, trades, token metadata, security…
Broad crypto market data from CoinGecko covering 13,000+ tokens. Global market stats,…
Cointegration testing for pairs trading using Engle-Granger, Johansen, and rolling stability…
Wallet evaluation, monitoring, and copy-trade strategy design for Solana DEX trading
Cross-asset correlation analysis including rolling correlation, hierarchical clustering, tail…