backtrader
Event-driven backtesting with bar-by-bar execution, complex order types, multiple analyzers,…
Reinforcement learning for trade execution optimization including order splitting, adaptive timing, and impact minimization
$ npx -y skills add agiprolabs/claude-trading-skills --skill rl-execution --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/rl-executionContext preview
The summary Claude sees to decide when to auto-load this skill.
Reinforcement learning for trade execution optimization including order splitting, adaptive timing, and impact minimization
name: rl-execution description: Reinforcement learning for trade execution optimization including order splitting, adaptive timing, and impact minimization
Reinforcement learning (RL) for trade execution teaches an agent to split and time large orders so that total market impact is minimized. Instead of following a fixed schedule (TWAP, VWAP), an RL agent observes real-time market state and adapts its trading rate on the fly.
Every trade has a cost beyond the quoted spread:
| Cost Component | Cause | Typical Magnitude | |---|---|---| | Spread cost | Crossing the bid-ask | 5-50 bps on DEXs | | Temporary impact | Consuming liquidity | Scales with trade rate | | Permanent impact | Information leakage | Scales with total size | | Timing risk | Price drifts while waiting | Scales with volatility and time |
A 100 SOL market buy on a thin pool can move the price 2-5%. Splitting it into ten 10 SOL slices over a few minutes can cut that cost by 30-60%. The question is **how** to split optimally — and that is where execution algorithms and RL come in.
The agent observes at each decision step:
state = [
remaining_qty, # How much is left to trade (0-1 normalized)
time_remaining, # Fraction of allowed horizon remaining
current_price, # Current mid-price (normalized to arrival price)
spread, # Current bid-ask spread
volatility, # Recent realized volatility
volume, # Recent trading volume (normalized)
]Discrete actions controlling how much to trade this step:
actions = [0%, 10%, 25%, 50%, 100%] # of remaining quantity
A small action space keeps the problem tractable. Each action represents the fraction of the remaining order to execute in the current time step.
The reward penalizes execution cost relative to a benchmark:
reward = -(execution_price - arrival_price) * quantity_traded
Summed over all steps, the total reward equals the negative implementation shortfall. The agent learns to minimize total cost.
One episode = one order from placement to completion:
1. Agent receives order: buy/sell Q units within T time steps 2. At each step, agent picks an action (trade amount) 3. Market simulator applies price impact and updates state 4. Episode ends when quantity is fully executed or time expires 5. Any remaining quantity at expiry is executed at market (penalty)
The simplest baseline — split the order equally across all time steps:
trade_per_step = total_quantity / num_steps
**Pros**: Simple, deterministic, easy to implement. **Cons**: Ignores market conditions entirely.
Split proportional to expected volume in each period:
trade_at_step_t = total_quantity * (expected_volume[t] / total_expected_volume)
**Pros**: Trades more when liquidity is available. **Cons**: Requires accurate volume forecasts; still non-adaptive.
The foundational analytical model. Minimizes a combination of execution cost and timing risk:
minimize: E[cost] + λ * Var[cost]
With linear impact assumptions, this yields a closed-form optimal trajectory. See `references/execution_algorithms.md` for the full derivation.
An RL agent (DQN, PPO, or similar) that learns the execution policy from simulated experience:
# Pseudocode training loop
for episode in range(num_episodes):
state = env.reset(order_qty=Q, horizon=T)
done = False
while not done:
action = agent.select_action(state)
next_state, reward, done, info = env.step(action)
agent.store_transition(state, action, reward, next_state, done)
agent.update()
state = next_state**Pros**: Adapts to current market conditions, can learn non-linear patterns. **Cons**: Requires realistic simulator, sim-to-real gap, training instability.
The simulator uses a standard two-component impact model:
temporary_impact = η * (trade_rate / avg_volume) permanent_impact = γ * (trade_rate / avg_volume)
The execution price for a trade of size `q` at time `t`:
exec_price = mid_price + permanent_impact + temporary_impact mid_price_next = mid_price + permanent_impact + noise
This skill is most valuable when:
For small retail orders (<$1,000 on liquid pairs), simple market orders or basic slippage limits are sufficient. See the `slippage-modeling` skill instead.
1. **Sim-to-real gap**: Simulated markets do not capture all real dynamics (queue position, adversarial flow, MEV). 2. **Non-stationarity**: Market microstructure changes over time; models trained on one regime may fail in another. 3. **DEX specifics**: On-chain execution has block-level granularity (~400ms on Solana), not continuous-time. Gas/priority fees add cost. 4. **Data requirements**: Training requires historical orderbook or trade data for realistic simulation.
| Skill | Integration | |---|---| | `slippage-modeling` | Provides impact estimates to calibrate the simulator | | `position-sizing` | Determines the total order size to execute | | `liquidity-analysis` | Assesses available liquidity for realistic
A comprehensive collection of 68 ready-to-use trading, DeFi, and quantitative finance Agent Skills. Works with Claude Code, Cursor, Codex, Gemini CLI, and 30+ other tools.
Repo: agiprolabs/claude-trading-skills
Event-driven backtesting with bar-by-bar execution, complex order types, multiple analyzers,…
Solana token market data via Birdeye — prices, OHLCV, trades, token metadata, security…
Broad crypto market data from CoinGecko covering 13,000+ tokens. Global market stats,…
Cointegration testing for pairs trading using Engle-Granger, Johansen, and rolling stability…
Wallet evaluation, monitoring, and copy-trade strategy design for Solana DEX trading
Cross-asset correlation analysis including rolling correlation, hierarchical clustering, tail…