Reinforcement Learning · Search · PyTorch
From search to self-play
Rebuilding the path from classical search to AlphaZero — environment, agents, and evaluation in one harness.
Minimax. Searches future states with alpha-beta pruning.
- Search depth
- 1 → 5
- Win rate vs heuristic
- 14.5% → 79%
- Search cost
- 7 → 2,022nodes/move
- Training
- 3,000DQN self-play episodes

