Zone Tag
A symmetric multi-agent reinforcement-learning environment in which two agents learn to pursue, evade, orient, and counter-adapt through self-play.
Can directional geometry shape richer competitive behavior?
The project asks whether two identically structured agents, without fixed chaser/evader roles, can learn meaningful pursuit-evasion strategies when tagging is determined not only by contact but by who is facing the opponent more directly at the moment of collision.
A dot-product comparison between each agent’s forward vector and the direction toward its opponent determines the successful tag. This turns orientation itself into part of the game mechanics and reward structure.
Core objectives
- Build a symmetric two-agent tag environment.
- Train both agents simultaneously with PPO.
- Use geometric alignment to resolve tagging.
- Study emergent strategies and counter-strategies.
- Test whether curriculum learning improves generalization.
Self-play, geometric rewards, and curriculum learning
Observation space
Each agent observes its own position and forward direction, the opponent’s position and orientation, direction and distance to the opponent, and a directional-alignment comparison.
Action space
Two continuous controls govern forward/backward motion and rotation, enabling fine-grained maneuvering rather than discrete movement.
Reward structure
A successful directional tag earns a positive reward, being tagged is penalized, and a small time penalty discourages stalling.

The interesting part was not just learning; it was adaptation.
Emergent exploit
In a compact arena, one agent discovered a spinning strategy that maximized directional coverage and initially dominated the opponent.
Counter-strategy
The opponent later adapted by charging directly into the predictable spinning behavior, reversing the performance gap. This was a useful demonstration of strategic co-adaptation in self-play.
Overfitting failure
With fixed spawn locations, both policies began exploiting positional regularities rather than learning general game dynamics. Randomized spawning and orientation were introduced to break those shortcuts.
Obstacle failure
When obstacles were added without being represented properly in the agents’ observations, behavior degraded. The agents could not reliably reason about objects they were not actually observing. This was a concrete lesson in state representation design.
Success was measured behaviorally as well as numerically.
The evaluation combined tag win rate, cumulative reward, episode length, periodic model snapshots, learning curves, and qualitative replay analysis. That combination mattered because an apparently high reward could still reflect an exploit or overfitting rather than robust strategic behavior.
Main metrics
- Tag win rate
- Average episode length
- Cumulative reward
- Behavioral replay analysis
- Performance against earlier policy snapshots
From capstone to research question
The project developed into a broader research direction on geometric awareness in agent-based tag, including how facing direction changes the learned behavior of autonomous agents.