Supervised Project · SCE · Computer Science · Ashdod

Zone Tag

A symmetric multi-agent reinforcement-learning environment in which two agents learn to pursue, evade, orient, and counter-adapt through self-play.

2024–2025Daniel Cohen · Gal ElhianiPPOUnity ML-AgentsSelf-playOutstanding Research Project · 2025
3M+training steps
17observation dimensions
2continuous actions
PPOself-play policy learning
Research question

Can directional geometry shape richer competitive behavior?

The project asks whether two identically structured agents, without fixed chaser/evader roles, can learn meaningful pursuit-evasion strategies when tagging is determined not only by contact but by who is facing the opponent more directly at the moment of collision.

A dot-product comparison between each agent’s forward vector and the direction toward its opponent determines the successful tag. This turns orientation itself into part of the game mechanics and reward structure.

Core objectives

  • Build a symmetric two-agent tag environment.
  • Train both agents simultaneously with PPO.
  • Use geometric alignment to resolve tagging.
  • Study emergent strategies and counter-strategies.
  • Test whether curriculum learning improves generalization.
System design

Self-play, geometric rewards, and curriculum learning

Observation space

Each agent observes its own position and forward direction, the opponent’s position and orientation, direction and distance to the opponent, and a directional-alignment comparison.

Action space

Two continuous controls govern forward/backward motion and rotation, enabling fine-grained maneuvering rather than discrete movement.

Reward structure

A successful directional tag earns a positive reward, being tagged is penalized, and a small time penalty discourages stalling.

Unity scene view of the Zone Tag arena: a walled square arena with two humanoid agents, one red and one blue, and cube obstacles around it.
The Zone Tag arena in the Unity editor, with the two competing agents. Select the image to open it full size.
Diagram of the Zone Tag training system: agents, geometric tag rule, PPO self-play and curriculum stages.
Conceptual view of the training loop and the geometric tag rule.
What the experiments revealed

The interesting part was not just learning; it was adaptation.

Emergent exploit

In a compact arena, one agent discovered a spinning strategy that maximized directional coverage and initially dominated the opponent.

Counter-strategy

The opponent later adapted by charging directly into the predictable spinning behavior, reversing the performance gap. This was a useful demonstration of strategic co-adaptation in self-play.

Overfitting failure

With fixed spawn locations, both policies began exploiting positional regularities rather than learning general game dynamics. Randomized spawning and orientation were introduced to break those shortcuts.

Obstacle failure

When obstacles were added without being represented properly in the agents’ observations, behavior degraded. The agents could not reliably reason about objects they were not actually observing. This was a concrete lesson in state representation design.

Evaluation

Success was measured behaviorally as well as numerically.

The evaluation combined tag win rate, cumulative reward, episode length, periodic model snapshots, learning curves, and qualitative replay analysis. That combination mattered because an apparently high reward could still reflect an exploit or overfitting rather than robust strategic behavior.

Main metrics

  • Tag win rate
  • Average episode length
  • Cumulative reward
  • Behavioral replay analysis
  • Performance against earlier policy snapshots
Project materials

Final report.

The student final report for this project, as a PDF.

Research continuation

From capstone to research question

The project developed into a broader research direction on geometric awareness in agent-based tag, including how facing direction changes the learned behavior of autonomous agents.